Job resource allocation method and device, electronic device, and storage medium
By using Prophet and FTRL models in computing clusters such as MapReduce, Spark, and Flink to predict business volume and resource metrics, automated resource allocation is achieved, solving the problem of resource allocation relying on human experience in existing technologies and improving the intelligence and accuracy of allocation.
Patent Information
- Application Number
- CN202110982573.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-25
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2041-08-25
AI Technical Summary
In existing computing clusters such as MapReduce, Spark, and Flink, resource allocation relies on human experience, resulting in low intelligence and difficulty in maintaining allocation rules. This leads to large resource estimation errors, affecting application processing capabilities and stability.
By collecting monitoring indicators from work sites, and using Prophet and FTRL models to predict workload and resource indicators, resources are automatically allocated, achieving precise allocation without human intervention.
It improves the intelligence and accuracy of resource allocation, reduces manual intervention, lowers prediction errors, and ensures application processing capacity and stability.
Smart Images

Figure CN115129441B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, in particular to a job resource allocation method and device, electronic equipment and computer storage medium. BACKGROUND
[0002] In the MapReduce, Spark, Flink and other computing clusters commonly used in the field of big data, a developer needs to deploy a large number of applications, and how to reasonably allocate resources to each application is a technical difficulty, because allocating too many resources will waste cluster resources, and allocating too few resources will cause the application processing capacity to decline or even crash.
[0003] In the prior art, the required resources of each application are usually predicted by manual experience, and the resources are manually allocated; or the resources are automatically adjusted according to rules generated by experience; or historical data of resource allocation is collected for manual annotation, and then a model is built using the manually annotated samples, and the model is used for resource prediction.
[0004] However, the above solutions are too dependent on manual experience or manual annotation, have low intelligence, and the allocation rules are difficult to maintain, manual annotation is not suitable for a large amount of logs, resulting in a large resource estimation error. SUMMARY
[0005] The technical problem solved by the present application is to provide a job resource allocation method and device, electronic equipment and computer storage medium to improve the intelligence and accuracy of resource allocation.
[0006] To solve the above technical problem, one technical solution adopted by the present application is to provide a job resource allocation method. The job resource allocation method comprises: collecting a current value of a job point monitoring index, wherein the monitoring index comprises a resource index and a business volume index; predicting a plurality of first prediction values of the business volume index based on the current value of the business volume index; predicting a plurality of second prediction values of the resource index based on the plurality of first prediction values; and allocating resources of the job point based on the plurality of second prediction values.
[0007] To solve the above technical problem, one technical solution adopted by the present application is to provide a job resource allocation device. The job resource allocation device comprises: a collection module for collecting a current value of a job point monitoring index, wherein the monitoring index comprises a resource index and a business volume index; a first prediction module connected to the collection module, for predicting a plurality of first prediction values of the business volume index based on the current value of the business volume index; a second prediction module connected to the first prediction module, for predicting a plurality of second prediction values of the resource index based on the plurality of first prediction values; and a resource allocation module connected to the second prediction module, for allocating resources of the job point based on the plurality of second prediction values.
[0008] To solve the above technical problems, one technical scheme adopted by the present application is to provide an electronic device. The electronic device comprises a processor and a memory connected with the processor, wherein the memory is used to store program data, and the processor is used to execute the program data to realize the job resource allocation method.
[0009] To solve the above technical problems, one technical scheme adopted by the present application is to provide a computer storage medium. The computer storage medium stores program data, and the program data can be executed to realize the job resource allocation method.
[0010] The beneficial effects of the embodiments of the present application are that the present application can predict a plurality of first prediction values of the business volume index according to the current values of the monitoring indexes of the job points, then predict a plurality of second prediction values of the resource index based on the plurality of first prediction values of the business volume index, and allocate the resources of the job points based on the plurality of second prediction values, so that the present application can automatically allocate the resources of the job points according to the collected monitoring indexes of the job points without manual experience guidance and manual annotation and other manual intervention, thereby improving the intelligent level and accuracy of resource allocation. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0012] Figure 1 is a flowchart of an embodiment of the job resource allocation method of the present application;
[0013] Figure 2 is Figure 1 is a specific flowchart of step S13 in the job resource allocation method of the embodiment;
[0014] Figure 3 is Figure 2 is a specific flowchart of step S24 in the embodiment;
[0015] Figure 4 is a flowchart of the partition training of the correlation model of the present application;
[0016] Figure 5 is a flowchart of real-time prediction of the resource index of the present application;
[0017] Figure 6 is Figure 1 is a specific flowchart of step S14 in the job resource allocation method of the embodiment;
[0018] Figure 7 is Figure 1 A specific flowchart of step S14 in the method for deploying job resources in an embodiment;
[0019] Figure 8 is Figure 1 A specific flowchart of step S14 in the method for deploying job resources in an embodiment;
[0020] Figure 9 is Figure 8 A specific flowchart of step S83 in an embodiment;
[0021] Figure 10 is a flowchart of an embodiment of the method for deploying job resources in the application;
[0022] Figure 11 is Figure 10 A flowchart of collecting monitoring indexes and forming training samples in the method for deploying job resources in an embodiment;
[0023] Figure 12 is a flowchart of an embodiment of the method for deploying job resources in the application;
[0024] Figure 13 is a structure diagram of a control chart of abnormal events of job resources in the application;
[0025] Figure 14 is a structure diagram of an embodiment of the device for deploying job resources in the application;
[0026] Figure 15 is a structure diagram of an embodiment of the electronic device in the application;
[0027] Figure 16 is a structure diagram of an embodiment of the computer storage medium in the application. DETAILED DESCRIPTION
[0028] The application will be further described below in conjunction with the drawings and embodiments. It is particularly pointed out that the following embodiments are only used to illustrate the application, but do not limit the scope of the application. Similarly, the following embodiments are only part of the embodiments of the application, but not all the embodiments of the application. All other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the application.
[0029] In the description of the embodiments of the present application, it should be noted that unless specifically defined and limited, the terms "connected", "connected" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected, it can be mechanically connected, or it can be electrically connected, it can be directly connected, or it can be indirectly connected through an intermediate medium. For those skilled in the art, the specific meaning of the above terms in the embodiments of the present application can be understood according to the specific circumstances.
[0030] In the embodiments of the present application, unless otherwise specifically defined and limited, the first feature is "on" or "under" the second feature, which can be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature can be above, above and above the second feature, or it can only mean that the horizontal height of the first feature is higher than that of the second feature. The first feature can be below, below and below the second feature, or it can only mean that the horizontal height of the first feature is less than that of the second feature.
[0031] Apache Flink is an open source stream processing framework developed by Apache Software Foundation, and its core is a distributed stream data flow engine written in Java and Scala. Flink executes any stream data program in a data parallel and pipeline manner, and the pipeline runtime system of Flink can execute batch processing and stream processing programs. In addition, the Flink runtime itself also supports the execution of iterative algorithms. Its memory processing and pipeline application in the online real-time processing scenario can provide real-time processing capability for business.
[0032] The following embodiments of the present application are based on the Flink real-time computing framework, and in other embodiments, Structed Streaming can be used instead of Flink.
[0033] In order to improve the intelligent level and accuracy of job point resource allocation in the cluster, the present application first proposes a job resource allocation method, as shown in Figure 1 Figure 1 is a flow diagram of an embodiment of the job resource allocation method of the present application. The job resource allocation method of the present application specifically includes the following steps:
[0034] Step S11: Collect the current value of the job point monitoring index, wherein the monitoring index includes resource index and business volume index.
[0035] In the cluster, a plurality of job points are usually provided, and the job point refers to a computing node in the cluster, which can be each physical or virtual computer, etc.
[0036] The Reporter of Flink periodically collects the monitoring indicators of each job point in the cluster.
[0037] The business volume indicator can include the data throughput and processing delay of the application of the job point at the collection time point, and represents the data volume or computing time consumption that the application can process at the given resource at the collection time point, or the resource required for processing the business volume. The business volume indicator can be represented by an x vector, such as x = |tps, lag, … |. The business volume indicator can be used as part of the feature independent variable of the training sample (i.e., the label sample) of the correlation model (regression model) in the following.
[0038] The resource indicator can be a monitoring indicator corresponding to an adjustable resource parameter configuration item, such as the JobManager process Jvm off-heap memory, the TaskManager process Slot number, etc. It is part of the context indicator itself. When it is necessary to predict a certain resource indicator and then guide the adjustment of the corresponding configuration item parameter, the resource indicator can be removed from the context indicator and used as the label dependent variable of the training sample (i.e., the label sample) of the correlation model (regression model) alone. The training sample does not need to be manually labeled and is directly obtained through monitoring indicator collection. The resource indicator can be represented by y, such as y = [slot | jobmanager.memory.off-heap.size | … ].
[0039] The current value of the job point monitoring indicator can be represented as x(t), y(t).
[0040] Step S12: predicting a plurality of first prediction values of the business volume indicator based on the current value of the business volume indicator.
[0041] The embodiment can predict the business volume of the job node at a plurality of future time points through the trend prediction model to obtain a plurality of first prediction values.
[0042] The trend prediction model can be represented by s, the current value of the business volume indicator at the current time t is x(t), and the first prediction value of the business volume indicator of the job point at the next time t+1 predicted by the trend prediction model is x(t+1) = s(x(t)).
[0043] The embodiment can use the open-source Prophet algorithm to implement the trend prediction model to realize the trend prediction of the business volume (the numerical value of the business volume indicator), and use the Flink window function to accumulate the training sample to update the Prophet model periodically. Since the running environment and application context of the job node do not change much in a short time, the embodiment can predict the resource indicator according to the first prediction values of the business volume indicator at a plurality of future time points, i.e., x(t+1), x(t+2), x(t+3), ….
[0044] In other embodiments, a trend prediction model such as ARIMA or LSTM can also be used.
[0045] Step S13: predicting a plurality of second prediction values of the resource index based on the plurality of first prediction values.
[0046] Optionally, to improve the accuracy of resource allocation, the collected monitoring indicators further include environmental information indicators, context indicators, and derived indicators.
[0047] The environmental information indicators can include indicators such as free memory, maximum available central processing unit (CPU), disk IO, and network port traffic of the job node at the collection time point; the environmental information indicators can be used as part of the feature independent variables of the training samples (i.e., label samples) of the correlation model (regression model), representing the environment in which the application or cluster runs; the environmental information indicators can be represented by a vector env, such as env = |Memory, CPU, IO, … |.
[0048] The context indicators can include features such as heap in / out memory, core number, parallelism, and computing framework output log text embedding word vector consumed by each Jvm process of the application runtime at the collection time point; the context indicators can be used as part of the feature independent variables of the training samples (i.e., label samples) of the correlation model (regression model), for describing the running state or real resource consumption value of the application near the collection time point; the context indicators can be represented by a vector ctx, such as ctx = |vCores, parallelism, embedding, … |.
[0049] The derived indicators are derived indicators of the above monitoring indicators, including monitoring indicators of each job node (such as Master and Worker in the master-slave mode) and different job granularities (such as App>Job>Task), and new indicators are constructed by calculating same / compared with the previous period, minimum value, maximum value, average value, and summary value, for enhancing the nonlinear fitting capability of the sample features; the derived indicators can be represented by a vector der.
[0050] The above traffic indicators, environmental information indicators, context indicators, and derived indicators constitute the entire sample feature independent variables, which can be represented as X = |x, env, ctx, der| = |tps, lag, Memory, CPU, IO, vCores, embedding, tpsMin, lagMax, MemoryAvg, vCoresSum, … |.
[0051] Optionally, the present embodiment can be implemented by, for example, Figure 2The method shown realizes step S13. The method of this embodiment includes steps S21 to S22.
[0052] Step S21: Obtain a correlation model between the resource indicator and the traffic volume indicator, the environment information indicator, the context indicator, and the derived indicator.
[0053] The feature independent variable plus the label dependent variable constitutes a complete training sample. Assuming that the relationship between X and y is a multivariate linear regression relationship, the weight coefficients of each feature independent variable are w1, w2, w3, …, and y can be represented as y = w1*tps + w2*lag + w3*Memory + w4*CPU + w5*IO + w6*vCores + w7*tpsMin, ….
[0054] Specifically, this embodiment can realize the correlation model through a regression algorithm, i.e., a regression model; the correlation model can be represented as f, and y = f(x, env, ctx, …), which indicates that the processing traffic volume x requires y resources under the same collection time point, context, and running environment.
[0055] The resource indicator y and the traffic volume indicator x, the running environment indicator env, and the environment information indicator ctx of this embodiment have a multivariate linear regression relationship (nonlinear is fitted through a derived indicator), so an important purpose of solving real-time resource prediction is to update the above weight coefficients in real time, and the second prediction value of the resource indicator is obtained by using the updated weight coefficients in the correlation model. To this end, this embodiment can use the FTRL model to realize the correlation model, which can update the feature weight coefficients online, and predict multiple second prediction values of the resource indicator according to multiple first prediction values of the traffic volume indicator.
[0056] Step S22: Use the correlation model to predict multiple second prediction values of the resource indicator corresponding to the multiple first prediction values.
[0057] Since the running environment (environment information indicator env) and the application running state (context indicator ctx) do not change much in a short time, the FTRL model can accurately predict the resource amount required by the job point at multiple future time points, i.e., multiple second prediction values of the resource indicator:
[0058] y(t+1) = f(x(t+1), env, ctx);
[0059] y(t+2) = f(x(t+2), env, ctx);
[0060] y(t+3) = f(x(t+3), env, ctx), ….
[0061] Optionally, this embodiment further includes steps S23 and S24.
[0062] Step S23: obtaining a training sample by using the current value of the resource indicator, the current value of the traffic indicator, the current value of the environment information indicator, the current value of the context indicator and the current value of the derived indicator.
[0063] The above monitoring indicators are assembled into a training sample composed of a feature independent variable and a label dependent variable by using Flink real-time processing, so as to train the FTRL model.
[0064] Step S24: training and updating the correlation model by using the training sample.
[0065] The model parameters of the correlation model are trained and updated in real time by using the above training sample, without storing a large amount of logs for manual labeling, so as to reduce the prediction error without human intervention, and achieve the effect of automatically optimizing the resources of the application deployed by the user.
[0066] Since the monitoring indicators are collected in real time as stream events, the correlation model needs to have the ability of online training, and the FTRL model used in the embodiment can assemble a training sample and update the weight coefficient once for each monitoring indicator log; the FTRL model can be continuously and automatically updated and optimized to maintain the optimal prediction ability.
[0067] Moreover, the model training and model prediction of the embodiment are parallel, the training link continuously optimizes and updates the FTRL model based on event driving, and the prediction link uses the latest (also the optimal) FTRL model to predict and output the resource prediction result. It can be known that the FTRL model is continuously self-optimized with the continuous collection of event sources, and the prediction effect is more stable and accurate.
[0068] Flink can realize the parameter update of the FTRL model by using the method:
[0069]
[0070]
[0071] The above code is the Java implementation of the following mathematical formula pseudocode (the weight w, the cumulative gradient g, and the FTRL model parameters such as delta, z and n are stored in the MapState of Flink with the feature as the key and the parameter value as the Value, and alpha, beta, lambda1 and lambda2 are transmitted from the configuration center as hyperparameters; on the basis of the prior art, the part in the box is modified to adapt to the stream gradient descent optimization of the multiple linear regression) :
[0072]
[0073] Optionally, the embodiment can be realized by using the Flink machine learning library as follows: Figure 3The method shown realizes step S24. The method of this embodiment includes steps S31 to S33.
[0074] Step S31: Obtain the application number of the job point.
[0075] Step S32: Select the correlation model corresponding to the application number from the preset plurality of correlation models.
[0076] Step S33: Train and update the correlation model corresponding to the application number using the training sample.
[0077] Steps S31 to S33 are collectively described as follows: Each job point can be applied in multiple applications, and a corresponding application number can be set for each application. Each application is based on the application number to correspond to an FTRL model. When Flink trains the FTRL model, it will route the monitoring indicators of each application to the task corresponding to the application number according to the application number partition, such as Figure 4 As shown, an FTRL model is independently trained for each application to capture the complexity of the application computing logic.
[0078] Therefore, this embodiment can realize fine-grained job monitoring and dynamic resource scheduling, that is, the monitoring and resource adjustment granularity of this embodiment is refined to a specific application or job (by dynamically partitioning the application by Flink, an FTRL model can be automatically trained for each application. Assuming there are ten thousand jobs or applications, then ten thousand FTRL models are automatically trained in ten thousand partition data).
[0079] The current values x(t) of the traffic indicators and the current values y(t) of the resource indicators, the environment information indicators env and the context indicators ctx collected by the current collection time point t of this embodiment, and the relationship between the models FTRL model and Prophet model updated in real time by the training link are as shown in Figure 5 As shown, the Prophet model obtains a plurality of first prediction values x(t+1)=s(x(t)), x(t+2)=s(x(t+1)), x(t+3)=s(x(t+2)), … of x(t). The sample alignment component is responsible for assembling feature independent variables. The assembled feature independent variable samples are as follows:
[0080] Sample 1: <x(t+1), env, ctx>;
[0081] Sample 2: <x(t+2), env, ctx>;
[0082] Sample 3: <x(t+3), env, ctx>; ….
[0083] Since the environmental information indicator evn and the context indicator ctx do not change much in a short period of time, the FTRL model can predict the second predicted values of the resource indicators of the job point at multiple future time points, i.e., the resource amounts required by the job point at multiple future time points: y(t+1) = f(x(t+1), env, ctx), y(t+2) = f(x(t+2), env, ctx), y(t+3) = f(x(t+3), env, ctx), ….
[0084] The Prophet model is mainly used to predict traffic trends to obtain relevant features, such as the data throughput (e.g., the amount of data produced by the business system to Kafka) and the processing delay (e.g., the number of Kafka message accumulations or the time length) of the application described above. The former represents the amount of data that needs to be processed by the application, and the latter represents how long it takes for the application to process one piece of data.
[0085] Step S14: allocating resources of the job point based on the multiple second predicted values.
[0086] Optionally, the step S14 can be implemented by the method as shown in Figure 6 The method of the embodiment includes steps S61 to S63.
[0087] Step S61: obtaining the maximum value and the minimum value from the multiple second predicted values.
[0088] Obtain Min(y(t+1), y(t+2), y(t+3), …) and Max(y(t+1), y(t+2), y(t+3), …).
[0089] Step S62: in response to the current value of the resource indicator being less than the minimum value, increasing the resource amount of the job point to the maximum value.
[0090] If y(t) < Min(y(t+1), y(t+2), y(t+3), …), the resource amount of the job point is increased to the maximum value Max(y(t+1), y(t+2), y(t+3), …) to meet the maximum resource demand amount at multiple future time points.
[0091] Step S63: in response to the current value of the resource indicator being greater than the maximum value, decreasing the resource amount of the job point to the minimum value.
[0092] If y(t) > Max(y(t+1), y(t+2), y(t+3), …), the resource amount of the job point is decreased to the minimum value Min(y(t+1), y(t+2), y(t+3), …) to save resources.
[0093] Through real-time prediction, the second prediction value corresponding to the resource index is continuously updated. The embodiment does not adjust the job resource in real time frequently, but triggers the operation of increasing the resource to Max(y(t+1), y(t+2), y(t+3), …) when it is determined that the second prediction values at future multiple time points are all greater than the current value y(t) at the current time, that is, y(t) < Min(y(t+1), y(t+2), y(t+3), …). The operation of reducing the resource to Min(y(t+1), y(t+2), y(t+3), …) is triggered when y(t) > Max(y(t+1), y(t+2), y(t+3), …). The resource is not adjusted in other cases. The once adjustment to the position can be realized, and the problem of frequent resource adjustment affecting application stability caused by traffic spikes can be avoided, so that the entire resource adjustment process is more smooth.
[0094] In another embodiment, step S14 can be implemented by a method as shown in Figure 7 The method of the embodiment includes steps S71 to S75.
[0095] Step S71: Obtain the maximum value and the minimum value from the plurality of second prediction values.
[0096] Obtain Min(y(t+1), y(t+2), y(t+3), …) and Max(y(t+1), y(t+2), y(t+3), …).
[0097] Step S72: In response to the current value of the resource index being less than the minimum value, obtain a first difference value between the minimum value and the current value of the resource index.
[0098] When y(t) < Min(y(t+1), y(t+2), y(t+3), …), obtain the first difference value [Min(y(t+1), y(t+2), y(t+3), …) - y(t)].
[0099] Step S73: In response to the first difference value being greater than a first threshold, increase the resource amount of the job point to the maximum value.
[0100] When [Min(y(t+1), y(t+2), y(t+3), …) - y(t)] > the first threshold, increase the resource amount of the job point to the maximum value Max(y(t+1), y(t+2), y(t+3), …).
[0101] Step S74: In response to the current value of the resource index being greater than the maximum value, obtain a second difference value between the current value of the resource index and the maximum value.
[0102] If y(t) > Max(y(t+1), y(t+2), y(t+3), …), a second difference value [y(t) - Max(y(t+1), y(t+2), y(t+3), …)] is obtained to meet the maximum resource demand of multiple future time points.
[0103] Step S75: In response to the second difference value being greater than the first threshold value, the resource amount of the job point is reduced to a minimum value.
[0104] If [y(t) - Max(y(t+1), y(t+2), y(t+3), …)] > the first threshold value, the resource amount of the job point is reduced to a minimum value Min(y(t+1), y(t+2), y(t+3), …) to save resources.
[0105] In other cases, no adjustment is made to prevent traffic spikes from interfering, avoiding frequent resource adjustments, and making the entire resource adjustment process smoother.
[0106] When Flink makes a resource allocation decision based on the resource prediction results of the Prophet model and the FTRL model, and queries from a resource manager such as Yarn that there is not enough resource to meet the resource allocation decision, if the resource allocation decision is still used for resource allocation, Yarn will make the resource allocation instruction command queue, and the resource allocation instruction will not take effect in time, and needs to wait for other applications to release resources before it has the opportunity to execute.
[0107] To solve the above problems, the embodiment can implement step S14 by the method as shown in Figure 8 The method of the embodiment includes steps S81 to S85.
[0108] Step S81: Obtain the maximum value and the minimum value from the plurality of second prediction values.
[0109] Similar to step S61, details are not repeated here.
[0110] Step S82: In response to the current value of the resource index being less than the minimum value, obtain the sum of the adjustment amounts of resources of all job points in the cluster that need to increase resources.
[0111] The current value of the resource index of the job point is less than the minimum value, and the resource of the job point needs to be increased. At this time, the sum of the resource adjustment amounts of all job points in the cluster that need to increase resources should be obtained, which represents the amount of resources that need to be increased in the future multiple time points in the cluster.
[0112] Step S83: In response to the sum of the adjustment amounts being greater than the available amount of resources of the cluster, reduce the adjustment amount of the job point that needs to increase resources.
[0113] If the sum of the adjustment amounts of all the job points that need to increase resources in the cluster is greater than the available amount of resources of the cluster, the resource allocation cannot be performed according to the adjustment amounts, and the adjustment amounts of the job points need to be reduced.
[0114] Optionally, the embodiment can implement step S83 by a method as shown in Figure 9 The method of the embodiment includes steps S91 to S94.
[0115] Step S91: divide the available amount of resources into different groups of preset amounts multiple times, where the sum of the multiple preset amounts in each group of preset amounts is the available amount of resources.
[0116] Step S92: predict the sum of the business volumes of the job points that need to increase resources after allocating the resources of each group of preset amounts, where the job points that need to increase resources are in one-to-one correspondence with the multiple preset amounts in each group of preset amounts.
[0117] Step S93: obtain the preset amount in the group of preset amounts corresponding to the maximum value of the sum of the business volumes as the actual adjustment amount of the resources of the job points that need to increase resources.
[0118] Step S94: adjust the resources of the job points that need to increase resources according to the actual adjustment amount.
[0119] The steps S91 to S94 are introduced together:
[0120] The embodiment trains a reverse multiple linear regression model x = f'(y, env, ctx) by using the FTRL model to predict the business volume x that can be processed by the job points under the condition of allocating y resources, for example: assuming that application 1 and application 2 both need to increase m memories and q CPUs, and the total resources of the cluster are only m memories and q CPUs (uniformly denoted as an r vector), the resource allocation can be performed on application 1 and application 2 in the following manner:
[0121] First, allocate all the resources to application 1, and the total business volume that can be processed by application 1 and application 2 is handling_ability = f'(y1+r, env, ctx) + f'(y2, env, ctx), and transfer the resources from application 1 to application 2 by a step s, and the total business volume that can be processed by application 1 and application 2 is in turn:
[0122] handling_ability1 = f'(y1+(r-1)*s, env, ctx) + f'(y2+1*s, env, ctx),
[0123] handling_ability2 = f'(y1+(r-2)*s, env, ctx) + f'(y2+2*s, env, ctx),
[0124] handling ability3 = f'(y1 + (r-3)*s, env, ctx) + f'(y2 + 3*s, env, ctx),… until r≤n*s.
[0125] Find the optimal allocation point i to maximize the sum of the business processing capabilities of application 1 and application 2, and the resource allocation scheme is: (r-i)*s resources for application 1 and i*s resources for application 2. In this way, the resource r neither exceeds the available resources of the cluster, nor can the overall business processing capability of the cluster be maximized under the condition of resource limitation.
[0126] Step S84: In response to the current value of the resource indicator being greater than the maximum value, the amount of resources of the job point is reduced to the minimum value.
[0127] Similar to step S63, details are not repeated here.
[0128] Step S85: In response to the sum of the adjustment amounts being less than or equal to the available amount of resources of the cluster, the amount of resources of the job point is increased to the maximum value.
[0129] The required resource amount of all job points in the cluster can be fully met, and the resource amount of each job point is directly increased to the maximum value to meet the maximum resource demand amount at multiple future time points.
[0130] The present application is directed to Figure 7 Embodiments can also be similarly improved, and details are not repeated here.
[0131] In the process of deploying Flink job resources, it is hoped that through training, the FTRL model will learn how much resource y is needed to process business volume x under certain operating environment and application state. Therefore, the training sample assembled by Flink must be in a corresponding relationship in time, such as:
[0132] (t+0) time: x(t+0), env(t+0), ctx(t+0), y(t+0)
[0133] (t+1) time: x(t+1), env(t+1), ctx(t+1), y(t+1)
[0134] (t+2) time: x(t+2), env(t+2), ctx(t+2), y(t+2)
[0135] (t+3) time: x(t+3), env(t+3), c(txt+3), y(t+3)…
[0136] But since the traffic index x and the resource index y are generated from different systems, the monitoring indexes collected from the heterogeneous systems are difficult to be guaranteed to be at the same time; if the traffic index x and the resource index y (other indexes change little in a short time) are directly assembled according to the collection time points, the FTRL model will be wrong.
[0137] To solve the above problems, the application further provides a job resource deployment method of another embodiment, as shown in the figure, the deployment method of the embodiment specifically includes the following steps: Figure 10
[0138] Step S101: Obtain global clock heartbeat information, and collect the current value of the job point monitoring index under the control of the heartbeat information.
[0139] The monitoring index can refer to the above embodiments.
[0140] As shown in the figure, the throughputReporter is responsible for collecting the traffic index x from the business system monitoring, the systemReporter is responsible for collecting the environment information index env of the job node, and the contextReporter is responsible for collecting the context index ctx of each jvm process of the big data cluster such as Hadoop, Spark and Flink, including the resource index y to be predicted. Figure 11
[0141] Step S102: Generate time identification information corresponding to the heartbeat information.
[0142] Generate the time identification information watermark corresponding to the heartbeat information timeStamp.
[0143] Step S103: Associate the time identification information with the current value of the monitoring index.
[0144] Associate the watermark with the current value of the monitoring index.
[0145] The global clock sends the heartbeat information to each Reporter according to the configured log collection period; each Reporter receives the heartbeat information timeStamp of the global clock, starts to collect the monitoring index responsible for by itself; then generates the watermark (i.e. time identification information) according to the unique number uuid distributed by the global clock and the heartbeat information timeStamp, and injects it into the Flink event stream, and sends it to the Flink time alignment component together with the collected monitoring index (associates the time identification information with the current value of the monitoring index); the time alignment component uses the watermark mechanism of Flink to process the time disorder problem, and realizes time alignment.
[0146] When Flink processes data, EventTime is selected to process data, and when the business data corresponding to EventTime is processed, the collection action is triggered. The running index of the big data cluster and the current application is collected and pushed to the contextReporter. Since the EventTime corresponding to the collection point is uniformly managed in the global clock, the index collected at the EventTime moment can be associated through the watermark generated by the global clock time stamp and uuid. Then, the monitoring index collected by each Reporter is aligned according to the watermark, and a training sample is assembled.
[0147] Step S104: A plurality of first prediction values of the traffic volume index are predicted based on the current value of the traffic volume index.
[0148] Similar to step S12, details are not repeated here.
[0149] Step S105: Obtain a correlation model between the resource index and the traffic volume index, the environmental information index, the context index and the derived index.
[0150] Similar to step S21, details are not repeated here.
[0151] Step S106: A plurality of second prediction values of the resource index corresponding to the plurality of first prediction values are predicted using the correlation model.
[0152] Similar to step S22, details are not repeated here.
[0153] Step S107: The current value of the resource index, the current value of the traffic volume index, the current value of the environmental information index, the current value of the context index and the current value of the derived index associated with the same time identifier information are formed into the same training sample.
[0154] In this way, the error of the training sample is avoided.
[0155] Step S108: The correlation model is trained and updated using the training sample.
[0156] Similar to step S24, details are not repeated here.
[0157] The present application further proposes another embodiment of the Flink job resource deployment method, as shown in Figure 12 The deployment method of the present embodiment specifically includes the following steps:
[0158] Step S121: Collect the current value of the monitoring index of the job point, wherein the monitoring index includes the resource index and the traffic volume index.
[0159] Step S122: Based on the current value of the business volume indicator, multiple first predicted values of the business volume indicator are obtained.
[0160] Step S123: Predict multiple second predicted values for resource indicators based on multiple first predicted values.
[0161] Step S124: Allocate resources for job points based on multiple second prediction values.
[0162] Steps S121 to S124 are similar to steps S11 to S14 above, and will not be repeated here.
[0163] Step S125: Calculate the difference between the current value of the resource indicator and the second predicted value.
[0164] Calculate the differences between y(t) and y(t+1), y(t+2), y(t+3), ... in sequence.
[0165] Step S126: An abnormal event is generated in response to a difference greater than the second threshold, a difference less than the third threshold, or multiple differences gradually increasing.
[0166] like Figure 13 As shown, if a difference B is greater than the second threshold +2R, or a difference A is less than the third threshold -2R, or if multiple differences D1-D7 show a consistent increasing trend, then the deviation between the FTRL model's prediction and the actual situation is considered to be surging. This is because the FTRL model's characteristic is to ensure that the weight gradient updated by new training samples does not deviate too far from the historical gradient. Therefore, the FTRL model's prediction results have a certain degree of "sluggishness." A sudden increase in the deviation between the predicted and actual values is likely due to an abnormal event such as a "sudden spike" in business volume. Whether it is an anomaly in the real event or an increase in the error of the FTRL model result, an alert should be generated for timely follow-up and verification.
[0167] Furthermore, in this embodiment, the mean of multiple second predicted values can be obtained, and then a second threshold and a third threshold can be set based on the mean.
[0168] Specifically, Flink calculates the mean μ(t) and extreme value R(t) of the monitoring index in the t-th sampling period, where μ(t) = [μ(t-1) + v] / t, the second threshold is +2R(t), and the third threshold is -2R(t).
[0169] This embodiment eliminates the need for manual configuration of monitoring rules and thresholds, enabling full-process adaptive and dynamic resource allocation, and providing finer-grained monitoring.
[0170] Furthermore, this embodiment can also use the Prophet model to predict the trends of various monitoring indicators, and can obtain the predicted values and confidence intervals of monitoring indicators for multiple future time periods.
[0171] The Prophet model of the embodiment is automatically trained by Flink real-time stream events after initialization, and the Prophet model is retrained when the predicted value exceeds the confidence interval, so as to realize self-learning of the Prophet model and improve the accuracy of the Prophet model.
[0172] The application further provides a job resource deployment device, as shown in Figure 14 The job resource deployment device of the embodiment includes a collection module 140, a first prediction module 141, a second prediction module 142, and a resource deployment module 143. The collection module 140 is configured to collect a current value of a job point monitoring index, wherein the monitoring index includes a resource index and a traffic index. The first prediction module 141 is connected to the collection module 140 and is configured to predict a plurality of first prediction values of the traffic index based on the current value of the traffic index. The second prediction module 142 is connected to the first prediction module 141 and is configured to predict a plurality of second prediction values of the resource index based on the plurality of first prediction values. The resource deployment module 143 is connected to the second prediction module and is configured to deploy resources of the job point based on the plurality of second prediction values.
[0173] Optionally, the deployment device of the embodiment can work based on a Flink stream processing framework, and the deployment device is provided with a Flink training application module 144 and a Flink prediction application module 145. The collection module 140 can be implemented by Reporter, the first prediction module 141 can be implemented by a Prophet model, the second prediction module 142 can be implemented by a FTRL model, and the resource deployment module 143 can be implemented by Yarn.
[0174] Reporter periodically collects monitoring indexes and sends stream events to the Flink training application module 144 for modeling after injecting a time stamp, and sends stream events to the Flink prediction application module 145 for resource prediction. In Step 2, the Flink training application module 144 processes monitoring indexes in real time to construct training samples, uses the latest batch of training samples accumulated by a window function to update a Prophet model, and continuously updates weight coefficients of a FTRL model for each event stream. The Flink prediction application module 145 takes out the Prophet model and the FTRL model trained by the Flink training application module 144 to perform resource prediction, and sends a command to Yarn to adjust application resources according to a decision result of the prediction.
[0175] The deployment device of the embodiment is also used to implement the above-described deployment method, and the specific working principle can be referred to the above-described deployment method.
[0176] The application further provides an electronic device, as shown inFigure 15 As shown in the figure, the image coding device 150 in this embodiment comprises a processor 151, a memory 152, an input and output device 153 and a bus 154.
[0177] The processor 151, the memory 152 and the input and output device 153 are connected with the bus 154 respectively, the memory 152 stores program data, and the processor 151 is used to execute the program data to realize the above-mentioned job resource allocation method.
[0178] The processor 151 also realizes the image coding method of the above-mentioned embodiment when executing the program data.
[0179] In this embodiment, the processor 151 can also be called CPU. The processor 151 can be an integrated circuit chip with signal processing capability. The processor 151 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 151 can also be any conventional processor or the like.
[0180] The present application further proposes a computer readable storage medium, such as Figure 16 As shown in the figure, the computer readable storage medium 160 in this embodiment is used to store the program instructions 161 of the above-mentioned embodiment, and the program instructions 161 can be executed by the above-mentioned job resource allocation method. The program instructions 161 have been described in detail in the above method embodiment, and will not be described here.
[0181] The computer readable storage medium 160 in this embodiment can be but is not limited to a U disk, an SD card, a PD optical drive, a mobile hard disk, a large-capacity floppy disk drive, a flash memory, a multimedia memory card, a server and the like.
[0182] Different from the prior art, the present application can predict a plurality of first prediction values of the service volume index according to the current value of the monitoring index of the job point, then predict a plurality of second prediction values of the resource index based on the plurality of first prediction values of the service volume index, and allocate the resources of the job point based on the plurality of second prediction values, so that the present application can automatically allocate the resources of the job point according to the collected monitoring index of the job point, without manual experience guidance and manual annotation and other manual intervention, so as to improve the intelligent level and accuracy of resource allocation.
[0183] Further, the application is based on an open source tool to solve the technical solutions of fine-grained job monitoring and early warning, abnormality detection and resource dynamic scheduling; and the application uses a Flink stream processing framework+FTRL model to realize millisecond-level model updating and prediction, and to refine the prediction granularity to below the job level (capture the code logic complexity characteristics of the job itself), and solves the time out-of-order problem.
[0184] In addition, the above functions, if realized in the form of software functions and sold or used as independent products, can be stored in a mobile terminal readable storage medium, that is, the application also provides a storage device storing program data, the program data can be executed to realize the method of the above embodiments, and the storage device can be, for example, a U disk, an optical disk, a server, etc. That is, the application can be embodied in the form of a software product, which includes a plurality of instructions for causing an intelligent terminal to execute all or part of the steps of the method described in each embodiment.
[0185] In the description of the application, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the application. In this specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in the specification and the features of different embodiments or examples without contradiction.
[0186] In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one feature. In the description of the application, the meaning of "multiple" is at least two, for example, two, three, etc., unless otherwise specifically limited.
[0187] Any process or method descriptions in flow charts or otherwise described herein can be understood as representing code modules, segments, or portions of code that include one or more executable instructions for implementing specific logic functions or other processes, and the various embodiments of the application include additional implementations in which the order of execution is different, in which other code modules are used, in which other structures are used, in which not all code modules are executed, and in which not all of the functions are performed, and in which the functions are performed in different orders, all of which are understood to be within the scope of the application.
[0188] The logic and / or steps represented in the flow diagrams and / or otherwise described herein, for example, can be embodied in non-transitory computer-readable media, executed by an instruction execution system, apparatus, or device, such as a personal computer, server, network device, or other computing / processing apparatuses that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. In this regard, the "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can comprise any one of the following: electric connections (electronic devices), a portable computer diskette (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium upon which the program is printed, as the program can be electronically captured, via for instance an optical scanner, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory.
[0189] The above description is merely illustrative of the application, and is not intended to limit the scope of the application. Any equivalent techniques not explicitly described herein are also intended to be within the scope of the application.
Claims
1. A method of allocating a job resource, characterized by, The method comprises: collecting a current value of a monitoring index of a job site, wherein the monitoring index comprises a resource index and a traffic index; predicting a plurality of first predicted values of the traffic index based on the current value of the traffic index; predicting a plurality of second predicted values of the resource index based on the plurality of first predicted values; allocating resources of the job site based on the plurality of second predicted values; wherein the allocating resources of the job site based on the plurality of second predicted values comprises: obtaining a maximum value and a minimum value from the plurality of second predicted values; in response to the current value of the resource index being less than the minimum value, obtaining a sum of adjustment amounts of resources of all the job sites that need to increase resources in a cluster; in response to the sum of the adjustment amounts being greater than an available amount of resources of the cluster, dividing the available amount of resources into different groups of preset amounts multiple times, wherein a sum of the preset amounts in each group of the preset amounts is the available amount of resources; predicting a sum of traffic amounts corresponding to the job sites that need to increase resources after each group of the preset amounts of resources is allocated, wherein the job sites that need to increase resources correspond to the preset amounts in each group of the preset amounts one by one; obtaining a preset amount in a group of the preset amounts corresponding to a maximum value of the sum of the traffic amounts as an actual adjustment amount of the resources of the job sites that need to increase resources; adjusting the resources of the job sites that need to increase resources according to the actual adjustment amount.
2. The method of claim 1, wherein, The allocating resources of the job site based on the plurality of second predicted values further comprises: in response to the current value of the resource index being greater than the maximum value, reducing the amount of resources of the job site to the minimum value.
3. The method of claim 2, wherein, The increasing the amount of resources of the job site to the maximum value in response to the current value of the resource index being less than the minimum value comprises: in response to the current value of the resource index being less than the minimum value, obtaining a first difference value between the minimum value and the current value of the resource index; in response to the first difference value being greater than a first threshold, increasing the amount of resources of the job site to the maximum value; The reducing the amount of resources of the job site to the minimum value in response to the current value of the resource index being greater than the maximum value comprises: in response to the current value of the resource index being greater than the maximum value, obtaining a second difference value between the current value of the resource index and the maximum value; in response to the second difference value being greater than the first threshold, reducing the amount of resources of the job site to the minimum value.
4. The method of claim 1, wherein, The monitoring index further comprises an environmental information index, a context index, and a derived index, and the predicting a plurality of second predicted values of the resource index based on the plurality of first predicted values comprises: obtaining a correlation model between the resource index and the traffic index, the environmental information index, the context index, and the derived index; using the correlation model to predict a plurality of second predicted values of the resource index corresponding to the plurality of first predicted values.
5. The method of claim 4, wherein, The predicting a plurality of second predicted values of the resource index based on the plurality of first predicted values further comprises: acquire a training sample by using the current value of the resource indicator, the current value of the traffic volume indicator, the current value of the environment information indicator, the current value of the context indicator and the current value of the derived indicator; train and update the correlation model by using the training sample.
6. The method of claim 5, wherein, The training and updating of the correlation model by using the training sample comprises: acquiring an application number of the job point; selecting a correlation model corresponding to the application number from a plurality of preset correlation models; training and updating the correlation model corresponding to the application number by using the training sample.
7. The method of claim 5, wherein, The acquisition of the current value of the job point monitoring indicator comprises: acquiring global clock heartbeat information, and acquiring the current value of the job point monitoring indicator under the control of the heartbeat information; generating time identification information corresponding to the heartbeat information; associating the time identification information with the current value of the monitoring indicator; The acquisition of the training sample by using the current value of the resource indicator, the current value of the traffic volume indicator, the current value of the environment information indicator, the current value of the context indicator and the current value of the derived indicator comprises: forming the current value of the resource indicator, the current value of the traffic volume indicator, the current value of the environment information indicator, the current value of the context indicator and the current value of the derived indicator associated with the same time identification information into a same training sample.
8. The method of formulating according to any one of claims 1 to 7, wherein, Further comprising: calculating the difference between the current value of the resource indicator and the second predicted value; in response to the difference being greater than a second threshold, the difference being less than a third threshold or a plurality of the differences gradually increasing, generating an abnormal event.
9. The method of claim 8, wherein, Further comprising: acquiring the mean of a plurality of the second predicted values; setting the second threshold and the third threshold based on the mean.
10. A device for allocating operational resources, characterized in that, Comprise: a collection module, configured to collect the current value of the job point monitoring indicator, wherein the monitoring indicator comprises a resource indicator and a traffic volume indicator; a first prediction module, connected with the collection module, configured to predict a plurality of first predicted values of the traffic volume indicator based on the current value of the traffic volume indicator; a second prediction module, connected with the first prediction module, configured to predict a plurality of second predicted values of the resource indicator based on the plurality of first predicted values; a resource allocation module, connected with the second prediction module, configured to allocate resources of the job point based on the plurality of second predicted values; wherein the allocation of the resources of the job point based on the plurality of second predicted values comprises: acquiring a maximum value and a minimum value from the plurality of second predicted values; in response to the current value of the resource indicator being less than the minimum value, acquiring the sum of the adjustment amounts of the resources of all the job points requiring resource increase in the cluster; in response to the sum of the adjustment amounts being greater than the available amount of resources of the cluster, dividing the available amount of resources into different groups of preset amounts multiple times, wherein the sum of a plurality of preset amounts in each group of preset amounts is the available amount of resources; predicting the sum of the traffic volumes corresponding to the job points requiring resource increase after the allocation of each group of preset amounts of resources, wherein the job points requiring resource increase correspond to a plurality of preset amounts in each group of preset amounts one by one; The preset quantity in the set of preset quantities corresponding to a maximum value of the sum of the service quantities is an actual adjustment quantity of the resource of the job point requiring an increase in resource; The resource of the job point requiring an increase in resource is adjusted according to the actual adjustment quantity.
11. An electronic device, comprising: The memory is configured to store program data, and the processor is configured to execute the program data to implement the method for allocating job resources according to any one of claims 1 to 9.
12. A computer storage medium, characterized in that, The memory is configured to store program data, and the processor is configured to execute the program data to implement the method for allocating job resources according to any one of claims 1 to 9. The memory is configured to store program data, and the processor is configured to execute the program data to implement the method for allocating job resources according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method and device for network adjustment
CN103491556A
Resource occupancy data prediction method, electronic device, and storage medium
CN109032914A