Resource quota adjustment method and device, electronic equipment and storage medium
By combining predictive models with historical QPS data and event information, the resource quotas of the LLM service platform are dynamically adjusted, which solves the problems of service crashes and resource waste caused by sudden traffic surges and achieves efficient and reasonable resource management.
Patent Information
- Application Number
- CN202511068983.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, privately deployed LLM service platforms are prone to service crashes when faced with sudden traffic surges due to either excessively low resource quotas leading to the rejection of legitimate requests or excessively high resource quotas leading to service failures. Furthermore, static quota mechanisms are difficult to cope with the impact of model response latency.
By utilizing a pre-trained prediction model, combined with historical QPS data and event characteristics, resource quotas are dynamically adjusted to predict QPS values in future time periods, thus achieving dynamic adjustment of resource quotas.
It effectively handles sudden traffic surges, avoids service crashes or resource waste, and achieves efficient and reasonable responses to sudden traffic scenarios, ensuring that services do not crash during peak periods and save resources during off-peak periods.
Smart Images

Figure CN120909790A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular, to a resource quota adjustment method and device, electronic equipment and storage medium. BACKGROUND
[0002] LLM (Large Language Model) service platform as the hub connecting enterprise business and AI capability, its core value lies in reducing AI engineering threshold, through visual tools to realize the whole process management from prompt word design, multi-model scheduling to service deployment. The platform is usually divided into two types of service mode: 1, workflow service: through drag and drop nodes to build complex logic chain, integrate conditional branch, model interaction, knowledge retrieval and other modules; 2, Agent service: define agent behavior based on structured Prompt, flexible call tool set or private knowledge base.
[0003] At present, the resource management of the LLM service platform deployed privately mainly adopts a static quota allocation mechanism. The mechanism mainly sets fixed resource quotas for each service of the LLM service platform. If the resource quota is set too low, it is easy to cause legitimate request rejection during the business peak period, and if the resource quota is set too high, it may cause service collapse. Moreover, due to the influence of model response delay, it is difficult to respond to sudden traffic by using the static quota mechanism. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a resource quota adjustment method, device, electronic equipment and storage medium, which dynamically adjusts the resource quota of the service to solve the problem of service collapse or resource waste caused by sudden traffic.
[0005] In order to achieve the above purpose, the technical solutions adopted by the embodiments of the present application are as follows: In a first aspect, the present application provides a resource quota adjustment method, which comprises: obtaining a QPS time sequence vector according to the query rate per second (QPS) of the service at each time point in a first time period in the past; obtaining an event feature vector according to the events associated with each time point of the service in the first time period in the past; processing the QPS time sequence vector and the event feature vector by using a pre-trained prediction model to obtain the QPS prediction value of the service at each time point in a second time period in the future; dynamically adjusting the resource quota of the service according to the QPS prediction value of the service at each time point in the second time period in the future.
[0006] In an optional implementation, the processing of the QPS time sequence vector and the event feature vector by using the pre-trained prediction model to obtain the QPS prediction value of each time point of the service in a second time period in the future includes: performing feature extraction on the QPS time sequence vector to obtain a QPS time sequence feature vector; performing feature transformation on the event feature vector to obtain an event feature transformation vector; performing feature fusion on the QPS time sequence feature vector and the event feature transformation vector, and predicting the QPS prediction value of each time point of the service in a second time period in the future according to the obtained fusion feature vector.
[0007] In an optional implementation, the processing of the QPS time sequence feature vector and the event feature transformation vector to obtain the QPS prediction value of each time point of the service in a second time period in the future includes: performing feature fusion on the QPS time sequence feature vector and the event feature transformation vector to obtain a fusion feature vector; extracting a time sequence dependency relationship in the fusion feature vector, obtaining a hidden state corresponding to each time point in the first time period based on the time sequence dependency relationship; mapping the hidden state corresponding to each time point in the first time period to the QPS prediction value of each time point of the service in a second time period in the future.
[0008] In an optional implementation, the obtaining of the event feature vector according to the events associated with each time point of the service in a first time period in the past includes: when the service platform where the service is located is in a cold start state or the type of the events associated with each time point of the service in a first time period in the past is a new type, generating an event feature vector according to a default feature value; when the service platform where the service is located is not in a cold start state and the type of the events associated with each time point of the service in a first time period in the past is not a new type, performing embedding processing on the type of the events associated with each time point of the service in a first time period in the past to obtain an event type vector; obtaining a time position vector according to a current time and the start and end times of the events associated with each time point of the service in a first time period in the past; the time position vector represents a relative position relationship between each time point in the first time period and the start and end times of the events; merging the event type vector and the time position vector to obtain an event feature vector.
[0009] In an optional implementation, the dynamically adjusting the resource quota of the service according to the QPS prediction values of the service at the time points in the second time period in the future comprises: For each service, determining a maximum value from the QPS prediction values of the service at the time points in the second time period in the future, and taking the maximum value as a predicted peak value; calculating a basic quota according to the predicted peak value and a preset coefficient; determining the resource quota of the service according to the basic quota, a preset lower limit value corresponding to the service, and a preset upper limit value corresponding to a service platform where the service is located.
[0010] In an optional implementation, the method further comprises: calculating an error between the QPS actual value and the QPS prediction value corresponding to the current time point; if the error is less than a first preset value, recording the QPS time sequence vector, the event feature vector, and the QPS actual value of the current time point; if the error is greater than or equal to the first preset value and less than a second preset value, determining the QPS time sequence vector, the event feature vector, and the QPS actual value corresponding to the current time point as incremental update data, and updating parameters of the prediction model according to the incremental update data obtained in each incremental update period; if the error is greater than or equal to the second preset value, updating the parameters of the prediction model according to the QPS time sequence vector, the event feature vector, and the QPS actual value obtained at the current time point and before the current time point.
[0011] In an optional implementation, the method further comprises: determining a maximum QPS prediction value from the QPS prediction values of the service at the time points in the second time period in the future; calculating a growth rate of the maximum QPS prediction value relative to a QPS actual value of a current time point; if the growth rate and the maximum QPS prediction value meet a first trigger condition, sending an early warning information; the first trigger condition is that the growth rate is greater than a third preset value, and the maximum QPS prediction value is greater than a first QPS threshold value; the first QPS threshold value is determined from the obtained QPS actual values; If the growth rate and the maximum QPS predicted value meet a second trigger condition for the first time, the number of requests processed per second of the service is limited according to a first throttling ratio to reduce the resource consumption of the service; if the growth rate and the maximum QPS predicted value meet the second trigger condition for multiple times in succession, the number of requests processed per second of the service is limited according to an adjusted throttling ratio each time, the adjusted throttling ratio being obtained by increasing a set ratio on the basis of a current throttling ratio; the second trigger condition is that the growth rate is greater than a fourth preset value and the maximum QPS predicted value is greater than a second QPS threshold; the second QPS threshold is determined from the obtained QPS actual value, the second QPS threshold being greater than the first QPS threshold, and the fourth preset value being greater than the third preset value; If the number of times of triggering of the second trigger condition reaches a preset number of times and no manual confirmation of traffic surge normal information is received, the number of requests processed per second of the service is limited according to a second throttling ratio.
[0012] In a second aspect, the present application provides a resource quota adjustment device, the device comprising: a data acquisition module configured to obtain a QPS time sequence vector according to the query rate per second (QPS) of a service at each time point in a past first time period, and obtain an event feature vector according to events associated with each time point of the service in the past first time period; a prediction module configured to process the QPS time sequence vector and the event feature vector by using a pre-trained prediction model to obtain QPS predicted values of the service at each time point in a future second time period; an execution module configured to dynamically adjust the resource quota of the service according to the QPS predicted values of the service at each time point in the future second time period.
[0013] In a third aspect, the present application provides an electronic device comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the resource quota adjustment method according to any one of the preceding embodiments.
[0014] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps of the resource quota adjustment method according to any one of the preceding embodiments.
[0015] The resource quota adjustment method, device, electronic equipment and storage medium provided by the embodiment of the present application, the method obtains a QPS time sequence vector according to the QPS of the service at each time point in a first time period in the past; obtains an event feature vector according to the events associated with each time point of the service in the first time period in the past; processes the QPS time sequence vector and the event feature vector by using a pre-trained prediction model to obtain the QPS prediction value of the service at each time point in a second time period in the future; and dynamically adjusts the resource quota of the service according to the QPS prediction value of the service at each time point in the second time period in the future. By combining the historical QPS data and event information, using the prediction model to predict the future QPS value in advance, and then dynamically adjusting the resource quota of the service according to the prediction result, efficient and reasonable response in the burst traffic scenario is realized, thereby effectively dealing with the service collapse or resource waste problem caused by the burst traffic.
[0016] In order to make the above objectives, characteristics and advantages of the present application more apparent, the following preferred embodiments are specifically described below, and the accompanying drawings are referred to for detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation to the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0018] Figure 1 Fig. 1 shows a flow diagram of the resource quota adjustment method provided by the embodiment of the present application; Figure 2 Fig. 2 shows a network structure diagram of the prediction model; Figure 3 Fig. 3 shows another flow diagram of the resource quota adjustment method provided by the embodiment of the present application; Figure 4 Fig. 4 shows another flow diagram of the resource quota adjustment method provided by the embodiment of the present application; Figure 5 Fig. 5 shows a functional module diagram of the resource quota adjustment device provided by the embodiment of the present application; Figure 6 Fig. 6 shows a block diagram of the electronic equipment provided by the embodiment of the present application.
[0019] Fig. 1 shows a block diagram of the electronic equipment provided by the embodiment of the present application. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations.
[0021] Therefore, the detailed description of the embodiments of the present application provided below in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0022] It should be noted that the relational terms such as "first" and "second" and the like are used only to distinguish one entity or operation from another, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus including a series of elements includes not only those elements, but also other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus including the element.
[0023] In view of the problem that the service collapse or resource waste caused by burst traffic cannot be coped with by using the static quota allocation mechanism in the prior art, embodiments of the present application provide a resource quota adjustment method and device, an electronic device and a storage medium. The method comprises the following steps: obtaining a QPS (Queries Per Second) time sequence vector according to QPSs of each time point of a service in a first time period in the past; obtaining an event feature vector according to events associated with each time point of the service in the first time period in the past; processing the QPS time sequence vector and the event feature vector by using a pre-trained prediction model to obtain QPS prediction values of each time point of the service in a second time period in the future; and dynamically adjusting a resource quota of the service according to the QPS prediction values of each time point of the service in the second time period in the future. By combining historical QPS data and event information, using a prediction model to predict the QPS value in the future in advance, and then dynamically adjusting the resource quota of the service according to the prediction result, efficient and reasonable response in a burst traffic scenario is realized, so that the problem of service collapse or resource waste caused by burst traffic can be effectively coped with.
[0024] Hereinafter, the embodiments of the present application will be described in detail with reference to the accompanying drawings.
[0025] Please refer to Figure 1 A flowchart of a resource quota adjustment method provided by an embodiment of the present application is shown in FIG. 1. It should be noted that the resource quota adjustment method of the present application is not limited to the order of the steps shown in FIG. 1. In other embodiments, the order of some steps of the resource quota adjustment method of the present application can be exchanged according to actual needs, or some steps can be omitted or deleted. The resource quota adjustment method can be applied to electronic devices such as notebook computers, tablet computers, personal computers (PCs), servers, etc. Hereinafter, the specific process shown in FIG. 1 will be described in detail. Figure 1 Figure 1
[0026] In step S101, a QPS time series vector is obtained according to the QPS of the service at each time point in the past first time period.
[0027] In this embodiment, the QPS can be understood as the number of requests processed per second by the service. The request processed by the service can be a service call request for calling the service, or other types of requests, which are not limited in this embodiment.
[0028] In this embodiment, the data collection period and the first time period can be set according to actual conditions. Assuming that the first time period is 24 hours and the data collection period is 5 minutes, the real-time QPS of the service will be collected every 5 minutes, so that the past 24 hours will correspond to the QPS actual values of 288 time points, thereby obtaining a QPS time series vector of 288 x 1 dimension.
[0029] In step S102, an event feature vector is obtained according to the events associated with each time point in the past first time period.
[0030] In this embodiment, the electronic device collects QPS at each time point and also collects event registration information, determines the events associated with each service at the time point according to the event registration information, and encodes the events associated with each time point in the past first time period to obtain an event feature vector. External events such as marketing activities, holiday promotions, sudden events (such as public opinion outbreak), various activities or events of live broadcast platforms, etc. will trigger traffic surge.
[0031] In step S103, the QPS time series vector and the event feature vector are processed by using a pre-trained prediction model to obtain the QPS prediction value of each time point in the future second time period.
[0032] In the embodiment, the QPS prediction value can be understood as a predicted number of requests processed per second of the service. The prediction model combines historical QPS data and event information to predict the future QPS value (QPS prediction value) in advance. For example, by processing the QPS data and event information of the past 288 time points (24 hours) in real time through the prediction model, the QPS prediction value of the future 12 time points (1 hour) is predicted.
[0033] In step S104, the resource quota of the service is dynamically adjusted according to the QPS prediction value of each time point of the service in the second time period in the future.
[0034] In the embodiment, the resource quota of the service can be understood as the upper limit of the resource usage of the service. The resource can include GPU (Graphics Processing Unit, graphics processor) computing resource, memory resource, network resource, etc. The greater the number of requests processed per second of the service, the greater the resource consumption of the service, that is, the number of requests processed per second of the service is positively correlated with the resource consumption of the service. By predicting the QPS prediction value of each time point of the service in the second time period in the future through the prediction model, the number of requests that the service can process at each time point in the second time period in the future can be known, so the resource quota of the service can be more accurately determined in combination with the QPS prediction value, and the dynamic adjustment of the resource quota is realized, so that the access pressure can be withstood during the traffic peak period, and the service collapse is avoided; during the traffic low peak period, the resource consumption is saved, and waste is avoided, and automatic and intelligent resource management is realized.
[0035] It can be seen that the resource quota adjustment method provided by the embodiment combines historical QPS data and event information, uses a prediction model to predict the future QPS value in advance, and then dynamically adjusts the resource quota of the service according to the prediction result, so that efficient and reasonable response to the burst traffic scenario is realized, thereby effectively dealing with the service collapse or resource waste problem caused by the burst traffic.
[0036] In an embodiment, the step S103 can include: performing feature extraction on the QPS time sequence vector to obtain a QPS time sequence feature vector; performing feature transformation on the event feature vector to obtain an event feature transformation vector; performing feature fusion on the QPS time sequence feature vector and the event feature transformation vector, and predicting the QPS prediction value of each time point of the service in the second time period in the future according to the obtained fusion feature vector.
[0037] In the embodiment, please refer to Figure 2 The prediction model includes a QPS processing module, an event processing module, and a fusion prediction module. The QPS processing module processes historical QPS data, and performs feature extraction on the QPS time sequence vector through the QPS processing module to obtain a QPS time sequence feature vector. The QPS processing module can include a first input layer, a convolution layer, a batch normalization layer, and a first activation layer; the QPS time sequence vector is received through the first input layer, and the QPS time sequence vector is input to the convolution layer; the QPS time sequence vector is extracted through the convolution layer to obtain a feature extraction result; the feature extraction result is normalized through the batch normalization layer to obtain a normalized processing result; and the normalized processing result is nonlinearly transformed through the first activation layer to obtain the QPS time sequence feature vector.
[0038] In an example, the first input layer receives 24 hours of historical QPS data (with an interval of 5 minutes, a total of 288 time points); local features are extracted through a one-dimensional convolution layer (a convolution kernel size of 5 and 16 filters), and the convolution operation adopts a "same" padding mode to keep the time dimension unchanged. The first activation layer uses a ReLU activation function to enhance the nonlinear expression ability. The batch normalization layer is used to stabilize the training process, accelerate the convergence, and reduce the sensitivity to the initial weight.
[0039] In the embodiment, the event processing module is used for processing event related information, and performs feature transformation on the event feature vector through the event processing module to obtain an event feature transformation vector. The event processing module can include a second input layer, a fully connected layer, and a second activation layer, the event feature vector is received through the second input layer, and the event feature vector is input to the fully connected layer; the event feature vector is transformed through the fully connected layer, and the transformed feature value is mapped to a preset range to obtain the final event feature transformation vector.
[0040] In one example, the second input layer receives a 5-dimensional event feature vector of 288 time points time-aligned with the QPS data. The 5-dimensional feature is mapped to a 16-dimensional space by a fully connected layer (Dense), and the second activation layer uses a tanh activation function to compress the feature values to the range of [-1, 1]. In addition, to enhance the generalization ability of the model, a spatial Dropout layer (SpatialDropout1D) can also be added to the structure of the event processing module to randomly discard the entire feature channel with a probability of 20% to prevent overfitting. It can be understood that the spatial Dropout layer is turned on during the model training stage; and the spatial Dropout layer is turned off during the model service stage.
[0041] In one embodiment, the above step of feature fusion of the QPS time series feature vector and the event feature transformation vector, and predicting the QPS prediction value of each time point of the service in the second time period in the future according to the obtained fusion feature vector can specifically include: The QPS time series feature vector and the event feature transformation vector are fused to obtain a fusion feature vector; the time series dependency in the fusion feature vector is extracted, and the corresponding hidden state of each time point in the first time period is obtained based on the time series dependency; and the corresponding hidden state of each time point in the first time period is mapped to the QPS prediction value of each time point of the service in the second time period in the future.
[0042] In this embodiment, the QPS time series feature vector and the event feature transformation vector can be fused by the fusion prediction module, and the QPS prediction value of each time point of the service in the second time period in the future can be predicted according to the obtained fusion feature vector. The fusion prediction module can include a feature fusion layer, a gated recurrent unit and a fully connected layer; the QPS time series feature vector and the event feature transformation vector are fused by the feature fusion layer to obtain a fusion feature vector; the time series dependency in the fusion feature vector is extracted by the gated recurrent unit, and the corresponding hidden state of each time point in the first time period is obtained based on the time series dependency; and the corresponding hidden state of each time point in the first time period is mapped to the QPS prediction value of each time point of the service in the second time period in the future by the fully connected layer.
[0043] In one example, when the feature fusion layer fuses the QPS time series feature vector and the event feature transformation vector, the QPS time series feature vector and the event feature transformation vector are fused by element-level addition, which retains the consistency of the feature space of the two channels, and allows the model to automatically learn the complementary relationship between the two features.
[0044] The fusion feature vector output by the feature fusion layer is input to a Gated Recurrent Unit (GRU) for modeling time sequence dependence. The GRU structure can effectively capture the time sequence dependence in the fusion feature vector. In the GRU, 288 16-dimensional features are input, the GRU includes 24 units, and each feature is processed in a loop, and each time 24-dimensional intermediate features are output. In addition, a 10% recurrent dropout is set, that is, 10% of the intermediate features are randomly set to zero during the loop processing to prevent overfitting. Finally, the complete time sequence is returned (return_sequences=True), that is, the intermediate features output each time are obtained, and finally a 288x24-dimensional feature (i.e., 288 time points correspond to hidden states) is obtained. In other embodiments, if the complete time sequence is not returned, the last obtained intermediate feature can be taken as the output of the GRU.
[0045] A dense full connection layer generates QPS prediction values for 12 time points within 1 hour in the future based on the output of the GRU (e.g., a 288x24-dimensional feature). The full connection layer can use a linear activation function to directly output the prediction values.
[0046] In an embodiment, the step S102 can include: when the service platform where the service is located is in a cold start state or the types of events associated with each time point of the service in the past first time period are new types, generating an event feature vector according to a default feature value; when the service platform where the service is located is not in a cold start state and the types of events associated with each time point of the service in the past first time period are not new types, performing embedding processing on the types of events associated with each time point of the service in the past first time period to obtain an event type vector; obtaining a time position vector according to the current time and the start and end times of the events associated with each time point of the service in the past first time period; the time position vector represents the relative position relationship between each time point in the first time period and the start and end times of the events; and merging the event type vector and the time position vector to obtain an event feature vector.
[0047] In the embodiment, the cold start state can be understood as the initial online operation of the service platform. The service platform can pre-store various event types, such as various anchor activities in the live broadcast platform, various Chinese events, various global events, and the like. If the type of the event associated with the service is different from the pre-stored event types, it is considered that the type of the event associated with the service is a new type. When the service platform is in a cold start state or a new type event (unknown event) occurs, the prediction model can not have complete prediction capability of event perception. At this time, the event processing branch can be initialized silently, that is, the event processing module is initialized to a "zero state" and a default feature value (such as 0) is output. In this way, the event feature vector obtained is a zero vector, so that the prediction model only uses historical QPS data for QPS prediction, which can not introduce additional influence, thereby ensuring the stability of the prediction model.
[0048] When the service platform where the service is located is not in a cold start state and the event associated with the service at each time point in the past first time period is not an unknown event, the event feature vector can be obtained based on the type and start and end time of the event.
[0049] For example, for the type of the associated event, a learnable embedding layer (Full Embedding) is used for embedding processing, so as to map the event type (such as the name of the event) into a 4-dimensional vector. For the start and end time of the associated event, a 1-dimensional time position vector is obtained according to "position = 2 x (current time - event start time) / event total duration - 1"; wherein the value range of position is [-5, 5], [-5, 0) represents a 5-time-length preheating phase before the start of the event, [0, 1) represents the activity, [1, 5) represents a 4-time-length tailing phase after the end of the event, and if the calculated position exceeds [-5, 5], zero value is filled. The 4-dimensional event type vector and the 1-dimensional time position vector are merged, and a 5-dimensional event feature vector is obtained. That is, through the above event space-time coding scheme, a corresponding 5-dimensional feature vector can be generated for each 5-minute time point in the past 24 hours, and finally a 288 x 5-dimensional vector is obtained.
[0050] In an embodiment, the step S104 can include: For each service, a maximum value is determined from the QPS prediction values of each time point of the service in the future second time period, and the maximum value is taken as a prediction peak value; a basic quota is calculated according to the prediction peak value and a preset coefficient; and the resource quota of the service is determined according to the basic quota, a preset lower limit value corresponding to the service, and a preset upper limit value corresponding to the service platform where the service is located.
[0051] In the embodiment, the preset coefficient, the preset lower limit value corresponding to the service, and the preset upper limit value corresponding to the service platform where the service is located can be set according to actual conditions, and are not limited in the embodiment.
[0052] For example, the dynamic adjustment process of the service quota can be represented as follows: Predicted peak value = max (QPS predicted values of future 12 time points); Basic quota = predicted peak value x 1.25 / number of requests supported by a single GPU; wherein the number of requests supported by a single GPU can be set by a platform user; Final resource quota = max (preset lower limit value corresponding to the service, min (basic quota, preset upper limit value corresponding to the service platform)).
[0053] In an implementation, referring to Figure 3 The resource quota adjustment method provided by the embodiment of the application can further include: Step S301, calculating an error between a QPS actual value and a QPS predicted value corresponding to a current time point.
[0054] In the embodiment, the error between the QPS actual value actual and the QPS predicted value predicted can be calculated in real time: ε = |predicted - actual|, and the value range of the error ε is determined. If the error ε is less than a first preset value, step S302 is performed; if the error ε is greater than or equal to the first preset value and less than a second preset value, step S303 is performed; and if the error ε is greater than or equal to the second preset value, step S304 is performed.
[0055] The first preset value and the second preset value can be set according to actual conditions. For example, the first preset value is 10%, and the second preset value is 25%.
[0056] Step S302, if the error is less than the first preset value, the QPS time sequence vector, the event feature vector, and the QPS actual value of the current time point are recorded.
[0057] For example, when the error ε is less than 10%, only the data is recorded.
[0058] Step S303, if the error is greater than or equal to the first preset value and less than the second preset value, the QPS time sequence vector, the event feature vector, and the QPS actual value corresponding to the current time point are determined as incremental update data, and the parameters of the prediction model are updated according to the incremental update data obtained in each incremental update period.
[0059] For example, when 10% ≤ ε < 25%, the QPS time sequence vector, the event feature vector and the QPS actual value corresponding to the current time point are determined as the incremental update data, and finally the parameters of the prediction model are updated according to the incremental update data obtained in each incremental update period. That is, the prediction model is incrementally learned using a small amount of error-related data, but the frequency of execution is not real-time, but at most once in an incremental update period. For example, once an hour.
[0060] In step S304, if the error is greater than or equal to the second preset value, the parameters of the prediction model are updated according to the QPS time sequence vector, the event feature vector and the QPS actual value obtained at the current time point and before the current time point.
[0061] For example, when the error ε ≥ 25%, the parameters of the prediction model are updated using the QPS time sequence vector, the event feature vector and the QPS actual value obtained at the current time point and before the current time point, that is, the prediction model is reinforced learning using a large amount of historical data.
[0062] It should be noted that in actual services, only historical data in a certain period of time (for example, one month) can be used for reinforcement learning of the prediction model, and all historical data do not have to be used.
[0063] It can be seen that the resource quota adjustment method provided by the embodiment of the application can calculate the absolute error between the predicted value and the actual value in real time during the service of the prediction model, and dynamically adjust the learning strategy according to the range of the error. If the error is less than 10%, only the data is recorded, and the model is not updated; if the error is between 10% and 25%, the prediction model is incrementally learned, and each incremental learning only processes the data related to the error, and at most once an hour; if the error exceeds 25%, the reinforcement learning process is immediately started, and the full amount of fine tuning is performed based on the historical data. The online learning mechanism drives model updating through error feedback, ensures that the prediction model can quickly adapt to new traffic patterns and event features, thereby improving the prediction accuracy and the rationality of quota adjustment, and enhancing the self-adaptive ability and stability of the prediction model.
[0064] In an embodiment, referring to Figure 4 The resource quota adjustment method provided by the embodiment of the application can further include: In step S401, the maximum QPS predicted value is determined from the QPS predicted values at each time point in a second time period in the future.
[0065] In step S402, the growth rate of the maximum QPS predicted value relative to the QPS actual value at the current time point is calculated.
[0066] For example, the growth rate = (the maximum QPS predicted value - the actual QPS value at the current time point) / the actual QPS value at the current time point.
[0067] In step S403, if the growth rate and the maximum QPS predicted value meet a first trigger condition, the warning information is sent; the first trigger condition is that the growth rate is greater than a third preset value and the maximum QPS predicted value is greater than a first QPS threshold; the first QPS threshold is determined from the acquired actual QPS values.
[0068] In the embodiment, the acquired actual QPS values include the actual QPS value acquired at the current time point and the actual QPS value acquired before the current time point. The first QPS threshold can be the actual QPS value at the first preset percentile point among the acquired actual QPS values. The third preset value and the first preset percentile point can be set according to actual conditions. For example, the third preset value is 50%, the first preset percentile point is 75%, after the growth rate is calculated, it is determined whether the growth rate is greater than 50%; the actual QPS value corresponding to the current time point and the actual QPS value corresponding to the past period (such as one month) of the current time point are sorted in ascending order, and the maximum QPS predicted value is compared with the actual QPS value at the 75% percentile point (the first QPS threshold) in the ascending order sorting result; if the growth rate is greater than 50% and the maximum QPS predicted value is greater than the actual QPS value at the 75% percentile point, the warning information is sent, so as to manually determine whether to perform the expansion processing.
[0069] In step S404, if the growth rate and the maximum QPS predicted value meet a second trigger condition for the first time, the number of requests processed per second of the service is limited according to a first throttling ratio, so as to reduce the resource consumption of the service; if the growth rate and the maximum QPS predicted value meet the second trigger condition for multiple times in succession, each time the set ratio is added on the basis of the current throttling ratio, and the number of requests processed per second of the service is limited according to the adjusted throttling ratio; the second trigger condition is that the growth rate is greater than a fourth preset value and the maximum QPS predicted value is greater than a second QPS threshold; the second QPS threshold is determined from the acquired actual QPS values, the second QPS threshold is greater than the first QPS threshold, and the fourth preset value is greater than the third preset value.
[0070] In the embodiment, the second QPS threshold can be a QPS actual value located at a second preset percentile among the acquired QPS actual values. The fourth preset value, the second preset percentile, the first throttling ratio, and the set ratio can be set according to actual conditions. For example, the fourth preset value is 90%, the second preset percentile is 85%, the first throttling ratio is 30%, and the set ratio is 20%. When the growth rate is greater than 90% and the maximum QPS predicted value is greater than the QPS actual value located at the 85% percentile for the first time, the number of requests processed per second of the service is limited according to the throttling ratio of 30%, for example, 30% of new requests are rejected, and only 70% of requests are allowed to pass. When the growth rate is greater than 90% and the maximum QPS predicted value is greater than the QPS actual value located at the 85% percentile for multiple times in succession, the number of requests processed per second of the service is limited according to the throttling ratio adjusted on the basis of the current throttling ratio each time, for example, when the second trigger condition is met for the second time, the number of requests processed per second of the service is limited according to the throttling ratio of 50% (i.e., 30%+20%).
[0071] In step S405, if the number of times of triggering the second trigger condition reaches the preset number of times and no traffic surge normal information is received, the number of requests processed per second of the service is limited according to the second throttling ratio.
[0072] In the embodiment, the traffic surge normal information can be understood as the confirmation information input after the artificial judgment that the current predicted traffic surge is normal. The preset number of times and the second throttling ratio can be set according to actual conditions. For example, the preset number of times is 3 and the second throttling ratio is 80%. When the growth rate and the maximum QPS predicted value meet the second trigger condition for 3 times in succession, i.e., the number of times of triggering the second trigger condition reaches 3, and no one confirms that the current predicted traffic surge is normal, i.e., no traffic surge normal information is received, the number of requests processed per second of the service is limited according to the throttling ratio of 80%.
[0073] It can be understood that, since the number of requests processed per second of the service is positively correlated with the resource consumption of the service, limiting the number of requests processed per second of the service according to different throttling ratios can reduce the resource consumption of the service to different degrees, effectively avoids the situation that the resource is overloaded or even exhausted due to too many requests processed, and achieves the purpose of resource protection.
[0074] It can be seen that the resource quota adjustment method provided by the embodiment of the application provides a three-level fuse mechanism, mainly including three levels of trigger conditions and corresponding response actions, namely, first-level early warning, second-level flow limiting and third-level protection. When the first-level early warning mechanism is triggered, early warning information is sent to prompt manual judgment of whether expansion is needed. When the second-level flow limiting mechanism is triggered, a gradual flow limiting strategy is started, that is, the first time flow limiting is 30%, and if the flow limiting is continuously triggered, 20% is added to the current flow limiting ratio each time. The third-level protection mechanism is the highest level of protection, which is started when the second-level flow limiting mechanism is continuously triggered for three times and no manual intervention (no one confirms that the traffic surge is normal), and the service is degraded and 80% of the requests are limited. The three-level fuse mechanism can effectively balance user experience and system stability when a burst traffic comes, avoid service collapse, minimize the rejection of legitimate requests, improve resource utilization efficiency, and realize dynamic resource regulation and control in response to burst traffic impact.
[0075] In order to perform the corresponding steps in the above-mentioned embodiments and various possible manners, an implementation manner of a resource quota adjustment device is given below. Please refer to Figure 5 A functional module diagram of the resource quota adjustment device 600 provided by the embodiment of the application is shown in the figure. It should be noted that the resource quota adjustment device 600 provided by the embodiment has the same basic principles and technical effects as the above-mentioned embodiments, and for brief description, the part not mentioned in this embodiment can be referred to the corresponding content in the above-mentioned embodiments. The resource quota adjustment device 600 includes a data acquisition module 610, a prediction module 620 and an execution module 630.
[0076] The data acquisition module 610 is configured to obtain a QPS time sequence vector according to the QPS of each time point of the service in a first time period in the past, and obtain an event feature vector according to the events associated with each time point of the service in the first time period in the past.
[0077] It can be understood that the data acquisition module 610 can perform the above-mentioned step S101.
[0078] The prediction module 620 is configured to process the QPS time sequence vector and the event feature vector by using a pre-trained prediction model to obtain the QPS prediction value of each time point of the service in a second time period in the future.
[0079] It can be understood that the prediction module 620 can perform the above-mentioned step S102.
[0080] The execution module 630 is configured to dynamically adjust the resource quota of the service according to the QPS prediction value of each time point of the service in the second time period in the future.
[0081] It can be understood that the execution module 630 can perform the above-mentioned step S103.
[0082] Optionally, the data collection module 610 is specifically configured to generate the event feature vector according to the default feature value when the service platform where the service is located is in a cold start state or the type of the event associated with each time point in the first time period in the past by the service is a new type; when the service platform where the service is located is not in a cold start state and the type of the event associated with each time point in the first time period in the past by the service is not a new type, perform embedding processing on the type of the event associated with each time point in the first time period in the past by the service to obtain an event type vector; obtain a time position vector according to the current time and the start and end times of the event associated with each time point in the first time period in the past; the time position vector represents the relative position relationship between each time point in the first time period and the start and end times of the event; and combine the event type vector and the time position vector to obtain the event feature vector.
[0083] Optionally, the prediction module 620 is specifically configured to perform feature extraction on the QPS time sequence vector to obtain a QPS time sequence feature vector; perform feature transformation on the event feature vector to obtain an event feature transformation vector; perform feature fusion on the QPS time sequence feature vector and the event feature transformation vector, and predict the QPS prediction value of each time point in the second time period in the future by the service according to the obtained fusion feature vector.
[0084] Optionally, the prediction module 620 is further specifically configured to perform feature fusion on the QPS time sequence feature vector and the event feature transformation vector to obtain a fusion feature vector; extract a time sequence dependency relationship in the fusion feature vector, obtain the hidden state corresponding to each time point in the first time period based on the time sequence dependency relationship; and map the hidden state corresponding to each time point in the first time period to the QPS prediction value of each time point in the second time period in the future by the service.
[0085] Optionally, the execution module 630 is specifically configured to determine, for each service, a maximum value from the QPS prediction value of each time point in the second time period in the future by the service, and take the maximum value as a predicted peak value; calculate a basic quota according to the predicted peak value and a preset coefficient; and determine the resource quota of the service according to the basic quota, a preset lower limit value corresponding to the service, and a preset upper limit value corresponding to the service platform where the service is located.
[0086] Optionally, the execution module 630 is further configured to calculate an error between the QPS actual value corresponding to the current time point and the QPS predicted value; record the QPS time sequence vector, the event feature vector and the QPS actual value of the current time point if the error is less than a first preset value; determine the QPS time sequence vector, the event feature vector and the QPS actual value corresponding to the current time point as incremental update data if the error is greater than or equal to the first preset value and less than a second preset value, and update the parameters of the prediction model according to the incremental update data obtained in each incremental update period; and update the parameters of the prediction model according to the QPS time sequence vector, the event feature vector and the QPS actual value obtained at the current time point and before the current time point if the error is greater than or equal to the second preset value.
[0087] It can be understood that the execution module 630 can also perform steps S301-S304 described above.
[0088] Optionally, the execution module 630 is further configured to determine a maximum QPS predicted value from the QPS predicted values of each time point in a second time period in the future, calculate a growth rate of the maximum QPS predicted value relative to the QPS actual value of the current time point, send a warning information if the growth rate and the maximum QPS predicted value meet a first trigger condition, the first trigger condition being that the growth rate is greater than a third preset value and the maximum QPS predicted value is greater than a first QPS threshold, the first QPS threshold being determined from the obtained QPS actual values, limit the number of requests processed per second of the service according to a first throttling ratio to reduce the resource consumption of the service if the growth rate and the maximum QPS predicted value meet a second trigger condition for the first time, increase the throttling ratio by a set ratio each time based on the current throttling ratio to limit the number of requests processed per second of the service according to the adjusted throttling ratio if the growth rate and the maximum QPS predicted value meet the second trigger condition for multiple times in succession, the second trigger condition being that the growth rate is greater than a fourth preset value and the maximum QPS predicted value is greater than a second QPS threshold, the second QPS threshold being determined from the obtained QPS actual values, the second QPS threshold being greater than the first QPS threshold, and the fourth preset value being greater than the third preset value, and limit the number of requests processed per second of the service according to a second throttling ratio if the number of times of triggering the second trigger condition reaches a preset number of times and no manual confirmation of traffic surge normal information is received.
[0089] It can be understood that the execution module 630 can also perform steps S401-S405 described above.
[0090] Please refer to Figure 6A block schematic diagram of an electronic device 100 according to an embodiment of the present application is shown. The electronic device 100 includes a memory 110, a processor 120 and a communication module 130. The memory 110, the processor 120 and the communication module 130 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, the elements can be electrically connected to each other through one or more communication buses or signal lines.
[0091] The memory 110 is configured to store programs or data. The memory 110 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electric erasable programmable read only memory (EEPROM), etc.
[0092] The processor 120 is configured to read / write the data or programs stored in the memory 110 and perform corresponding functions. For example, when the computer program stored in the memory 110 is executed by the processor 120, the resource quota adjustment method disclosed in the above embodiments can be realized.
[0093] The communication module 130 is configured to establish a communication connection between the electronic device 100 and other devices through a network and to receive / transmit data through the network.
[0094] It should be understood that, Figure 6 The structure shown is only a structural schematic diagram of the electronic device 100. The electronic device 100 can further include more or less components than those shown in the above embodiments or have a different configuration from that shown in the above embodiments. Figure 6 The components shown in the above embodiments can be realized in hardware, software or a combination thereof. Figure 6 Figure 6 The embodiments of the present application further provide a computer readable storage medium having a computer program stored thereon. The computer program is executed by the processor 120 to realize the resource quota adjustment method disclosed in the above embodiments.
[0095] The embodiments of the present application further provide a computer readable storage medium having a computer program stored thereon. The computer program is executed by the processor 120 to realize the resource quota adjustment method disclosed in the above embodiments.
[0096] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can also be implemented by other manners. The apparatus embodiments described above are merely illustrative, for example, the flowcharts and block diagrams in the drawings show the possible implementation architecture, function and operation of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logic function. It should also be noted that in some alternative implementation manners, the functions noted in the blocks can also occur in different order from that noted in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can also be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for executing the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0097] In addition, each functional module in the embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0098] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0099] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for adjusting resource quotas, characterized in that, The method comprises: obtaining a QPS time sequence vector according to a query per second (QPS) of the service at each time point in a first time period in the past; obtaining an event feature vector according to events associated with the service at each time point in the first time period in the past; processing the QPS time sequence vector and the event feature vector by using a pre-trained prediction model to obtain a QPS prediction value of the service at each time point in a second time period in the future; dynamically adjusting a resource quota of the service according to the QPS prediction value of the service at each time point in the second time period in the future.
2. The method of claim 1, wherein, The processing of the QPS time sequence vector and the event feature vector by using the pre-trained prediction model to obtain the QPS prediction value of the service at each time point in the second time period in the future comprises: performing feature extraction on the QPS time sequence vector to obtain a QPS time sequence feature vector; performing feature transformation on the event feature vector to obtain an event feature transformation vector; performing feature fusion on the QPS time sequence feature vector and the event feature transformation vector, and predicting the QPS prediction value of the service at each time point in the second time period in the future according to a fusion feature vector obtained.
3. The method of claim 2, wherein, The performing of the feature fusion on the QPS time sequence feature vector and the event feature transformation vector, and the predicting of the QPS prediction value of the service at each time point in the second time period in the future according to the fusion feature vector comprises: performing feature fusion on the QPS time sequence feature vector and the event feature transformation vector to obtain a fusion feature vector; extracting a time sequence dependency relationship in the fusion feature vector, obtaining a hidden state corresponding to each time point in the first time period based on the time sequence dependency relationship; and mapping the hidden state corresponding to each time point in the first time period to the QPS prediction value of the service at each time point in the second time period in the future.
4. The method of claim 1, wherein, The obtaining of the event feature vector according to the events associated with the service at each time point in the first time period in the past comprises: when the service platform where the service is located is in a cold start state or the type of the events associated with the service at each time point in the first time period in the past is a new type, generating an event feature vector according to a default feature value; when the service platform where the service is located is not in a cold start state and the type of the events associated with the service at each time point in the first time period in the past is not a new type, performing embedding processing on the type of the events associated with the service at each time point in the first time period in the past to obtain an event type vector; obtaining a time position vector according to a current time and start and end times of the events associated with the service at each time point in the first time period in the past; the time position vector represents a relative position relationship between each time point in the first time period and the start and end times of the events; merging the event type vector and the time position vector to obtain an event feature vector.
5. The method of claim 1, wherein, The dynamically adjusting of the resource quota of the service according to the QPS prediction value of the service at each time point in the second time period in the future comprises: determining a maximum value from QPS predicted values of each service at different time points in a second time period in the future, and taking the maximum value as a predicted peak value; calculating a basic quota according to the predicted peak value and a preset coefficient; determining a resource quota of the service according to the basic quota, a preset lower limit value corresponding to the service, and a preset upper limit value corresponding to a service platform where the service is located.
6. The method of claim 1-5, wherein, The method further comprises: calculating an error between a QPS actual value and a QPS predicted value corresponding to a current time point; if the error is less than a first preset value, recording the QPS time sequence vector, the event feature vector, and the QPS actual value of the current time point; if the error is greater than or equal to the first preset value and less than a second preset value, determining the QPS time sequence vector, the event feature vector, and the QPS actual value corresponding to the current time point as incremental update data, and updating parameters of the prediction model according to the incremental update data obtained in each incremental update period; if the error is greater than or equal to the second preset value, updating the parameters of the prediction model according to the QPS time sequence vector, the event feature vector, and the QPS actual value obtained at the current time point and before the current time point.
7. The method of claim 1-5, wherein, The method further comprises: determining a maximum QPS predicted value from QPS predicted values of each service at different time points in a second time period in the future; calculating a growth rate of the maximum QPS predicted value relative to a QPS actual value of a current time point; if the growth rate and the maximum QPS predicted value meet a first trigger condition, sending an early warning information; the first trigger condition is that the growth rate is greater than a third preset value, and the maximum QPS predicted value is greater than a first QPS threshold value; the first QPS threshold value is determined from the obtained QPS actual values; if the growth rate and the maximum QPS predicted value meet a second trigger condition for the first time, limiting the number of requests processed per second of the service according to a first throttling ratio to reduce the resource consumption of the service; if the growth rate and the maximum QPS predicted value meet the second trigger condition for multiple times in succession, increasing a set ratio on the basis of a current throttling ratio each time, and limiting the number of requests processed per second of the service according to the adjusted throttling ratio; the second trigger condition is that the growth rate is greater than a fourth preset value, and the maximum QPS predicted value is greater than a second QPS threshold value; the second QPS threshold value is determined from the obtained QPS actual values, the second QPS threshold value is greater than the first QPS threshold value, and the fourth preset value is greater than the third preset value; if the number of times of triggering the second trigger condition reaches a preset number of times, and no artificial confirmation of traffic surge normal information is received, limiting the number of requests processed per second of the service according to a second throttling ratio.
8. A resource quota adjustment apparatus characterized by comprising: The device comprises: a data collection module, configured to obtain a QPS time sequence vector according to a query rate per second (QPS) of a service at different time points in a first time period in the past, and obtain an event feature vector according to events associated with the service at the different time points in the first time period in the past; a prediction module, configured to process the QPS time sequence vector and the event feature vector by using a pre-trained prediction model to obtain QPS prediction values of the service at each time point in a second time period in the future; an execution module, configured to dynamically adjust the resource quota of the service according to the QPS prediction values of the service at each time point in the second time period in the future.
9. An electronic device, comprising: a processor, a memory, and a computer program stored on the memory and executable on the processor, and the computer program, when executed by the processor, implements the steps of the resource quota adjustment method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, a computer program is stored on the computer readable storage medium, and the computer program, when executed by the processor, implements the steps of the resource quota adjustment method according to any one of claims 1-7.