Method and device for adjusting number of service instances, electronic equipment and computer program product

By obtaining real-time performance indicators of service instances and dynamically adjusting the number of service instances based on elastic scaling strategies, the problem of resource waste in traditional K8S deployment is solved, and resource utilization efficiency and system flexibility are improved.

CN120803610APending Publication Date: 2025-10-17JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510942463.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In traditional Kubernetes (K8S) deployments, the interface service module suffers from resource waste due to the fixed number of instances, which cannot be dynamically adjusted according to real-time load conditions, affecting service quality and resource utilization efficiency.

Method used

By obtaining real-time performance indicators of service instances, such as the number of call requests, CPU usage, and memory usage, we determine whether they comply with the elastic scaling policy and dynamically adjust the number of service instances according to the policy, introducing a dynamic elastic scaling mechanism based on real-time performance data.

Benefits of technology

It enables flexible adjustment of service instances according to load conditions, improves resource utilization efficiency and system flexibility, ensures interface service quality, and avoids resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803610A_ABST
    Figure CN120803610A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for adjusting the number of service instances, electronic equipment and a computer program product, and relates to the field of computers.The method comprises the steps that real-time performance indexes of a plurality of service instances of a current interface are obtained, and the real-time performance indexes at least comprise the number of call requests for the service instances, the CPU utilization rate and the memory utilization rate; according to the real-time performance index, whether the current interface accords with an elastic scaling strategy is determined, and the elastic scaling strategy is used for adjusting the number of service instances of the current interface; when it is determined that the current interface conforms to the elastic scaling strategy, the number of service instances of the current interface is adjusted according to the elastic scaling strategy; by adopting the scheme, the problem of resource waste caused by the fixed instance number of the interface service module in the traditional K8S deployment in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of storage, and in particular to a service instance quantity adjustment method and device, electronic equipment and computer program product. BACKGROUND

[0002] With the popularity of cloud computing, Kubernetes (K8S) has become a de facto standard for deploying and managing cloud-native applications. However, in K8S, traditional interface service modules often use a fixed number of instances for deployment. This approach is not sufficient when facing fluctuating loads. When the number of service instances is fixed, they cannot dynamically adjust according to real-time load conditions, resulting in idle resources during low load periods and resource shortages during high load periods, affecting service quality and user experience.

[0003] There are certain barriers between business modules and K8S, making service adjustment less flexible and immediate. Therefore, the current cloud platform service deployment method has obvious shortcomings in resource utilization efficiency and responding to real-time performance requirements.

[0004] In view of the problem of resource waste caused by the fixed number of instances of the interface service module in the traditional K8S deployment in the related art, no effective solution has been proposed so far. SUMMARY

[0005] The present application provides a service instance quantity adjustment method and device, electronic equipment and computer program product to at least solve the problem of resource waste caused by the fixed number of instances of the interface service module in the traditional K8S deployment in the related art.

[0006] The present application provides a service instance quantity adjustment method, comprising: obtaining real-time performance indicators of a plurality of service instances of a current interface, wherein the real-time performance indicators at least include: the number of call requests to the service instances, CPU usage rate, and memory usage rate; determining whether the current interface meets an elastic scaling policy according to the real-time performance indicators, wherein the elastic scaling policy is used to adjust the number of service instances of the current interface; and adjusting the number of service instances of the current interface according to the elastic scaling policy in the case that it is determined that the current interface meets the elastic scaling policy.

[0007] The application further provides a service instance quantity adjustment device, comprising: an acquisition module, configured to acquire real-time performance indexes of a plurality of service instances of a current interface, wherein the real-time performance indexes at least comprise: a number of calling requests to the service instances, CPU usage, and memory usage; a determination module, configured to determine whether the current interface meets an elastic scaling policy according to the real-time performance indexes, wherein the elastic scaling policy is used to adjust the quantity of service instances of the current interface; and an adjustment module, configured to adjust the quantity of service instances of the current interface according to the elastic scaling policy in a case where it is determined that the current interface meets the elastic scaling policy.

[0008] The application further provides an electronic device, comprising: a memory, configured to store a computer program; and a processor, configured to execute the computer program to implement the steps of any of the service instance quantity adjustment methods.

[0009] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of any of the service instance quantity adjustment methods.

[0010] The application further provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the steps of any of the service instance quantity adjustment methods.

[0011] According to the application, real-time performance indexes of a plurality of service instances of a current interface are acquired in real time, wherein the real-time performance indexes at least comprise: a number of calling requests to the service instances, CPU usage, and memory usage; whether the current interface meets an elastic scaling policy is determined according to the real-time performance indexes, wherein the elastic scaling policy is used to adjust the quantity of service instances of the current interface; and the quantity of service instances of the current interface is dynamically adjusted according to the elastic scaling policy in a case where it is determined that the current interface meets the elastic scaling policy; by using the above scheme, a dynamic elastic scaling mechanism based on real-time performance data is introduced, so that the service instances can be flexibly adjusted according to the load condition, thereby ensuring the service quality of the interface, and improving the resource utilization efficiency and the overall flexibility of the system; and the problem of resource waste caused by the fixed instance quantity of the interface service module in the traditional K8S deployment in the related art is further solved. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort.

[0013] Figure 1 is a hardware structure block diagram of a cloud server according to the service instance quantity adjustment method of the embodiment of the application;

[0014] Figure 2 is a flow chart of the service instance quantity adjustment method according to the embodiment of the application;

[0015] Figure 3 is a flow chart of the cloud platform interface service dynamic elastic scaling method according to the embodiment of the application;

[0016] Figure 4 is a structure block diagram of the service instance quantity adjustment device according to the embodiment of the application. DETAILED DESCRIPTION

[0017] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the application.

[0018] It should be noted that, in the description of the application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0019] In order for those skilled in the art to better understand the application, the application will be further described in detail below with reference to the drawings and specific embodiments.

[0020] In combination with the specific application environment architecture or specific hardware architecture on which the service instance quantity adjustment method is executed, the specific application environment architecture or specific hardware architecture is described here.

[0021] The method embodiments provided in the embodiments of the application can be executed in a cloud server or similar computing device. Taking the case of running on a cloud server, Figure 1 is a hardware structure block diagram of a cloud server according to the service instance quantity adjustment method of the embodiment of the application. As Figure 1 shown, the cloud server can include one or more Figure 1The cloud server shown in the figure can include, but is not limited to, a processor 102 (the processor 102 can include, but is not limited to, a processing device such as a microprocessor MPU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the cloud server can further include a transmission device 106 for communication function and an input and output device 108. Figure 1 The structure shown in the figure is only schematic and does not limit the structure of the cloud server. For example, the cloud server can further include more or fewer components than those shown in the figure, or have a different configuration from that shown in the figure. Figure 1 The cloud server shown in the figure can include, but is not limited to, a processor 102 (the processor 102 can include, but is not limited to, a processing device such as a microprocessor MPU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the cloud server can further include a transmission device 106 for communication function and an input and output device 108. Figure 1 The cloud server shown in the figure can include, but is not limited to, a processor 102 (the processor 102 can include, but is not limited to, a processing device such as a microprocessor MPU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the cloud server can further include a transmission device 106 for communication function and an input and output device 108.

[0022] The memory 104 can be used to store computer programs, such as software programs of application software and modules, such as computer programs corresponding to the booting method of the operating system in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above-mentioned method. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, and these remote memories can be connected to the cloud server through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0023] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include a wireless network provided by a communication service provider of the cloud server. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC for short), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, RF for short) module, which is used to communicate with the Internet in a wireless manner.

[0024] In the embodiments of the present application, a service instance quantity adjustment method is provided, which can be applied to a cloud server, but is not limited thereto. Figure 2 The flow chart of the service instance quantity adjustment method according to the embodiments of the present application is shown in FIG. 2, which includes the following steps S202-S206. Figure 2

[0025] In step S202, real-time performance indicators of a plurality of service instances of a current interface are acquired, wherein the real-time performance indicators at least include: a number of calling requests to the service instances, a CPU usage rate, and a memory usage rate.​

[0026] It should be noted that the real-time performance indicators can also include request delay, request error number and the like.

[0027] In step S204, it is determined whether the current interface meets the elastic scaling policy according to the real-time performance indicators, wherein the elastic scaling policy is used to adjust the number of service instances of the current interface.

[0028] In step S206, the number of service instances of the current interface is adjusted according to the elastic scaling policy when it is determined that the current interface meets the elastic scaling policy.

[0029] It should be noted that the determination of whether the current interface meets the elastic scaling policy can be understood as the determination of whether the state of the current interface meets the pre-judgment condition of the elastic scaling policy, that is, the judgment condition corresponding to the expansion threshold and the shrinkage threshold. The adjustment of the number of service instances according to the elastic scaling policy can be understood as the adjustment of the number of service instances according to the specific adjustment strategy of the elastic scaling policy.

[0030] Through the above steps, the real-time performance indicators of the plurality of service instances of the current interface are obtained in real time, wherein the real-time performance indicators at least include the number of service instance calls, CPU usage and memory usage. It is determined whether the current interface meets the elastic scaling policy according to these real-time performance indicators, wherein the elastic scaling policy is used to adjust the number of service instances of the current interface. When it is determined that the current interface meets the elastic scaling policy, the number of service instances of the current interface is dynamically adjusted according to the elastic scaling policy. By using the above scheme, the dynamic elastic scaling mechanism based on real-time performance data is introduced, and it is ensured that the service instance can be flexibly adjusted according to the load condition, so as to improve the resource utilization efficiency and the overall flexibility of the system while ensuring the quality of interface service. Further, the problem of resource waste caused by the fixed number of instances of the interface service module in the traditional K8S deployment in the related art is solved.

[0031] In one example embodiment, before determining whether the current interface meets the elastic scaling policy according to the real-time performance indicators, the method further comprises: obtaining historical request data of the plurality of service instances in a first time period, wherein the first time period is before the current time, and the historical request data is used to indicate the calling request data of the service instances; determining the relative calling frequencies of the plurality of service instances according to a plurality of the historical request data, and determining the calling proportions of the plurality of service instances according to a plurality of the relative calling frequencies; constructing a stress test set according to the calling proportions, wherein the stress test set comprises a plurality of test requests, the proportions of different categories of test requests in the plurality of test requests are the calling proportions, and the stress test set is used to simulate the actual workloads of the plurality of service instances; performing stress testing on the plurality of service instances through the stress test set to obtain test performance indicators of the plurality of service instances, wherein the test performance indicators are of the same categories as the performance indicators contained in the real-time performance indicators; determining the maximum QPS values of each service instance according to a plurality of the test performance indicators, wherein the maximum QPS values are used to indicate the maximum number of calling requests that can be processed simultaneously by the service instances under normal performance range; and determining the expansion threshold and the shrinkage threshold of the performance indicators of a plurality of categories according to the maximum QPS values, wherein the expansion threshold and the shrinkage threshold are used to determine the elastic scaling policy.

[0032] The embodiment proposes a fine elastic scaling policy making method, aiming to improve the response efficiency and resource utilization rate of the cloud platform interface service. Specifically, the following steps are included:

[0033] 1. Obtain service instance historical request data, which covers the actual calling situation in a period of time, and is used to reveal the calling frequency of each instance and further determine the calling proportion. The calling proportion reveals the calling frequency distribution of each service instance in actual work, and is the basis for constructing the stress test set.

[0034] 2. The stress test set is constructed according to the above-mentioned calling proportion, which contains a plurality of test requests similar to the actual business requests, and the proportion is consistent, so as to ensure that the stress test result can truly reflect the load situation of the service instance. The simulation of the stress test set is close to the actual work load, which improves the accuracy and practicability of the test.

[0035] 3. Perform stress testing on the service instances through the stress test set, collect test performance indicators, and cover multi-dimensional indicators of request processing capacity, which are matched with the real-time performance indicators. The test performance indicators include but are not limited to the number of requests, delay, error rate, etc., which are the key data for evaluating the carrying capacity of the service instance.

[0036] 4、Based on the stress test results, determine the maximum QPS value of each service instance, that is, the maximum number of requests that can be processed simultaneously within the normal performance range. The maximum QPS value is a hard indicator of service instance processing capacity, providing data support for elastic scaling strategies.

[0037] 5、Further, according to the maximum QPS value, set the expansion and contraction threshold of the performance index, and guide the elastic scaling operation. The threshold setting needs to be accurate, both to avoid resource waste and to ensure timely response and high availability of the service.

[0038] It should be noted that QPS (Queries Per Second) is the query rate per second, which is an important performance indicator for measuring the request processing capacity of a server or service instance.

[0039] It should be noted that historical request data is a record of invocation requests to the service instance within a specific time range before the current time, used to analyze the load pattern and invocation frequency of the service instance.

[0040] It should be noted that the stress test set is a set of test requests used to simulate actual workloads, with a composition ratio matching the invocation ratio in actual business scenarios, used to evaluate the performance limit and stability of the service instance.

[0041] It should be noted that the maximum QPS value refers to the maximum number of requests that can be processed under normal operation of the service instance, and is an important reference for formulating elastic scaling strategies.

[0042] This embodiment provides an accurate elastic scaling strategy formulation method by combining the use of historical request data and stress test sets. This not only ensures that the service remains stable and efficient under high load, but also dynamically adjusts resources according to the actual workload of the service instance, avoiding excessive or insufficient resource configuration, significantly improving resource utilization efficiency and cost-effectiveness. In addition, through the calculation of the maximum QPS value, the expansion and contraction thresholds can be set more accurately, enhancing the scientificity and practicality of the elastic scaling strategy. Overall, this embodiment aims to achieve dynamic, intelligent and efficient management of cloud platform interface services.

[0043] It should be noted that the performance index data (including the above real-time performance index and historical performance index) is collected by the cloud platform monitoring system, which is the data basis of this application. The data synchronization and calculation module obtains the basic performance data of all interface services from Prometheus and performs data calculation to obtain the final index number, as shown in Table 1:

[0044] Table 1

[0045]

[0046]

[0047] Optionally, determining whether the current interface complies with the elastic scaling policy based on the real-time performance indicators includes: obtaining historical performance indicators of the multiple service instances within a second time period, wherein the second time period is before the current moment and the second time period is adjacent to the current moment; determining whether the current interface complies with the elastic scaling policy within a third time period based on the real-time performance indicators and the historical performance indicators, wherein the current moment is the start time of the third time period.

[0048] This embodiment focuses on an elastic scaling decision-making mechanism that combines real-time and historical performance data. It emphasizes that when evaluating whether a service instance needs to be scaled, not only the current real-time performance indicators should be considered, but also the historical performance indicators from the immediate past should be considered. Real-time performance indicators reflect the instantaneous working status, while historical performance indicators provide insight into the recent behavioral trends of the service instance. Combining these two allows for a more comprehensive assessment of the health and potential demand of the service instance, ensuring the accuracy and foresight of elastic scaling operations.

[0049] Specifically, this embodiment collects historical performance metrics from multiple service instances over a short period of time immediately following the current moment. These metrics include, but are not limited to, the number of requests, latency, error rates, and resource usage. This historical data is then analyzed alongside real-time performance metrics to determine whether the service instances need to be scaled up or down to accommodate load changes in the next time period, known as the third time period.

[0050] This embodiment achieves a comprehensive evaluation of the service instance load by integrating real-time and historical performance indicators, ensuring the accuracy and timeliness of elastic scaling decisions. Such a mechanism can not only quickly respond to sudden high-load events, but also predict future load requirements based on historical data, so as to make adjustments in advance and avoid service delays and interruptions. At the same time, it can also optimize resource allocation according to the actual operation of the service instance, effectively avoiding excessive or insufficient resource reservation, thereby improving the overall operating efficiency and resource utilization of the system and reducing operation and maintenance costs. In this way, this embodiment significantly enhances the elasticity and reliability of the cloud platform interface service, ensuring service quality and user experience under various load conditions.

[0051] Optionally, determining whether the current interface meets the elastic scaling policy in the third time period according to the real-time performance indicators and the historical performance indicators comprises: performing linear regression on performance indicators of the plurality of service instances according to the real-time performance indicators and the historical performance indicators to obtain linear regression models of the plurality of service instances; predicting performance indicator change data of the plurality of service instances in the third time period according to the plurality of linear regression models; and determining whether the current interface meets the elastic scaling policy in the third time period according to the plurality of performance indicator change data.

[0052] The core feature of the embodiment is to use linear regression analysis to predict future performance and then determine whether to start elastic scaling. First, the collected real-time and historical performance indicator data of the service instances are combined, and linear regression modeling is used to build a performance prediction model for each service instance. Then, the possible trends of various performance indicators, including request quantity, response delay, error rate and other key indicators, in the upcoming third time period are predicted through these models. Finally, based on the predicted performance data, it is evaluated whether the preset elastic scaling conditions are met to determine whether the number of service instances needs to be increased or decreased, ensuring that the service can cope with future high loads and save resources when demand decreases.

[0053] This embodiment makes the elastic scaling decision more accurate through forward-looking performance prediction, avoiding resource waste and service interruption. It can predict performance bottlenecks in advance and timely expand resources, while releasing excess instances during the trough period, maintaining the high efficiency and low cost operation of the system. This data-driven dynamic adjustment strategy greatly improves the adaptive ability and user satisfaction of the cloud platform interface service.

[0054] Optionally, determining whether the current interface meets the elastic scaling policy in the third time period according to the plurality of performance indicator change data comprises: in a case where a first value in first performance indicator change data is greater than an expansion threshold in the elastic scaling policy, determining that a first service instance in the plurality of service instances meets an expansion strategy in the elastic scaling policy in the third time period, wherein the first service instance corresponds to the first performance indicator change data, and the plurality of performance indicator change data includes the first performance indicator change data; and in a case where second performance indicator change data are all less than or equal to a contraction threshold in the elastic scaling policy, determining that a second service instance in the plurality of service instances meets a contraction strategy in the elastic scaling policy in the third time period, wherein the second service instance corresponds to the second performance indicator change data, and the plurality of performance indicator change data includes the second performance indicator change data.

[0055] In this embodiment, the execution of the elastic scaling strategy relies on the analysis of the performance indicator change data of the service instances. If the first performance indicator prediction data of a service instance shows that the first value exceeds the preset expansion threshold, it indicates that the service instance (labeled as the first service instance) is likely to encounter an overload beyond the normal in the future third time period, and thus meets the criteria for executing the expansion strategy. The expansion threshold is a boundary. Once the prediction data reaches or exceeds the expansion threshold, the expansion mechanism is triggered to increase the number of instances to cope with the upcoming demand peak.

[0056] On the contrary, if the second performance indicator change data of all service instances shows that it is lower than or equal to the contraction threshold, it means that their loads tend to be stable or decrease. At this time, the second service instance (corresponding to the second performance indicator change data) will be considered to meet the condition for executing the contraction strategy. The contraction threshold is used to determine whether the service instance is in a low load state, allowing the number of instances to be reduced to optimize resource allocation and avoid unnecessary resource consumption.

[0057] It should be noted that:

[0058] Expansion threshold: a preset upper limit value of the performance indicator. When the predicted performance indicator exceeds this threshold, the system considers it necessary to increase the number of service instances.

[0059] Contraction threshold: a set lower limit value of the performance indicator, used to determine whether the number of service instances can be reduced to save resources.

[0060] It should be noted that the above expansion threshold and contraction threshold are respectively set for each type of performance indicator. In an optional embodiment, examples of the expansion threshold and the contraction threshold are shown in Table 2:

[0061] Table 2

[0062]

[0063] Based on Table 2, when any single indicator triggers the expansion threshold, the expansion operation (i.e., increasing the service instances) is performed. When all indicators trigger the contraction threshold, the contraction operation is generated.

[0064] Optionally, the following limiting conditions can be added:

[0065] When the CPU usage of the host in the cluster and the memory usage of the host exceed 80% (a preset threshold, not limited to 80%, but also 90%, 70%, etc.), the expansion is no longer performed. When the number of interface service instances is 1, the contraction is no longer performed.

[0066] The embodiment ingeniously uses the comparison between the performance index change data and the preset threshold value to realize intelligent expansion and contraction of the service instance. The prediction-based expansion and contraction strategy can not only respond to the changing load demand in a timely manner to ensure high availability and response speed of the service, but also save resources during the low load period to improve the resource use efficiency and economic benefits of the entire system. By dynamically adjusting the number of service instances, the system can meet the service quality and performance while achieving the optimal resource allocation state.

[0067] Optionally, in a case where it is determined that the current interface meets the elastic scaling policy, the number of service instances of the current interface is adjusted according to the elastic scaling policy, including: in a case where it is determined that the first service instance meets the expansion policy in the third time period, a first number is determined according to a second value in the first performance index change data, and the number of third service instances in the current interface is increased to the first number, wherein the second value is a maximum value in the first performance index change data, and the first service instance and the third service instance are of the same type; and in a case where it is determined that the second service instance meets the contraction policy in the third time period, a second number is determined according to a third value in the second performance index change data, and the number of fourth service instances in the current interface is reduced to the second number, wherein the third value is a maximum value in the second performance index change data, and the second service instance and the fourth service instance are of the same type.

[0068] The embodiment further discusses how to accurately adjust the number of instances based on the performance index change data when it is confirmed that the service instance of the current interface needs to follow the elastic scaling policy. When it is detected that the first service instance faces a demand surge and is expected to be expanded in the third time period, the system determines the specific scale of expansion, i.e., the first number, according to the highest demand value, i.e., the second value, reflected in the first performance index change data. Then, the number of third service instances of the same type as the first service instance is increased to the first number to ensure that it can cope with the upcoming high load.

[0069] Similarly, if the performance index change of the second service instance shows that it is in a low load state and is expected to perform the contraction policy in the third time period, the number of instances after contraction, i.e., the second number, is determined according to the maximum value, i.e., the third value, in the second performance index change data. Then, the system adjusts the number of fourth service instances of the same type as the second service instance to the second number to reasonably reduce the instances and avoid resource waste.

[0070] Through the embodiment, the system can accurately adjust the number of service instances of the same type when predicting load changes, ensure that the service always matches the demand, effectively cope with high load, reasonably reduce resources, provide sufficient service during peak periods, avoid resource idling during low peak periods, and achieve intelligent allocation of resources and efficient operation of cloud platform interface services.

[0071] Based on the above steps, after adjusting the number of service instances of the current interface according to the elastic scaling policy, the method further comprises: in the case where it is determined that the number of service instances of the current interface has been adjusted, prohibiting adjustment of the number of service instances of the current interface again within a fourth time period, wherein the start time of the fourth time period is the time when the adjustment of the number of service instances of the current interface is completed.

[0072] After implementing the elastic scaling policy and adjusting the number of service instances of the current interface, the embodiment introduces a prohibition period, i.e., a fourth time period, during which the system will prevent any further adjustment of the number of instances of the interface. The setting of the prohibition period starts from the time when the adjustment of the number of service instances is completed, ensuring that the system state is stable for a period of time after adjustment, avoiding resource fluctuations and performance instability caused by frequent scaling operations.

[0073] Once the number of service instances is adjusted according to the elastic scaling policy, whether it is expanded or contracted, the system immediately enters the fourth time period, during which the number of service instances is locked from changing. The purpose of this design is to give the adjusted system state enough time to stabilize and prevent unnecessary resource reallocation and performance disturbance caused by continuous scaling operations.

[0074] By setting the prohibition mechanism of the fourth time period, the embodiment effectively balances the relationship between real-time response and system stability, avoids excessive frequent scaling actions, ensures that each adjustment is made after sufficient evaluation, helps to maintain the continuity and high performance of the service, while reducing the complexity of operation and maintenance, and optimizing the overall resource planning and utilization efficiency.

[0075] Obviously, the above-described embodiments are only part of the embodiments of the present application, not all. In order to better understand the above method, the above process is described in combination with the embodiments as follows, but not used to limit the technical solutions of the embodiments of the present application, specifically:

[0076] In an optional embodiment, the present application provides a method for dynamic elastic scaling of cloud platform interface services, which realizes that the cloud platform interface services formulate an elastic scaling policy according to real-time performance indicators and historical performance indicators, dynamically adjust the number of interface service instances, and realize the dynamic, real-time and forward-looking of elastic scaling services. As shown in Figure 3 the specific implementation process is as follows:

[0077] 1. Multi-dimensional performance indicators and calculation:

[0078] 1) Upper limit of preset request number: the upper limit of QPS of each interface service instance is calculated based on the stress test, wherein the proportion of each interface is determined according to the specific function of the interface in the stress test.

[0079] 2) Performance data: the performance indicator data is collected by the cloud platform monitoring system, which is the data basis of the present application, and the specific data is shown in Table 1.

[0080] 2. Scaling strategy (i.e. the above-mentioned elastic scaling strategy): the basis for determining the strategy is shown in Table 2.

[0081] The specific determination rules include:

[0082] 1) Generate a capacity expansion operation when any single indicator triggers the capacity expansion threshold;

[0083] 2) Generate a capacity reduction operation when all indicators trigger the capacity reduction threshold;

[0084] 3) Do not expand when the CPU usage rate of the host and the memory usage rate of the host exceed 80% in the cluster;

[0085] 4) Do not reduce capacity when the number of interface service instances is 1;

[0086] 5) Do not perform the same operation within ten minutes (i.e. the fourth time period mentioned above) after performing the capacity expansion and capacity reduction operations.

[0087] 3. Process scheduling:

[0088] 1) The task scheduling engine schedules the entire elastic scaling process with a period of 1 minute;

[0089] 2) Access Prometheus to obtain performance data;

[0090] 3) Calculate the average value and the K value of the one-dimensional linear regression equation of each service type and the dimensions of the host;

[0091] 4) Compare the performance data of each service type with the strategy threshold to obtain the elastic scaling result;

[0092] 5) For the interface service that needs to be scaled, call the K8S interface to send a scheduling request.

[0093] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment.

[0094] The embodiment of the application further provides a service instance quantity adjusting device, Figure 4 The embodiment of the application further provides a service instance quantity adjusting device, Figure 4 The device comprises:

[0095] The acquisition module 42 is configured to acquire real-time performance indexes of a plurality of service instances of a current interface, wherein the real-time performance indexes at least comprise a number of calling requests for the service instances, a CPU usage rate and a memory usage rate.

[0096] The determination module 44 is configured to determine whether the current interface meets an elastic scaling policy according to the real-time performance indexes, wherein the elastic scaling policy is used to adjust the quantity of service instances of the current interface.

[0097] The adjustment module 46 is configured to adjust the quantity of service instances of the current interface according to the elastic scaling policy when it is determined that the current interface meets the elastic scaling policy.

[0098] According to the above device, the real-time performance indexes of a plurality of service instances of a current interface are acquired in real time, wherein the real-time performance indexes at least comprise a number of calling requests for the service instances, a CPU usage rate and a memory usage rate; whether the current interface meets an elastic scaling policy is determined according to the real-time performance indexes, wherein the elastic scaling policy is used to adjust the quantity of service instances of the current interface; and the quantity of service instances of the current interface is dynamically adjusted according to the elastic scaling policy when it is determined that the current interface meets the elastic scaling policy. According to the above scheme, a dynamic elastic scaling mechanism based on real-time performance data is introduced, so that the service instances can be flexibly adjusted according to the load, thereby ensuring the quality of interface services, improving the resource utilization efficiency and the overall flexibility of the system, and further solving the problem of resource waste caused by the fixed quantity of service instances in the interface service module in the traditional K8S deployment in the related art.

[0099] Optionally, the determination module 44 is further used to obtain historical request data of the multiple service instances within a first time period, wherein the first time period is before the current moment, and the historical request data is used to indicate call request data for the service instance; determine the relative call frequency of the multiple service instances based on the multiple historical request data, and determine the call ratio of the multiple service instances based on the multiple relative call frequencies; construct a stress test set based on the call ratio, wherein the stress test set includes multiple test requests, and the ratio of test requests of different categories in the multiple test requests is the call ratio, and the stress test set is used to simulate the multiple service instances actual workload; stress testing the multiple service instances through the stress testing set to obtain test performance indicators of the multiple service instances, wherein the test performance indicators are of the same category as the performance indicators included in the real-time performance indicators; determining the maximum QPS value of each service instance according to the multiple test performance indicators, wherein the maximum QPS value is used to indicate the maximum number of call requests that the service instance can process simultaneously under a normal performance range; determining the expansion threshold and the contraction threshold of multiple categories of performance indicators according to the maximum QPS value, wherein the expansion threshold and the contraction threshold are used to determine the elastic scaling strategy.

[0100] Optionally, the above-mentioned determination module 44 is also used to obtain historical performance indicators of the multiple service instances in a second time period, wherein the second time period is before the current moment and the second time period is adjacent to the current moment; and determine whether the current interface complies with the elastic scaling policy in a third time period based on the real-time performance indicators and the historical performance indicators, wherein the current moment is the start moment of the third time period.

[0101] Optionally, the above-mentioned determination module 44 is used to perform linear regression on the performance indicators of the multiple service instances based on the real-time performance indicators and the historical performance indicators to obtain linear regression models of the multiple service instances; predict the performance indicator change data of the multiple service instances within the third time period based on the multiple linear regression models; and determine whether the current interface complies with the elastic scaling strategy within the third time period based on the multiple performance indicator change data.

[0102] Optionally, the determining module 44 is further configured to determine that a first service instance in the plurality of service instances conforms to a scaling-out policy in the elastic scaling policy in the third time period, in a case where the first value in the first performance index change data is greater than a scaling-out threshold in the elastic scaling policy, the first service instance corresponding to the first performance index change data, the plurality of performance index change data including the first performance index change data; and determine that a second service instance in the plurality of service instances conforms to a scaling-in policy in the elastic scaling policy in the third time period, in a case where the second value in the second performance index change data is less than or equal to a scaling-in threshold in the elastic scaling policy, the second service instance corresponding to the second performance index change data, the plurality of performance index change data including the second performance index change data.

[0103] Optionally, the adjusting module 46 is further configured to determine a first quantity according to a second value in the first performance index change data and increase a quantity of a third service instance in the current interface to the first quantity, in a case where it is determined that the first service instance conforms to the scaling-out policy in the third time period, the second value being a maximum value in the first performance index change data, the first service instance and the third service instance being of the same type; and determine a second quantity according to a third value in the second performance index change data and reduce a quantity of a fourth service instance in the current interface to the second quantity, in a case where it is determined that the second service instance conforms to the scaling-in policy in the third time period, the third value being a maximum value in the second performance index change data, the second service instance and the fourth service instance being of the same type.

[0104] Optionally, the adjusting module 46 is further configured to prohibit adjustment of the quantity of the service instances in the current interface again in a fourth time period, in a case where it is determined that the quantity of the service instances in the current interface has been adjusted, a start time of the fourth time period being a time when the adjustment of the quantity of the service instances in the current interface is completed.

[0105] The features of the embodiments of the service instance quantity adjustment apparatus can be understood with reference to the related descriptions of the embodiments of the service instance quantity adjustment method, which will not be repeated here.

[0106] Embodiments of the present application further provide an electronic device including a memory and a processor, the memory storing a computer program, and the processor being configured to run the computer program to perform the steps in any of the embodiments of the service instance quantity adjustment method.

[0107] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is configured to execute the steps in the service instance quantity adjustment method embodiments.

[0108] In an example embodiment, the computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0109] The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in the service instance quantity adjustment method embodiments.

[0110] The embodiment of the present application further provides another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in the service instance quantity adjustment method embodiments.

[0111] The skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0112] The above provides a detailed description of the service instance quantity adjustment method and device, electronic equipment and computer program product. The principle and implementation of the present application are described by applying specific examples. The above description of the examples is only used to help understand the method and its core idea. It should be pointed out that for ordinary skilled in the art, without departing from the principle of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for adjusting the number of service instances, characterized in that: include: Obtaining real-time performance indicators of multiple service instances of the current interface, wherein the real-time performance indicators include at least: the number of call requests to the service instance, CPU usage, and memory usage; Determining whether the current interface complies with an elastic scaling policy based on the real-time performance indicator, wherein the elastic scaling policy is used to adjust the number of service instances of the current interface; When it is determined that the current interface complies with the elastic scaling policy, the number of service instances of the current interface is adjusted according to the elastic scaling policy.

2. The method for adjusting the number of service instances according to claim 1, wherein: Before determining whether the current interface complies with the elastic scaling policy according to the real-time performance indicator, the method further includes: Obtaining historical request data of the multiple service instances within a first time period, wherein the first time period is before a current moment, and the historical request data is used to indicate call request data for the service instances; Determining relative call frequencies of the plurality of service instances according to the plurality of historical request data, and determining call ratios of the plurality of service instances according to the plurality of relative call frequencies; Constructing a stress testing set according to the call ratio, wherein the stress testing set includes multiple test requests, the ratio of different types of test requests in the multiple test requests is the call ratio, and the stress testing set is used to simulate the actual workload of the multiple service instances; Performing stress testing on the multiple service instances using the stress testing set to obtain test performance indicators of the multiple service instances, wherein the test performance indicators are of the same category as the performance indicators included in the real-time performance indicators; Determine a maximum QPS value for each service instance based on the multiple test performance indicators, wherein the maximum QPS value is used to indicate the maximum number of call requests that the service instance can process simultaneously under a normal performance range; The expansion thresholds and the reduction thresholds of the performance indicators of multiple categories are determined according to the maximum QPS value, wherein the expansion thresholds and the reduction thresholds are used to determine the elastic scaling policy.

3. The method for adjusting the number of service instances according to claim 1, wherein: Determining whether the current interface complies with the elastic scaling policy according to the real-time performance indicator includes: Obtain historical performance indicators of the multiple service instances in a second time period, wherein the second time period is before the current moment and the second time period is adjacent to the current moment; Determine whether the current interface complies with the elastic scaling policy within a third time period according to the real-time performance indicator and the historical performance indicator, wherein the current time is a start time of the third time period.

4. The method for adjusting the number of service instances according to claim 3, wherein: Determining whether the current interface complies with the elastic scaling policy within a third time period according to the real-time performance indicator and the historical performance indicator includes: Performing linear regression on the performance indicators of the multiple service instances according to the real-time performance indicators and the historical performance indicators to obtain linear regression models of the multiple service instances; Predicting the performance indicator change data of the multiple service instances in the third time period respectively according to the multiple linear regression models; Determine whether the current interface complies with the elastic scaling policy within the third time period according to the plurality of performance indicator change data.

5. The method for adjusting the number of service instances according to claim 4, characterized in that: Determining whether the current interface complies with the elastic scaling policy within the third time period according to the plurality of performance indicator change data includes: When a first value in the first performance indicator change data is greater than a capacity expansion threshold in the elastic scaling policy, determining that a first service instance among the multiple service instances complies with the capacity expansion policy in the elastic scaling policy within the third time period, wherein the first service instance corresponds to the first performance indicator change data, and the multiple performance indicator change data include the first performance indicator change data; When the second performance indicator change data are all less than or equal to the shrinking threshold in the elastic scaling policy, determine that the second service instance among the multiple service instances complies with the shrinking policy in the elastic scaling policy within the third time period, wherein the second service instance corresponds to the second performance indicator change data, and the multiple performance indicator change data include the second performance indicator change data.

6. The method for adjusting the number of service instances according to claim 5, characterized in that: When it is determined that the current interface complies with the elastic scaling policy, adjusting the number of service instances of the current interface according to the elastic scaling policy includes: If it is determined that the first service instance complies with the capacity expansion policy within the third time period, determine a first quantity according to a second value in the first performance indicator change data, and increase the number of third service instances in the current interface to the first quantity, wherein the second value is a maximum value in the first performance indicator change data, and the first service instance and the third service instance are of the same category; When it is determined that the second service instance complies with the scaling-down strategy within the third time period, a second quantity is determined based on the third value in the second performance indicator change data, and the number of fourth service instances in the current interface is reduced to the second quantity, wherein the third value is the maximum value in the second performance indicator change data, and the second service instance is of the same category as the fourth service instance.

7. The method for adjusting the number of service instances according to any one of claims 1 to 6, characterized in that: After adjusting the number of service instances of the current interface according to the elastic scaling policy, the method further includes: When it is determined that the number of service instances of the current interface has been adjusted, it is prohibited to adjust the number of service instances of the current interface again within a fourth time period, wherein the start time of the fourth time period is the time when the adjustment of the number of service instances of the current interface is completed.

8. A device for adjusting the number of service instances, characterized in that: include: An acquisition module is used to obtain real-time performance indicators of multiple service instances of the current interface, wherein the real-time performance indicators include at least: the number of call requests to the service instance, CPU usage, and memory usage; a determination module, configured to determine, based on the real-time performance indicator, whether the current interface complies with an elastic scaling policy, wherein the elastic scaling policy is configured to adjust the number of service instances of the current interface; The adjustment module is configured to adjust the number of service instances of the current interface according to the elastic scaling policy when it is determined that the current interface complies with the elastic scaling policy.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method for adjusting the number of service instances as described in any one of claims 1 to 7 when executing the computer program.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for adjusting the number of service instances according to any one of claims 1 to 7 are implemented.