Scaling Control Method, Device, Electronic Device and Program Product
By obtaining and analyzing the index thresholds and dynamic call requirements of each node of the service call link, and automatically adjusting the number of services, the problem of low manual expansion and capacity efficiency in the existing technology is solved, and efficient automatic expansion and capacity control is achieved.
Patent Information
- Application Number
- CN202411147988.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-08-20
AI Technical Summary
When the traffic of existing information systems suddenly increases or decreases, they need to manually submit expansion and reduction work orders, and it is difficult to ensure that the underlying services support the call of upstream services in a single service.
By obtaining the index thresholds and dynamic call requirements of each node in the call link, adjust the number of services on each node to achieve scaling control. The specific steps include obtaining available resources and customer call requirements information, traversing the service call link, determining the number of service updates based on node indicators and call requirements, and updating available resource information.
It realizes automated expansion and capacity control to ensure that the service call link meets customer call needs, improves operation and maintenance efficiency, and reduces operation and maintenance difficulties.
Smart Images

Figure CN118869707B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of computer networks, and specifically to a scaling control method. In addition, the present disclosure also relates to related devices, electronic devices, and program products. Background Art
[0002] When existing information systems face sudden increases or decreases in traffic, developers often need to manually submit scaling-out and scaling-in work orders, and then operation and maintenance personnel manually control the scaling-out and scaling-in of services. In some scenarios, metrics (CPU usage, memory occupancy, network traffic, etc.) can be monitored, and individual services can be scaled out or in when needed. However, few existing systems provide services independently. Generally, through long service call chains, scaling out a single service is very likely to fail to ensure that the underlying services support the calls of upstream services.
[0003] The content described in this background art is only for facilitating the understanding of related technologies in this field and is not regarded as an admission of existing technologies.
[0004] Application Content
[0005] Therefore, embodiments of the present disclosure provide a scaling control method, which is intended to adjust the number of services on each node according to the metric thresholds of each node on the call chain and dynamic call requirements, so as to achieve scaling control. Specifically, embodiments of the present disclosure provide a scaling control method, including the following steps:
[0006] Obtain available resource information and customer call requirement information of the target service system, where the target service system includes a service call chain, services are deployed on each call node on the service call chain, each service is configured with a resource occupancy amount and an upstream call metric threshold, each call node is configured with a node metric threshold, and call nodes on the service call chain except the most downstream call node are also configured with downstream call requirements;
[0007] Traverse each call node on the service call chain from the most upstream call node to the most downstream call node of the service call chain, and determine the updated number of services on each call node and update the available resource information according to the node metric threshold, call requirement information, number of services, resource occupancy amount of a single service, and upstream call metric threshold of each call node;
[0008] Traverse each call node on the service call chain from the most upstream call node to the most downstream call node of the service call chain, and determine the updated number of services on each call node and update the available resource information according to the node metric threshold, call requirement information, number of services, resource occupancy amount of a single service, and upstream call metric threshold of each call node;
[0009] Scale the calling nodes in the service call link up or down according to the update quantity of the services in each calling node on the service call link.
[0010] In some embodiments of the present disclosure, determining the update quantity of the services of each calling node and updating the available resource information includes:
[0011] Traverse each calling node on the service call link from the most upstream calling node to the most downstream calling node of the service call link, and for each calling node, perform the following steps:
[0012] Determine the predicted quantity of the services of the current calling node according to the node metric threshold, call demand information, quantity of services, resource occupancy of a single service, and the upstream call metric threshold of the current calling node;
[0013] In response to the current available resource information meeting the requirements of the predicted quantity and resource occupancy of the services of the current calling node, determine the update quantity of the services of the current calling node and update the available resource information.
[0014] In some embodiments of the present disclosure, determining the call demand information of the current calling node includes:
[0015] In response to the current calling node being the most upstream calling node on the service call link, use the customer call demand information as the call demand information of the current calling node;
[0016] In response to the current calling node being a calling node after the most upstream calling node on the service call link, use the product of the downstream call demand of the services of the upstream calling node of the current calling node and the update quantity of the services of the upstream calling node of the current calling node as the call demand information of the current calling node.
[0017] In some embodiments of the present disclosure, the upstream call metric threshold includes an expansion metric threshold, and determining the predicted quantity of the services of the current calling node according to the node metric threshold, call demand information, quantity of services, resource occupancy of a single service, and the upstream call metric threshold includes:
[0018] When the call demand information of the current calling node meets the node metric threshold requirement and the call demand information is greater than the service expansion metric of the current calling node, increase the quantity of the services of the current calling node so that the service expansion metric of the current calling node is greater than the call demand information, where the service expansion metric is obtained according to the quantity of the services and the expansion metric threshold of a single service;
[0019] When the call demand information of the current calling node does not meet the node metric threshold requirement, generate a warning message.
[0020] In some embodiments of the present disclosure, the upstream - called metric threshold includes a scaling - down metric threshold. Determining the expected number of services of the current call node according to the node metric threshold of the current call node, the call demand information, the number of services, the resource occupancy of a single service, and the upstream - called metric threshold includes:
[0021] When the call demand information of the current call node meets the node metric threshold requirement and the call demand information is less than the service scaling - down metric of the current call node, reduce the number of services of the current call node so that the service scaling - down metric is less than the call demand information, where the service scaling - down metric is obtained according to the number of services and the scaling - down metric threshold of a single service;
[0022] When the call demand information of the current call node does not meet the node metric threshold requirement, generate a warning message.
[0023] In some embodiments of the present disclosure, determining the updated number of services of the current call node and updating the available resource information in response to the current available resource information meeting the requirements of the expected number and resource occupancy of the services of the current call node includes:
[0024] When the current available resource information is greater than the expected resource occupancy of the services of the current call node, use the expected number of services of the current call node as the updated number of services of the current call node, and update the available resource information with the current available resource information and the expected resource occupancy of the services of the current call node, where the expected resource occupancy of the services is obtained according to the expected number of services and the resource occupancy.
[0025] In some embodiments of the present disclosure, determining the updated number of services of each call node and updating the available resource information further includes:
[0026] When the current available resource information does not meet the requirements of the expected number and resource occupancy of the services of the current call node, generate a warning message.
[0027] In some embodiments of the present disclosure, scaling the call nodes in the service call link up or down according to the updated number of services in each call node in the service call link includes:
[0028] For each call node in the service call link, perform the following steps:
[0029] When the updated number of services in the call node is less than the number before the update, start timing for delayed execution according to the scaling - down configuration of the call node;
[0030] When the delay time reaches the setting of the scaling - down configuration delay, close the services in the call node according to the updated number and recycle the resources corresponding to the closed services.
[0031] In some embodiments of the present disclosure, scaling the calling nodes in the service call link according to the update quantity of the services in each calling node in the service call link includes:
[0032] For each calling node in the service call link, the following steps are performed:
[0033] When the update quantity of the service in the calling node is greater than the quantity before the update, the resources corresponding to the available resource information are called, and the service in the calling node is deployed incrementally according to the update quantity of the service in the calling node.
[0034] In some embodiments of the present disclosure, it further includes:
[0035] Count the number of scaling operations and the quantity of services of each calling node within a statistical period, and update the node metric thresholds, the preset quantity of services in each calling node, and the upstream call metric thresholds of the services in each calling node according to the number of scaling operations and the quantity of services of each calling node.
[0036] On the other hand, an embodiment of the present disclosure further provides a scaling control device, including an acquisition module, a scaling plan formation module, and an execution module, where
[0037] The acquisition module is configured to acquire available resource information and customer call demand information of a target service system, where the target service system includes a service call link, services are deployed on each calling node of the service call link, each service is configured with a resource occupancy quantity and an upstream call metric threshold, each calling node is configured with a node metric threshold, and the calling nodes except the most downstream calling node on the service call link are further configured with downstream call demands;
[0038] The scaling plan formation module is configured to traverse each calling node on the service call link from the most upstream calling node to the most downstream calling node of the service call link, and determine the update quantity of the services of each calling node and update the available resource information according to the node metric thresholds, call demand information, quantity of services, resource occupancy quantity of a single service, and upstream call metric thresholds of each calling node;
[0039] The execution module is configured to scale the calling nodes in the service call link according to the update quantity of the services in each calling node in the service call link.
[0040] In an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, where the program, when executed by a processor, implements the scaling control method of any embodiment of the present disclosure.
[0041] In an embodiment of the present disclosure, an electronic device is provided, including: a processor and a memory storing a computer program, and the processor is configured to execute the scaling control method of any embodiment of the present disclosure when running the computer program.
[0042] The embodiment of the present disclosure provides a scaling control method and apparatus, which traverse each call node on the service call link from the upstream to the downstream, obtain the call requirements of each call node, the existing number of services in the node, the call metric threshold of the services in the node, and the metric threshold of the node itself, and gradually update the number of services in the call node from the upstream to the downstream. When the number of services is updated upstream, the downstream can also adjust the number of services in the call node according to the dynamic number of services, so as to ensure that the entire service call link meets the customer call requirements. In the embodiment of the present disclosure, the metrics of each node include the node metric threshold related to the call node and the upstream call metric threshold related to the service, which can fully cover the evaluation conditions for scaling, thereby ensuring the reliability of automatic scaling.
[0043] Some optional features and technical effects of the embodiment of the present disclosure are described below, and some can be understood by reading this article. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Hereinafter, the embodiments of the present disclosure will be described in detail with reference to the drawings. The elements shown are not limited by the scale shown in the drawings, and the same or similar reference numerals in the drawings represent the same or similar elements, where:
[0045] Figure 1 FIG. shows an interaction schematic diagram between the scaling control system and the target service system according to an embodiment of the present disclosure;
[0046] Figure 2 FIG. shows an exemplary flowchart of the scaling control method according to an embodiment of the present disclosure;
[0047] Figure 3 FIG. shows an architecture schematic diagram of a service call link in an embodiment of the present disclosure;
[0048] Figure 4 FIG. shows another architecture schematic diagram of a service call link in an embodiment of the present disclosure;
[0049] Figure 5 FIG. shows an exemplary flowchart of determining the updated number of services and updated available resource information in the scaling control method according to an embodiment of the present disclosure;
[0050] Figure 6a FIG. shows an exemplary flowchart of the judgment process for increasing the number of services in the scaling control method according to an embodiment of the present disclosure;
[0051] Figure 6bShows an exemplary flowchart of the judgment process for reducing the number of services in the scaling control method according to an embodiment of the present disclosure;
[0052] Figure 7 Shows a schematic architecture diagram of a target service system according to an embodiment of the present disclosure;
[0053] Figure 8 Shows an exemplary structural diagram of a scaling control device according to an embodiment of the present disclosure; and
[0054] Figure 9 Shows an exemplary structural diagram of an electronic device capable of implementing the method according to an embodiment of the present disclosure. Detailed implementation manners
[0055] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to the detailed implementation manners and the accompanying drawings. Here, the exemplary implementation manners of the present disclosure and their descriptions are used to explain the present disclosure, but do not limit the present disclosure.
[0056] In the embodiments of the present disclosure, "Kubernetes (k8s)" refers to a container-based distributed architecture technology; "Message Queue (MQ)" refers to a message queue.
[0057] As Figure 1 shown, some embodiments of the present disclosure provide a scaling control system 100. The scaling control system 100 can collect the metric thresholds of each call node 211 in the service call link 210 of the target service system 200, perform scaling calculation of the number of services according to the metric thresholds, and can perform resource scheduling according to the scaling plan. The scaling control system 100 in the embodiments of the present disclosure runs a scaling control method. As Figure 2 shown, the specific steps are as follows:
[0058] S310: Obtain the available resource information and the customer call demand information of the target service system 200.
[0059] Among them, the target service system 200 includes a service call link 210. Services are deployed on each call node 211 on the service call link 210. Each service is configured with a resource occupancy amount and an upstream call metric threshold, and each call node is configured with a node metric threshold. Call nodes on the service call link except the most downstream call node are also configured with downstream call demands. In the embodiments of the present disclosure, as Figure 3 shown, the uppermost call node of the service call link 210 can be a server, the lowermost call node can be a bottom middleware, and the downstream call node of the server can be a sub-service. In Figure 3In the illustrated example, the server and the sub-service can also be configured with downstream call requirements.
[0060] In the embodiments of the present disclosure, each call node is configured with a node metric threshold. The node metric threshold is associated with the call node and does not affect the scaling calculation process of the service call link as the number of services configured in the call node changes. Even if the number of services in the call node changes, it will not affect the node metric threshold associated with the call node. As the number of services changes, the service scaling metric of the call node also changes accordingly. For example, as the number of services in the call node changes, the total resource occupancy of the services included in the call node and the upstream call metric threshold will change, which can affect the scaling calculation process of the service call link.
[0061] In the embodiments of the present disclosure, a service can be an application, a thread, or a process associated with specific resource occupancy (such as the number of hardware servers or virtual machine servers, the amount of memory, the number of CPU cores, etc.). The service can be deployed on one server or multiple servers. A call node includes one or more services. The available resource information can be available physical resources, including the number of CPU cores, the amount of memory, etc. The customer call requirement information can be the call requirement for the entire system. For example, it can be the customer call concurrency.
[0062] In some embodiments of the present disclosure, the node metric threshold can be any one or a combination of the following: a piece of logic, a rule, used to describe or limit service parameters and / or node parameters, such as the limitation of certain value ranges used by the same service deployed on different machines, a constant that cannot be scaled due to certain reasons, the delay time for triggering scaling, etc. For example: Since the downstream service or the underlying middleware needs to provide services to different systems, it will limit the number of connections to the upstream service call to ensure that services can be provided to each system. The node metric threshold of the downstream service or the underlying middleware is the limitation of the number of connections to the upstream service call. Based on this limitation, the total number of calls to the underlying service or middleware after the upstream service is scaled out does not exceed the agreed number of connections (i.e., the limited number of connections).
[0063] In some embodiments of the present disclosure, the upstream - called metric threshold is a range of upstream - called metrics. When the node metric of the calling node does not reach the node metric threshold, it is possible to determine whether to scale in or out according to the upstream - called metric value and the upstream - called metric threshold: when the upstream - called metric value is less than the minimum value of the upstream - called metric range, scaling in may be required; when the upstream - called metric value is greater than the maximum value of the upstream - called metric range, scaling out may be required; when the upstream - called metric value is within the upstream - called metric range, it can remain unchanged. In some examples, the node metric threshold is the maximum number of connections, and the upstream - called metric threshold can be the threshold of the available concurrency. The concurrency threshold can be a range of concurrency. When the called concurrency is less than the minimum value of the concurrency range, scaling in may be required; when the called concurrency is greater than the maximum value of the concurrency range, scaling out may be required; when the called concurrency is within the concurrency range, it can remain unchanged. For example, the maximum number of connections of the calling node is 20, and the concurrency threshold is (5, 10). When the number of upstream connections of the calling node does not exceed the maximum number of connections, when the called concurrency exceeds 10, scaling out is required, and when the called concurrency is less than 5, scaling in is required.
[0064] In some embodiments of the present disclosure, the node metric threshold is the maximum concurrency, and the upstream - called metric threshold can also be set correspondingly according to the calling node. For example, when the calling node is the message middleware MQ, the node metric threshold of the calling node is the maximum concurrency of 20, and the upstream - called metric threshold can be the threshold of the message backlog value. The message backlog value threshold is a range of backlog values (1000, 10000). When the upstream concurrency of the calling node does not exceed the maximum concurrency, when the message backlog value is greater than 10,000, scaling out is required, and when the message backlog value is less than 1000, scaling in is required.
[0065] In the embodiments of the present disclosure, the node metric threshold can be set as needed, and the above description is not regarded as a limitation on the node metric threshold.
[0066] In the embodiments of the present disclosure, each service on the service call chain is configured with the amount of occupied resources. For example, Figure 4As shown in the figure, the resource occupancy of Service 1 called by Invocation Node 1 is 1G of memory and 10 millicores of CPU. The upstream call metric thresholds of Service 1 include: scale out when the concurrency is greater than 10, and scale in when the concurrency is less than 5. The downstream call requirement of Service 1 is a concurrency of 5. The number of Service 1 in Invocation Node 1 is 1. The node metric threshold configured for Invocation Node 1 is a maximum connection number of 50. The downstream invocation node of Invocation Node 1 is Invocation Node 2. Invocation Node 2 is configured with 2 Service 2s. The resource occupancy of each Service 2 is 1G of memory and 5 millicores of CPU. The upstream call metric thresholds of Service 2 include: scale out when the concurrency is greater than 10, and scale in when the concurrency is less than 5. The downstream call requirement of Service 2 is a concurrency of 5. The node metric threshold configured for Invocation Node 2 is a maximum connection number of 200. The downstream invocation node of Invocation Node 2 is Invocation Node 3. Invocation Node 3 is configured with 3 Service 3s. The resource occupancy of each Service 3 is 1G of memory and 5 millicores of CPU. The upstream call metric thresholds of Service 3 include: scale out when the concurrency is greater than 10, and scale in when the concurrency is less than 5. The node metric threshold configured for Invocation Node 3 is a maximum connection number of 1000. Service 3 is a service in the underlying invocation node and has no downstream call requirement.
[0067] In some embodiments of the present disclosure, the initial available resources are set to 16G of memory and 800 millicores of CPU. After allocation according to the above configuration, the available resources are: 16 - 1 - 1*2 - 1*3 = 10G of memory, 800 - 10 - 5*2 - 5*3 = 765 millicores of CPU. At the initial runtime, the customer call demand information is a concurrency of 8 and a maximum connection number of 20, and the scale in / out conditions of Invocation Node 1 are not triggered. After running for a period of time, when the concurrency in the customer call demand information increases to 12 and the maximum connection number is 40, Service 1 in Invocation Node 1 reaches the scale out condition and scale out calculation is performed. The specific process is described in the following steps.
[0068] In the embodiments of the present disclosure, the downstream call requirement of a service can be configured with a specific value or can be configured to be related to the call requirement of its upstream invocation node.
[0069] S320: Traverse each invocation node on the service call chain from the most upstream invocation node to the most downstream invocation node of the service call chain. Determine the updated number of services of each invocation node and update the available resource information according to the node metric threshold, call demand information, number of services, resource occupancy of a single service, and upstream call metric threshold of each invocation node. By traversing from upstream to downstream, the number of services in the invocation node is gradually updated.
[0070] In some embodiments of the present disclosure, it is determined whether the updated number can be executed based on the available resource information, such as Figure 5As shown, determine the update quantity of the services of each call node and update the available resource information, including:
[0071] S321: Traverse each call node on the service call chain from the most upstream call node to the most downstream call node of the service call link. For each call node, perform the following steps:
[0072] S322: Determine the expected quantity of the services of the current call node according to the node metric threshold, call demand information, quantity of the services, resource occupancy of a single service, and upstream call metric threshold of the current call node.
[0073] S323: In response to the current available resource information meeting the requirements of the expected quantity and resource occupancy of the services of the current call node, determine the update quantity of the services of the current call node and update the available resource information.
[0074] In some embodiments of the present disclosure, determining the call demand information of the current call node includes:
[0075] In response to the current call node being the most upstream call node on the service call link, use the customer call demand information as the call demand information of the current call node;
[0076] In response to the current call node being a call node after the most upstream call node on the service call link, use the product of the downstream call demand of the services of the upstream call node of the current call node and the update quantity of the services of the upstream call node of the current call node as the call demand information of the current call node.
[0077] For example, as Figure 4 shown in the call link, the call demand information of call node 1 is the customer call demand information, the call demand information of call node 2 is the product of the quantity of the services of call node 1 and the downstream call demand of the services of call node 1, and the call demand information of call node 3 is the product of the quantity of the services of call node 2 and the downstream call demand of the services of call node 2.
[0078] In the embodiments of the present disclosure, the service call link is traversed from the most upstream call node to the most downstream call node to evaluate whether each call node triggers the capacity expansion condition. If the scaling condition is met, pre-scaling is performed. In one example, it is set that the capacity expansion condition for each call node is to expand the capacity when the concurrency exceeds 20. Traverse from the most upstream call node to the downstream call nodes, and pre-update the service quantity of the call node according to the call demand information, service quantity, and the threshold of the upstream call index of the call node. After pre-updating the service quantity, determine whether the available resources can support it according to the pre-updated service quantity and the resource occupancy of a single service. If it supports, determine the service quantity in the call node according to the pre-updated service quantity, and then gradually evaluate whether the downstream call node triggers the capacity expansion condition according to the service quantity determined by the upstream call node and the downstream call demand until the most downstream call node is determined. During the above traversal process, if it is determined that the available resources cannot support the total resource demand of the pre-updated service quantity according to the pre-updated service quantity and the resource occupancy of a single service, a warning message is generated, and manual intervention is performed to evaluate and allocate resources from the new resource environment to increase the available resource quantity.
[0079] In some embodiments of the present disclosure, the upstream call index threshold includes a capacity expansion index threshold, such as Figure 6a shown, determining the expected quantity of the service of the current call node according to the node index threshold, call demand information, service quantity, resource occupancy of a single service, and upstream call index threshold of the current call node includes:
[0080] S3221a: Obtain the call demand information of the current call node.
[0081] S3222a: When the call demand information of the current call node meets the node index threshold requirement and the call demand information is greater than the service capacity expansion index of the current call node, increase the service quantity of the current call node so that the service capacity expansion index of the current call node is greater than the call demand information, where the service capacity expansion index is obtained according to the service quantity and the capacity expansion index threshold of a single service. For example, the service capacity expansion index is obtained by multiplying the service quantity by the capacity expansion index threshold of a single service.
[0082] S3223a: When the call demand information of the current call node does not meet the node index threshold requirement, a warning message is generated.
[0083] connect Figure 4For example, the resource occupancy of Service 1 called by Invocation Node 1 is 1G of memory and 10 millicores of CPU. The upstream call metric thresholds for Service 1 include: scaling out when the concurrency number is greater than 10, and scaling in when the concurrency number is less than 5. The downstream call requirement for Service 1 is a concurrency number of 5. The number of Service 1 in Invocation Node 1 is 1, and the node metric threshold configured for Invocation Node 1 is a maximum connection number of 50. When the concurrency number in the customer call requirement information grows to 12 (greater than 10) and the maximum connection number is 40, Service 1 in Invocation Node 1 reaches the scaling-out condition, and scaling-out calculation is performed to increase the number of Service 1 until the service scaling metric meets the call requirement information. Specifically, the expected number of Service 1 in Invocation Node 1 is 2.
[0084] According to the upstream call metric thresholds of Service 1 {scaling-in metric threshold: scaling-in condition is triggered when the concurrency number is less than 5, scaling-out metric threshold: scaling-out condition is triggered when the concurrency number is greater than 10}, after the number of services increases to 2, the maximum concurrency number that the invocation node can support is 20. The maximum concurrency number of 20 meets the call requirement, and the maximum connection number of 40 is within the range of the maximum connection number of 100 set for Invocation Node 1. The available resources become 10 - 1 = 9G of memory and 765 - 10 = 755 millicores of CPU, and the available resources are sufficient. After the above analysis, 2 instances of Service 1 after scaling out can meet the customer call requirement. Therefore, it is determined that the updated number of Service 1 is 2, and the available resource information is updated to 9G of memory and 755 millicores of CPU.
[0085] The call requirement of Invocation Node 1 for Invocation Node 2 becomes: the concurrency number is 10. The upstream call metric thresholds for Service 2 in Invocation Node 2 are: scaling-in is triggered when the concurrency number is less than 10, and scaling-out is triggered when the concurrency number is greater than 20. The number of Service 2 in Invocation Node 2 is 2. The upstream call metric thresholds for Service 2 include: scaling-out when the concurrency number is greater than 10, and scaling-in when the concurrency number is less than 5. Then the service scaling metric for Invocation Node 2 is: scaling-out when the concurrency number is greater than 20. The call requirement of Invocation Node 1 for Invocation Node 2 does not trigger the service scaling metric. Then the number of services in Invocation Node 2 already meets the call requirement of Invocation Node 1 for Service 2. Subsequently, continue to calculate the number of services in Invocation Node 3, which also meets the requirements. Finally, it is determined that the number of services in Invocation Node 1 needs to be changed.
[0086] If after running for a period of time, the customer's call requirement changes to a maximum connection number of 60, which exceeds the node metric threshold of Invocation Node 1: the maximum connection number of 50, Invocation Node 1 cannot meet the customer call requirement, and a warning message is generated.
[0087] In some embodiments of the present disclosure, the upstream call metric thresholds include a scaling-in metric threshold, such as Figure 6bAs shown, determining the estimated quantity of services of the current call node according to the node metric threshold of the current call node, the call demand information, the quantity of services, the resource occupancy of a single service, and the upstream call metric threshold includes:
[0088] S3221b: Obtain the call demand information of the current call node.
[0089] S3222b: When the call demand information of the current call node meets the node metric threshold requirement and the call demand information is less than the service scaling-down metric of the current call node, reduce the quantity of services of the current call node so that the service scaling-down metric of the current call node is less than the call demand information, where the service scaling-down metric is obtained according to the quantity of services and the scaling-down metric threshold of a single service. For example, the service scaling-up metric is obtained by multiplying the quantity of services by the scaling-up metric threshold of a single service.
[0090] S3223b: When the call demand information of the current call node does not meet the node metric threshold requirement, generate a warning message.
[0091] Received Figure 4For example, the resource consumption of Service 1 called by Invocation Node 1 is 1G of memory and 10 millicores of CPU. The upstream call metric thresholds for Service 1 include: scaling out when the concurrency is greater than 10 and scaling in when the concurrency is less than 5. The downstream call requirement for Service 1 is a concurrency of 5. The number of Service 1 in Invocation Node 1 is 1, and the node metric threshold configured for Invocation Node 1 is a maximum connection number of 50. Invocation Node 2 is configured with 2 instances of Service 2. The resource consumption of each Service 2 is 1G of memory and 5 millicores of CPU. The upstream call metric thresholds for Service 2 include: scaling out when the concurrency is greater than 10 and scaling in when the concurrency is less than 5. The downstream call requirement for Service 2 is a concurrency of 5. The node metric threshold configured for Invocation Node 2 is a maximum connection number of 200. The downstream invocation node of Invocation Node 2 is Invocation Node 3. Invocation Node 3 is configured with 3 instances of Service 3. The resource consumption of each Service 3 is 1G of memory and 5 millicores of CPU. The upstream call metric thresholds for Service 3 include: scaling out when the concurrency is greater than 10 and scaling in when the concurrency is less than 5. The node metric threshold configured for Invocation Node 3 is a maximum connection number of 1000. Service 3 is a service in the underlying invocation node and has no downstream call requirement. The initial available resources are 16G of memory and 800 millicores of CPU. After allocation according to the above configuration, the current available resources are: 16 - 1 - 1×2 - 1×3 = 10G of memory, 800 - 10 - 5×2 - 5×3 = 765 millicores of CPU. When the concurrency in the customer call demand information decreases to 4 and the maximum connection number is 40, the number of Service 1 should be reduced. However, the existing number of Service 1 is already the minimum, so no reduction is needed. For Invocation Node 2, the upstream call requirement is a concurrency of 5. The number of services in Invocation Node 2 is 2. The service scaling-in metric for Invocation Node 2 is: scaling in when the concurrency is less than 10. Since the concurrency has triggered the service scaling-in metric, the expected number of Service 2 is determined to be 1. If it is determined that the available resources meet the conditions, the updated number of Service 2 is determined to be 1. For Invocation Node 3, the upstream call requirement is updated from a concurrency of 10 to a concurrency of 5. The service scaling-in metric for Invocation Node 3 is: scaling in when the concurrency is less than 15. Since the service scaling-in metric is triggered, the number of Service 3 is reduced. When the number of Service 3 is reduced to 1, it reaches the minimum and cannot be further reduced. So the updated number of Service 3 is determined to be 1. After the available resource information is updated, it is 16 - 1 - 1 - 1 = 13G of memory, 800 - 10 - 5 - 5 = 780 millicores of CPU.
[0092] In an embodiment of the present disclosure, the number of services of each invocation node is calculated, and the available resource information is updated. Specifically, in response to the current available resource information meeting the requirements of the expected number and resource consumption of the services of the current invocation node, determining the updated number of the services of the current invocation node and updating the available resource information includes:
[0093] When the current available resource information is greater than the expected resource consumption of the service of the current calling node, the expected number of the service of the current calling node is used as the updated number of the service of the current calling node, and the available resource information is updated according to the current available resource information and the expected resource consumption of the service of the current calling node, where the expected resource consumption of the service is obtained according to the expected number and the resource consumption of the service.
[0094] In some embodiments of the present disclosure, the expected resource consumption of the service is obtained by multiplying the expected number of the service by the resource consumption, and the current available resource information is subtracted from the expected resource consumption of the service to obtain the latest value of the available resource information, and the latest value is used to update the available resource information.
[0095] In some embodiments of the present disclosure, the control method further includes: generating a warning message in response to the current available resource information not meeting the requirements of the expected number and the resource consumption of the service of the current calling node.
[0096] For example, the current available resource information is 1G of memory, the expected number of the service of the current calling node increases by 2, and the resource consumption of the service of the current calling node is 1G. Then, 1G of memory of the current available resource information is less than 1 * 2G of memory, so the current available resource information cannot meet the requirements, and a warning message is generated.
[0097] In the embodiments of the present disclosure, the generated warning message can include the reason for the warning, so that the operation and maintenance personnel can determine the problem point.
[0098] S330: According to the updated number of the services in each calling node on the service call link, scale the calling nodes in the service call link up or down. The operation of scaling the calling nodes up or down in the embodiments of the present disclosure is a process of allocating resources, and the allocation is performed according to the resource consumption of each service.
[0099] In some embodiments of the present disclosure, for the operation of scaling down, it can be executed with a delay. The delay time can be configured for each calling node, or the delay time can be set for the service call link uniformly. By delaying the execution of scaling down, the scaling down operations in a short time are reduced, and the situation of blocked service calls is avoided. The step of scaling the calling nodes in the service call link up or down according to the updated number of the services in each calling node on the service call link includes:
[0100] For each calling node on the service call link, the following steps are executed:
[0101] When the number of updated services in the calling node is less than the number before the update, perform delayed execution timing according to the scaling-down configuration of the calling node; when the delay time reaches the setting of the scaling-down configuration delay, close the services in the calling node according to the number of updated services, and recycle the resources corresponding to the closed services.
[0102] When the number of updated services in each calling node on the service call link becomes smaller compared to the number before the update, perform delay according to the scaling-down configuration of each calling node on the service call link. After the delay, according to the number of updated services in each calling node on the service call link, close the services in each calling node in the service call link, and recycle the resources corresponding to the closed services. Delay the closing of the redundant services with a number greater than the number of updated services. For example, if the number of services in calling node 2 before the update is 2 and the number of updated services is 1, then close 1 service and recycle the resources corresponding to this 1 service, and the available resources increase by 1G of memory and 5 millicores of CPU.
[0103] In some embodiments of the present disclosure, for the scaling-up operation, immediate scaling-up can be performed. The scaling-up and scaling-down of each calling node in the service call link according to the number of updated services in each calling node in the service call link includes:
[0104] For each calling node on the service call link, perform the following steps:
[0105] When the number of updated services in the calling node is greater than the number before the update, call the resources corresponding to the available resource information, and increase the deployment of the services in the calling node according to the number of updated services in the calling node.
[0106] When the number of updated services in each calling node on the service call link becomes larger, call the resources corresponding to the available resource information, and increase the deployment of the services of each calling node on the service call link according to the number of updated services in each calling node on the service call link. For example, if the number of services in calling node 1 before the update is 1 and the number of updated services is 2, then call the available resources to newly deploy 1 service, and the available resources decrease by 1G of memory and 5 millicores of CPU.
[0107] In the early stage, each index threshold in the embodiments of the present disclosure can be set by experience. After the system in the embodiments of the present disclosure runs for a period of time, the index threshold can be updated according to the recorded scaling-up and scaling-down operation records to improve the accuracy of controlling scaling-up and scaling-down. The method in the embodiments of the present disclosure further includes:
[0108] Count the number of times of execution for scaling in and out and the number of services of each calling node within the statistical period. According to the number of times of scaling in and out and the number of services of each calling node, update the node metric thresholds of each calling node, the preset number of services in each calling node, and the upstream call metric thresholds of the services in each calling node.
[0109] For example, when the number of services of an execution node remains at 3 for a long time or frequently, update the preset number of services to 3. Another example is that when the number of services of an execution node is often 1 and the concurrency is less than the scaling-in condition, the concurrency value for triggering scaling in can be reduced.
[0110] In some embodiments of the present disclosure, the initial available resource information is as follows: Cluster hardware conditions: 2 4-core CPUs, that is, 8000 millicores (m), and 16G of memory. The operation process of the system and method in the embodiments of the present disclosure is as follows:
[0111] As Figure 7 shown in the cluster resource architecture, taking the call link of service 3 in customer call system 1 as an example, one service 3 is deployed on the first calling node, one sub-service 2 is deployed on the second calling node, and three underlying middleware 5 are deployed on the third calling node. The first calling node calls the second calling node, and the second calling node calls the third calling node. The node metric threshold configured for the first calling node is a maximum of 20 concurrencies. The resource occupancy of service 3 is 1G of memory and 10 millicores of CPU. The upstream call metric thresholds of service 3 include: when the concurrency is greater than 10, the scaling-out condition is reached for scaling out; when the concurrency is less than 5, the scaling-in condition is reached for scaling in, and it is set to execute 1 hour after the scaling-in condition is reached. The downstream call requirement of service 3 is a concurrency of 5. The node metric threshold configured for the second calling node is a maximum of 20 concurrencies. The resource occupancy of sub-service 2 is 1G of memory and 10 millicores of CPU. The upstream call metric thresholds of sub-service 2 include: when the concurrency is greater than or equal to 10, the scaling-out condition is reached for scaling out; when the concurrency is less than or equal to 5, the scaling-in condition is reached for scaling in, and it is set to execute 1 hour after the scaling-in condition is reached. The downstream call requirement of sub-service 2 is a concurrency of 5. The node metric thresholds of the third calling node include: a maximum of 40 concurrencies and a maximum of 40 connections. The resource occupancy of the underlying middleware 5 is 4G of memory and 10 millicores of CPU. The upstream call metric thresholds of the underlying middleware 5 include: scaling in and out are not allowed. After the metric thresholds on the above link are determined, service 3 directly provides an interface to the customer. After the customer calls the service 3 interface, service 3 needs to call sub-service 2. After sub-service 2 calls the underlying middleware, the final result is summarized to service 3 and then returned to the customer.
[0112] Expansion plan: When a customer calls the interface of Service 3, the agreed concurrency is 5, and the CPU and memory utilization rates reach 20% and 50% respectively. At this time, a second customer also wants to call the interface of Service 3 with the same concurrency of 5. At this time, Service 3 can provide services normally, but the concurrency of 10 has reached the service expansion index of the first call node where Service 3 is located. According to the upstream call index threshold of Service 3 on the first call node, increase the deployment quantity of Service 3 until the service expansion index of the first call node meets the customer call requirements, and then judge whether the occupied resource amount corresponding to the increased Service 3 exceeds the available resources, that is: the available resources before update are 2G of memory and 7950 millicores of CPU. Adding a new Service 3 occupies 1G of memory and 10 millicores of CPU, and the available resources meet the new requirements; determine that the update quantity of Service 3 is 2, and the downstream call requirement of the first call node after update is a concurrency of 10. There is 1 sub-service 2 deployed on the second call node, and the expansion service index of the second call node is: the concurrency of 10 triggers the expansion condition. The downstream call requirement of the first call node triggers the expansion condition of the second call node, and the downstream call requirement of the first call node does not trigger the node index threshold of the second call node. It is necessary to expand the sub-service 2. After the quantity of the sub-service 2 is updated to 2, the expansion service index of the second call node is: the concurrency greater than or equal to 20 triggers the expansion condition. The downstream call requirement of the first call node does not trigger the expansion condition of the updated second call node. At this time, the available resources before update are 1G of memory and 7940 millicores of CPU, which meet the resource occupancy requirements of the newly added sub-service 2. The available resources after update are: 0G of memory and 7930 millicores of CPU. The downstream call requirement of the second call node after update is a concurrency of 10, and the updated one does not trigger the node threshold index of the third call node, and the underlying middleware 5 does not scale in or out, so the quantity of the middleware 5 does not need to be updated. Determine that the expansion plan is executable, and the available resources meet the overall expansion plan requirements. However, if there are still new customers calling Service 3 with an agreed concurrency of 10, which reaches the expansion service index of the first call node (the concurrency greater than or equal to 20 triggers the expansion condition), it will be found that there is not enough memory to allocate during the expansion calculation at this time, and the expansion operation cannot be completed. At this time, the scaling control system will trigger the warning function and display the memory shortage encountered during the pre-expansion calculation, which can quickly locate which specific resources are insufficient.
[0113] Shrinkage plan: After sub-service 2 is scaled up to 2 services, as the customer calls end, the concurrency of sub-service 2 gradually decreases to within 2×5 = 10. At this time, the scaling control system determines the shrinkage condition for reaching sub-service 2 from the collected metrics. At this time, the shrinkage plan is executed with a delay (i.e., executed after meeting the delay condition): start timing one hour from the time when the shrinkage condition for reaching sub-service 2 is touched. If the customer call concurrency does not reach 10 or more within one hour, then trigger the shrinkage operation and record the touched condition and the shrinkage result. If the customer call concurrency reaches 10 or more within one hour, then continue to maintain the existing deployment plan. The shrinkage operation of service 3 is the same as that of sub-service 2. And the underlying middleware 3 does not perform scaling operations because it has customized metrics for not scaling.
[0114] In the systems and methods of the present disclosure embodiments, by collecting the metric thresholds of each call node, according to the call requirements of customers, the number of services in the call nodes is gradually updated from upstream to downstream. During the gradual update process, it is judged that the available resources meet the conditions, thereby ensuring the reliability and stability of scaling for the entire link, realizing automatic scaling control for the call link, improving the operation and maintenance efficiency, and reducing the operation and maintenance difficulty. Using the systems and methods of the present disclosure embodiments, development and operation and maintenance personnel can reduce the attention on services caused by traffic increase or decrease, effectively reducing the communication cost between personnel; realizing flexible scheduling, and can better release service resources during the business off-peak period for other business lines to use, and making more efficient use of service resources.
[0115] On the other hand, the present disclosure also provides a scaling control device 800, as Figure 8 shown. The control device 800 includes an acquisition module 810, a scaling plan formation module 820, and an execution module 830. Among them,
[0116] The acquisition module 810 is configured to acquire available resource information and customer call requirement information of the target service system. Among them, the target service system includes a service call link, services are deployed on each call node on the service call link, each service is configured with an occupied resource amount and an upstream call metric threshold, each call node is configured with a node metric threshold, and call nodes other than the most downstream call node on the service call link are also configured with downstream call requirements;
[0117] The scaling plan formation module 820 is configured to traverse each call node on the service call link from the most upstream call node to the most downstream call node of the service call link, and determine the updated number of services of each call node and update the available resource information according to the node metric threshold, call requirement information, number of services, occupied resource amount of a single service, and upstream call metric threshold of each call node.
[0118] The execution module 830 is configured to scale the call nodes in the service call link up or down according to the update quantity of the services in each call node on the service call link.
[0119] In some embodiments of the present disclosure, the scaling plan forming module 820 is further configured to traverse each call node on the service call link from the most upstream call node to the most downstream call node of the service call link, and for each call node, perform the following steps:
[0120] Determine the predicted quantity of the services of the current call node according to the node metric threshold, call demand information, quantity of the services, resource occupancy of a single service, and upstream call metric threshold of the current call node;
[0121] In response to the current available resource information meeting the requirements of the predicted quantity and resource occupancy of the services of the current call node, determine the update quantity of the services of the current call node and update the available resource information.
[0122] In some embodiments of the present disclosure, the scaling plan forming module 820 is further configured to:
[0123] In response to the current call node being the most upstream call node on the service call link, use the customer call demand information as the call demand information of the current call node;
[0124] In response to the current call node being a call node after the most upstream call node on the service call link, use the product of the downstream call demand of the services of the upstream call node of the current call node and the update quantity of the services of the upstream call node of the current call node as the call demand information of the current call node.
[0125] In some embodiments of the present disclosure, the upstream call metric threshold includes an expansion metric threshold, and the scaling plan forming module 820 is further configured to:
[0126] When the call demand information of the current call node meets the node metric threshold requirement and the call demand information is greater than the service expansion metric of the current call node, increase the quantity of the services of the current call node so that the service expansion metric of the current call node is greater than the call demand information, where the service expansion metric is obtained according to the quantity of the services and the expansion metric threshold of a single service;
[0127] When the call demand information of the current call node does not meet the node metric threshold requirement, generate a warning message.
[0128] In some embodiments of the present disclosure, the upstream call metric threshold includes a scaling-down metric threshold, and the scaling plan forming module 820 is further configured to:
[0129] When the call requirement information of the current call node meets the requirements of the node metric threshold and the call requirement information is less than the service scaling-down metric of the current call node, reduce the number of services of the current call node so that the service scaling-down metric of the current call node is less than the call requirement information, where the service scaling-down metric is obtained based on the number of services and the scaling-down metric threshold of a single service;
[0130] When the call requirement information of the current call node does not meet the requirements of the node metric threshold, generate a warning message.
[0131] In some embodiments of the present disclosure, the scaling scheme forming module 820 is further configured to:
[0132] When the current available resource information is greater than the product of the expected number of services of the current call node and the occupied resource amount, use the expected number of services of the current call node as the updated number of services of the current call node, and update the available resource information according to the current available resource information and the expected resource occupancy of the services of the current call node, where the expected resource occupancy of the services is obtained based on the expected number of services and the occupied resource amount.
[0133] In some embodiments of the present disclosure, the scaling scheme forming module 820 is further configured to: in response to the current available resource information not meeting the requirements of the expected number of services and the occupied resource amount of the current call node, generate a warning message.
[0134] In some embodiments of the present disclosure, the execution module 830 is further configured to:
[0135] For each call node on the service call link, perform the following steps:
[0136] When the updated number of services in the call node is less than the number before the update, perform a delay execution timing according to the scaling-down configuration of the call node;
[0137] When the delay time reaches the setting of the scaling-down configuration delay, close the services in the call node according to the updated number, and recycle the resources corresponding to the closed services.
[0138] In some embodiments of the present disclosure, the execution module 830 is further configured to:
[0139] For each call node on the service call link, perform the following steps:
[0140] When the updated number of services in the call node is greater than the number before the update, call the resources corresponding to the available resource information, and increase the deployment of the services in the call node according to the updated number of services in the call node.
[0141] In some embodiments of the present disclosure, a threshold metric update module 840 is further included, and the threshold metric update module 840 is configured to: count the number of times of execution of scaling and the number of services of each calling node within a statistical period, and update the node metric thresholds of each calling node, the preset number of services in each calling node, and the upstream call metric thresholds of the services in each calling node according to the number of times of scaling and the number of services of each calling node.
[0142] In some embodiments, the device may incorporate the features of the method of any one of the embodiments, and vice versa, which will not be elaborated herein.
[0143] In an embodiment of the present disclosure, an electronic device is provided, including: a processor and a memory storing a computer program, and the processor is configured to execute the method of any one of the embodiments of the present disclosure when running the computer program.
[0144] Figure 9 A schematic diagram showing an electronic device 900 that can implement the method of the embodiments of the present disclosure or implement the embodiments of the present disclosure is shown. In some embodiments, there may be more or fewer electronic devices than shown in the figure. In some embodiments, a single or multiple electronic devices may be utilized for implementation. In some embodiments, cloud or distributed electronic devices may be utilized for implementation.
[0145] As Figure 9 shown, the electronic device 900 includes a central processing unit (CPU) 901, which can perform various appropriate operations and processes according to the programs and / or data stored in the read-only memory (ROM) 902 and / or the programs and / or data loaded from the storage section 908 into the random access memory (RAM) 903. The CPU 901 may be a multi-core processor or may include multiple processors. In some embodiments, the CPU 901 may include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a neural network processing unit (NPU), a digital signal processor (DSP), and so on. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The CPU 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.
[0146] The above-mentioned processor and the memory are jointly used to execute the program stored in the memory, and when the program is executed by a computer, it can implement the steps or functions of the methods described in the above embodiments.
[0147] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, a touch screen, etc.; an output section 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 910 as needed so that a computer program read therefrom is installed into the storage section 908 as needed. Figure 9 Only some components are schematically shown, and it does not mean that the computer system 900 only includes Figure 9 the components shown.
[0148] In some embodiments, the electronic device 900 refers to a mobile terminal or a computer, including a mobile phone, an in-vehicle terminal, a smart TV, etc. Taking the mobile phone as an example, the electronic device 900 further includes a display screen with touch function, an external speaker, a gyroscope, a camera, 4G / 5G antennas and other device modules.
[0149] The systems, devices, modules or units illustrated in the above embodiments can be implemented by a computer or its associated components. The computer can be, for example, a mobile terminal, a smart phone, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server or a combination thereof
[0150] Although not shown, in the embodiments of the present disclosure, a storage medium is provided, and the storage medium stores a computer program, and the computer program is configured to execute the method of any embodiment of the present disclosure when being run.
[0151] Although not shown, in the embodiments of the present disclosure, a program product is provided, and the program product includes a computer program, wherein the computer program implements the method of any embodiment of the present disclosure when being executed by a processor.
[0152] The storage media of the embodiments of the present disclosure include permanent and non-permanent, removable and non-removable articles that can implement information storage by any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0153] The methods, programs, systems, devices, etc. of the embodiments of the present disclosure can be executed or implemented in a single or multiple networked computers, and can also be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks can be executed by remote processing devices connected through a communication network.
[0154] Those skilled in the art should understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, those skilled in the art can think that the implementation of the functional modules / units or controllers and related method steps clarified in the above embodiments can be achieved in a manner combining software, hardware, and soft / hardware.
[0155] Unless explicitly stated, the actions or steps of the methods and programs described according to the embodiments of the present disclosure do not necessarily have to be executed in a specific order and can still achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be beneficial.
[0156] In this article, multiple embodiments of the present disclosure are described. However, for the sake of brevity, the descriptions of each embodiment are not exhaustive, and the same or similar features or parts between the various embodiments may be omitted. In this article, "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean applicable to at least one embodiment or example according to the present disclosure, rather than all embodiments. The above terms do not necessarily refer to the same embodiment or example. Without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0157] Exemplary systems and methods of the present disclosure have been specifically shown and described with reference to the above embodiments, which are merely examples of the best mode for implementing the systems and methods. Those skilled in the art will understand that various changes can be made to the embodiments of the systems and methods described herein when implementing the systems and / or methods without departing from the spirit and scope of the present disclosure as defined in the appended claims.
Claims
1. A method for controlling expansion and contraction, characterized in that: The steps include: Obtaining available resource information and customer call demand information of the target service system, wherein the target service system includes a service call link, each call node on the service call link is deployed with a service, each service is configured with an occupied resource amount and an upstream call index threshold, the upstream call index threshold is a concurrent number threshold available for call, each call node is configured with a node index threshold, and the call nodes other than the most downstream call node on the service call link are also configured with downstream call demands; Traverse each call node on the service call link from the most upstream call node to the most downstream call node, and determine the updated number of services of each call node and update the available resource information according to the node indicator threshold, call demand information, number of services, resource usage of a single service and upstream call indicator threshold of each call node; According to the updated number of services in each calling node on the service calling link, each calling node in the service calling link is expanded or reduced in capacity.
2. The method according to claim 1, characterized in that The step of determining the update quantity of services of each calling node and updating available resource information includes: Traverse each call node on the service call chain from the most upstream call node to the most downstream call node, and perform the following steps for each call node: Determine the estimated number of services of the current calling node based on the node indicator threshold of the current calling node, the calling demand information, the number of services, the amount of resources occupied by a single service, and the upstream calling indicator threshold; In response to the current available resource information satisfying the estimated quantity and occupied resource amount requirements of the service of the current calling node, the updated quantity of the service of the current calling node is determined and the available resource information is updated.
3. The method according to claim 2, characterized in that Determining the calling requirement information of the current calling node includes: In response to the current calling node being the most upstream calling node on the service calling link, using the client calling demand information as the calling demand information of the current calling node; In response to the current calling node being the calling node after the most upstream calling node on the service calling link, the product of the downstream calling demand of the service of the upstream calling node of the current calling node and the updated number of the service of the upstream calling node of the current calling node is used as the calling demand information of the current calling node.
4. The method according to claim 2, characterized in that: The upstream call index threshold includes a capacity expansion index threshold, and the estimated number of services of the current calling node is determined according to the node index threshold of the current calling node, call demand information, the number of services, the amount of resources occupied by a single service, and the upstream call index threshold, including: When the call demand information of the current call node meets the node index threshold requirement, and the call demand information is greater than the service expansion index of the current call node, increase the number of services of the current call node so that the service expansion index of the current call node is greater than the call demand information, wherein the service expansion index is obtained according to the number of services and the expansion index threshold of a single service; When the calling demand information of the current calling node does not meet the node indicator threshold requirements, a warning message is generated.
5. The method according to claim 2, characterized in that: The upstream call index threshold includes a shrinkage index threshold, and the estimated number of services of the current calling node is determined according to the node index threshold of the current calling node, call demand information, the number of services, the amount of resources occupied by a single service, and the upstream call index threshold, including: When the call demand information of the current call node meets the node index threshold requirement, and the call demand information is less than the service shrinkage index of the current call node, the number of services of the current call node is reduced so that the service shrinkage index is less than the call demand information, wherein the service shrinkage index is obtained according to the number of services and the shrinkage index threshold of a single service; When the calling demand information of the current calling node does not meet the node indicator threshold requirements, a warning message is generated.
6. The method according to claim 2, characterized in that In response to the current available resource information satisfying the estimated number of services of the current calling node and the requirements of the occupied resource amount, determining the updated number of services of the current calling node and updating the available resource information includes: When the current available resource information is greater than the estimated amount of resources occupied by the services of the current calling node, the estimated number of services of the current calling node is used as the updated number of services of the current calling node, and the available resource information is updated based on the current available resource information and the estimated amount of resources occupied by the services of the current calling node, wherein the estimated amount of resources occupied by the services is obtained based on the estimated number and occupied resources of the services.
7. The method according to claim 1, characterized in that The step of scaling each calling node in the service calling link according to the updated number of services in each calling node on the service calling link includes: For each call node on the service call link, perform the following steps: When the updated number of services in the calling node is less than the number before the update, delay execution timing according to the scaling-down configuration of the calling node; When the delay time reaches the setting of the shrinking configuration delay, the service in the calling node is closed according to the update quantity, and the resources corresponding to the closed service are recovered.
8. The method according to claim 1, characterized in that The step of scaling each calling node in the service calling link according to the updated number of services in each calling node on the service calling link includes: For each call node on the service call link, perform the following steps: When the updated number of services in the calling node is greater than the number before the update, the resources corresponding to the available resource information are called, and the services in the calling node are additionally deployed according to the updated number of services in the calling node.
9. A capacity expansion and contraction control device, characterized in that: It includes an acquisition module, a scaling plan forming module and an execution module, wherein: The acquisition module is configured to acquire available resource information and customer call demand information of the target service system, wherein the target service system includes a service call link, each call node on the service call link is deployed with a service, each service is configured with an occupied resource amount and an upstream call index threshold, the upstream call index threshold is a concurrent number threshold available for call, each call node is configured with a node index threshold, and the call nodes on the service call link except the most downstream call node are also configured with downstream call demands; The expansion and contraction plan forming module is configured to traverse each calling node on the service calling link from the most upstream calling node to the most downstream calling node of the service calling link, and determine the updated number of services of each calling node and update the available resource information according to the node indicator threshold, calling demand information, number of services, occupied resources of a single service and upstream calling indicator threshold of each calling node; The execution module is configured to scale up or down each calling node in the service calling link according to the updated number of services in each calling node on the service calling link.
10. An electronic device, characterized in that: include: A processor and a memory storing a computer program, wherein the processor is configured to execute the method according to any one of claims 1 to 8 when running the computer program.
11. A program product comprising a computer program, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Service resource control method and device, equipment and storage medium
CN110196767A