Elastic Resource Scheduling Method and System for Power Grid Data Middleware Service
The method and system for elastic resource allocation in electric grid data platforms address inefficiencies by analyzing service call relationships to optimize resource use, reducing waste and ensuring stable operation.
Patent Information
- Application Number
- CN202111332780.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-11
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-11
AI Technical Summary
In traditional elastic resource scheduling, in power grid data, the implicit dependence between services is not fully considered, resulting in insufficient resource allocation and waste.
By analyzing the request log, using the breadth-first search algorithm BFS to obtain the service's request path and call relationship, and combining the remaining resources of the server, elastic resources are accurately allocated.
Accurate resource allocation under the conditions of satisfying the service basic QPS, reducing resource waste and improving resource utilization efficiency.
Smart Images

Figure CN115951990B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer systems, and particularly to a method and system for elastic resource scheduling for power grid data middle - platform services. Background Art
[0002] The power grid data middle - platform contains a large number of services. There are implicit calls between different services. The request path formed by a request invoking each service can form a directed acyclic graph (DAG). Traditional elastic resource scheduling is to detect that a service is in the peak period and then allocate more resources to it. However, there are many implicit dependency relationships between services in the power grid data middle - platform. Simply and crudely expanding the capacity for a certain service will cause unnecessary waste and the resource allocation is not accurate enough. Summary of the Invention
[0003] To solve the above problems, this application provides a method for elastic resource scheduling for power grid data middle - platform services, including:
[0004] Obtain the request logs of power grid data middle - platform services;
[0005] Parse the request logs to obtain the request paths of the services;
[0006] Through the request paths, obtain the call relationships between each service;
[0007] According to the remaining resources of the server and the call relationships, allocate elastic resources to the services under the condition of meeting the basic queries - per - second condition of the services.
[0008] Preferably, parsing the request logs to obtain the request paths of the services includes:
[0009] Parse the request logs through the breadth - first search algorithm (BFS) to obtain the request paths formed by each request invoking each service, and the nodes included in the request paths;
[0010] Use characters to identify each node;
[0011] Form a string with the nodes passed by each path;
[0012] Use the string to identify the request paths of the services.
[0013] Preferably, it further includes:
[0014] Obtain the number of times each node is accessed and the number of times the request path is accessed.
[0015] Preferably, through the request paths, obtaining the call relationships between each service includes:
[0016] Obtain the intersecting nodes among various services according to the request path;
[0017] Obtain the call relationships among various services through the intersecting nodes.
[0018] Preferably, according to the resource occupancy and call relationships, elastic resources are allocated to services under the condition of meeting the basic queries per second (QPS) of the services, including:
[0019] When the remaining resources of the server are sufficient, calculate the remaining resources of the server and the recyclable resources;
[0020] Recycle the excessive service resources of the server;
[0021] Allocate the excessive service resources to the nodes of each service to which resources are to be allocated and the nodes of the services having call relationships with the services to which resources are to be allocated, where the excessive service resources are greater than or equal to the service resources required for half of the basic queries per second (QPS) of their nodes;
[0022] When the remaining resources of the server are insufficient, recycle the resources of services during the off-peak period;
[0023] Calculate the proportion of the resources required by services during the peak period in the total resources required by services during the peak period;
[0024] Obtain the number of nodes in the request path of the services during the peak period to be allocated, and allocate the remaining resources according to the proportion of the number of machines required according to the node number in the total number of machines required.
[0025] This application also provides an elastic resource scheduling system for power grid data center services, including:
[0026] A request log acquisition module, configured to acquire the request logs of power grid data center services;
[0027] A request path acquisition module, configured to parse the request logs to obtain the request paths of services;
[0028] A call relationship acquisition module, configured to obtain the call relationships among various services through the request paths;
[0029] A resource allocation module, configured to allocate elastic resources to the services under the condition of meeting the basic queries per second condition of the services according to the remaining resources of the server and the call relationships.
[0030] Preferably, the request log acquisition module includes:
[0031] A parsing sub-module, configured to parse the request logs through the breadth-first search (BFS) algorithm to obtain the request paths formed by each request calling various services; and each node included in the request paths;
[0032] An identification sub-module, used to identify each of the nodes using characters;
[0033] A string composition sub-module, used to compose the nodes passed by each path into a string;
[0034] An identification subunit, used to identify the request path of the service using the string.
[0035] Preferably, it further includes:
[0036] An access times acquisition sub-module, used to acquire the number of times each node is accessed and the number of times the request path is accessed.
[0037] Preferably, the call relationship acquisition module includes:
[0038] An intersection node acquisition sub-module, used to acquire the intersecting nodes between each service according to the request path;
[0039] A call relationship acquisition sub-module, used to acquire the call relationship between each service through the intersecting nodes.
[0040] Preferably, the resource allocation module includes:
[0041] A first resource calculation sub-module, used to calculate the remaining resources of the server and the recyclable resources when the remaining resources of the server are sufficient;
[0042] A first recycling sub-module, used to recycle the excess service resources of the server;
[0043] A first resource allocation sub-module, used to allocate the excess service resources to the nodes of each service waiting for resource allocation and the nodes of the services having a call relationship with the service waiting for resource allocation, and the excess service resources are greater than or equal to half of the service resources required by the basic queries per second (QPS) of its nodes;
[0044] A second recycling sub-module, used to recycle the resources of low-peak services when the remaining resources of the server are insufficient;
[0045] A ratio calculation sub-module, used to calculate the ratio of the resources required by the peak services to the total resources required by the peak services;
[0046] A second resource allocation sub-module, used to acquire the number of nodes of the request path of the peak service waiting for resource allocation, and allocate the remaining resources according to the ratio of the number of machines required by the node demand to the total number of machines required. Brief Description of the Drawings
[0047] Figure 1 It is a schematic flowchart of a method for elastic resource scheduling for power grid data middle platform services provided by this application;
[0048] Figure 2 It is a simplified flowchart of the elastic resource scheduling method involved in this application;
[0049] Figure 3 It is a schematic diagram of log collection involved in this application;
[0050] Figure 4 It is a flowchart of converting a request path into a unique identifier involved in this application;
[0051] Figure 5 It is the service classification when resources are insufficient involved in this application;
[0052] Figure 6 It is a schematic diagram of experimental service invocation involved in this application;
[0053] Figure 7 It is an application flowchart of the elastic resource scheduling method for the power grid data center platform service involved in this application;
[0054] Figure 8 It is a schematic diagram of the structure of an elastic resource scheduling system for the power grid data center platform service provided by this application. Detailed implementation manners
[0055] Many specific details are set forth in the following description in order to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this application. Therefore, this application is not limited by the specific implementations disclosed below.
[0056] Figure 1 It is a schematic flowchart of an elastic resource scheduling method for the power grid data center platform service provided by this application. The method includes the following steps:
[0057] Step S101: Obtain the request logs of the power grid data center platform service.
[0058] The elastic resource scheduling method for the power grid data center platform service provided by this application allocates elastic resources based on real-time log parsing. This method obtains the call relationships between services through log parsing for elastic resource allocation. Its simplified flowchart is as Figure 2 shown and includes:
[0059] 1) Log collection.
[0060] 2) Log parsing.
[0061] 3) Elastic resource allocation.
[0062] To improve the efficiency of log collection and logs, a Redis cluster is used to store logs, and log details are obtained through the Redis cluster for log parsing. When the logs expire, the system will retrieve them from the Redis cluster and encapsulate them into files for storage in the HDFS cluster.
[0063] Log collection is the basis of log collection. To reflect the dependency relationships between services, Redis is used in this article to record logs. Each request contains a unique RequestID, and the logs of this request will be recorded in Redis and identified by this RequestID. All nodes on the request path insert the logs into the content with the key RequestID. Figure 3 is an example of log collection.
[0064] Step S102, parse the request log to obtain the request path of the service.
[0065] Parse the request log through the breadth-first search algorithm BFS to obtain the request path formed by each request calling each service; and each node included in the request path; use characters to identify each node; form a string from the nodes passed by each path; use the string to identify the request path of the service. Also, obtain the number of times each node is accessed and the number of times the request path is accessed.
[0066] Log parsing is the basis for elastic resource allocation. Elastic resource allocation is to allocate resources for all nodes on the request path. To distinguish different request paths, the BFS algorithm is used in this article to parse the request logs, and each request path will form a DAG (directed acyclic graph). As Figure 4 shown, the request path of the log contains 4 nodes, namely A, B, C, and D. Node A accesses B and C respectively, and nodes B and C access node D respectively. Node D is the end point of the request path. After parsing the request, the DAG will be stored in the form of a string, and this string is unique and used to identify the request path. As shown in the right part of the figure, the list of child nodes is in the parentheses and sorted in alphabetical order, and the nodes are separated by ";".
[0067] The process of log parsing is divided into the following steps:
[0068] 1. Use the BFS algorithm to parse the request log and record the nodes in alphabetical order as a unique identifier.
[0069] 2. Record the number of times the unique identifier is accessed.
[0070] 3. Record the number of times each node is accessed respectively.
[0071] Step S103: Obtain the call relationships between various services through the said request path;
[0072] According to the said request path, obtain the intersecting nodes between various services; through the intersecting nodes, obtain the call relationships between various services.
[0073] Step S104: Allocate elastic resources to the services under the condition of meeting the basic queries per second rate condition of the services according to the remaining resources of the server and the said call relationships.
[0074] Elastic resource scheduling is divided into two cases, namely elastic resource allocation when resources are sufficient and elastic resource allocation when resources are insufficient. Additionally, it should be noted that when the service goes online, the maximum QPS (queries per second) that each node can withstand is tested, which is called the basic QPS of the service.
[0075] When the remaining resources of the server are sufficient, calculate the remaining resources of the server and the recyclable resources; recycle the excess service resources of the server; allocate the excess service resources to each node of the service waiting for resource allocation and the nodes of the services that have call relationships with the service waiting for resource allocation, and the excess service resources are greater than or equal to the service resources required for half of the basic queries per second rate QPS of its node;
[0076] When the remaining resources of the server are insufficient, recycle the resources of the off-peak services; calculate the proportion of the resources required by the peak services in the total resources required by the peak services; obtain the number of nodes in the request path of the peak services waiting for resource allocation, and allocate the remaining resources according to the proportion of the number of machines required by the nodes in the total number of machines required. For example, the resources of 50 off-peak services are recycled, and there are 50 remaining resources of the current remaining resources, with a total of 100 remaining resources of the server. Now there are three services, A, B, and C, that need to allocate resources. According to the node data of the services, it is determined that A needs 300 machines, B needs 150 machines, and C needs 150 machines. The existing 100 servers definitely cannot meet the requirements of the three services for servers. Therefore, according to the demand ratio of the three services for servers, which is 2:1:1, the remaining resources are allocated. Then A is allocated 50 machines, B is allocated 25 machines, and C is allocated 25 machines.
[0077] The specific allocation method is that when resources are sufficient, twice the server resources will be allocated to each service, so that the average QPS of each node of the service is maintained at half of the basic QPS of the service. For the stability of the service, when the QPS of the service is about 80% of the basic QPS, the system will allocate resources to the service. When the QPS of the service is lower than 50% of the basic QPS, the system will recycle the excess resources of the service.
[0078] The elastic resource allocation process when resources are sufficient is divided into the following steps:
[0079] 1) Calculate the remaining server resources and recyclable server resources of the service.
[0080] 2) Determine whether the resource allocation is a situation of sufficient resources.
[0081] 3) Recycle the resources of the service with excessive service resources.
[0082] 4) Add each service that needs to allocate resources to the allocation queue.
[0083] 5) Allocate resources to each service in sequence.
[0084] Elastic resource allocation when resources are insufficient is somewhat meaningless. At this time, the entire data center should be expanded. There are two cases of insufficient resources. First, some services in the system are in the peak period and a small part of the services are in the off-peak period. At this time, the service resources in the off-peak period can be allocated to the peak period. Second, the vast majority of the services in the service center are in the peak period. At this time, elastic resource allocation is meaningless. When resources are insufficient, this article defines that the average QPS lower than 80% of the basic QPS is the off-peak period service, and the peak period service is the service with QPS greater than or equal to the basic QPS. The service classification when resources are insufficient is as Figure 5 shown.
[0085] The elastic resource allocation when resources are insufficient is divided into two steps:
[0086] 1) Calculate the remaining service resources of the service and the recyclable server resources;
[0087] 2) Determine whether the resource allocation is a situation of sufficient resources;
[0088] 3) Recycle the resources of the off-peak period service. The recycling strategy is to maintain the average QPS of the off-peak period service at about 80% of the basic QPS, and recycle the remaining resources.
[0089] 4) Calculate the resources required for each service in the peak period service;
[0090] 5) Calculate the proportion of the resources required by each peak service in the total resources required by the peak period services;
[0091] 6) Package the peak period service and the proportion of its allocated resources into a task and add it to the queue;
[0092] 7) Perform elastic resource allocation for the services in the queue.
[0093] To verify the effect of the present invention, the following tests are carried out in two experiments.
[0094] Experiment 1, elastic resource allocation under sufficient resources.
[0095] Experiment 2: Elastic resource allocation under resource shortage.
[0096] In the experiment, the services are built using Django. Each service only includes requests and forwarding, and there is a computing module in the middle. The computing workload of the computing module of each service is different, so that different services can be simulated. The experimental environment is a server with 16G of memory, 8 cores, and 100G of storage. Kubernetes with version number 1.15 and Docker with version number 1.7 are deployed on the server. A total of 10 services are deployed in the experiment, and each service is deployed using Service and Deployment in Kubernetes. The dependency relationships of the 10 services are as Figure 6 shown.
[0097] As Figure 6 can be seen, services 1, 2, and 3 are external services, and the remaining services are basic services. Services 1, 4, 5, and 8 form a request path. Services 2, 5, 6, and 9 form a request path. Services 3, 6, 7, and 10 form a request path. For the convenience of testing, the default QPS of services 1-10 in this paper is set to 50.
[0098] Experiment 1:
[0099] To simulate a real service scenario, this paper gives basic traffic with a QPS of 50 to services 1, 2, and 3 respectively. When the system resources are sufficient, services 1, 2, 3, 4, and 7 each contain two working nodes in the cluster, and services 5, 6, 8, 9, and 10 each contain 4 nodes. The experiment performs stress testing on the request path of service 1 with different QPSs, and records the number of nodes when the services on the request path are stable and the time taken for the services to return to a stable state. The stress testing is carried out on the basis of the basic traffic.
[0100] Table 1 Test results of Experiment 1
[0101]
[0102] As can be seen from the above table, when stress testing the request path where service 1 is located, the service reaches a stable state in only about 2.5S. The number of nodes when services 1, 4, 5, and 8 return to the normal state meets the expectations. Therefore, the scheduling under sufficient resources applied by this institute meets the expectations.
[0103] Experiment 2:
[0104] To simulate elastic resource allocation under real resource shortage conditions, the experiment imposed an upper limit on the total service resource volume, and a maximum of 45 containers could be built in the cluster. As in the previous experiment, the experiment respectively gave the basic traffic with a QPS of 50 to the request paths corresponding to Services 1, 2, and 3. Then, Services 1, 2, 3, 4, and 7 each contained two working nodes in the cluster, and Services 5, 6, 8, 9, and 10 each contained four nodes. The experiment respectively performed stress testing on the request paths corresponding to Services 1, 2, and 3 to ensure that the QPS on each request path was approximately 200. At this time, the resources of the system could not meet all services. The experiment recorded the number of nodes finally allocated to each service and the time taken for the service to recover stability.
[0105] Table 2 Experimental service call examples
[0106]
[0107]
[0108] In the case of resource shortage in the cluster, the resource allocation is to allocate the remaining machines according to the ratio of the number of machines required by each service to the total required number of machines. From the test results in the above table, it can be seen that the node allocation for each service is correct, and it only takes about 4.3S for the service to recover stability. Therefore, the elastic resource allocation under resource shortage also meets the expectations.
[0109] The specific implementation steps are as Figure 7 shown:
[0110] Step 1: When a node is accessed, the node will record the log in the Redis cluster. The log information includes information such as RequestID and parent node information.
[0111] Step 2: Set a scheduled task in the log parsing. Every minute, the log information will be parsed, different request paths will be parsed into a unique identifier, and at the same time, the number of times each node is accessed and the number of times the request path is accessed will be recorded.
[0112] Step 3: Elastic resource scheduling is also a scheduled task. It is executed in 6 sub-steps.
[0113] 1. Calculate the remaining resources of the cluster.
[0114] 2. Calculate the recyclable resources.
[0115] 3. Calculate the resources that need to be allocated to peak services.
[0116] 4. Determine whether it is resource-sufficient allocation or resource-insufficient allocation during resource allocation.
[0117] 5. Recycle the resources of off-peak services.
[0118] Allocate resources for peak-hour services according to different patterns.
[0119] Based on the same inventive concept, this application also provides an elastic resource scheduling system 800 for grid data center services, as Figure 8 shown, including:
[0120] A request log acquisition module 810, configured to acquire request logs of grid data center services;
[0121] A request path acquisition module 820, configured to parse the request logs to acquire request paths of the services;
[0122] A call relationship acquisition module 830, configured to acquire call relationships between services through the request paths;
[0123] A resource allocation module 840, configured to allocate elastic resources for the services under the condition of meeting the basic queries per second condition of the services according to the remaining resources of the servers and the call relationships.
[0124] Preferably, the request log acquisition module includes:
[0125] A parsing sub-module, configured to parse the request logs through a breadth-first search algorithm (BFS) to acquire request paths formed by each request calling each service; and each node included in the request paths;
[0126] An identification sub-module, configured to identify each of the nodes using characters;
[0127] A string composition sub-module, configured to form a string from the nodes passed by each path;
[0128] An identification sub-unit, configured to identify the request paths of the services using the strings.
[0129] Preferably, it further includes:
[0130] An access times acquisition sub-module, configured to acquire the number of times each node is accessed and the number of times the request paths are accessed.
[0131] Preferably, the call relationship acquisition module includes:
[0132] An intersection node acquisition sub-module, configured to acquire intersection nodes between services according to the request paths;
[0133] A call relationship acquisition sub-module, configured to acquire call relationships between services through the intersection nodes.
[0134] Preferably, the resource allocation module includes:
[0135] The first resource calculation sub-module is used to calculate the remaining resources of the server and the recyclable resources when the remaining resources of the server are sufficient;
[0136] The first recycling sub-module is used to recycle the excess service resources of the server;
[0137] The first resource allocation sub-module is used to allocate the excess service resources to each node of the service waiting for resource allocation and the nodes of the service having a call relationship with the service waiting for resource allocation, and the excess service resources are greater than or equal to the service resources required for half of the basic queries per second (QPS) of its node;;
[0138] The second recycling sub-module is used to recycle the resources of the low-peak period service when the remaining resources of the server are insufficient;
[0139] The proportion calculation sub-module is used to calculate the proportion of the resources required by the peak period service in the total resources required by the peak period service;
[0140] The second resource allocation sub-module is used to obtain the number of nodes in the request path of the peak period service to be allocated, and allocate the remaining resources according to the proportion of the number of machines required by the number of nodes in the total number of machines required;
[0141] An elastic resource scheduling method and system for power grid data center services provided by this application can detect the load of the service and the dependency relationship of the service at the same time, and can allocate resources more accurately. The power grid data center saves a lot of hardware costs. It solves the problem that blindly expanding capacity for a certain service will cause unnecessary waste and the resource allocation is not accurate enough.
[0142] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of this application rather than limit the scope of its protection. Although this application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: after reading this application, those skilled in the art can still make various changes, modifications or equivalent replacements to the specific implementation manners of the application. These changes, modifications or equivalent replacements are all within the scope of the claims of this application pending approval.
Claims
1. An elastic resource scheduling method for power grid data center services, characterized in that, Including: Obtain the request logs of the power grid data middle platform service; Parse the request logs to obtain the request paths of the services; Through the request paths, obtain the call relationships between various services; According to the remaining resources of the server and the call relationships, under the condition of meeting the basic queries per second rate of the services, allocate elastic resources for the services, including: When the remaining resources of the server are sufficient, calculate the remaining resources of the server and the recyclable resources; Recycle the excess service resources of the server; Allocate the excess service resources to the nodes of each service waiting for resource allocation and the nodes of the services having call relationships with the services waiting for resource allocation, where the excess service resources are greater than or equal to the service resources required for half of the basic queries per second rate QPS of their nodes; When the remaining resources of the server are insufficient, recycle the resources of the services during the low peak period; Calculate the proportion of the resources required by the services during the peak period in the total resources required by the services during the peak period; Obtain the number of nodes in the request paths of the services during the peak period to be allocated, and allocate the remaining resources according to the proportion of the number of machines required by the nodes in the total number of machines required; 2. The method according to claim 1, wherein Parse the request logs to obtain the request paths of the services, including: Parse the request logs through the breadth-first search algorithm BFS to obtain the request paths formed by each request calling various services; and the various nodes included in the request paths; Use characters to identify the various nodes; Form a string with the nodes passed by each path; Use the string to identify the request paths of the services.
3. The method according to claim 2, wherein Also including: Obtain the number of times each node is accessed and the number of times the request paths are accessed.
4. The method according to claim 1 or 3, characterized in that, Through the request paths, obtain the call relationships between various services, including: According to the request paths, obtain the intersecting nodes between various services; Through the intersecting nodes, obtain the call relationships between various services.
5. An elastic resource scheduling system for power grid data middle platform services, characterized in that, Including: Request log acquisition module, used to obtain the request logs of the power grid data middle platform service; Request path acquisition module, used to parse the request logs to obtain the request paths of the services; Call relationship acquisition module, used to obtain the call relationships between various services through the request paths; Resource allocation module, used to allocate elastic resources for the services according to the remaining resources of the server and the call relationships under the condition of meeting the basic queries per second rate of the services; Resource allocation module, including: First resource calculation sub-module, used to calculate the remaining resources of the server and the recyclable resources when the remaining resources of the server are sufficient; First recycling sub-module, used to recycle the excess service resources of the server; First resource allocation sub-module, used to allocate the excess service resources to the nodes of each service waiting for resource allocation and the nodes of the services having call relationships with the services waiting for resource allocation, where the excess service resources are greater than or equal to the service resources required for half of the basic queries per second rate QPS of their nodes; Second recycling sub-module, used to recycle the resources of the services during the low peak period when the remaining resources of the server are insufficient; Proportion calculation sub-module, used to calculate the proportion of the resources required by the services during the peak period in the total resources required by the services during the peak period; The second resource allocation sub-module is used to obtain the request paths of the peak-period services to be allocated.
6. The system according to claim 5, wherein The request log acquisition module includes: The parsing sub-module is used to parse the request log through the breadth-first search algorithm BFS to obtain the request paths formed by each request calling each service; and each node included in the request path; The identification sub-module is used to identify each of the nodes with characters; The string composition sub-module is used to compose the nodes passed by each path into a string; The identification subunit is used to identify the request path of the service with the string.
7. The system according to claim 6, wherein It further includes: The access times acquisition sub-module is used to obtain the number of times each node is accessed and the number of times the request path is accessed.
8. The system according to claim 5 or 7, characterized in that, The call relationship acquisition module includes: The intersecting node acquisition sub-module is used to obtain the intersecting nodes between each service according to the request path; The call relationship acquisition sub-module is used to obtain the call relationship between each service through the intersecting nodes.
Citation Information
Patent Citations
Micro-service request response method and device, equipment and storage medium
CN113419852A
Task calling method and device, electronic equipment and storage medium
CN113535363A