Node capacity expansion and contraction method and system, computing device and computer readable storage medium
By obtaining the basic and operation and maintenance indicator data of the data processing nodes and combining the node expansion strategy, the automatic elastic expansion and capacity of the Kubernetes platform in long-connected scenarios is achieved, solving the node load management problems that are difficult to achieve in the existing technology, and improving the stability and flexibility of the system.
Patent Information
- Application Number
- CN202311543633.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
The prior art is difficult to achieve automatic elastic scaling of Kubernetes platform in long-connected scenarios, especially when data processing node load needs to be considered.
By obtaining the basic indicator data and operation and maintenance indicator data of the data processing node, and combining the node expansion strategy, dynamically add or offline data processing nodes to achieve automatic expansion and capacity of the node.
It realizes efficient node scaling in long-connected scenarios, avoids service abnormalities caused by overloading of data processing nodes, and effectively reduces the impact of node scaling on projects.
Smart Images

Figure CN120021201A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technologies, and particularly to a method and system for node scaling, a computing device, and a computer-readable storage medium. Background Art
[0002] With the evolution of cloud native, more and more traditional services have migrated from the deployment methods of virtual machines and physical machines to the deployment method of Kubernetes (a container orchestration system) to enjoy the advantages of elastic scaling, high availability, automated scheduling, multi-platform support, etc. brought by Kubernetes.
[0003] However, currently, most services that implement elastic scaling based on the Kubernetes platform are stateless services, that is, short connections are established between the client and the data processing nodes in Kubernetes, and the specific load conditions of each data processing node do not need to be considered. In the case of long connections established between the client and the data processing nodes in Kubernetes, the specific load conditions of each data processing node need to be considered. How to achieve automatic elastic scaling in this case by means of Kubernetes is a major problem that plagues the industry. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a method for node scaling. One or more embodiments of this specification also relate to a node scaling system, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a method for node scaling is provided, including:
[0006] Obtaining the basic metric data and operation and maintenance metric data of the data processing node, where the data processing node is a node that establishes a long connection with the corresponding client;
[0007] Increasing or taking offline data processing nodes according to the basic metric data, the operation and maintenance metric data, and the node scaling policy.
[0008] According to the second aspect of the embodiments of this specification, a node scaling system is provided, including an elastic scaling component, a metric monitoring component, and a node scaling component, where:
[0009] The metric monitoring component is configured to obtain the basic metric data and operation and maintenance metric data of the data processing node, where the data processing node is a node that establishes a long connection with the corresponding client, and send the basic metric data and the operation and maintenance metric data to the elastic scaling component;
[0010] The elastic scaling component is used to determine a node scaling strategy based on the basic metric data and the operation and maintenance metric data, and send the node scaling strategy to the node scaling component;
[0011] The node scaling component is used to add or take offline data processing nodes according to the node scaling strategy.
[0012] According to the third aspect of the embodiments of this specification, a computing device is provided, including:
[0013] A memory and a processor;
[0014] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above node scaling method are implemented.
[0015] According to the fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above node scaling method are implemented.
[0016] According to the fifth aspect of the embodiments of this specification, a computer program is provided. When the computer program is executed on a computer, the computer is made to execute the steps of the above node scaling method.
[0017] The node scaling method provided by an embodiment of this specification obtains the basic metric data and the operation and maintenance metric data of a data processing node, where the data processing node is a node that establishes a long connection with a corresponding client; according to the basic metric data, the operation and maintenance metric data, and the node scaling strategy, add or take offline data processing nodes.
[0018] Specifically, this node scaling method accurately selects a corresponding node scaling strategy according to the obtained basic metric data and operation and maintenance metric data of the data processing node, so as to reasonably add or take offline data processing nodes according to the node scaling strategy, efficiently implement node scaling, avoid service anomalies caused by overloading of data processing nodes, and effectively avoid directly scaling down data processing nodes, and can effectively reduce the impact of node scaling down on the project according to the node scaling strategy. Description of the Drawings
[0019] Figure 1 is a schematic diagram of a scenario of a node scaling method provided by an embodiment of this specification;
[0020] Figure 2 is a flowchart of a node scaling method provided by an embodiment of this specification;
[0021] Figure 3 It is a schematic diagram showing a user establishing a connection with a specific data processing node in a gateway service component provided by an embodiment of this specification;
[0022] Figure 4 It is a schematic diagram showing a user establishing a connection with a newly added data processing node provided by an embodiment of this specification;
[0023] Figure 5 It is a timing diagram showing a user establishing a TCP long connection with a gateway service and corresponding information transmission provided by an embodiment of this specification;
[0024] Figure 6 It is a flowchart of the processing procedure for node expansion in a node scaling method provided by an embodiment of this specification;
[0025] Figure 7 It is a flowchart of the processing procedure for node contraction in a node scaling method provided by an embodiment of this specification;
[0026] Figure 8 It is a schematic structural diagram of a node scaling system provided by an embodiment of this specification;
[0027] Figure 9 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0028] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0029] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0030] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0031] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0032] First, the noun terms involved in one or more embodiments of this specification are explained.
[0033] Kubernetes: A mechanism for managing containerized applications on multiple hosts in a cloud platform. The goal of Kubernetes is to make the deployment of containerized applications simple and efficient. Kubernetes provides a mechanism for application deployment, planning, updating, and maintenance.
[0034] Pod: The smallest unit that can be created and managed in the Kubernetes system. A pod has multiple containers, and one application program runs in one container.
[0035] Elastic scaling: A common method in cloud computing. Through this method, the amount of computing resources in the server pool (usually measured according to the number of available servers) will be dynamically scaled according to the load in the server pool.
[0036] TCP: Transmission Control Protocol, a connection-oriented, reliable, byte-stream-based transport layer communication protocol.
[0037] Long connection: In application services, a long connection usually refers to a persistent connection established between the client and the server. Different from a short connection (i.e., a one-time request-response), a long connection can remain open for a period of time and perform data exchange multiple times. The usage scenarios of long connections include real-time communication, push services, real-time data synchronization, etc.
[0038] HPA: Horizontal Pod Autoscaler, which is a component in Kubernetes for automatically horizontally scaling the number of pod replicas to adjust the capacity of an application automatically according to load requirements.
[0039] API: Application Programming Interface, which refers to some predefined functions. The purpose is to provide applications and developers with the ability to access a set of routines based on a certain software or hardware, without the need to access the source code or understand the details of the internal working mechanism.
[0040] Resource Metrics API: Resource Metrics API. Kubernetes defines a set of standardized API interfaces, Resource Metrics API, to facilitate client applications (such as HPA) to obtain performance data of target resource objects, such as CPU and memory usage data of containers.
[0041] Prometheus: An open-source service monitoring system and time series database. Metric collection is an important part of the Prometheus monitoring system. It is mainly responsible for obtaining various measurement data (such as CPU usage rate, memory occupancy rate, etc.) from target objects (such as containers, hosts, services, etc.) and providing this data to Prometheus for storage and processing.
[0042] Controller Manager: The component in Kubernetes that manages the elastic scaling of the cluster.
[0043] Scale: The component used to actually control the elastic scaling of nodes.
[0044] API Server: The component used to configure the corresponding elastic threshold.
[0045] CPU: The Central Processing Unit (CPU) is the operation and control core of a computer system and the final execution unit for information processing and program operation.
[0046] QPS: Queries-Per-Second, which is a measure of the amount of traffic processed by a specific query server within a specified time. In the embodiments of this specification, it can be understood as traffic.
[0047] Thread: A thread is the smallest unit that an operating system can schedule for computing. It is contained within a process and is the actual operating unit within the process. A thread refers to a single sequential control flow within a process. Multiple threads can run concurrently within a process, and each thread executes different tasks in parallel.
[0048] In this specification, a method for scaling nodes in and out is provided. This specification also relates to a system for scaling nodes in and out, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0049] See Figure 1 , Figure 1 which shows a schematic diagram of a scenario of a method for scaling nodes in and out provided according to an embodiment of this specification.
[0050] Specifically, in the embodiments of this specification, taking the method for scaling nodes in and out applied to a system for scaling nodes in and out for the scenario of scaling nodes in and out of a gateway service for long connections as an example, a detailed description will be given.
[0051] Specifically, the system for scaling nodes in and out includes: a configuration policy component 102, an elastic scaling component 104, a metric monitoring component 106, a node scaling component 108, and a gateway service component 110.
[0052] In the scenario of scaling nodes in and out, the system for scaling nodes in and out can be understood as a Kubernetes cluster. The configuration personnel send the node scaling policy to the elastic scaling component 104 that manages the Kubernetes cluster through the configuration policy component 102. The elastic scaling component 104 that manages the Kubernetes cluster receives the node scaling policy sent by the configuration personnel through the configuration policy component 102 and configures the node scaling policy. The node scaling policy includes a node expansion policy and a node contraction policy. The node expansion policy includes information such as a metric expansion threshold and a node expansion rule, and the node contraction policy includes information such as an operation and maintenance metric contraction threshold and a node contraction rule.
[0053] The metric monitoring component 106 cyclically obtains the basic metric data and operation and maintenance metric data of the data processing nodes at a preset time interval. For example, the preset time interval can be 20 seconds, and the basic metric data and operation and maintenance metric data of the pod nodes are obtained every 20 seconds. Among them, the data processing nodes can be understood as pod nodes, which are used to process data processing requests sent by users through the gateway service component 110. The basic metric data can include metrics such as CPU and memory. The operation and maintenance metric data can be understood as metric data related to the project, such as the query per second rate (QPS, which can also be understood as traffic), the number of threads, and other metric data.
[0054] The metric monitoring component 106 sends the obtained basic metric data and operation and maintenance metric data to the elastic scaling component 104. Then, in the elastic scaling component 104, the basic metric data, the operation and maintenance metric data are compared with the thresholds corresponding to the node scaling policy. Specifically, when the basic metric data or the operation and maintenance metric data is greater than or equal to the metric expansion threshold in the node expansion policy, the number of pod nodes to be expanded is determined according to the node expansion step in the node expansion policy, and the pod nodes are added by the node scaling component 108 according to the number of pod nodes to be expanded.
[0055] When the operation and maintenance metric data is less than the operation and maintenance metric contraction threshold in the node contraction policy, the newly added pod nodes added within a preset time period are obtained. For example, the pod nodes added within 30 minutes are obtained, and these pod nodes are sorted according to data such as traffic and the number of threads. Then, the number of pod nodes to be contracted is determined according to the node contraction step, so that the node scaling component 108 determines and takes offline the pod nodes to be contracted from the sorted newly added pod nodes according to the number of pod nodes to be contracted.
[0056] Alternatively, the intersection of two sets can be used to determine and take offline the pod nodes to be contracted. Specifically, the set of pod nodes to be contracted determined in this loop is used as the current set of pod nodes to be contracted, and it is judged whether there is a historical set of pod nodes to be contracted. The determination methods of the historical set of pod nodes to be contracted and the current set of pod nodes to be contracted are the same. If there is a historical set of pod nodes to be contracted, the intersection of the historical set of pod nodes to be contracted and the current set of pod nodes to be contracted is taken, and the node scaling component 108 takes offline the pod nodes to be contracted in this intersection. If there is no historical set of pod nodes to be contracted, the current set of pod nodes to be contracted obtained in this loop is used as the historical set of pod nodes to be contracted, and the intersection is taken with the current set of pod nodes to be contracted obtained in the next loop to complete the determination of the pod nodes to be contracted.
[0057] After the pod nodes to be contracted are determined, if there are still long connections between the pod nodes to be contracted and the customer, the steps of disconnecting the long connection and four-way handshake are executed to ensure that the long connection corresponding to the pod nodes to be contracted is completely disconnected. Then, the node scaling component 108 takes offline the pod nodes to be contracted, effectively avoiding the bad experience brought to users due to the sudden interruption of the long connection.
[0058] After the expansion or contraction of the pod nodes is completed, the gateway service component 110 receives a data processing request sent by a user, determines a target pod node according to the minimum connection number policy, and establishes a long connection between the user and the target pod node. Among them, based on the minimum connection number policy, it can be ensured that when the pod nodes are expanded, the client corresponding to the user establishes a long connection with the newly added pod nodes, avoiding overloading the previous pod nodes.
[0059] Based on this, the node expansion and contraction method accurately selects the corresponding node expansion and contraction strategy according to the obtained basic index data and operation and maintenance index data of the data processing nodes, and reasonably adds or takes offline the data processing nodes according to the node expansion and contraction strategy, efficiently realizing node expansion and contraction, avoiding service exceptions caused by overloading of data processing nodes, and effectively reducing the impact of node contraction on the project according to the node expansion and contraction strategy.
[0060] See Figure 2 , Figure 2 shows a flowchart of a node expansion and contraction method provided by an embodiment of this specification, which specifically includes the following steps.
[0061] Step 202: Obtain the basic index data and operation and maintenance index data of the data processing nodes, where the data processing nodes are the nodes that establish long connections with the corresponding clients.
[0062] Among them, the data processing nodes can be understood as the pod nodes in the above embodiments, which are used to process the data processing requests of the clients; the basic index data can be understood as the index data of the hardware related to the data processing nodes, such as CPU and memory data, and the operation and maintenance index data can be understood as the index data describing the load situation of the data processing nodes, such as the query rate per second and the number of threads, where the number of threads can also be understood as the number of connections between a pod node and a user.
[0063] Taking the node expansion and contraction scenario of the stateful gateway service applied to long connections as an example, combined with the elastic scaling ability of Kubernetes itself, this node expansion and contraction method will be described in detail.
[0064] Specifically, the pod nodes can be understood as the specific pod nodes in the Kubernetes cluster bound under the gateway service; obtain the CPU utilization rate, memory utilization rate, traffic, number of threads, interface response success rate, etc. of the pod nodes.
[0065] In one or more embodiments, there are multiple pod nodes in the cluster. The obtained basic index data refers to the average value of the basic monitoring data of the pod nodes, and the operation and maintenance index data refers to the operation and maintenance index data of a single pod node. The specific implementation method is as follows:
[0066] Obtaining the basic metric data and operation and maintenance metric data of the data processing node includes:
[0067] Obtaining the basic monitoring data of the data processing node, and using the average value of the basic monitoring data as the basic metric data, and obtaining the operation and maintenance metric data of each data processing node.
[0068] Among them, the basic monitoring data can be understood as the CPU and memory data of the pod node.
[0069] Specifically, the metric monitoring component obtains the CPU and memory data of the pod node, calculates the average value of the sum of the CPU and memory data of the pod node, uses this average value as the basic metric data, and obtains the operation and maintenance metric data of each pod node. In practical applications, the CPU and memory data are used as the basic metric data to determine the load condition of the pod node. In the case of determining that the pod node is overloaded, the expansion operation can be directly performed; in the long connection scenario, the traffic and the number of threads related to the actual project are used as the operation and maintenance metric data, which is more referenceable.
[0070] Among them, the metric monitoring component can be Resource Metrics API, Prometheus, which is not limited here.
[0071] For example, currently there are 3 pod nodes that establish long connections with the client. Through the Resource Metrics API, it is obtained that the CPU and memory occupancy rates in pod node 1 are 50%, the CPU and memory occupancy rates in pod node 2 are 40%, and the CPU and memory occupancy rates in pod node 3 are 60%. Therefore, the average value of the CPU and memory occupancy rates of these 3 pod nodes is 50%, that is, the basic metric data is 50%.
[0072] The node scaling method provided by the embodiments of this specification accurately obtains the basic metric data by calculating the average value of the basic monitoring data of the data processing node, and obtains the operation and maintenance metric data of each pod node, so as to reasonably increase or take offline the data processing node when the operation and maintenance metric data of each subsequent pod node all exceed or are less than the operation and maintenance metric threshold.
[0073] In one or more embodiments, before obtaining the basic metric data and the operation and maintenance metric data, it is also necessary to receive the node scaling policy set by the configuration personnel, so as to reasonably perform node expansion or node scaling according to the node scaling policy after obtaining the basic metric data and the operation and maintenance metric data. The specific implementation method is as follows:
[0074] Before obtaining the basic metric data and operation and maintenance metric data of the data processing node, the following steps are further included:
[0075] Receive the node scaling policy sent by the configuration personnel, and configure the node scaling policy for the data processing node.
[0076] Among them, the node scaling policy includes a node expansion policy and a node contraction policy, and the node expansion policy includes a node expansion rule, and the node contraction policy includes a node contraction rule.
[0077] Specifically, when the node scaling method is applied to a node scaling system based on the Kubernetes cluster, the elastic scaling component receives the node scaling policy sent by the configuration personnel through the configuration policy component, and configures the node scaling policy for the pod node; among them, the elastic scaling component can be understood as the HPA in the ControllerManager in the Kubernetes cluster, and the configuration policy component can be understood as the API Server.
[0078] After the HPA obtains the basic metric data and operation and maintenance metric data of the pod node sent by the Resource Metrics API, it compares the obtained basic metric data and operation and maintenance metric data with the basic metric expansion threshold or operation and maintenance metric contraction threshold sent by the API Server, so as to determine whether to expand according to the node expansion policy or to contract according to the node contraction policy.
[0079] The node scaling method provided by the embodiments of this specification can customize the node scaling policy according to the actual situation when receiving and configuring the node scaling policy, and can flexibly select the corresponding node expansion policy according to the obtained basic metric data and operation and maintenance metric data to efficiently complete the expansion operation, or efficiently complete the contraction operation according to the node contraction policy.
[0080] In one or more embodiments, the basic metric data and operation and maintenance metric data of the data processing node can be cyclically obtained at intervals of a period of time, so as to be able to continuously perform expansion or contraction operations according to the actual situation of the data processing node. The specific implementation method is as follows:
[0081] The obtaining of the basic metric data and operation and maintenance metric data of the data processing node includes:
[0082] Cyclically obtain the basic metric data and operation and maintenance metric data of the data processing node according to a preset time interval.
[0083] Among them, the preset time interval can be understood as a pre-set time interval, such as 20 seconds, 30 seconds.
[0084] Specifically, taking a preset time interval of 30 seconds as an example for detailed description, the Resource Metrics API can obtain the basic metric data and operation and maintenance metric data of the data processing node every 30 seconds.
[0085] The node scaling method provided by the embodiments of this specification cyclically obtains the basic metric data and operation and maintenance metric data according to the preset time interval, and then performs node scaling processing, rather than real-time processing, which can achieve the effects of reducing the workload of the metric monitoring component and avoiding resource occupation.
[0086] Step 204: Add or take offline data processing nodes according to the basic metric data, the operation and maintenance metric data, and the node scaling policy.
[0087] Specifically, after the HPA obtains the basic metric data and operation and maintenance metric data, it compares the obtained basic metric data and operation and maintenance metric data with the basic metric expansion threshold or operation and maintenance metric contraction threshold sent by the API Server to determine whether to expand according to the node expansion policy or contract according to the node contraction policy. After determining the node scaling policy, the Scale component adds data processing nodes according to the node expansion policy or takes offline data processing nodes according to the node contraction policy.
[0088] In one or more embodiments, the basic metric data, or the operation and maintenance metric data is compared with the metric expansion threshold or the operation and maintenance metric contraction threshold, so that in the case of needing to expand, the data processing nodes are reasonably added according to the node expansion step, or in the case of needing to contract, the data processing nodes are taken offline according to the node contraction rule. The specific implementation method is as follows:
[0089] The node scaling policy includes a node expansion policy and a node contraction policy, and the node expansion policy includes a node expansion rule, and the node contraction policy includes a node contraction rule;
[0090] The adding or taking offline data processing nodes according to the basic metric data, the operation and maintenance metric data, and the node scaling policy includes:
[0091] In the case where it is determined that the basic metric data, or the operation and maintenance metric data is greater than or equal to the metric expansion threshold, determine the number of data processing nodes to be expanded according to the node expansion rule, and add data processing nodes according to the number of data processing nodes to be expanded;
[0092] Or
[0093] When it is determined that the operation and maintenance metric data is less than the operation and maintenance metric scaling-down threshold, determine the data processing nodes to be scaled down according to the data processing nodes added within a preset time period and the node scaling-down rule, and take offline the data processing nodes to be scaled down.
[0094] Among them, the node scaling-up rule includes determining the number of data processing nodes to be scaled up according to a preset node scaling-up step size, or determining the number of data processing nodes to be scaled up according to the data processing requests of the client, the basic metric data, and the operation and maintenance metric data.
[0095] The node scaling-down rule includes determining the number of data processing nodes to be scaled down according to a preset node scaling-down step size, or determining the number of data processing nodes to be scaled down according to the basic metric data and the operation and maintenance metric data of the data processing nodes added within the preset time period.
[0096] Based on obtaining the basic metric data and operation and maintenance metric data of the data processing nodes, the metric scaling-up threshold includes the basic metric scaling-up threshold corresponding to the basic metric data and the operation and maintenance metric scaling-up threshold corresponding to the operation and maintenance metric data; the operation and maintenance metric scaling-down threshold can be understood as the operation and maintenance metric scaling-down threshold corresponding to the operation and maintenance metric data.
[0097] Moreover, when the basic metric data includes CPU and memory data, the basic metric scaling-up threshold includes the CPU scaling-up threshold corresponding to the CPU and the memory scaling-up threshold corresponding to the memory data; when the operation and maintenance metric data includes traffic and the number of connections, the operation and maintenance metric scaling-up threshold includes the traffic scaling-up threshold corresponding to the traffic and the connection number scaling-up threshold corresponding to the number of connections; the operation and maintenance metric scaling-down threshold includes the traffic scaling-down threshold corresponding to the traffic and the connection number scaling-down threshold corresponding to the number of connections.
[0098] For example, the basic metric scaling-up threshold can be set to 70% for CPU and memory utilization; when there can be 500 connections in a pod node, the connection number scaling-up threshold in the operation and maintenance metric scaling-up threshold can be set to 400 connections, and the connection number scaling-down threshold can be set to 100 connections; the traffic scaling-up threshold in the operation and maintenance metric scaling-up threshold can be set to 200G (gigabytes), and the traffic scaling-down threshold can be set to 20G.
[0099] The node scaling-up step size can be understood as the number of data processing nodes that can be added in one scaling-up, such as it can be 1 or 2, and is not limited here; the node scaling-down step size can be understood as the number of data processing nodes that can be taken offline in one scaling-down, such as it can be 1 or 2, and is not limited here.
[0100] The data processing node to be expanded can be understood as the data processing node to be added during the expansion operation; the data processing node to be scaled down can be understood as the data processing node to be taken offline during the scaling-down operation.
[0101] The preset time period can be understood as a pre-set time period. For example, if the preset time period is 20 minutes, the data processing nodes added within 20 minutes are obtained.
[0102] For example, when the node expansion step size is 1, when it is determined that the CPU and memory utilization rate of the pod node exceeds 70%, or the connection number of the pod node exceeds 400, or the traffic of the pod node is greater than 200G, according to the node expansion step size, the number of pod nodes to be expanded can be determined to be 1, so as to add 1 pod node.
[0103] Or it is also possible to determine the data processing node to be expanded according to the connection number, traffic, and specific conditions such as the basic index data and operation and maintenance index data of the pod node carried in the data processing request of the client.
[0104] For example, when the connection number carried in the data processing request of the client is 700 and the traffic is 30G, the number of pod nodes to be expanded is determined to be 2, so as to add 2 pod nodes.
[0105] For the case of scaling down the data processing node, when it is determined that the connection number of the pod node is less than 400 and the traffic is less than 20G, according to the pod nodes added within the preset time period and the node scaling-down step size, the data processing node to be scaled down is determined and taken offline.
[0106] The node expansion and scaling method provided by the embodiments of this specification compares the basic index data or the operation and maintenance index data with the corresponding basic index expansion threshold or operation and maintenance index scaling threshold, and then performs expansion and scaling according to the node expansion and scaling step size; the specific index expansion threshold and node expansion and scaling rules can be custom-set according to actual needs, so as to reasonably perform expansion or scaling.
[0107] In one or more embodiments, when the index expansion threshold includes the basic index expansion threshold and the operation and maintenance index expansion threshold, the basic index data or the operation and maintenance index data is compared with its corresponding expansion threshold, so as to quickly and accurately determine whether to perform expansion. The specific implementation method is as follows:
[0108] Determining that the basic index data or the operation and maintenance index data is greater than or equal to the index expansion threshold includes:
[0109] Determining that the basic index data is greater than or equal to the basic index expansion threshold;
[0110] Or
[0111] Determine that the basic metric data is less than the basic metric expansion threshold, but the operation and maintenance metric data is greater than or equal to the operation and maintenance metric expansion threshold.
[0112] Specifically, when the basic metric data is greater than or equal to the basic metric expansion threshold, it is considered that the CPU and memory are overloaded, and there is no need to compare the operation and maintenance metric data with the operation and maintenance metric expansion threshold to determine whether the basic metric data or the operation and maintenance metric data is greater than or equal to the metric expansion threshold.
[0113] Or when the basic metric data is less than the basic metric expansion threshold, the CPU and memory are not overloaded, and it is necessary to compare the operation and maintenance metric data with the operation and maintenance metric expansion threshold. When it is determined that the operation and maintenance metric data is greater than or equal to the operation and maintenance metric expansion threshold, determine that the basic metric data or the operation and maintenance metric data is greater than or equal to the metric expansion threshold.
[0114] The node scaling method provided by the embodiments of this specification can quickly and accurately determine that the basic metric data or the operation and maintenance metric data is greater than or equal to the metric expansion threshold when the basic metric data is greater than or equal to the basic metric expansion threshold, thus accelerating the expansion processing flow.
[0115] In one or more embodiments, when it is determined that pod node scaling down is to be performed, the added pod nodes within a preset time are sorted according to the operation and maintenance metric data, and then the pod nodes to be scaled down are determined from the sorted newly added data processing nodes according to the node scaling down rule. The specific implementation method is as follows:
[0116] Determining the data processing node to be scaled down according to the added data processing nodes within a preset time period and the node scaling down rule includes:
[0117] Obtain the newly added data processing nodes added within a preset time period, and sort the newly added data processing nodes according to the operation and maintenance metric data of the newly added data processing nodes;
[0118] Determine the data processing node to be scaled down according to the node scaling down rule and the sorted newly added data processing nodes.
[0119] Specifically, obtain the newly added pod nodes added within a preset time period, and sort the newly added data processing nodes in ascending order according to data such as the connection number and traffic of the newly added data processing nodes, and determine the data processing node to be scaled down from the sorted newly added data processing nodes according to the node scaling down step.
[0120] For example, the preset time period is 20 minutes. Obtain the newly added pod nodes within 20 minutes: pod1, pod2, pod3. Among them, the number of connections of pod1 is 70, the number of connections of pod2 is 70, and the number of connections of pod3 is 80; the traffic of pod1 is 15G, the traffic of pod2 is 10G, and the traffic of pod3 is 10G; thus, the sorting of the pod nodes from small to large is pod2, pod1, pod3; when the node scaling-down step size is 2, it can be determined that the pod nodes to be scaled down are pod2 and pod1.
[0121] The node scaling method provided by the embodiments of this specification can more efficiently and accurately determine the data processing nodes to be scaled down by obtaining the newly added data processing nodes within the preset time period and based on the operation and maintenance metric data and the node scaling-down step size.
[0122] In one or more embodiments, the number of data processing nodes to be scaled down in one scaling-down operation can be determined according to the node scaling rule, so that when performing node scaling down, the data processing nodes to be scaled down are determined according to the number of data processing nodes to be scaled down. The specific implementation method is as follows:
[0123] Determining the data processing nodes to be scaled down according to the node scaling rule and the sorted newly added data processing nodes includes:
[0124] Determine the number of data processing nodes to be scaled down according to the node scaling rule, and determine the data processing nodes to be scaled down from the sorted newly added data processing nodes according to the number of data processing nodes to be scaled down.
[0125] Among them, the node scaling rule includes determining the number of data processing nodes to be scaled down according to a preset node scaling-down step size, or determining the number of data processing nodes to be scaled down according to the basic metric data and the operation and maintenance metric data of the data processing nodes added within the preset time period.
[0126] Specifically, the number of pod nodes to be scaled down can be determined according to the set node scaling-down step size. For example, when the node scaling-down step size is set to 1, determine that the number of pod nodes to be scaled down is 1, and determine the first pod node from the sorted newly added pod nodes in ascending order as the pod node to be scaled down; when the node scaling-down step size is set to 2, determine that the number of pod nodes to be scaled down is 2, and determine the first 2 pod nodes from the sorted newly added pod nodes in ascending order as the pod nodes to be scaled down.
[0127] Alternatively, the number of pod nodes to be scaled down can also be determined according to the basic metric data of the data processing nodes and the operation and maintenance metric data of the newly added pod nodes within the preset time.
[0128] For example, calculate the average value of the operation and maintenance metric data of the newly added pod nodes within a preset time, count the number of newly added pod nodes lower than the average value, and determine the number of data processing nodes to be scaled down as this number.
[0129] The node scaling method provided by the embodiments of this specification can determine the number of data processing nodes to be scaled down according to the node scaling step, avoiding the problem of low efficiency where only one node can be scaled down one by one in case of urgent need for scaling down. That is, it can more flexibly determine the data processing nodes to be scaled down from the sorted newly added data processing nodes according to the set node scaling step.
[0130] In one or more embodiments, to ensure that the data processing nodes to be scaled down will not be applied within a certain period of time later, the data processing nodes to be scaled down obtained in one cycle are stored in the candidate set, and the intersection of the double candidate sets is selected to determine the data processing nodes to be scaled down. The specific implementation method is as follows:
[0131] Determining the data processing nodes to be scaled down according to the node scaling rule and the sorted newly added data processing nodes includes:
[0132] Determine the number of data processing nodes to be scaled down according to the node scaling rule, and determine the current set of data processing nodes to be scaled down from the sorted newly added data processing nodes according to the number of the data processing nodes to be scaled down;
[0133] Judge whether there is a set of historical data processing nodes to be scaled down,
[0134] If so, determine the data processing nodes to be scaled down according to the current set of data processing nodes to be scaled down and the set of historical data processing nodes to be scaled down,
[0135] If not, determine the current set of data processing nodes to be scaled down as the set of historical data processing nodes to be scaled down, and continue to execute the step of determining the data processing nodes to be scaled down according to the data processing nodes added within the preset time period and the node scaling rule and taking offline the data processing nodes to be scaled down in the case where it is determined that the operation and maintenance metric data is less than the operation and maintenance metric scaling threshold.
[0136] Among them, the current set of data processing nodes to be scaled down can be understood as the candidate set of the data processing nodes to be scaled down obtained in this detection cycle; the set of historical data processing nodes to be scaled down can be understood as the candidate set of the data processing nodes to be scaled down obtained in the previous detection cycle.
[0137] Specifically, following the above example, it can be obtained that the set S1 of data processing nodes to be scaled down currently includes pod1 and pod2. Determine whether there is a set of historical data processing nodes to be scaled down. If there is, for example, the set S2 of historical data processing nodes to be scaled down includes pod2 and pod3. Then, take the intersection S of the current set S1 of data processing nodes to be scaled down and the set S2 of historical data processing nodes to be scaled down. Among them, the intersection S includes pod2, and pod2 is determined as the data processing node to be scaled down.
[0138] If there is no set of historical data processing nodes to be scaled down, then use the current set S1 of data processing nodes to be scaled down as the set of historical data processing nodes to be scaled down, and wait to take the intersection with the current set S3 of data processing nodes to be scaled down obtained in the next detection cycle to determine the data processing node to be scaled down.
[0139] The node scaling method provided by the embodiments of this specification can ensure that the determined data processing node to be scaled down will not be used within a certain period of time in the future by taking the intersection of the current set of data processing nodes to be scaled down and the set of historical data processing nodes to be scaled down, and the determined data processing node to be scaled down is more accurate and reasonable.
[0140] In one or more embodiments, after adding or taking offline data processing nodes, when receiving a data processing request sent by a client, the target data processing node can be determined according to the minimum connection number policy, and a long connection between the client and the target data processing node is established. The specific implementation method is as follows:
[0141] After adding or taking offline data processing nodes according to the basic metric data, the operation and maintenance metric data, and the node scaling policy, it further includes:
[0142] Receive a data processing request sent by a client, determine the target data processing node according to the minimum connection number policy, and establish a long connection between the client and the target data processing node.
[0143] Among them, the minimum connection number policy can be understood as distributing the data processing requests of the client to the data processing node with the fewest connections to the client among the current data processing nodes; the target data processing node can be understood as the data processing node with the fewest connections to the user among the current data processing nodes; the long connection can be understood as a data path that can continuously send and receive messages after being established based on one or more protocols such as TCP / UDP / QUIC / WebSocket.
[0144] Specifically, users at the application layer establish a long connection with specific pod nodes in the cluster bound under the gateway service component through the gateway service component. For details, please refer to Figure 3 , Figure 3It shows a schematic diagram of a user establishing a connection with a specific data processing node in a gateway service component; among them, Figure 3 The implementation part in it is the real network request address. User 1, User 2, and User 3 establish connections with the pod nodes through the gateway service component; the dotted part is the one-to-one correspondence between the user and the specific pod node. For example, User 1 corresponds to pod1, and User 2 and User 3 correspond to pod2. It can be seen here that on the same long connection, the binding relationship between the user and the pod node is determined.
[0145] In actual applications, there is no strong binding relationship between the user and the pod node. It can be one-to-one, one-to-many, or many-to-one. However, once the connection between the user and a certain pod node is established, in the case of a long connection, the relationship between the user and the pod node is bound.
[0146] After receiving the data processing request sent by the client corresponding to the user, determine the data processing node with the least number of connections to the client among the current data processing nodes according to the minimum connection number policy in the gateway service component, and establish a long connection between the client and this data processing node.
[0147] Specifically in implementation, in the case of adding a data processing node, the connection number of the newly added data processing node is 0. When receiving the data processing request of the client, according to the minimum connection number policy, it can be ensured that the client will establish a long connection with the newly added data processing node.
[0148] Specifically, reference can be made to Figure 4 , Figure 4 It shows a schematic diagram of a user establishing a connection with the newly added data processing node; among them, pod1 and pod2 are the historically existing pod nodes, podN is the newly added pod node, User 1 and User 2 are the historically existing users, and User 3 is a new user. In the case where User 3 sends a data processing request, the gateway service component makes a specified request allocation according to the minimum connection number policy and establishes a long connection between User 3 and podN; if there is subsequent node expansion and podN + 1 is added, the new User 4 will connect to the newly added podN + 1.
[0149] The node scaling method provided by the embodiments of this specification, according to the minimum connection number policy, can ensure that in the case of adding a data processing node, when receiving the user data processing request, this user will establish a long connection with the newly added data processing node, so that the newly added data processing node processes the user's data processing request, avoiding overloading the previous data processing nodes.
[0150] In one or more embodiments, a TCP long connection between a user and a target data processing node based on the CMPP (China Mobile Peer to Peer) protocol or the SMPP (Short Message Mobile Terminated) protocol can be established through a three-way handshake connection establishment method. The specific implementation method is as follows:
[0151] Establishing the long connection between the user and the target data processing node includes:
[0152] Establishing a long connection between the user and the target data processing node based on a preset protocol through a three-way handshake connection establishment method, where the preset protocol includes the CMPP protocol or the SMPP protocol.
[0153] Specifically, a TCP long connection between a user and a target data processing node based on the CMPP protocol or the SMPP protocol is established through a TCP three-way handshake connection establishment method; specifically, refer to Figure 5 , Figure 5 which shows a timing diagram of establishing a TCP long connection between a user and a gateway service and the corresponding information transfer.
[0154] Step 1: TCP three-way handshake.
[0155] The user and the gateway service establish a connection through a TCP three-way handshake.
[0156] Step 1.1: Corresponding response.
[0157] The gateway service sends a corresponding response to the user.
[0158] Step 2: Establish a TCP long connection.
[0159] Create a stable TCP long connection.
[0160] At this point, the long connection between the user and the gateway service is established. The following is the process of information transfer, where the information transmitted in the CMPP or SMPP protocol is the information for short message delivery.
[0161] Step 3: Login verification.
[0162] This step is the user connection establishment stage. The user sends the username and password to the gateway service to complete the verification of the user's identity.
[0163] Step 3.1: Result response.
[0164] The gateway service verifies the user's identity and sends a response result to the user.
[0165] Step 4: Short message sending.
[0166] This step is the user submission stage. The user sends the SMS information to the gateway service through a TCP long connection. The information includes the client's seq Id (network serial number), mobile phone number, content, etc.
[0167] Step 4.1: SMS sending response.
[0168] This step is the submission response stage. The gateway service makes a corresponding response (response failure or response success) to the SMS request submitted by the user's corresponding client. The response includes the client's seq Id (network serial number), the seq Id generated by the gateway service, and the response status, etc.
[0169] Step 5: SMS receipt.
[0170] This step is the receipt sending stage. The gateway service sends the SMS receipt to the user through a TCP long connection. Specifically, the gateway service sends a receipt report on the final sending result of the SMS submitted by the user, including the mobile phone number, msgId (message identifier), and receipt status, etc.
[0171] Step 5.1: Receipt confirmation.
[0172] This step is for the user to send a confirmation for the request of the SMS receipt status sent by the gateway service, including msgId, etc.
[0173] Step 6: Disconnection of the TCP long connection.
[0174] Step 7: Four-way TCP handshake.
[0175] The disconnection of the TCP long connection is completed through Step 6 and Step 7.
[0176] When scaling down the above-mentioned embodiment, that is, taking offline the data processing node to be scaled down, if there are still a small number of CMPP or SMPP long connections between the data processing node to be scaled down and the user, the steps of TCP long connection disconnection and four-way TCP handshake will be executed on the stable TCP long connection to ensure that the TCP long connection is completely disconnected, and then the scaling down of the data processing node to be scaled down is completed.
[0177] The node scaling method provided by the embodiments of this specification establishes a TCP long connection between the user and the target data processing node based on the CMPP protocol or SMPP protocol through a connection establishment method of three-way handshake, which can save resources and prevent connection chaos problems caused by old and repeated connections at the same time.
[0178] Based on this, the node scaling method provided by an embodiment of this specification accurately selects the corresponding node scaling strategy according to the obtained basic index data and operation and maintenance index data of the data processing node, so as to reasonably add or take offline the data processing node according to the node scaling strategy, efficiently implement node scaling, avoid service anomalies caused by overloading of the data processing node, and effectively avoid directly scaling down the data processing node, thereby avoiding a bad experience for users caused by the sudden interruption of long connections.
[0179] See Figure 6 , Figure 6 shows the process flow chart of node expansion in a node scaling method provided by an embodiment of this specification.
[0180] Step 602: Loop check.
[0181] The controller of Kubernetes (such as HPA in the above embodiment) triggers an interval loop check based on the configuration rules. It can be configured to check once every 30 seconds or once every minute, which is not limited here.
[0182] Step 604: Check CPU and memory data.
[0183] Specifically, the CPU and memory data are collected by the metric monitoring component, that is, the basic metric data in the above embodiment, and the metric monitoring component can send the CPU and memory data to HPA, and HPA compares the CPU and memory data with the basic metric expansion threshold.
[0184] Step 606: Determine whether the CPU and memory data are overloaded. If so, execute step 612; if not, execute step 608.
[0185] Specifically, after HPA compares the CPU and memory data with the basic metric expansion threshold, when the CPU and memory data are less than the basic metric expansion threshold, it is determined that the CPU and memory data are not overloaded, and step 608 is executed; when the CPU and memory data are greater than or equal to the basic metric expansion threshold, it is determined that the CPU and memory data are overloaded, and step 612 is executed.
[0186] Step 608: Check the queries per second rate and the number of threads of a single pod node.
[0187] Similarly, the query rate per second and the number of threads of the pod node are collected by the metric monitoring component, that is, the operation and maintenance metric data in the above embodiments, and the metric monitoring component can send the query rate per second and the number of threads of the pod node to the HPA, and the HPA compares the query rate per second and the number of threads of the pod node with the operation and maintenance metric expansion threshold; and because the user accesses the gateway service component through the load balancer, the query rate per second and the number of threads of each pod node are basically the same, so only a single pod node needs to be checked.
[0188] Step 610: Determine whether the query rate per second and the number of threads of the pod node have reached the expansion threshold. If so, execute Step 612; if not, do nothing.
[0189] Among them, the expansion threshold can be understood as the operation and maintenance metric expansion threshold in the above embodiments.
[0190] After the HPA compares the query rate per second and the number of threads of the pod node with the operation and maintenance metric expansion threshold, when the query rate per second and the number of threads of the pod node are greater than or equal to the operation and maintenance metric expansion threshold, it is determined that the expansion condition is met, and Step 612 is executed.
[0191] Step 612: Expand.
[0192] Specifically, the pod node to be expanded can be determined according to the node expansion step size in the above embodiments, and the pod node is added through Scale to complete the expansion operation. For details, please refer to the above embodiments and will not be elaborated here.
[0193] Step 614: Update the gateway load candidate, and execute Step 602.
[0194] Specifically, the newly added pod node is updated to the gateway load candidate, so that after receiving the user's data processing request, the newly added pod node can be selected from the gateway load candidate for connection according to the least connection number policy.
[0195] The node expansion method provided by the embodiments of this specification expands according to the basic metric data and the operation and maintenance metric data, avoiding service anomalies caused by too many connections of the current data processing nodes; at the same time, based on the least connection number policy for specified request allocation, it can control the user to access the newly expanded pod node directionally, avoiding overloading the pod nodes before expansion. Based on this, minute-level expansion can be achieved to smoothly handle traffic peaks.
[0196] See Figure 7 , Figure 7 shows the process flow chart of node contraction in a node expansion and contraction method provided by an embodiment of this specification.
[0197] Step 702: Loop check.
[0198] The Kubernetes controller (such as the HPA in the above embodiments) triggers an interval loop check based on the configuration rules. The loop check can be configured to occur once every 30 seconds or once every minute, and there is no limitation here.
[0199] Step 704: Determine whether the operation and maintenance metric data has reached the scaling-down threshold. If so, execute Step 706; if not, execute Step 702.
[0200] Among them, the scaling-down threshold can be understood as the operation and maintenance metric scaling-down threshold in the above embodiments.
[0201] Specifically, the metric monitoring component collects the traffic (i.e., queries per second) and the number of connections (i.e., the number of threads) of the pod node, which are the operation and maintenance metric data in the above embodiments. And the metric monitoring component can send the traffic and the number of connections of the pod node to the HPA, and the HPA compares the traffic and the number of connections of the pod node with the operation and maintenance metric scaling-down threshold.
[0202] After that, in the case where the traffic and the number of connections of the pod node are less than or equal to the operation and maintenance metric scaling-down threshold, it is determined that the traffic and the number of connections of the pod node have reached the scaling-down threshold, and Step 706 is executed; in the case where the traffic and the number of connections of the pod node are greater than the operation and maintenance metric scaling-down threshold, it is determined that the scaling-down threshold has not been reached, and Step 702 is executed.
[0203] Step 706: Obtain the newly scaled-up pod nodes within a preset time period.
[0204] Among them, the newly scaled-up pod nodes can be understood as the newly added pod nodes in the above embodiments. For the specific implementation, reference can be made to the above embodiments and will not be elaborated here.
[0205] Step 708: Sort according to the number of connections.
[0206] The newly scaled-up pod nodes obtained above can be sorted in ascending order according to the number of connections.
[0207] Step 710: Sort according to the traffic.
[0208] For the newly scaled-up pod nodes with the same number of connections, they are further sorted in ascending order according to the traffic.
[0209] Step 712: Select the first N pod nodes and store them in the current candidate set.
[0210] Among them, N can be understood as the node scaling-down step size set in the above embodiments; the current candidate set can be understood as the current set of pod nodes to be scaled down in the above embodiments.
[0211] Select the first N pod nodes from the newly expanded and sorted pod nodes and store them in the current candidate set. For specific details, refer to the above embodiments and will not be elaborated here.
[0212] Step 714: Determine whether there is a historical candidate set. If so, execute Step 718; if not, execute Step 716.
[0213] Among them, the historical candidate set can be understood as the set of historical pod nodes to be scaled down in the above embodiments.
[0214] Specifically, in the case of an existing set of historical pod nodes to be scaled down, execute Step 718; if not, execute Step 716.
[0215] Step 716: Take the current candidate set as the historical candidate set.
[0216] In the case where there is no historical candidate set, take the current candidate set as the historical candidate set, and return to execute Step 704, waiting to take the intersection with the current candidate set obtained in the next round of loop to determine the pod nodes to be scaled down.
[0217] Step 718: Select the intersection of the double candidate sets.
[0218] In the case of an existing historical candidate set, take the intersection of the current candidate set and the historical candidate set to determine the pod nodes to be scaled down.
[0219] Step 720: Scale down and execute Step 702.
[0220] Specifically, even during the low-traffic period, a long connection state is maintained between the user and the pod node, and directly disconnecting will affect the user. After ensuring that the long connections on the pod nodes to be scaled down are completely disconnected, take offline the pod nodes to be scaled down to complete the scaling down of the pod nodes to be scaled down.
[0221] The node scaling-down method provided in the embodiments of this specification, based on the method of sorting by traffic and the number of connections, performs candidate set processing on the pod nodes to be scaled down. By adopting a double-candidate set scheme based on time series, it selects the machines with the lowest number of connections and traffic in the recent period for scaling down, effectively avoiding the bad impact on users caused by the sudden interruption of long connections. The scaling-down scheme based on the double strategies of the number of connections and traffic and the double-candidate set realizes minute-level scaling down during the low peak period of the project, effectively reducing the impact on users caused by the scaling down of the application cluster.
[0222] Corresponding to the above method embodiments, this specification also provides embodiments of a node scaling system, Figure 8 showing the structural schematic diagram of a node scaling system provided by an embodiment of this specification. As Figure 8As shown in the figure, the system includes an index monitoring component 802, an elastic scaling component 804, and a node scaling component 808, where:
[0223] The index monitoring component 802 is used to obtain the basic index data and operation and maintenance index data of the data processing node. Among them, the data processing node is the node that establishes a long connection with the corresponding client, and sends the basic index data and the operation and maintenance index data to the elastic scaling component 804;
[0224] The elastic scaling component 804 is used to determine the node scaling strategy according to the basic index data and the operation and maintenance index data, and send the node scaling strategy to the node scaling component 808;
[0225] The node scaling component 808 is used to add or take offline data processing nodes according to the node scaling strategy.
[0226] The system further includes a configuration policy component, where:
[0227] The elastic scaling component 804 is used to receive the node scaling strategy sent by the configurator through the configuration policy component, and configure the node scaling strategy for the data processing node.
[0228] The system further includes a gateway service component, where:
[0229] The gateway service component is used to receive the data processing request sent by the user, determine the target data processing node according to the minimum connection number strategy, and establish a long connection between the user and the target data processing node.
[0230] Optionally, the elastic scaling component 804 is further used for:
[0231] In the case where it is determined that the basic index data or the operation and maintenance index data is greater than or equal to the index expansion threshold, determine the number of data processing nodes to be expanded according to the node expansion rule, and add data processing nodes according to the number of data processing nodes to be expanded;
[0232] In the case where it is determined that the operation and maintenance index data is less than the operation and maintenance index shrinkage threshold, determine the data processing nodes to be shrunk according to the data processing nodes added within the preset time period and the node shrinkage rule, and take offline the data processing nodes to be shrunk.
[0233] Optionally, the elastic scaling component 804 is further used for:
[0234] Obtain the newly added data processing nodes added within the preset time period, and sort the newly added data processing nodes according to the operation and maintenance index data of the newly added data processing nodes;
[0235] Determine the data processing nodes to be scaled down according to the node scaling-down rule and the sorted newly added data processing nodes.
[0236] Optionally, the elastic scaling component 804 is further configured to:
[0237] Determine the number of data processing nodes to be scaled down according to the node scaling-down rule, and determine the data processing nodes to be scaled down from the sorted newly added data processing nodes according to the number of the data processing nodes to be scaled down.
[0238] Optionally, the elastic scaling component 804 is further configured to:
[0239] Determine the number of data processing nodes to be scaled down according to the node scaling-down rule, and determine the current set of data processing nodes to be scaled down from the sorted newly added data processing nodes according to the number of the data processing nodes to be scaled down;
[0240] Determine whether there is a set of historical data processing nodes to be scaled down,
[0241] If so, determine the data processing nodes to be scaled down according to the current set of data processing nodes to be scaled down and the set of historical data processing nodes to be scaled down,
[0242] If not, determine the current set of data processing nodes to be scaled down as the set of historical data processing nodes to be scaled down, and continue to execute the step of determining the data processing nodes to be scaled down according to the data processing nodes added within a preset time period and the node scaling-down rule and taking offline the data processing nodes to be scaled down in the case that the operation and maintenance metric data is less than the operation and maintenance metric scaling-down threshold.
[0243] Optionally, the metric monitoring component 802 is further configured to:
[0244] Obtain the basic monitoring data of the data processing nodes, and use the average value of the basic monitoring data as the basic metric data, and obtain the operation and maintenance metric data of each data processing node.
[0245] The node scaling system provided by the embodiments of the present specification accurately selects the corresponding node scaling strategy according to the basic metric data and the operation and maintenance metric data of the obtained data processing nodes, so as to reasonably add or take offline data processing nodes according to the node scaling strategy, efficiently implement node scaling, avoid service anomalies caused by overloading of data processing nodes, and effectively reduce the impact of node scaling-down on the project according to the node scaling strategy.
[0246] The above is a schematic solution of a node scaling system according to this embodiment. It should be noted that the technical solution of this node scaling system and the technical solution of the above node scaling method belong to the same concept. For the details not described in the technical solution of the node scaling system, reference can be made to the description of the technical solution of the above node scaling method.
[0247] Figure 9 FIG. shows a structural block diagram of a computing device 900 according to an embodiment of this specification. The components of the computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 through a bus 930, and a database 950 is used to store data.
[0248] The computing device 900 further includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface.
[0249] In an embodiment of this specification, the above components of the computing device 900 and Figure 9 other components not shown in Figure 9 may also be connected to each other, for example, through a bus. It should be understood that
[0250] The computing device 900 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 900 can also be a mobile or stationary server.
[0251] Wherein, the processor 920 is configured to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above node scaling method are implemented.
[0252] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above node scaling method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above node scaling method.
[0253] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the above node scaling method are implemented.
[0254] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above node scaling method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above node scaling method.
[0255] An embodiment of this specification also provides a computer program, wherein when the computer program is executed on a computer, the computer is made to execute the steps of the above node scaling method.
[0256] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above node scaling method belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the description of the technical solution of the above node scaling method.
[0257] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0258] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or system capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0259] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential for the embodiments of this specification.
[0260] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0261] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and utilize this specification well. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A node expansion and contraction method, comprising: Obtaining basic indicator data and operation and maintenance indicator data of a data processing node, wherein the data processing node is a node that establishes a persistent connection with a corresponding client; According to the basic indicator data, the operation and maintenance indicator data and the node expansion and contraction strategy, data processing nodes are added or offline.
2. According to the node expansion and contraction method of claim 1, the node expansion and contraction strategy includes a node expansion strategy and a node contraction strategy, and the node expansion strategy includes a node expansion rule, and the node contraction strategy includes a node contraction rule; Accordingly, according to the basic indicator data, the operation and maintenance indicator data and the node expansion and contraction strategy, adding or offline data processing nodes includes: When it is determined that the basic indicator data or the operation and maintenance indicator data is greater than or equal to the indicator expansion threshold, the number of data processing nodes to be expanded is determined according to the node expansion rule, and data processing nodes are added according to the number of data processing nodes to be expanded; or When it is determined that the operation and maintenance indicator data is less than the operation and maintenance indicator shrinking threshold, the data processing nodes to be shrunk are determined according to the data processing nodes added within the preset time period and the node shrinking rule, and the data processing nodes to be shrunk are taken offline.
3. The node expansion and contraction method according to claim 2, wherein the determining the data processing nodes to be contracted according to the data processing nodes added within a preset time period and the node contraction rule comprises: Acquire newly added data processing nodes added within a preset time period, and sort the newly added data processing nodes according to the operation and maintenance indicator data of the newly added data processing nodes; The data processing nodes to be shrunk are determined according to the node scaling-down rule and the sorted newly added data processing nodes.
4. The node expansion and contraction method according to claim 3, wherein determining the data processing nodes to be contracted according to the node contraction rule and the sorted newly added data processing nodes comprises: The number of data processing nodes to be shrunk is determined according to the node shrinking rule, and the data processing nodes to be shrunk are determined from the sorted newly added data processing nodes according to the number of data processing nodes to be shrunk.
5. The node expansion and contraction method according to claim 3, wherein determining the data processing nodes to be contracted according to the node contraction rule and the sorted newly added data processing nodes comprises: Determine the number of data processing nodes to be shrunk according to the node shrinking rule, and determine the current set of data processing nodes to be shrunk from the sorted newly added data processing nodes according to the number of data processing nodes to be shrunk; Determine whether there is a set of historical data processing nodes to be shrunk. If yes, determine the data processing nodes to be shrunk according to the current set of data processing nodes to be shrunk and the historical set of data processing nodes to be shrunk, If not, the current set of data processing nodes to be shrunk is determined as the historical set of data processing nodes to be shrunk, and the step of determining the data processing nodes to be shrunk is continued according to the data processing nodes added within the preset time period and the node shrinkage rules when it is determined that the operation and maintenance indicator data is less than the operation and maintenance indicator shrinkage threshold.
6. The node expansion and contraction method according to claim 1, after adding or removing data processing nodes according to the basic indicator data, the operation and maintenance indicator data and the node expansion and contraction strategy, further comprising: A data processing request sent by a client is received, a target data processing node is determined according to a minimum connection number strategy, and a long connection between the client and the target data processing node is established.
7. The node expansion and contraction method according to claim 1, before obtaining the basic indicator data and operation and maintenance indicator data of the data processing node, further comprising: Receive the node expansion strategy sent by the configuration personnel, and configure the node expansion strategy for the data processing node.
8. According to the node expansion and contraction method of claim 1, the step of obtaining basic indicator data and operation and maintenance indicator data of the data processing node comprises: The basic monitoring data of the data processing node is obtained, and the average value of the basic monitoring data is used as the basic indicator data, and the operation and maintenance indicator data of each data processing node is obtained.
9. According to the node expansion and contraction method according to any one of claims 1 to 8, the basic indicator data includes processor indicator data and memory indicator data, and the operation and maintenance indicator data includes query rate per second and number of threads.
10. The node expansion and contraction method according to claim 2, wherein the node expansion rule comprises determining the number of data processing nodes to be expanded according to a preset node expansion step, or determining the number of data processing nodes to be expanded according to the data processing request of the client, the basic indicator data, and the operation and maintenance indicator data; The node scaling-down rule includes determining the number of data processing nodes to be scaled down according to a preset node scaling-down step, or determining the number of data processing nodes to be scaled down according to the basic indicator data and operation and maintenance indicator data of the data processing nodes added within the preset time period.
11. A node expansion and contraction system, comprising an elastic expansion and contraction component, an indicator monitoring component, and a node expansion and contraction component, wherein: The indicator monitoring component is used to obtain basic indicator data and operation and maintenance indicator data of the data processing node, wherein the data processing node is a node that establishes a long connection with the corresponding client, and sends the basic indicator data and the operation and maintenance indicator data to the elastic expansion component; The elastic scaling component is used to determine the node scaling strategy according to the basic indicator data and the operation and maintenance indicator data, and send the node scaling strategy to the node scaling component; The node expansion component is used to add or remove data processing nodes according to the node expansion strategy.
12. The node expansion and contraction system according to claim 11, further comprising a configuration policy component, wherein: The elastic scaling component is used to receive the node scaling policy sent by the configuration personnel through the configuration policy component, and configure the node scaling policy for the data processing node.
13. The node expansion and contraction system according to claim 11, further comprising a gateway service component, wherein: The gateway service component is used to receive a data processing request sent by a client, determine a target data processing node according to a minimum connection number strategy, and establish a long connection between the client and the target data processing node.
14. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the node scaling method described in any one of claims 1 to 9 are implemented.
15. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the node scaling method described in any one of claims 1 to 9.