A distributed cluster control method and device
By diversion of service requests based on the status of the host and service interface in a distributed cluster, the service request failure caused by the host in a normal state but the service interface is abnormal, and fine-grained monitoring of the distributed cluster and system stability are improved.
Patent Information
- Application Number
- CN201911182842.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2039-11-27
AI Technical Summary
When the amount of service requests received per unit time is too large, it may cause too much pressure on the host and some service interfaces cannot provide services to the outside world. However, Nginx still determines that the host is alive, causing service requests to flow into the abnormal service interface, causing service requests to fail.
By determining a host corresponding to the target service system and in a normal state in a distributed cluster, and selecting a target service interface with an exception request number less than the first threshold based on the number of exception requests of the service interface in the host, the service request is sent to the corresponding second host to ensure that the service request is diverted to the interface that can provide services normally.
It realizes that when the host is in a normal state, it quickly senses the abnormal situation of some service interfaces, avoids service requests from flowing into the abnormal interface, improves the system stability of the distributed cluster, and reduces economic losses.
Smart Images

Figure CN112860505B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a distributed cluster control method and device. Background Art
[0002] With the development of Internet technology, the number of daily visits per second has gradually increased, and the deployment of distributed clusters has emerged. In order to ensure the system stability of distributed clusters, real-time monitoring of distributed clusters is required so that user service requests can be processed as soon as possible.
[0003] At present, Nginx load balancing is usually used to achieve real-time monitoring of the host. Specifically, the host IP bound to the domain name is accessed in real time through a scheduled task, and the host is accessed through the Ping command to obtain the host Pong feedback. When the host Pong feedback is returned normally and the parsed pong data is normal, it is determined that the current host is alive and can provide services normally; if the pong feedback cannot be obtained or the pong feedback data is abnormal, the current host is determined to be down.
[0004] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0005] When the number of service requests received per unit time is too large, the host may be overloaded, causing some of the host's service interfaces to be unable to provide services to the outside world. However, if Nginx detects that the host is still alive and the host feedback is normal, the service request will still flow into the abnormal service interface of the current host, causing the service request to fail. Summary of the invention
[0006] In view of this, an embodiment of the present invention provides a control method for a distributed cluster, which can divert service requests to service interfaces that can provide services normally when the host is in a normal state, and ensure that the service interface in an abnormal state will not receive service requests, thereby realizing fine-grained monitoring of the distributed cluster to ensure that service requests can be processed quickly and successfully.
[0007] To achieve the above objective, according to one aspect of an embodiment of the present invention, a distributed cluster control method is provided.
[0008] A distributed cluster control method according to an embodiment of the present invention includes: receiving a service request, wherein the service request indicates a target service system and a requested service type;
[0009] Determine one or more first hosts corresponding to the target service system and in a normal state from the distributed cluster; wherein the distributed cluster includes one or more service systems, the service systems correspond to one or more hosts, and the hosts correspond to one or more service interfaces for providing services;
[0010] In the case where it is determined that the first host exists, according to the number of abnormal requests of the service interface corresponding to the service type in the first host, selecting a target service interface whose number of abnormal requests is less than a first threshold for the service request; wherein the number of abnormal requests indicates the number of historical times when the service interface has abnormalities;
[0011] The service request is sent to a second host corresponding to the target service interface in the first host, so that the second host provides a service corresponding to the service type through the target service interface.
[0012] Optionally, the method further comprises:
[0013] Obtaining a request result regarding the service request returned by the second host;
[0014] When the request result is abnormal, the number of abnormal requests of the target service interface in the second host is incremented.
[0015] Optionally, determining one or more first hosts corresponding to the target service system and in a normal state from the distributed cluster includes:
[0016] Receive the configuration information of the host in normal state and the service interface corresponding to the host in the distributed cluster, and form a distribution map corresponding to the distributed cluster according to the configuration information; wherein the configuration information of the host indicates the service system corresponding to the host; and determine the first host according to the distribution map.
[0017] Optionally,
[0018] When a plurality of first hosts are determined, the second host is determined from the plurality of first hosts in a polling manner.
[0019] Optionally, the determining the second host from the plurality of first hosts in a polling manner includes:
[0020] The following steps are executed in a loop until it is determined that the number of the second host or the number of the polled first host is greater than the second threshold:
[0021] Determine the number of abnormal requests of the service interface corresponding to the service type among the service interfaces corresponding to the current host among the plurality of first hosts, and determine whether the number of abnormal requests is less than a first threshold;
[0022] If yes, taking the current host as the second host, and taking the service interface corresponding to the service type as the target service interface;
[0023] If not, a first host that has not been selected is selected from the plurality of hosts as the current host.
[0024] Optionally, the method further comprises:
[0025] The service interface whose number of abnormal requests is not less than the first threshold is called by using the historical service request with successful request. When the call is successful, the number of abnormal requests of the service interface whose number of abnormal requests is not less than the first threshold is decremented; when the call fails and the number of failed calls is greater than a third threshold, it is determined that the service interface whose number of abnormal requests is not less than the first threshold is in an abnormal state.
[0026] Optionally, the method further comprises:
[0027] In the absence of the first host or the service interface being in an abnormal state, an alarm message is output.
[0028] To achieve the above objective, according to another aspect of an embodiment of the present invention, a distributed cluster control device is provided.
[0029] A distributed cluster control device according to an embodiment of the present invention includes: a request receiving module, a host determination module, an interface selection module and a processing module; wherein:
[0030] The request receiving module is used to receive a service request, wherein the service request indicates a target service system and a requested service type;
[0031] The host determination module is used to determine one or more first hosts corresponding to the target service system and in a normal state from the distributed cluster; wherein the distributed cluster includes one or more service systems, the service systems correspond to one or more hosts, and the hosts correspond to one or more service interfaces for providing services;
[0032] The interface selection module is configured to select, when it is determined that the first host exists, a target service interface whose number of abnormal requests is less than a first threshold for the service request according to the number of abnormal requests of the service interface corresponding to the service type in the first host; wherein the number of abnormal requests indicates the number of historical times when the service interface has been abnormal;
[0033] The processing module is used to send the service request to a second host corresponding to the target service interface in the first host, so that the second host provides a service corresponding to the service type through the target service interface.
[0034] Optionally,
[0035] The processing module is further configured to obtain a request result regarding the service request returned by the second host, and if the request result is abnormal, increment the number of abnormal requests of the target service interface in the second host.
[0036] Optionally, the device further includes: a configuration module; wherein,
[0037] The configuration module is used to receive configuration information of a host in a normal state in the distributed cluster and a service interface corresponding to the host, and form a distribution map corresponding to the distributed cluster according to the configuration information; wherein the configuration information of the host indicates the service system corresponding to the host.
[0038] To achieve the above objective, according to another aspect of an embodiment of the present invention, a server for regulating and controlling a distributed cluster is provided.
[0039] A server for controlling a distributed cluster according to an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement a method for controlling a distributed cluster according to an embodiment of the present invention.
[0040] To achieve the above objective, according to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided.
[0041] A computer-readable storage medium according to an embodiment of the present invention stores a computer program, and when the program is executed by a processor, a distributed cluster control method according to an embodiment of the present invention is implemented.
[0042] One embodiment of the above invention has the following advantages or beneficial effects: when a service request is received, one or more hosts corresponding to the target service system indicated by the service request and in a normal state are determined from the distributed cluster, and then according to the number of abnormal requests of the service interface corresponding to the service type indicated by the service request in the first host, a target service interface with an abnormal request number less than a first threshold is selected for the service request, and then the service request is sent to the second host corresponding to the target service interface, so that the second host provides the service corresponding to the service request through the target service interface, thereby diverting the service request to the service interface that can provide the service normally, and ensuring that the service interface in an abnormal state will not receive the service request, so that the service request can be successfully processed. Thus, according to the monitoring of both the host and the service interface, it is ensured that when the host is in a normal state, the abnormal situation of some service interfaces can be quickly perceived, which solves the problem of service request failure caused by the host being in a normal state but the service interface being abnormal, realizes fine-grained monitoring of the distributed cluster, improves the system stability of the distributed cluster, and avoids economic losses caused by failure to monitor the abnormality of the service system in a timely manner.
[0043] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with the specific implementation manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.
[0045] Figure 1 is a schematic diagram of main steps of a distributed cluster control method according to an embodiment of the present invention;
[0046] Figure 2 is a schematic diagram of main modules of a distributed cluster control device according to an embodiment of the present invention;
[0047] Figure 3 is a schematic diagram of a distribution diagram of a distributed cluster according to an embodiment of the present invention;
[0048] Figure 4 is a schematic diagram of a distribution diagram of another distributed cluster according to an embodiment of the present invention;
[0049] Figure 5 is a schematic diagram of main steps of a polling process in a distributed cluster control method according to an embodiment of the present invention;
[0050] Figure 6 It is a schematic diagram of the main steps of updating the number of abnormal requests in a distributed cluster control method according to an embodiment of the present invention;
[0051] Figure 7 is a schematic diagram of main modules of another distributed cluster control device according to an embodiment of the present invention;
[0052] Figure 8 is an exemplary system architecture diagram to which embodiments of the present invention may be applied;
[0053] Fig. 9 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0055] It should be pointed out that the embodiments of the present invention and the technical features therein may be combined with each other without conflict.
[0056] The significance of distributed systems lies not only in the comprehensive utilization of resources distributed in various places and the transfer of loads from a single node to multiple nodes, which can double the system throughput, but also in the convenience of rapid capacity expansion, ensuring the stability of the system under force majeure. However, when problems occur on some hosts or service interfaces in a distributed cluster, it takes a relatively long time to locate and solve the problem, which can cause huge losses. In addition, when the number of service requests received per unit time is too large, it may cause excessive pressure on the host, which may cause some of the host's service interfaces to be unable to provide services to the outside world. However, if Nginx detects that the host is still alive and the host feedback is normal, the service request will still flow into the abnormal service interface of the current host, resulting in the failure of the service request.
[0057] In order to solve the above problems, an embodiment of the present invention provides a control method for a distributed cluster, which realizes fine-grained monitoring of the distributed cluster, expands on the basis of Nginx monitoring host dimension, and realizes efficient monitoring of the distributed cluster host in the form of distributed locks. And provides monitoring of the service interface dimension to solve the economic losses caused by the current host survival but service interface abnormality. Through dual monitoring of host survival monitoring and service interface abnormality capture, it is ensured that when the host is normal, some interfaces have abnormal conditions and the system can quickly perceive them. And the present invention will rely on the basis of dual monitoring. When the interface service has an abnormality exceeding the specified number of times, it will quickly perceive and quickly divert the traffic to ensure that the current abnormal interface service will not receive any requests.
[0058] Specifically, Figure 1 As shown, a distributed cluster control method according to an embodiment of the present invention mainly includes the following steps:
[0059] Step S101: receiving a service request, wherein the service request indicates a target service system and a requested service type.
[0060] A distributed cluster control method provided by an embodiment of the present invention can be based on the following Figure 2 The control device shown is implemented, and the control device may include: a resource configuration module, a flow control gateway, a load balancing module and a message alarm module. Among them, the resource configuration module can be used to uniformly manage the service system in the distributed cluster, the host corresponding to the service system, and the service interface corresponding to the host, and form a distribution map of the service system, the host and the service interface. The load balancing module can synchronize data in real time according to the subscription message of the resource configuration module, update the distribution map corresponding to the distributed cluster to the local memory, and when receiving the request of the flow control gateway, return the corresponding host IP address list to the flow control gateway according to the distribution map stored in the memory data. The flow control gateway can integrate service information according to the host IP address list returned by the load balancing module, and provide interface services to the outside. The message alarm module can alarm in time when the system is in an abnormal state, so as to prompt the operation and maintenance personnel to deal with it as soon as possible. The functions of the above modules and their implementation methods will be further explained in the following embodiments.
[0061] The resource configuration module can receive the configuration information of the host in the distributed cluster that is in a normal state and the service interface corresponding to the host, and form a distribution map corresponding to the distributed cluster according to the configuration information; wherein the configuration information of the host indicates the service system corresponding to the host. For example, the host sends its own IP as configuration information to the resource configuration module, so that the resource configuration module forms a corresponding distribution map according to the host IP.
[0062] Specifically, the resource configuration module can be supported by the Zookeeper distributed architecture. By using the Znode temporary node method, the host IP corresponding to multiple service systems in the distributed cluster and the information of multiple service interfaces corresponding to multiple hosts are stored in the Znode to form a distribution map of the distributed cluster. The distribution map is as follows: Figure 3 The tree data structure shown realizes efficient monitoring of the distributed cluster in the form of distributed locks.
[0063] It can be understood that a host that can establish a long link with the resource configuration module is a host in a normal state.
[0064] like Figure 3As shown, each service system corresponds to one or more hosts. In the resource configuration module, the service system serves as a parent node and binds its corresponding host IP as a child node. When a long link is established between the host and the resource configuration module, the host IP registration is completed. Then the information of all service interfaces provided by the registered host can be traversed to register one or more service interfaces corresponding to the host to the resource configuration module. When the service interface is registered with the resource configuration module, the service interface data is registered under the temporary node of the corresponding host IP, and the service interface serves as a child node of the corresponding host IP. When the service interface is registered with the corresponding host IP, the corresponding number of exception requests will be initialized, and the initialized number of exception requests is generally set to 0.
[0065] Take the inventory service system of a specific e-commerce platform as an example. Figure 4 As shown, the inventory service system is the parent node, under which host IP A and host IP B are registered. Under host IP A, a service interface for providing product names and a service interface for providing product pictures are registered. The initial number of abnormal requests for the two service interfaces is 0.
[0066] Among them, ZooKeeper is an open source distributed application coordination service, an open source implementation of Google's Chubby, and an important component of Hadoop and Hbase. It is a software that provides consistency services for distributed applications, and provides functions including: configuration maintenance, domain name services, distributed synchronization, group services, etc. ZooKeeper implements the namespace in a tree data structure similar to the file system. Each node in the namespace is a znode. The znode is different from the path of the file system. In the file system, the path is just a name and does not contain data. The znode is not only a path, but also carries data.
[0067] When deploying a distributed cluster, the host corresponding to each service system in the distributed cluster will establish a long link with the resource configuration module, and each host will register the host IP and service interface related information to the resource configuration mode in the form of a temporary node. When the host is down, that is, when the host is in an abnormal state, the host will disconnect the long link with the resource configuration module, and the temporary node will automatically disappear. Correspondingly, the distribution map in the resource configuration module will also be updated accordingly, that is, the host IP in the abnormal state will be deleted from the distribution map. After the temporary node disappears, Zookeeper informs the flow control gateway, load balancing module, and message alarm module through the subscription mode that the temporary node corresponding to the abnormal host disappears. That is to say, when the host is in an abnormal state, the resource configuration module pushes the corresponding abnormal message to the flow control gateway, load balancing module, and message alarm module. At this time, the load balancing module can update the distribution map stored in the memory in real time according to the subscription message of the resource configuration module. There is no host IP in the abnormal state in the updated distribution map.
[0068] It can be understood that the operations of registering the service system, host and service interface in the resource configuration module can be performed when deploying the distributed cluster. After the distributed cluster is deployed, the service request can be received through the flow control gateway. The service request includes the domain name of the target service system to indicate the target service system that provides the target service and indicates the requested service type, for example, indicating that the requested service type is to provide a product name or provide a product picture.
[0069] Step S102: Determine one or more first hosts in the distributed cluster that correspond to the target service system and are in a normal state; wherein the distributed cluster includes one or more service systems, the service systems correspond to one or more hosts, and the hosts correspond to one or more service interfaces for providing services.
[0070] When receiving a service request, the flow control gateway parses the domain name of the service request to determine the target service system indicated by the service request, and parses the information related to the service interface corresponding to the service type. Then, a load request is sent according to the service request to request the corresponding host list from the load balancing module according to the target service system indicated by the service request. The load balancing module reads the memory data according to the request of the flow control gateway to traverse the first host IP under the parent node corresponding to the target service system indicated by the service request according to the distribution map corresponding to the distributed cluster, generate the first host IP list, and feed the first host IP list back to the flow control gateway.
[0071] It is worth mentioning that the host that maintains a long link with the resource configuration module is a host in a normal state, and the memory data of the load balancing module is updated according to the subscription message of the resource configuration module. Therefore, in the first host IP list generated by the load balancing module, the first hosts corresponding to each first host IP are all hosts in a normal state.
[0072] Step S103: when it is determined that the first host exists, according to the number of abnormal requests of the service interface corresponding to the service type in the first host, a target service interface whose number of abnormal requests is less than a first threshold is selected for the service request.
[0073] Step S104: Send the service request to a second host corresponding to the target service interface in the first host, so that the second host provides a service corresponding to the service type through the target service interface; wherein the number of abnormal requests indicates the historical number of times the service interface has abnormalities.
[0074] After receiving the first host IP list feedback from the load balancing module, the flow control gateway first determines whether the first host IP list is empty. If it is empty, it means that there is no first host in normal state or all gates of all first hosts under the target service system indicated by the service request, that is, there is no first host that can provide services. At this time, the flow control gateway notifies the message alarm module, so that the message alarm module quickly issues an alarm message to prompt the operation and maintenance personnel to quickly deal with the abnormal situation and minimize the economic losses caused by the abnormal situation.
[0075] In addition, when the first host IP list is not empty, it indicates that there is a first host in a normal state, that is, there is a first host that can provide services. When there is only one first host that can provide services, determine whether the number of abnormal requests for the service interface corresponding to the service type indicated by the service request in the first host is less than the first threshold. If so, send the service request to this unique first host, so that the first host provides the service requested by the service request through the target service interface whose number of abnormal requests is less than the first threshold. When the number of abnormal requests for the service interface in the unique first host is not less than the first threshold, the flow control gateway can notify the message alarm module to alarm.
[0076] In addition, when there are multiple first host IPs in the first host IP list, it means that there are multiple first hosts in a normal state. At this time, a polling method can be used to determine the second host that can provide services according to the service request from the multiple first hosts. For example, the second host can be determined by polling one by one or by weighted polling. Specifically, the following steps can be performed repeatedly until it is determined that the number of second hosts or polled first hosts is greater than the second threshold: determine the number of abnormal requests of the service interface corresponding to the service type in the service interface corresponding to the current host among the multiple first hosts, and determine whether the number of abnormal requests is less than the first threshold; if so, use the current host as the second host, and use the service interface corresponding to the service type as the target service interface; if not, select a first host that has not been selected from the multiple hosts as the current host.
[0077] Here, when there are multiple first hosts in a normal state, the flow control gateway can hit the current host among the multiple first hosts (the current host can be any one of the multiple first hosts) by modulo algorithm polling or weighted polling. Then the flow control gateway reads the data synchronized by the resource configuration module, finds the parent node of the host IP corresponding to the current host from the distribution map corresponding to the distributed cluster, and then determines the service interface corresponding to the service type indicated by the service request under the parent node of the host IP, and determines the service interface corresponding to the service request by obtaining the name information of the service interface, and determines the number of abnormal requests of the service interface according to the child node information of the service interface.
[0078] For example, when the distribution graph of the distributed cluster is Figure 4 When the target service system indicated by the service request is the inventory system and the service type indicated by the service request is to obtain product images, the first host IP list returned by the load balancing module to the flow control gateway includes the A host IP and the B host IP. In this case, the A host can be used as the current host first, and then the product image interface under the A host can be determined as the service interface corresponding to the service request. Figure 4 From the tree structure shown, it is known that the number of abnormal requests for the service interface (product picture interface) is 0, and the first threshold is a value greater than 0. For example, when the first threshold is 3, the number of abnormal requests for the product picture interface is less than the first threshold, then the service request can be sent to host A, so that host A can provide corresponding services to the user through the product picture interface.
[0079] In this example, if the number of abnormal requests for the product image interface under host A is not less than the first threshold, it means that the service interface corresponding to the service request in the current host has an abnormality and cannot provide services to the outside. At this time, the flow control gateway will re-hit the first host that has not been selected in the first host IP list as the current host through polling or weighted polling algorithm. For example, in this example, re-hit host B, use host B as the current host, and query whether the service interface under host B has a service interface corresponding to the service request and the number of abnormal requests for the service interface is less than the first threshold. This cycle is repeated until a second host that can provide services according to the service request is determined from multiple first hosts or the number of polled first hosts is greater than the second threshold, then polling is stopped. In addition, when there is no service interface corresponding to the service type under the current host, the flow control gateway will also re-select a first host that has not been selected from the multiple hosts as the current host and continue polling.
[0080] The number of the second threshold can be determined according to the number of first hosts in the first host IP list, for example, half of the number of first hosts in the first host IP list is determined as the second threshold. When more than half of the first hosts among the multiple first hosts are polled, a second host that can provide services according to the service request is still not determined, indicating that an abnormality may have occurred in the distributed cluster. In order to shorten the response time and perform timely maintenance on the cluster, the polling is stopped and the message alarm module is notified to alarm.
[0081] Reference below Figure 5 , the process of the flow control gateway receiving the service request and selecting the second host to provide the service for the service request is described in detail, such as Figure 5 As shown, the process may include the following steps:
[0082] Step S501: receiving a service request, and requesting a host IP list from a load balancing module according to a target service system indicated by the service request.
[0083] Step S502: Receive the host IP list fed back by the load balancing module.
[0084] Step S503: Determine whether the host IP list is empty, if so, execute step S504, otherwise execute step S505.
[0085] Step S504: the notification message alarm module generates an alarm and ends the current process.
[0086] Step S505: Select the current host from the first host corresponding to the host IP list.
[0087] Step S506: Determine the number of abnormal requests of the service interface corresponding to the service type indicated by the service request in the service interface corresponding to the current host.
[0088] Step S507: Determine whether the number of abnormal requests is less than the first threshold, if so, execute step S508, otherwise execute step S509.
[0089] Step S508: The current host is used as the second host for providing services for the service request, and the current process ends.
[0090] Step S509: Determine whether there is an unselected host in the first host corresponding to the host IP list. If yes, execute step S510; otherwise, end the current process.
[0091] Step S510: determine whether the number of selected first hosts is greater than a second threshold; if so, end the current process; otherwise, execute step S511.
[0092] Step S511: Select a host that has not been selected from the first host corresponding to the host IP list as the current host, and execute step S506.
[0093] After determining the second host that provides the service request, the service request can be sent to the second host, so that the second host provides the service corresponding to the service type through the target service interface whose abnormal request number is less than the first threshold. Then, the request result returned by the second host regarding the service request is obtained; if the request result is abnormal, the abnormal request number of the target service interface in the second host is incremented.
[0094] After determining the target service interface that can provide services normally, the flow control gateway can forward the service request to the second host corresponding to the target service interface to obtain the request result of the second host for the service request. If the request result is normal, it means that the user's service request has been successfully requested. If the request result is abnormal, it means that the target service interface has an abnormality. The flow control gateway will update the number of abnormal requests of the relevant service interface in the resource configuration module. Specifically, the flow control gateway will increase the number of abnormal requests of the target service interface in the resource configuration module by 1.
[0095] That is, after determining that the number of abnormal requests of the target service interface in the first host is less than the first threshold, the distributed cluster control method provided by the embodiment of the present invention may further include: Figure 6 The steps shown, the following steps can be performed by the flow control gateway.
[0096] Step S601: Send a service request to a second host corresponding to a target service interface.
[0097] Step S602: Receive a request result of the second host for the service request.
[0098] Step S603: determine whether the request result is normal, if yes, end the current process, otherwise execute step S604;
[0099] Step S604: Incrementally increase the number of abnormal requests of the target service interface.
[0100] In the implementation mode of the present invention, the number of abnormal requests of the target service interface is incremented by +1. It is understandable that when each service interface is registered with the resource configuration center, since the number of abnormal requests indicates the historical number of abnormalities of the service interface, the initial number of abnormal requests of each service interface is 0. During the operation of the service system, when an abnormality occurs in a service interface, its number of abnormal requests increases.
[0101] In addition, when the host is not disconnected from the resource configuration module and the temporary node corresponding to the host in the resource configuration module has not disappeared, it means that the host is in a normal state. If the number of abnormal requests of the service interface under this host is not less than the first threshold at this time, the flow control gateway will perform a survival check on the service interface. When performing the survival check, the service interface with the number of abnormal requests not less than the first threshold is called using the historical service request with successful requests. When the call is successful, the number of abnormal requests of the service interface with the number of abnormal requests not less than the first threshold is decremented; when the call fails and the number of failed calls is greater than the third threshold, it is determined that the service interface with the number of abnormal requests not less than the first threshold is in an abnormal state.
[0102] For example, a historical service request that can normally call other service interfaces is captured, and the historical service request is used to periodically call a service interface whose number of abnormal requests is not less than a first threshold (hereinafter referred to as the abnormal service interface). If the abnormal service interface returns a normal request result, it is determined that the abnormal service interface can provide services normally. The previously increased number of abnormal requests may be due to system misjudgment, such as a misjudgment caused by abnormal data of the service request. At this time, the number of abnormal requests of the abnormal service interface is reduced by 1 to achieve accurate monitoring of the service interface.
[0103] If the request result returned by the abnormal service interface is abnormal or the historical service request fails to call the abnormal service interface, the abnormal service interface will be called again according to the period set by the scheduled task. If the call still fails or the request result returned by the abnormal service interface is abnormal, and the number of failures or exceptions is greater than the third threshold (for example, greater than 2 times), it means that the abnormal service interface cannot provide services normally, and the flow control gateway determines that the abnormal service interface is dead. At this time, the flow control gateway notifies the message alarm module to alarm.
[0104] In summary, in the distributed cluster control method provided in the embodiment of the present invention, the following three abnormal requests can be monitored:
[0105] 1. When the host is gated (in an abnormal state), the long link between the host and the resource configuration module is disconnected, and the temporary data node corresponding to the host disappears, the first host IP list returned by the load balancing module to the flow control gateway will no longer obtain the IP address of the gated host, ensuring that the flow control gateway cannot hit the gated host, so that the host in an abnormal state will no longer receive any service requests.
[0106] 2. All service systems connected to the distributed cluster control device need to inject the resource configuration module into the system through Spring Aop. The distributed cluster control device configures the path to be monitored through the exception aspect method in Spring Aop. If an exception occurs in the service interface, the distributed cluster control device will update the resource configuration module, and the number of exception requests of the current service interface node will increase by 1.
[0107] 3. When the flow control gateway sends a service request to the host, it determines whether the service request is successful through the request result (such as the return code). If it is unsuccessful, the number of abnormal requests of the service interface is updated.
[0108] According to the control method of the distributed cluster of the embodiment of the present invention, it can be seen that when a service request is received, one or more hosts corresponding to the target service system indicated by the service request and in a normal state are determined from the distributed cluster, and then according to the number of abnormal requests of the service interface corresponding to the service type indicated by the service request in the first host, a target service interface with an abnormal request number less than a first threshold is selected for the service request, and then the service request is sent to the second host corresponding to the target service interface, so that the second host provides the service corresponding to the service request through the target service interface, thereby diverting the service request to the service interface that can provide the service normally, and ensuring that the service interface in an abnormal state will not receive the service request, so that the service request can be successfully processed. Therefore, according to the monitoring of both the host and the service interface, it is ensured that when the host is in a normal state, the abnormal situation of some service interfaces can be quickly perceived, which solves the problem of service request failure caused by the host being in a normal state but the service interface being abnormal, realizes fine-grained monitoring of the distributed cluster, improves the system stability of the distributed cluster, and avoids economic losses caused by failure to monitor the abnormality of the service system in a timely manner.
[0109] Figure 7 It is a schematic diagram of main modules of a distributed cluster control device according to an embodiment of the present invention.
[0110] like Figure 7 As shown, a distributed cluster control device 700 according to an embodiment of the present invention includes: a request receiving module 701, a host determination module 702, an interface selection module 703 and a processing module 704; wherein,
[0111] The request receiving module 701 is used to receive a service request, wherein the service request indicates a target service system and a requested service type;
[0112] The host determination module 702 is used to determine one or more first hosts corresponding to the target service system and in a normal state from the distributed cluster; wherein the distributed cluster includes one or more service systems, the service systems correspond to one or more hosts, and the hosts correspond to one or more service interfaces for providing services;
[0113] The interface selection module 703 is used to select, when it is determined that the first host exists, a target service interface whose number of abnormal requests is less than a first threshold for the service request according to the number of abnormal requests of the service interface corresponding to the service type in the first host; wherein the number of abnormal requests indicates the number of historical times when the service interface has been abnormal;
[0114] The processing module 704 is configured to send the service request to a second host corresponding to the target service interface in the first host, so that the second host provides a service corresponding to the service type through the target service interface.
[0115] In one embodiment of the present invention, the processing module 704 is further used to obtain a request result regarding the service request returned by the second host, and if the request result is abnormal, increment the number of abnormal requests of the target service interface in the second host.
[0116] Continue to refer Figure 7 In one embodiment of the present invention, the control device further includes: a configuration module 705; wherein,
[0117] The configuration module 705 is used to receive the configuration information of the host in the distributed cluster that is in a normal state and the service interface corresponding to the host, and form a distribution map corresponding to the distributed cluster according to the configuration information; wherein the configuration information of the host indicates the service system corresponding to the host.
[0118] In one embodiment of the present invention, the host determination module 702 is used to determine the second host from the multiple first hosts by polling when multiple first hosts are determined.
[0119] In one embodiment of the present invention, the host determination module 702 is used to loop the following steps until it is determined that the number of second hosts or polled first hosts is greater than a second threshold: determine the number of abnormal requests of the service interface corresponding to the service type among the service interfaces corresponding to the current host among the multiple first hosts, and judge whether the number of abnormal requests is less than the first threshold; if so, use the current host as the second host, and use the service interface corresponding to the service type as the target service interface; if not, select a first host that has not been selected from the multiple hosts as the current host.
[0120] In one embodiment of the present invention, the processing module 704 is also used to use the historical service request with successful requests to call the service interface with the number of abnormal requests not less than the first threshold, and when the call is successful, decrement the number of abnormal requests of the service interface with the number of abnormal requests not less than the first threshold; when the call fails and the number of failed calls is greater than a third threshold, determine that the service interface with the number of abnormal requests not less than the first threshold is in an abnormal state.
[0121] like Figure 7 As shown, in one embodiment of the present invention, the control device further includes an alarm module 706, wherein the alarm module 706 is used to output alarm information when the first host or service interface is not in an abnormal state.
[0122] In summary, the above embodiments of the present invention have at least the following beneficial effects:
[0123] 1. The problem that the host is in a normal state but the service interface is abnormal in the existing method is solved. The embodiment of the present invention performs abnormal aspect monitoring through AOP injection, realizes the abnormal capture of request failures caused by code problems, middleware network abnormalities, etc., and uses request results and return codes to be compatible with the business logic abnormalities of the host service interface, and monitors the host IP survival status through a distributed lock temporary node. This solves the problem that only monitoring the host in the existing technology cannot perform fine-grained monitoring.
[0124] 2. The embodiment of the present invention verifies the number of abnormal requests of the service interface through the flow control gateway, performs host re-hitting, ensures that the abnormal service interface no longer provides services, and through the Zookeeper subscription method, when the number of abnormal requests of the service interface is not less than the first threshold, the control device will quickly feedback and divert. This solves the problem that the monitoring system in the prior art only issues an alarm after detecting an error, resulting in losses in the process of the operation and maintenance personnel handling the abnormality.
[0125] Figure 8An exemplary system architecture 800 is shown to which the distributed cluster control method or distributed cluster control apparatus according to the embodiments of the present invention can be applied.
[0126] like Figure 8 As shown, system architecture 800 may include terminal devices 801, 802, 803, a network 804, and a server 805. Network 804 is used to provide a medium for communication links between terminal devices 801, 802, 803 and server 805. Network 804 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0127] Users can use terminal devices 801, 802, and 803 to interact with server 805 through network 804 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 801, 802, and 803, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0128] The terminal devices 801 , 802 , and 803 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0129] The server 805 may be a server that provides various services, such as a backend management server that provides support for shopping websites browsed by users using the terminal devices 801, 802, and 803. The backend management server may analyze and process the received data such as product information query requests, and feed back the processing results (such as target push information and product information) to the terminal device.
[0130] It should be noted that the distributed cluster control method provided in the embodiment of the present invention is generally executed by the server 805 , and accordingly, the distributed cluster control device is generally disposed in the server 805 .
[0131] It should be understood that Figure 8 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0132] Reference below Fig. 9 , which shows a schematic diagram of the structure of a computer system 900 of a terminal device suitable for implementing an embodiment of the present invention. Fig. 9 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0133] like Fig. 9As shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the system 900 are also stored. The CPU 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0134] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed, so that a computer program read therefrom is installed into the storage section 908 as needed.
[0135] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above-mentioned functions defined in the system of the present invention are executed.
[0136] It should be noted that the computer-readable medium shown in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0137] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0138] The modules involved in the embodiments of the present invention may be implemented by software or hardware. The modules described may also be set in a processor. For example, they may be described as follows: a processor includes a request receiving module, a host determination module, an interface selection module, and a processing module. The names of these modules do not, in some cases, constitute limitations on the modules themselves. For example, the request receiving module may also be described as a "module for receiving service requests."
[0139] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiment; or may exist independently and not assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device includes: receiving a service request, the service request indicates a target service system and a requested service type; determining one or more first hosts corresponding to the target service system and in a normal state from the distributed cluster; wherein the distributed cluster includes one or more service systems, the service systems correspond to one or more hosts, and the hosts correspond to one or more service interfaces for providing services; when it is determined that the first host exists, according to the number of abnormal requests of the service interface corresponding to the service type in the first host, selecting a target service interface whose number of abnormal requests is less than a first threshold for the service request; wherein the number of abnormal requests indicates the number of historical abnormalities of the service interface; sending the service request to a second host corresponding to the target service interface in the first host, so that the second host provides a service corresponding to the service type through the target service interface.
[0140] According to the technical solution of the embodiment of the present invention, when a service request is received, one or more hosts corresponding to the target service system indicated by the service request and in a normal state are determined from the distributed cluster, and then according to the number of abnormal requests of the service interface corresponding to the service type indicated by the service request in the first host, a target service interface with an abnormal request number less than a first threshold is selected for the service request, and then the service request is sent to the second host corresponding to the target service interface, so that the second host provides the service corresponding to the service request through the target service interface, thereby diverting the service request to the service interface that can provide the service normally, and ensuring that the service interface in an abnormal state will not receive the service request, so that the service request can be successfully processed. Thus, according to the monitoring of both the host and the service interface, it is ensured that when the host is in a normal state, the abnormal situation of some service interfaces can be quickly perceived, which solves the problem of service request failure caused by the host being in a normal state but the service interface being abnormal, realizes fine-grained monitoring of the distributed cluster, improves the system stability of the distributed cluster, and avoids economic losses caused by failure to monitor the abnormality of the service system in a timely manner.
[0141] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A distributed cluster control method, characterized in that: receiving a service request, the service request indicating a target service system and a requested service type; Determine one or more first hosts corresponding to the target service system and in a normal state from the distributed cluster; wherein the distributed cluster includes one or more service systems, the service systems correspond to one or more hosts, and the hosts correspond to one or more service interfaces for providing services; In the case where it is determined that the first host exists, according to the number of abnormal requests of the service interface corresponding to the service type in the first host, selecting a target service interface whose number of abnormal requests is less than a first threshold for the service request; wherein the number of abnormal requests indicates the number of historical times when the service interface has abnormalities; Sending the service request to a second host corresponding to the target service interface in the first host, so that the second host provides a service corresponding to the service type through the target service interface; When the number of the first hosts determined is multiple, the following steps are executed in a loop until the number of the second hosts or the number of the polled first hosts is greater than a second threshold: Determine the number of abnormal requests of the service interface corresponding to the service type among the service interfaces corresponding to the current host among the plurality of first hosts, and determine whether the number of abnormal requests is less than a first threshold; If yes, taking the current host as the second host, and taking the service interface corresponding to the service type as the target service interface; If not, a first host that has not been selected is selected from the plurality of hosts as the current host.
2. The method according to claim 1, characterized in that Also includes: Obtaining a request result regarding the service request returned by the second host; When the request result is abnormal, the number of abnormal requests of the target service interface in the second host is incremented.
3. The method according to claim 1, characterized in that: The determining, from the distributed cluster, one or more first hosts corresponding to the target service system and in a normal state includes: Receive the configuration information of the host in the distributed cluster that is in a normal state and the service interface corresponding to the host, and form a distribution map corresponding to the distributed cluster according to the configuration information; wherein the configuration information of the host indicates the service system corresponding to the host; The first host is determined according to the distribution map.
4. The method according to claim 1, characterized in that Also includes: Using a successful historical service request to call the service interface whose number of abnormal requests is not less than the first threshold, when the call is successful, decrementing the number of abnormal requests of the service interface whose number of abnormal requests is not less than the first threshold; When the call fails and the number of call failures is greater than a third threshold, it is determined that the service interface whose number of abnormal requests is not less than the first threshold is in an abnormal state.
5. The method according to any one of claims 1 to 4, characterized in that: Also includes: In the absence of the first host or the service interface being in an abnormal state, an alarm message is output.
6. A distributed cluster control device, characterized in that: include: request receiving module, host determination module, interface selection module and processing module; wherein, The request receiving module is used to receive a service request, wherein the service request indicates a target service system and a requested service type; The host determination module is used to determine one or more first hosts corresponding to the target service system and in a normal state from the distributed cluster; wherein the distributed cluster includes one or more service systems, the service systems correspond to one or more hosts, and the hosts correspond to one or more service interfaces for providing services; The interface selection module is configured to select, when it is determined that the first host exists, a target service interface whose number of abnormal requests is less than a first threshold for the service request according to the number of abnormal requests of the service interface corresponding to the service type in the first host; wherein the number of abnormal requests indicates the number of historical times when the service interface has been abnormal; The processing module is used to send the service request to a second host corresponding to the target service interface in the first host, so that the second host provides a service corresponding to the service type through the target service interface; When the number of the first hosts determined is multiple, the processing module is used to loop through the following steps until the number of the second hosts or the polled first hosts is determined to be greater than a second threshold: Determine the number of abnormal requests of the service interface corresponding to the service type among the service interfaces corresponding to the current host among the plurality of first hosts, and determine whether the number of abnormal requests is less than a first threshold; If yes, taking the current host as the second host, and taking the service interface corresponding to the service type as the target service interface; If not, a first host that has not been selected is selected from the plurality of hosts as the current host.
7. The device according to claim 6, characterized in that The processing module is further configured to obtain a request result regarding the service request returned by the second host, and if the request result is abnormal, increment the number of abnormal requests of the target service interface in the second host.
8. The device according to claim 6, characterized in that Also includes: Configuration module; where, The configuration module is used to receive configuration information of a host in a normal state in the distributed cluster and a service interface corresponding to the host, and form a distribution map corresponding to the distributed cluster according to the configuration information; wherein the configuration information of the host indicates the service system corresponding to the host.
9. A server, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
10. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Distributed task scheduling operation and maintenance monitoring system and method
CN106484530A
Distributed scheduling automated testing platform and method
CN106844198A