Load balancing method, device and product
By using sliding window technology to dynamically monitor and fault analysis of service instances in distributed systems, dynamically adjusting load balancing, the problem of inability to adjust the status of service instances in the existing technology is solved, and the stability and reliability of the system are improved.
Patent Information
- Application Number
- CN202510361506.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-22
AI Technical Summary
The existing load balancing technology cannot adjust the working status of service instances in the server cluster in a timely manner, resulting in the distributed system being unable to perform load balancing adjustments in time when service instances are abnormal, affecting system stability and reliability.
By obtaining the working status information of the service instances in the distributed system, using sliding window technology to intercept the status information to be analyzed, performing fault analysis, dividing time periods and assigning weight values, determining the target weight value based on the fault analysis results, dynamically adjusting the load balancing task, and giving priority to selecting service instances with high stability for load balancing.
Real-time load balancing adjustment of service instances is realized, the stability and reliability of distributed systems are improved, traffic is allocated to service instances with low failure frequency, and overall system performance is improved.
Smart Images

Figure CN120358241A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to a load balancing method, device, and product. Background Art
[0002] To ensure the stable operation of backend servers, the server cluster structure is designed as a distributed deployment structure.
[0003] In a distributed system, to improve system performance and enhance reliability, a load balancing technology is adopted to achieve the stable and reliable operation of the distributed system. In the existing load balancing technology, load balancing adjustment is usually performed according to traffic after startup. However, the working status of each server and service instance is not understood, and adjustment cannot be made in a timely manner when a service instance has an exception. Therefore, a solution that can achieve timely and accurate load balancing adjustment is needed. Summary of the Invention
[0004] The present disclosure provides a load balancing method, device, and product.
[0005] According to a first aspect of the present disclosure, a load balancing processing method is provided. The method is applied to a server, and specifically includes: obtaining the working status information of multiple service instances in a distributed system; where the working status information includes: normal service response and / or abnormal service response; intercepting the status information to be analyzed from the working status information by using a sliding window with a specified time length; and performing a fault analysis by using the number of normal service responses and / or the number of abnormal service responses in the status information to be analyzed to obtain a fault analysis result; dividing the status information to be analyzed in the sliding window into multiple time periods, and assigning a weight value to each time period; determining a target weight value of each service instance from multiple weight values based on the fault analysis result of the status information to be analyzed; and determining a target service instance for performing a load balancing task among multiple service instances according to each target weight value.
[0006] Based on the above, by dynamically monitoring the working status of service instances. During the monitoring, the working status information of service instances in multiple servers is collected by using the sliding window technology, and the probability of service exception occurrence for each service instance is statistically analyzed. In this way, the situation of service instance exception can be found in a timely manner, and dynamic adjustment can be made according to the real-time service request status, and the traffic of service requests can be adjusted to service instances with a low recent failure frequency (that is, a high target weight value). Thus, real-time load balancing adjustment for service instances is achieved, and the overall stability of the distributed system is improved.
[0007] According to at least one embodiment of the present disclosure, determining a target service instance for performing a load balancing task among a plurality of service instances according to respective target weight values includes: summing the target weight values of the service instances in the sliding window to obtain a total weight value; and determining, in a total weight range constructed based on the total weight value, a target service instance for performing the load balancing task by using a random number.
[0008] According to at least one embodiment of the present disclosure, determining, in a total weight range constructed based on the total weight value, a target service instance for performing a load balancing task by using a random number includes: dividing, according to the target weight value corresponding to the service instance, a corresponding sub-range in the total weight range; determining a random number by using a random number algorithm, where the random number is an integer greater than zero and less than or equal to the total weight value; determining the target service instance corresponding to the sub-range containing the random number; and performing the load balancing task by using the target service instance.
[0009] According to at least one embodiment of the present disclosure, allocating a plurality of weight values to a sliding window includes: dividing the status information to be analyzed in the sliding window into a plurality of time periods, and allocating a weight value to each time period.
[0010] According to at least one embodiment of the present disclosure, dividing the status information to be analyzed in the sliding window into a plurality of time periods includes: sorting the status information to be analyzed in reverse order according to the generation order; and evenly dividing the status information to be analyzed after reverse sorting into a plurality of time periods according to a specified time interval.
[0011] According to at least one embodiment of the present disclosure, allocating a weight value to each time period includes: allocating a first weight value to a first time period and allocating a second weight value to a second time period, where the first time period is closer to the current moment than the second time period, and the first weight value is less than the second weight value.
[0012] According to at least one embodiment of the present disclosure, determining a target weight value of a service instance from a plurality of weight values based on a fault analysis result of the status information to be analyzed includes: counting a first number of anomalies of abnormal service responses in the status information to be analyzed, and counting a total number of service responses of normal service responses and abnormal service responses in the first time period; calculating a first anomaly rate based on a ratio of the first number of anomalies and the total number of service responses; comparing a magnitude relationship between the first anomaly rate and a fault threshold; and when a comparison result of the magnitude relationship is that the first anomaly rate is greater than the fault threshold, determining a fault weight of the service instance as the target weight value corresponding to the first time period.
[0013] According to at least one embodiment of the present disclosure, it further includes: when the comparison result of the magnitude relationship is that the first abnormal rate is not greater than the failure threshold, calculating the second abnormal rate within the second time period; comparing the magnitude relationship between the second abnormal rate and the failure threshold; when the comparison result of the magnitude relationship is that the second abnormal rate is greater than the failure threshold, determining the failure weight of the service instance as the target weight value corresponding to the second time period; when the comparison result of the magnitude relationship is that the second abnormal rate is greater than the failure threshold, and the second time period is the last time period in the sliding window, determining the failure weight of the service instance as the default weight value; wherein, the default weight value is greater than the target weight value.
[0014] According to at least one embodiment of the present disclosure, the method for statistically analyzing the first abnormal number of service response anomalies in the status information to be analyzed includes: respectively statistically analyzing the abnormal sub-numbers corresponding to each time period in the status information to be analyzed; respectively determining the corresponding magnification factors according to the chronological order of the time periods; and summing the products of the magnification factors and the abnormal sub-numbers to obtain the first abnormal number.
[0015] According to at least one embodiment of the present disclosure, the method for intercepting the status information to be analyzed from the working status information by using a sliding window with a specified time length includes: determining the service instance as the key; intercepting the status information to be analyzed from the working status information by using a sliding window with a specified time length as the value; and storing the service instance and the sliding window in the hash table in the form of a key-value pair.
[0016] According to a second aspect of the present disclosure, there is provided an electronic device, including: a memory storing execution instructions; and a processor that executes the execution instructions stored in the memory, such that the processor executes the method according to the first aspect of any one of the embodiments of the present disclosure.
[0017] According to a third aspect of the present disclosure, there is provided a readable storage medium storing execution instructions, and when the execution instructions are executed by a processor, they are used to implement the method according to the first aspect of any one of the embodiments of the present disclosure.
[0018] According to a fourth aspect of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, it implements the method according to the first aspect of any one of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.
[0020] Figure 1 It is a schematic diagram of the overall process of the load balancing method according to an embodiment of the present disclosure.
[0021] Figure 2 Flow schematic diagram of the weight allocation method provided by the present disclosure.
[0022] Figure 3 Schematic diagram of weight allocation illustrated by the present disclosure.
[0023] Figure 4 Schematic diagram of the sliding window illustrated by the present disclosure.
[0024] Figure 5 Schematic diagram of abnormal rate statistics illustrated by the present disclosure.
[0025] Figure 6 Flow schematic diagram of the method for determining a target instance provided by the present disclosure.
[0026] Figure 7 Schematic diagram of multiple service instances in the sliding window illustrated by the present disclosure.
[0027] Figure 8 Structural schematic block diagram of a load balancing device according to an embodiment of the present disclosure.
[0028] Figure 9 Structural schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0029] The present disclosure will be further described in detail below with reference to the accompanying drawings and examples. It can be understood that the specific examples described herein are only for explaining the relevant content, rather than limiting the present disclosure. Additionally, it should be noted that for the sake of convenience of description, only parts related to the present disclosure are shown in the accompanying drawings.
[0030] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the accompanying drawings and embodiments.
[0031] With the development of computer technology, distributed systems have been widely applied. For example, distributed architectures are adopted in cloud computing scenarios, Internet of Things scenarios, and big data processing scenarios. In order to enable distributed systems to work more stably and exhibit better performance, load balancing technologies are used to help coordinate tasks. However, when implementing existing load balancing solutions, a traffic warm-up strategy is usually adopted. In the case of service instance failures, traditional load balancing solutions cannot discover and coordinate the traffic of each service instance. Therefore, a solution that can achieve timely and accurate load balancing adjustment of service instances is needed.
[0032] Term explanation.
[0033] For ease of description and to make the technical solutions of the specific embodiments of the present disclosure easier to understand, before describing the load balancing method implemented in the present disclosure, the technical terms involved in the specific embodiments of the present disclosure are explained as follows.
[0034] Distributed system: It is composed of multiple computers, which are geographically dispersed and can be spread within an organization, a city, a country, or even globally. They are interconnected through a communication network and cooperate to complete tasks. During the task execution process, if a server malfunctions, the task can be adjusted to other servers for execution. This ensures that the distributed system can execute tasks stably.
[0035] Service instance: It refers to a specific running entity that can provide a specific service function. Each service instance is usually a separately deployed copy of an application, which can be a process or a group of processes in a physical server, a virtual machine, a container (such as a Docker container), or even an execution unit in a serverless architecture (such as an AWS Lambda function). These instances work together to process requests from clients. Specifically, when implementing load balancing, multiple service instances are configured as a cluster or a set of accessible service endpoints. A load balancer (which can be a hardware device or a software component) is located between the client and these service instances, and its main responsibility is to distribute incoming requests to different service instances according to a certain policy or algorithm (such as round-robin, least connections, hashing, etc.). This not only improves the availability and reliability of the system but also effectively allocates resources and avoids a single service instance becoming a performance bottleneck.
[0036] Sliding Window: It is a data processing and analysis technique, usually used to create a movable window with a fixed size (or dynamically adjusted according to conditions) in a series of data. It moves on the data sequence, processes a certain number of data points each time, and then moves to the next position to continue processing the next set of data points for various calculation and analysis operations.
[0037] Load balancing: It is used to distribute the traffic of a network or an application to multiple servers or resources to optimize the response time, maximize the throughput, reduce the latency, improve the scalability and reliability of the system. It ensures that each server or resource can process requests in an efficient and balanced manner and avoids single-point overload.
[0038] Figure 1 It is a schematic diagram of the overall process of the load balancing method according to an embodiment of the present disclosure. As Figure 1 shown, the method includes steps 101 to 105. Among them, the method can be executed by an electronic device such as a server (local server or cloud server).
[0039] Specifically, Figure 1 The method shown includes: Step 101: Obtain the working status information of multiple service instances in a distributed system; wherein, the working status information includes: normal service response and / or abnormal service response.
[0040] In a distributed system, many servers are distributed (for example, an Elastic Compute Service (ECS) is a server that provides scalable computing power), and one or more service instances are running on each server.
[0041] These service instances can cooperate with each other. When a certain service instance has a running failure, the load balancing technology means can be used to coordinate and transfer the related tasks of this service instance to other service instances with surplus computing power to continue running. Thus, the stable operation of the distributed system can be ensured.
[0042] The working status information mentioned here can be understood as using a performance monitoring system to collect the working status information of each service instance, and can capture and analyze request response data in real time, including key indicators such as response time, throughput, error rate, etc.
[0043] In addition, a logging application can also be obtained, and a logging management tool is used to parse the generated log files (such as ELK Stack, Splunk, etc.) to collect, store, and analyze log data. These logging management tools provide powerful query, search, and visualization functions, which help to quickly collect which services have normal responses and which services have abnormal responses.
[0044] Abnormal service response can be various response error situations such as service non-response or service response timeout resulting in failure. Normal service response is to be able to correctly execute a service request.
[0045] It should be noted that when collecting the working status information of multiple service instances, it is necessary to ensure that the multiple service instances have a unified time correspondence relationship. For example, use the same time identification tool to mark timestamps for the working status information of multiple service instances. So as to have a unified time standard when calculating the exception rate and determining the target weight value subsequently.
[0046] Step 102: Use a sliding window with a specified time length to intercept the status information to be analyzed from the working status information of each service instance; and use the number of normal service responses and / or the number of abnormal service responses in the status information to be analyzed for fault analysis to obtain a fault analysis result.
[0047] The working status information is a set of consecutive historical information of multiple service instances. When selecting the working status information required for fault analysis from the historical working status information, a sliding window method can be used to intercept a part of the useful status information to be analyzed from the working status information. The sliding window slides in real time according to the status information generation speed, ensuring that the information intercepted by each sliding window is the latest (closest to the current moment) status information. The fault status of the service instance is judged using the status information to be analyzed within the sliding window.
[0048] During the analysis of multiple service instances, once the specified time length of the sliding window is determined, it will not be adjusted or changed anymore. This specified time length can be determined by the user according to the general fault recovery time of the service instance and can be an empirical value.
[0049] Step 103: Divide the status information to be analyzed in the sliding window into multiple time periods and assign weight values to each time period.
[0050] When assigning weight values to the sliding window, it can be assigned according to time periods (specifically, refer to the relevant embodiments in the following text), or it can be assigned according to the number of anomalies within the window without dividing time periods. For example, the higher the number of anomalies, the smaller the corresponding weight value. After analyzing the status information to be analyzed within the sliding window, according to different analysis results, determine the weight value corresponding to this sliding window. In the subsequent embodiments, examples will be given for the implementation method of assigning weight values to the sliding window, and it will not be repeated here.
[0051] Step 104: Based on the fault analysis results of the status information to be analyzed, determine the target weight value of each service instance from multiple weight values.
[0052] In practical applications, after obtaining the status information to be analyzed, the number of times of normal service response and the number of times of abnormal service response in it will be statistically analyzed. According to the statistical analysis results of the service response, determine the target weight value corresponding to each service instance within this sliding window.
[0053] It should be noted that the status information to be analyzed of multiple service instances is intercepted simultaneously in a sliding window. When analyzing the status information to be analyzed, it is necessary to distinguish and analyze different service instances separately; in other words, each service instance uses its own status information to be analyzed to determine its own target weight value.
[0054] The target weight value mentioned here can be understood as a weight value selected by a service instance from multiple pre-assigned weight values according to the fault analysis results within a certain sliding window time range. The selection principle is that the worse the stability of the service instance, the smaller the selected target weight value; the better the stability of the service instance, the larger the selected target weight value.
[0055] Step 105: Determine the target service instances for performing load balancing tasks among multiple service instances according to the respective target weight values.
[0056] After determining the respective target weight values corresponding to each service instance in the manner described above, the load balancing task is performed according to the magnitudes of the target weight values of each service instance. Simply put, the service instance with a larger target weight value is preferentially selected as the target service instance for performing the load balancing task.
[0057] The worse the fault analysis result of the service instance (that is, if a fault has occurred recently or faults occur frequently, the fault analysis result is considered bad), the smaller the corresponding target weight value; conversely, the better the fault analysis result (that is, if no fault has occurred recently or faults occur at a low frequency, the fault analysis result is considered good), the larger the corresponding target weight value.
[0058] When performing the load balancing task, traffic will be preferentially allocated to the service instance with a larger target weight value. The larger the target weight value, the more stable the service instance runs, the lower the probability of failure, and the more request tasks it can receive and process.
[0059] Conversely, the service instance with a smaller target weight value is allocated as few request tasks as possible, or the probability of distributing request tasks to the service instance with a smaller target weight value is reduced. Therefore, the service instance with a larger target weight value can be directly selected as the target service instance to perform the load balancing task after comparing the magnitude relationship of the target weight values.
[0060] Based on the above embodiments, the working status of the service instance is dynamically monitored. During the monitoring, the sliding window technology is used to collect the working status information of the instances in multiple servers, and when it is found that a service exception occurs in the service instance, it can be dynamically adjusted according to the real-time service request status, and the traffic is adjusted to the service instance with a low recent fault frequency. Thus, real-time load balancing adjustment for the service instance is achieved.
[0061] In one or more embodiments of the present disclosure, a weight allocation scheme is further provided. As Figure 2 is a schematic flowchart of the weight allocation method provided by the present disclosure. As can be seen from Figure 2 it, multiple weight values are assigned to the sliding window, including: Step 201: Divide the status information to be analyzed in the sliding window into multiple time periods. Step 202: And assign weight values to each time period.
[0062] When dividing into multiple time periods, there is no limit on the number and length of the divided time periods. However, generally, the duration of the divided time period is the duration required to eliminate most service response exceptions.
[0063] When assigning weight values to each time period, the time period closer to the current moment is assigned a smaller weight value, and the time period farther from the current moment is assigned a larger weight value.
[0064] Such as Figure 3 is a schematic diagram of weight allocation illustrated for this disclosure. From Figure 3 it can be seen that assuming the specified time length of the sliding window is 60 seconds, when dividing the time period, it is divided into segments of 15 seconds each. That is, the sliding window is divided into 4 segments, namely the recent 15S, the recent 30S, the recent 45S, and the recent 60S. Among them, the recent 15S refers to the time range from the current moment to the 15th second, the recent 30S refers to the time range from the 15th second to the 30th second, the recent 45S refers to the time range from the 30th second to the 45th second, and the recent 60S refers to the time range from the 45th second to the 60th second.
[0065] When assigning weight values, the weight value assigned to the recent 15S is 1, the weight value assigned to the recent 30S is 4, the weight value assigned to the recent 45S is 16, and the weight value assigned to the recent 60S is 64. Here, it is assumed that the default weight value is 100. Of course, this is only an example. In actual applications, the staff can set the size of the weight value and the size of the default weight value according to actual needs. However, it should be noted that it is necessary to ensure that the default weight value is significantly greater than the assigned weight value.
[0066] It should be noted that in the above-mentioned time period division scheme, the number of service requests in each time period is not necessarily exactly the same. For example, service requests are frequently responded to in the first 15S time period, and there are no service requests in the last 60S time period. Therefore, when marking the weight value, it is marked according to the size of the time difference from the current moment. The smaller the time difference from the current moment for the time period when a service response exception occurs, the more unstable the working state of the current service instance is. On the contrary, after a service response exception occurs in the time period with a larger time difference from the current moment, the impact on the working state of the current service instance is smaller, and there is a possibility that the current service instance reaches a stable working state. Therefore, when assigning weight values, the time period closer to the current moment (that is, the time period with a smaller time difference) is assigned a smaller weight value to reduce the probability of this service instance being selected. On the contrary, the time period farther from the current moment (that is, the time period with a larger time difference) is assigned a larger weight value to increase the probability of this service instance being selected.
[0067] Through the above-mentioned disclosed scheme, the status information to be analyzed in the sliding window is divided into multiple time periods, and corresponding weight values are assigned in chronological order, so as to more truly reflect the service status of the service instance through the weight values and provide a basis for subsequent load balancing.
[0068] As described in step 201, the status information to be analyzed in the sliding window is divided into multiple time periods, including: Step 2011: Reverse sort the status information to be analyzed according to the generation order. Step 2012: Divide the reverse-sorted status information to be analyzed into multiple time periods at specified time intervals.
[0069] As Figure 3 shown, the sliding window divides the status information to be analyzed into 4 segments in the reverse sorting order of 60S, 45S, 30S, and 15S, and the specified time interval for each time period is 15S. Through this reverse sorting method, it is convenient for subsequent division of time periods for newly added status information to be analyzed and the sliding window. As Figure 4 is the schematic diagram of the sliding window illustrated in this disclosure. As can be seen from Figure 4 it, after adding 1 15S, another time period is added. And for the 60S corresponding to the previous time period ①, it becomes 75S after adding 15S; for the 45S corresponding to the previous time period ②, it becomes 60S after adding 15S; for the 30S corresponding to the previous time period ③, it remains 30S after adding 15S; for the 15S corresponding to the previous time period ④, it becomes 30S after adding 15S; the newly added 15S is time period ⑤. The sliding window also moves to time periods ② to ⑤. In this way, after adding, it is ensured that the status information to be analyzed intercepted by the sliding window is the latest working status information.
[0070] Of course, the reverse order method may not be adopted. For example, sort in ascending order according to the generation order. The sliding direction of the corresponding sliding window is also adjusted accordingly to ensure that the sliding direction of the sliding window is always consistent with the direction of the newly added time period.
[0071] As described in step 202, weight values are assigned to each time period, including: Step 2021: Assign a first weight value to the first time period. Step 2022: And assign a second weight value to the second time period; where the first time period is closer to the current moment than the second time period, and the first weight value is less than the second weight value.
[0072] It should be noted that the status information to be analyzed in the sliding window can be divided into multiple time periods. The first time period and the second time period mentioned here are any two time periods among the divided multiple time periods. When assigning weight values to time periods, different weight values need to be assigned according to the order of time periods. The closer the time period is to the current moment, the smaller the assigned weight value, and the farther the time period is from the current moment, the larger the assigned weight value. For example, as Figure 3 shown, assume that the first time period is time period ④ and the second time period is time period ③. It can be seen that the weight value of the first time period is 1 and the weight value of the second time period is 4.
[0073] As can be seen from the above disclosed solution, the reason for allocating corresponding weight values according to the time difference between the time period and the current moment is that the closer the time period is to the current moment, the greater the impact on the current situation after a service response anomaly occurs, and the farther the time period is from the current moment, the smaller the impact on the current situation after a service response anomaly occurs. Through the above allocation scheme, it is possible to better assist in selecting a better (with a lower probability of service response anomaly) service instance to perform the load balancing task.
[0074] In an alternative solution, when allocating multiple weight values to the sliding window, it is also possible not to divide into multiple time periods, but according to the proportional range of service response anomalies or the quantity range of service response anomalies. For example, it is set that when the number of service response anomalies is from 1 to 5, the corresponding weight value is 64; when the number of service response anomalies is from 5 to 10, the corresponding weight value is 16; when the number of service response anomalies is from 10 to 20, the corresponding weight value is 4; when the number of service response anomalies is greater than 20, the corresponding weight value is 1; if there is no service response anomaly, the weight value is the default value of 100. Thus, the more times a service instance has a service response anomaly, the more unstable the service instance is, and the smaller the corresponding weight value. On the contrary, the fewer times a service instance has a service response anomaly, the more stable the service instance is, and the greater the corresponding weight value, which also means that this service instance will be preferentially selected to perform the load balancing task.
[0075] Or, after dividing the sliding window into multiple time periods (that is, dividing the status information to be analyzed in the sliding window into multiple time periods), count the number of anomalies occurring in each time period. Allocate corresponding weight values according to the number of service response anomalies in each time period. Generally speaking, the more times of service response anomalies, the greater the corresponding weight value; on the contrary, the fewer times of service response anomalies, the smaller the corresponding weight value.
[0076] In one or more embodiments of the present disclosure, based on the fault analysis result of the status information to be analyzed, determining the target weight value of the service instance from multiple weight values includes: counting the first number of anomalies of service response in the status information to be analyzed, and counting the total number of service responses of normal service response and service response anomaly in the first time period; calculating the first anomaly rate based on the ratio of the first number of anomalies and the total number of service responses; comparing the size relationship between the first anomaly rate and the fault threshold; when the comparison result of the size relationship is that the first anomaly rate is greater than the fault threshold, determining the fault weight of the service instance as the target weight value corresponding to the first time period.
[0077] As Figure 5 is the schematic diagram of anomaly rate statistics illustrated for the present disclosure. Taking a service instance in the sliding window as an example, from Figure 5As can be seen, the status information to be analyzed in the sliding window is divided into 4 time periods, namely the recent 60S time period ①, the recent 45S time period ②, the recent 30S time period ③, and the recent 15S time period ④.
[0078] Taking the calculation of the first exception rate corresponding to the first time period as an example, the following is an elaboration. Here, it is assumed that the first time period is the recent 15S time period ④ (of course, it can also be other time periods in the same sliding window such as the recent 30S time period ③). Since when calculating the exception rate, if there are multiple time periods, the polling calculation method is adopted, and the calculation processes are similar.
[0079] Here, it is assumed that within the recent 60S time period ① (which can be called the fourth time period), the number of abnormal sub - times A1 of service response exception and the number of normal service response times B1. Within the recent 45S time period ② (which can be called the third time period), the number of abnormal sub - times A2 of service response exception and the number of normal service response times B2. Within the recent 30S time period ③ (which can be called the second time period), the number of abnormal sub - times A3 of service response exception and the number of normal service response times B3. Within the recent 15S time period ④ (which can be called the first time period), the number of abnormal sub - times A4 of service response exception and the number of normal service response times B4.
[0080] When calculating the first time period (that is, within the recent 15S), the first exception number of service response exception in the status information to be analyzed refers to the sum of all service response exception numbers in the 4 time periods of the sliding window. Then the first exception number A = A1 + A2 + A3 + A4.
[0081] When counting the total number of service responses C of normal service response and service response exception within the first time period, C = A4 + B4.
[0082] Therefore, the calculated first exception rate D = A / C = (A1 + A2 + A3 + A4) / (A4 + B4).
[0083] Assume that the fault threshold is 1. Next, compare the size relationship between the first exception rate and the fault threshold. If the comparison result shows that the first exception rate is greater than the fault threshold, then the weight value corresponding to the first time period is used as the target weight value of this sliding window.
[0084] If the comparison result shows that the first exception rate is not greater than the fault threshold, it is necessary to calculate the exception rate for the next time period. The specific method is as follows. When the comparison result of the magnitude relationship is that the first exception rate is not greater than the fault threshold, calculate the second exception rate within the second time period; compare the magnitude relationship between the second exception rate and the fault threshold; when the comparison result of the magnitude relationship is that the second exception rate is greater than the fault threshold, determine that the fault weight of the service instance is the target weight value corresponding to the second time period; when the comparison result of the magnitude relationship is that the second exception rate is greater than the fault threshold, and the second time period is the last time period in the sliding window, determine that the fault weight of the service instance is the default weight value; where the default weight value is greater than the target weight value.
[0085] When calculating the exception rate, the polling calculation method is adopted. That is, when the first exception rate is not greater than the fault threshold, next calculate the second exception rate for the next 30S time period ③ (i.e., the second time period) until a time period with an exception rate greater than the fault threshold is calculated, and use the weight value corresponding to this time period as the target weight value of this sliding window.
[0086] If the second time period is not the last time period in the sliding window, when the calculated second exception rate is less than the fault threshold, the next time period will be calculated in the polling order.
[0087] If the second time period is the last time period in the sliding window (such as the Figure 5 in the last 60S time period ①), if the calculation result shows that the exception rates of all time periods in this sliding window are not greater than the fault threshold, the default weight value will be used as the target weight value of this sliding window. In this case, it means that the service response status of this service instance is very good within the sliding window, no service response exception occurs, or the probability of service response exception is very low and does not affect the normal operation of the service instance.
[0088] It should be noted that when calculating the exception rate of other time periods, the number of exceptions A is the sum of the number of exceptions corresponding to multiple time periods before the time period for which the exception rate needs to be calculated within the sliding window. Continuing to assume that the second time period is the last time period of the sliding window, then the number of exceptions is A1, and the total number of service responses C = A1 + B1. Therefore, the calculated exception rate D = A / C = A1 / (A1 + B1).
[0089] The above calculation process is illustrated with one service instance. The calculation process of the target weight value of other service instances within the sliding window is similar to the above content and will not be elaborated here. For details, please refer to the above embodiments.
[0090] In the above manner, calculate the abnormality rates for each time period in the sliding window, and determine a target weight value that can represent the stability of the service instance within the sliding window based on the relationship between the abnormality rate and the fault threshold. So as to perform load balancing on the service instance according to the target weight value.
[0091] In an alternative solution, to improve the perception ability of the abnormality rate, the above calculation method can be further optimized. Count the first abnormal times of service response abnormalities in the status information to be analyzed, including: separately count the abnormal sub-times corresponding to each time period in the status information to be analyzed; determine the corresponding magnification factors according to the chronological order of the time periods; and sum the products of the magnification factors and the abnormal sub-times to obtain the first abnormal times.
[0092] The abnormal sub-times mentioned here refer to the number of service response abnormalities occurring within each time period. When setting the magnification factors for each time period, the corresponding magnification factors can be set in the chronological order of the time periods, and the magnification factors corresponding to different time periods are different.
[0093] Continue to assume that within the 60S time period ① (which can be called the fourth time period), the abnormal sub-times A1 of service response abnormalities and the normal times B1 of service response are counted. Within the 45S time period ② (which can be called the third time period), the abnormal sub-times A2 of service response abnormalities and the normal times B2 of service response are counted. Within the 30S time period ③ (which can be called the second time period), the abnormal sub-times A3 of service response abnormalities and the normal times B3 of service response are counted. Within the 15S time period ④ (which can be called the first time period), the abnormal sub-times A4 of service response abnormalities and the normal times B4 of service response are counted.
[0094] Assume that the magnification factor for the 60S time period ① is 64, the magnification factor for the 45S time period ② is 16, the magnification factor for the 30S time period ③ is 4, and the magnification factor for the 15S time period ④ is 1. Here, the magnification factors are set in a stepped manner according to the chronological order of the time periods. In actual applications, users can choose an equivalent setting according to actual needs, that is, the same magnification factor can be set for all time periods in the sliding window.
[0095] When calculating the first time period (that is, within the recent 15S), the first abnormal times of service response abnormalities in the status information to be analyzed refer to the sum of all service response abnormal times in the 4 time periods of the sliding window. Then the first abnormal times A = 64 * A1 + 16 * A2 + 4 * A3 + A4.
[0096] In the first time period, the total number of service responses C for normal and abnormal service responses is C = A4 + B4. It should be noted that the total number of service responses mentioned here refers to the total number of service requests in this time period, including the number of abnormal service responses and the number of normal service responses.
[0097] Therefore, the calculated first exception rate D = A / C = (64 * A1 + 16 * A2 + 4 * A3 + A4) / (A4 + B4).
[0098] Based on the above disclosed solution, when calculating the first number of exceptions, a preset magnification factor can be multiplied, so that the exception rate is more likely to be greater than the fault threshold. In terms of the usage effect, it makes the exception perception more sensitive.
[0099] Such as Figure 6 This is a schematic flow chart of the method for determining the target instance provided by the present disclosure. In one or more embodiments of the present disclosure, according to the target weight value, a target service instance for performing the load balancing task is selected from multiple service instances. The method specifically includes the following steps: Step 601: Sum the target weight values of the service instances in the sliding window to obtain the total weight value. Step 602: In the total weight interval constructed based on the total weight value, use a random number to determine the target service instance for performing the load balancing task.
[0100] When calculating the total weight value, the target weight values of all service instances in the sliding window (that is, the target weight values of the service instances determined in the above embodiments) are summed. The total weight value obtained after summation is used as the total weight interval. For example, if the total weight value is 300, then the total weight interval is from 1 to 300, and this total weight interval is used to select the target service instance for performing the load balancing. When using the random number algorithm to select a random number, it is necessary to ensure that the random number is any number within the range of the total weight interval (the random number is a number, not a numerical range or numerical interval). Whichever service instance's corresponding interval range this random number is in, that service instance is selected as the target service instance.
[0101] Such as Figure 7 This is a schematic diagram of multiple service instances in the sliding window illustrated by the present disclosure. From Figure 7As can be seen, there are a total of 5 service instances in the sliding window, namely service instance 1, service instance 2, service instance 3, service instance 4, and service instance 5. Four time periods are divided in the sliding window. Further assume that after the exception rates of the 5 service instances are statistically analyzed and evaluated, it is determined that the target weight value corresponding to service instance 1 in the current sliding window is 64, the target weight value corresponding to service instance 2 is 16, the target weight value corresponding to service instance 3 is 1, the target weight value corresponding to service instance 4 is 1, and the target weight value corresponding to service instance 5 is 100 (i.e., the default weight value). Next, sum up the target weight values. The total weight value = 64 + 16 + 1 + 1 + 100 = 182. That is, the total weight range is 182. When using a random number to select the target service instance, ensure that the random number is any number within the range of 1 to 182. The following will explain the specific implementation process of selecting the target service instance using a random number.
[0102] In one or more embodiments of the present disclosure, as described in step 602, in the total weight range constructed based on the total weight value, using a random number to determine the target service instance for performing the load balancing task includes: Step 6021: According to the target weight value corresponding to the service instance, divide the corresponding sub-range from the total weight range. Step 6022: Use a random number algorithm to determine a random number; where the random number is an integer greater than zero and less than or equal to the total weight value. Step 6023: Determine the target service instance corresponding to the sub-range containing the random number. Step 6024: Use the target service instance to perform the load balancing task.
[0103] For the convenience of understanding the total weight range, it is represented by a line segment in Figure 7 It should be noted that in the line segment of the total weight range shown in Figure 7 the order of each service instance can be adjusted arbitrarily. When dividing the corresponding sub-range from the line segment of the total weight range, divide it according to the target weight value of each service instance (although the order is not limited, but the size of the sub-range should be the same as the target weight value).
[0104] The probability that the random number falls into the sub-range of which service instance depends on the size of the target weight value of that service instance. The larger the target weight value, the larger the sub-range occupied by the service instance in the total weight range, and the greater the probability that the random number falls into the sub-range of that service instance.
[0105] As mentioned above, the larger the target weight value, the better the stability of the service instance within the range of this sliding window and the lower the failure rate, so it is more suitable to be selected as the target service instance for performing the load balancing task.
[0106] Based on the above - disclosed solution, after determining the total weight range, a random selection method using random numbers is employed to find the target service instance suitable for performing the load - balancing task. Through the above - mentioned method, service instances with good stability (large target weight values) have a greater probability of being selected, better meeting the load - balancing requirements.
[0107] In an alternative solution, in addition to selecting the target service instance using random numbers after obtaining the total weight value, the target weight values of multiple service instances can be sorted in ascending order. When performing the load - balancing task, the service instance with a larger target weight value is preferentially selected to perform the load - balancing task. During the process of performing the load - balancing task, the working - state information can be obtained in real - time and the fault - state analysis of each service instance can be carried out. According to the fault - analysis results, the load - balancing task can be dynamically adjusted. Thus, the load - balancing task is preferentially assigned to the target service instance with a larger target weight value.
[0108] In one or more embodiments of the present disclosure, a sliding window with a specified time length is used to intercept the state information to be analyzed from the working - state information, including: determining the service instance as the primary key; using the sliding window with the specified time length to intercept the state information to be analyzed from the working - state information as the value; storing the service instance and the sliding window in the form of a key - value pair in a hash table.
[0109] A large hash table is set up. The primary key of this large table is the service name, and the value is a small hash table for the state information to be analyzed in the sliding window. The primary key of the small hash table is the service instance, and the value is the state information in the sliding window.
[0110] In the large table, there are many services and their corresponding service names. The same service name can correspond to at least one small hash table. And in the small hash table, the primary key is the service instance, and the corresponding value is the sliding window. Here, the correspondence between the service instance and the sliding window is a one - to - many correspondence, that is, the same sliding window corresponds to multiple service instances participating in load balancing.
[0111] Based on the above - disclosed solution, after intercepting the state information to be analyzed using the sliding window, it can be stored in the corresponding hash table. This is convenient for subsequent load - balancing - related calculations, so as to quickly find the target service instance suitable for performing the load - balancing task.
[0112] For the sake of easy understanding, the implementation process of load balancing will be illustrated by a complete embodiment below.
[0113] First, collect the basic data. For the convenience of management, the relevant data of each service instance can be stored in a hash table. During subsequent calculations, the hash table can be called through an interface. The hash table includes: a large hash table, the primary key of which is the service name, and the value is a small hash table. The primary key of the small hash table is the service instance, and the value is a sliding window and the status information to be analyzed.
[0114] Sliding window: Each sliding window represents the relevant data of the service instance in the recent x minutes (such as 1 minute). The window is divided into y time periods (such as 4 time periods), representing the recent 15 seconds, recent 30 seconds, recent 45 seconds, and recent 60 seconds respectively. And the fault weights corresponding to each time period are (1, 4, 16, 64), and the default weight is 100. Generally, the default weight should be greater than the weights assigned to each time period.
[0115] Service instance fault evaluation criteria: For example, taking the request response timeout of each service instance as the exception index of the service instance, count the number of normal service responses and the number of abnormal service responses of the corresponding instance of the current interface, and put them into the corresponding time periods in the sliding window.
[0116] The fault evaluation algorithm is as follows: Polling window: Poll the four windows in reverse order from the current window, and execute the following two steps.
[0117] Calculate the exception rate: Input the total number of exceptions in the sliding window and the total number of requests in the specified time period. At the same time, divide the amplified number of exceptions (amplified by 4 times, 16 times, and 64 times respectively) by the total number to get the exception rate of the current time period. The purpose of amplification is to make the number of exceptions more sensitive to detection.
[0118] Next, compare the exception rate: According to the business-set fault threshold (default is 1), if the exception rate > the fault threshold, then return the weight value of the current time period as the target weight value; otherwise, continue to poll and calculate the exception rates of other remaining time periods in the sliding window until all time periods in the sliding window have been calculated and no exception rate greater than the fault threshold is found, then return the default value 100.
[0119] Load balancing algorithm: Statistically sum up the total target weights corresponding to all service instances in the sliding window, and then randomly poll a number within this total weight range. Whichever service instance's target weight range the random number falls into, that service instance is the finally selected target service instance.
[0120] Based on any of the above embodiments, the present disclosure also provides a load balancing device.
[0121] Figure 8 It is a structural schematic block diagram of the load balancing device according to an embodiment of the present disclosure.
[0122] As Figure 8As shown in the figure, the load balancing device includes: an acquisition module 81 that acquires the working status information of multiple service instances in a distributed system; where the working status information includes: normal service response and / or abnormal service response.
[0123] An interception module 82, which is used to intercept the status information to be analyzed from the working status information of each service instance by using a sliding window with a specified time length; and perform fault analysis by using the number of normal service responses and / or the number of abnormal service responses in the status information to be analyzed to obtain a fault analysis result.
[0124] An allocation module 83, which is used to divide the status information to be analyzed in the sliding window into multiple time periods and assign weight values to each time period.
[0125] A determination module 84, which is used to determine the target weight value of each service instance from multiple weight values based on the fault analysis result of the status information to be analyzed.
[0126] A selection module 85, which is used to determine the target service instance for performing the load balancing task among multiple service instances according to each target weight value.
[0127] A selection module 85, which is used to sum up the target weight values of the service instances in the sliding window to obtain a total weight value; in the total weight range constructed based on the total weight value, use a random number to determine the target service instance for performing the load balancing task.
[0128] A selection module 85, which is used to divide the corresponding sub-range from the total weight range according to the target weight value corresponding to the service instance; use a random number algorithm to determine a random number; where the random number is an integer greater than zero and less than or equal to the total weight value; determine the target service instance corresponding to the sub-range containing the random number; use the target service instance to perform the load balancing task.
[0129] An allocation module 83, which is used to reverse the order of the status information to be analyzed according to the generation order; and evenly divide the status information to be analyzed in reverse order into multiple time periods at a specified time interval.
[0130] An allocation module 83, which is used to assign a first weight value to the first time period and a second weight value to the second time period; where the first time period is closer to the current moment than the second time period, and the first weight value is less than the second weight value.
[0131] A determination module 84 is configured to count a first number of anomalies in service response anomalies in the status information to be analyzed, and count the total number of service responses in which the service response is normal and the service response is abnormal within a first time period; calculate a first anomaly rate based on the ratio of the first number of anomalies to the total number of service responses; compare the magnitude relationship between the first anomaly rate and a failure threshold; when the comparison result of the magnitude relationship is that the first anomaly rate is greater than the failure threshold, determine that the failure weight of the service instance is the target weight value corresponding to the first time period.
[0132] The determination module 84 is configured to, when the comparison result of the magnitude relationship is that the first anomaly rate is not greater than the failure threshold, calculate a second anomaly rate within a second time period; compare the magnitude relationship between the second anomaly rate and the failure threshold; when the comparison result of the magnitude relationship is that the second anomaly rate is greater than the failure threshold, determine that the failure weight of the service instance is the target weight value corresponding to the second time period; when the comparison result of the magnitude relationship is that the second anomaly rate is greater than the failure threshold, and the second time period is the last time period in the sliding window, determine that the failure weight of the service instance is the default weight value; wherein, the default weight value is greater than the target weight value.
[0133] The determination module 84 is configured to respectively count the sub - numbers of anomalies corresponding to each time period in the status information to be analyzed; respectively determine the corresponding magnification factors according to the chronological order of the time periods; and obtain the first number of anomalies by summing the products of the magnification factors and the sub - numbers of anomalies.
[0134] An interception module 82 is configured to determine the service instance as the primary key; intercept the status information to be analyzed from the working status information as the value using a sliding window with a specified time length; and store the service instance and the sliding window in the hash table in the form of key - value pairs.
[0135] The execution subject of the load balancing method in the specific implementation manner of the present disclosure may be an electronic device such as a server (including a local server or a cloud server).
[0136] Therefore, based on any of the above - mentioned embodiments, the present disclosure further provides an electronic device, and this electronic device can execute the load balancing method of any of the above - described embodiments of the present disclosure.
[0137] Figure 9 It is a structural schematic block diagram of an electronic device 1000 according to an embodiment of the present disclosure.
[0138] The hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnecting buses and bridges, depending on the specific application of the hardware and the overall design constraints. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.
[0139] The bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Component (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one connecting line is shown in this figure, but it does not mean that there is only one bus or one type of bus.
[0140] The present disclosure also provides a readable storage medium. A computer program is stored in the readable storage medium and is used to implement the above method when executed by a processor. The "readable storage medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0141] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part.
[0142] Computer programs or instructions can be stored in a readable storage medium or transmitted from one readable storage medium to another. For example, the computer programs or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any available medium that can be accessed or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.
[0143] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, system, or computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0144] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable load balancing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable load balancing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0145] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable load balancing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0146] These computer program instructions can also be loaded onto a computer or other programmable load balancing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process or multiple processes and / or blocks. Figure 1 One process or multiple processes and / or blocks Figure 1 Steps for implementing the functions specified in one block or multiple blocks.
[0147] In the description of this specification, the descriptions referring to terms such as "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc. mean that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.
[0148] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of such features. In the description of the present disclosure, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0149] Those skilled in the art should understand that the above embodiments are only for clearly explaining the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications can be made based on the above disclosure, and these changes or modifications are still within the scope of the present disclosure.
Claims
1. A load balancing method, characterized in that, The method includes: Obtaining the working status information of multiple service instances in a distributed system; wherein, the working status information includes: normal service response and / or abnormal service response; Using a sliding window with a specified time length to intercept the status information to be analyzed from the working status information of each service instance; and performing fault analysis using the number of normal service responses and / or the number of abnormal service responses in the status information to be analyzed to obtain a fault analysis result; Dividing the status information to be analyzed in the sliding window into multiple time periods, and assigning weight values to each time period; Based on the fault analysis result of the status information to be analyzed, determining the target weight value of each service instance from multiple weight values; According to each target weight value, determining the target service instance among the multiple service instances for performing the load balancing task.
2. The method according to claim 1, wherein The determining the target service instance among the multiple service instances for performing the load balancing task according to each target weight value includes: Summing up the target weight values of the service instances in the sliding window to obtain a total weight value; In the total weight interval constructed based on the total weight value, using a random number to determine the target service instance for performing the load balancing task.
3. The method according to claim 2, wherein The using a random number to determine the target service instance for performing the load balancing task in the total weight interval constructed based on the total weight value includes: According to the target weight value corresponding to the service instance, dividing the corresponding sub-interval range from the total weight interval; Using a random number algorithm to determine a random number; wherein, the random number is an integer greater than zero and less than or equal to the total weight value; Determining the target service instance corresponding to the sub-interval range containing the random number; Using the target service instance to perform the load balancing task.
4. The method according to claim 1, wherein Assigning weight values to each time period includes: Assigning a first weight value to the first time period and a second weight value to the second time period; wherein, the first time period is closer to the current moment than the second time period, and the first weight value is less than the second weight value.
5. The method according to claim 4, wherein The determining the target weight value of each service instance from multiple weight values based on the fault analysis result of the status information to be analyzed includes: Counting the first abnormal number of abnormal service responses in the status information to be analyzed, and counting the total number of service responses of normal service responses and abnormal service responses in the first time period; Calculating a first abnormal rate based on the ratio of the first abnormal number and the total number of service responses; Comparing the size relationship between the first abnormal rate and the fault threshold; When the comparison result of the size relationship is that the first abnormal rate is greater than the fault threshold, determining the fault weight of the service instance as the target weight value corresponding to the first time period.
6. The method according to claim 5, characterized in that, It further includes: When the comparison result of the size relationship is that the first abnormal rate is not greater than the fault threshold, calculating the second abnormal rate in the second time period; Comparing the size relationship between the second abnormal rate and the fault threshold; When the comparison result of the size relationship is that the second exception rate is greater than the fault threshold, determine that the fault weight of the service instance is the target weight value corresponding to the second time period; When the comparison result of the size relationship is that the second exception rate is greater than the fault threshold, and the second time period is the last time period in the sliding window, determine that the fault weight of the service instance is the default weight value; wherein, the default weight value is greater than the target weight value.
7. The method according to claim 5, characterized in that, The step of counting the first number of exceptions of service response exceptions in the status information to be analyzed includes: Count the number of sub-exceptions corresponding to each time period in the status information to be analyzed respectively; Determine the corresponding magnification factors according to the chronological order of the time periods; Use the sum of the products of the magnification factors and the number of sub-exceptions to obtain the first number of exceptions.
8. The method according to claim 1, characterized in that, The step of intercepting the status information to be analyzed from the working status information of each service instance by using a sliding window with a specified time length includes: Determine the service instance as the primary key; Use a sliding window with a specified time length to intercept the status information to be analyzed from the working status information as the value; Store each service instance and the corresponding sliding window in the form of key-value pairs into the hash table.
9. An electronic device, characterized in that, Comprising: A memory that stores execution instructions; And A processor that executes the execution instructions stored in the memory, so that the processor executes the method according to any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 9.