Service capacity risk determination method and device, electronic equipment and storage medium
By acquiring and analyzing the multiple indicator data of the service components and monitoring components, calculating weighted statistical values and issuing alarm information, the problems of inefficiency and insufficient accuracy in service capacity risk analysis and management in the prior art are solved, and more efficient and accurate risk determination is achieved.
Patent Information
- Application Number
- CN202510134390.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art has problems of inefficiency and insufficient accuracy in service capacity risk analysis and management, especially when dealing with special circumstances, it is difficult to distinguish different types of capacity risks.
By obtaining the CPU, memory, response time and query rate per second data of each target metric data, based on these data, the weighted statistical values are calculated as the target water level value. When the target water level value exceeds the preset water level value, determine whether the preset alarm conditions are met and issue corresponding alarm information.
It improves the efficiency and accuracy of service capacity risk determination, can more accurately distinguish between service capacity risk and capacity risk caused by monitoring components themselves, reduces manual inspection costs, and improves the accuracy of capacity management decisions.
Smart Images

Figure CN120029869A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to a service capacity risk determination method, device, electronic device and storage medium. Background Art
[0002] Currently, online service capacity risk analysis and management systems mainly evaluate service capacity by monitoring multiple indicators such as CPU usage, number of requests per second (QPS), and request duration (Cost). These systems produce a capacity watermark by fitting these data to indicate the current capacity usage. Recent technological advances have enabled the system to produce a risk list through statistical data to help the operation and maintenance team identify potential capacity issues.
[0003] In related technologies, manual intervention is often required to screen and confirm real risk points, which is a time-consuming and inefficient process. Existing systems perform poorly in handling special situations, such as difficulty in distinguishing high CPU usage caused by the monitoring component itself, water level fluctuations caused by non-main process services or batch tasks, and short-term capacity water level anomalies caused by data glitches. These problems lead to a large amount of manual troubleshooting costs and may cause incorrect capacity management decisions. That is, the efficiency and accuracy of risk determination are low. Summary of the invention
[0004] In view of this, embodiments of the present invention provide a service capacity risk determination method, device, electronic device and storage medium to improve the efficiency and accuracy of risk determination.
[0005] According to one aspect of the present invention, a method for determining service capacity risk is provided, the method comprising:
[0006] Obtain each target indicator data, wherein the target indicator data includes service component data of each service component running in the instance and the container and monitoring component data of each monitoring component; the service component data and the monitoring component data both include CPU data, memory data, response time, and query rate per second;
[0007] Based on each of the target indicator data and the weight assigned to each of the target indicator data, a weighted statistical value corresponding to each of the target indicator data is calculated as a target water level value;
[0008] In the case where the target water level value exceeds the preset water level value, determining whether each of the target indicator data meets the preset alarm conditions for the data type to which the target indicator data belongs, wherein the data type includes CPU data, memory data, response time, and query rate per second; each of the alarm conditions includes different threshold conditions corresponding to the target indicator data;
[0009] In the case that there is target indicator data that meets the alarm condition, alarm information is issued based on the target indicator data in an alarm manner corresponding to the threshold condition met by the target indicator data.
[0010] In a possible embodiment, the obtaining of each target indicator data includes:
[0011] Collect the CPU usage, memory usage, response time, and query rate per second of each instance at a preset collection frequency; the instances include service instances and monitoring instances; and,
[0012] The CPU usage, memory usage, response time, and query rate per second of each service running in each pod container are collected according to the preset collection frequency, and the CPU quota, CPU usage, memory usage, response time, and query rate per second of the monitoring components running in each pod container are collected.
[0013] In a possible embodiment, when the target indicator data is CPU data, determining whether each target indicator data meets a preset alarm condition for the data type to which the target indicator data belongs includes:
[0014] Determine whether the target indicator data exceeds a first CPU threshold, and if the target indicator data exceeds the first CPU threshold, compare the target indicator data with historical CPU data, wherein the historical CPU data is CPU data collected before a time point when the target indicator data is collected;
[0015] When the increment of the target indicator data relative to the historical CPU data exceeds the first increment threshold, the overall CPU utilization of the server corresponding to the target indicator data is determined, and when the overall CPU utilization exceeds the second CPU threshold, it is determined that the CPU data meets the alarm condition and a CPU data alarm is sent, wherein the second CPU threshold is higher than the first CPU threshold.
[0016] In a possible embodiment, the method further includes: when the target indicator data does not exceed the first CPU threshold, determining whether the CPU data in the monitoring component data exceeds the CPU data of the service component, and when the CPU data in the monitoring component data exceeds the CPU data of the service component, determining that the CPU data meets the alarm condition, and sending an alarm on the CPU occupancy of the monitoring component;
[0017] When the increment of the target indicator data relative to the historical CPU data does not exceed a first increment threshold, it is determined that the CPU data meets the alarm condition, and a continuous high load alarm is sent.
[0018] In a possible embodiment, when the target indicator data is a response time, determining whether each target indicator data meets a preset alarm condition for a data type to which the target indicator data belongs includes:
[0019] Determine whether the target data exceeds a first response time threshold. If the target data exceeds the first response time threshold, determine whether the target data exceeds a second response time threshold. If the target data exceeds the second response time threshold, determine that the target indicator data meets the alarm condition and send a timeout alarm message, wherein the second response time threshold is higher than the first response time threshold.
[0020] In a possible embodiment, when the target indicator data is a query rate per second, determining whether each target indicator data meets a preset alarm condition for a data type to which the target indicator data belongs includes:
[0021] Determine whether the target indicator data exceeds a preset QPS threshold, and if the target indicator data exceeds the preset QPS threshold, query historical QPS data within a preset time period;
[0022] Based on the historical QPS data and the target indicator data, determine the growth form of the query rate per second, and if the growth form is a short-term fluctuation, determine that the target indicator data meets the alarm condition, and send a short-term fluctuation alarm;
[0023] In the case where the growth form is continuous growth, it is determined that the target indicator data meets the alarm condition, and a continuous growth alarm is sent.
[0024] In a possible embodiment, the method further includes: based on the target indicator data, outputting an indicator data prediction value through a pre-trained risk prediction model, and sending an alarm message based on the indicator data prediction value.
[0025] According to another aspect of the present invention, a service capacity risk determination device is provided, the device comprising:
[0026] An acquisition module is used to acquire target indicator data, wherein the target indicator data includes service component data of each service component running in the instance and the container and monitoring component data of each monitoring component; the service component data and the monitoring component data both include CPU data, memory data, response time, and query rate per second;
[0027] A calculation module, used for calculating the weighted statistical value corresponding to each target indicator data as the target water level value based on each target indicator data and the weight assigned to each target indicator data;
[0028] A determination module, used for determining whether each of the target indicator data meets the preset alarm conditions for the data type to which the target indicator data belongs when the target water level value exceeds the preset water level value, wherein the data type includes CPU data, memory data, response time and query rate per second; each of the alarm conditions includes different threshold conditions corresponding to the target indicator data;
[0029] The alarm module is used to issue an alarm message based on the target indicator data in an alarm mode corresponding to the threshold condition satisfied by the target indicator data when there is target indicator data satisfying the alarm condition.
[0030] According to another aspect of the present invention, there is provided an electronic device, comprising:
[0031] Processor; and
[0032] Memory for storing programs,
[0033] The program includes instructions, which, when executed by the processor, cause the processor to execute any of the above-mentioned service capacity risk determination methods.
[0034] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute any of the above-mentioned service capacity risk determination methods.
[0035] One or more technical solutions provided in the embodiments of the present invention obtain target indicator data, which includes CPU data, memory data, response time and query rate per second data of each service component and monitoring component. Based on the target indicator data and the weights assigned to each target indicator data, the weighted statistical value of the target indicator data is calculated. When the weighted statistical value is higher than the preset water level value, the target indicator data that meets the alarm condition is determined and an alarm message is issued. By applying the embodiments of the present invention, risk determination is performed by collecting monitoring component data, which helps to distinguish between service capacity risk and capacity risk caused by the monitoring component itself, avoids misjudgment of capacity risk caused by the monitoring component, and improves the accuracy of risk determination. Furthermore, by determining whether capacity risk occurs based on the weighted statistical value of each indicator data, further capacity risk judgment is performed for each type of indicator data, and capacity risk judgment is performed through multi-level ranges and multi-dimensional indicators. While improving the efficiency of risk determination, the accuracy of risk judgment is improved. In addition, the scheme provided in this application does not require manual participation throughout the process, which greatly improves the efficiency of risk determination. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Further details, features and advantages of the invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0037] Figure 1 A schematic diagram of a flow chart of a method for determining service capacity risk provided by an embodiment of the present invention;
[0038] Figure 2 A schematic diagram of a process for determining whether CPU data meets a preset alarm condition in an embodiment of the present invention;
[0039] Figure 3 A schematic diagram of a process for determining whether response time data meets a preset alarm condition in an embodiment of the present invention;
[0040] Figure 4 A schematic diagram of a process for determining whether QPS data meets a preset alarm condition in an embodiment of the present invention
[0041] Figure 5 A schematic diagram of another flow chart of a method for determining service capacity risk provided by an embodiment of the present invention;
[0042] Figure 6 A schematic diagram of another flow chart of a method for determining service capacity risk provided by an embodiment of the present invention;
[0043] Figure 7 A schematic diagram of the structure of a device for determining service capacity risk provided by an embodiment of the present invention;
[0044] Figure 8A block diagram of an exemplary electronic device that can be used to implement an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0045] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.
[0046] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0047] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0048] It should be noted that the modifications of "one" and "plurality" mentioned in the present invention are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0049] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes, and are not used to limit the scope of these messages or information.
[0050] In order to improve the efficiency and accuracy of risk screening in big data systems, embodiments of the present invention provide a service capacity risk determination method, device, electronic device, and storage medium. The service capacity risk determination method provided by the embodiments of the present invention can be applied to any electronic device with a service capacity risk determination function, which can be a server, a computer, or a mobile terminal. The following describes the solution of the present invention with reference to the accompanying drawings:
[0051] Figure 1A flow chart of a method for determining service capacity risk provided by an embodiment of the present invention may include the following steps:
[0052] S101. Obtain target indicator data, wherein the target indicator data includes service component data of each service component running in the instance and the container and monitoring component data of each monitoring component; the service component data and the monitoring component data both include CPU data, memory data, response time, and query rate per second;
[0053] S102, based on each of the target indicator data and the weight assigned to each of the target indicator data, calculating a weighted statistical value corresponding to each of the target indicator data as a target water level value;
[0054] S103, when the target water level value exceeds the preset water level value, determining whether each of the target indicator data meets the preset alarm condition for the data type to which the target indicator data belongs, wherein the data type includes CPU data, memory data, response time, and query rate per second; each of the alarm conditions includes different threshold conditions corresponding to the target indicator data;
[0055] S104: When there is target indicator data that meets the alarm condition, issue an alarm message based on the target indicator data in an alarm manner corresponding to the threshold condition met by the target indicator data.
[0056] In an embodiment of the present invention, target indicator data is obtained, and the target indicator data includes CPU data, memory data, response time and query rate per second data of each service component and monitoring component. Based on the target indicator data and the weight assigned to each target indicator data, the weighted statistical value of the target indicator data is calculated. When the weighted statistical value is higher than the preset water level value, the target indicator data that meets the alarm condition is determined and an alarm message is issued. By applying the embodiment of the present invention, risk determination is performed by collecting monitoring component data, which helps to distinguish between service capacity risk and capacity risk caused by the monitoring component itself, avoids misjudgment of capacity risk caused by the monitoring component, and improves the accuracy of risk determination. In addition, by determining whether capacity risk occurs based on the weighted statistical value of each indicator data, further capacity risk judgment is performed for each type of indicator data, and capacity risk judgment is performed through multi-level ranges and multi-dimensional indicators. While improving the efficiency of risk determination, the accuracy of risk judgment is improved, and the scheme provided by the present application does not require manual participation throughout the process, which greatly improves the efficiency of risk determination.
[0057] The above S101-S104 are exemplarily described below:
[0058] In S101, the target indicator data can be obtained through a big data monitoring component or through a system log. The monitoring component can be prometheus, noah or other possible big data monitoring components. The monitoring component is used to monitor various data in each instance and pod container in the distributed cluster, including CPU data, memory data, response time, and query rate per second (QPS). Among them, an instance refers to a specific object created according to a class template, which has the attributes and behaviors of the class. A pod container is the smallest deployable object in Kubernetes. It is a combination of one or more containers, and various instances can be run in the container. The system log can record the resource usage data of each component, including the resource usage data of the above instance, pod container and monitoring component. The resource may include CPU and memory.
[0059] In a possible embodiment, the CPU usage, memory usage, response time, and query rate per second of each instance may be collected at a preset collection frequency; the instance includes a service instance and a monitoring instance; and,
[0060] The CPU usage, memory usage, response time, and query rate per second of each service running in each pod container are collected according to the preset collection frequency, and the CPU quota, CPU usage, memory usage, response time, and query rate per second of the monitoring components running in each pod container are collected.
[0061] The above-mentioned collection frequency can be flexibly set according to the actual application scenario, such as collecting data every two minutes to ensure the real-time and continuity of the data. The above-mentioned instance refers to the instance directly deployed in the server, which may specifically include a physical machine instance and a virtual machine instance. The physical machine instance and the virtual machine instance may include a service instance and a monitoring instance, wherein the service instance refers to the service instance running each network service in the big data cluster, and the monitoring instance refers to the instance of the monitoring component. The above-mentioned CPU utilization rate refers to the proportion of the CPU used by the instance in the CPU allocated to the instance. The above-mentioned response time refers to the time taken from the request end initiating the request to the request end receiving the return from the server end. The query rate per second (QPS) refers to the number of query requests completed per second.
[0062] The above pod container runs services and monitoring components. As a possible implementation, the CPU usage, memory usage, response time cost, and query rate per second (QPS) of each service component running in the pod container can be collected. The overall CPU usage of the container used to run the service component in the pod can also be collected. As a possible implementation, the CPU usage, CPU quota, memory usage, response time, and query rate per second of the monitoring container running in the pod can be collected, where the CPU quota refers to the maximum CPU usage limit allocated to a specific service or application, which is used to control and manage the allocation of system resources.
[0063] After obtaining the above-mentioned target indicator data, the target indicator data can be weighted calculated based on the target indicator data and the preset weights for each type of target indicator data. Among them, the types of target indicator data can include CPU data, memory data, response time and query rate per second. The weights set for the above-mentioned various types of target indicator data can be set according to the actual application scenario, and it is only necessary to ensure that the sum of each weight is 1. As a possible implementation method, the weights of CPU data, memory, response time and query rate per second can be set to 25%. In another possible implementation method, part of the target indicator data can also be selected for calculation, such as CPU data, response time and query rate per second can be selected for weighted calculation, and the weight of CPU data is set to 50%, and the weight of response time and query rate per second is set to 25%.
[0064] In a possible embodiment, a weighted statistical value can be obtained based on the sum of the product of the target indicator data and the weight, that is, the weighted statistical value can be obtained by CPU data × weight 1 + memory data × weight 2 + response time × weight 3 + query rate per second × weight 4. In a possible embodiment, the above weighted statistical value can also be obtained by the following formula:
[0065] CPU data × weight 1 + memory data × weight 2 + response time × weight 3 / maximum response time + query rate per second × weight 4 / maximum query rate per second.
[0066] The above maximum response time / maximum query rate per second may be the maximum value of the historical response time / historical query rate per second, and the historical response time / historical query rate per second may be the maximum value of the response time data and QPS data collected in the past 24 hours, or the maximum value of the response time data and QPS data collected in the past week, and the present invention does not make specific limitations on this. For ease of description, the above weighted statistical value will be referred to as the target water level value below. The above target water level value calculation may be performed for one or more services, or for one or more containers, and the present invention does not make specific limitations on this.
[0067] In a possible embodiment, the target water level value and the preset water level value may be compared to determine whether the target water level value is greater than the preset water level value. The preset water level value may be obtained during the stress test, or may be obtained by predicting the threshold of each indicator data through a pre-trained regression model and performing weighted calculation based on the threshold. Exemplarily, the preset water level value may be 70%. When the target water level value exceeds 70%, it indicates that the capacity currently allocated to the service may be too small.
[0068] The above regression model can be pre-trained based on a large amount of indicator data, and is used to output the corresponding indicator data thresholds when the capacity limit is output based on the indicator data. Exemplarily, the regression model can be trained based on the indicator data and the corresponding capacity risk label, and the regression model can be trained based on the difference between the risk prediction result output by the regression model based on the predicted indicator data threshold and the risk label. In the present invention, any feasible method can be used to train the above regression model.
[0069] If the target water level value exceeds the preset water level value, the target indicator data can be further analyzed to determine the indicator data causing capacity risk. If the target water level value does not exceed the preset water level value, the indicator data can be continuously monitored.
[0070] In the present invention, each target indicator data may be detected based on the alarm condition preset for each type of target indicator data. In a possible embodiment, when the target indicator data is CPU data, determining whether each target indicator data meets the alarm condition preset for the data type to which the target indicator data belongs includes:
[0071] Determine whether the target indicator data exceeds a first CPU threshold, and if the target indicator data exceeds the first CPU threshold, compare the target indicator data with historical CPU data, wherein the historical CPU data is CPU data collected before a time point when the target indicator data is collected;
[0072] When the increment of the target indicator data relative to the historical CPU data exceeds a first increment threshold, the overall CPU utilization of the server corresponding to the target indicator data is determined; when the overall CPU utilization exceeds a second CPU threshold, it is determined that the CPU data meets the alarm condition, and a CPU data alarm is sent, wherein the second CPU threshold is higher than the first CPU threshold.
[0073] The first CPU threshold, the second CPU threshold and the first increment threshold can all be set according to the actual application scenario, such as the first CPU threshold can be set to 35% CPU usage, the second CPU threshold can be set to 60% CPU usage, and the first increment threshold can be set to 50%. The increment of the target indicator data and the historical CPU data can be obtained by (target indicator data-historical CPU data) / historical CPU data.
[0074] The above-mentioned historical data may be multiple CPU data collected before the target indicator data collection time point. As a possible implementation method, the first increment between the historical CPU data and the increment between the historical CPU data whose collection time is closest to the target indicator data collection time and the target indicator data can also be calculated in chronological order as the second increment. When the difference between the second increment and the first increment exceeds the preset difference, it can also be determined that the CPU data meets the alarm condition.
[0075] In a possible embodiment, when the target indicator data does not exceed the first CPU threshold, it is determined whether the CPU data in the monitoring component data exceeds the CPU data of the service component, and when the CPU data in the monitoring component data exceeds the CPU data of the service component, it is determined that the CPU data meets the alarm condition, and an alarm of the CPU occupancy of the monitoring component is sent;
[0076] When the increment of the target indicator data relative to the historical CPU data does not exceed a first increment threshold, it is determined that the CPU data meets the alarm condition, and a continuous high load alarm is sent.
[0077] For example, Figure 2 As shown, Figure 2 A flow chart for determining whether CPU data meets preset alarm conditions in an embodiment of the present invention, specifically, first check whether the CPU usage rate exceeds 35%, if it exceeds, determine whether the CPU share of the monitoring component exceeds the service. If it exceeds, it means that the monitoring component may be the cause of the high target water level value, so the CPU occupancy of the monitoring component can be notified, if it does not exceed, compare the CPU usage rate with the historical data, and determine whether it is a sudden increase in CPU based on its increment, if it is a sudden increase (increment exceeds 50%), the overall CPU utilization of the server can be determined, if it exceeds the second threshold (60%), a capacity alarm is issued, if it exceeds the second threshold, only record it, and do not issue an alarm; if it is not a sudden increase, a continuous high load alarm can be issued.
[0078] In a possible embodiment, when the target indicator data is a response time, determining whether each target indicator data meets a preset alarm condition for a data type to which the target indicator data belongs includes:
[0079] Determine whether the target data exceeds a first response time threshold. If the target data exceeds the first response time threshold, determine whether the target data exceeds a second response time threshold. If the target data exceeds the second response time threshold, determine that the target indicator data meets the alarm condition and send a timeout alarm message, wherein the second response time threshold is higher than the first response time threshold.
[0080] The above-mentioned first response time threshold and second response time threshold can also be obtained through the above-mentioned regression model. For example, the response time in the past 24 hours can be input into the regression model, and the response time threshold output by the regression model can be obtained. The response time threshold can be used as the second response time threshold. The first response time threshold can be obtained by subtracting the preset value from the second response time threshold. For example, the second response time threshold can be 100%, and the first response time threshold can be 30%. The percentage is the proportion of the preset response time. The preset response time can be set according to the actual application scenario, such as 0.5s, 0.1s, etc.
[0081] like Figure 3 As shown, Figure 3 A flow chart of determining whether the response time data meets the preset alarm condition in an embodiment of the present invention. Specifically, when the response time cost is lower than 30%, it is possible to evaluate whether the service performance has reached a bottleneck, and output prompt information for reducing the package or expanding the instance. When the cost is between 30% and 100%, the system load can be checked, and prompt information for evaluating whether the code needs to be optimized or resources need to be increased can be output. When the cost is higher than 100%, an alarm message is output to enable relevant personnel to check whether there is a timeout problem and perform emergency performance diagnosis.
[0082] In a possible embodiment, when the target indicator data is a query rate per second, determining whether each target indicator data meets a preset alarm condition for a data type to which the target indicator data belongs includes:
[0083] Determine whether the target indicator data exceeds a preset QPS threshold, and if the target indicator data exceeds the preset QPS threshold, query historical QPS data within a preset time period;
[0084] Based on the historical QPS data and the target metric data, determine the growth pattern of the queries per second. In the case where the growth pattern is short-term fluctuation, determine that the target metric data meets the alarm condition and send a short-term fluctuation alarm.
[0085] In the case where the growth pattern is continuous growth, determine that the target metric data meets the alarm condition and send a continuous growth alarm.
[0086] The above QPS threshold can be obtained through the above regression model. Exemplarily, the QPS data of the past 24 hours can be input into the regression model so that the regression model outputs the highest QPS value that the service can withstand under normal conditions. If it exceeds this value, it indicates that there is a capacity risk. At this time, the risk type can be further judged.
[0087] In a possible embodiment, the QPS growth pattern can be determined based on the historical QPS data and the QPS data in the target metric data. The growth rate between the QPS data with adjacent collection times can be calculated in chronological order. If the difference between the growth rates is less than the preset difference threshold, it is determined that the QPS is in continuous growth. If there is a growth rate greater than the preset difference threshold, it can be determined that the QPS is in short-term fluctuation.
[0088] As Figure 4 shown, Figure 4 This is a schematic flowchart for determining whether the QPS meets the preset alarm condition in the embodiment of the present invention. Determine whether the QPS data exceeds the limit regression value. If it exceeds, query the data of the previous ten minutes to judge the QPS growth pattern. If it is short-term fluctuation, evaluate the short-term impact of the fluctuation on the system and check whether additional resources need to be temporarily added. If it is continuous growth, evaluate the growth trend and analyze the impact on the long-term performance of the system.
[0089] In a possible embodiment, the above method may further include outputting a predicted value of the metric data based on the target metric data through a pre-trained risk prediction model, and sending an alarm message based on the predicted value of the metric data.
[0090] In a possible embodiment, the above alarm message may include a risk level, specific values, intelligent annotation results, duration, and scope of influence. The above risk level can be determined based on CPU data. Exemplarily, the first-level risk can be set as the CPU usage rate ≥ 70%, which means that the system may not be able to withstand the load during the traffic peak period; the second-level risk is 50% ≤ CPU usage rate < 70%, which means that the system can run during the traffic peak period but cannot perform unilateral switching; the third-level risk is 35% < CPU usage rate < 50%, which means that the system can perform unilateral traffic switching during the traffic peak period but there is a certain risk.
[0091] The above intelligent labeling results can be the above alarm types. For example, if the CPU usage exceeds the threshold for three consecutive sampling periods, it can be labeled [Continuous High Load] to identify non-incidental water level rises; MONITOR_CPU_USED> preset threshold (such as 10%) can be labeled [Monitor CPU Utilization] to indicate water level abnormalities caused by monitoring components; CPU usage surges during non-peak hours can be labeled [Non-main process / batch task] to identify temporary high water levels caused by specific tasks; high water levels in services in hbf / hbg computer rooms can be labeled [Specific computer room] to distinguish the situation in special computer rooms. The above impact range can be determined based on the service instances running in the server, or it can be output through the above risk prediction model.
[0092] like Figure 5 As shown, Figure 5 Another flow chart of a method for determining service capacity risk provided by an embodiment of the present invention may include the following steps:
[0093] S501. Collect CPU, QPS, and Cost data.
[0094] S502, determine whether the target water level is higher than 70%, if not, execute S503; if higher, execute S504.
[0095] S503, check whether the single indicator is abnormal, if so, execute S504, if not, execute S507.
[0096] S504. Analyze the reasons for the high water level. Specifically, it may include CPU analysis, Cost analysis, and QPS analysis. Among them, CPU analysis includes checking CPU usage, comparing with historical data, and determining whether it is monitored occupancy; Cost analysis includes comparing with the limit regression value and analyzing the increase; QPS analysis includes comparing with the limit regression value and checking whether it is a short-term fluctuation.
[0097] S505. Comprehensive risk assessment. Specifically, the risk level can be judged based on the high water level reasons analyzed above.
[0098] S506. Generate risk reports and recommendations.
[0099] S507. Continuous monitoring.
[0100] In a possible embodiment, Figure 6 As shown, Figure 6Another flow chart of the service capacity risk determination method provided in the embodiment of the present invention, specifically, a set of product lines can be obtained through a preset script query, and the product lines can be polled to process the resource collection process of each product line. For different product lines selected by the user in the capacity assessment platform, initialize and execute specific capacity collection tasks. Based on the service management and service list in the capacity assessment platform, obtain the product line BNS (Business Networking Services) list and the current module limit assessment data of the product line, then traverse the BNS list and specifically analyze a single BNS worker.
[0101] The traversal process includes: obtaining the monitoring collection items corresponding to the BNS based on the RDS table (usually refers to the database table created or managed when using the relational database service (RDS)), and obtaining the monitoring data corresponding to the BNS based on the collection items. Specifically, the instance set corresponding to the BNS and the monitoring item collection data corresponding to each logical computer room can be obtained. The above data can be obtained through the Noah interface. Evaluate the original monitoring capacity data of a single BNS, then normalize the data, and check whether the current monitoring data reaches the alarm threshold to determine whether to trigger an alarm. Specifically, it can be determined whether the current capacity alarm needs to be sent to the preset user group, traverse the monitoring data and obtain the relevant indicator data of the corresponding computer room. Then obtain the capacity assessment limit value of the product line BNS computer room, and encapsulate the relevant indicator values of the corresponding computer room.
[0102] After obtaining the above values, load the historical information of the number of instances from the local file to determine whether to notify the instance reduction, and perform data processing and analysis to determine whether to trigger the capacity water level alarm. Use the relevant algorithm to determine whether the relevant data of the current module is abnormal. Specifically, it can determine whether the latest sampling point of the computer room corresponding to the current module is an abnormal value, and whether the average response time is abnormal. Based on the judgment result, the traffic status is judged and the number of instances recommended for expansion is determined. Specifically, it can determine whether the latest QPS data of each computer room of BNS is normal.
[0103] After determining that an alarm is needed, check the alarm history to determine whether an alarm is needed, notify the relevant personnel and user groups of the alarm content, and store the alarm data. The capacity assessment platform can generate risk reports based on the alarm data.
[0104] In the embodiment of the present invention, accurate risk assessment is achieved through an intelligent risk grading algorithm combined with multi-dimensional indicators and historical data analysis. The innovative adaptive alarm labeling mechanism and multi-dimensional anomaly diagnosis method greatly improve the efficiency of problem identification and resolution. Dynamic threshold adjustment and predictive capacity management model enable this method to adapt to different business scenarios and be forward-looking. Resource utilization efficiency evaluation and intelligent alarm suppression strategies optimize resource management and operation and maintenance response processes. Cross-service correlation analysis methods help prevent system-level capacity problems. The application of the embodiment of the present invention significantly improves the accuracy, foresight and operability of capacity management, and provides strong technical support for the stable operation of enterprise-level systems. This method not only solves multiple pain points in traditional capacity management, but also opens up new directions for future intelligent operation and maintenance.
[0105] Based on the same inventive concept, the embodiment of the present invention also provides a service capacity risk determination device, such as Figure 7 As shown, the apparatus 700 may include:
[0106] The acquisition module 701 is used to acquire various target indicator data, wherein the target indicator data includes service component data of each service component running in the instance and the container and monitoring component data of each monitoring component; the service component data and the monitoring component data both include CPU data, memory data, response time, and query rate per second;
[0107] A calculation module 702 is used to calculate a weighted statistical value corresponding to each target indicator data as a target water level value based on each target indicator data and a weight assigned to each target indicator data;
[0108] The determination module 703 is used to determine whether each of the target indicator data meets the preset alarm conditions for the data type to which the target indicator data belongs when the target water level value exceeds the preset water level value, wherein the data type includes CPU data, memory data, response time, and query rate per second; each of the alarm conditions includes different threshold conditions corresponding to the target indicator data;
[0109] The alarm module 704 is used to issue an alarm message based on the target indicator data in an alarm mode corresponding to the threshold condition satisfied by the target indicator data when there is target indicator data satisfying the alarm condition.
[0110] Among them, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the present invention are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0111] The exemplary embodiment of the present invention further provides an electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to enable the electronic device to perform a method according to an embodiment of the present invention when executed by the at least one processor.
[0112] Exemplary embodiments of the present invention also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present invention.
[0113] An exemplary embodiment of the present invention further provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor of a computer, the computer is used to enable the computer to perform a method according to an embodiment of the present invention.
[0114] refer to Figure 8 , a block diagram of an electronic device 800 that can be used as a server or client of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0115] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0116] A plurality of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 may be any type of device capable of inputting information to the electronic device 800, and the input unit 806 may receive input digital or character information, and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 807 may be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 may include, but is not limited to, a disk, an optical disk. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0117] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above. For example, in some embodiments, any of the above-described service capacity risk determination methods may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. In some embodiments, the computing unit 801 may be configured to perform any of the above-described service capacity risk determination methods in any other appropriate manner (e.g., by means of firmware).
[0118] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0119] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0120] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0121] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0122] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0123] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.
Claims
1. A method for determining service capacity risk, characterized in that: The method comprises: Obtain each target indicator data, wherein the target indicator data includes service component data of each service component running in the instance and the container and monitoring component data of each monitoring component; the service component data and the monitoring component data both include CPU data, memory data, response time, and query rate per second; Based on each of the target indicator data and the weight assigned to each of the target indicator data, a weighted statistical value corresponding to each of the target indicator data is calculated as a target water level value; In the case where the target water level value exceeds the preset water level value, determining whether each of the target indicator data meets the preset alarm conditions for the data type to which the target indicator data belongs, wherein the data type includes CPU data, memory data, response time, and query rate per second; each of the alarm conditions includes different threshold conditions corresponding to the target indicator data; In the case that there is target indicator data that meets the alarm condition, alarm information is issued based on the target indicator data in an alarm manner corresponding to the threshold condition met by the target indicator data.
2. The method according to claim 1, characterized in that The obtaining of target indicator data includes: Collect the CPU usage, memory usage, response time, and query rate per second of each instance at a preset collection frequency; the instances include service instances and monitoring instances; and, The CPU usage, memory usage, response time, and query rate per second of each service running in each pod container are collected according to the preset collection frequency, and the CPU quota, CPU usage, memory usage, response time, and query rate per second of the monitoring components running in each pod container are collected.
3. The method according to claim 1, characterized in that In the case where the target indicator data is CPU data, determining whether each target indicator data meets a preset alarm condition for the data type to which the target indicator data belongs includes: Determine whether the target indicator data exceeds a first CPU threshold, and if the target indicator data exceeds the first CPU threshold, compare the target indicator data with historical CPU data, wherein the historical CPU data is CPU data collected before a time point when the target indicator data is collected; When the increment of the target indicator data relative to the historical CPU data exceeds a first increment threshold, the overall CPU utilization of the server corresponding to the target indicator data is determined; when the overall CPU utilization exceeds a second CPU threshold, it is determined that the CPU data meets the alarm condition, and a CPU data alarm is sent, wherein the second CPU threshold is higher than the first CPU threshold.
4. The method according to claim 3, characterized in that The method further comprises: When the target indicator data does not exceed the first CPU threshold, determine whether the CPU data in the monitoring component data exceeds the CPU data of the service component. When the CPU data in the monitoring component data exceeds the CPU data of the service component, determine that the CPU data meets the alarm condition, and send an alarm on the CPU occupancy of the monitoring component. When the increment of the target indicator data relative to the historical CPU data does not exceed the first increment threshold, it is determined that the CPU data meets the alarm condition, and a continuous high load alarm is sent.
5. The method according to claim 1, characterized in that: In the case where the target indicator data is a response time, determining whether each target indicator data meets a preset alarm condition for a data type to which the target indicator data belongs includes: Determine whether the target data exceeds a first response time threshold. If the target data exceeds the first response time threshold, determine whether the target data exceeds a second response time threshold. If the target data exceeds the second response time threshold, determine that the target indicator data meets the alarm condition and send a timeout alarm message, wherein the second response time threshold is higher than the first response time threshold.
6. The method according to claim 1, characterized in that In the case where the target indicator data is a query rate per second, determining whether each target indicator data meets a preset alarm condition for a data type to which the target indicator data belongs includes: Determine whether the target indicator data exceeds a preset QPS threshold, and if the target indicator data exceeds the preset QPS threshold, query historical QPS data within a preset time period; Based on the historical QPS data and the target indicator data, determine the growth form of the query rate per second, and if the growth form is a short-term fluctuation, determine that the target indicator data meets the alarm condition, and send a short-term fluctuation alarm; In the case where the growth form is continuous growth, it is determined that the target indicator data meets the alarm condition, and a continuous growth alarm is sent.
7. The method according to claim 1, characterized in that The method further comprises: Based on the target indicator data, a predicted value of the indicator data is output through a pre-trained risk prediction model, and an alarm message is sent based on the predicted value of the indicator data.
8. A service capacity risk determination device, characterized in that: The device comprises: An acquisition module is used to acquire target indicator data, wherein the target indicator data includes service component data of each service component running in the instance and the container and monitoring component data of each monitoring component; the service component data and the monitoring component data both include CPU data, memory data, response time, and query rate per second; A calculation module, used for calculating the weighted statistical value corresponding to each target indicator data as the target water level value based on each target indicator data and the weight assigned to each target indicator data; A determination module, used for determining whether each of the target indicator data meets the preset alarm conditions for the data type to which the target indicator data belongs when the target water level value exceeds the preset water level value, wherein the data type includes CPU data, memory data, response time and query rate per second; each of the alarm conditions includes different threshold conditions corresponding to the target indicator data; The alarm module is used to issue an alarm message based on the target indicator data in an alarm mode corresponding to the threshold condition satisfied by the target indicator data when there is target indicator data satisfying the alarm condition.
9. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1-7.