A method for scoring service system risks
By establishing a service system risk scoring method, classifying risk levels and assigning weights, the problem of difficulty in quantifying the stability of microservice systems is solved, and quantitative assessment of system health and risk quantification are achieved, providing a basis for governance.
Patent Information
- Application Number
- CN202210915353.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2042-07-29
AI Technical Summary
The stability of microservice systems is difficult to quantify and assess, leading to governance difficulties and a lack of scientific decision-making basis.
Establish a service system risk scoring method, which comprehensively assesses the health of the system by classifying risk levels and assigning weights to each level, combined with operation and maintenance experience and data mining.
It enables objective and quantitative assessment of the system, provides scientific evidence to support decision-makers in system maintenance, quantifies risks, and provides data support for governance.
Smart Images

Figure CN115269330B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of big data technology, specifically relating to a method for scoring the risk of a service system. Background Technology
[0002] With the rise of microservices, service stability has become increasingly important, and the development and governance of microservices have become more difficult and harder to quantify. Therefore, we need to conduct a comprehensive assessment of the system's functionality, performance, reliability, and stability, establish a service risk governance model, and perform an objective and quantitative comprehensive evaluation of the system's operational status to provide decision-makers with a scientific basis for system operation and maintenance. Summary of the Invention
[0003] In view of this, the present invention provides a service system risk scoring method to achieve a comprehensive evaluation of the system's functions, performance, reliability, and stability, and to provide an objective and quantitative comprehensive assessment of the system's operating status, thereby providing a basis for decision-making regarding system operation and maintenance.
[0004] The technical solution adopted in this invention is as follows:
[0005] A method for scoring service system risk includes the following steps:
[0006] Step 1: Establish different risk levels;
[0007] Step 2: Based on the administrator's historical experience and the system's historical issues, determine the sources of risk in the system and establish the corresponding risk levels for different sources of risk;
[0008] Step 3: Score the system based on the various risk levels it contains and the weight of each individual risk level, and determine the health status of the system based on the score.
[0009] Preferably, the sources of risk mentioned in step 2 include: changes, monitoring warnings, performance inspections, and capacity inspections.
[0010] Preferably, in step 1, the different risk levels are specifically divided into five risk levels: Level 1 is slightly risky, Level 2 is moderate risky, Level 3 is significant risky, Level 4 is highly risky, and Level 5 is extremely dangerous.
[0011] Preferably, step 2, which involves determining the risk level for the changes, specifically includes:
[0012] The risk level is determined based on the number of emergency changes made to the same system within a week; the more emergency changes made, the higher the risk level.
[0013] The risk level is determined based on the number of times the same system is put into production within a week; the more times it is put into production, the higher the risk level.
[0014] Preferably, step 2, which involves determining the risk level for the changes, further includes:
[0015] Emergency change orders that are based on the same system and are not related to production windows are classified as high-risk.
[0016] If the same production schedule is executed multiple times with intervals of A hours, it is considered to be slightly risky.
[0017] Preferably, step 2, which involves determining the risk level for the monitoring warning, specifically includes:
[0018] Error warnings that persist for more than C hours within B days are considered significant risks.
[0019] An emergency warning that remains in effect for more than C hours within B days is classified as high-risk.
[0020] Preferably, step 2, which involves determining the risk level for the performance inspection, specifically includes:
[0021] If an application has an average daily CPU usage exceeding 80% for at least E days in the past D days, it is considered a significant risk (E). <D);
[0022] If an application has used more than 80% of its disk I / O capacity on average for at least E days in the past D days, it is considered a significant risk (E). <D);
[0023] If the peak TPS changes by less than 10% compared to yesterday, and the application causes a CPU utilization increase of more than 20%, it is considered a significant risk.
[0024] If the peak TPS changes by less than 10% compared to yesterday, but the application causes a memory increase of more than 20%, it is considered a significant risk.
[0025] If, within F hours, the database causes CPU utilization to exceed 70%, the number of database connections to exceed 70%, and the number of database connections to exceed E minutes, it is considered a significant risk.
[0026] If, out of G days, CPU utilization exceeds 80% for more than H days due to caching, it is considered a significant risk.
[0027] If, out of G days, memory usage exceeds 80% for H days or more due to caching, then it is considered a significant risk (H). <G)。
[0028] Preferably, step 2, which involves determining the risk level for the performance inspection, further includes:
[0029] The risk level is determined based on the number of times the interface was slow yesterday; the more times it was slow, the higher the risk level.
[0030] Preferably, step 2, which involves determining the risk level for the capacity inspection, further includes:
[0031] If disk usage consistently exceeds 80% within a consecutive day, it is considered a significant risk.
[0032] If debug-level logs exist in the system, they are considered a significant risk.
[0033] If the daily log volume of a subsystem exceeds 10GB, it will be classified as a general risk.
[0034] If a single log entry exceeds 1MB, it is classified as a general risk.
[0035] Preferably, step 3 above specifically includes the following steps:
[0036] Step 3.1: Establish the weight of each risk level, specifically: slightly risky is 5%, moderate risk is 8%, significant risk is 12%, high risk is 15%, and extremely dangerous is 60%.
[0037] Step 3.2: Analyze all risk sources in the system and their corresponding risk levels;
[0038] Step 3.3: Score the system according to Step 3.1 and Step 3.2, specifically: Total system score = 100 - the sum of the weights of each risk level included in the system multiplied by 100 respectively;
[0039] Step 3.4: Determine the system's health status based on the total system score from Step 3.3, specifically as follows:
[0040] When the total system score is greater than or equal to 70, the system status is healthy, and the system service risk is low.
[0041] When the total system score is greater than or equal to 50 and less than 70, the status is sub-healthy, and the system service risk is high.
[0042] When the total system score is less than 50, the status is unhealthy, and the system service risk is extremely high.
[0043] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0044] This model, based on fundamental operational tools and data mining techniques, organically combines information from various dimensions of the application system with the operational experience and incident summaries of system administrators. By digitizing the operational data according to certain weights, it generates numerical feedback on the application system in a specific dimension. Subsequently, a comprehensive score is awarded based on various dimensions of the business system, including change events, resource capacity, monitoring environment, transaction data, network traffic, and log information. This creates a digital profile of the application system's operational health. Finally, by combining practical operational experience to set weights for each dimension, a health score for the business system is obtained. Based on this score, the model determines whether the business system has risks, quantifies these risks, and provides data support for subsequent service risk management. Attached Figure Description
[0045] The present invention will be described by way of example and with reference to the accompanying drawings, wherein:
[0046] Figure 1 This is a schematic diagram of the process structure of the present invention;
[0047] Figure 2 This is a schematic diagram of the modified structure of the present invention;
[0048] Figure 3 This is a schematic diagram of the monitoring and alarm structure of the present invention;
[0049] Figure 4 This is a schematic diagram of the performance inspection structure of the present invention;
[0050] Figure 5 This is a schematic diagram of the capacity inspection structure of the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0052] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0053] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other.
[0054] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0055] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0056] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0057] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.
[0058] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other.
[0059] Example
[0060] This invention discloses a method for scoring the risk of a service system, comprising the following steps:
[0061] A method for scoring service system risk includes the following steps:
[0062] Step 1: Establish different risk levels;
[0063] The risk levels are specifically divided into five levels: Level 1 is slightly risky (1), Level 2 is moderately risky (2), Level 3 is significantly risky (3), Level 4 is highly risky (4), and Level 5 is extremely dangerous (5).
[0064] Step 2: Based on the administrator's historical experience and the system's historical issues, determine the sources of risk in the system and establish the corresponding risk levels for different sources of risk; that is, the administrator collects information on the sources of risk in the system based on their own experience and the system's historical issues; the sources of risk mentioned in step 2 specifically include: changes, monitoring warnings, performance inspections, and capacity inspections.
[0065] Step 2 is as follows:
[0066] 2.1: For example Figure 1 As shown, the risk level is established for each change, specifically including:
[0067] 2.11: Risk level is determined based on the number of emergency changes to the same system within a week. The more emergency changes, the higher the risk level. Specifically: 2.111: If there is 1 emergency change to the same system within a week, it is marked as significant risk; 2.112: If there are ≥2 emergency changes to the same system within a week, it is marked as high risk.
[0068] 2.12: The risk level is determined based on the number of times the same system is put into production within a week. The more times the system is put into production, the higher the risk level. Specifically: 2.121: If the number of times the system is put into production within a week is 2, it is marked as slightly risky; 2.122: If the number of times the system is put into production within a week is 3, it is marked as moderate risk (2); 2.123: If the number of times the system is put into production within a week is more than 3, it is marked as significant risk.
[0069] 2.13: Furthermore, step 2, in which the risk level is determined for the aforementioned changes, specifically includes:
[0070] 2.131: Emergency change orders that are based on the same system and involve changes outside the production window and are unrelated are classified as high-risk.
[0071] 2.132: If the same production schedule is executed multiple times with an interval of A (A=2) hours, it is considered to be slightly risky (1).
[0072] 2.2: For example Figure 2As shown, specifically, step 2, determining the risk level for the monitoring warning situation, includes:
[0073] 2.21: Error warnings that remain unclosed for more than C hours (C=12 hours) within B days (B=2 days) are considered significant risks;
[0074] 2.22: An emergency warning that remains in effect for more than C hours (C=12 hours) within B (B=2) days is classified as high risk.
[0075] 2.3: For example Figure 3 As shown, specifically, step 2, determining the risk level for the performance inspection situation, includes:
[0076] 2.31: If an application has an average daily CPU usage rate exceeding 80% for at least E (E=4) days in the past D (D=7) days, it is considered a significant risk.
[0077] 2.32: If an application has an average daily disk I / O usage rate exceeding 80% for at least E days (E=4 days) in the past D (D=7) days, it is considered a significant risk.
[0078] 2.32: If the peak TPS changes by less than 10% compared to yesterday, and the application causes a CPU utilization increase of more than 20%, it is considered a significant risk.
[0079] 2.33: If the peak TPS changes by less than 10% compared to yesterday, but the application causes an increase in memory usage of more than 20%, it is considered a significant risk.
[0080] 2.34: If, within F (F=24) hours, the database causes CPU utilization to exceed 70%, the number of database connections to exceed 70%, and the number of database connections to exceed E (E=60) minutes, it is considered a significant risk.
[0081] 2.35: If, out of G (G=14) days, H (H=4) or more days have CPU utilization exceeding 80% due to caching, then it is considered a significant risk;
[0082] 2.36: If, within G (G=14) days, H (H=4) or more days experience memory usage exceeding 80% due to caching, then this is considered a significant risk (H). <G)。
[0083] 2.37: Step 2, determining the risk level for the performance inspection results, further includes: determining the risk level based on the number of slow interfaces yesterday; the more slow interfaces, the higher the risk level. Specifically, this includes: 2.371: If the number of slow interfaces yesterday exceeds 5% of the total call volume of the structure, it is considered slightly risky; 2.372: If the number of slow interfaces yesterday exceeds 10% of the total call volume of the structure, it is considered moderately risky; 2.373: If the number of slow interfaces yesterday exceeds 15% of the total call volume of the structure, it is considered significantly risky; 2.374: If the number of slow interfaces yesterday exceeds 20% of the total call volume of the structure, it is considered highly risky; 2.375: If the number of slow interfaces yesterday exceeds 25% of the total call volume of the structure, it is considered extremely risky.
[0084] 2.4: For example Figure 4 As shown, step 2, determining the risk level for the capacity inspection, further includes:
[0085] 2.41: If disk usage consistently exceeds 80% for one consecutive day (I=3), it is considered a significant risk.
[0086] 2.42: If debug-level logs exist in the system, they are considered a significant risk;
[0087] 2.43: If the daily log volume of a subsystem exceeds 10GB, it is considered a general risk.
[0088] 2.44: If a single log entry exceeds 1MB, it is considered a general risk.
[0089] Once the sources of risk and their levels are determined, proceed to step 3.
[0090] Step 3: Score the system based on the various risk levels it contains and the weight of each individual risk level, and determine the health status of the system based on the score.
[0091] Step 3 specifically includes the following steps:
[0092] Step 3.1: Establish the weight of each risk level, specifically: slightly risky is 5%, moderate risk is 8%, significant risk is 12%, high risk is 15%, and extremely dangerous is 60%.
[0093] Step 3.2: Compile statistics on all sources of risk in the system and the corresponding risk level for each source of risk;
[0094] Step 3.3: Score the system according to Step 3.1 and Step 3.2, specifically: Total system score = 100 - the sum of the weights of each risk level included in the system multiplied by 100 respectively;
[0095] Step 3.4: Determine the system's health status based on the total system score from Step 3.3, specifically as follows:
[0096] When the total system score is greater than or equal to 70, the system status is healthy, and the system service risk is low.
[0097] When the total system score is greater than or equal to 50 and less than 70, the status is sub-healthy, and the system service risk is high.
[0098] When the total system score is less than 50, the status is unhealthy, and the system service risk is extremely high.
[0099] In the specific implementation process, the administrators first retrieve data from the system to obtain information on the sources of risk, as shown in the table below:
[0100]
[0101] The system is scored, and the total system score is calculated as follows: 100 - (0.15 x 100 + 0.12 x 100 + 0.15 x 100) = 58
[0102] The system's health status was determined by its total score of 58, which falls between 50 and 70, indicating a sub-healthy state.
[0103] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for scoring service system risk, applicable to microservice systems, characterized in that, Includes the following steps: Step 1: Establish different risk levels; In step 1, the different risk levels are specifically divided into five risk levels: Level 1 is slightly risky, Level 2 is moderate risky, Level 3 is significant risky, Level 4 is high risky, and Level 5 is extremely dangerous. Step 2: Based on the administrator's historical experience and the system's historical issues, determine the sources of risk in the system and establish the corresponding risk levels for different sources of risk; The sources of risk mentioned in step 2 specifically include: changes, monitoring warnings, performance inspections, and capacity inspections. Step 2, which involves determining the risk level for the aforementioned changes, specifically includes: The risk level is determined based on the number of emergency changes made to the same system within a week; the more emergency changes made, the higher the risk level. The risk level is determined based on the number of times the same system is put into production within a week; the more times it is put into production, the higher the risk level. Step 2, in which the risk level is determined for the aforementioned changes, specifically includes: Emergency change orders that are based on the same system and are not related to production windows are classified as high-risk. Based on the same production schedule, if it is executed multiple times with an interval of A hours, it is judged to be slightly risky. Step 3: Score the system based on the various risk levels it includes and the weight of each individual risk level, and determine the system's health status based on the scores. Step 3 specifically includes the following steps: Step 3.1: Establish the weight of each risk level, specifically: slightly risky at 5%, moderate risk at 8%, significant risk at 12%, high risk at 15%, and extremely dangerous at 60%; Step 3.2: Analyze all risk sources in the system and their corresponding risk levels; Step 3.3: Score the system according to Step 3.1 and Step 3.2, specifically: Total system score = 100 - the sum of the weights of each risk level included in the system multiplied by 100 respectively; Step 3.4: Determine the system's health status based on the total system score from Step 3.3, specifically as follows: When the total system score is greater than or equal to 70, the system status is healthy, and the system service risk is low. When the total system score is greater than or equal to 50 and less than 70, the status is sub-healthy, and the system service risk is high. When the total system score is less than 50, the status is unhealthy, and the system service risk is extremely high.
2. The service system risk scoring method according to claim 1, characterized in that, Step 2, which involves determining the risk level for the monitoring warning, specifically includes: Error warnings that persist for more than C hours within B days are considered significant risks. An emergency warning that remains in effect for more than C hours within B days is classified as high-risk.
3. The service system risk scoring method according to claim 1, characterized in that, Step 2, which involves determining the risk level for the performance inspection, specifically includes: If an application has an average daily CPU usage rate exceeding 80% for at least E days in the past D days, it is considered a significant risk, where E < D. If an application has an average daily disk I / O usage exceeding 80% for at least E days in the past D days, it is considered a significant risk, where E < D. If the peak TPS changes by less than 10% compared to yesterday, and the application causes a CPU utilization increase of more than 20%, it is considered a significant risk. If the peak TPS changes by less than 10% compared to yesterday, but the application causes a memory increase of more than 20%, it is considered a significant risk. If, within F hours, the database causes CPU utilization to exceed 70%, the number of database connections to exceed 70%, and the number of database connections to exceed E minutes, it is considered a significant risk. If, out of G days, CPU utilization exceeds 80% for more than H days due to caching, it is considered a significant risk. If, out of G days, more than H days have memory usage exceeding 80% due to caching, then it is considered a significant risk, where H < G.
4. The service system risk scoring method according to claim 3, characterized in that, Step 2, which involves determining the risk level for the performance inspection, also includes: The risk level is determined based on the number of times the interface was slow yesterday; the more times it was slow, the higher the risk level.
5. The service system risk scoring method according to claim 1, characterized in that, Step 2, which involves determining the risk level for the capacity inspection, specifically includes: If disk usage consistently exceeds 80% within a consecutive day, it is considered a significant risk. If debug-level logs exist in the system, they are considered a significant risk. If the daily log volume of a subsystem exceeds 10GB, it will be classified as a general risk. If a single log entry exceeds 1MB, it is classified as a general risk.
Citation Information
Patent Citations
Method and system for detecting the health degree of data center and storage medium
CN111475377A