A plasma control system health judgment method based on decision table and analytic hierarchy process
By combining decision tables and the analytic hierarchy process, the key performance indicators of the plasma control system are monitored in real time, which solves the problem of accurate assessment of the system health status in existing technologies, realizes global health assessment and fault diagnosis of complex systems, and improves the stability and safety of the system.
Patent Information
- Application Number
- CN202510381300.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Existing plasma control system health assessment methods are difficult to accurately judge the system health status in real time in a complex multi-factor environment. Traditional single-indicator monitoring and rule-based diagnosis methods cannot fully reflect the health status of the system and have difficulty dealing with the interactive effects of multiple modules.
Combining decision tables and the analytic hierarchy process, through real-time monitoring of key performance indicators, a decision table is established and thresholds are set. The system health is calculated using the analytic hierarchy process, the interdependence between modules is considered, and a time factor is introduced to adjust the health status judgment.
It realizes the global quantitative evaluation of the plasma control system, improves the sensitivity of fault detection and the accuracy of diagnosis, prevents the spread of single-point faults, and improves system stability and safety.
Smart Images

Figure CN120234210B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of nuclear fusion tokamak device control systems, and in particular to a plasma control system health judgment method based on a decision table and a hierarchical analysis method. Background Art
[0002] With the increasing complexity of modern industrial control systems, especially in environments with high precision and high stability requirements such as plasma control systems (PCS), real-time assessment of system health, identification of fault sources, and implementation of effective countermeasures are crucial for ensuring stable system operation and experimental safety. Plasma control systems typically involve multiple key subsystems and modules with complex interdependencies. An anomaly or failure in any component can lead to performance degradation or even system crashes of the entire system. Therefore, traditional fault detection and health assessment methods face numerous challenges in dealing with complex system behaviors. Especially in environments with multiple factors acting in parallel and complex component interactions, how to accurately determine system health in real time and take timely measures has become an urgent problem to be solved.
[0003] Existing health assessment technologies fall into two main categories: single-metric fault detection and rule-based fault diagnosis. Single-metric methods assess system health by monitoring key runtime metrics (such as CPU usage and memory utilization). These methods offer advantages in simplicity and computational efficiency, but they typically only detect surface-level issues and struggle to identify potential fault sources arising from the interactions between multiple components in the system. Consequently, when a system fault occurs, monitoring based on a single metric can lead to misjudgments or missed detections, failing to fully reflect the system's health. Rule-based fault diagnosis methods typically use predefined rules to comprehensively assess various system metrics. Decision tables are a common rule-based expression method. They perform logical judgments on multiple system operational metrics and output corresponding fault or health status. While decision tables can quickly locate faults through clear rules and provide effective guidance for further troubleshooting, they also face limitations. First, they rely on manually defined rules and are difficult to handle in complex, multi-dimensional fault modes in a system. Second, when dealing with the interactions of multiple modules, decision tables may overlook the interactions between modules, resulting in inaccurate fault location.
[0004] Therefore, this application proposes a plasma control system health judgment method based on decision table and hierarchical analysis method. Summary of the Invention
[0005] The purpose of the present invention is to address the problem of how to accurately judge the health status of a plasma control system in real time and take timely measures in the background technology, and to propose a plasma control system health judgment method based on a decision table and a hierarchical analysis method.
[0006] The technical solution of the present invention is a method for determining the health of a plasma control system based on a decision table and a hierarchical analysis method, comprising the following steps:
[0007] S1. Determine the health status of the plasma control system based on the decision table;
[0008] S2. Determine the health status of the plasma control system based on the analytic hierarchy process;
[0009] S3. The system health judgment method based on the decision table is combined with the health judgment method based on the hierarchical analysis method to comprehensively judge the health status of the plasma control system.
[0010] Optionally, the step S1 specifically includes the following steps:
[0011] S101. Determine the system operation and maintenance key performance indicators that affect the health of the plasma control system. The English abbreviation KPI is used below to represent the key performance indicators. The KPIs include CPU usage, memory occupancy, system I / O, CPU temperature, disk capacity, and file descriptor utilization:
[0012] S102. Setting KPI thresholds through long-term operation data collection, data waveform analysis, combining operation and maintenance logs with experience, experience-based threshold adjustment, and dynamic adjustment and optimization of thresholds;
[0013] S103. Different KPIs are classified into healthy, subhealthy, and unhealthy states based on set thresholds. Different values of each KPI are defined as input conditions for a decision table. A decision table is established, where the input conditions for the decision table include different health state values of CPU usage, memory utilization, system I / O performance, CPU temperature, disk capacity, and file descriptor utilization. The decision table is used to quickly determine the system health state based on the values of each KPI.
[0014] S104. During system operation, various operation and maintenance KPIs are collected in real time, and the health of the system is judged in combination with the decision table. When the system is in an unhealthy state, an alarm is issued.
[0015] Optionally, in S102,
[0016] Long-term operation data collection: In the early stages of system operation, the system's KPI indicators are monitored over a long period of time. Multi-dimensional data, including CPU usage, memory usage, disk space, and network bandwidth, is collected and analyzed to determine the fluctuation range and change trend of each indicator during normal operation, providing a basis for threshold setting.
[0017] Data waveform analysis: After collecting a large amount of real-time monitoring data, the data waveform is analyzed to determine the normal fluctuation range of each KPI indicator and set corresponding critical values as alarm trigger points. When the indicator fluctuates abnormally or remains under high load for a long time, the system is judged to be in a health risk state;
[0018] Combine operation and maintenance logs with experience: During long-term system monitoring, operation and maintenance logs are used to record system abnormal events, error messages, warnings, and alarm data. These data are then combined with changes in KPI indicators to determine the correlation between the performance of various system indicators and failures or unhealthy conditions.
[0019] Experience-based threshold adjustment: After initially setting the threshold through data analysis and waveform judgment, the threshold is fine-tuned based on the experience of the operation and maintenance personnel and actual work conditions to avoid unnecessary alarms caused by overly strict threshold settings;
[0020] Dynamic adjustment and optimization of thresholds: Based on the continuous operation of the system, changes in business load and hardware environment, thresholds are adjusted in a timely manner to ensure the accuracy of health assessment standards.
[0021] Optionally, in S2, judging the health status of the plasma control system based on the analytic hierarchy process specifically includes the following steps:
[0022] S201, selecting CPU usage, memory occupancy, system I / O, CPU temperature, disk capacity and system file descriptor utilization as indicators for health calculation;
[0023] S202. Quantify each server operation and maintenance KPI into parameters , by calculating the maximum eigenvalue and corresponding eigenvector of the comparison matrix, and normalizing the eigenvector to obtain the weight of each indicator;
[0024] S203, performing a consistency test, wherein the consistency index and consistency ratio are calculated. When the consistency ratio is less than 0.1, the consistency requirement is met;
[0025] S204. Calculate the system health based on the weight of each indicator and the distance from the critical value, combined with the time factor. When the health is less than 90, the system is determined to be in a sub-healthy state. When the health is less than 60, the system is determined to be in an unhealthy state.
[0026] Optionally, in S202, the comparison matrix is:
[0027] ,
[0028] By calculation, the maximum eigenvalue of the comparison matrix is obtained: ,
[0029] The corresponding eigenvectors are: ,
[0030] Normalize the eigenvectors: ,
[0031] The weights of each indicator are as follows: .
[0032] Optional, in S203, consistency index: ,
[0033] in, is the maximum eigenvalue of the comparison matrix, n is the number of operation and maintenance KPIs, ;
[0034] RI is the random consistency index, and its values are shown in the following table: ,
[0035] Calculate the available consistency ratio: ,
[0036] A consistency ratio less than 0.1 indicates that consistency is met and the health of the system is: ,
[0037] in, is a normalized eigenvector, 1≤i≤6, each i corresponds to an indicator, Represents the unhealthy state of a single indicator. When the indicator is less than the critical value and keeps approaching or greater than the critical value, the healthiness decreases. When the healthiness is less than 90, the system is considered to be in a sub-healthy state. At this time, the time factor It starts to work, and increases by 1 every cycle. When calculating, the time factor is multiplied by the indicator weight. , when the indicator is less than the critical value and far enough away from the critical value, the time factor , otherwise the system is considered to be in sub-healthy state.
[0038] Optionally, in S3, the comprehensive judgment of the health status of the plasma control system specifically includes the following steps:
[0039] S301. When the system is running, first use a decision table-based method to determine the system health status. If the system is healthy, continue to run, and use the analytic hierarchy process to calculate the system health. If the system is sub-healthy, log the information, continue to run, and use the analytic hierarchy process to calculate the system health. If the system is unhealthy, an alarm is issued and the operation and maintenance personnel are notified.
[0040] S302. Calculate the system health using the AHP method. If the AHP method determines that the system is unhealthy, an alarm is generated and the operation and maintenance personnel are notified. If the AHP method determines that the system is sub-healthy, a system log is recorded and the system continues to operate. If the AHP method determines that the system is healthy, no log is recorded and the system continues to operate.
[0041] Compared with the prior art, this application has at least one of the following beneficial technical effects:
[0042] By introducing a hierarchical analysis algorithm, the interdependencies and impacts of each module in the system are comprehensively considered, and the module's contribution to the overall health status is quantified. This addresses the shortcomings of traditional single-metric monitoring and decision tables in global health assessment. It enables a global quantitative assessment of system health and effectively reflects the interactive impacts between modules.
[0043] Decision tables are used to quickly locate single faults and provide clear fault diagnosis results. A hierarchical analysis algorithm is used for global health assessment, combining dynamic indicators and weight changes to adjust health status judgment criteria in real time. This combination not only improves fault detection sensitivity but also provides a flexible diagnostic tool for system health assessment in complex environments.
[0044] This invention introduces a time factor as a dynamic weight in health assessment, reflecting the changing trend of system health over the duration of an anomaly. This approach gradually reduces system health as system metrics continuously approach or exceed critical values, significantly improving the ability to detect dynamic anomalies.
[0045] In view of the multi-module and highly interactive operating characteristics of the plasma control system, the present invention can simulate the association relationships between modules and analyze the impact of these relationships on the overall health of the system through a hierarchical analysis algorithm, thereby achieving effective health assessment of complex interactive scenarios and preventing system crashes caused by the spread of single-point failures.
[0046] By combining the advantages of decision tables and hierarchical analysis methods, the present invention overcomes the limitations of existing technologies in single indicator monitoring, regularized diagnosis, inter-module interaction analysis, and dynamic response capabilities, and realizes real-time health assessment and accurate fault diagnosis of complex distributed control systems. It significantly improves the stability, reliability, and safety of the plasma control system (PCS), providing strong support for the efficient operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a flow chart of a plasma control system health judgment method based on decision table and hierarchical analysis method;
[0048] Figure 2 This is a diagram of the health calculation test results of this embodiment. DETAILED DESCRIPTION
[0049] The technical solution of the present invention is further described below with reference to the accompanying drawings and specific embodiments. Example
[0050] During the operation of a plasma control system, real-time assessment of system health and rapid identification of fault sources are key to ensuring system stability and experimental safety. To achieve this goal, the present invention introduces a system health assessment algorithm based on decision tables and a hierarchical analysis method. The decision table classifies and judges system operating data using predefined rules, providing fast and efficient preliminary fault location. The hierarchical analysis method, based on the relationships between system components, calculates the health impact weight of each component, achieving a global quantitative assessment of the system's health status. This algorithmic design, combining regularization and weighting, not only improves the accuracy of fault detection but also provides a solution for health assessment of complex distributed systems that balances real-time and comprehensiveness. This solution is described in detail below.
[0051] System health judgment method based on decision table:
[0052] Step 1: Identify system operation and maintenance KPIs (key performance indicators) that affect server health
[0053] When determining system operation and maintenance KPIs that impact server health, we selected the following six key indicators as the basis for health calculation. These indicators directly reflect server resource usage and operational stability: CPU utilization, memory usage, system I / O (input / output), CPU temperature, disk capacity, and file descriptors. We then determined thresholds for these KPIs; when these thresholds are reached, the system is considered unhealthy.
[0054] When assessing system health, simply monitoring changes in individual KPIs in real time cannot fully reflect the system's status, especially in complex production environments where the system may experience varying workloads and operating modes. To accurately determine system health, appropriate thresholds must be set. When individual KPIs reach or exceed these critical values, potential health risks can be promptly identified, preventing undetected system failures.
[0055] Steps and basis for threshold setting:
[0056] (1) Long-term operation data collection
[0057] Data collection and monitoring are crucial during the early stages of system operation. To ensure the accuracy of thresholds, it's often necessary to monitor system KPIs over a long period of time. By collecting and analyzing multi-dimensional data such as CPU usage, memory utilization, disk space, and network bandwidth, we can assess system performance under varying loads and operating conditions. This long-term data collection allows us to capture the fluctuation range and changing trends of various indicators during normal operation, providing a scientific basis for threshold setting.
[0058] (2) Data waveform analysis
[0059] Data waveforms are an important indicator of system operating status. After collecting large amounts of real-time monitoring data, analyzing the waveforms can help operations personnel better understand the fluctuation patterns of various indicators under normal operating conditions. For example, CPU usage and memory utilization typically increase with increasing system load, but significant fluctuations or prolonged periods of high load may indicate that the system is entering a health risk state. Waveform analysis can determine the normal fluctuation range of each KPI indicator and set appropriate critical values as alarm triggers.
[0060] (3) Combining operation and maintenance logs with experience
[0061] Operations and maintenance logs are another important data source that cannot be ignored. When monitoring a system over a long period of time, logs are often used to record system abnormal events, error messages, warnings, alerts, and other data. This log data, combined with changes in KPI indicators, can help operations and maintenance personnel better understand how the performance of various system indicators is associated with failures or unhealthy conditions in specific situations. For example, when disk space utilization consistently exceeds 80%, the system may trigger a "disk full" warning. Combined with data such as I / O performance and system load at this time, it can effectively determine whether the anomaly is caused by a system bottleneck or hardware failure.
[0062] (4) Threshold adjustment based on experience
[0063] While initial threshold settings can be accomplished through data analysis and waveform analysis, final threshold settings often require adjustment based on the operator's experience and actual workload. For example, in certain application scenarios, abnormal fluctuations in certain KPIs may occur without immediately causing system failure. In these cases, operators need to fine-tune thresholds based on their experience to avoid unnecessary alerts caused by overly strict threshold settings. Accumulated experience enables operators to better determine which fluctuations are acceptable and which represent potential risks that could lead to system failure.
[0064] (5) Dynamic adjustment and optimization of thresholds
[0065] The system's operating status changes dynamically, so threshold settings are not static. As the system continues to operate, business loads change, and the hardware environment evolves, the system's health assessment criteria also need to be adjusted appropriately. For example, if a business load increases, the system's CPU usage and memory utilization thresholds may need to be raised appropriately, while after a system hardware upgrade, the disk I / O threshold may need to be lowered accordingly. Therefore, threshold setting is a dynamic optimization process that requires regular review and adjustment based on the system's operating status and historical data.
[0066] Step 2: Create a decision table based on the selected server operation and maintenance KPI and the set threshold
[0067] Building a decision table based on the selected server operation and maintenance KPIs and set thresholds is a crucial step in the health assessment process. This table helps the system quickly and accurately determine the health status of each KPI and make appropriate operational decisions based on this health status. To achieve this goal, it's necessary to consider the different value ranges of the KPIs and the specific thresholds for each KPI, ultimately forming a regularized decision table.
[0068] (1) Determine KPI indicators and their value ranges
[0069] First, according to the KPI thresholds selected in the second step, different KPIs are divided into healthy state, sub-healthy state and unhealthy state.
[0070] (2) Define the input conditions of the decision table
[0071] The core of a decision table is to combine the different values of multiple KPIs with preset thresholds to quickly determine the health status of the system. Based on this, the different values of each KPI can be defined as the input conditions of the decision table.
[0072] CPU usage There are three values: healthy, sub-healthy, and unhealthy;
[0073] Memory usage There are three values: healthy, sub-healthy, and unhealthy;
[0074] System I / O performance There are two values: healthy and unhealthy;
[0075] CPU temperature There are three values: healthy, sub-healthy, and unhealthy;
[0076] Disk capacity There are three values: healthy, sub-healthy, and unhealthy;
[0077] File descriptor utilization There are three values: healthy, sub-healthy, and unhealthy;
[0078] These conditions are combined into different rules to determine system health.
[0079] In a decision table, the input conditions list the values for each KPI, and the decision result is an assessment of the system's health based on these values. Assuming the output in the decision table is "Healthy" or "Unhealthy," the system's health status will be determined as Healthy, Sub-Healthy, or Unhealthy based on the values of each KPI.
[0080] Rule Number CPU usage (C1) Memory usage (C2) I / O performance (C3) CPU temperature (C4) Disk capacity (C5) File descriptor (C6) Health status assessment 1 healthy healthy healthy healthy healthy healthy healthy 2 healthy healthy healthy Sub-health healthy healthy healthy 3 healthy Sub-health healthy healthy healthy healthy healthy 4 healthy healthy healthy unhealthy healthy healthy unhealthy 5 healthy healthy unhealthy healthy healthy healthy unhealthy 6 healthy Sub-health healthy healthy healthy healthy healthy 7 unhealthy healthy healthy healthy healthy healthy unhealthy 8 Sub-health healthy healthy healthy healthy healthy unhealthy 9 unhealthy unhealthy healthy healthy unhealthy healthy unhealthy 10 healthy healthy healthy healthy unhealthy healthy unhealthy ;
[0081] By establishing a decision table, we can quickly determine the system health status based on the values of various KPIs, providing a clear basis for subsequent fault diagnosis. This threshold- and rule-based judgment method enables the system to respond to various anomalies in real time, issuing alerts and handling them promptly, thereby improving system reliability and operation and maintenance efficiency.
[0082] The third step is to collect various operation and maintenance KPIs in real time during system operation and judge the system health status based on the decision table.
[0083] Based on the above judgment method, when the system is running, the system can collect various operation and maintenance KPIs and combine them with the decision table to judge the health status of the system. If the system is in an unhealthy state, an alarm will be issued.
[0084] System health judgment method based on hierarchical analysis:
[0085] Decision tables can quickly locate single faults and provide clear judgments, but they struggle to assess the combined impact of interactions among multiple modules on the overall health of a complex system. To address this, the Analytic Hierarchy Process (AHP) health calculation method dynamically analyzes the dependencies and state changes between system modules to calculate global health, providing a more comprehensive health assessment.
[0086] The steps to calculate system health using the AHP method are as follows:
[0087] In the first step, CPU usage, memory occupancy, system I / O, CPU temperature, disk capacity, and system file descriptor utilization are selected as indicators for health calculation.
[0088] The second step is to quantify the various operation and maintenance KPIs of the server into parameters , the comparison matrix is: ,
[0089] The third step is to calculate the maximum eigenvalue of the comparison matrix: ,
[0090] The corresponding eigenvectors are: ,
[0091] Normalize the eigenvectors: ,
[0092] The weights of each indicator are shown in Table 1: ,
[0093] In the fourth step, since the importance is subjectively assessed by experts and there are certain objective errors, a consistency test is needed.
[0094] Consistency indicators: ,
[0095] in, is the maximum eigenvalue of the comparison matrix, n is the number of operation and maintenance KPIs, ;
[0096] RI is the random consistency index, and its values are shown in the following table: ,
[0097] Calculate the available consistency ratio: ,
[0098] A consistency ratio less than 0.1 indicates that consistency is met and the health of the system is: ,
[0099] Among them, 1≤i≤6, each i corresponds to an indicator, Represents the unhealthiness of a single indicator. When the indicator is less than the critical value, the closer the indicator is to the critical value, the greater the unhealthiness of the single indicator. Conversely, when the indicator is greater than the critical value, the farther the indicator is from the critical value, the greater the unhealthiness of the single indicator. When calculating, the time factor is multiplied by the indicator weight. , when the indicator is less than the critical value and far enough away from the critical value, Otherwise, the system is considered to be in a sub-healthy state. . .
[0100] During the five-step operation, the health of the system is calculated according to the above method. If the health is less than 90, the system is considered to be in a sub-healthy state. If it is less than 60, the system is considered to be in an unhealthy state.
[0101] Combining the system health judgment method based on decision table with the health judgment method based on analytic hierarchy process:
[0102] During actual operation, the system collects the above operation and maintenance KPIs and inputs these data into the system health judgment method based on the decision table. This method quickly determines whether any indicators exceed the threshold. If there are unhealthy indicators, an alarm will be issued to notify the operation and maintenance personnel. If they are healthy or sub-healthy,
[0103] The KPI is fed into the hierarchical analysis method and the health is calculated. If the hierarchical analysis method determines that the system is unhealthy, an alarm is issued. If the system is sub-healthy, a log is recorded and the system continues to run. If the system is healthy, no log is recorded and the system continues to run. This cycle continues until the end of the run. The overall flow chart is as follows: Figure 1 shown.
[0104] In order to verify the technical effect of the present invention, the embodiment is verified. A piece of memory leak code is deployed in the system to make the memory occupancy rate gradually deviate from the normal value. At the same time, the disk space occupancy rate is increased. The health waveform of each indicator and the entire system is calculated as follows: Figure 2 As shown. Figure 2 It can be seen that after the disk space occupancy rate is greater than 90%, the system is in a sub-healthy state, and the time factor begins to work. The longer the time, the more severe the reduction in the health of the disk space occupancy rate. Similarly, with the occurrence of memory leaks, the health of the memory occupancy rate gradually decreases with the increase of time. The two work together to make the overall health of the system less than 90 points when the memory occupancy rate is greater than 90%, and continue to decrease with the increase of the failure time, and finally less than 60 points, indicating that the system has entered a dangerous state. The above experiments show that the health assessment algorithm that introduces the time factor can effectively reflect the dynamic changes in the health of the system. When a single indicator (such as memory occupancy rate or disk space occupancy rate) gradually deviates from the normal value, the time factor The cumulative effect of these metrics can significantly increase sensitivity to sub-health conditions. Results show that when key indicators exceed critical values, the rate of decline in system health accelerates over time, ultimately signaling a potential system crisis. This algorithm not only accurately quantifies the health of a single indicator but also uses multiple indicators to evaluate the evolution of the system's overall operational status through comprehensive health calculations, providing timely fault warnings to operations and maintenance personnel.
[0105] The above specific embodiments are merely several optional embodiments of the present invention. Based on the technical solutions of the present invention and the relevant inspirations of the above embodiments, those skilled in the art may make various alternative improvements and combinations to the above specific embodiments.
Claims
1. A plasma control system health judgment method based on decision table and hierarchical analysis method, characterized in that: The following steps are involved: S1. Determine the health status of the plasma control system based on the decision table; S2. Determine the health status of the plasma control system based on the analytic hierarchy process; S3. Combine the system health judgment method based on the decision table with the health judgment method based on the analytic hierarchy process to comprehensively judge the health status of the plasma control system; In S2, judging the health status of the plasma control system based on the analytic hierarchy process specifically includes the following steps: S201, selecting CPU usage, memory occupancy, system I / O, CPU temperature, disk capacity and system file descriptor utilization as indicators for health calculation; S202. Quantify each server operation and maintenance KPI into parameters , by calculating the maximum eigenvalue and corresponding eigenvector of the comparison matrix, and normalizing the eigenvector to obtain the weight of each indicator; S203, performing a consistency test, wherein the consistency index and consistency ratio are calculated. When the consistency ratio is less than 0.1, the consistency requirement is met; S204. Calculate the system health based on the weight of each indicator and the distance from the critical value, combined with the time factor. If the health is less than 90, the system is considered to be in a sub-healthy state. If the health is less than 60, the system is considered to be in an unhealthy state. In S203, the consistency index is: , in, is the largest eigenvalue of the comparison matrix, is the number of operation and maintenance KPIs, ; RI is the random consistency index, and its values are shown in the following table: Calculate the available consistency ratio: , A consistency ratio less than 0.1 indicates that consistency is satisfied, and the health of the system is: , in, is the normalized eigenvector, 1≤i≤6, each For an indicator, Represents the unhealthy state of a single indicator. When the indicator is less than the critical value and keeps approaching or greater than the critical value, the healthiness decreases. When the healthiness is less than 90, the system is considered to be in a sub-healthy state. At this time, the time factor It starts to work, and increases by 1 every cycle. When calculating, the time factor is multiplied by the indicator weight. , when the indicator is less than the critical value and far enough away from the critical value, the time factor , otherwise the system is considered to be in a sub-healthy state; In S3, the comprehensive judgment of the health status of the plasma control system specifically includes the following steps: S301. When the system is running, first use a decision table-based method to determine the system health status. If the system is healthy, continue to run, and use the analytic hierarchy process to calculate the system health. If the system is sub-healthy, log the information, continue to run, and use the analytic hierarchy process to calculate the system health. If the system is unhealthy, an alarm is issued and the operation and maintenance personnel are notified. S302. Calculate the system health using the AHP method. If the AHP method determines that the system is unhealthy, an alarm is generated and the operation and maintenance personnel are notified. If the AHP method determines that the system is sub-healthy, a system log is recorded and the system continues to operate. If the AHP method determines that the system is healthy, no log is recorded and the system continues to operate.
2. The plasma control system health judgment method based on decision table and hierarchical analysis method according to claim 1 is characterized in that: Said S1 specifically includes the following steps: S101. Determine system operation and maintenance key performance indicators that affect the health of the plasma control system. The English abbreviation KPI is used below to represent key performance indicators. The KPIs include CPU usage, memory occupancy, system I / O, CPU temperature, disk capacity, and file descriptor utilization. S102. Setting KPI thresholds through long-term operation data collection, data waveform analysis, combining operation and maintenance logs with experience, experience-based threshold adjustment, and dynamic adjustment and optimization of thresholds; S103. Different KPIs are classified into healthy, subhealthy, and unhealthy states based on set thresholds. Different values of each KPI are defined as input conditions for a decision table. A decision table is established, where the input conditions for the decision table include different health state values of CPU usage, memory utilization, system I / O performance, CPU temperature, disk capacity, and file descriptor utilization. The decision table is used to quickly determine the system health state based on the values of each KPI. S104. During system operation, various operation and maintenance KPIs are collected in real time, and the health of the system is judged in combination with the decision table. When the system is in an unhealthy state, an alarm is issued.
3. The plasma control system health judgment method based on decision table and hierarchical analysis method according to claim 2 is characterized in that: In the above S102, Long-term operation data collection: In the early stage of system operation, the system operation KPI indicators are monitored for a long time, and multi-dimensional data including CPU usage, memory usage, disk space, and network bandwidth are collected and analyzed; Data waveform analysis: After collecting a large amount of real-time monitoring data, the data waveform is analyzed to determine the normal fluctuation range of each KPI indicator and set corresponding critical values as alarm trigger points. When the indicator fluctuates abnormally or remains under high load for a long time, the system is judged to be in a health risk state; Combine operation and maintenance logs with experience: During long-term system monitoring, operation and maintenance logs are used to record system abnormal events, error messages, warnings, and alarm data. These data are then combined with changes in KPI indicators to determine the correlation between the performance of various system indicators and failures or unhealthy conditions. Experience-based threshold adjustment: After initially setting the threshold through data analysis and waveform judgment, the threshold is fine-tuned based on the experience of the operation and maintenance personnel and actual work conditions; Dynamic adjustment and optimization of thresholds: timely adjust thresholds based on the system's continuous operation, business load changes, and hardware environment changes.
4. The plasma control system health judgment method based on decision table and hierarchical analysis method according to claim 1 is characterized in that: In S202, the comparison matrix is: , By calculation, the maximum eigenvalue of the comparison matrix is obtained: , The corresponding eigenvectors are: , Normalize the eigenvectors: , The weights of each indicator are as follows: 。
Citation Information
Patent Citations
Electric power information system health degree assessment method and assessment system thereof
CN110070461A