Application health assessment method and device, electronic equipment and storage medium
By acquiring and normalizing multi-dimensional monitoring data, and combining application profiling and intelligent analysis, the problem of data silos in traditional IT service management has been solved, enabling accurate assessment and real-time reflection of application health, and improving operational efficiency and intelligence.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PICC INFORMATION TECH CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-12
Smart Images

Figure CN122019325A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of application operation and maintenance, and in particular to a method, apparatus, electronic device and storage medium for health assessment of an application. Background Technology
[0002] Traditional IT service management systems primarily revolve around predefined, standardized processes for application operation and maintenance. Within this system, application monitoring data typically originates from multiple heterogeneous and independent tool platforms such as Zabbix, ELK, and APM. Common operational solutions focus on passive alerting and post-incident manual investigation, essentially amounting to a simple aggregation of various monitoring tools and data.
[0003] However, due to the lack of effective data linkage and information fusion between different tool platforms, operations and maintenance personnel often need to perform data correlation analysis across multiple systems, which is inefficient and makes it difficult to form a comprehensive and timely application health status assessment. At the same time, existing operations and maintenance technical solutions not only exhibit partiality but also show significant passive response and result lag, making it impossible for operations and maintenance teams to make real-time and comprehensive health assessments of the system. Furthermore, health assessments mainly rely on human experience to judge various technical indicators, making it difficult to comprehensively cover and objectively reflect the true health status centered on user experience. Summary of the Invention
[0004] In view of this, embodiments of this application provide a health assessment method, apparatus, electronic device, and storage medium for applications, which improves the accuracy, practicality, and intelligence of application health assessment, and reduces reliance on human labor and operational risks.
[0005] This application mainly includes the following aspects: In a first aspect, embodiments of this application provide an applied health assessment method, the health assessment method comprising: Obtain monitoring indicator data for the target application across multiple health dimensions; For each health dimension, based on the baseline parameters corresponding to that health dimension, the monitoring indicator data corresponding to that health dimension is normalized to obtain the health sub-score corresponding to that health dimension. Based on the application profile of the target application, determine the dynamic weight corresponding to each health dimension; The health score of the target application is determined based on the health sub-score and dynamic weight corresponding to each health dimension.
[0006] The application profile includes: static attributes and dynamic attributes of the target application; the determination of the dynamic weight corresponding to each health dimension based on the application profile of the target application includes: Based on the static attributes, the basic weights corresponding to each health dimension are matched from the weight template library; Based on the priority of the dynamic attributes, the adjustment operator for the basic weight corresponding to each health dimension is determined; The basic weights corresponding to each health dimension are adjusted according to the corresponding adjustment operators to obtain the dynamic weights corresponding to each health dimension.
[0007] Furthermore, the adjustment operator for determining the basic weights corresponding to each health dimension based on the priority of the dynamic attributes includes: Based on the priority of the dynamic attributes, determine the first adjustment value of the basic weight corresponding to each health dimension; The historical health data and historical event record data of the target application within the first preset time period are associated according to the preset association rules to generate event diagnosis conclusions of the target application within the first preset time period. Determine whether an event anomaly instruction is detected that indicates the event diagnosis conclusion is abnormal; If the event abnormality command is detected, then in response to the event abnormality command, key features are determined from the event diagnosis conclusion; Based on the aforementioned key features, a second adjustment value for the basic weight corresponding to each health dimension is determined; Based on the first and second adjustment values of the basic weights corresponding to each health dimension, determine the adjustment operator for the basic weights corresponding to each health dimension; If no abnormal event instruction is detected, the first adjustment value of the basic weight corresponding to each health dimension is determined as the adjustment operator of the basic weight corresponding to each health dimension.
[0008] Furthermore, the health assessment method also includes: In response to a score anomaly instruction used to indicate an abnormality in health status, the magnitude of change in historical monitoring indicator data corresponding to each health dimension within a second preset time period is determined. Based on the magnitude of the change, determine the health dimensions that cause the total health score to be abnormal; Based on a preset adjustment step size, a new benchmark parameter corresponding to the health dimension that causes the health status to be abnormal is determined, and the original benchmark parameter corresponding to the health dimension that causes the health status to be abnormal is replaced with the new benchmark parameter.
[0009] Furthermore, the health assessment method also includes: The historical health data, the historical events, and the event diagnostic conclusions are rendered to generate the first view in the visualization interface corresponding to the target application.
[0010] Furthermore, the health assessment method also includes: The health data of the target application is rendered to generate a second view in the corresponding visualization interface of the target application.
[0011] Furthermore, the health data includes: the health score of the target application, the health sub-score corresponding to each health dimension, and the dynamic weight corresponding to each health dimension.
[0012] Secondly, embodiments of this application also provide an application-specific health assessment device, the health assessment device comprising: The acquisition module is used to acquire monitoring indicator data for the target application across multiple health dimensions. The processing module is used to normalize the monitoring indicator data corresponding to each health dimension based on the benchmark parameters corresponding to that health dimension, and obtain the health sub-score corresponding to that health dimension. The dynamic weight determination module is used to determine the dynamic weight corresponding to each health dimension based on the application profile of the target application. The health determination module is used to determine the health of a target application based on the health sub-score and dynamic weight corresponding to each health dimension.
[0013] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the machine-readable instructions are executed by the processor to perform the steps of the health assessment method of the application described in the first aspect or any possible implementation of the first aspect.
[0014] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the health assessment method of the application described in the first aspect or any possible implementation of the first aspect.
[0015] This application provides a method, apparatus, electronic device, and storage medium for assessing the health of an application. The method acquires monitoring indicator data corresponding to a target application across multiple health dimensions. For each health dimension, based on the baseline parameters corresponding to that health dimension, the monitoring indicator data is normalized to obtain a health sub-score for that health dimension. Based on the application profile of the target application, the dynamic weight corresponding to each health dimension is determined. Based on the health sub-score and dynamic weight corresponding to each health dimension, the health level of the target application is determined.
[0016] This improves the accuracy, practicality, and intelligence of health assessment applications, while reducing reliance on human labor and operational risks.
[0017] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart of a health assessment method provided in an embodiment of this application is shown; Figure 2 This paper illustrates an architecture diagram of a health assessment method provided in an embodiment of this application. Figure 3 This invention provides a schematic diagram of the structure of a health assessment device for an application according to an embodiment of this application. Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0021] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0022] The methods, apparatus, electronic devices, or computer-readable storage media described in this application embodiment can be applied to any scenario requiring application maintenance. This application embodiment does not limit specific application scenarios, and any scheme using the health assessment methods and apparatus provided in this application embodiment is within the protection scope of this application.
[0023] It's worth noting that traditional IT service management systems primarily revolve around predefined, standardized processes for application operation and maintenance. Within this system, application monitoring data typically originates from multiple heterogeneous and independent tool platforms such as Zabbix, ELK, and APM. Common operation and maintenance solutions focus on passive alerts and post-incident manual investigation, essentially amounting to a simple aggregation of various monitoring tools and data. However, due to the lack of effective data linkage and information fusion between different tool platforms, operations personnel often need to perform data correlation analysis across multiple systems, resulting in low efficiency and difficulty in forming a comprehensive and timely application health status assessment. Furthermore, existing operation and maintenance solutions not only exhibit partiality but also demonstrate significant passive response and delayed results, making it impossible for operations teams to make real-time, comprehensive health assessments of the system. In addition, health assessments rely heavily on human experience in judging various technical indicators, making it difficult to comprehensively cover and objectively reflect the true health status centered on user experience.
[0024] To address the aforementioned issues, this application proposes an application health assessment method, apparatus, electronic device, and storage medium, which improves the accuracy, practicality, and intelligence of application health assessment, while reducing reliance on human labor and operational risks.
[0025] To facilitate understanding of this application, the technical solutions provided in this application will be described in detail below with reference to specific embodiments.
[0026] In this application embodiment, against the backdrop of the increasing popularity of cloud-native and microservice architectures, traditional IT service management systems are facing significant challenges. Monitoring data within these systems is isolated and difficult to correlate. This data state makes the assessment of application system health often one-sided, passive, and lagging, failing to meet the operational needs of modern dynamic and distributed environments. Currently, common technical solutions in the industry mostly focus on passive alarms and post-incident investigations, essentially remaining a simple accumulation of various tools and data, failing to build a health management system centered on applications and covering their entire lifecycle. Specifically: First, alarm methods based on fixed thresholds and rules, while capable of static monitoring of single indicators such as CPU utilization, are prone to triggering alarms due to isolated and uncorrelated data, and cannot achieve fault prediction; Second, while horizontally aggregating data from multiple systems and displaying it on a unified dashboard provides a comprehensive view, it is limited to data presentation and lacks in-depth analysis and root cause localization capabilities; Third, while introducing dynamic baselines and time-series data analysis can identify anomalies deviating from historical patterns, it is mostly limited to technical indicators, with weak correlation to business health, and problem localization still heavily relies on manual experience. Therefore, current operation and maintenance practices generally suffer from the following prominent problems: severe data silos, with systems such as basic monitoring, logs, application performance management, IT service management, and cost data operating independently, forming data silos and resulting in low efficiency in information association; subjective and one-sided health assessments, relying on limited indicators and experience-based judgments, failing to fully reflect the real application status centered on user experience; lack of unified quantitative health indicators and forward-looking analysis capabilities, making preventative intervention difficult; and when a failure occurs, performance anomalies are disconnected from recent changes and configurations, resulting in extremely high costs for locating and reviewing the problem.
[0027] Please see Figure 1 , Figure 1 This is a flowchart of a health assessment method provided in an embodiment of this application.
[0028] like Figure 1 As shown in the embodiments of this application, the health assessment method for the application includes the following steps: Step S101: Obtain monitoring indicator data for the target application across multiple health dimensions.
[0029] Here, as Figure 2 As shown, health dimensions may include, but are not limited to: security dimension, performance dimension, resource dimension, business impact dimension, and availability dimension.
[0030] Monitoring metrics for the security dimension include: number of vulnerabilities, security incidents, etc.; monitoring metrics for the performance dimension include: response time, throughput, error rate, error logs, exception stack traces, etc.; monitoring metrics for the resource dimension include: queries per second, number of slow queries, execution time, connection pool utilization, number of waiting queries, server CPU, memory, and disk utilization, etc.; monitoring metrics for the business impact dimension include: number of business transactions per second and number of calls to business interfaces, etc.; and monitoring metrics for the availability dimension include: request success rate and response time, etc.
[0031] Specifically, monitoring metrics data for all health dimensions are obtained through the following tools: emergency response platform, APM (Application Performance Management) tool, ELK (Elasticsearch, Logstash, Kibana) tool, DPM (Database Performance Monitoring) tool, Zabbix (enterprise-grade distributed monitoring solution) tool, BPC (Business Traffic Monitoring) tool, and availability probing tool. The emergency response platform obtains data such as the number of vulnerabilities and security incidents; APM obtains data such as response time, throughput, and error rate; ELK obtains data such as error logs and exception stack traces; Zabbix obtains data such as server CPU, memory, and disk usage; BPC obtains data such as the number of business transactions per second and the number of business interface calls; and availability probing tool obtains data such as request success rate and response time. The availability probing tool measures request success rate and response time by periodically sending simulated requests (such as HTTP(S) / TCP / Ping) to the application from external network nodes.
[0032] Step S102: For each health dimension, based on the benchmark parameters corresponding to that health dimension, normalize the monitoring indicator data corresponding to that health dimension to obtain the health sub-score corresponding to that health dimension.
[0033] Here, the monitoring indicator data corresponding to each health dimension are normalized and calibrated to a range of 0-100 points.
[0034] For the security dimension, the baseline parameter is the event deduction value. In the security dimension, points are deducted based on the number of security events of different levels until the score drops to 0. Specifically, the monitoring indicator data corresponding to the security dimension is normalized as follows: Within the set evaluation period, firstly, a list of security events for the target application is obtained. Each security event record must include at least the event type and event severity level (high-risk, medium-risk, low-risk). Then, according to predefined deduction rules, each security event is mapped to a specific deduction value. For example, high-risk events (e.g., data breaches, core system intrusions) deduct 30 points, medium-risk events (e.g., malicious scanning, sensitive information leakage risks) deduct 10 points, and low-risk events (e.g., low-risk vulnerability alerts) deduct 3 points. Subsequently, the deduction values of all security events within the evaluation period are summed to obtain the cumulative deduction. Then, the formula Ssec=max(0,100) is applied. (Cumulative deductions) Calculate the health sub-score corresponding to the safety dimension, ensuring the score is not lower than 0, where Ssec is the health sub-score corresponding to the safety dimension. For example, if one high-risk event and two medium-risk events occur recently, the cumulative deduction = 1 × 30 + 2 × 10 = 50 points, and the safety score Ssec = max(0, 100). 50) = 50 points.
[0035] For the performance dimension, the benchmark parameters corresponding to the performance dimension are SLO (Service Level Objective) and SLA (Service Level Agreement). In the performance dimension, SLO and service SLA are set for each monitoring metric data corresponding to the performance dimension, and linear scoring is performed according to the interval where the monitoring metric data is located. Specifically, the monitoring metric data corresponding to the performance dimension is normalized in the following way: within the set evaluation period, first, for monitoring metrics where the smaller the value, the better (such as response time, error rate, etc.), if the monitoring metric data ≤ SLO value, the score = 100; if the monitoring metric data ≥ SLA value, the score = 0; if SLO value < monitoring metric data < SLA value, the score is calculated by linear interpolation: score = 100 × [(SLA value - monitoring metric data) / (SLA value - SLO value)]. For monitoring metrics where the larger the value, the better (such as throughput, etc.), if the monitoring metric data ≥ SLO value, the score = 100; if the monitoring metric data ≤ SLA value, the score = 0; if SLA value < monitoring metric data < SLO value, the score is calculated by linear interpolation: score = 100 × [(monitoring metric data - SLA value) / (SLO value - SLA value)]. Then, each monitoring metric data corresponding to the performance dimension corresponds to a weight, and the scores of each monitoring metric data corresponding to the performance dimension are weighted and summed to obtain the health sub-score corresponding to the performance dimension.
[0036] For the resource dimension, the corresponding baseline parameters are the resource deduction value, warning value, and safety value. Within the resource dimension, warning and danger values are set for each monitoring indicator data point. When the monitoring indicator data is within the safety range, no points are deducted. Once it exceeds the warning value, points are deducted linearly. If it exceeds the danger value, all scores for the monitoring indicator data are reduced to 0. Specifically, the monitoring indicator data for the resource dimension is normalized as follows: Within the set evaluation period, firstly, if the monitoring indicator data ≤ the warning value, then the deduction = 0; if the warning value < the monitoring indicator data ≤ the danger value, then the deduction is calculated using linear interpolation: Deduction = (Monitoring indicator data - Warning value) / (Danger value - Warning value) × the maximum deduction corresponding to the monitoring indicator data; if the monitoring indicator data > the danger value, then the deduction = the maximum deduction corresponding to the monitoring indicator data. Then, the deductions for all monitoring indicator data are summed to obtain the cumulative deduction. Subsequently, the formula Sr=MAX(0, 100 - total deductions) is applied to calculate the health sub-score corresponding to the resource dimension, where Sr is the health sub-score corresponding to the security dimension. As an example, taking memory and CPU utilization as examples, the warning value for both memory and CPU utilization is set to 80%, the danger value to 95%, and the maximum deduction to 30 points for both. Currently, the CPU utilization is 85%, and the memory utilization is 90%. The CPU deduction is calculated as follows: CPU deduction = (85-70) / (90-70)×30 = 22.5 points. The memory deduction is calculated as follows: (90-80) / (95-80)×30 = 20 points. Sr = 100 - (22.5 + 20) = 57.5 points.
[0037] In the business impact dimension, the baseline parameter corresponding to the resource dimension is the reduction threshold. The health sub-score corresponding to the business impact dimension is determined by the reduction magnitude of the monitoring indicator data. Specifically, the monitoring indicator data corresponding to the business impact dimension is normalized as follows: Within a set evaluation period, firstly, the baseline value of the target application's business indicator is obtained, where the baseline value is calculated by calculating the monitoring indicator data within the past first predetermined period (e.g., two weeks). Then, the current business indicator value is determined, where the current business indicator value is calculated by calculating the monitoring indicator data within the past second predetermined period (e.g., 5 minutes). Subsequently, the health sub-score corresponding to the business impact dimension is calculated using the formula Sbus = monitoring indicator data / business indicator baseline value × 100, where Sbus is the health sub-score corresponding to the business impact dimension. In this embodiment, a reduction threshold is set. When 1 - monitoring indicator data / business indicator baseline value > reduction threshold, Sbus = 0, indicating that the application business has been severely affected.
[0038] In the availability dimension, the health sub-score corresponding to the business impact dimension is determined by the decline in monitored metric data. Specifically, the monitored metric data corresponding to the availability dimension is normalized as follows: Within a set evaluation period, firstly, the probe success rate is determined, where the probe success rate is the success rate of external nodes accessing the target application. Then, the health sub-score corresponding to the availability dimension is calculated using the formula Sa = probe success rate × 100, where Sa is the health sub-score corresponding to the availability dimension. In this embodiment, if the number of consecutive failures of the target application's internal health check exceeds a set threshold (e.g., 3 times), the Sa of the target application is set to 0; this rule has the highest priority.
[0039] Step S103: Based on the application profile of the target application, determine the dynamic weight corresponding to each health dimension.
[0040] Here, the application profile includes both static and dynamic attributes of the target application. Static attributes include: application type (e.g., Web, Computing, Database) and business criticality (e.g., core, important, general). Dynamic attributes include: business scenario (e.g., peak season, critical security period) and operational events (e.g., recent serious failure). The following section explains in detail how to determine the dynamic weights corresponding to each health dimension based on the application profile of the target application.
[0041] Regarding step S103, as an example in specific implementation, it may include the following steps: Step S1031: Based on the static attributes, match the basic weights corresponding to each health dimension from the weight template library.
[0042] Here, based on the combination of static attributes, the most matching one is selected from a predefined weight template library. Specifically, as an example, if the application type is a web service and business criticality is the core, then the basic weight of template A is selected. The basic weight of template A is: basic weight (Wp) = 0.35 for the performance dimension, basic weight (Wbus) = 0.30 for the business dimension, basic weight (Wa) = 0.20 for the availability dimension, basic weight (Wr) = 0.10 for resources, and basic weight (Wsec) = 0.05 for the security dimension. If the application type is a database, then the basic weight of template B is selected. The basic weight of template B is: basic weight (Wp) = 0.4 for the performance dimension, basic weight (Wbus) = 0.05 for the business dimension, basic weight (Wa) = 0.15 for the availability dimension, basic weight (Wr) = 0.3 for resources, and basic weight (Wsec) = 0.1 for the security dimension.
[0043] Step S1032: Based on the priority of the dynamic attributes, determine the adjustment operator of the basic weight corresponding to each health dimension.
[0044] The following section explains in detail how to determine the adjustment operator for the basic weight corresponding to each health dimension based on the priority of the dynamic attributes.
[0045] Regarding step S1032, as an example in specific implementation, it may include the following steps: Step S10321: Based on the priority of the dynamic attributes, determine the first adjustment value of the basic weight corresponding to each health dimension.
[0046] Here, the base weights are adjusted based on the priority of dynamic attributes, with higher priority attributes overriding lower priority ones. Specifically, business scenarios are set to high priority. For example, if the current scenario is the peak of the first quarter (peak business traffic in the first quarter), Wbus is increased by 0.15, Wa by 0.10, Wsec by 0.05, and Wr by 0.10; if the current scenario is during a security attack / defense period, Wsec is increased by 0.20, Wp by 0.10, and Wbus by 0.05. Event feedback is set to low priority. Specifically, if the root cause of the last three failures is related to the database, Wp is increased by 0.05. If a serious security incident occurred last week, Wsec is increased by 0.10. As an example, suppose the insurance policy issuance service (core web service) triggers a health calculation during the first quarter peak: The application is identified as a core web service, and a basic weight template is selected: Wp=0.35, Wbus=0.30, Wa=0.20, Wr=0.10, Wsec=0.05. If the current scenario is the first quarter peak, then Wp remains unchanged, Wbus increases by 0.15, Wa increases to 0.1, Wsec decreases by 0.1, and Wr decreases to 0.1. During the first quarter peak, the application's health will be more sensitive to the impact of business traffic and availability to align with the prevailing protection goals.
[0047] Step S10322: Associate the historical health data and historical event record data of the target application within the first preset time period according to the preset association rules to generate the event diagnosis conclusion of the target application within the first preset time period.
[0048] Here, health data includes: the target application's health score, the health sub-score for each health dimension, and the dynamic weight for each health dimension. The health score represents a quantitative value of the target application's overall health status at a specific point in time. Event log data includes: service request tickets, alarm event tickets, and CI / CD push change release log tickets. The health data and event log data provide conclusions about changes in the health data according to preset association rules. For example, if the application's health score drops by 15 points at 14:30, the main reason is a decrease in the performance dimension score (-12 points), specifically due to a database issue. It should be noted that the basic weights of the target health dimensions that need adjustment can be determined based on the event diagnosis conclusions.
[0049] In this embodiment, historical health data, historical events, and event diagnostic conclusions are rendered to generate a first view in the visualization interface corresponding to the target application. Specifically, this application generates a spatiotemporal correlation view to display changes in health and related events in a linked manner. The first view combines a timeline and multiple data panels, displaying the historical change curve of the application's health score through a central timeline. Multiple data panels synchronized with the timeline respectively display monitoring indicator data curves, alarm event lists, and change release record lists within the same time period. In the visualization interface, not only is the data displayed, but also insights are provided in natural language description based on preset association rules. For example, the first view shows: the application's health score dropped by 15 points at 14:30, mainly due to a 12-point drop in the performance dimension score. A database change was detected at 14:25, and the number of slow queries in the database increased by 300%. It is recommended to roll back this change and check the related SQL statements.
[0050] Step S10323: Determine whether an event abnormality instruction used to indicate that the event diagnosis conclusion is abnormal has been detected.
[0051] Here, when operations and maintenance personnel click to confirm the event diagnosis conclusion as abnormal on the visualization interface, an event exception command is generated to indicate that the event diagnosis conclusion is abnormal. As an example, the visualization interface shows a sharp drop in the health of the insurance policy issuance service, presumably due to slow database queries (affecting the performance dimension Sp). The operations and maintenance personnel click to confirm that this prediction is correct. This transforms manually verified expert experience into executable configuration parameters for the system, enabling more accurate and intelligent behavior when encountering similar scenarios in the future.
[0052] Step S10324: If the event abnormality command is detected, then in response to the event abnormality command, key features are determined from the event diagnosis conclusion.
[0053] Here, the key features corresponding to the recorded event diagnosis conclusion are identified as the anomalies. For example, the key features include: the application type is a core web service, the root cause dimension is the performance dimension, and the specific cause is a database problem.
[0054] Step S10325: Based on the key features, determine the second adjustment value of the basic weight corresponding to each health dimension.
[0055] Here, as an example, for core web services, database issues affect the health sub-score corresponding to the performance dimension. Therefore, the weight corresponding to the performance dimension needs to be increased, with the increase being the second adjustment value. After multiple similar feedbacks, the system will automatically learn the evaluation priorities for database-dependent web services. When evaluating similar applications in the future, the system will automatically increase the weight of the performance dimension, making the health score more sensitive to such issues and thus making the alerts more accurate. Specifically, if the application has recently experienced related anomalies, the base weight will be temporarily adjusted; if the application has not recently experienced related anomalies, the base weight will be adjusted, and the sum of the current second adjustment value and the current first adjustment value will be determined as the first adjustment value of the base weight for subsequent health score calculations.
[0056] Step S10326: Determine the adjustment operator for the basic weights corresponding to each health dimension based on the first and second adjustment values of the basic weights corresponding to each health dimension.
[0057] Here, the adjustment operator is determined by the sum of the first adjustment value and the second adjustment value.
[0058] Step S1037: If no abnormal event instruction is detected, the first adjustment value of the basic weight corresponding to each health dimension is determined as the adjustment operator of the basic weight corresponding to each health dimension.
[0059] Step S1033: Correct the basic weights corresponding to each health dimension according to the corresponding adjustment operators to obtain the dynamic weights corresponding to each health dimension.
[0060] In one possible implementation, the health assessment method further includes: Step S11: In response to the score anomaly instruction used to indicate the health abnormality, determine the change range of historical monitoring indicator data corresponding to each health dimension within the second preset time period.
[0061] Here, when the complaint rate increases but the health level is within the normal range, a score abnormality instruction is generated to indicate that the health level is abnormal.
[0062] Step S12: Based on the magnitude of the change and the feedback data from the target application, determine the health dimension that causes the total health score to be abnormal.
[0063] Analyzing the changes in historical monitoring metrics helps pinpoint the abnormal baseline parameters. For example, frequent complaints from business personnel and an increased business complaint rate, yet the health score shows 85 (good), indicates that the current Sbus calculation formula is not sensitive enough to the actual business situation. The problem may stem from an overly lenient threshold for the decline. Assuming the original rule only assigns an Sbus score of 0 when transaction volume drops by more than 50%, this results in a score of 60 even with a 40% drop, failing to trigger a timely and severe alert. Analysis of historical data reveals that when transaction volume drops by more than 30%, the business complaint rate increases significantly. Therefore, the system automatically adjusts the decline threshold for the business impact dimension Sbus from 50% to 30%. In the future, when transaction volume drops by 35%, Sbus will directly be calculated as 0, lowering the overall health score to the danger zone and triggering a higher-level alert, more accurately reflecting the actual extent of business damage. By adjusting the baseline parameters based on feedback, the health sub-scores for each health dimension more accurately reflect the severity of the business, thereby improving accuracy and reliability. The optimized strategies and baseline parameters will automatically take effect for the next evaluation, ensuring that the system's response is more accurate and sensitive.
[0064] Step S13: Based on the preset adjustment step size, determine the new benchmark parameter corresponding to the health dimension that causes the health level to be abnormal, and replace the original benchmark parameter corresponding to the health dimension that causes the health level to be abnormal with the new benchmark parameter.
[0065] Return to reference Figure 1 Step S104: Determine the health score of the target application based on the health sub-score and dynamic weight corresponding to each health dimension.
[0066] Here, the result of weighted summation of the health sub-scores corresponding to each health dimension and their respective dynamic weights is determined as the health score of the target application.
[0067] In one possible implementation, the health assessment method further includes: rendering the health data of the target application into a second view in the corresponding visualization interface of the target application. Here, an application service map or list is used as the entry point, and the health score, sub-scores, and their weights are presented in a visual form (such as a radar chart, trend chart, or topology map). Each application is presented as a visualization element (such as a card or node), and its color (red / yellow / green) and internal fill level directly represent its health score. Drill-down analysis allows users to click on any application from the overview page and drill down to the application's details page to display its health score composition.
[0068] In this embodiment, corresponding operational measures can be automatically triggered based on the absolute value, trend, or prediction results of the health status. Specifically, when the health status falls below a set threshold, the system automatically creates an alarm event ticket. This ensures that relevant teams are quickly notified and can take necessary countermeasures when application performance degrades or problems occur. When a decline in health status is detected to be strongly correlated with recent CI / CD changes, the system not only automatically suggests performing a rollback operation but also creates a problem management record. This helps to quickly restore service stability and records events for future analysis and improvement. When health status trend predictions indicate a potential failure in the near future, the system will automatically allocate resources for capacity expansion or send early warning notifications to operations personnel. This early warning mechanism allows the operations team to prepare in advance, thereby reducing the impact of potential service interruptions or performance degradation.
[0069] This application provides a refined and intelligent operation and maintenance (O&M) system for complex distributed application systems throughout their entire lifecycle. Topological associations are established through application profiles and architectural dependencies, combined with agent analysis, to achieve a comprehensive, real-time, and proactive assessment of application health status. Visualization technology presents health data in a spatiotemporal relation to contextual information in O&M work order processing, providing O&M personnel with intuitive insights and decision support. Based on the policy context of application profiles, health assessment standards can be dynamically adjusted according to the application's identity, scenario, and stage. Furthermore, this application implements a unified indicator system that moves from technology silos to business integration, quantifying the impact of technical failures on business and effectively resolving communication barriers between technical O&M and business management. This application employs an intelligent evolution mechanism from an open-loop system to closed-loop learning, collecting manual confirmation information through visualization and feedback layers, and feeding this feedback back to the application profile, forming a reinforcement learning closed loop of assessment -> diagnosis -> feedback -> optimization. This mechanism enables the system to continuously self-optimize, achieving a level of intelligence far exceeding that of traditional static systems. Finally, the system can automatically trigger precise operation and maintenance operations, such as creating work orders, performing rollbacks, and sending alerts, which reduces the workload of operation and maintenance personnel, reduces the operational risks caused by human negligence, and promotes the evolution of operation and maintenance response from manual to intelligent.
[0070] This application provides a method for assessing the health of an application, which improves the accuracy, practicality, and intelligence of application health assessment, while reducing reliance on human labor and operational risks.
[0071] Based on the same application concept, this application also provides a health assessment device for an application corresponding to the health assessment method of the application provided in the above embodiments. Since the principle of the device in this application is similar to the health assessment method of the application in the above embodiments of this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0072] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a health assessment device provided in an embodiment of this application.
[0073] like Figure 3 As shown in the illustration, the health assessment device 310 provided in this application embodiment includes: The acquisition module 311 is used to acquire monitoring indicator data of the target application under multiple health dimensions; Processing module 312 is used to normalize the monitoring indicator data corresponding to each health dimension based on the benchmark parameters corresponding to that health dimension, and obtain the health sub-score corresponding to that health dimension. The dynamic weight determination module 313 is used to determine the dynamic weight corresponding to each health dimension based on the application profile of the target application. The health determination module 314 is used to determine the health of the target application based on the health sub-score and dynamic weight corresponding to each health dimension.
[0074] This application provides a health assessment device for applications, which improves the accuracy, practicality, and intelligence of application health assessment, and reduces reliance on human labor and operational risks.
[0075] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0076] like Figure 4 As shown, the electronic device 400 includes a processor 410, a memory 420, and a bus 430.
[0077] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 is running, the processor 410 communicates with the memory 420 via the bus 430. When the machine-readable instructions are executed by the processor 410, they can perform the operations described above. Figure 1 and Figure 2 The steps of the health assessment method applied in the method embodiment shown are described in detail in the method embodiment, and will not be repeated here.
[0078] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described actions. Figure 1 and Figure 2 The steps of the health assessment method applied in the method embodiment shown are described in detail in the method embodiment, and will not be repeated here.
[0079] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0080] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0081] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0082] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0083] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for health assessment in application, characterized in that, The health assessment methods include: Obtain monitoring indicator data for the target application across multiple health dimensions; For each health dimension, based on the baseline parameters corresponding to that health dimension, the monitoring indicator data corresponding to that health dimension is normalized to obtain the health sub-score corresponding to that health dimension. Based on the application profile of the target application, determine the dynamic weight corresponding to each health dimension; The health score of the target application is determined based on the health sub-score and dynamic weight corresponding to each health dimension.
2. The health assessment method according to claim 1, characterized in that, The application profile includes: static attributes and dynamic attributes of the target application; the determination of the dynamic weight corresponding to each health dimension based on the application profile of the target application includes: Based on the static attributes, the basic weights corresponding to each health dimension are matched from the weight template library; Based on the priority of the dynamic attributes, the adjustment operator for the basic weight corresponding to each health dimension is determined; The basic weights corresponding to each health dimension are adjusted according to the corresponding adjustment operators to obtain the dynamic weights corresponding to each health dimension.
3. The health assessment method according to claim 2, characterized in that, The adjustment operator for determining the basic weights corresponding to each health dimension based on the priority of the dynamic attributes includes: Based on the priority of the dynamic attributes, determine the first adjustment value of the basic weight corresponding to each health dimension; The historical health data and historical event record data of the target application within the first preset time period are associated according to the preset association rules to generate event diagnosis conclusions of the target application within the first preset time period. Determine whether an event anomaly instruction is detected that indicates the event diagnosis conclusion is abnormal; If the event abnormality command is detected, then in response to the event abnormality command, key features are determined from the event diagnosis conclusion; Based on the aforementioned key features, a second adjustment value for the basic weight corresponding to each health dimension is determined; Based on the first and second adjustment values of the basic weights corresponding to each health dimension, determine the adjustment operator for the basic weights corresponding to each health dimension; If no abnormal event instruction is detected, the first adjustment value of the basic weight corresponding to each health dimension is determined as the adjustment operator of the basic weight corresponding to each health dimension.
4. The health assessment method according to claim 3, characterized in that, The health assessment method also includes: In response to a score anomaly instruction used to indicate an abnormality in health status, the magnitude of change in historical monitoring indicator data corresponding to each health dimension within a second preset time period is determined. Based on the magnitude of the change, determine the health dimensions that cause the total health score to be abnormal; Based on a preset adjustment step size, a new benchmark parameter corresponding to the health dimension that causes the health status to be abnormal is determined, and the original benchmark parameter corresponding to the health dimension that causes the health status to be abnormal is replaced with the new benchmark parameter.
5. The health assessment method according to claim 3, characterized in that, The health assessment method also includes: The historical health data, the historical events, and the event diagnostic conclusions are rendered to generate the first view in the visualization interface corresponding to the target application.
6. The health assessment method according to claim 1, characterized in that, The health assessment method also includes: The health data of the target application is rendered to generate a second view in the corresponding visualization interface of the target application.
7. The health assessment method according to claim 6, characterized in that, The health data includes: the health score of the target application, the health sub-score for each health dimension, and the dynamic weight for each health dimension.
8. A health assessment device for application, characterized in that, The health assessment device includes: The acquisition module is used to acquire monitoring indicator data for the target application across multiple health dimensions. The processing module is used to normalize the monitoring indicator data corresponding to each health dimension based on the benchmark parameters corresponding to that health dimension, and obtain the health sub-score corresponding to that health dimension. The dynamic weight determination module is used to determine the dynamic weight corresponding to each health dimension based on the application profile of the target application. The health determination module is used to determine the health of a target application based on the health sub-score and dynamic weight corresponding to each health dimension.
9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the machine-readable instructions are executed by the processor to perform the steps of the health assessment method of the application as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the health assessment method for the application as described in any one of claims 1 to 7.