Health management method and system based on multi-element heterogeneous monitoring data
By using a health management method that dynamically adjusts weights and self-optimizes, the problem of false alarms and missed alarms in traditional monitoring methods under dynamic and changing environments is solved, achieving efficient and accurate system health assessment and early warning, and reducing operation and maintenance costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-28
- Publication Date
- 2026-05-29
AI Technical Summary
Traditional system health monitoring methods cannot adapt to dynamic changes, resulting in frequent false alarms and missed alarms. They lack self-calibration capabilities, making it difficult to achieve proactive prevention and unable to analyze the correlation and impact between multi-source and multi-level indicators from a holistic system perspective.
A health management method based on multivariate heterogeneous monitoring data is adopted. The weights are dynamically adjusted through an online health assessment model, and self-optimization is performed by combining historical operation and maintenance feedback data. A health prediction model is constructed and the parameters are updated by minimizing the loss function to achieve self-adaptation and self-learning.
It significantly reduces the false alarm rate, improves the accuracy of assessment and the precision of early warning, reduces operation and maintenance costs, can adapt to system changes and improve itself by following new failure modes, and maintains the durability and reliability of assessment effectiveness.
Smart Images

Figure CN122111739A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer system monitoring and operation and maintenance automation technology, and in particular to a health management method and system based on diverse heterogeneous monitoring data. Background Technology
[0002] With the deepening of digital transformation, enterprise IT infrastructure and business systems are becoming increasingly complex, exhibiting characteristics of heterogeneity (hybrid cloud, various industrial equipment, and different application architectures) and dynamism (elastic scaling, continuous delivery). Traditional system health monitoring methods mainly have the following limitations:
[0003] Static and rigid alerts often use fixed thresholds or static weights, which cannot adapt to dynamic changes in the system's operating environment (such as peak business periods or resource pressure), leading to frequent false alarms and missed alarms.
[0004] Local and isolated: Monitoring is usually done on a single device or service, lacking analysis of the correlation and impact between multiple sources and multiple levels of indicators from a system-wide perspective, making it difficult to accurately diagnose the root cause.
[0005] Delay and passivity: Most methods only issue alarms after a failure occurs, lacking the ability to quantitatively assess and provide early warnings of the system's "sub-healthy" state, leaving operations and maintenance in a passive response mode.
[0006] Lack of self-evolution: The evaluation model parameters rely on manual experience to set, and cannot use historical operation and maintenance feedback data for self-calibration and optimization, making it difficult to cope with the continuous evolution of the system and new failure modes.
[0007] Therefore, there is an urgent need for a health assessment method that can adapt to dynamic changes, integrate multi-dimensional information, quantify health trends, and have self-learning capabilities, in order to achieve the transformation of intelligent operation and maintenance from "passive firefighting" to "proactive prevention". Summary of the Invention
[0008] The purpose of this invention is to disclose a health management method and system based on diverse heterogeneous monitoring data, so as to improve the intelligence and reliability of operation and maintenance.
[0009] To achieve the above objectives, this invention discloses a health management method based on diverse heterogeneous monitoring data, comprising: Step S1: The health assessment model performs an online health assessment process based on the current value of the parameter vector to be optimized, including the following sub-steps: the parameter vector to be optimized is a set consisting of at least two parameters to be optimized; Sub-step S11: Collect multi-source heterogeneous monitoring data and obtain the indicator status matrix through preprocessing; Sub-step S12: Based on the current state of each indicator in the indicator state matrix, and combined with the basic weight, abnormal timeliness, running context and indicator coupling relationship, calculate the dynamic weight vector of each indicator in the current evaluation period to reflect the importance of different indicators in the current system state. Sub-step S13: Based on the indicator state matrix and dynamic weight vector, comprehensively evaluate the current operating status of the system and output the health prediction result; Step S2: Based on the health prediction results generated during the online health assessment process, and combined with subsequent actual operation and maintenance events, fault handling results, and business impact records, construct a "prediction result - actual status" paired sample set; Step S3: Determine whether the number of paired samples in the paired sample set meets the update condition. If not, return to step S1; if yes, proceed to step S4. Step S4: Form a residual vector based on each paired sample, and then, with the goal of minimizing the loss function, update the parameter vector to be optimized based on the parameter sensitivity matrix, the weights of each sample pair, and the residual vector, and then loop back to step S1; wherein, the parameter sensitivity matrix is used to characterize the response relationship of the predicted output to parameter changes.
[0010] Preferably, the indicator state matrix Specifically: ;in: For the first Each indicator at time The standardized value; This represents the rate of change of the indicator within a preset time window; This refers to the duration of the indicator since it entered the current abnormal level; Mark the current abnormal level of this indicator; Rate the data quality; The scene coefficients corresponding to the current context-aware results; This is the overall correlation strength value between this indicator and other abnormal indicators.
[0011] Preferably, all indicators are classified and a set of parameters to be optimized is shared among indicators of the same type.
[0012] Preferably, the parameter vector to be optimized Specifically: ;in: Indicates the first Time decay coefficient of similar indicators; and These represent the emergency weight multipliers for this type of indicator in the warning and severe states, respectively; and These represent the influence intensity parameter and trigger threshold parameter of the Sigmoid function in the calculation of the correlation influence factor, respectively; and They represent the first Class indicators in the first The membership function center value and width parameter under each health level; The scoring threshold parameters represent the boundaries between different levels of health, sub-health, warning, and malfunction. The Sigmoid function is defined as follows: Indicator coupling correlation factor Defined as: ;in, Indicates the first The indicator category to which each indicator belongs. For the first Each indicator at time The weighted average association strength, wherein the weighted average association strength is defined as: ;in: In order to be with the first A set of related indicators that exhibit abnormal correlation; For the first The relevant indicator for the first The overall correlation of the indicators; As the anomaly factor, when the first When each indicator is in a normal, warning, or severe state, take the following values respectively: , and ; The time decay term is defined as: ;in, For the first The duration of each indicator since entering the current anomaly level; and the comprehensive correlation degree of each indicator is calculated based on the comprehensive anomaly co-occurrence correlation degree, statistical trend correlation degree, and causal inference correlation degree. Only when the comprehensive correlation degree exceeds the set threshold is the abnormal correlation between the indicator pairs confirmed.
[0013] Preferably, the first Each indicator at time Dynamic weights The calculation formula is: ;in: Basic weights; The time factor is used to characterize the combined impact of anomaly duration and anomaly level on the weights. As a context factor, it is used to characterize the impact of business time period, load level, environmental state and special operating scenario on the weight; For the coupling correlation factor of the indicator; For normalization processing; the time factor is defined as: ;in: For the first The duration of each indicator since it entered the current abnormal level; For the first Each indicator at time The abnormality level multiplier, when the first When each indicator is in a normal, warning, or severe state, Take respectively , and .
[0014] Preferably, the process of forming a residual vector based on each paired sample, and then updating the parameter vector to be optimized based on the parameter sensitivity matrix, the weights of each sample pair, and the residual vector with the goal of minimizing the loss function, includes the following sub-steps: Sub-step S41: Define the loss function Specifically: ;in: The vector of parameters to be optimized is At that time, the first Predicted health score output for each sample; Indicates the first The true target score corresponding to each sample; Indicates sample weights; This represents the previous parameter vector to be optimized used by the current online model; Indicates the smoothing coefficient of the parameter; The L2 norm constraint term represents the parameter update magnitude; Sub-step S42: Construct the parameter sensitivity matrix Specifically: ;in: This represents the number of paired samples; for The dimension; the first dimension in the matrix line, number Column elements Indicates the first The health score of the sample is related to the first The sensitivity of each parameter is calculated using the finite difference method: ;in: Indicates the first A tiny perturbation applied to each parameter; Indicates the first The unit basis vectors corresponding to each parameter; Sub-step S43, calculate the first The parameter vector to be optimized. The specific calculation formula is as follows: ;in, Indicates the learning rate; Indicates by sample weights The resulting diagonal weighted matrix; This represents the current predicted score vector. Represents the true target score vector. Represents the residual vector; This indicates that the updated parameters will be projected onto the valid parameter constraint domain. Inside; This represents the approximate gradient of the loss function with respect to the parameter vector to be optimized.
[0015] Preferably, the total number of specific indicators participating in the health assessment at a certain assessment time is . The total number of indicator categories is The total number of health levels is ;in, Indicates the first A specific indicator, Indicates the first Class indicators, Indicates the first Each health level, any specific indicator Belonging to a certain indicator category At any moment ,all The states of specific indicators are combined row by row to form an indicator state matrix. ,and The dynamic weight vector consists of the dynamic weights corresponding to each specific indicator. ; Sub-step S13 specifically involves: first targeting all Each specific indicator is calculated separately for its effect on... The membership degree of each health level is used to obtain the single-index membership degree matrix. Then, the system-level fuzzy evaluation vector is obtained by weighted aggregation of the dynamic weight vector and the single-index membership matrix: Then, based on the system-level fuzzy evaluation vector, defuzzification is performed to obtain the system-level health score. .
[0016] Preferably, let: ; Then we have: ; ; in, This is the Hadamard product operator; This represents a matrix composed of the dynamic weight vectors of each indicator. The value of and the parameter vector to be optimized Related; To form a basic weight matrix that integrates all indicators, The values of the elements in the time factor matrix, which integrates all indicators, are related to the time decay coefficient of the corresponding category and the emergency weight multiplier in the warning and severe states. This indicates that the values of the elements in the matrix of coupled correlation factors of all indicators are related to the influence intensity parameter and trigger threshold parameter of the corresponding category. Indicates from The matrix formed by the scene coefficients corresponding to the context-aware results of all extracted indicators; ;in, express The value of and the parameter vector to be optimized Related, Indicates a single indicator Membership calculation and this index The value at the current moment and corresponding category indicators in the first The membership function center value and width parameter are related under each health level; ;in, express The value of and the parameter vector to be optimized Related; ;in, This indicates the representative score corresponding to each health level; express Corresponding health level The possible values of ; express The value of and the parameter vector to be optimized Related; ;in, Indicates according to The divided scoring intervals will A system-level health score obtained by integrating all indicators at all times Mapped to the corresponding health level The function.
[0017] Preferably, the parameter constraints when performing update processing on the parameter vector to be optimized include: ; ; ;and To ensure that the system-level health score is greater than or equal to The output level is healthy, greater than or equal to and less than The output level is sub-healthy, greater than or equal to... and less than The output level is warning, less than... The output level is fault.
[0018] To achieve the above objectives, the present invention also discloses a health management system based on diverse heterogeneous monitoring data, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.
[0019] The present invention has the following beneficial effects: 1. By dynamically adjusting weights, the assessment focus is on the most critical anomalies in real time, significantly reducing the false alarm rate caused by rigid thresholds or environmental changes; the assessment accuracy is high and the early warning is more precise.
[0020] 2. Strong adaptability and low operation and maintenance costs: The algorithm can automatically adapt to changes in system load, business cycle, and other contexts, and perform self-optimization using paired sample sets. This significantly reduces reliance on manual adjustments of rules and thresholds by operation and maintenance experts, thus lowering long-term operation and maintenance costs. Based on iterative updates of the parameter vector to be optimized, it can self-improve as the system evolves and new failure modes emerge, maintaining the durability and reliability of the evaluation effectiveness.
[0021] 3. Strong engineering feasibility and easy integration: The algorithm features a clear modular design and standardized data interfaces, enabling easy integration with existing monitoring platforms and operation and maintenance systems. Parameter updates and knowledge evolution processes are controllable and auditable, meeting the stability and manageability requirements of enterprise-level applications.
[0022] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description
[0023] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a schematic diagram of the health management method based on multi-source heterogeneous monitoring data disclosed in an embodiment of the present invention. Detailed Implementation
[0024] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings, but the present invention can be implemented in many different ways as defined and covered by the claims.
[0025] Example 1 This embodiment discloses a health management method based on diverse heterogeneous monitoring data, such as... Figure 1 As shown, it includes: Step S1: The health assessment model performs an online health assessment process based on the current values of the parameter vector to be optimized.
[0026] In this embodiment, the parameter vector to be optimized is a set consisting of at least two parameters to be optimized, and this step includes the following sub-steps S11 to S13.
[0027] Sub-step S11: Collect multi-source heterogeneous monitoring data and obtain the indicator status matrix through preprocessing.
[0028] During the execution of this sub-step, the following processing is required: A. System Configuration and Data Source Connection: Define the various data sources to be monitored, including device sensors, system performance counters, application log files, database audit logs, network traffic probes, etc. Configure connection parameters for each data source (such as IP address, port, authentication information, and sampling frequency).
[0029] B. Data Parsing and Format Standardization: The corresponding protocol parser or adapter is invoked to extract key fields from the raw data stream. The following core fields are extracted and standardized from various heterogeneous data sources to form the original data records: Data source identifier: Uniquely identifies the data source, in the format: {system type}_{location / hostname}_{data source type}; Collection timestamp: The precise time when the data was collected; Original metric name: The original metric name provided by the data source; Original value: The original numerical value or string corresponding to the metric; Unit: The unit of the original value, used for subsequent normalization conversion.
[0030] C. Data cleaning and preliminary outlier filtering: Apply statistical (such as the 3σ principle) or rule-based filters to remove obviously erroneous collected values (such as values that exceed the physical range).
[0031] D. Raw Indicator Calculation and Extraction: Preliminary calculations are performed on the raw data to generate basic indicators that can be used for evaluation. These include: device / service online rate, communication link quality, power / network status, CPU utilization, memory utilization, transaction processing rate, log error frequency, sensor accuracy degradation, and success rate of critical business transactions.
[0032] E. Output standardized indicator set: Convert the cleaned indicator data into internal standard data objects to obtain the indicator status matrix for use by subsequent modules.
[0033] For example, the indicator state matrix Specifically: ;in: For the first Each indicator at time The standardized value; This represents the rate of change of the indicator within a preset time window; This refers to the duration of the indicator since it entered the current abnormal level. Mark the current abnormal level of this indicator; Rate the data quality; The scene coefficients corresponding to the current context-aware results; This is the combined correlation strength value between this indicator and other abnormal indicators.
[0034] Optionally, anomalies based on a single metric can be categorized into two states: warning and critical. A warning indicates a critical point requiring attention but not immediate action, while a critical state indicates a significant anomaly requiring immediate investigation. For example, CPU utilization threshold settings: Warning threshold: 75% (slight overload, observable); Critical threshold: 85% (significant overload, requires investigation).
[0035] Sub-step S12: Based on the current state of each indicator in the indicator state matrix, and combined with the basic weight, abnormal timeliness, running context and indicator coupling relationship, calculate the dynamic weight vector of each indicator in the current evaluation period to reflect the importance of different indicators in the current system state.
[0036] In this sub-step, for example: the first Each indicator at time Dynamic weights The calculation formula is: ;in: Basic weights; The time factor is used to characterize the combined impact of anomaly duration and anomaly level on the weights. As a context factor, it is used to characterize the impact of business time period, load level, environmental state and special operating scenario on the weight; For the coupling correlation factor of the indicator; This is for normalization purposes.
[0037] The time factor is defined as follows: ; For the first The duration of each indicator since it entered the current abnormal level; For the first Each indicator at time The anomaly level multiplier, when the first When each indicator is in a normal, warning, or severe state, Take respectively , and Therefore, under normal circumstances The dynamic weights of the indicators remain consistent with the basic weights or are adjusted under the influence of other factors. In abnormal situations, the time factor amplifies or reduces the basic weights in a multiplicative manner to avoid the basic weights being calculated repeatedly in the dynamic weight formula.
[0038] In this sub-step, the indicators are coupled with correlation factors. It can be defined as: ;in, Indicates the first The indicator category to which each indicator belongs. For the first Each indicator at time The weighted average correlation strength, and They represent the first The calculation of the correlation influence factors corresponding to each indicator includes the influence intensity parameter and trigger threshold parameter of the Sigmoid function. The classification of all indicators is primarily to provide supporting mechanisms for subsequently using similar indicators to share a set of parameter vectors to be optimized, thereby simplifying the system's complexity.
[0039] The Sigmoid function is defined as follows: ; Let be the independent variable of the function. The weighted average association strength is defined as: ; In order to be with the first A set of related indicators that exhibit abnormal correlation; For the first The relevant indicator for the first The overall correlation of the indicators; As the anomaly factor, when the first When each indicator is in a normal, warning, or severe state, take the following values respectively: , and , and They represent the first The emergency weight multiplier of the corresponding class indicator in the warning and severe states.
[0040] The time decay term is defined as: ;in, For the first The duration of each indicator since it entered the current abnormal level.
[0041] In this embodiment, an abnormal correlation between an indicator pair is confirmed only when the comprehensive correlation exceeds a set threshold. Preferably, each indicator is calculated using a comprehensive correlation of co-occurrence of anomalies, statistical trend correlation, and causal inference correlation. Specifically, the co-occurrence of anomalies correlation is calculated based on historical anomaly events, determining the conditional probability that when indicator i is anomaly, indicator j is also anomaly within a specific time window; the statistical trend correlation is used to analyze the similarity of numerical change trends of indicators during anomaly development; and the causal inference correlation can apply statistical methods such as the Granger causality test to analyze the causal direction and strength between indicators. All three are existing technologies and will not be elaborated upon further.
[0042] It is worth noting that this was used in both the indicator correlation influencing factor and the time factor. However, the two are different: one is used to characterize indicators. The impact of the duration of anomalies on ontology weights, and a metric for characterizing related anomalies. The duration of the effect on its contribution to coupling propagation; the two act on objects at two different levels, and are not a repeated calculation of the same amount of influence.
[0043] Sub-step S13: Based on the indicator state matrix and dynamic weight vector, comprehensively evaluate the current operating status of the system and output the health prediction result.
[0044] This step can use fuzzy mathematics theory to handle the uncertainty in the assessment, and can specifically include: constructing a membership function for each assessment indicator and defining four fuzzy levels: "healthy", "sub-healthy", "warning", and "faulty"; calculating the membership degree of all indicators to each level based on the real-time values of the indicators and the dynamic weights obtained in the previous step; and integrating global assessment information through weighted aggregation and centroid method to finally output the system's accurate health score (0-100 points) and health level.
[0045] Step S2: Based on the health prediction results generated during the online health assessment process, and combined with subsequent actual operation and maintenance events, fault handling results, and business impact records, construct a "prediction result - actual status" paired sample set.
[0046] In the specific execution process, this step can generate an evaluation snapshot based on the actual abnormal events recorded by the work order system, event management system, fault handling system, and business operation system, denoted as: ;in: Identify actual events; and These are the actual start time and recovery time of the event, respectively. The severity level of the actual incident; For affected system components, business links, or sets of metrics; This refers to the root cause category or the conclusion of the action taken in response to the incident.
[0047] To generate paired samples that can be used for supervised optimization, the system aligns the predicted evaluation data with the actual operational ground truth data over time. Specifically, for any given evaluation moment... A pre-set warning window is then set up. Internally, a search is performed to determine if any actual abnormal events exist. The preferred timeframe is 60 minutes. When an abnormal event is detected within this window, the evaluation snapshot is considered to have a supervised association with that event; when no abnormal event is detected, the evaluation snapshot is considered to correspond to a sample with no abnormalities or stable operation. In this way, a spatiotemporal pairing relationship between "predicted results and actual state" can be established.
[0048] To transform actual operational feedback into calculable calibration targets, the system constructs a calibration sample set based on the aforementioned spatiotemporal alignment results: ;in: For the first The index state matrix corresponding to each sample; The model's predicted health score for this sample; This is the actual target score obtained by mapping subsequent actual operation and maintenance events; The sample weights are used to distinguish the importance of events of different severity levels to the calibration. The true target score is also included. Generate according to preset mapping rules, for example, the following mapping method can be used: If no abnormal event occurs within the warning window, the sample is marked as a healthy sample and taken. If a minor performance degradation event occurs within the warning window but does not cause critical business interruption, the sample will be marked as a sub-healthy sample and taken. If an abnormal event requiring manual intervention occurs within the warning window, or if key system indicators enter a significantly abnormal state, the sample will be marked as a warning sample and taken. If a critical service becomes unavailable, a serious failure occurs, or an event triggers emergency maintenance within the warning window, the sample will be marked as a fault sample and retrieved. .
[0049] To reflect the differences in importance among different samples, sample weights The weighting is related to the severity, scope, and duration of the actual event. Preferably, the weighting can be set as follows: severe fault samples have a greater weight than warning samples, warning samples have a greater weight than sub-healthy samples, and sub-healthy samples have a greater weight than healthy samples. For example, a healthy sample weight of 1, a sub-healthy sample weight of 2, a warning sample weight of 4, and a fault sample weight of 6 can be set. This allows the optimization process to prioritize reducing the risk of severe missed detections.
[0050] Step S3: Determine whether the number of paired samples in the paired sample set meets the update condition. If not, return to step S1; if yes, execute step S4.
[0051] In this step, the paired sample sets corresponding to different iterations of the parameter vector to be optimized need to be stored separately to avoid the historical paired samples in the old paired sample sets affecting the statistics of the number of the latest paired samples.
[0052] Step S4: Form a residual vector based on each paired sample. Then, with the goal of minimizing the loss function, update the parameter vector to be optimized based on the parameter sensitivity matrix, the weights of each sample pair, and the residual vector, and then loop back to step S1.
[0053] In this step, the parameter sensitivity matrix is used to characterize the response relationship of the predicted output to parameter changes; preferably, all indicators are classified and the same type of indicators share a set of the parameter vectors to be optimized.
[0054] For example, the parameter vector to be optimized Specifically: ;in: Indicates the first Time decay coefficient of similar indicators; and These represent the emergency weight multipliers for this type of indicator in the warning and severe states, respectively; and These represent the influence intensity parameter and trigger threshold parameter of the Sigmoid function in the calculation of the correlation influence factor, respectively; and They represent the first Class indicators in the first The membership function center value and width parameter under each health level; The scoring threshold parameters represent the boundaries between the levels of health, sub-health, warning, and malfunction.
[0055] Furthermore, this step can be implemented by the following sub-steps.
[0056] Sub-step S41: Define the loss function Specifically: ;in: The vector of parameters to be optimized is At that time, the first Predicted health score output for each sample; Indicates the first The true target score corresponding to each sample; Indicates sample weights; This represents the previous parameter vector to be optimized used by the current online model; Indicates the smoothing coefficient of the parameter; The L2 norm constraint term represents the magnitude of parameter update.
[0057] The loss function described above consists of two parts: the first part is the prediction error term, which measures the degree of fit of the current model to the actual operation and maintenance status; the second part is the smoothing update term, which limits the parameter update magnitude and avoids drastic changes in parameters due to short-term sample fluctuations, thus affecting the stability of online evaluation.
[0058] In a preferred embodiment, for samples that experience severe underreporting, the sample weights can be adjusted. An additional false negative enhancement coefficient is applied to the existing model; for samples that generate false alarms for an extended period without causing actual faults, a false alarm penalty coefficient can be added. These methods further reflect the operational goal of "prioritizing the suppression of false alarms and moderately controlling false alarms."
[0059] Sub-step S42: Construct the parameter sensitivity matrix Specifically: ;in: This represents the number of paired samples; for The dimension of the matrix; the first dimension of the matrix. line, number Column elements Indicates the first The health score of the sample is related to the first The sensitivity of each parameter is calculated using the finite difference method: ;in: Indicates the first A tiny perturbation applied to each parameter; Indicates the first The unit basis vectors corresponding to each parameter.
[0060] Sub-step S43, calculate the first... The parameter vector to be optimized. The specific calculation formula is as follows: ;in, Indicates the learning rate; Indicates by sample weights The resulting diagonal weighted matrix; This represents the current predicted score vector. Represents the true target score vector. Represents the residual vector; This indicates that the updated parameters will be projected onto the valid parameter constraint domain. Inside; This represents the approximate gradient of the loss function with respect to the parameter vector to be optimized.
[0061] In the above process, the residual vector represents the current deficiency to be optimized, and the parameter sensitivity matrix represents how the parameter will change this deficiency. The two together determine how to modify the parameter.
[0062] To facilitate a deeper understanding of the specific implementation of the above steps by those skilled in the art, detailed examples are provided below.
[0063] Let the total number of specific indicators participating in the health assessment at a certain assessment time be . The total number of indicator categories is The total number of health levels is ;in, Indicates the first A specific indicator, Indicates the first Class indicators, Indicates the first Each health level, any specific indicator Belonging to a certain indicator category At any moment ,all The states of specific indicators are combined row by row to form an indicator state matrix. ,and The dynamic weight vector consists of the dynamic weights corresponding to each specific indicator. .
[0064] Sub-step S13 specifically involves: first targeting all Each specific indicator is calculated separately for its effect on... The membership degree of each health level is used to obtain the single-index membership degree matrix. Then, the system-level fuzzy evaluation vector is obtained by weighted aggregation of the dynamic weight vector and the single-index membership matrix: Then, based on the system-level fuzzy evaluation vector, defuzzification is performed to obtain the system-level health score. .
[0065] Furthermore, let: ; Then we have: ; .
[0066] in, This is the Hadamard product operator; This represents a matrix composed of the dynamic weight vectors of each indicator. The value of and the parameter vector to be optimized Related; To form a basic weight matrix that integrates all indicators, The values of the elements in the time factor matrix, which integrates all indicators, are related to the time decay coefficient of the corresponding category and the emergency weight multiplier in the warning and severe states. This indicates that the values of the elements in the matrix of coupled correlation factors of all indicators are related to the influence intensity parameter and trigger threshold parameter of the corresponding category. Indicates from The matrix is formed by the scene coefficients corresponding to the context-aware results of all extracted indicators.
[0067] ;in, express The value of and the parameter vector to be optimized Related, Indicates a single indicator Membership calculation and this index The value at the current moment and corresponding category indicators in the first The membership function center value and width parameter are related under each health level.
[0068] ;in, express The value of and the parameter vector to be optimized Related.
[0069] ;in, This indicates the representative score corresponding to each health level; express Corresponding health level The possible values of ; express The value of and the parameter vector to be optimized Related.
[0070] ;in, Indicates according to The divided scoring intervals will A system-level health score obtained by integrating all indicators at all times Mapped to the corresponding health level The function.
[0071] Furthermore, in a specific computational example, the parameter constraints when performing update processing on the above-mentioned parameter vector to be optimized include: ; ; ;and To ensure that the system-level health score is greater than or equal to The output level is healthy, greater than or equal to and less than The output level is sub-healthy, greater than or equal to... and less than The output level is warning, less than... The output level is fault.
[0072] When performing updates based on paired sample sets, the calibration sample set is first divided into a training set and a validation set. After updating the parameters using the training set, the combined performance of the new and old parameter models is compared on the validation set. Preferably, step S4 is only allowed to proceed to the release process if the new parameter model meets the following conditions: The weighted loss on the validation set decreased by more than the preset minimum improvement threshold; the false negative rate of severely faulty samples decreased by more than the preset proportion; the false positive rate did not increase by more than the preset upper limit; and the output score distribution did not show any obvious abnormal drift.
[0073] For example, when the false negative rate decreases by at least 5% and the false positive rate increases by no more than 2%, the system will designate the new parameter version as a candidate version and gradually replace the old model parameters using blue-green deployment or canary rollout. Specifically, the system maintains both the old and new parameter versions simultaneously. For some evaluation requests, the new parameter model is first imported for parallel evaluation, and the differences in output and online event feedback are compared. After the system stabilizes, all traffic is then switched to the new parameter model.
[0074] Furthermore, the system can version-record each parameter update process, with the recorded information including at least: calibration cycle range, number of samples participating in calibration, differences before and after parameter update, verification results, release time, and rollback version number. Through these methods, traceability, auditability, and rollback capability of model parameter updates can be achieved.
[0075] Example 2 This embodiment discloses a health management system based on diverse heterogeneous monitoring data, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method of Embodiment 1 above. The underlying logic is the same as that of Embodiment 1 and will not be described again.
[0076] In summary, the health management methods and systems based on diverse heterogeneous monitoring data disclosed in the above embodiments of the present invention have the following beneficial effects: 1. By dynamically adjusting weights, the assessment focus is on the most critical anomalies in real time, significantly reducing the false alarm rate caused by rigid thresholds or environmental changes; the assessment accuracy is high and the early warning is more precise.
[0077] 2. Strong adaptability and low operation and maintenance costs: The algorithm can automatically adapt to changes in system load, business cycle, and other contexts, and perform self-optimization using paired sample sets. This significantly reduces reliance on manual adjustments of rules and thresholds by operation and maintenance experts, thus lowering long-term operation and maintenance costs. Based on iterative updates of the parameter vector to be optimized, it can self-improve as the system evolves and new failure modes emerge, maintaining the durability and reliability of the evaluation effectiveness.
[0078] 3. Strong engineering feasibility and easy integration: The algorithm features a clear modular design and standardized data interfaces, enabling easy integration with existing monitoring platforms and operation and maintenance systems. Parameter updates and knowledge evolution processes are controllable and auditable, meeting the stability and manageability requirements of enterprise-level applications.
[0079] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A health management method based on multi-source heterogeneous monitoring data, characterized in that, include: Step S1: The health assessment model performs an online health assessment process based on the current value of the parameter vector to be optimized, including the following sub-steps: the parameter vector to be optimized is a set consisting of at least two parameters to be optimized; Sub-step S11: Collect multi-source heterogeneous monitoring data and obtain the indicator status matrix through preprocessing; Sub-step S12: Based on the current state of each indicator in the indicator state matrix, and combined with the basic weight, abnormal timeliness, running context and indicator coupling relationship, calculate the dynamic weight vector of each indicator in the current evaluation period to reflect the importance of different indicators in the current system state. Sub-step S13: Based on the indicator state matrix and dynamic weight vector, comprehensively evaluate the current operating status of the system and output the health prediction result; Step S2: Based on the health prediction results generated during the online health assessment process, and combined with subsequent actual operation and maintenance events, fault handling results, and business impact records, construct a "prediction result - actual status" paired sample set; Step S3: Determine whether the number of paired samples in the paired sample set meets the update condition. If not, return to step S1. If so, proceed to step S4; Step S4: Form a residual vector based on each paired sample, and then, with the goal of minimizing the loss function, update the parameter vector to be optimized based on the parameter sensitivity matrix, the weights of each sample pair, and the residual vector, and then loop back to step S1; wherein, the parameter sensitivity matrix is used to characterize the response relationship of the predicted output to parameter changes.
2. The health management method based on multi-source heterogeneous monitoring data according to claim 1, characterized in that, The indicator state matrix Specifically: ;in: For the first Each indicator at time The standardized value; This represents the rate of change of the indicator within a preset time window; This refers to the duration of the indicator since it entered the current abnormal level. Mark the current abnormal level of this indicator; Rate the data quality; The scene coefficients corresponding to the current context-aware results; This is the combined correlation strength value between this indicator and other abnormal indicators.
3. The health management method based on multi-source heterogeneous monitoring data according to claim 2, characterized in that, All indicators are categorized and those of the same type share a set of parameter vectors to be optimized.
4. The health management method based on multi-source heterogeneous monitoring data according to claim 3, characterized in that, The parameter vector to be optimized Specifically: ;in: Indicates the first Time decay coefficient of similar indicators; and These represent the emergency weight multipliers for this type of indicator in the warning and severe states, respectively; and These represent the influence intensity parameter and trigger threshold parameter of the Sigmoid function in the calculation of the correlation influence factor, respectively; and They represent the first Class indicators in the first The membership function center value and width parameter under each health level; The scoring threshold parameters represent the boundaries between different levels of health, sub-health, warning, and malfunction. The Sigmoid function is defined as follows: Indicator Coupling Correlation Factor Defined as: ;in, Indicates the first The indicator category to which each indicator belongs. For the first Each indicator at time The weighted average association strength, wherein the weighted average association strength is defined as: ;in: In order to be with the first A set of related indicators that exhibit abnormal correlation; For the first The relevant indicator for the first The overall correlation of the indicators; As the anomaly factor, when the first When each indicator is in a normal, warning, or severe state, take the following values respectively: , and ; The time decay term is defined as: ;in, For the first The duration of each indicator since entering the current anomaly level; and the comprehensive correlation degree of each indicator is calculated based on the comprehensive anomaly co-occurrence correlation degree, statistical trend correlation degree, and causal inference correlation degree. Only when the comprehensive correlation degree exceeds the set threshold is the abnormal correlation between the indicator pairs confirmed.
5. The health management method based on multi-source heterogeneous monitoring data according to claim 4, characterized in that, No. Each indicator at time Dynamic weights The calculation formula is: ;in: Basic weights; The time factor is used to characterize the combined impact of anomaly duration and anomaly level on the weights. As a context factor, it is used to characterize the impact of business time period, load level, environmental state and special operating scenario on the weight; For the coupling correlation factor of the indicator; For normalization processing; the time factor is defined as: ;in: For the first The duration of each indicator since it entered the current abnormal level; For the first Each indicator at time The anomaly level multiplier, when the first When each indicator is in a normal, warning, or severe state, Take respectively , and .
6. The health management method based on multi-source heterogeneous monitoring data according to claim 5, characterized in that, Based on the paired samples, a residual vector is formed. Then, with the goal of minimizing the loss function, the parameter vector to be optimized is updated according to the parameter sensitivity matrix, the weights of each sample pair, and the residual vector. This includes the following sub-steps: Sub-step S41: Define the loss function Specifically: ;in: The vector of parameters to be optimized is At that time, the first Predicted health score output for each sample; Indicates the first The true target score corresponding to each sample; Indicates sample weights; This represents the previous parameter vector to be optimized used by the current online model; Indicates the smoothing coefficient of the parameter; The L2 norm constraint term represents the magnitude of parameter update; Sub-step S42: Construct the parameter sensitivity matrix Specifically: ;in: This represents the number of paired samples; for The dimension; the first dimension in the matrix line, number Column elements Indicates the first The health score of the sample is related to the first The sensitivity of each parameter is calculated using the finite difference method: ;in: Indicates the first A tiny perturbation applied to each parameter; Indicates the first The unit basis vectors corresponding to each parameter; Sub-step S43, calculate the first The parameter vector to be optimized. The specific calculation formula is as follows: ;in, Indicates the learning rate; Indicates by sample weights The resulting diagonal weighted matrix; This represents the current predicted score vector. Represents the true target score vector. Represents the residual vector; This indicates that the updated parameters will be projected onto the valid parameter constraint domain. Inside; This represents the approximate gradient of the loss function with respect to the parameter vector to be optimized.
7. The health management method based on multi-source heterogeneous monitoring data according to claim 6, characterized in that, Let the total number of specific indicators participating in the health assessment at a certain assessment time be . The total number of indicator categories is The total number of health levels is ;in, Indicates the first A specific indicator, Indicates the first Class indicators, Indicates the first Each health level, any specific indicator Belonging to a certain indicator category At any moment ,all The states of specific indicators are combined row by row to form an indicator state matrix. ,and The dynamic weight vector consists of the dynamic weights corresponding to each specific indicator. ; Sub-step S13 specifically involves: first targeting all Each specific indicator is calculated separately for its effect on... The membership degree of each health level is used to obtain the single-index membership degree matrix. Then, the system-level fuzzy evaluation vector is obtained by weighted aggregation of the dynamic weight vector and the single-index membership matrix: Then, based on the system-level fuzzy evaluation vector, defuzzification is performed to obtain the system-level health score. .
8. The health management method based on multi-source heterogeneous monitoring data according to claim 7, characterized in that, set up: ; Then we have: ; ; in, This is the Hadamard product operator; This represents a matrix composed of the dynamic weight vectors of each indicator. The value of and the parameter vector to be optimized Related; To form a basic weight matrix that integrates all indicators, The values of the elements in the time factor matrix, which integrates all indicators, are related to the time decay coefficient of the corresponding category and the emergency weight multiplier in the warning and severe states. This indicates that the values of the elements in the matrix of coupled correlation factors of all indicators are related to the influence intensity parameter and trigger threshold parameter of the corresponding category. Indicates from The matrix formed by the scene coefficients corresponding to the context-aware results of all extracted indicators; ;in, express The value of and the parameter vector to be optimized Related, Indicates a single indicator Membership calculation and this index The value at the current moment and corresponding indexes in the first The membership function center value and width parameter are related under each health level; ;in, express The value of and the parameter vector to be optimized Related; ;in, This indicates the representative score corresponding to each health level; express Corresponding health level The possible values of ; express The value of and the parameter vector to be optimized Related; ;in, Indicates according to The divided scoring intervals will A system-level health score obtained by integrating all indicators at all times Mapped to the corresponding health level The function.
9. The health management method based on multi-source heterogeneous monitoring data according to claim 8, characterized in that, The parameter constraints when performing update processing on the parameter vector to be optimized include: ; ; ;and To ensure that the system-level health score is greater than or equal to The output level is healthy, greater than or equal to and less than The output level is sub-healthy, greater than or equal to... and less than The output level is warning, less than... The output level is fault.
10. A health management system based on diverse heterogeneous monitoring data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method described in any one of claims 1 to 9.