A method and system for data center failure prediction and health management

CN122653883APending Publication Date: 2026-08-28SHAANXI ZHONGGUANG TELECOM HI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610814959.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0003]现有数据中心健康管理技术通常将服务器资源占用率、供电电流、环境温湿度、网络吞吐量、磁盘读写延迟和告警记录作为健康评估依据,通过阈值判定、趋势预测或异常检测模型生成设备健康评分,并根据评分结果执行迁移、限流、隔离、重启或维护提醒,该类技术能够对运行指标偏离进行识别,也能够为异常处置提供控制依据;但在预测性运维动作提前介入的场景中,异常指标回落的来源可能并不唯一,源设备异常指标在迁移、限流或制冷补偿后回落时,若缺少对控制动作生效区间、迁移承接对象和目标侧延期风险的关联记录,相关运行片段可能被归入正常恢复经验,而目标宿主机、目标机柜、目标供电回路、目标网络端口或目标存储池承接的新增压力,则可能因尚未超过即时告警阈值而未被同步纳入健康基线更新判断,由此,在数据中心受控运维场景下,仍需要一种能够将异常前兆、受控干预、迁移承接、延期确认和基线回写许可进行关联处理的故障预测与健康管理方式,以降低近失效证据被截断后造成样本归属偏差的风险,并提高后续维护决策依据的一致性

Benefits of technology

1.本发明通过以受控运维事件链作为统一承接对象,先限定源侧受控对象的前兆持续偏离状态,再校核受控运维动作归属,并在同一事件链内比较控制介入前后的标准化偏离回落证据,将自然恢复、受控运维动作截断故障演化、采样中断以及普通运行波动等容易混淆的来源限定在同一受控运维事件链下进行归属、筛除和准入,解决了现有数据中心健康管理中容易将被控制动作压下去的近失效风险误写成普通健康恢复样本的问题,由此,本发明能够把被干预抑制的近失效片段从正常样本中剥离出来,使临界风险样本池保留原本会被受控运维动作掩盖的风险证据,从源头上提高历史健康基线的样本纯度和故障预测依据的真实性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653883A_ABST
    Figure CN122653883A_ABST
Patent Text Reader

Abstract

The application discloses a kind of data center failure prediction and health management method and system, it is related to failure prediction and health management technical field;Its technical points are: event anchor point time correction and standardization deviation conversion are executed to operation event data, controlled operation and maintenance event chain is constructed after precursor sustained deviation state of source side controlled object appears, and target bearing object boundary is determined;Parse control action effective interval, identify near failure segment that is inhibited by intervention together with working condition natural fall-back range;Target side takes over pressure is extracted based on target bearing topology constraint graph, and migration risk debt segment is formed;With the time when migration load is stabilized, delay confirmation window is started, source side recurrence state and target side debt release state are checked, and health baseline write permission constraint result is generated, the present application can avoid controlled operation and maintenance risk miswriting as normal sample, improve failure prediction and baseline update credibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fault prediction and health management technology, specifically to a data center fault prediction and health management method and system. Background Technology

[0002] In data centers hosting server hosting, virtual hosting, and distributed storage services, servers, power supply circuits, cooling units, switching equipment, and storage nodes typically operate in a continuous and collaborative manner. Fluctuations in the operation of a single device can be transmitted to other resource domains through service scheduling, network paths, and rack environments. In actual operation and maintenance, when host load continues to rise, rack intake air temperature rises abnormally, current fluctuates in rack head cabinets or power distribution units, storage node read / write latency increases, or switching ports show signs of congestion, the health management platform usually triggers proactive measures based on real-time monitoring data. These measures include migrating virtual hosts out of high-risk hosts, switching service traffic to backup paths, redistributing storage replicas, limiting or isolating abnormal nodes, and coordinating with the cooling system to increase fan speed or adjust air supply parameters. These controlled operation and maintenance actions can reduce the abnormal indicators of the source device before a fault occurs, resulting in a decrease in load, a drop in temperature, a reduction in latency, or the elimination of alarms on the monitoring interface. Therefore, these actions are often used as an important reference for the effectiveness of operation and maintenance in short-term observations.

[0003] Existing data center health management technologies typically use server resource utilization, power supply current, ambient temperature and humidity, network throughput, disk read / write latency, and alarm records as health assessment criteria. They generate device health scores through threshold determination, trend prediction, or anomaly detection models, and then execute migration, rate limiting, isolation, restart, or maintenance alerts based on the score results. These technologies can identify deviations in operational indicators and provide a control basis for anomaly handling. However, in scenarios where predictive maintenance actions intervene early, the source of anomaly indicator decline may not be unique. When the source device's anomaly indicator declines after migration, rate limiting, or cooling compensation, if there is a lack of effective control action intervals... The correlation records between the migration recipient and the target side delay risk may be included in the normal recovery experience. However, the new pressure on the target host, target cabinet, target power supply circuit, target network port or target storage pool may not be included in the health baseline update judgment because it has not yet exceeded the real-time alarm threshold. Therefore, in the controlled operation and maintenance scenario of data center, there is still a need for a fault prediction and health management method that can correlate abnormal precursors, controlled intervention, migration acceptance, delay confirmation and baseline write-back permission to reduce the risk of sample attribution bias caused by the truncation of near-failure evidence and improve the consistency of subsequent maintenance decision-making. Summary of the Invention

[0004] To achieve the above objectives, the present invention provides the following technical solution: A data center fault prediction and health management method includes: S1 performs event anchor time calibration on the operational event data generated during the controlled operation and maintenance process of the data center, performs standardized deviation conversion on the source side controlled object, and constructs the controlled operation and maintenance event chain based on the controlled operation and maintenance action attribution verification results after the source side controlled object shows a precursor continuous deviation state, and determines the boundary of the target carrying object. S2, within the controlled operation and maintenance event chain, analyze the effective interval of control actions, compare the standardized deviation and fallback evidence before and after control intervention, eliminate natural recovery by using the natural fallback range under the same working conditions extracted from the historical health baseline, identify the operation segment whose fault evolution is truncated by control actions as the near failure segment suppressed by intervention, and write it into the critical risk sample pool. S3, generate a target bearing topology constraint diagram according to the boundary of the target bearing object, extract the target side bearing pressure exceeding the upper limit of the natural pressure increment in the associated resource domain of the actual landing point of the migration load, and form a migration risk debt fragment based on the temporal bearing evidence and standardized evidence bearing relationship of the source side standardized deviation fallback evidence according to the target side bearing pressure. S4 initiates the delayed confirmation window when the migration load stabilizes. Within the delayed confirmation window, the recurrence status on the source side and the debt release status on the target side are checked. Based on the check results, a health baseline write-back permission constraint result is generated. The health baseline write-back permission constraint result is used to limit the status of the controlled operation and maintenance event chain entering the normal sample pool, the critical risk sample pool, the controlled successful sample record, and the record to be reviewed.

[0005] This invention also discloses a data center fault prediction and health management system, comprising: The controlled operation and maintenance event chain construction module performs event anchor time calibration on the running event data generated during the controlled operation and maintenance process of the data center, performs standardized deviation conversion on the source side controlled objects, and constructs the controlled operation and maintenance event chain based on the controlled operation and maintenance action attribution verification results after the source side controlled objects show a precursor continuous deviation state, and determines the boundary of the target carrying object. The near-failure segment identification module analyzes the effective interval of control actions within the controlled operation and maintenance event chain, compares the standardized deviation and fallback evidence before and after control intervention, eliminates natural recovery by using the natural fallback range under the same working conditions extracted from the historical health baseline, and identifies the operation segments whose fault evolution is truncated by control actions as near-failure segments that have been suppressed by intervention, and writes them into the critical risk sample pool. The migration risk debt tracking module generates a target bearing topology constraint diagram according to the boundary of the target bearing object. It extracts the target-side bearing pressure exceeding the upper limit of the natural pressure increment in the associated resource domain of the actual landing point of the migration load. Based on the target-side bearing pressure, it forms a migration risk debt fragment by combining the temporal bearing evidence and standardized evidence bearing relationship of the standardized deviation fallback evidence of the source side. The write-back permission constraint module initiates a delayed confirmation window when the migration load stabilizes. Within the delayed confirmation window, it verifies the recurrence status on the source side and the debt release status on the target side. Based on the verification results, it generates a health baseline write-back permission constraint result. The health baseline write-back permission constraint result is used to limit the status of the controlled operation and maintenance event chain entering the normal sample pool, the critical risk sample pool, the controlled successful sample record, and the record to be reviewed.

[0006] This invention provides a method and system for data center fault prediction and health management, which has the following beneficial effects: 1. This invention uses a controlled operation and maintenance event chain as a unified receiving object. It first defines the precursory continuous deviation state of the controlled object on the source side, then verifies the attribution of controlled operation and maintenance actions, and compares the standardized deviation fallback evidence before and after control intervention within the same event chain. It limits easily confused sources such as natural recovery, controlled operation and maintenance action-interrupted fault evolution, sampling interruption, and normal operational fluctuations to the same controlled operation and maintenance event chain for attribution, screening, and admission. This solves the problem in existing data center health management where near-failure risks suppressed by controlled actions are easily miswritten as ordinary health recovery samples. Thus, this invention can separate the near-failure fragments suppressed by intervention from normal samples, allowing the critical risk sample pool to retain risk evidence that would otherwise be covered by controlled operation and maintenance actions, thereby improving the sample purity of historical health baselines and the authenticity of fault prediction basis from the source.

[0007] 2. This invention generates a target-bearing topology constraint diagram according to the boundary of the target-bearing object. It extracts the target-side bearing pressure exceeding the upper limit of the natural pressure increment only within the associated resource domain of the actual landing point of the migrated load. It limits the source-side standardized deviation fallback evidence, target-side pressure changes, actual landing point of the migrated load, and target-bearing topology relationship to the same controlled operation and maintenance event chain for topology screening, natural fluctuation elimination, and standardized evidence acceptance verification. This solves the problem in existing operation and maintenance analysis that only focuses on the fallback of source-side indicators while ignoring the risk migration destination, or misattributes unrelated resource fluctuations on the target side to migration consequences. In this way, it can form migration risk debt fragments, making the source side appear to recover and establishing a verifiable debt relationship with the target-side bearing pressure. This avoids the migration action creating a false sense of health on the source side, thus giving the subsequent target-side debt release status verification clear object boundaries, time boundaries, and evidence constraints.

[0008] 3. This invention initiates a delayed confirmation window at the moment the migration load stabilizes, jointly verifying the recurrence status on the source side, the debt release status on the target side, the confidence level of evidence, and the data integrity status. It restricts the permission judgment of limited write-back permissions, write-back prohibition results, and write-back results awaiting review to the same controlled maintenance event chain. This solves the problems of premature write-back after controlled maintenance actions are completed, low-confidence samples entering the normal sample pool, and updating the health baseline before the migration risk debt is released. Therefore, this invention isolates controlled successful sample records from the normal sample pool, making the health baseline write-back subject to the joint constraints of near-failure tags, migration risk debt, and delayed confirmation results, thereby improving the reliability of subsequent fault prediction, successor selection, and health baseline updates. Attached Figure Description

[0009] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a system framework diagram of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0011] Example 1: Please see Figure 1 This embodiment provides a data center fault prediction and health management method, including the following specific steps: S1 Perform event anchor time calibration on the operational event data generated during the controlled operation and maintenance process of the data center, and perform standardized deviation conversion on the source-side controlled object; After the source-side controlled object shows a precursory continuous deviation state, construct a controlled operation and maintenance event chain based on the controlled operation and maintenance action attribution verification result, and determine the target bearing object boundary. The controlled operation and maintenance event chain serves as the input for S2 to parse the effective range of control actions and identify the near-failure segments that have been intervened and suppressed. The target bearing object boundary serves as the input for S3 to generate the target bearing topology constraint diagram and extract the target-side bearing pressure.

[0012] The system acquires operational event data generated during the controlled operation and maintenance (O&M) process of the data center. This data includes operational observation records, controlled O&M action records, resource scheduling landing point records, O&M handling records, and bearer relationship records. Operational observation records serve as input for standardized deviation conversion, characterizing the resource occupancy, network links, storage performance, and environmental status of the source-side controlled objects. Controlled O&M action records serve as input for the verification of controlled O&M action attribution, characterizing the execution and completion times of migration, rate limiting, service traffic switching, storage replica redistribution, and cooling compensation. Resource scheduling landing point records serve as the basis for confirming the target bearer objects, characterizing the actual landing points of migrated loads, service traffic, and storage replicas after controlled O&M actions. O&M handling records serve as auxiliary verification basis for the action reasons and handling objects. Bearer relationship records serve as the basis for determining the boundaries of the target bearer objects, characterizing the host machine, rack, power supply circuit, cooling zone, network path, and storage access relationships.

[0013] A common event timeline is established using the execution and completion times of controlled maintenance actions as basic event anchors. When a migration load stability confirmation record has been formed in the resource scheduling landing point record, the migration load stability time is written into the common event timeline as an auxiliary event anchor. When the migration load stability time has not yet been formed, the event anchor time of S1 is synchronized only with the execution and completion times. The migration load stability time is formed before subsequent delayed confirmation. The common event timeline is used to unify the time base of operation observation records, controlled maintenance action records, and resource scheduling landing point records without changing the original sampling granularity. When the sampling periods of different data sources are inconsistent, time window aggregation is only performed during the window statistics stage, without directly splicing the original physical quantities. The sampling period for server resource indicators is preferably 5 to 60 seconds. The sampling period for network link indicators is preferably 5 to 30 seconds, the sampling period for storage performance indicators is preferably 10 to 60 seconds, and the sampling period for rack temperature and cooling status indicators is preferably 30 to 300 seconds. The time synchronization error of the event anchor point in the operation observation record is preferably controlled within half of the corresponding indicator's sampling period. The time synchronization error of action execution record, action completion record, and resource scheduling landing point record is preferably no more than 5 seconds. The aforementioned value range is determined by the monitoring sampling configuration, resource scheduling log accuracy, and historical operation and maintenance review records. When the time error exceeds the corresponding range, a low confidence marker is added to the relevant operation event data. When the missing timestamp makes it impossible to distinguish between the pre-action window and the post-action window, the relevant operation event data is not used for the precursor access judgment, and the corresponding event is transferred to the pending review process.

[0014] After completing the event anchor time calibration, available historical segments qualified for baseline updates are extracted from the historical health baseline. Available historical segments are operational segments that have not been excluded by the health baseline write-back permission constraint results of the historical stage. Event chains that have generated write-back prohibition results, event chains corresponding to write-back results pending review, event chains with undetermined target carrier boundaries, and controlled operation and maintenance event chains that have not obtained limited write-back permission are not included in the available historical segments. This process ensures that the source-side standardized deviation caused by the early intervention of controlled operation and maintenance actions is not directly written into the historical health baseline, reducing the risk of the historical health baseline being contaminated by controlled operation and maintenance events.

[0015] Using the workload range as the basis for operational condition stratification, available historical segments are merged into samples based on the same operational condition, taking into account equipment roles, environmental status, and maintenance status. The workload range is formed by the workload distribution in the historical health baseline and can be calibrated using peak, off-peak, and low-peak periods from data center operation and maintenance procedures. Equipment roles include host machines, storage nodes, network switching equipment, and rack cooling-related objects. Environmental status includes rack intake air temperature range, cooling unit operating mode, and external environmental fluctuations. Maintenance status includes normal operation, pre-planned maintenance, and post-planned maintenance stable status. Each sample of samples preferably contains no fewer than 20 available historical segments, and the effective sampling ratio of a single available historical segment is preferably no less than 80%. When the aforementioned conditions are met, a baseline for the same operating condition is formed. If there are fewer than 20 samples for the same operating condition but no fewer than 5, a temporary baseline is formed using available historical fragments of the same equipment role, and a low-confidence flag is added to the precursor access result formed by the temporary baseline. If there are fewer than 5 samples for the same operating condition, an independent baseline for the current source-side controlled object is not formed, and the corresponding event is transferred to the pending review process. The low-confidence flag does not block the construction of the controlled operation and maintenance event chain, but restricts the automatic write-back of the subsequent health baseline. The controlled operation and maintenance event chain with the low-confidence flag can enter S2 for control action effective interval analysis, but it cannot be directly entered into the historical health baseline as a normal sample. When the low-confidence flag involves the boundary of the target bearing object, S3 does not extract the target-side bearing pressure, and the corresponding event chain is transferred to the pending review process.

[0016] The baseline central value and natural fluctuation scale of the corresponding index of the source-side controlled object are extracted from the same operating condition sample. The baseline central value is formed by the median level of the effective observations of the corresponding index in the same operating condition sample. The natural fluctuation scale is formed by the robust dispersion of the corresponding index around the baseline central value in the same operating condition sample. The robust dispersion is obtained by converting at least one of the median absolute deviation, interquartile range and similar robust dispersion quantities, and is converted into a fluctuation scale with the same dimension as the observations of the corresponding index. This processing is used to reduce the impact of short-term load changes, temporary flow switching and local temperature disturbances on the standardized deviation conversion.

[0017] Standardized deviation conversion is performed on the current observations of the controlled objects on the source side. Standardized deviation conversion is used to convert processor utilization, memory usage, rack temperature, power supply current, storage read / write latency, network packet loss rate, link latency, remaining capacity, available bandwidth, and available storage space into dimensionless deviation evidence in the same risk direction. During the conversion, the offset of the current observation relative to the baseline center value under the same operating condition is first processed to be consistent according to the risk direction of the indicator, and then processed to be dimensionless using the natural fluctuation scale of the same operating condition to form a standardized deviation. An increase in processor utilization, rack temperature, power supply current, storage read / write latency, network packet loss rate, and link latency represents an increase in risk; a decrease in remaining capacity, available bandwidth, and available storage space represents an increase in risk. When the conversion result does not point to the risk direction, the corresponding standardized deviation is recorded as 0.

[0018] The standardized deviation conversion is expressed as follows: Dsq(t) = rq × [xsq(t) - Msqc(t)] / [Ssqc(t) + e]; where s represents the object identifier of the source-side controlled object, which comes from the source-side controlled object identifier in the running event data; q represents the indicator identifier, which comes from the precursor indicator set; t represents the sampling time on the common event time axis, which comes from the time mapping result after the event anchor point is calibrated; c(t) represents the same working condition category corresponding to time t, which comes from the same working condition merging result formed by business load range, equipment role, environmental status and maintenance status; Dsq(t) represents the standardized deviation of the source-side controlled object s at time t relative to indicator q, which is the dimensionless result output by this formula, used to form the precursor admission result, and as the subsequent standardized deviation return. The basis for this is evidence; xsq(t) represents the observed value of the source-side controlled object s for index q at time t, derived from operational observation records; Msqc(t) represents the baseline central value of index q of the source-side controlled object s under the same operating condition category c(t), derived from samples under the same operating condition; Ssqc(t) represents the natural fluctuation scale of index q of the source-side controlled object s under the same operating condition category c(t), derived from the robust discrete conversion result of the corresponding index in the samples under the same operating condition; rq represents the risk direction coefficient of index q, derived from data center operation and maintenance procedures, equipment manuals, and historical fault review records. When the index increases, it is taken as +1 when the risk increases, and when the index decreases, it is taken as -1 when the risk increases; e represents the zero-prevention parameter, derived from monitoring sampling configuration and historical health baseline, and has the same dimensions as the observed value of index q.

[0019] The formula first characterizes the deviation of the current operating state from the benchmark by the offset between the current observed value and the benchmark value under the same operating condition. Then, it unifies the upward and downward risk indicators into the same risk direction through the risk direction coefficient. Subsequently, it performs dimensionless processing using the natural fluctuation scale, so that indicators with different dimensions can enter the same precursor access judgment chain. The natural fluctuation scale and the current observed value have the same dimension, and the zero-prevention parameter and the natural fluctuation scale have the same dimension, so the denominator is not zero, and the fractional result is a dimensionless result. The risk direction coefficient is a dimensionless direction coefficient and is not used as a weight; the zero-prevention parameter is only used to prevent the natural fluctuation scale. A value of 0 or too small will cause conversion distortion and will not be used as a precursor entry boundary. When the conversion result is less than 0, the standardized deviation will be recorded as 0, and only the deviation evidence in the risk direction will be retained. This formula serves the identification of the continuous deviation state of precursors and the construction of the controlled operation and maintenance event chain, so that subsequent steps can compare the standardized deviation fallback evidence before and after control intervention under a unified evaluation scale, and reduce the risk of misjudgment caused by the direct mixing of indicators of different dimensions. The zero-prevention parameter is preferably taken as 0.1 to 1 times the minimum acquisition resolution of the corresponding monitoring indicator. When the monitoring platform does not provide a minimum acquisition resolution, the zero-prevention parameter is determined by the minimum non-zero fluctuation amplitude of the same type of indicator in the historical health baseline.

[0020] Based on standardized deviations, precursory admission results are generated by first reading the precursory indicator set. This set originates from data center operation and maintenance procedures, historical fault replay records, and the device roles of controlled objects on the source side. For the host machine, the precursory indicator set includes at least processor utilization, memory usage, virtual machine load, and host temperature. For storage nodes, it includes at least read / write latency, queue depth, storage pool load, and replica migration status. For network paths, it includes at least port occupancy, packet loss rate, and link latency. For rack cooling-related objects, it includes at least intake air temperature, return air temperature, and cooling... For unit operating status, the standardized deviation is compared with the precursor access boundary. The precursor access boundary is taken from the 90% to 97.5% quantile record of the risk direction deviation distribution of the healthy samples under the same operating conditions, and is combined with the calibration of the pre-alarm range in the operation and maintenance procedures. For network link and storage performance objects, the precursor access boundary preferably adopts the higher quantile record in this range; the precursor access boundary of temperature, power supply and cooling status objects is combined with the calibration of the pre-alarm range in the equipment operation and maintenance procedures; the precursor access boundary of newly launched equipment, equipment with few samples and equipment that has been running for a short period of time after maintenance is formed using a temporary benchmark of the same equipment role, and a low confidence mark is added.

[0021] To determine whether a controlled object on the source side has entered a state of precursory persistent deviation, the persistence condition is used to determine whether the deviation has entered a state of precursory persistent deviation. For indicators such as server resources, network links, and storage performance, the standardized deviation must reach the precursory admission boundary within 2 to 5 consecutive sampling periods. For indicators such as temperature and cooling status, the standardized deviation must reach the precursory admission boundary within 2 to 3 consecutive valid sampling points. The duration is defined as the time difference between the first valid sampling moment that meets the precursory admission boundary and the last valid sampling moment that consecutively meets the precursory admission boundary. The duration is between 2 and 10 minutes. The values ​​of the persistence condition are based on the sampling period of the corresponding indicator, physical response lag, and historical fault replay records. When a sustained deviation from the precursor state is established, the precursor admission result is output. The precursor admission result includes at least the source-side controlled object identifier, the time of precursor occurrence, the precursor indicator set, the intensity of the abnormal precursor, the same operating condition category, the data availability status, and the low confidence marker status. The intensity of the abnormal precursor is formed by the standardized deviation exceeding the precursor admission boundary, the duration of the sustained deviation state, and the data availability status. First, the indicators that reach the precursor admission boundary in the precursor indicator set are extracted, and then the degree of exceedance of the corresponding standardized deviation relative to the precursor admission boundary is calculated. Then, the intensity of the abnormal precursor is formed by combining the duration and the effective sampling ratio, which is used to determine the strength of risk evidence before control action intervention in subsequent steps.

[0022] After the precursor access results are formed, the moment when the source-side controlled object shows a continuous deviation from the precursor state is taken as the starting point for event retrieval. Controlled operation and maintenance action records are read on the common event timeline, and controlled operation and maintenance action attribution verification is performed. Controlled operation and maintenance action attribution verification includes object coverage verification, response timing verification, and action interpretation verification. Object coverage verification is used to confirm whether the object recorded in the controlled operation and maintenance action covers the source-side controlled object. When the object recorded in the controlled operation and maintenance action is confirmed to be in at least one of the following object coverage relationships with the source-side controlled object in the resource scheduling record, configuration management database, or operation and maintenance handling record: same host machine, same business instance, same network path, same storage pool, same cabinet cooling association range, the response timing verification is performed. If there is only similar name, similar work order description, or the same large management area, the object coverage verification is not considered to have passed. Response timing verification is used to confirm whether the action execution time is within the respondable period of the same type of action relative to the moment when the continuous deviation from the precursor state occurs.

[0023] The responsive timeframe is formed from historical operation and maintenance review records and resource scheduling action logs. The responsive timeframe for virtual machine migration actions is 1 to 30 minutes after the occurrence of a precursor; for rate limiting actions, it is 10 seconds to 10 minutes after the occurrence of a precursor; for cooling compensation actions, it is 1 to 20 minutes after the occurrence of a precursor; and for network switching actions, it is 10 seconds to 5 minutes after the occurrence of a precursor. Action interpretation verification is used to confirm whether there is an action interpretation relationship between the action type and the precursor indicator set. This action interpretation relationship is only used for the attribution verification of controlled operation and maintenance actions and does not replace the standardized deviation return in S2. For the determination of evidence, migration actions can explain controlled maintenance interventions that deviate from the correlation between source-side load, storage read / write latency, and temperature; rate limiting actions can explain controlled maintenance interventions that deviate from the correlation between source-side load and network path occupancy; cooling compensation actions can explain controlled maintenance interventions that deviate from the correlation between temperature; and network switching actions can explain controlled maintenance interventions that deviate from the correlation between link latency, packet loss rate, and path congestion. When restart actions and isolation actions cause source-side controlled object sampling interruption, service instance reconstruction, or monitoring objects temporarily leaving the normal operating state, they are not directly considered as valid action explanations, and the corresponding events are transferred to the pending review process.

[0024] When the object coverage verification, response timing verification, and action interpretation verification all pass, a controlled operation and maintenance action attribution verification result is generated. The action completion record is read to confirm that the controlled operation and maintenance action has been executed. Then, based on the resource scheduling landing point record, the target bearer object that actually undertakes the migration load is confirmed. For virtual machine migration, the target bearer object is the target host machine and its associated bearer resources that actually run the corresponding virtual machine after the migration is completed. For business traffic switching, the target bearer object is the target network path and its associated switching resources after the traffic switching. For storage replica redistribution, the target bearer object is the target storage pool and target storage node that undertake the replica write and read / write access pressure. For cooling compensation, the boundary of the target bearer object is based on the cooling bearing range corresponding to the compensated cabinet. Cooling compensation events that have not occurred only form the boundary of the target bearer object within the cooling bearing range and are not used as the triggering basis for subsequent migration risk debt fragments. When the action completion record and the resource scheduling landing point record are inconsistent, the actual landing point in the business instance running location, the actual landing point of the storage replica, and the actual landing point in the network path switching record shall prevail. When the actual landing point cannot be confirmed, the boundary of the target bearer object is uncertain, and the corresponding controlled operation and maintenance event chain is transferred to the pending review process.

[0025] Based on the bearing relationships recorded in the configuration management database, the infrastructure bearing domain corresponding to the target bearing object is written into the target bearing object boundary. The infrastructure bearing domain includes at least the rack where the target host is located, the corresponding power supply circuit, the corresponding cooling area, the target network path, the target storage pool, and the target storage node. The target bearing object boundary is based on the actual resource landing point that undertakes the migration load and its direct bearing relationship, and is not based solely on the work order object name and the scheduling plan object name. By limiting the boundary, it is possible to avoid including the observed changes of irrelevant resource objects in the subsequent target-side load extraction process.

[0026] At the end of S1, the precursor admission results, controlled operation and maintenance action records, controlled operation and maintenance action attribution verification results, control action execution time, action completion time, source-side controlled object identifier, target carrier object boundary, data availability status, and low confidence flag status are written into the same controlled operation and maintenance event chain. The controlled operation and maintenance event chain serves as the input for S2 to parse the effective range of control actions and identify near-failure segments suppressed by intervention. The target carrier object boundary serves as the input for S3 to generate the target carrier topology constraint diagram and extract the target-side bearing pressure. Through this process, a traceable event chain relationship is established between the precursor continuous deviation state of the source-side controlled object, controlled operation and maintenance actions, and the target carrier object, enabling subsequent health baseline write-back permission judgment to distinguish between natural recovery, control action truncation fault evolution, and target-side bearing pressure transfer.

[0027] S2 is used to analyze the effective interval of control actions within the controlled operation and maintenance event chain, compare the standardized deviation fallback evidence before and after control intervention, eliminate natural recovery by using the natural fallback range of the same working condition extracted from the historical health baseline, identify the operation segment whose fault evolution was truncated by the control action as the near failure segment suppressed by intervention, and write it into the critical risk sample pool. This step does not directly identify the fallback of the standardized deviation of the source side controlled object as a healthy recovery, but rather, within the object range and time range defined by the controlled operation and maintenance event chain, it judges whether the standardized deviation fallback evidence exceeds the natural fallback range of the same working condition under the scenario without control actions, so as to distinguish between natural recovery candidate segments and near failure segments suppressed by intervention.

[0028] Read the controlled operation and maintenance event chain output by S1. The controlled operation and maintenance event chain includes at least the source-side controlled object identifier, the precursor admission result, the precursor indicator set, the controlled operation and maintenance action record, the controlled operation and maintenance action attribution verification result, the control action execution time, the action completion time, the target carrier object boundary, the standardized deviation, the data availability status, and the low confidence mark status. If the precursor admission result is not established, the identification of the near failure segment that has been intervened and suppressed will not be performed. When the controlled operation and maintenance event chain is marked as pending review due to missing timestamps, undetermined target carrier object boundaries, or the source-side controlled object temporarily leaving the normal operation state, the corresponding event chain will not be written into the critical risk sample pool.

[0029] The effective interval of control actions is determined based on the action type and action execution records in the controlled operation and maintenance event chain. For migration actions, the migration completion time is used as the response baseline time, and the effective interval of control actions is formed by combining the migration load stability confirmation record, the response lag of source-side load indicators, and the response lag of source-side storage read / write latency. The response lag of the indicators corresponding to the migration action is preferably 1 to 15 minutes, and the value is based on the resource scheduling action log and historical operation and maintenance review records. When the migration load stability confirmation record has not yet been formed, a temporary effective interval of control actions is formed by the migration completion time and the continuous effective sampling results on the source side, and a low confidence mark is attached. The effective interval of temporary control actions with a low confidence mark does not support the near-failure fragments that have been intervened and suppressed to be written into the critical risk sample pool. It needs to participate in the subsequent admission judgment after the migration load stability confirmation record is supplemented.

[0030] For rate limiting actions, the execution time of the rate limiting command is used as the response baseline time, and the effective interval of the control action is formed by combining the confirmation time of the rate limiting command and the response lag of the indicator consistent with the resource affected by the rate limiting policy. When the rate limiting policy applies to computing resources, the source-side load response lag is used; when the rate limiting policy applies to network paths, the network path occupancy response lag is used; and when the rate limiting policy applies to storage access, the storage access pressure response lag is used. The preferred response lag of the indicator corresponding to the rate limiting action is 10 seconds to 10 minutes, and the value is based on the rate limiting policy log, sampling period, and historical operation and maintenance review. For cooling compensation actions, the effective time of the cooling parameters is used as the response benchmark time. Combined with the air supply adjustment response time and the cabinet temperature response cycle, the effective range of the control action is formed. The response lag of the indicators corresponding to the cooling compensation action is preferably 2 to 20 minutes. The values ​​are based on the cooling unit operation log, the cabinet temperature sampling cycle, and historical operation and maintenance review records. When restart actions and isolation actions cause the source-side controlled object sampling to be interrupted, the service instance to be rebuilt, or the monitored object to be temporarily out of normal operation, the corresponding controlled operation and maintenance event chain will be transferred to the pending review process and will not enter the near failure suppression identification.

[0031] A precursor retention window is constructed before the control action takes effect, and an effective observation window is constructed within the control action's effective period. The precursor retention window characterizes the persistent deviation state of the controlled object on the source side that already exists before the control action intervention. The effective observation window characterizes the standardized deviation change of the controlled object on the source side after the control action takes effect. The length of the precursor retention window covers 2 to 5 effective sampling periods and is not shorter than the minimum persistence requirement for the persistent deviation state. The length of the effective observation window covers no less than 3 effective sampling points within the control action's effective period. For temperature and cooling status indicators, the effective observation window also covers at least 1 cabinet temperature response period. The window length is determined based on the corresponding indicator's sampling period, control action response lag, and historical maintenance review records. If an action is executed too early, a sufficient length of precursor maintenance window cannot be formed before the control action takes effect, thus failing to establish a pre-control deviation level. The corresponding event chain will then be transferred to the pending review process. If the effective observation window has fewer than three valid sampling points, standardized deviation fallback evidence will not be formed, and the corresponding event chain will be transferred to the pending review process. If the effective sampling ratio is less than 80% in any window, the corresponding indicator will not participate in the formation of standardized deviation fallback evidence. If the proportion of valid indicators in the precursor indicator set is less than 50%, or if the precursor indicator set contains only one indicator that is unavailable in any window, the corresponding event chain will be transferred to the pending review process. Both the effective sampling ratio and the proportion of valid indicators are dimensionless ratios, and their values ​​are based on the monitoring sampling configuration, the size of the precursor indicator set, and historical operation and maintenance review records.

[0032] The pre-control deviation level is formed based on the standardized deviation within the precursor maintenance window, and the post-control deviation level is formed based on the standardized deviation within the effective observation window. The pre-control deviation level is formed by the median standardized deviation of the same precursor indicator within the precursor maintenance window, and the post-control deviation level is formed by the median standardized deviation of the same precursor indicator within the effective observation window. The pre-control deviation level and the post-control deviation level are compared separately for the same precursor indicator. The decrease in the pre-control deviation level relative to the post-control deviation level is considered as the candidate decline portion of that precursor indicator, and only the downward direction that can be controlled by the controlled movement is retained. The candidate fallback portions of the maintenance action explanation form standardized deviation fallback evidence. For migration actions, the fallback portions related to source-side load, storage read / write latency, network path occupancy, and temperature are retained. For rate limiting actions, the fallback portions related to source-side load, network path occupancy, and storage access pressure are retained. For cooling compensation actions, the fallback portions related to temperature-related deviations are retained. Fallbacks of indicators that do not match the action type are not included in the standardized deviation fallback evidence. This process ensures that the standardized deviation fallback evidence is consistent with the target and mechanism of the controlled maintenance actions, avoiding the inclusion of irrelevant indicator fluctuations in the near-failure suppression judgment.

[0033] Extracting uncontrolled operation segments from the historical health baseline under the same operating conditions, a natural decline range for these segments is formed. These uncontrolled operation segments refer to available historical segments that did not undergo migration, rate limiting, business traffic switching, storage copy redistribution, cooling compensation, restart, isolation, or maintenance work order processing within the corresponding time period. Based on the same precursor indicator, the same equipment role, and the same operating condition category, the standardized deviation natural decline records in these historical segments are paired to form the natural decline distribution of the corresponding precursor indicator. The upper quantile record of the risk-direction decline distribution is then used to form the natural decline range for these segments. For source-side controlled objects with at least 20 uncontrolled operation segments, the natural decline range is taken from the 90% to 97.5% quantile records. For network link and storage performance indicators, the higher percentile record within this range is used. For temperature and cooling status indicators, the calibration is performed by combining the cabinet temperature response cycle and the pre-alarm range in the operation and maintenance procedures. For source-side controlled objects with fewer than 20 but no fewer than 5 operating segments under the same operating conditions and with historical segments of the same equipment role, the historical segments of the same equipment role are used to form a temporary natural fallback range, and a low confidence mark is added. When there are fewer than 5 operating segments under the same operating conditions, and the historical segments of the same equipment role are still insufficient to form a temporary natural fallback range, the identification of near-failure segments suppressed by intervention is not performed, and the corresponding event chain is transferred to the pending review process. The aforementioned value range is determined by the historical health baseline, the number of samples under the same operating conditions, and the historical operation and maintenance review records.

[0034] The standardized deviation decline evidence is compared with the natural decline range under the same operating conditions. When the standardized deviation decline evidence does not exceed the natural decline range under the same operating conditions, and the continuous deviation state of the precursor subsides within the natural decline range under the same operating conditions, a natural recovery candidate fragment is generated. The natural recovery candidate fragment is not written into the critical risk sample pool, is not written into the historical health baseline as a controlled successful sample record, is not used to reduce the precursor admission boundary of the same precursor indicator, and does not enter the migration risk debt fragment formation process. When the standardized deviation decline evidence exceeds the natural decline range under the same operating conditions, and the decline occurs within the effective period of the control action, and there is an action interpretation relationship between the decline indicator and the controlled operation and maintenance action type, near failure suppression identification is further performed.

[0035] The near-failure suppression identification function reads standardized deviation fallback evidence, the natural fallback range under the same operating conditions, the controlled maintenance action attribution verification results, the control action effective interval matching results, the precursor admission results, and the data availability status to form an intervention suppression confidence score. The intervention suppression confidence score is a dimensionless result ranging from 0 to 1. The near-failure suppression identification function is implemented using interpretable rules. The state where standardized deviation fallback evidence exceeds the natural fallback range under the same operating conditions is converted into an over-boundary rule item; the control action effective interval matching state is converted into a time-series rule item; the controlled maintenance action attribution verification state is converted into an attribution rule item; and the data availability status of the precursor admission results is converted into an availability rule item. Each rule item is categorized as satisfied, not satisfied, and low confidence. Three types of status records are recorded, and a rule table formed from historical operation and maintenance review samples is mapped to the intervention and suppression confidence level. When any key input has a low confidence marker, the intervention and suppression confidence level is lowered according to the rule table. The admission requirements for intervention and suppression confidence level are formed by historical operation and maintenance review samples. When there are no less than 20 historical confirmed controlled suppression segments and natural recovery segments, the 10% to 25% lower quantile records of the intervention and suppression confidence level distribution of historical confirmed controlled suppression segments are used to form the admission boundary, and the natural recovery segments are used to check the admission boundary for false admission. When there are not enough historical operation and maintenance review samples to form the admission boundary, a low confidence marker is output, and the corresponding event chain is transferred to the pending review process. Parameterized training results are not automatically generated.

[0036] When the evidence of standardized deviation decline exceeds the natural decline range under the same operating conditions, the matching result of the effective interval of the control action is valid, the verification result of the controlled operation and maintenance action is valid, and the confidence level of intervention and suppression meets the admission requirements, the corresponding operation segment is identified as a near-failure segment that has been intervened and suppressed. The near-failure segment that has been intervened and suppressed includes at least the controlled operation and maintenance event chain identifier, the source-side controlled object identifier, the time of occurrence of the precursor, the effective interval of the control action, the set of precursor indicators, the evidence of standardized deviation decline, the natural decline range under the same operating conditions, the confidence level of intervention and suppression, the data availability status, and the low confidence mark status. The near-failure segment that has been intervened and suppressed is written into the critical risk sample pool and a sample attribution mark prohibiting it from entering the ordinary normal sample pool is attached to it. This process is used to prevent the decline of the source-side standardized deviation caused by the early intervention of controlled operation and maintenance actions from being identified as normal health recovery, thereby reducing the risk of the historical health baseline being contaminated by near-failure segments.

[0037] When the evidence of standardized deviation decline does not exceed the natural decline range under the same working conditions, and there are no situations such as missing target carrying object boundaries, insufficient effective window sampling, or invalid action attribution, a natural recovery candidate segment is generated. The natural recovery candidate segment is only used as a historical review record and does not enter the migration risk debt segment formation process. When the source-side standardized deviation decline occurs outside the effective range of control actions, the decline direction cannot be explained by controlled operation and maintenance actions, the control action record is incomplete, the target carrying object boundary is not determined, the effective window sampling is insufficient, the historical samples are insufficient, or the near failure suppression identification function outputs a low confidence mark, a segment to be reviewed is generated. The segment to be reviewed does not enter the critical risk sample pool and is not used as a normal sample to update the historical health baseline.

[0038] At the end of S2, the near-failure segments suppressed by intervention, natural recovery candidate segments, and segments to be reviewed are output. Among them, the near-failure segments suppressed by intervention with the actual landing point of the migration load and the boundary of the target bearing object enter the migration risk debt segment formation process of S3. The near-failure segments suppressed by intervention, natural recovery candidate segments, and segments to be reviewed that have not undergone migration load transfer do not enter the main process of S3. Through the above processing, a verifiable suppression relationship is established between the precursor continuous deviation state of the source side controlled object, the controlled operation and maintenance action, and the effective interval of the control action, so that the formation of subsequent migration risk debt segments is based on near-failure evidence within the controlled operation and maintenance event chain, rather than based on the decline of ordinary indicators.

[0039] S3 is used to generate a target load topology constraint graph according to the boundary of the target load object, extract the target side bearing pressure exceeding the upper limit of the natural pressure increment in the associated resource domain of the actual landing point of the migration load, and form a migration risk debt fragment based on the temporal bearing evidence and standardized evidence bearing relationship of the source side standardized deviation fallback evidence according to the target side bearing pressure. This step is used to determine whether the source side standardized deviation fallback evidence is accompanied by the formation of bearing pressure of the target load object, so that the target side bearing pressure enters the subsequent healthy baseline write-back permission constraint.

[0040] Read the near-failure segments that have been intervened and suppressed from the output of S2, and the target carrier object boundary formed by S1. The near-failure segments that have been intervened and suppressed include at least the controlled operation and maintenance event chain identifier, the source-side controlled object identifier, the set of precursor indicators, the standardized deviation fallback evidence, the effective interval of the control action, the intervention and suppression confidence level, and the data availability status. The target carrier object boundary includes at least the actual landing point of the migration load and the infrastructure resource objects that have a direct carrier relationship with the actual landing point. The migration risk debt segment formation process is executed only when the near-failure segments that have been intervened and suppressed have the actual landing point of the migration load and the target carrier object boundary. Near-failure segments that have been intervened and suppressed without migration load transfer do not enter the migration risk debt segment formation process.

[0041] Verify the integrity of the target bearer object boundary. Verification inputs include resource scheduling landing point records, action completion records, bearer relationship records in the configuration management database, and controlled operation and maintenance event chain identifiers. Compare the actual landing point in the resource scheduling landing point record with the completed object in the action completion record. If they match, the actual landing point is confirmed as the target bearer object. If they do not match, the actual landing point in the business instance running location, storage copy actual landing point, and network path actual switching record shall prevail. If the actual landing point cannot be confirmed and the target bearer object boundary lacks a direct bearer relationship record, no target bearer topology constraint diagram is generated, and the corresponding controlled operation and maintenance event chain is transferred to the pending review process. This process is used to avoid inferring the target bearer object based solely on the scheduling plan, work order object name, and the same management area.

[0042] After the target load boundary passes integrity verification, a target load topology constraint graph is constructed. The actual landing point of the migrated load is used as the core node of the topology. Resource objects with direct load relationships with the core node are read from the configuration management database. Infrastructure resource objects with direct power supply, direct cooling, direct network path, and direct storage access relationships with the core node are included in the topology constraint scope. Infrastructure resource objects include the rack where the target host is located, the corresponding power supply circuit, the corresponding cooling area, the target network path, the target storage pool, and the target storage node. Topology nodes represent the target load and the infrastructure resource objects with direct load relationships with them. Topology edges represent power supply, cooling, network path, and storage access relationships. Resource objects whose observed changes on the target side are not included in the target load topology constraint graph do not participate in the target side load extraction and do not participate in the formation of migration risk debt fragments.

[0043] The target-side load observation records of each target resource object within the target load topology constraint diagram are read and used as input for the formation of target-side pressure status. The target-side load observation records include target host load, target rack temperature, target power supply circuit current, target network path occupancy, target storage pool load, target storage node read / write latency, and target storage node queue depth. The standardized deviation conversion logic in S1 is reused to convert the target-side load observation records into target-side standardized pressure deviation. This conversion only reuses risk direction consistency and dimensionless processing. The input objects are the target resource objects and their load indicators within the target load topology constraint diagram. No new formulas, weights, or models are introduced. During the conversion, both the increase in target-side pressure and the decrease in target-side margin are converted into dimensionless deviation evidence in the same risk direction.

[0044] Within the target load-bearing topology constraint diagram, construct a pre-load-bearing window and a post-load-bearing window. The pre-load-bearing window is located before the migrating load enters the target load-bearing object and is used to characterize the pressure state of the target load-bearing object before it takes on the migrating load. The post-load-bearing window is located after the migrating load has stabilized and is used to characterize the pressure state of the target load-bearing object after it takes on the migrating load. The pre-load-bearing window covers 2 to 5 effective sampling periods before the start of migration; the post-load-bearing window covers 3 to 8 effective sampling periods after the migrating load has stabilized and is used to form the post-load-bearing pressure state. The observable window in the subsequent time-series acceptance evidence can use the same time range as the target acceptance post-acceptance window, but it is used to verify the sequential relationship between the target-side acceptance pressure and the source-side standardized deviation fallback evidence. For temperature and cooling status resource objects, both the target acceptance pre-acceptance window and the target acceptance post-acceptance window cover at least one cabinet temperature response cycle. The window length is determined by the monitoring sampling configuration, migration load stability confirmation record, and historical operation and maintenance review record. When the migration load stability moment is missing, the target acceptance post-acceptance window is not formed, and the corresponding controlled operation and maintenance event chain is transferred to the pending review process.

[0045] When the effective sampling ratio in the window before or after the target acceptance is less than 80%, the corresponding resource object will not participate in the target-side pressure extraction. When the proportion of effective resource objects with acceptance relationships with source-side precursor indicators in the target-bearing topology constraint diagram is less than 50%, the corresponding controlled operation and maintenance event chain will be transferred to the pending review process. The effective sampling ratio and the proportion of effective resource objects are both dimensionless ratios, and their values ​​are based on the monitoring sampling configuration, the number of observable resource objects in the target-bearing topology constraint diagram, and historical operation and maintenance review records.

[0046] The pre-acceptance pressure state is formed based on the standardized pressure deviation of the target side within the pre-acceptance window, and the post-acceptance pressure state is formed based on the standardized pressure deviation of the target side within the post-acceptance window. The pre-acceptance pressure state is formed by the median level of the standardized pressure deviation of the target side for the same resource object and the same carrying capacity index within the pre-acceptance window; the post-acceptance pressure state is formed by the median level of the standardized pressure deviation of the target side for the same resource object and the same carrying capacity index within the post-acceptance window. The pre-acceptance pressure state and the post-acceptance pressure state are compared separately for the same resource object and the same carrying capacity index. When the pre-acceptance pressure state has reached the corresponding precursor access boundary, the continued change of the post-acceptance pressure state is not attributed solely to the migration load acceptance, and the target side pressure state of the corresponding resource object is marked as pending review. When the pre-acceptance pressure state has not reached the corresponding precursor access boundary, the excess portion of the post-acceptance pressure state relative to the pre-acceptance pressure state is taken as a candidate acceptance pressure change. This process is used to ensure that the target side pressure comparison occurs within the same resource object, the same index caliber, and the same topological constraints.

[0047] Historical segments of the same operating condition target side without migration load transfer coverage are extracted from the historical health baseline to form the upper bound of natural pressure increment. Historical segments of the same operating condition target side refer to available historical segments within the target bearer topology constraint diagram that contain the same resource object and the same bearer index under the same business load range, device role, environmental state, and maintenance state, and that have not experienced migration, rate limiting, business traffic switching, storage replica redistribution, cooling compensation, restart, isolation, or maintenance work order processing within the corresponding time period. Natural pressure change records in these historical segments are paired according to the same resource object, the same bearer index, and the same operating condition category to form the natural pressure change distribution of the corresponding resource object and bearer index. The upper quantile record of the pressure change distribution in the risk direction is taken to form the upper bound of natural pressure increment. For the same operating condition target side... For resource objects with at least 20 historical segments, the upper bound of the natural pressure increment is taken from the 90th to 97.5th percentile records. For network path and storage performance resource objects, the higher percentile records within this range are used. For temperature, power supply, and cooling status resource objects, calibration is performed by combining the cabinet temperature response cycle, power supply circuit sampling cycle, and the pre-alarm range in the operation and maintenance procedures. For resource objects with fewer than 20 but no fewer than 5 historical segments on the target side under the same operating condition, the historical segments of the same equipment role are used to form a temporary upper bound of the natural pressure increment, and a low confidence flag is added. When there are fewer than 5 historical segments on the target side under the same operating condition, and the historical segments of the same equipment role are still insufficient to form a temporary upper bound of the natural pressure increment, the target side pressure is not extracted, and the corresponding controlled operation and maintenance event chain is transferred to the pending review process.

[0048] The candidate load change is compared with the upper limit of the natural load increment. The portion exceeding the upper limit of the natural load increment is retained as the target-side load. The target-side load includes at least the target resource object identifier, the load index identifier, the target pre-load window, the target post-load window, the pre-load pressure status, the post-load pressure status, the upper limit of the natural load increment, the pressure portion exceeding the upper limit of the natural load increment, and the data availability status. The target-side load only indicates the load exceeding the natural fluctuation after the target load is migrated, and does not indicate that the target load has failed.

[0049] After the target-side bearing pressure is formed, a migration risk debt fragment is formed based on the temporal bearing evidence and standardized evidence bearing relationship of the source-side standardized deviation fall evidence. The temporal bearing evidence is used to verify whether the target-side bearing pressure occurs after the source-side standardized deviation fall evidence and whether it is within the observable window after the migration load stabilizes. The observable window starts from the moment the migration load stabilizes and covers 3 to 8 effective sampling periods. For resource objects such as temperature and cooling status, the observable window also covers at least 1 cabinet temperature response period. If the target-side bearing pressure is earlier than the source-side standardized deviation fall evidence or occurs outside the observable window, the corresponding target-side bearing pressure does not participate in the formation of the migration risk debt fragment.

[0050] The standardized evidence acceptance relationship is used to verify whether the meaning of the target-side acceptance pressure indicators can accept the source of the standardized deviation fallback evidence on the source side, and to verify whether the target-side acceptance pressure relative to the source-side standardized deviation fallback evidence meets the minimum acceptance requirements formed by historical migration review records. The source-side load precursor fallback is accepted by at least one of the target-side acceptance pressures: target host load increase, target power supply circuit current increase, and target cabinet temperature increase. The source-side storage read / write latency fallback is accepted by at least one of the target-side acceptance pressures: target storage pool load increase, target storage node read / write latency increase, and target storage node queue depth increase. The source-side network path occupancy and link latency fallback is accepted by at least one of the target-side acceptance pressures: target network path occupancy increase and target network path latency increase. Target-side pressure changes whose indicator meanings do not correspond are not included in the standardized evidence acceptance relationship.

[0051] Both the target-side load and the source-side standardized deviation fallback evidence are limited to standardized evidence quantities and compared to form a standardized evidence acceptance ratio, which is a dimensionless ratio from 0 to 1. When the ratio of the target-side load to the source-side standardized deviation fallback evidence exceeds 1, the standardized evidence acceptance ratio is counted as 1. When the source-side standardized deviation fallback evidence is missing, zero, or has a low confidence marker, no standardized evidence acceptance ratio is formed, and the corresponding controlled operation and maintenance event chain is transferred to the pending review process. Direct acceptance by a single target bearer object means that only one type of bearer resource object in the target bearer topology constraint diagram has an indicator meaning correspondence with the source-side precursor indicators. In this scenario, the corresponding target-side load is considered as the bearing... The proportion of receiving pressure relative to the standardized deviation and fallback evidence on the source side forms the standardized evidence acceptance ratio. Multi-target resource sharing acceptance refers to the situation where there are two or more types of carrying resource objects in the target carrying topology constraint diagram that have an index meaning correspondence with the source-side precursor indicators. In this scenario, the target-side receiving pressure is first divided into sub-items according to the type of carrying resource. Then, the sub-item that can explain the standardized deviation and fallback evidence on the source side is taken as the main acceptance evidence, and it is compared with the standardized deviation and fallback evidence on the source side to form the standardized evidence acceptance ratio. The standardized evidence acceptance ratio is used to represent the degree of acceptance of the source-side standardized deviation and fallback evidence by the target-side receiving pressure, and does not indicate that there is a conservation relationship between the source-side physical quantity and the target-side physical quantity.

[0052] The minimum acceptance requirement is formed from historical migration review records. For migration events involving multi-target resource sharing, the minimum acceptance requirement is 30% to 50%; for migration events directly undertaken by a single target, the minimum acceptance requirement is 50% to 80%. When there are no fewer than 20 historical migration review records, the minimum acceptance requirement for the corresponding scenario is formed based on the standardized evidence acceptance ratio distribution of confirmed migration risk debt fragments. The minimum acceptance requirement is taken from the lower quantile of this distribution, ranging from 10% to 25%. When there are fewer than 20 historical migration review records but no fewer than [number missing], the minimum acceptance requirement is determined based on the distribution of the acceptance ratio of the confirmed migration risk debt fragments. When there are 5 historical migration review records, a temporary minimum acceptance requirement is formed based on the historical migration review records of the same equipment role or the same business type, and a low confidence mark is attached. When there are fewer than 5 historical migration review records, no minimum acceptance requirement is formed, and the corresponding controlled operation and maintenance event chain is transferred to the pending review process. When any key link in the target acceptance window, the upper limit of the natural pressure increment, or the minimum acceptance requirement has a low confidence mark, a migration risk debt fragment is not directly generated, but a pending review fragment is generated. After the historical review records are supplemented, the migration risk debt fragment formation process is re-executed.

[0053] When the target-side pressure has temporal evidence of acceptance and satisfies the standardized evidence acceptance relationship, a migration risk debt fragment is generated. The migration risk debt fragment includes at least the controlled operation and maintenance event chain identifier, the near-failure fragment identifier that has been intervened and suppressed, the source-side controlled object identifier, the target carrier object boundary, the target carrier topology constraint diagram, the target-side pressure, temporal evidence of acceptance, standardized evidence acceptance relationship, minimum acceptance requirements, data availability status, and low confidence marker status. The migration risk debt fragment does not indicate that the target carrier object has failed, but rather that there is evidence of the source-side standardized deviation fallback evidence being accepted by the target-side pressure. This fragment serves as the input for verifying the target-side debt release status within the S4 extension confirmation window.

[0054] When the target-side load observation record is missing, the target-side load object boundary is incomplete, the target-side load topology constraint diagram cannot be generated, the sampling of the target-side pre-load window or target-side post-load window is insufficient, the upper bound of the natural pressure increment cannot be formed, the temporal load evidence is not valid, or the standardized evidence load relationship is not valid, no migration risk debt fragment is generated, and the corresponding controlled operation and maintenance event chain is transferred to the pending review process. If the target-side load pressure does not exceed the upper bound of the natural pressure increment, but the target-side load object boundary and the target-side load topology constraint diagram are complete, no debt candidate fragment is generated. The no debt candidate fragment is not used as the basis for confirming the target-side debt release status, and does not enter the main process of S4 target-side debt release status verification, but is only used as a historical review record.

[0055] At the end of S3, the system outputs migration risk debt segments, undiscovered debt candidate segments, and segments awaiting review. Among these, migration risk debt segments enter the delayed confirmation window of S4 as input for verifying the target-side debt release status. Undiscovered debt candidate segments and segments awaiting review are not used as the basis for confirming the target-side debt release status. Through the above processing, the actual landing point of the migration load, the target bearing topology constraint diagram, the target-side bearing pressure, the timing bearing evidence, and the standardized evidence bearing relationship are bound within the same controlled operation and maintenance event chain. This ensures that the near-failure segments that are suppressed by intervention on the source side continue to be subject to the target-side bearing pressure constraints, preventing the return of the source-side standardized deviation from being directly used for health baseline write-back permission judgment.

[0056] S4 is used to initiate the delayed confirmation window when the migration load stabilizes. Within the delayed confirmation window, the recurrence status on the source side and the debt release status on the target side are checked. Based on the check results, a health baseline write-back permission constraint result is generated to indicate the entry status of the controlled operation and maintenance event chain sample. This step uses the near-failure fragments that have been intervened and suppressed in S2, the migration risk debt fragments formed in S3, the recurrence status on the source side, the debt release status on the target side, and the data integrity status as the basis for write-back permission judgment, so as to restrict the controlled operation and maintenance event chain from entering the historical health baseline.

[0057] Read the controlled operation and maintenance event chain, the near-failure fragments suppressed by intervention output from S2, the migration risk debt fragments output from S3, the debt candidate fragments not found, and the fragments to be reviewed, and read the target carrier object boundary formed by S1. If S2 does not form a near-failure fragment suppressed by intervention, the health baseline write-back permission constraint judgment is not performed; if S3 outputs a fragment to be reviewed, a write-back result to be reviewed is generated, but a limited write-back permission is not generated. Before the delayed confirmation window is completed, the controlled operation and maintenance event chain does not participate in the update of the normal sample pool, does not reduce the precursor access boundary of the same type of abnormal precursor, and does not change the subsequent acceptance restriction status of the target carrier object.

[0058] A delayed confirmation window is established, starting from the moment the target bearer object reaches a stable migration load state. This stable migration load state is confirmed by resource scheduling landing point records, service instance running locations, actual storage replica landing points, actual network path switching records, and continuous valid sampling results on the target side. If the target bearer object has not reached a stable migration load state, the delayed confirmation window is not initiated, and a write-back result awaiting review is generated. The length of the delayed confirmation window is determined by the historical response lag of the bearer resources within the target bearer object's boundary, the service stabilization period, and the monitoring sampling period. When the target bearer object's boundary contains multiple types of bearer resources, if migration risk debt fragments exist, the resource object with the longest historical response lag among those with standardized evidence of connection to the migration risk debt fragments is considered. The length of the delayed confirmation window is determined by class. If no debt candidate fragment is found in the S3 output, the length of the delayed confirmation window is determined by the class of resource objects with longer historical response lags that actually undertake the migration load within the boundary of the target bearing object. The delayed confirmation window length for computing resources and network path objects is 15 to 60 minutes, the delayed confirmation window length for storage access objects is 30 to 120 minutes, and the delayed confirmation window length for temperature, power supply, and cooling status objects is 60 to 240 minutes. The window length value is derived from historical operation and maintenance review records, resource scheduling logs, monitoring sampling configurations, and data center operation and maintenance procedures. When historical samples are insufficient, a more conservative observation period in the operation and maintenance procedures is used to perform delayed confirmation and generate a write-back result to be reviewed. Limited write-back permissions are not supported.

[0059] Within the delayed confirmation window, the same type of abnormal precursors corresponding to the near-failure segments suppressed by intervention are reviewed in the source-side controlled objects to form a source-side recurrence state. The source-side controlled object identifier, precursor indicator set, standardized deviation, precursor admission boundary, and data availability status within the delayed confirmation window are read. Subsequently, based on the same source-side controlled object, the same precursor indicator, and the same operating condition category, it is determined whether the standardized deviation within the delayed confirmation window has again reached the precursor admission boundary. Combined with persistence conditions, a source-side recurrence state is formed. For server resources, network paths, and storage performance indicators, if the standardized deviation reaches the precursor admission boundary within 2 to 5 consecutive effective sampling periods, the same type of abnormal precursor is considered to have reappeared. For temperature and cooling status indicators, if the standardized deviation reaches the precursor admission boundary within 2 to 3 consecutive effective sampling points, and the duration is between 2 and 10 minutes, the same type of abnormal precursor is considered to have reappeared. If a similar anomaly reappears, the persistence condition follows the judgment criteria for the continuous deviation of the anomaly status in S1. If the effective sampling ratio of the source side is less than 80% within the delayed confirmation window, or the proportion of effective indicators in the anomaly indicator set is less than 50%, the source side recurrence status is not directly considered invalid. Instead, a write-back result pending review is generated. The invalidity of the source side recurrence status only indicates that the same type of anomaly anomaly has not been observed again within the delayed confirmation window, and is not a sufficient condition for limited write-back permission alone. When the source side business load is removed from the business load range used to form the anomaly admission result in S1 due to migration actions, and there are no comparable samples of the same working conditions in the historical health baseline, a low confidence flag is added to the source side recurrence status. When the target side debt release status is established, a write-back prohibition result is generated. When the target side debt release status is not established but the source side recurrence status is still low confidence, a write-back result pending review is generated, and limited write-back permission is not generated.

[0060] When S3 outputs a migration risk debt fragment, the target-side debt release status is checked within the delayed confirmation window. The target-bearing object boundary and the target-bearing topology constraint diagram are used as the object constraint range. Only target resource objects that fall within this range and have a standardized evidence connection with the source-side precursor indicators are read. Target-side observation changes that do not fall within the target-bearing topology constraint diagram are not included in the target-side debt release status judgment. The target-side bearing observation records within the delayed confirmation window are converted into target-side standardized pressure deviation according to the risk direction consistency and dimensionless logic in S1. Then, using the pre-acceptance pressure status in S3 as a comparison benchmark, a pressure increment is formed relative to the pre-acceptance pressure status of the target-side standardized pressure deviation within the delayed confirmation window. This pressure increment is compared with the upper bound of the natural pressure increment formed in S3. The continuous increment exceeding the upper bound of the natural pressure increment is taken as candidate evidence of the target-side debt release status.

[0061] The target persistence result is formed based on the candidate evidence for target-side debt release. This result is determined by the ratio between the number of valid samples meeting the conditions for target-side debt release candidate evidence within the extension confirmation window and the total number of valid samples for the corresponding target resource object within the extension confirmation window. This ratio is a dimensionless ratio from 0 to 1. For network path and storage performance objects, the target persistence ratio threshold is 30% to 50%; for computing resources, power supply, temperature, and cooling status objects, the threshold is 50% to 70%. The ratio threshold is formed by the sample distribution of target-side pressure continuously released as risky debt in historical migration review records. There are no fewer than 2 historical migration review records. When there are 0 historical migration review records, a threshold is formed using the 10% to 25% lower quantile of the persistent proportion distribution of the confirmed target-side debt release samples. When there are fewer than 20 but no fewer than 5 historical migration review records, a temporary threshold is formed using historical migration review records of the same device role or the same business type, with a low confidence flag attached. The temporary threshold is only used to form the write-back result to be reviewed and does not support limited write-back permission. When there are fewer than 5 historical migration review records, no target-side debt release status is formed, and a write-back result to be reviewed is generated. The target-side debt release status is considered established only when the candidate evidence for target-side debt release exceeds the upper bound of the natural pressure increment and the target persistence result meets the corresponding proportion threshold.

[0062] When S3 outputs no debt candidate fragments, it still reviews the target-side pressure status within the target bearing object boundary within the extension confirmation window. If the target-side pressure status does not exceed the upper limit of the natural pressure increment within the extension confirmation window, and the data integrity status meets the write-back admission requirements, the target-side debt release status is deemed not established. If the target-side pressure status exceeds the upper limit of the natural pressure increment within the extension confirmation window and meets the persistence requirements, the target-side debt release status is deemed established. If the target-side sampling integrity is insufficient, a write-back result pending review is generated. When S3 does not form a migration risk debt fragment but forms a debt candidate fragment that was not found, the write-back permission decision function does not read the evidence confidence of the migration risk debt fragment, but reads the data integrity status of the debt candidate fragment that was not found and the target-side debt release status. If the data integrity status of the debt candidate fragment that was not found does not meet the write-back admission requirements, a write-back result pending review is generated.

[0063] When migration risk debt fragments exist, the evidence confidence level for these fragments is determined by converting the integrity of the target bearing object boundary, the integrity of the target bearing topology constraint diagram, temporal inheritance evidence, standardized evidence inheritance relationship, upper bound source of natural pressure increment, target-side sampling integrity, and low-confidence marker state into rule items. Each rule item is recorded according to three states: satisfied, not satisfied, and low confidence. The evidence confidence level is then mapped based on historical write-back admission rules. The evidence confidence level is a dimensionless result ranging from 0 to 1, serving only as the basis for write-back admission judgment of the healthy baseline write-back permission constraint results. The write-back admission boundary of the evidence confidence level is formed by historical write-back admission rules. When both historically permitted write-back samples and historically prohibited write-back samples are no less than 20, the write-back admission is determined based on historically permitted write-back... The 10% to 25% lower quantile records of the evidence confidence distribution of the written samples form the write-back admission boundary, and the historical write-back prohibited samples are used for false admission verification. The historical permitted write-back samples and historical write-back prohibited samples are derived from the health baseline write-back permission constraint results of the historical period. When there are fewer than 20 but no fewer than 5 historical samples, a temporary write-back admission boundary is formed and a low confidence mark is attached. The temporary write-back admission boundary is only used to form the write-back result to be reviewed and does not support limited write-back permission. When there are fewer than 5 historical samples, no write-back admission boundary is formed, and a write-back result to be reviewed is generated. When any key input is missing, the low confidence mark is not replaced by subsequent complete evidence, or the target carrying topology constraint map is incomplete, the evidence confidence does not meet the write-back admission requirements, and a write-back result to be reviewed is generated.

[0064] The data integrity status of the delayed confirmation window is formed. The data integrity status includes the source-side data integrity status and the target-side data integrity status. The source-side data integrity status is formed by the effective sampling ratio, the proportion of effective indicators, and the timestamp continuity of the source-side precursor indicator set within the delayed confirmation window. The target-side data integrity status is formed by the effective sampling ratio, the proportion of effective resource objects, and the timestamp continuity of the target resource objects that have a relationship with the source-side precursor indicators within the target carrying topology constraint diagram. If the effective sampling ratio of the source side or the target side is less than 80%, the proportion of effective indicators or effective resource objects is less than 50%, or the timestamp is missing and the delayed confirmation window cannot be distinguished, the data integrity status does not meet the write-back admission requirements, and a write-back result pending review is generated. If any key item in the delayed confirmation window length, target persistence ratio threshold, evidence confidence write-back admission boundary, or data integrity status is formed by temporary samples and has a low confidence marker, a limited write-back permission is not generated.

[0065] The write-back permission decision function generates healthy baseline write-back permission constraints. This function reads the source-side recurrence status, target-side debt release status, and data integrity status of the delayed confirmation window. If a migration-risk debt fragment exists, it reads the evidence confidence level of that fragment; if no debt candidate fragment exists, it reads the data integrity status of that fragment. Then, it generates healthy baseline write-back permission constraints according to historical write-back admission rules. The write-back permission decision function uses an interpretable rule table. The inputs to the interpretable rule table are the source-side recurrence status, target-side debt release status, evidence confidence level, and data integrity status. The outputs are limited write-back permission, write-back prohibition, or write-back pending review. Historical write-back admission rules are formed from historical operation and maintenance review records, data center operation and maintenance procedures, and historical healthy baseline write-back permission constraints. If historical write-back admission rules are insufficient, write-back admission is not automatically relaxed; instead, a write-back pending review result is generated.

[0066] Limited write-back permission is generated when the source-side recurrence status is not established, the target-side debt release status is not established, the data integrity status of the delayed confirmation window meets the write-back admission requirements, and the evidence confidence of the migration risk debt fragment meets the write-back admission requirements when migration risk debt fragments exist, and the data integrity status of the debt candidate fragments that have not been found meets the write-back admission requirements when no debt candidate fragments exist, and there are no low-confidence markers that have not been replaced by subsequent complete evidence. Limited write-back permission allows the controlled operation and maintenance event chain to be written into the controlled successful sample record that retains the near-failure label. The controlled successful sample record is used to update the control action effectiveness record, the target carrier object acceptance record, and the historical operation and maintenance review record, but the corresponding event chain is not allowed to enter the ordinary normal sample pool, the precursor admission boundary of the same type of abnormal precursor is not allowed to be lowered, and the near-failure fragments and migration risk debt fragments that have been intervened and suppressed are not allowed to be covered.

[0067] When the source-side recurrence status is established, a write-back prohibition result is generated. The write-back prohibition result blocks the controlled operation and maintenance event chain from entering the normal sample pool, and writes the source-side controlled object identifier, similar abnormal precursor, source-side recurrence status, the near-failure segment identifier that has been intervened and suppressed, and the controlled operation and maintenance event chain identifier into the historical operation and maintenance review record. This process is used to mark that the source-side risk has not been stably suppressed, so as to avoid the subsequent health baseline from mistakenly taking this type of event as a successful control sample.

[0068] When the target-side debt release status is established, a write-back prohibition result is generated, and the boundary of the target carrier object is written into the subsequent acceptance restriction record. The subsequent acceptance restriction record includes at least the target carrier object boundary, the target carrier topology constraint diagram, the target-side debt release status, the target-side pressure status, the migration risk debt fragment identifier, similar anomaly precursors, and the data availability status. When a controlled operation and maintenance event corresponding to a similar anomaly precursor occurs, the subsequent acceptance restriction record is read, and the acceptance priority of the corresponding target carrier object or the same target resource domain is reduced. If the corresponding target carrier object still needs to participate in the migration acceptance, the subsequent postponement confirmation window is extended, and a write-back result pending review is generated. The subsequent acceptance restriction record serves as the constraint basis for subsequent migration acceptance and is not used as a conclusion of target carrier object failure.

[0069] When both the source-side recurrence status and the target-side debt release status are established simultaneously, a write-back prohibition result is generated and simultaneously written to the source-side recurrence record and subsequent acceptance restriction record, preventing the corresponding controlled operation and maintenance event chain from entering the controlled success sample record.

[0070] When the event chain attribution evidence is incomplete, the target carrying object boundary is incomplete, the migration risk debt fragment has a low confidence marker that has not been replaced by subsequent complete evidence, the sampling integrity of the delayed confirmation window is insufficient, the source-side recurrence status cannot be determined, or the target-side debt release status cannot be determined, a pending review write-back result is generated. The pending review write-back result blocks the controlled operation and maintenance event chain from entering the normal sample pool and blocks its automatic update of the health baseline and subsequent acceptance restriction records until the source-side recurrence evidence, target-side debt release evidence, and event chain attribution evidence are supplemented, and then the health baseline write-back permission constraint judgment is re-executed.

[0071] At the end of S4, the health baseline write-back permission constraint result is output. The health baseline write-back permission constraint result includes one of the following: limited write-back permission, write-back prohibition result, and write-back result pending review. Through the above processing, the near-failure segments that have been intervened and suppressed, migration risk debt segments, source-side recurrence status, target-side debt release status, and data integrity status are all included in the health baseline write-back permission judgment. This avoids risk segments that are truncated by controlled operation and maintenance actions being mistakenly written as ordinary normal samples, and ensures that subsequent migration is subject to the constraint of the target-side debt release status.

[0072] Example 2: like Figure 2 As shown, based on Embodiment 1, this embodiment also provides a data center fault prediction and health management system, including: a controlled operation and maintenance event chain construction module, which performs event anchor time calibration on the running event data formed by the controlled operation and maintenance process of the data center, performs standardized deviation conversion on the source side controlled object, and constructs a controlled operation and maintenance event chain based on the controlled operation and maintenance action attribution verification result after the source side controlled object shows a precursor continuous deviation state, and determines the boundary of the target carrying object; The near-failure segment identification module analyzes the effective interval of control actions within the controlled operation and maintenance event chain, compares the standardized deviation and fallback evidence before and after control intervention, eliminates natural recovery by using the natural fallback range under the same working conditions extracted from the historical health baseline, and identifies the operation segments whose fault evolution is truncated by control actions as near-failure segments that have been suppressed by intervention, and writes them into the critical risk sample pool. The migration risk debt tracking module generates a target bearing topology constraint diagram according to the boundary of the target bearing object. It extracts the target-side bearing pressure exceeding the upper limit of the natural pressure increment in the associated resource domain of the actual landing point of the migration load. Based on the target-side bearing pressure, it forms a migration risk debt fragment by combining the temporal bearing evidence and standardized evidence bearing relationship of the standardized deviation fallback evidence of the source side. The write-back permission constraint module initiates a delayed confirmation window when the migration load stabilizes. Within the delayed confirmation window, it verifies the recurrence status on the source side and the debt release status on the target side. Based on the verification results, it generates a health baseline write-back permission constraint result. The health baseline write-back permission constraint result is used to limit the status of the controlled operation and maintenance event chain entering the normal sample pool, the critical risk sample pool, the controlled successful sample record, and the record to be reviewed.

[0073] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0074] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0075] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for data center fault prediction and health management, characterized in that, include: Perform event anchor time calibration on the operational event data generated during the controlled operation and maintenance process of the data center, perform standardized deviation conversion on the source side controlled object, and construct the controlled operation and maintenance event chain based on the controlled operation and maintenance action attribution verification results after the source side controlled object shows a precursor continuous deviation state to determine the boundary of the target carrying object. Within the controlled operation and maintenance event chain, analyze the effective interval of control actions, compare the standardized deviation and fallback evidence before and after control intervention, eliminate natural recovery by using the natural fallback range under the same working conditions extracted from the historical health baseline, identify the operation segment whose fault evolution is truncated by control actions as the near failure segment suppressed by intervention, and write it into the critical risk sample pool. A target-bearing topology constraint diagram is generated according to the boundary of the target-bearing object. The target-side bearing pressure exceeding the upper limit of the natural pressure increment is extracted in the associated resource domain of the actual landing point of the migration load. Based on the target-side bearing pressure, the temporal bearing evidence and standardized evidence bearing relationship of the source-side standardized deviation fallback evidence are used to form a migration risk debt fragment. The delayed confirmation window is initiated when the migration load stabilizes. Within the delayed confirmation window, the recurrence status on the source side and the debt release status on the target side are checked. Based on the check results, a health baseline write-back permission constraint result is generated. The health baseline write-back permission constraint result is used to limit the status of the controlled operation and maintenance event chain entering the normal sample pool, the critical risk sample pool, the controlled successful sample record, and the record to be reviewed.

2. The data center fault prediction and health management method according to claim 1, characterized in that, The specific operation of performing standardized deviation conversion is as follows: extract available historical segments that are eligible for baseline updates from the historical health baseline. Available historical segments are running segments that have not been excluded by the health baseline write-back permission constraint results of the historical stage; use the business load range as the basis for working condition layering, and combine the equipment role, environmental status and maintenance status to merge the available historical segments into the same working condition to form the same working condition sample. The baseline center value and natural fluctuation scale of the corresponding indicators of the source-side controlled objects are extracted from the samples under the same working conditions. The deviation of the current observation value of the source-side controlled objects relative to the baseline center value is converted into a uniform value according to the direction of indicator risk. Then, the dimensionless value is processed by the natural fluctuation scale to form a standardized deviation. When there are insufficient samples of the same operating condition to form a benchmark for the same operating condition, a temporary benchmark is formed using available historical fragments of the same equipment role, and the precursor admission results formed based on the temporary benchmark are marked as low confidence markers.

3. The data center fault prediction and health management method according to claim 1, characterized in that, The specific operation of constructing a controlled operation and maintenance event chain is as follows: take the moment when the source-side controlled object shows a precursor to a continuous deviation state as the starting point for event retrieval, and read the controlled operation and maintenance action records on the common event timeline; first verify whether the object recorded in the controlled operation and maintenance action covers the source-side controlled object; after the object coverage is established, verify whether the action execution time is within the respondable period of the same type of action relative to the moment when the precursor to a continuous deviation state appears; after the response sequence is satisfied, verify whether there is an action interpretation relationship between the action type and the precursor indicator set; After the controlled operation and maintenance action is approved, the action completion record is read to confirm that the controlled operation and maintenance action has been executed. Then, the target bearer object that actually undertakes the migrated load is confirmed according to the resource scheduling landing point record. Based on the bearer relationship recorded in the configuration management database, the infrastructure bearer domain corresponding to the target bearer object is written into the boundary of the target bearer object.

4. The data center fault prediction and health management method according to claim 1, characterized in that, The specific steps for determining the effective range of control actions are as follows: Read the action type and action execution record in the controlled operation and maintenance event chain, and select the response reference time according to the action type; for migration actions, the migration completion time is used as the response reference time, for flow limiting actions, the flow limiting command execution time is used as the response reference time, and for cooling compensation actions, the cooling parameter effectiveness time is used as the response reference time; based on the indicator response lag and sampling period in the historical operation and maintenance review records of similar actions, form the effective range of control actions extending from the response reference time to the indicator response lag end time; When restart or isolation actions cause the source-side controlled object sampling to be interrupted, service instance to be rebuilt, or the monitored object to temporarily leave the normal operating state, the corresponding controlled operation and maintenance event chain will be transferred to the pending review process and will not enter the near failure suppression identification.

5. The data center fault prediction and health management method according to claim 1, characterized in that, The specific steps for identifying the near-failure segment suppressed by intervention are as follows: truncate the precursor holding window before the control action takes effect, and truncate the effective observation window within the control action takes effect; form the pre-control deviation level based on the standardized deviation within the precursor holding window, form the post-control deviation level based on the standardized deviation within the effective observation window, and then form evidence of standardized deviation decline based on the decrease in the pre-control deviation level relative to the post-control deviation level. Read out the same operating conditions without controlled actions from the historical health baseline, and extract the fallback range of similar abnormal precursors under natural operating conditions; When the evidence of standardization deviation falls back exceeds the fall back range under natural operating conditions, and the fall back direction can be explained by controlled operation and maintenance actions, the near failure suppression identification function forms an intervention suppression confidence level based on historical operation and maintenance review samples. Once the intervention suppression confidence level meets the admission requirements, the corresponding operational segment is written into the critical risk sample pool to form a near-failure segment that has been intervened and suppressed. When the historical operation and maintenance review samples are insufficient to form the intervention suppression confidence level, a low confidence marker is generated and the process is transferred to the review process.

6. The data center fault prediction and health management method according to claim 1, characterized in that, The specific steps for constructing the target bearer topology constraint graph are as follows: Using the boundary of the target bearer object as the topology entry point, read resource objects in the configuration management database that have a direct bearer relationship with the target bearer object; using the actual landing point of the migrated load as the topology core node, include infrastructure resource objects whose bearer relationships point to the topology core node within the topology constraint scope. Infrastructure resource objects include power supply resources, cooling resources, network path resources, and storage access resources; generate the target bearer topology constraint graph based on the bearer relationships; resource objects corresponding to changes observed on the target side that do not fall into the target bearer topology constraint graph are not included in the target side's load extraction and are not included in the formation of migration risk debt fragments.

7. The data center fault prediction and health management method according to claim 1, characterized in that, The specific operation for forming a migration risk debt fragment is as follows: Within the scope defined by the target bearing topology constraint diagram, read the target-side pressure state before the migration load enters the target bearing object, and then read the target-side pressure state after the migration load stabilizes; when the pressure state before acceptance has reached the corresponding precursor access boundary, transfer the target-side pressure state of the corresponding resource object to the pending review process, and do not separately regard the continued change of the pressure state after acceptance as the target-side acceptance pressure; when the pressure state before acceptance has not reached the corresponding precursor access boundary, compare the excess part of the pressure state after acceptance relative to the pressure state before acceptance with the upper limit of the natural pressure increment under the same working condition, and retain the part exceeding the upper limit of the natural pressure increment as the target-side acceptance pressure. Within the same controlled operation and maintenance event chain, verify whether the pressure received by the target side is later than the standardized deviation and fallback evidence on the source side, verify whether the indicator meaning of the pressure received by the target side can accept the source of the standardized deviation and fallback evidence on the source side, and then verify whether the pressure received by the target side relative to the standardized deviation and fallback evidence on the source side meets the minimum acceptance requirements formed by the historical migration review record; when all verification results are satisfied, generate a migration risk debt fragment.

8. The data center fault prediction and health management method according to claim 1, characterized in that, The specific operation for determining the delayed confirmation window is as follows: The starting point of the delayed confirmation window is the moment when the target load reaches a stable state during load migration; the historical response lag of the load resources within the target load boundary is read, and the length of the delayed confirmation window is formed by combining the business stability cycle and the monitoring sampling cycle; within the delayed confirmation window, the similar abnormal precursors corresponding to the near-failure segments suppressed by intervention in the source-side controlled objects are reviewed to form a source-side recurrence state; within this window, the target-side pressure state of the target resource objects within the target load topology constraint diagram that have a standardized evidence-based relationship with the source-side precursor indicators is read, and the target-side pressure state before the load migration enters the target load is taken as the pre-acceptance pressure state, and the target-side pressure increment within the delayed confirmation window is formed using the pre-acceptance pressure state as a comparison benchmark; pressure increments exceeding the upper limit of natural pressure increments and meeting the target continuity requirements are confirmed as target-side debt release states; when the sampling integrity within the delayed confirmation window is insufficient, limited write-back permission is not generated, and the write-back result to be reviewed is output.

9. A data center fault prediction and health management method according to claim 1, characterized in that, The specific operation for generating the health baseline write-back permission constraint result is as follows: the write-back permission decision function reads the data integrity status of the source side recurrence status, the target side debt release status, and the extension confirmation window. When there is a migration risk debt segment, it reads the evidence confidence of the migration risk debt segment. When there is no discovered debt candidate segment, it reads the data integrity status of the no discovered debt candidate segment. The health baseline write-back permission constraint result is generated according to the historical write-back admission rules. If the source-side recurrence status is not established, the target-side debt release status is not established, the data integrity status of the extension confirmation window meets the write-back admission requirements, and the evidence confidence of the migration risk debt fragment meets the write-back admission requirements when there is a migration risk debt fragment, and the data integrity status of the debt candidate fragment does not meet the write-back admission requirements when there is no discovered debt candidate fragment, and there are no low confidence markers that have not been replaced by subsequent complete evidence, a limited write-back permission is generated, so that the controlled operation and maintenance event chain is written to the controlled successful sample record that retains the near-failure label; When the recurrence status is established on the source side, a write-back prohibition result is generated to block the controlled operation and maintenance event chain from entering the normal sample pool; when the debt release status is established on the target side, a write-back prohibition result is generated to write the boundary of the target bearing object into the subsequent acceptance restriction record. When the evidence of event chain attribution is incomplete, the boundary of the target object is incomplete, there are debt fragments with migration risk and the confidence level of the evidence for the debt fragments with migration risk does not meet the write-back admission requirements, there are no debt candidate fragments found and the data integrity status of the no debt candidate fragments does not meet the write-back admission requirements, or the data integrity status of the delayed confirmation window does not meet the write-back admission requirements, a write-back result pending review is generated.

10. A data center fault prediction and health management system, employing the data center fault prediction and health management method as described in any one of claims 1-9, characterized in that, include: The controlled operation and maintenance event chain construction module performs event anchor time calibration on the running event data generated during the controlled operation and maintenance process of the data center, performs standardized deviation conversion on the source side controlled objects, and constructs the controlled operation and maintenance event chain based on the controlled operation and maintenance action attribution verification results after the source side controlled objects show a precursor continuous deviation state, and determines the boundary of the target carrying object. The near-failure segment identification module analyzes the effective interval of control actions within the controlled operation and maintenance event chain, compares the standardized deviation and fallback evidence before and after control intervention, eliminates natural recovery by using the natural fallback range under the same working conditions extracted from the historical health baseline, and identifies the operation segments whose fault evolution is truncated by control actions as near-failure segments that have been suppressed by intervention, and writes them into the critical risk sample pool. The migration risk debt tracking module generates a target bearing topology constraint diagram according to the boundary of the target bearing object. It extracts the target-side bearing pressure exceeding the upper limit of the natural pressure increment in the associated resource domain of the actual landing point of the migration load. Based on the target-side bearing pressure, it forms a migration risk debt fragment by combining the temporal bearing evidence and standardized evidence bearing relationship of the standardized deviation fallback evidence of the source side. The write-back permission constraint module initiates a delayed confirmation window when the migration load stabilizes. Within the delayed confirmation window, it verifies the recurrence status on the source side and the debt release status on the target side. Based on the verification results, it generates a health baseline write-back permission constraint result. The health baseline write-back permission constraint result is used to limit the status of the controlled operation and maintenance event chain entering the normal sample pool, the critical risk sample pool, the controlled successful sample record, and the record to be reviewed.