Intelligent operation and maintenance monitoring method and system of data center
By building an inter-layer data sharing model and outlier scanning in the data center, the problem of difficulty in tracing the root cause of faults in data center operation and maintenance is solved, enabling rapid fault location and automatic repair, and improving operation and maintenance efficiency and reliability.
Patent Information
- Application Number
- CN202511106326.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-08
AI Technical Summary
The lack of a systematic hierarchical division and collaboration mechanism in data center operation and maintenance makes it difficult to quickly trace the root cause of faults, resulting in slow response speed and low overall management efficiency.
By collecting data from the physical device layer, virtual resource layer, and application service layer of the data center, an inter-layer data sharing model is constructed, format conversion and correlation identification are performed, a real-time mapping table of cross-layer data interaction is obtained, anomaly scanning is performed, the location and propagation path of the fault point are determined, a fault location report is generated, and parameters are adjusted according to the repair priority sequence.
It enables rapid fault location and automatic repair, improving the operational efficiency and reliability of data centers.
Smart Images

Figure CN120602308B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data center operation and maintenance technology, and in particular to an intelligent operation and maintenance monitoring method and system for data centers. Background Technology
[0002] Whether it's cloud computing services or internal enterprise systems, the operational status of data centers must be precisely controlled to avoid huge economic losses and social impacts caused by failures.
[0003] Current data center operations and maintenance face numerous challenges. Due to a lack of clear hierarchical division, operations and monitoring are often fragmented, with data from devices and systems at different levels failing to form effective correlations, making it difficult to quickly trace the root cause of faults. This fragmentation further exacerbates the difficulty of inter-layer collaborative analysis, making it impossible to promptly link other layers for comprehensive judgment even if a problem is discovered at one level, ultimately leading to delays and inaccuracies in operations and maintenance decisions.
[0004] Therefore, current data center operation and maintenance methods lack a systematic hierarchical division and collaborative mechanism, resulting in difficulties in problem localization, slow response speed, and low overall management efficiency. Summary of the Invention
[0005] This invention provides an intelligent operation and maintenance monitoring method and system for data centers to achieve a reliable operation and maintenance system.
[0006] Firstly, in order to solve the above-mentioned technical problems, the present invention provides an intelligent operation and maintenance monitoring method for data centers, comprising:
[0007] Collect data from the physical device layer, virtual resource layer, and application service layer of the data center, obtain the raw dataset of the operating status of each layer, obtain a structured set of hierarchical monitoring information and perform format conversion, and obtain the correlation identifier between the data of each layer;
[0008] Based on the correlation identifier, an inter-layer data sharing model is constructed to obtain a real-time mapping table for cross-layer data interaction, thereby obtaining the basic dataset for inter-layer collaborative analysis.
[0009] Anomaly scanning is performed on the operating status of each layer based on the basic dataset. If the monitoring index of one layer exceeds the preset index threshold, the location information of the fault point is obtained. The real-time mapping table is called based on the location information, and the propagation path of the fault point is obtained based on the real-time mapping table.
[0010] A preliminary location report is generated based on the propagation path. The preliminary location report is then verified in multiple dimensions. If the verification is successful, a final location basis is generated, and a repair priority sequence is determined based on the final location basis.
[0011] Based on the repair priority sequence, the parameters of the level and equipment type corresponding to the fault point are adjusted, and the adjusted equipment is restored to the normal operating range.
[0012] Preferably, the data collection process involves acquiring data from the physical device layer, virtual resource layer, and application service layer of the data center to obtain raw datasets of the operational status of each layer, resulting in a structured set of hierarchical monitoring information, including:
[0013] Set corresponding monitoring indicators and indicator ranges for the physical device layer, the virtual resource layer and the application service layer respectively, and obtain the initial data of the running status from each layer to obtain the original dataset;
[0014] The original dataset is subjected to noise reduction and format unification processing to obtain a set of structured information;
[0015] If the monitoring indicators in the structured information set exceed the range of the indicators, they are marked as abnormal data;
[0016] The operational status of each layer is adjusted in real time based on the abnormal data to obtain a set of hierarchical monitoring information that meets the monitoring requirements of each layer.
[0017] Preferably, the step of performing format conversion to obtain the correlation identifier between data layers includes:
[0018] Based on the pre-established hierarchical information specifications, the format of the hierarchical monitoring information set is adjusted to obtain an initial set of data with unified data.
[0019] The data fields at different levels in the initial set are matched to determine the identification information of the inter-level connections;
[0020] If the identification information does not match the preset identification threshold, it is checked to obtain the corrected association identification.
[0021] Preferably, the step of constructing an inter-layer data sharing model based on the correlation identifier, obtaining a real-time mapping table for cross-layer data interaction, and obtaining the basic dataset for inter-layer collaborative analysis includes:
[0022] The data dependencies of each layer are obtained from the pre-established interaction mapping rules, and the data flow of each layer is obtained. After obtaining the initial cross-layer data interaction set, the data flow is compared layer by layer. If the data flow of one layer does not match the preset traffic threshold, the corresponding data flow is adjusted.
[0023] The adjusted inter-layer data streams are classified and processed to obtain the basis for inter-layer collaborative analysis after classification, resulting in a structured real-time mapping table.
[0024] The data dependencies of each layer are archived according to the structured real-time mapping table. It is then determined whether the archived data conforms to the inter-layer collaborative analysis criteria. If so, the final basic dataset is obtained.
[0025] Preferably, the step of scanning for outliers in the operating status of each layer based on the basic dataset, and obtaining the location information of the fault point if the monitoring indicator of one layer exceeds a preset indicator threshold, includes:
[0026] Obtain monitoring metrics for each layer, determine whether the monitoring metrics deviate from preset metric thresholds, and obtain a preliminary status monitoring dataset;
[0027] If the monitoring indicators of one layer exceed the threshold of the indicator, it is marked as abnormal and identified as a key object for anomaly scanning;
[0028] If the anomaly of the key object persists, an alarm signal is generated to determine the preliminary location information of the fault point. The preliminary location information is then confirmed to determine the final distribution of the fault point.
[0029] Preferably, the step of calling the real-time mapping table based on the location information and obtaining the propagation path of the fault point based on the real-time mapping table includes:
[0030] Obtain the hierarchical distribution information corresponding to the fault point from the real-time mapping table to obtain the preliminary list of associated layers;
[0031] By performing a deep query on the historical operating status of each layer through the preliminary list of related layers, the status monitoring records related to the fault point are obtained, and the starting point of the propagation path of the fault point is determined.
[0032] If the starting point of the propagation path deviates from the preset position, the interaction records of the inter-layer data are verified layer by layer to obtain the spread trajectory of the fault point in the hierarchical distribution and determine the direction of the propagation path.
[0033] Based on the propagation path direction, the status monitoring records are compared and analyzed with the historical operating status of the associated layer to obtain the distribution of the fault impact of the fault point among each level and determine the propagation path of the fault point.
[0034] Preferably, generating a preliminary location report based on the propagation path includes:
[0035] Obtain the hierarchical interaction information corresponding to the propagation path from the pre-established cross-layer data repository to determine the initial path direction;
[0036] Based on the initial path direction, the data is filtered layer by layer. If the state fluctuation exceeds the preset fluctuation threshold, the corresponding cross-layer data is marked as a key focus object.
[0037] Analyze the key targets of concern to determine the boundaries of their impact.
[0038] Each node of the propagation path is verified one by one according to the boundary to generate a preliminary fault location report.
[0039] Preferably, the step of performing multi-dimensional verification on the preliminary location report, and if the verification is successful, generating a final location basis and determining a repair priority sequence based on the final location basis, includes:
[0040] The verification dimensions of the preliminary positioning report are checked one by one according to the pre-established decision rules to obtain the status tracking data associated with the level. If the status tracking data meets the predefined conditions, it is marked as a valid basis and a preliminary verification result set is obtained.
[0041] The preliminary verification result set is compared with the decision rule to obtain the key matching items corresponding to the fault point and confirm the verification consistency distribution.
[0042] If the key matching item satisfies the predefined conditions, then the final location basis of the fault point is obtained;
[0043] The repair priority sequence is dynamically adjusted based on the final positioning criteria to determine the final repair priority sequence.
[0044] Preferably, the step of adjusting the parameters of the level and equipment type corresponding to the fault point according to the repair priority sequence, and determining that the adjusted equipment has been restored to the normal operating range, includes:
[0045] Based on the pre-established repair priority list, the corresponding automated script library is called to obtain the script execution instructions associated with the level, determine the specific order of script calls and target devices, and obtain the initial execution plan table;
[0046] The parameters of the target device are adjusted using the initial execution plan table to obtain real-time operating status data and determine whether it meets the expected range.
[0047] If the conditions are met, adjustment basis data is generated, a list of equipment parameters to be optimized is determined, the contents of the equipment parameter list are adjusted according to the adjustment basis data, the latest feedback data is obtained from the adjusted equipment operation log, and the adjusted equipment is restored to the normal operation range.
[0048] Secondly, the present invention also provides an intelligent operation and maintenance monitoring system for data centers, the system comprising:
[0049] The detection end is used to collect data from the physical device layer, virtual resource layer and application service layer of the data center, obtain the raw dataset of the operating status of each layer, obtain a structured set of hierarchical monitoring information and perform format conversion, and obtain the correlation identifier between the data of each layer.
[0050] The processing end is used to construct an inter-layer data sharing model based on the correlation identifier, obtain a real-time mapping table for cross-layer data interaction, and obtain a basic dataset for inter-layer collaborative analysis; perform anomaly scanning on the operating status of each layer based on the basic dataset; if the monitoring indicators of one layer exceed a preset indicator threshold, obtain the location information of the fault point; call the real-time mapping table based on the location information and obtain the propagation path of the fault point based on the real-time mapping table; generate a preliminary location report based on the propagation path; perform multi-dimensional verification on the preliminary location report; if the verification is successful, generate the final location basis and determine the repair priority sequence based on the final location basis.
[0051] The adjustment terminal is used to adjust the parameters of the level and equipment type corresponding to the fault point according to the repair priority sequence, and determine that the adjusted equipment is restored to the normal operating range.
[0052] Compared to existing technologies, this invention provides an intelligent operation and maintenance monitoring method and system for data centers. This method establishes a hierarchical monitoring framework, performing layered data collection and standardized processing across three levels: physical devices, virtual resources, and application services, and constructing an inter-layer data sharing model. A pre-defined fault detection algorithm scans the operational status of each layer for anomalies; once an anomaly is detected, an alarm is triggered, and the propagation path of the fault's impact is determined through data backtracking. Subsequently, an association rule mining algorithm is used to analyze the causal relationship between the anomaly and cross-layer data, combined with a decision support rule base for multi-dimensional verification, ultimately generating fault location conclusions and repair priorities. This invention can also automatically call response scripts to adjust parameters, achieving rapid fault location and automatic repair, effectively improving the operation and maintenance efficiency and reliability of data centers. Attached Figure Description
[0053] Figure 1 This is a flowchart of an intelligent operation and maintenance monitoring method for a data center provided by an embodiment of the present invention;
[0054] Figure 2 This is a flowchart of another intelligent operation and maintenance monitoring method for data centers provided by an embodiment of the present invention;
[0055] Figure 3 This is a flowchart of another intelligent operation and maintenance monitoring method for data centers provided by an embodiment of the present invention;
[0056] Figure 4 This is a flowchart of another intelligent operation and maintenance monitoring method for data centers provided by an embodiment of the present invention;
[0057] Figure 5 This is a flowchart of another intelligent operation and maintenance monitoring method for data centers provided by an embodiment of the present invention;
[0058] Figure 6 This is a flowchart of another intelligent operation and maintenance monitoring method for data centers provided by an embodiment of the present invention;
[0059] Figure 7 This is a flowchart of another intelligent operation and maintenance monitoring method for data centers provided by an embodiment of the present invention;
[0060] Figure 8 This is a flowchart of another intelligent operation and maintenance monitoring method for data centers provided by an embodiment of the present invention;
[0061] Figure 9 This is a flowchart of another intelligent operation and maintenance monitoring method for data centers provided by an embodiment of the present invention;
[0062] Figure 10 This is a schematic diagram of the structure of an intelligent operation and maintenance monitoring system for a data center provided in an embodiment of the present invention. Detailed Implementation
[0063] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0064] Reference Figure 1 The first embodiment of the present invention provides a flowchart of an intelligent operation and maintenance monitoring method for a data center, including the following steps:
[0065] S1: Collect data from the physical device layer, virtual resource layer, and application service layer of the data center, obtain the raw dataset of the operating status of each layer, obtain a structured set of hierarchical monitoring information and perform format conversion, and obtain the correlation identifier between the data of each layer.
[0066] S2, Construct an inter-layer data sharing model based on the correlation identifier, obtain a real-time mapping table for cross-layer data interaction, and obtain the basic dataset for inter-layer collaborative analysis;
[0067] S3, perform anomaly scanning on the operating status of each layer based on the basic dataset. If the monitoring index of one layer exceeds the preset index threshold, obtain the location information of the fault point, call the real-time mapping table based on the location information, and obtain the propagation path of the fault point based on the real-time mapping table.
[0068] S4. Generate a preliminary location report of the fault root cause based on the propagation path, perform multi-dimensional verification on the preliminary location report, and if the verification is successful, generate the final location basis and determine the repair priority sequence based on the final location basis.
[0069] S5, adjust the parameters of the level and equipment type corresponding to the fault point according to the repair priority sequence, and determine that the adjusted equipment is restored to the normal operating range.
[0070] In step S1, data from the physical device layer, virtual resource layer, and application service layer of the data center are collected to obtain the original dataset of the operating status of each layer, resulting in a structured set of hierarchical monitoring information and format conversion, and obtaining the correlation identifier between the data of each layer.
[0071] In the data center field, hierarchical monitoring information sets refer to the sum of monitoring data collected, integrated, and managed through a centralized platform, which categorizes equipment, systems, and environmental parameters at different levels within the data center according to logical or physical hierarchies. Its core purpose is to achieve full-stack visualization, fault warning, and performance optimization from infrastructure to application services.
[0072] A structured hierarchical monitoring information set refers to a unified data system formed by classifying, standardizing, and integrating various monitoring data within a data center according to a predefined logical or physical hierarchical model. Its core is to transform scattered monitoring information into standardized data that can be efficiently analyzed and automatically processed through a layered and modular architecture, thereby supporting operational decisions and intelligent management.
[0073] refer to Figure 2 S1 includes:
[0074] S11, set corresponding monitoring indicators and indicator ranges for the physical device layer, the virtual resource layer and the application service layer respectively, and obtain the initial data of the running status from each layer to obtain the original dataset;
[0075] S12, Denoise and format unification processing is performed on the original dataset to obtain a set of structured information;
[0076] S13, If the monitoring indicators in the structured information set exceed the range of the indicators, they are marked as abnormal data;
[0077] S14, Adjust the operating status of each layer in real time according to the abnormal data, and obtain a set of hierarchical monitoring information that meets the monitoring requirements of each layer.
[0078] For example, setting monitoring metrics and scopes for the three levels of physical devices, virtual resources, and application services is crucial to ensuring comprehensive system coverage. At the physical device level, hardware temperature, CPU utilization, and disk space can be monitored, with metrics such as temperature controlled between 20-50 degrees Celsius and CPU utilization not exceeding 80%. At the virtual resource level, focus on virtual machine memory usage and network bandwidth, with memory usage set between 30%-70%. At the application service level, monitor response time and error rate, with response time ideally less than 200 milliseconds. These metrics allow for the initial collection of operational status data, forming a raw dataset. This process lays the foundation for subsequent analysis, ensuring the comprehensiveness and accuracy of the data sources.
[0079] Specifically, data cleaning is a crucial step in processing the raw dataset. For example, if outliers exist in the collected physical device temperature data—such as a device displaying a temperature of 100 degrees Celsius, clearly exceeding the reasonable range—these outliers are marked as noise and removed. Simultaneously, the data format is standardized, such as adjusting the timestamp format to a uniform standard, facilitating subsequent analysis. This process not only improves data quality but also provides a reliable basis for generating structured information sets, helping to reduce the risk of misjudgment.
[0080] In one possible implementation, if the monitoring metrics in the structured information set exceed a preset range, such as virtual machine memory usage reaching 85%, the anomaly is flagged and categorized as a virtual resource-level issue. Detailed log recording helps quickly pinpoint the root cause of the problem; for example, recording the time of the anomaly and the specific virtual machine number provides precise guidance for subsequent adjustments. This step improves problem tracing efficiency and reduces manual troubleshooting costs.
[0081] If the application service response time reaches 300 milliseconds, exceeding the preset range, the script can automatically increase the number of service instances or optimize the load balancing strategy. After adjustment, the response time is reduced to 150 milliseconds, meeting the requirements. The adjusted status data is updated to the information set, ensuring the dynamic adaptability of the monitoring system. This automation significantly improves system stability and reduces the latency of manual intervention.
[0082] It's important to note that the updated information from hierarchical monitoring is not merely data aggregation, but also a basis for system optimization. Through continuous monitoring and adjustments, such as automatically reducing load when physical equipment temperatures are abnormal, hardware failures can be effectively prevented. This mechanism is invaluable in ensuring system continuity and provides managers with clear decision support, achieving closed-loop management across the entire chain from data acquisition to problem resolution.
[0083] refer to Figure 3 S1 also includes:
[0084] S15, According to the pre-established hierarchical information specifications, the format of the hierarchical monitoring information set is adjusted to obtain an initial set of unified data;
[0085] S16, Match the data fields of different levels in the initial set to determine the identification information of inter-level connections;
[0086] S17, If the identification information does not match the preset identification threshold, then a verification is performed to obtain the set of related identifications after verification.
[0087] For example, when processing heterogeneous formats in monitoring data, it's crucial to unify the format of raw records from the physical layer, virtual resources, and application services. Physical layer records might be stored as text logs, containing device numbers and operating parameters; virtual resource data might be in tabular form, recording memory and bandwidth information; and application service information might be structured database entries, including fields such as response time. Adjusting these different formats to a unified structured format—for example, changing all time fields to the standard "year-month-day hour:minute:second" format—ensures consistency in subsequent processing.
[0088] For example, the application of field mapping rules is a crucial step when dealing with the content in the initial set. Assuming the physical layer's device number field is "Device ID," the corresponding field at the virtual resource layer is "Host Number," and the application service layer's field is "Service Identifier," field mapping rules can link these three fields together to form unified identification information. This mapping helps clarify the dependencies between layers; for instance, a device ID may correspond to multiple host numbers, and each host number may be associated with a specific service identifier, thus revealing a complete hierarchical relationship.
[0089] Suppose the preset threshold requires a device ID to host ID association ratio of 1:3, but the actual mapping result shows 1:5, clearly exceeding the range. In this case, the mapping result is checked to see if there are duplicate associations or incorrect matches. For example, if a device ID is found to be incorrectly associated with an unrelated host ID, invalid data is removed after checking, generating an accurate set of association identifiers. This method improves data reliability and lays the foundation for subsequent integration.
[0090] For example, monitoring data from different layers can be categorized and stored using a set of associated identifiers. Physical-level data such as temperature and CPU utilization can be stored in a hardware monitoring library, virtual resource memory usage data in a resource allocation library, and application service response time data in a service quality library. Through the set of associated identifiers, a clear correspondence can be established between these libraries. For instance, a device ID can be used to find the corresponding host number, and then further traced to the specific service identifier, forming a complete hierarchical information association structure. This structured storage method facilitates rapid retrieval and analysis.
[0091] For example, in a real-world scenario, suppose a physical device displays a temperature of 45 degrees Celsius, corresponding to 60% virtual resource memory usage and an application service response time of 180 milliseconds. Through a hierarchical information association structure, it's possible to quickly determine whether the device and its related resources and services are in a normal state. If an anomaly is detected, administrators can quickly pinpoint the source of the problem based on the association structure, such as tracing from the device ID to a specific service identifier, and then take targeted measures. This approach significantly improves the efficiency of problem investigation.
[0092] For example, the construction of hierarchical information association structures can be extended to dynamic update scenarios. Suppose the memory usage of a virtual resource increases from 60% to 75%, the records in the repository can be updated in real time, and the set of associated identifiers can be adjusted synchronously to ensure the timeliness of the hierarchical structure. This dynamic adjustment capability helps maintain the accuracy of monitoring data and provides continuous support for system management.
[0093] In step S2, an inter-layer data sharing model is constructed based on the correlation identifier, a real-time mapping table for cross-layer data interaction is obtained, and the basic dataset for inter-layer collaborative analysis is obtained.
[0094] In the data center field, the Inter-layer Data Sharing Model refers to a standardized data interaction framework used to achieve efficient flow and collaborative utilization of monitoring, management, and operational data between different layers of the data center (such as infrastructure, network, server, and application layers). Its core objective is to break down data silos and support cross-layer comprehensive analysis, automated decision-making, and unified operation and maintenance through standardized interfaces and protocols.
[0095] refer to Figure 4 S2 includes:
[0096] S21. Obtain the data dependencies of each layer from the pre-established interaction mapping rules, and obtain the data flow of each layer. After obtaining the initial cross-layer data interaction set, compare it layer by layer. If the data flow of one layer does not match the preset traffic threshold, adjust the corresponding data flow.
[0097] S22, classify the adjusted inter-layer data stream, obtain the basis for inter-layer collaborative analysis after classification, and obtain a structured real-time mapping table;
[0098] S23, archive the data dependencies of each layer according to the structured real-time mapping table, determine whether the archived data conforms to the inter-layer collaborative analysis criteria, and if so, obtain the final basic dataset.
[0099] For example, in the field of data center operations and maintenance technology, interaction mapping rules refer to predefined logic or strategies used to describe the data interaction relationships (such as calls, transmissions, and dependencies) between different layers of the data center (e.g., physical layer, virtualization layer, application layer, and service layer). Data dependencies refer to the information transmission and influence chains formed between layers during operation. For instance, the operating status of physical devices directly affects the allocation of virtual resources, and the performance of virtual resources in turn affects the performance of application services. Key data flow information is extracted from each layer to form an initial set of cross-layer data interactions. Assuming the data flow from the physical device layer includes device runtime and temperature information, the virtual resource layer records CPU allocation ratios, and the application service layer includes request processing times, this information can be aggregated to form a set containing multiple layers of data flows.
[0100] For example, for the initial cross-layer data interaction set, one can start by comparing dynamic link relationships. Dynamic link relationships refer to the real-time changes in correlation between layers; for instance, a temperature increase in a device might lead to adjustments in virtual resource allocation. Suppose a preset threshold stipulates that the correlation between device temperature and CPU allocation ratio should not exceed 20%, but the actual data stream shows a 30% change in CPU allocation ratio when the temperature increases by 10 degrees Celsius, this would be flagged as an anomaly. The cause of the anomaly would then be analyzed, which could be due to data entry errors or outdated device status, leading to adjustments in the data stream to ensure the accuracy of the interaction mapping dataset.
[0101] For example, when classifying data flows between layers, a specific method could be to categorize the data flows according to their layer characteristics. Data flows from the physical device layer could be categorized under hardware status, those from the virtual resource layer under resource scheduling, and those from the application service layer under service performance. Assuming a device temperature of 50 degrees Celsius corresponds to a virtual resource CPU allocation of 70%, and an application service request processing time of 200 milliseconds, a structured real-time mapping table can be generated, clearly showing the correspondence between data from each layer. This classification method facilitates subsequent analysis of inter-layer collaboration.
[0102] In the field of data center operations and maintenance technology, collaborative analysis requirements refer to the standardized and interactive requirements for data integration, process collaboration, and tool interoperability in order to achieve efficient, cross-team, or cross-system operations and maintenance management. Its core objective is to break down information silos and improve the efficiency of fault diagnosis, performance optimization, and resource scheduling through multi-party collaboration.
[0103] For example, structured real-time mapping tables can be archived according to hierarchical dependencies. Suppose that after archiving, it is found that data from a certain device does not correspond effectively to virtual resource data; this could be due to missing data streams or broken links. In this case, it is necessary to determine whether the requirements for inter-layer collaborative analysis are met. If not, missing data needs to be supplemented or links corrected to ultimately form the basic dataset. This archiving method helps ensure data integrity and consistency, providing a reliable basis for subsequent analysis.
[0104] In step S3, anomaly scanning is performed on the operating status of each layer based on the basic dataset. If the monitoring index of one layer exceeds the preset index threshold, the location information of the fault point is obtained, and the real-time mapping table is called based on the location information to obtain the propagation path of the fault point.
[0105] refer to Figure 5 S3 includes:
[0106] S31, obtain the monitoring indicators of each layer, determine whether the monitoring indicators deviate from the preset indicator threshold, and obtain a preliminary status monitoring dataset;
[0107] S32, if the monitoring index of one layer exceeds the index threshold, it is marked as abnormal and identified as a key object for abnormal value scanning;
[0108] S33, if the abnormality of the key object persists, an alarm signal is generated, the preliminary location information of the fault point is determined, the preliminary location information is confirmed a second time, and the final distribution of the fault point is determined.
[0109] For example, in scenarios involving the construction of inter-layer data sharing models, real-time metric information can be obtained through pre-defined rules to monitor the operational status of the physical device layer, virtual resource layer, and application service layer. The monitoring rules are formulated based on the operational characteristics of each layer; for instance, the physical device layer focuses on hardware temperature and runtime, the virtual resource layer focuses on resource allocation ratios, and the application service layer focuses on service response time. Assuming the monitoring rules specify a device temperature threshold of 60 degrees Celsius, a resource allocation ratio fluctuation range of ±10%, and a response time not exceeding 300 milliseconds, these metrics will be captured in real time to form a preliminary status monitoring dataset.
[0110] For example, for a preliminary status monitoring dataset, a layer-by-layer verification approach can be adopted. Suppose a device in the physical device layer reaches a temperature of 65 degrees Celsius, exceeding the threshold by 5 degrees, it will be immediately marked as an anomaly and listed as a key target for scanning. Similarly, if the allocation ratio of a node in the virtual resource layer fluctuates by more than 15%, it will also be marked. This layer-by-layer verification method can quickly identify potential problem points, providing a clear direction for subsequent analysis.
[0111] For example, after identifying key targets in outlier scanning, in-depth analysis becomes crucial. Suppose a device's temperature consistently exceeds a threshold, accompanied by abnormal fluctuations in the virtual resource layer allocation ratio. The correlation between the two is analyzed to determine if hardware overheating is causing resource scheduling issues, and an alarm signal is generated. Initial location information might point to a server in the physical device layer, but the specific fault point still needs further confirmation. This in-depth analysis helps narrow down the problem scope and avoids blind troubleshooting.
[0112] For example, a secondary confirmation of initial location information can be achieved by comparing historical and real-time data. Suppose historical data indicates that the server did not experience resource fluctuations under similar temperature conditions, but current data shows anomalies. Further investigation would be conducted to determine if the problem is caused by sensor data deviation or adjustments to resource scheduling strategies. Ultimately, if the fault is confirmed to be in the server's cooling module, the hardware component can be precisely located. This secondary confirmation method improves the accuracy of location and provides a reliable basis for subsequent processing.
[0113] For example, throughout the process, collaborative work demonstrates the advantages of the inter-layer data sharing model. It ensures the comprehensiveness of indicator information, quickly filters anomalies, deeply explores the root causes of problems, and further pinpoints the problem location. Suppose that the above process reveals that a device's overheating problem is causing uneven virtual resource allocation, thus affecting application service response time. After timely location and handling of the fault point, the operational status of each layer can be restored to normal. This complete chain from monitoring to location not only improves the efficiency of problem discovery but also provides solid data support for inter-layer collaborative analysis.
[0114] refer to Figure 6 S3 also includes:
[0115] S34, Obtain the hierarchical distribution information corresponding to the fault point from the real-time mapping table to obtain the preliminary list of associated layers;
[0116] S35, By performing a deep query on the historical operating status of each layer through the preliminary list of associated layers, the status monitoring records related to the fault point are obtained, and the starting point of the propagation path of the fault point is determined.
[0117] S36, if there is a deviation between the starting point of the propagation path and the preset position, the interaction record of the inter-layer data is verified layer by layer to obtain the diffusion trajectory of the fault point in the hierarchical distribution and to determine the direction of the propagation path.
[0118] S37. Based on the propagation path direction, the status monitoring record is compared and analyzed with the historical operating status of the associated layer to obtain the fault impact distribution of the fault point among each level and determine the propagation path of the fault point.
[0119] For example, in inter-layer data analysis scenarios, to extract hierarchical distribution information of abnormal signals, relevant data is obtained from a real-time mapping table, focusing on the correlation between the physical device layer, virtual resource layer, and application service layer. Assuming the real-time mapping table records the data interaction relationships between each layer, the distribution information of potentially affected layers is filtered based on the initial alarm range of the abnormal signal. For instance, a temperature anomaly signal from a physical device layer might map to scheduling records in the virtual resource layer, and then be associated with response latency data in the application service layer. This layer-by-layer extraction method can quickly identify the initial impact range of abnormal signals, laying the foundation for subsequent analysis.
[0120] For example, after generating an initial list of related layers, in-depth querying becomes particularly important. Suppose the list shows that a server in the physical device layer is related to scheduling anomalies in the virtual resource layer. We would then trace back the operational status records of the past 24 hours to check if the server temperature was consistently high and if resource scheduling was frequently adjusted. If we find that the temperature reached 62 degrees Celsius multiple times in the past 6 hours, exceeding the threshold by 2 degrees, and the resource scheduling records show a fluctuation in allocation ratio of 12%, we can preliminarily determine that the origin of the anomaly signal may lie in the physical device layer. This kind of backtracking query helps to clarify the initial point of the fault's impact, providing a basis for subsequent path tracing.
[0121] For example, when there is a discrepancy between the starting point and location information of the propagation path, layer-by-layer verification becomes crucial. Suppose the initial location points to a specific server, but virtual resource layer data shows the anomaly has spread to multiple nodes. Through inter-layer data exchange records, the system analyzes how the abnormal signal propagated from the physical device layer to the virtual resource layer. For instance, if records show that after an abnormal server temperature, resource scheduling commands were frequently issued within 10 minutes, causing an imbalance in the allocation ratio across multiple nodes, the propagation trajectory can be traced, determining that the propagation direction was from lower to higher layers. This verification method clearly shows the propagation path of the abnormal signal, avoiding location errors.
[0122] For example, further analysis of the propagation path involves comparing status monitoring records with the operational status of related layers. Assuming the path indicates the anomaly spreads from the physical device layer to the application service layer, integrating data from each layer reveals that abnormal temperature at the physical device layer causes resource scheduling imbalance, consequently increasing the application service layer response time from 200 milliseconds to 350 milliseconds. Through comparative analysis, the distribution of the fault's impact across each layer can be determined, ultimately forming a complete picture of the propagation path. This integrated analysis comprehensively reveals the hierarchical distribution of the anomaly's impact, providing precise guidance for subsequent handling.
[0123] Throughout the process, collaborative work demonstrates the advantages of inter-layer data analysis. It involves defining the initial scope, tracing back to the starting point, clarifying the direction, and integrating and refining the overall picture. This collaborative approach effectively improves the comprehensiveness and accuracy of anomaly signal analysis, providing strong support for rapid response and handling of fault impacts.
[0124] In step S4, a preliminary location report of the fault root cause is generated based on the propagation path. The preliminary location report is then verified in multiple dimensions. If the verification is successful, a final location basis is generated, and a repair priority sequence is determined based on the final location basis.
[0125] refer to Figure 7 S4 includes:
[0126] S41, Obtain the hierarchical interaction information corresponding to the propagation path from the pre-established cross-layer data repository to determine the preliminary path direction;
[0127] S42, perform layer-by-layer filtering based on the initial path direction. If the state fluctuation exceeds the preset fluctuation threshold, mark the corresponding cross-layer data as a key focus object.
[0128] S43, Analyze the key objects of concern and determine the boundaries of the scope of influence;
[0129] S44, each node of the propagation path is verified one by one according to the boundary to generate a preliminary fault location report.
[0130] For example, in scenarios involving anomaly signal analysis based on cross-layer data repositories, related methods can be understood and implemented from multiple perspectives. Regarding the step of retrieving hierarchical interaction information from the repository, assuming the repository stores historical interaction data from the physical device layer, virtual resource layer, and application service layer, interaction records between relevant layers are filtered based on the alarm range of the anomaly signal. For instance, a power anomaly signal from a physical device layer might be associated with allocation adjustment records in the virtual resource layer; relevant data from the past 48 hours would be extracted, initially determining the path direction as diffusion from lower to upper layers. This extraction method helps to quickly focus on key information.
[0131] For example, the process of filtering historical records layer by layer can be understood as a hierarchical verification mechanism. Suppose that in the status fluctuation data, the allocation ratio of the virtual resource layer fluctuates by 15%, exceeding a preset threshold of 10%, then that layer will be marked as a key focus. This filtering method can effectively narrow down the scope of analysis and pinpoint potential problem nodes. For instance, if the response time fluctuation of the application service layer is found to be within acceptable limits, it can be temporarily excluded as a primary influencing factor.
[0132] For example, when analyzing the distribution of cross-layer data and abnormal signals, the focus is on uncovering key data distribution characteristics. Suppose the analysis reveals unstable voltage at a node in the physical device layer, highly correlated with frequent scheduling adjustments in the virtual resource layer. Further analysis would extract the node's operational logs to determine potential sources of the fault. For instance, if the logs show the voltage repeatedly falling below the standard value by 18% over the past 12 hours, combined with a 20% increase in scheduling adjustment frequency, the boundaries of the affected area can be preliminarily defined. This matching analysis provides precise guidance for subsequent troubleshooting.
[0133] For example, the step of verifying the propagation path layer by layer can be seen as a refined verification of the initial path direction. Assuming the boundary of the affected area points to a certain physical device node, historical records are used to analyze whether its state fluctuations are consistent with anomalies at other levels. For instance, if multiple nodes in the virtual resource layer experience scheduling imbalances within 10 minutes of discovering a voltage anomaly at this node, and the application service layer response time increases to 300 milliseconds, then this can be considered evidence to confirm it as the root cause of the fault. This layer-by-layer verification method ensures the reliability of the location. Through the above multi-faceted implementation methods, from data extraction to path verification, each step is closely linked, jointly supporting the complete process of anomaly signal analysis. The refined operation of each step effectively improves the targeting of the analysis, providing a solid foundation for subsequent fault handling.
[0134] refer to Figure 8 S4 also includes:
[0135] S45, the verification dimensions of the preliminary positioning report are checked one by one according to the pre-established decision rules to obtain the status tracking data associated with the level. If the status tracking data meets the predefined conditions, it is marked as a valid basis and a preliminary verification result set is obtained.
[0136] S46, compare the preliminary verification result set with the decision rule to obtain the key matching item corresponding to the fault point and confirm the verification consistency distribution;
[0137] S47, If the key matching item satisfies the predefined condition, then the final location basis of the fault point is obtained;
[0138] S48, The repair priority sequence is dynamically adjusted according to the final positioning criteria to determine the final repair priority sequence.
[0139] For example, in fault location scenarios based on cross-layer data analysis, a pre-established decision rule base can be understood in principle as a knowledge base containing multi-level association rules, used to guide the verification and location of fault root causes. Assuming this rule base stores state thresholds and association conditions for the physical device layer, virtual resource layer, and application service layer, it will extract state tracking data from relevant layers based on the preliminary location report. For instance, if the report indicates an abnormal temperature at a certain physical device node, the temperature data of that node over the past 24 hours will be checked. If it is found to have consistently exceeded a preset threshold of 45 degrees Celsius for more than 6 hours, it will be marked as valid evidence, forming a preliminary verification result set.
[0140] For example, data integration can be understood as a process of deep matching the preliminary verification result set with the criteria in the decision rule base. Suppose the verification result set shows an abnormal fluctuation of 12% in the resource allocation ratio of the virtual resource layer, while the threshold defined in the rule base is 10%. This item would be marked as a key matching item, and its correlation with abnormal temperature in the physical equipment layer would be analyzed in conjunction with multi-dimensional verification requirements to determine the final verification consistency distribution. This approach helps to focus on core issues.
[0141] For example, the key lies in combining verification consistency distribution and state tracking data to generate conclusions. Suppose a key matching item shows a high correlation between abnormal physical device layer temperature and fluctuations in virtual resource layer allocation, with the fluctuation times differing by only 2 hours. This would be determined to meet a predefined threshold condition, generating the final basis for locating the root cause of the fault. This layer-by-layer verification method effectively improves the accuracy of fault location.
[0142] For example, prioritization can be viewed as a process of dynamically adjusting the repair order. Suppose the generated data indicates that abnormal temperature at the physical device layer is the primary root cause of the failure. Based on its impact, the repair priority will be adjusted to address this node first, followed by issues related to the virtual resource layer, ultimately generating a repair priority list. This dynamic adjustment ensures efficient resource allocation and prioritizes resolving critical issues.
[0143] For example, in specific business scenarios, if an abnormal temperature in the physical equipment layer of a data center causes frequent adjustments to the virtual resource layer scheduling, cooling measures will be prioritized based on the location of the problem, while simultaneously monitoring resource allocation status. This processing logic, which addresses the core issue and its related impacts, effectively prevents the problem from escalating and ensures system stability.
[0144] It should be noted that the above-mentioned steps are closely linked, forming a complete closed loop from data verification to priority adjustment, ensuring the efficiency and reliability of fault handling.
[0145] In step S5, the parameters of the level and equipment type corresponding to the fault point are adjusted according to the repair priority sequence, and it is determined that the adjusted equipment is restored to the normal operating range.
[0146] refer to Figure 9 S5 includes:
[0147] S51, based on the pre-established repair priority list, call the corresponding automated script library, obtain the script execution instructions associated with the level, determine the specific order of script calls and the target device, and obtain the initial execution plan table;
[0148] S52, through the initial execution plan table, adjust the parameters of the target device, obtain real-time operating status data, and determine whether it meets the expected range;
[0149] S53, if the conditions are met, generate adjustment basis data, determine the list of equipment parameters to be optimized, adjust the contents of the equipment parameter list according to the adjustment basis data, obtain the latest feedback data from the adjusted equipment operation log, and determine that the adjusted equipment has been restored to the normal operation range.
[0150] For example, in a data center fault repair scenario, a pre-established repair priority list can be understood in principle as a task allocation framework based on hierarchy and device type. Assuming that temperature control at the physical device layer is listed as a top priority task in the list, the automation script library will filter out script instructions specifically for adjusting device temperature based on device type and hierarchy, clarifying that these instructions should be prioritized for specific server nodes, thus forming an initial execution plan. This approach ensures targeted task allocation.
[0151] For example, in script-based applications, this can be understood as a process of adjusting device parameters driven by a schedule. Suppose the initial execution schedule specifies increasing the speed of a server node's cooling fan. Real-time data on the adjusted fan's operating status is collected; for instance, if the speed reaches the preset 8000 RPM, it is compared with a preset threshold range of 7000-9000 RPM to determine if it meets the expected range, generating preliminary feedback results. This real-time data collection and comparison method helps to quickly identify deviations in parameter adjustments.
[0152] The key is to verify the initial feedback results against the recovery range requirements. For example, if the feedback shows that the fan speed is within the expected range, but the device temperature remains at 48 degrees Celsius, exceeding the 45-degree Celsius upper limit of the recovery range, adjustment data will be generated to determine if further optimization of the fan control strategy or inspection of other heat dissipation components is needed. This verification process can accurately pinpoint the parameters that are not meeting the standards, providing a basis for subsequent optimization.
[0153] This can be analyzed from the perspective of secondary adjustments. Assuming the adjustment data indicates that the fan speed needs further increase, and the coolant circulation status needs to be checked, the speed will be prioritized to 8500 rpm based on the equipment parameter list, and coolant flow monitoring will be initiated. The latest feedback data will be obtained from the adjusted operating log, such as the temperature dropping to 44 degrees Celsius, indicating it has entered the recovery range, ultimately forming a status monitoring conclusion. This secondary adjustment mechanism can effectively address situations where the initial adjustment does not meet expectations.
[0154] For example, the generation of status monitoring conclusions can be understood as a comprehensive judgment of the overall system recovery status. Assuming the latest feedback data indicates that the temperature has stabilized at 44 degrees Celsius, and other related parameters such as resource allocation ratios have recovered to 8% of the normal range, it would confirm that the system has entered the recovery phase. This multi-dimensional verification approach ensures the comprehensiveness and reliability of the repair process, reducing the risk of problem recurrence.
[0155] From the perspective of the overall process, the above-mentioned steps, from script invocation to the formation of the final conclusion, constitute a closed-loop mechanism. Assuming that in a data center, after abnormal temperatures at the physical equipment layer are prioritized for handling, the scheduling stability of the virtual resource layer also improves, thereby enhancing the overall system efficiency. This closed-loop mechanism ensures the efficient progress of repair tasks while providing data support and reference for subsequent maintenance.
[0156] In summary, this invention discloses an intelligent operation and maintenance monitoring method for data centers. This method establishes a hierarchical monitoring framework, performing layered data collection and standardized processing across three levels: physical devices, virtual resources, and application services, and constructing an inter-layer data sharing model. A pre-defined fault detection algorithm scans the operational status of each layer for anomalies, triggering alarms upon detection and determining the propagation path of the fault's impact through data backtracking. Subsequently, an association rule mining algorithm is used to analyze the causal relationship between the anomaly and cross-layer data, combined with a decision support rule base for multi-dimensional verification, ultimately generating fault location conclusions and repair priorities. This invention can also automatically invoke response scripts to adjust parameters, achieving rapid fault location and automatic repair, effectively improving the operation and maintenance efficiency and reliability of data centers.
[0157] like Figure 10 As shown, this embodiment of the invention provides an intelligent operation and maintenance monitoring system for data centers, including:
[0158] The detection end is used to collect data from the physical device layer, virtual resource layer and application service layer of the data center, obtain the raw dataset of the operating status of each layer, obtain a structured set of hierarchical monitoring information and perform format conversion, and obtain the correlation identifier between the data of each layer.
[0159] The processing end is used to construct an inter-layer data sharing model based on the correlation identifier, obtain a real-time mapping table for cross-layer data interaction, and obtain a basic dataset for inter-layer collaborative analysis; perform anomaly scanning on the operating status of each layer based on the basic dataset; if the monitoring indicators of one layer exceed a preset indicator threshold, obtain the location information of the fault point; call the real-time mapping table based on the location information and obtain the propagation path of the fault point based on the real-time mapping table; generate a preliminary location report based on the propagation path; perform multi-dimensional verification on the preliminary location report; if the verification is successful, generate the final location basis and determine the repair priority sequence based on the final location basis.
[0160] The adjustment terminal is used to adjust the parameters of the level and equipment type corresponding to the fault point according to the repair priority sequence, and determine that the adjusted equipment is restored to the normal operating range.
[0161] It should be noted that the intelligent operation and maintenance monitoring system for data centers provided in this embodiment of the invention is used to execute all the process steps of the intelligent operation and maintenance monitoring method for data centers in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0162] This invention also provides a terminal device. The terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a data center intelligent operation and maintenance monitoring program. When the processor executes the computer program, it implements the steps described in the above-described embodiments of the intelligent operation and maintenance monitoring methods for data centers, for example... Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above system embodiments.
[0163] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.
[0164] The terminal device may be a desktop computer, laptop, handheld computer, or smart tablet, etc. The terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above components are merely examples of terminal devices and do not constitute a limitation on the terminal device. It may include more or fewer components than described above, or a combination of certain components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.
[0165] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.
[0166] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0167] Wherein, if the modules / units integrated in the terminal device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or system capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0168] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0169] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for intelligent operation and maintenance monitoring of a data center, characterized in that, include: Collect data from the physical device layer, virtual resource layer, and application service layer of the data center, obtain the raw dataset of the operating status of each layer, obtain a structured set of hierarchical monitoring information and perform format conversion, and obtain the correlation identifier between the data of each layer; Based on the correlation identifier, an inter-layer data sharing model is constructed to obtain a real-time mapping table for cross-layer data interaction, thereby obtaining the basic dataset for inter-layer collaborative analysis. Anomaly scanning is performed on the operating status of each layer based on the basic dataset. If the monitoring index of one layer exceeds the preset index threshold, the location information of the fault point is obtained. The real-time mapping table is called based on the location information, and the propagation path of the fault point is obtained based on the real-time mapping table. A preliminary location report is generated based on the propagation path. The preliminary location report is then verified in multiple dimensions. If the verification is successful, a final location basis is generated, and a repair priority sequence is determined based on the final location basis. Based on the repair priority sequence, the parameters of the level and equipment type corresponding to the fault point are adjusted, and the adjusted equipment is restored to the normal operating range. The step of constructing an inter-layer data sharing model based on the correlation identifier, obtaining a real-time mapping table for cross-layer data interaction, and obtaining the basic dataset for inter-layer collaborative analysis includes: The data dependencies of each layer are obtained from the pre-established interaction mapping rules, and the data flow of each layer is obtained. After obtaining the initial cross-layer data interaction set, the data flow is compared layer by layer. If the data flow of one layer does not match the preset traffic threshold, the corresponding data flow is adjusted. The adjusted inter-layer data streams are classified and processed to obtain the basis for inter-layer collaborative analysis after classification, resulting in a structured real-time mapping table. The data dependencies of each layer are archived according to the structured real-time mapping table. It is then determined whether the archived data conforms to the inter-layer collaborative analysis criteria. If so, the final basic dataset is obtained. Among them, the data dependency relationship refers to the information transmission and influence chain formed between each layer during operation; the dynamic link relationship refers to the real-time changing correlation between layers; the data flow of the physical device layer is classified into the hardware status category, the virtual resource layer into the resource scheduling category, and the application service layer into the service performance category. The process of performing multi-dimensional verification on the preliminary location report, and generating a final location basis and determining a repair priority sequence based on the final location basis if the verification is successful, includes: The verification dimensions of the preliminary positioning report are checked one by one according to the pre-established decision rules to obtain the status tracking data associated with the level. If the status tracking data meets the predefined conditions, it is marked as a valid basis and a preliminary verification result set is obtained. The preliminary verification result set is compared with the decision rule to obtain the key matching items corresponding to the fault point and determine the verification consistency distribution. If the key matching item satisfies the predefined conditions, then the final location basis of the fault point is obtained; The repair priority sequence is dynamically adjusted based on the final positioning criteria to determine the final repair priority sequence.
2. The intelligent operation and maintenance monitoring method for data centers according to claim 1, characterized in that, The process involves collecting data from the physical device layer, virtual resource layer, and application service layer of the data center to obtain raw datasets of the operational status of each layer, resulting in a structured set of hierarchical monitoring information, including: Set corresponding monitoring indicators and indicator ranges for the physical device layer, the virtual resource layer and the application service layer respectively, and obtain the initial data of the running status from each layer to obtain the original dataset; The original dataset is subjected to noise reduction and format unification processing to obtain a set of structured information; If the monitoring indicators in the structured information set exceed the range of the indicators, they are marked as abnormal data; The operational status of each layer is adjusted in real time based on the abnormal data to obtain a set of hierarchical monitoring information that meets the monitoring requirements of each layer.
3. The intelligent operation and maintenance monitoring method for data centers according to claim 1, characterized in that, The process of format conversion and obtaining the correlation identifiers between data layers includes: Based on the pre-established hierarchical information specifications, the format of the hierarchical monitoring information set is adjusted to obtain an initial set of data with unified data. The data fields at different levels in the initial set are matched to determine the identification information of the inter-level connections; If the identification information does not match the preset identification threshold, it is checked to obtain the corrected association identification.
4. The intelligent operation and maintenance monitoring method for data centers according to claim 1, characterized in that, In the process of scanning for outliers in the operating status of each layer based on the basic dataset, if the monitoring indicator of one layer exceeds a preset indicator threshold, the location information of the fault point is obtained, including: Obtain monitoring metrics for each layer, determine whether the monitoring metrics deviate from preset metric thresholds, and obtain a preliminary status monitoring dataset; If the monitoring indicators of one layer exceed the threshold of the indicator, it is marked as abnormal and identified as a key object for anomaly scanning; If the anomaly of the key object persists, an alarm signal is generated to determine the preliminary location information of the fault point. The preliminary location information is then confirmed to determine the final distribution of the fault point.
5. The intelligent operation and maintenance monitoring method for data centers according to any one of claims 1-4, characterized in that, The step of calling the real-time mapping table based on the location information and obtaining the propagation path of the fault point based on the real-time mapping table includes: Obtain the hierarchical distribution information corresponding to the fault point from the real-time mapping table to obtain the preliminary list of associated layers; By performing a deep query on the historical operating status of each layer through the preliminary list of related layers, the status monitoring records related to the fault point are obtained, and the starting point of the propagation path of the fault point is determined. If the starting point of the propagation path deviates from the preset position, the interaction records of the inter-layer data are verified layer by layer to obtain the spread trajectory of the fault point in the hierarchical distribution and determine the direction of the propagation path. Based on the propagation path direction, the status monitoring records are compared and analyzed with the historical operating status of the associated layer to obtain the distribution of the fault impact of the fault point among each level and determine the propagation path of the fault point.
6. The intelligent operation and maintenance monitoring method for data centers according to any one of claims 1-4, characterized in that, The step of generating a preliminary location report based on the propagation path includes: Obtain the hierarchical interaction information corresponding to the propagation path from the pre-established cross-layer data repository to determine the initial path direction; Based on the initial path direction, the data is filtered layer by layer. If the state fluctuation exceeds the preset fluctuation threshold, the corresponding cross-layer data is marked as a key focus object. Analyze the key targets of concern to determine the boundaries of their impact. Each node of the propagation path is verified one by one according to the boundary to generate a preliminary fault location report.
7. The intelligent operation and maintenance monitoring method for data centers according to any one of claims 1-4, characterized in that, The step of adjusting parameters for the level and equipment type corresponding to the fault point according to the repair priority sequence, and determining that the adjusted equipment has been restored to the normal operating range, includes: Based on the pre-established repair priority list, the corresponding automated script library is called to obtain the script execution instructions associated with the level, determine the specific order of script calls and target devices, and obtain the initial execution plan table; The parameters of the target device are adjusted using the initial execution plan table to obtain real-time operating status data and determine whether it meets the expected range. If the conditions are met, adjustment basis data is generated, a list of equipment parameters to be optimized is determined, the contents of the equipment parameter list are adjusted according to the adjustment basis data, the latest feedback data is obtained from the adjusted equipment operation log, and the adjusted equipment is restored to the normal operation range.
8. An intelligent operation and maintenance monitoring system for a data center, characterized in that, The system is used to implement the intelligent operation and maintenance monitoring method for data centers as described in any one of claims 1-4, the system comprising: The detection end is used to collect data from the physical device layer, virtual resource layer and application service layer of the data center, obtain the raw dataset of the operating status of each layer, obtain a structured set of hierarchical monitoring information and perform format conversion, and obtain the correlation identifier between the data of each layer. The processing end is used to construct an inter-layer data sharing model based on the correlation identifier, obtain a real-time mapping table for cross-layer data interaction, and obtain a basic dataset for inter-layer collaborative analysis; perform anomaly scanning on the operating status of each layer based on the basic dataset; if the monitoring indicators of one layer exceed a preset indicator threshold, obtain the location information of the fault point; call the real-time mapping table based on the location information and obtain the propagation path of the fault point based on the real-time mapping table; generate a preliminary location report based on the propagation path; perform multi-dimensional verification on the preliminary location report; if the verification is successful, generate the final location basis and determine the repair priority sequence based on the final location basis. The adjustment terminal is used to adjust the parameters of the level and equipment type corresponding to the fault point according to the repair priority sequence, and determine that the adjusted equipment is restored to the normal operating range.
Citation Information
Patent Citations
Cloud monitoring service operation and maintenance dynamic optimization system and method based on AI intelligent agent
CN120223501A
Charging service real-time monitoring method and device based on service probe
CN120416093A