Intelligent operation and maintenance management method, system and storage medium based on data center
By introducing automated monitoring architecture and self-healing architecture in the data center, the monitoring and failure prediction problems of the data center hardware and software environment are solved, and the operation and maintenance efficiency and fault prediction accuracy of the data center are improved.
Patent Information
- Application Number
- CN202510600034.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing technology cannot conduct distribution monitoring and analysis of the hardware and software environment of the data center, cannot reasonably deploy data centers, and cannot make failure predictions.
The intelligent operation and maintenance management system based on data center is adopted, including an automated monitoring architecture, self-healing architecture and data operation fault prediction module. Through the automated monitoring architecture, the hardware environment is monitored in real time, the self-healing architecture monitors the software environment, and fault prediction is carried out in combination with historical operation data.
Real-time monitoring and fault prediction of the data center hardware environment are realized, improving the operation and maintenance efficiency of the data center and the accuracy of fault prediction, and reducing the risk of operation failures.
Smart Images

Figure CN120104432B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data center operation and maintenance technology, and specifically to an intelligent operation and maintenance management method, system and storage medium based on a data center. Background Art
[0002] In the digital age, data has become one of an enterprise's most critical assets. As the core hub for data storage, processing, and distribution, data centers are of undeniable importance to modern enterprises. They support various core business systems of enterprises, covering everything from daily office automation and customer relationship management to complex e-commerce transactions, supply chain management, and big data analysis.
[0003] Patent number CN112818403A discloses a container data center operation and maintenance system, including: the image construction module is used to obtain the data stream in the data pool and construct the image text according to the data stream; the container publishing module is used to obtain the image text and construct the data container according to the image text; the service authorization module is used to receive a service request sent from any user terminal, and generate a verification request corresponding to the data container according to the service request.
[0004] However, in existing technologies, data centers are unable to conduct distributed monitoring and analysis of their hardware and software environments through architectural settings, and are unable to compare and analyze the environment with actual data center operations based on algorithms, making it impossible to reasonably deploy data centers and complete data center operation and maintenance management. In addition, data center failures cannot be predicted.
[0005] In view of the above technical defects, a solution is now proposed. Summary of the Invention
[0006] The purpose of the present invention is to solve the above-mentioned problems and to propose an intelligent operation and maintenance management method, system and storage medium based on a data center.
[0007] The objectives of the present invention can be achieved through the following technical solutions: an intelligent operation and maintenance management system based on a data center, including an automated monitoring architecture, a self-healing architecture and a data operation fault prediction module; the automated monitoring architecture is constructed by four underlying logics of monitoring, discovery, deployment and troubleshooting, and is used to monitor the hardware environment of the data center in real time; when the automated monitoring architecture is in operation, the self-healing architecture operates synchronously, and the self-healing architecture is divided into three underlying logics of resource distribution, parameter collection and fault detection, and the function of the self-healing architecture is to monitor the software environment of the data center; the hardware environment and software environment are controlled and automatically monitored through the automated monitoring architecture and the self-healing architecture, and the automated monitoring architecture and the self-healing architecture complete the underlying logic operation, and after the fault is detected and eliminated, the data operation fault prediction module compares the historical operation of the data center with the real-time operation, and predicts operation performance faults through operation data comparison.
[0008] As a preferred embodiment of the present invention, the underlying logical operation process of the automated monitoring architecture is as follows: the operating environment of the data center is monitored, and the operating environment is the hardware environment; the environmental parameters of the hardware environment are divided into operation-affecting parameters and operation-unaffecting parameters; the real-time environmental parameters are monitored and compared with the set parameter thresholds, and if the real-time monitored environmental parameters exceed the set parameter thresholds, the current period is marked as a parameter growth period; conversely, if the real-time monitored environmental parameters do not exceed the set parameter thresholds, the current period is marked as a parameter stability period.
[0009] As a preferred embodiment of the present invention, the control buffer time of the corresponding environmental parameters before and after the hardware's own operation influencing parameters fluctuate is obtained; before the fluctuation occurs, the control buffer time exceeds the set time threshold, then the environmental parameter control is abnormal, and the environmental control system of the data center is redeployed, specifically deployed to control the environmental parameters in the hardware's surrounding operating environment to ensure that the environmental parameters of the hardware's surrounding operating environment are within the set range; after the fluctuation occurs, the control buffer time exceeds the set time threshold, then the hardware's own operation influencing parameters are controlled abnormally, and the hardware operation operation that causes the operation influencing parameters to fluctuate is deployed and controlled to reduce the real-time workload of the corresponding operation operation, and give priority to the current operation when the same type of operation influencing parameters in the environmental parameters are at a valley value, and the amount of hardware operation influencing parameters generated is controlled, that is, the deviation of the amount of operation influencing parameters generated in the same workload operation period within different commissioning stages of the same type of hardware is within the set deviation range.
[0010] As a preferred embodiment of the present invention, after the parameters are discovered and deployed, the environmental parameters are analyzed as a whole, the environmental parameters corresponding to the parameter growth period and the parameter stability period are collected, and the growth period parameter curve and the stability period parameter curve are constructed respectively; the growth period parameter curve and the stability period parameter curve are analyzed, and the growth period parameter curve exceeds the risk level line, and the corresponding angle is formed, wherein the risk level line is represented by the horizontal line corresponding to the value of the environmental parameter set threshold, and is in the same coordinate system as the curve; if the corresponding angle formed by the curve exceeds the set angle threshold, it is inferred that the environmental parameter control during the parameter growth period is abnormal, that is, the environmental parameter cannot be quickly reduced after exceeding the threshold and the parameter control speed is slow when the environmental parameter is on a downward trend, then the control performance of the environmental control system cannot match the current hardware operation scenario, and the operation process of the environmental control system in the corresponding operation period is detected and debugged to perform troubleshooting; if the corresponding angle formed by the curve does not exceed the set angle threshold, it is inferred that the environmental parameter control during the parameter growth period is normal.
[0011] As a preferred embodiment of the present invention, the critical values corresponding to the growth trend and the decrease trend of the curve in the parameter curve during the stable period are collected, and the reciprocating floating frequency of the curve corresponding to the critical value is collected in real time when the critical value is continuously updated during the operation stage; if the reciprocating floating frequency of the curve corresponding to the critical value exceeds the floating frequency threshold, it is inferred that the fluctuation of the environmental parameter cannot be effectively controlled, so the control logic of the environmental control system is rectified, and in addition to the control logic for the floating control of the environmental parameter value, a control logic is added, that is, the speed threshold is set for the growth rate of the environmental parameter, and after exceeding the speed threshold, the increase in the environmental parameter value under two unit times of running at the current speed is limited. If the limited increase value exceeds the limited increase value, the environmental control system is executed when the growth rate of the environmental parameter exceeds the threshold. Otherwise, the environmental control system will not run temporarily to continuously monitor the growth rate of the environmental parameter.
[0012] As a preferred embodiment of the present invention, the underlying logic operation process of the self-healing architecture is as follows: resources are distributed according to the processing resources of the data center, network terminals with the same distributed resources are monitored, the distributed resource occupancy rates when the demand data volume of the same type of network terminals is met and the distributed resource occupancy rates when the demand data volume is not met are collected, and the occupancy rates are analyzed: when the demand data volume is met, if the distributed resource occupancy rate deviation of the same type of network terminals does not exceed the occupancy rate deviation threshold, it is inferred that the distributed resource distribution of the same type of network terminals is reasonable, and resource distribution is performed according to the current resource distribution logic; otherwise, if the distributed resource occupancy rate deviation of the same type of network terminals exceeds the occupancy rate deviation threshold, it is inferred that the distributed resource distribution of the same type of network terminals is unreasonable, and resource distribution is adjusted using the current resource distribution logic as the adjustment starting point; if the distributed resource occupancy rates of the same type of network terminals are all within the resource occupancy rate range, it is inferred that the distributed resource utilization of the same type of network terminals is qualified, and data processing is performed with the current distributed resources; otherwise, if the distributed resource occupancy rates of the same type of network terminals are not all within the resource occupancy rate range, it is inferred that the distributed resource utilization of the same type of network terminals is unqualified, and the distributed resources are redistributed.
[0013] As a preferred embodiment of the present invention, when the required data volume is not met, if the distributed resource occupancy rates of the same type of network terminals are all lower than the set occupancy rate threshold, it is inferred that the distributed resources are not suitable for the data processing of the current network terminal; the current network terminals of the same type are finely divided, that is, the data processing requirements are synchronized and the processing requirements are used as the type division standard to ensure the executability of data processing; if the distributed resource occupancy rates of the same type of network terminals are all higher than the set occupancy rate threshold, it is inferred that the distributed resources are suitable for the data processing of the current network terminal.
[0014] As a preferred embodiment of the present invention, the process of the data operation fault prediction module is as follows: obtain the fault period of the data center; obtain the hardware operation parameters corresponding to the distributed resources of the data center according to the fault period; obtain the floating trend of the hardware operation parameters when the fault occurs according to the fault period and mark it as a fault trend, collect the excess increase span of the parameter growth rate of the hardware operation parameters in the current operation process that is in the fault trend relative to the peak value of the parameter growth rate of the historical period, and at the same time obtain the increase span value of the parameter growth span of the hardware operation parameters in the current operation period that is not in the fault trend relative to the historical period; if the excess increase span exceeds the increase span threshold, or the increase span value exceeds the increase span threshold, then it is inferred that the fault prediction of the operation process is high risk; if the excess increase span does not exceed the increase span threshold, and the increase span value does not exceed the increase span threshold, then it is inferred that the fault prediction of the operation process is low risk.
[0015] As a preferred embodiment of the present invention, an intelligent operation and maintenance management method based on a data center is as follows: automated monitoring, based on the four underlying logic constructions of monitoring, discovery, deployment and troubleshooting, to perform real-time monitoring of the hardware environment of the data center; self-healing fault detection, based on the three underlying logics of resource distribution, parameter collection and fault detection, to monitor the software environment of the data center; automated monitoring and self-healing fault detection complete the underlying logic operation, and after the fault is detected and eliminated, the historical operation of the data center is compared with the real-time operation, and operation performance faults are predicted through operation data comparison.
[0016] An intelligent operation and maintenance management method based on a data center, such as the intelligent operation and maintenance management system based on a data center as described in any one of the above implementation methods.
[0017] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0018] Compared with the prior art, the beneficial effects of the present invention are: 1. In the present invention, the hardware environment of the data center is monitored in real time, and the hardware environment of the current data center is detected through an automated monitoring architecture to ensure the operational safety of the data center and prevent the operating environment from affecting the actual data processing efficiency when the data center is operating. The influence of the characteristics of the data center itself is not considered, and only external influences are controlled, and timely discovery and redeployment are carried out to eliminate faults; through software environment monitoring combined with the analysis of the operating parameters of the data center, the operating risk faults of the data center itself are detected and controlled in a timely manner to reduce the operating fault efficiency of the data center. The current operating fault is different from the troubleshooting of the automated monitoring architecture. The current operating fault is a network data operation fault of the data center, while the troubleshooting of the automated monitoring architecture is a hardware environment fault of the data center; 2. In the present invention, the historical operation and real-time operation of the data center are compared, and the operating performance fault is predicted by comparing the operating data to complete the fault prediction of the data center. The accuracy of fault prediction is improved based on the operating data, and the operation and maintenance efficiency of the data center is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0020] Figure 1 This is the overall architecture diagram of the present invention; Figure 2 Flowchart of the automated monitoring architecture of the present invention. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0023] See also Figure 1 As shown, the intelligent operation and maintenance management system based on the data center includes an automated monitoring architecture, a self-healing architecture and a data operation fault prediction module; it should be explained that the automated monitoring architecture, the self-healing architecture and the data operation fault prediction module cooperate with each other to perform intelligent monitoring of the data center, so as to improve the efficiency of operation and maintenance management; the thresholds used in this application are obtained by personnel in this field by collecting corresponding data and averaging the historical operation and maintenance process of the data center, and are artificially set parameter values; the automated monitoring architecture is constructed by four underlying logics of monitoring, discovery, deployment and troubleshooting, which is used to monitor the hardware environment of the data center in real time, and detect the hardware environment of the current data center through the automated monitoring architecture to ensure the operational safety of the data center and avoid the operating environment affecting the actual data processing efficiency when the data center is in operation. It does not consider the influence of the characteristics of the data center itself, only controls external influences, and promptly discovers and redeploys them to troubleshoot faults; please refer to Figure 2As shown, the operating environment of the data center is monitored, and the operating environment is the hardware environment; the environmental parameters of the hardware environment are divided into operation-affecting parameters and operation-unaffecting parameters, wherein the operation-affecting parameters are parameters that will cause the environmental parameters to rise when the data center hardware is operating, such as temperature; the operation-unaffecting parameters are parameters that will not cause the environmental parameters to rise when the data center hardware is operating, such as dust concentration; the real-time environmental parameters are monitored and compared with the set parameter thresholds, and if the real-time monitored environmental parameters exceed the set parameter thresholds, the current period is marked as a parameter growth period; otherwise, if the real-time monitored environmental parameters do not exceed the set parameter thresholds, the current period is marked as a parameter stability period; at the same time, the hardware operation-affecting parameters are obtained. The control buffer time of the corresponding environmental parameters before and after the fluctuation occurs, where the control buffer time is expressed as the buffer time for the environmental parameters to be controlled within the set threshold range after the fluctuation occurs; before the fluctuation occurs, the control buffer time exceeds the set time threshold, then the environmental parameter control is abnormal, and the environmental control system of the data center is redeployed, where the environmental control system is represented by the environmental control system, such as the cooling system, ventilation system, etc.; the specific deployment is to control the environmental parameters in the operating environment around the hardware to ensure that the environmental parameters in the operating environment around the hardware are within the set range; after the fluctuation occurs, the control buffer time exceeds the set time threshold, then the hardware itself operates abnormally, and the hardware operation that causes the operation-affecting parameter fluctuation is deployed and controlled. Control, reduce the real-time workload of the corresponding operation, and give priority to the current operation when the same type of operation-affecting parameters in the environmental parameters are at valley values, and control the amount of hardware operation-affecting parameters generated, that is, the deviation of the amount of operation-affecting parameters generated in the same workload operation period within the same type of hardware at different stages of commissioning is within the set deviation range; after the parameters are discovered and deployed, an overall analysis of the environmental parameters is performed, the corresponding environmental parameters of the parameter growth period and the parameter stability period are collected, and the growth period parameter curve and the stability period parameter curve are constructed respectively; the growth period parameter curve and the stability period parameter curve are analyzed, and the angle formed by the curve of the collected growth period parameter curve exceeding the risk level line, where the risk level line is represented by the environmental parameter setting The horizontal line of the corresponding value of the fixed threshold is in the same coordinate system as the curve; the angle formed by the curve is expressed as the part of the curve above the horizontal line of the red line value, corresponding to the angle formed by the different trends of the curve when the curve shows a downward trend after reaching the peak value; if the corresponding angle formed by the curve exceeds the set angle threshold, it is inferred that the environmental parameter control during the parameter growth period is abnormal, that is, the environmental parameter cannot be quickly reduced after exceeding the threshold and the parameter control speed is slow when the environmental parameter shows a downward trend, then the control performance of the environmental control system cannot match the current hardware operation scenario, and the operation process of the environmental control system in the corresponding operation period will be tested and debugged to perform troubleshooting; if the corresponding angle formed by the curve does not exceed the set angle threshold, it is inferred that the environmental parameter control during the parameter growth period is normal;Collect the critical values corresponding to the growth trend and the reduction trend of the curve in the stable period parameter curve, and collect the reciprocating floating frequency of the curve corresponding to the critical value in real time when the critical value is continuously updated in the operation stage; if the reciprocating floating frequency of the curve corresponding to the critical value exceeds the floating frequency threshold, it is inferred that the control of the environmental parameters in the stable power transmission is inefficient, that is, the fluctuation of the environmental parameters cannot be effectively controlled, but the control is carried out after the environmental parameters show an increasing trend. Although it does not exceed the risk level line, when the fluctuation of the environmental parameters has an impact on the hardware operation, the control logic of the environmental control system is rectified. In addition to the control logic of the environmental parameter numerical floating control, an additional control logic is added, that is, the speed threshold of the environmental parameter growth rate is set. After exceeding the speed threshold, the increase in the environmental parameter value under two unit times of running at the current speed is limited. If the limited increase value is exceeded, the environmental control system will be executed when the growth rate of the environmental parameter exceeds the threshold. Otherwise, the environmental control system will not run temporarily and will continue to monitor the growth rate of the environmental parameter. When the automated monitoring architecture is running, the self-healing architecture will run synchronously. Among them, the self-healing architecture is divided into three underlying logics: resource distribution, parameter collection, and fault detection. The role of the self-healing architecture is to monitor the software environment of the data center. Through software environment monitoring combined with the analysis of the operating parameters of the data center, the risk faults of the data center itself are detected and controlled in a timely manner to reduce the operation of the data center. Fault transfer efficiency: the current operation fault is different from the troubleshooting of the automated monitoring architecture. The current operation fault is the network data operation fault of the data center, while the troubleshooting of the automated monitoring architecture is the hardware environment fault of the data center. The resources are distributed according to the processing resources of the data center, where the processing resources are represented by data processing memory, algorithm computing power pool and other resources. Parameters are collected according to the operation process of the network terminals covered by the data center. The network terminals with the same distribution resources are monitored, and the distribution resource occupancy rate of the same type of network terminals when the demand data volume is met and the distribution resource occupancy rate when the demand data volume is not met are collected, and the occupancy rate is analyzed: when the demand data volume is met, if the same type of network terminals are If the distribution resource occupancy rate deviation of the same type of network terminal does not exceed the occupancy rate deviation threshold, it is inferred that the distribution resource distribution of the same type of network terminal is reasonable, and resource distribution is performed according to the current resource distribution logic; on the contrary, if the distribution resource occupancy rate deviation of the same type of network terminal exceeds the occupancy rate deviation threshold, it is inferred that the distribution resource distribution of the same type of network terminal is unreasonable, and the current resource distribution logic is used as the adjustment starting point to adjust the resource distribution, reduce the occupancy rate deviation, and prevent excess data processing resources; if the distribution resource occupancy rates of the same type of network terminals are all within the resource occupancy rate range, it is inferred that the distribution resource utilization of the same type of network terminals is qualified, and data processing is performed according to the current distribution resources;On the contrary, if the distribution resource occupancy rates of the same type of network terminals are not all within the resource occupancy rate range, it is inferred that the distribution resource utilization of the same type of network terminals is unqualified, and the distribution resources are redistributed, and the resource distribution is carried out in combination with the data complexity of the network terminals, where the data complexity is expressed as the number of non-same type data, the frequency of continuous data processing, etc.; when the required data volume is not met, if the distribution resource occupancy rates of the same type of network terminals are all lower than the set occupancy rate threshold, it is inferred that the distribution resources are not suitable for the data processing of the current network terminals, such as inconsistent data transmission protocols, incompatible network hardware equipment, and inconsistent data transmission formats; for the current The network terminals of the same type are divided into detailed categories, that is, the data processing requirements are synchronized and the processing requirements are used as the type classification standard to ensure the executability of data processing; if the distributed resource occupancy rates of the network terminals of the same type are all higher than the set occupancy rate threshold, it is inferred that the distributed resources are suitable for the data processing of the current network terminals; the same type of network terminals are represented by the deviation of the required data volume within the set range, and the resource demand deviation is not large; the required data volume is represented by the amount of data that the network terminal can produce, and the data needs to be processed through distributed resources; the hardware environment and software environment are controlled and automatically monitored through the automated monitoring architecture and self-healing architecture to reduce manual labor Risks are checked to minimize the impact of external interference on the data center hardware storage location; the automated monitoring architecture and self-healing architecture complete the underlying logic operation, and after fault detection and elimination, completion signals are generated and sent to the data operation fault prediction module; the data operation fault prediction module is used to compare the historical operation of the data center with the real-time operation, and predict the operation performance fault by comparing the operation data to complete the data center fault prediction. Based on the operation data, the accuracy of the fault prediction is improved, and the operation and maintenance efficiency of the data center is improved; according to the historical operation process, the failure period of the data center is obtained, where the failure period includes all failure periods, such as equipment failure caused by hardware environment anomalies and equipment processing failure caused by software environment anomalies; according to the failure period, the hardware operation parameters corresponding to the distributed resources of the data center are obtained, such as hard disk read and write speed, temperature and other parameters; according to the failure period, the floating trend of the hardware operation parameters when the fault occurs is obtained and marked as a failure trend, and the excess increase span of the parameter growth rate of the hardware operation parameters in the current operation process that is within the failure trend relative to the peak value of the parameter growth rate in the historical period is collected, and at the same time, the increase span value of the average parameter growth span of the hardware operation parameters in the current operation period that is not within the failure trend relative to the historical period is obtained;And compare the excess increase span of the parameter growth rate of the hardware operating parameters in the fault trend during the current operation process relative to the peak value of the parameter growth rate in the historical period, and the increase span value of the average parameter growth span of the hardware operating parameters in the current operation period that is not in the fault trend relative to the historical period with the increase span threshold and the increase span threshold respectively: if the excess increase span of the parameter growth rate of the hardware operating parameters in the fault trend during the current operation process relative to the peak value of the parameter growth rate in the historical period exceeds the increase span threshold, or the increase span value of the average parameter growth span of the hardware operating parameters in the current operation period that is not in the fault trend relative to the historical period exceeds the increase span threshold , then the fault prediction for the operation process is inferred to be high risk, a high-risk signal is generated, and the high-risk signal is sent to the administrator terminal. After receiving the signal, the administrator terminal monitors the operating parameters of the data center and promptly controls the trend growth. If the hardware operating parameters in the current operation process are within the fault trend, the parameter growth rate relative to the peak value of the parameter growth rate in the historical period does not exceed the increase span threshold, and the hardware operating parameters in the current operation period are not within the fault trend, and the average parameter growth span relative to the increase span value in the historical period does not exceed the increase span threshold, then the fault prediction for the operation process is inferred to be low risk, a low-risk signal is generated, and the low-risk signal is sent to the administrator terminal.
[0024] The intelligent operation and maintenance management method based on the data center is as follows: automated monitoring, based on the four underlying logics of monitoring, discovery, deployment and troubleshooting, to conduct real-time monitoring of the hardware environment of the data center; self-healing fault detection, based on the three underlying logics of resource distribution, parameter collection and fault detection, to monitor the software environment of the data center; automated monitoring and self-healing fault detection complete the underlying logic operation, and after the fault is detected and eliminated, the historical operation of the data center is compared with the real-time operation, and the operation performance fault is predicted through the comparison of operation data.
[0025] Additionally, a storage medium stores a computer program thereon, which, when executed by a processor, implements the above-mentioned intelligent operation and maintenance management method based on a data center.
[0026] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods.
[0027] Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory, non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory, and volatile memory may include random access memory (RAM) or external cache memory.
[0028] By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0029] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0030] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. Intelligent operation and maintenance management system based on data center, characterized by: It includes an automated monitoring architecture, a self-healing architecture, and a data operation fault prediction module. The automated monitoring architecture is built with four underlying logics: monitoring, discovery, deployment, and troubleshooting, and is used to monitor the data center's hardware environment in real time. When the automated monitoring architecture is operating, the self-healing architecture operates synchronously. The self-healing architecture is divided into three underlying logics: resource distribution, parameter collection, and fault detection. The function of the self-healing architecture is to monitor the software environment of the data center. The hardware environment and software environment are controlled and automatically monitored through the automated monitoring architecture and the self-healing architecture. After the automated monitoring architecture and the self-healing architecture complete the underlying logic operation and the fault is detected and eliminated, the data operation fault prediction module compares the historical operation of the data center with the real-time operation, and predicts the operation performance fault through the operation data comparison. The underlying logic operation process of the automated monitoring architecture is as follows: the operating environment of the data center is monitored, and the operating environment is the hardware environment; the environmental parameters of the hardware environment are divided into operation-affecting parameters and operation-unaffecting parameters; the real-time environmental parameters are monitored and compared with the set parameter thresholds, and if the real-time monitored environmental parameters exceed the set parameter thresholds, the current period is marked as a parameter growth period; conversely, if the real-time monitored environmental parameters do not exceed the set parameter thresholds, the current period is marked as a parameter stability period; Obtain the control buffer time of the corresponding environmental parameters before and after the hardware's own operation-affecting parameters fluctuate; before the fluctuation occurs, if the control buffer time exceeds the set time threshold, the environmental parameter control is abnormal, and the data center's environmental control system is redeployed, specifically deployed to control the environmental parameters in the operating environment around the hardware to ensure that the environmental parameters of the operating environment around the hardware are within the set range; after the fluctuation occurs, if the control buffer time exceeds the set time threshold, the hardware's own operation-affecting parameter control is abnormal, and the hardware operation operation that causes the operation-affecting parameter fluctuation is deployed and controlled to reduce the real-time workload of the corresponding operation operation, and give priority to the current operation when the same type of operation-affecting parameters in the environmental parameters are at a valley value, and control the amount of hardware operation-affecting parameters generated, that is, the deviation of the amount of operation-affecting parameters generated in the same workload operation period within the same type of hardware at different commissioning stages is within the set deviation range; after the parameters are discovered and deployed, conduct an overall analysis of the environmental parameters, collect the environmental parameters corresponding to the parameter growth period and the parameter stability period, and construct the growth period parameter curve and the stability period parameter curve respectively; Perform curve analysis on the parameter curves during the growth period and the parameter curves during the stable period, and collect the angles formed by the curves where the parameter curves during the growth period exceed the risk level line. The risk level line is the horizontal line corresponding to the value of the set threshold of the environmental parameter and is in the same coordinate system as the curve. If the angle formed by the curves exceeds the set angle threshold, it is inferred that the environmental parameter control during the parameter growth period is abnormal, that is, the environmental parameter cannot be quickly reduced after exceeding the threshold, and the parameter control speed is slow when the environmental parameter is on a downward trend. In this case, the control performance of the environmental control system cannot match the current hardware operation scenario. The operation process of the environmental control system in the corresponding operation period will be tested and debugged to perform troubleshooting. If the angle formed by the curve does not exceed the set angle threshold, it is inferred that the environmental parameter control is normal during the parameter growth period; Collect the critical values corresponding to the growth trend and the decrease trend of the curve in the parameter curve during the stable period, and collect the reciprocating floating frequency of the curve corresponding to the critical value in real time when the critical value is continuously updated during the operation stage; if the reciprocating floating frequency of the curve corresponding to the critical value exceeds the floating frequency threshold, it is inferred that the fluctuation of the environmental parameters cannot be effectively controlled, so the control logic of the environmental control system is rectified. In addition to the control logic for the floating control of the environmental parameter values, an additional control logic is added, that is, the speed threshold is set for the growth rate of the environmental parameter, and after exceeding the speed threshold, the increase in the environmental parameter value under two unit times of running at the current speed is limited. If the limited increase value exceeds the limited increase, the environmental control system will be executed when the growth rate of the environmental parameter exceeds the threshold. Otherwise, the environmental control system will not run temporarily to continuously monitor the growth rate of the environmental parameter.
2. The intelligent operation and maintenance management system based on data center according to claim 1 is characterized in that: The underlying logic of the self-healing architecture operates as follows: Resources are distributed based on the processing resources of the data center, network terminals with the same distributed resources are monitored, and the distributed resource occupancy rates of the same type of network terminals when the required data volume is met and when the required data volume is not met are collected. The occupancy rates are then analyzed. When the required data volume is met, if the deviation of the distributed resource occupancy rates of the same type of network terminals does not exceed the occupancy rate deviation threshold, the distributed resource distribution of the same type of network terminals is inferred to be reasonable, and resource distribution is performed based on the current resource distribution logic. Conversely, if the deviation of the distributed resource occupancy rates of the same type of network terminals exceeds the occupancy rate deviation threshold, the distributed resource distribution of the same type of network terminals is inferred to be unreasonable, and resource distribution is adjusted based on the current resource distribution logic. If the distributed resource occupancy rates of the same type of network terminals are all within the resource occupancy rate range, the distributed resource utilization of the same type of network terminals is inferred to be qualified, and data processing is performed using the current distributed resources. Conversely, if the distributed resource occupancy rates of the same type of network terminals are not all within the resource occupancy rate range, the distributed resource utilization of the same type of network terminals is inferred to be unqualified, and the distributed resources are redistributed.
3. The intelligent operation and maintenance management system based on data center according to claim 2 is characterized in that: When the required data volume is not met, if the distributed resource occupancy rates of network terminals of the same type are all lower than the set occupancy rate threshold, it is inferred that the distributed resources are not suitable for the data processing of the current network terminal; the current network terminals of the same type are finely divided, that is, the data processing requirements are synchronized and the processing requirements are used as the type division standard to ensure the executability of data processing; if the distributed resource occupancy rates of network terminals of the same type are all higher than the set occupancy rate threshold, it is inferred that the distributed resources are suitable for the data processing of the current network terminal.
4. The intelligent operation and maintenance management system based on data center according to claim 1 is characterized in that: The process of the data operation fault prediction module is as follows: obtain the fault period of the data center; obtain the hardware operation parameters corresponding to the distributed resources of the data center according to the fault period; obtain the floating trend of the hardware operation parameters when the fault occurs according to the fault period and mark it as a fault trend, collect the excess increase span of the parameter growth rate of the hardware operation parameters in the current operation process that is in the fault trend relative to the peak value of the parameter growth rate of the historical period, and at the same time obtain the increase span value of the parameter growth span of the hardware operation parameters in the current operation period that is not in the fault trend relative to the historical period; if the excess increase span exceeds the increase span threshold, or the increase span value exceeds the increase span threshold, then it is inferred that the fault prediction of the operation process is high risk; if the excess increase span does not exceed the increase span threshold, and the increase span value does not exceed the increase span threshold, then it is inferred that the fault prediction of the operation process is low risk.
Citation Information
Patent Citations
Container data center operation and maintenance system
CN112818403A
Data center environment monitoring method and system, electronic equipment and storage medium
CN113065293A
Network line intelligent operation and maintenance monitoring management system and method
CN119835143A