Intelligent operation and maintenance management method and system based on data center and storage medium
By adopting intelligent operation and maintenance management systems in the data center, including automated monitoring and self-healing architecture, the problem of inability to monitor and manage the data center environment in the existing technology is solved, real-time monitoring and fault prediction of the hardware and software environment are achieved, and operation and maintenance efficiency is improved.
Patent Information
- Application Number
- CN202510600034.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing technology cannot conduct distributed monitoring and analysis of the hardware and software environment of the data center through architectural settings, cannot reasonably deploy and operate and maintain data centers, and cannot make fault predictions.
It adopts an intelligent operation and maintenance management system based on data center, including an automated monitoring architecture, self-healing architecture and data operation fault prediction module. The automated monitoring architecture monitors the hardware environment in real time through the logic of monitoring, discovery, deployment and troubleshooting. The self-healing architecture monitors the software environment through the logic of resource distribution, parameter acquisition and fault detection. After fault detection and troubleshooting, fault prediction is carried out through historical and real-time operation data comparison.
Real-time monitoring and automated control of the data center hardware and software environment is realized, reducing the risk of operation and maintenance failures, and improving the operation and maintenance efficiency and fault prediction accuracy of the data center.
Smart Images

Figure CN120104432A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data center operation and maintenance technology, and specifically to an intelligent operation and maintenance management method, system and storage medium based on a data center. Background Art
[0002] In the digital age, data has become one of the most critical assets of an enterprise. As the core hub for data storage, processing and distribution, the importance of data centers to modern enterprises is self-evident. It supports various core business systems of enterprises, covering everything from daily office automation and customer relationship management to complex e-commerce transactions, supply chain management and big data analysis.
[0003] The patent with patent number CN112818403A discloses a container data center operation and maintenance system, including: the image construction module is used to obtain the data flow in the data pool and construct the image text according to the data flow; the container publishing module is used to obtain the image text and construct the data container according to the image text; the business authorization module is used to receive the business request sent from any user terminal, and generate a request to be verified corresponding to the data container according to the business request.
[0004] However, in the existing technology, data centers are unable to perform distributed monitoring and analysis of their hardware and software environments through architectural settings, and are unable to compare and analyze the environment with actual data center operations based on algorithms, so that data centers cannot be reasonably deployed and the operation and maintenance management of data centers cannot be completed. In addition, faults in data centers cannot be predicted.
[0005] In view of the above technical defects, a solution is now proposed. Summary of the invention
[0006] The purpose of the present invention is to solve the above-mentioned problems and to propose an intelligent operation and maintenance management method, system and storage medium based on a data center.
[0007] The purpose of the present invention can be achieved through the following technical solutions: an intelligent operation and maintenance management system based on a data center, including an automated monitoring architecture, a self-healing architecture and a data operation fault prediction module; the automated monitoring architecture is constructed by four underlying logics of monitoring, discovery, deployment and troubleshooting, and is used to perform real-time monitoring of the hardware environment of the data center; when the automated monitoring architecture is in operation, the self-healing architecture operates synchronously, and the self-healing architecture is divided into three underlying logics of resource distribution, parameter collection and fault detection, and the function of the self-healing architecture is to monitor the software environment of the data center; the hardware environment and the software environment are controlled and automatically monitored through the automated monitoring architecture and the self-healing architecture, the automated monitoring architecture and the self-healing architecture complete the underlying logic operation, and after the fault is detected and eliminated, the data operation fault prediction module compares the historical operation of the data center with the real-time operation, and predicts the operation performance fault through the operation data comparison.
[0008] As a preferred embodiment of the present invention, the underlying logical operation process of the automated monitoring architecture is as follows: monitor the operating environment of the data center, and the operating environment is the hardware environment; divide the environmental parameters of the hardware environment into operation-affecting parameters and operation-unaffecting parameters; monitor the real-time environmental parameters and compare them with the set parameter thresholds, and if the real-time monitored environmental parameters exceed the set parameter thresholds, the current period is marked as a parameter growth period; otherwise, if the real-time monitored environmental parameters do not exceed the set parameter thresholds, the current period is marked as a parameter stability period.
[0009] As a preferred embodiment of the present invention, the control buffer time of the corresponding environmental parameters before and after the hardware's own operation influencing parameters fluctuate is obtained; before the fluctuation occurs, if the control buffer time exceeds the set time threshold, the environmental parameter control is abnormal, and the environmental control system of the data center is redeployed, specifically deployed to control the environmental parameters in the operating environment around the hardware to ensure that the environmental parameters in the operating environment around the hardware are within the set range; after the fluctuation occurs, if the control buffer time exceeds the set time threshold, the hardware's own operation influencing parameters are controlled abnormally, and the hardware operation operation that causes the operation influencing parameters to fluctuate is deployed and controlled to reduce the real-time workload of the corresponding operation operation, and give priority to the current operation when the same type of operation influencing parameters in the environmental parameters are at a valley value, and the amount of hardware operation influencing parameters generated is controlled, that is, the deviation of the amount of operation influencing parameters generated in the same workload operation period in different commissioning stages of the same type of hardware is within the set deviation range.
[0010] As a preferred embodiment of the present invention, after the parameters are discovered and deployed, the environmental parameters are analyzed as a whole, the environmental parameters corresponding to the parameter growth period and the parameter stability period are collected, and the growth period parameter curve and the stability period parameter curve are constructed respectively; the growth period parameter curve and the stability period parameter curve are analyzed, and the growth period parameter curve exceeds the risk level line, and the corresponding angle is formed, wherein the risk level line is represented as a horizontal line corresponding to the value of the environmental parameter set threshold, and is in the same coordinate system as the curve; if the corresponding angle formed by the curve exceeds the set angle threshold, it is inferred that the environmental parameter control during the parameter growth period is abnormal, that is, the environmental parameter cannot be rapidly reduced after exceeding the threshold and the parameter control speed is slow when the environmental parameter is on a downward trend, then the control performance of the environmental control system cannot match the current hardware operation scenario, and the operation process of the environmental control system in the corresponding operation period is detected and debugged to perform troubleshooting; if the corresponding angle formed by the curve does not exceed the set angle threshold, it is inferred that the environmental parameter control during the parameter growth period is normal.
[0011] As a preferred embodiment of the present invention, the critical values corresponding to the growth trend and the decrease trend of the curve in the parameter curve during the stable period are collected, and the reciprocating floating frequency of the curve corresponding to the critical value is collected in real time when the critical value is continuously updated during the operation stage; if the reciprocating floating frequency of the curve corresponding to the critical value exceeds the floating frequency threshold, it is inferred that the floating of the environmental parameters cannot be effectively controlled, so the control logic of the environmental control system is rectified, and after the control logic of the floating control of the environmental parameter values, a control logic is added, that is, the speed threshold is set for the growth rate of the environmental parameters, and after exceeding the speed threshold, the increase in the environmental parameter value under two unit times of running at the current speed is limited, and if the limited increase exceeds the limited increase, the environmental control system is executed when the growth rate of the environmental parameters exceeds the threshold, otherwise, the environmental control system will not run for the time being to continuously monitor the growth rate of the environmental parameters.
[0012] As a preferred implementation of the present invention, the underlying logic operation process of the self-healing architecture is as follows: resources are distributed according to the processing resources of the data center, network terminals with the same distributed resources are monitored, the distributed resource occupancy rate when the demand data volume of the same type of network terminals is met and the distributed resource occupancy rate when the demand data volume is not met are collected, and the occupancy rate is analyzed: when the demand data volume is met, if the distribution resource occupancy rate deviation of the same type of network terminals does not exceed the occupancy rate deviation threshold, it is inferred that the distribution of the distributed resources of the same type of network terminals is reasonable, and the resources are distributed according to the current resource distribution logic; on the contrary, if the distribution resource occupancy rate deviation of the same type of network terminals exceeds the occupancy rate deviation threshold, it is inferred that the distribution of the distributed resources of the same type of network terminals is unreasonable, and the current resource distribution logic is used as the adjustment starting point to adjust the resource distribution; if the distribution resource occupancy rates of the same type of network terminals are all within the resource occupancy rate range, it is inferred that the distribution resource utilization of the same type of network terminals is qualified, and data processing is performed with the current distributed resources; on the contrary, if the distribution resource occupancy rates of the same type of network terminals are not all within the resource occupancy rate range, it is inferred that the distribution resource utilization of the same type of network terminals is unqualified, and the distributed resources are redistributed.
[0013] As a preferred implementation mode of the present invention, when the required data volume is not met, if the occupancy rates of the distributed resources of the same type of network terminals are all lower than the set occupancy rate threshold, it is inferred that the distributed resources are not suitable for the data processing of the current network terminal; the current network terminals of the same type are finely divided, that is, the data processing requirements are synchronized and the processing requirements are used as the type division standard to ensure the executability of data processing; if the occupancy rates of the distributed resources of the same type of network terminals are all higher than the set occupancy rate threshold, it is inferred that the distributed resources are suitable for the data processing of the current network terminal.
[0014] As a preferred implementation of the present invention, the process of the data operation fault prediction module is as follows: obtain the fault time period of the data center; obtain the hardware operation parameters corresponding to the distributed resources of the data center according to the fault time period; obtain the floating trend of the hardware operation parameters when the fault occurs according to the fault time period and mark it as a fault trend, collect the excess increase span of the parameter growth rate of the hardware operation parameters in the fault trend during the current operation process relative to the peak value of the parameter growth rate during the historical time period, and at the same time obtain the increase span value of the parameter growth span of the hardware operation parameters in the current operation period that are not in the fault trend relative to the historical time period; if the excess increase span exceeds the increase span threshold, or the increase span value exceeds the increase span threshold, then it is inferred that the fault prediction of the operation process is high risk; if the excess increase span does not exceed the increase span threshold, and the increase span value does not exceed the increase span threshold, then it is inferred that the fault prediction of the operation process is low risk.
[0015] As a preferred embodiment of the present invention, an intelligent operation and maintenance management method based on a data center is as follows: automated monitoring, based on the four underlying logic constructions of monitoring, discovery, deployment and troubleshooting, real-time monitoring of the hardware environment of the data center; self-healing fault detection, monitoring of the software environment of the data center based on the three underlying logics of resource distribution, parameter collection and fault detection; automated monitoring and self-healing fault detection complete the underlying logic operation, and after the fault is detected and eliminated, the historical operation of the data center is compared with the real-time operation, and operation performance faults are predicted by comparing the operation data.
[0016] An intelligent operation and maintenance management method based on a data center, such as an intelligent operation and maintenance management system based on a data center as described in any one of the above implementation modes.
[0017] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0018] Compared with the prior art, the beneficial effects of the present invention are: 1. In the present invention, the hardware environment of the data center is monitored in real time, and the hardware environment of the current data center is detected through an automated monitoring architecture to ensure the operational safety of the data center and avoid the operating environment affecting the actual data processing efficiency when the data center is in operation. The influence of the characteristics of the data center itself is not considered, and only the external influence is controlled, and timely discovery and redeployment are carried out to eliminate faults; through software environment monitoring combined with the analysis of the operating parameters of the data center, the operating risk faults of the data center itself are detected and controlled in time to reduce the operating fault efficiency of the data center. The current operating fault is different from the troubleshooting of the automated monitoring architecture. The current operating fault is a network data operating fault of the data center, and the troubleshooting of the automated monitoring architecture is a hardware environment fault of the data center; 2. In the present invention, the historical operation of the data center is compared with the real-time operation, and the operating performance fault is predicted by comparing the operating data to complete the fault prediction of the data center. The accuracy of fault prediction is improved based on the operating data, and the operation and maintenance efficiency of the data center is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0020] Figure 1 It is the overall architecture diagram of the present invention; Figure 2 Flow chart of the automated monitoring architecture of the present invention. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0023] See also Figure 1 As shown, an intelligent operation and maintenance management system based on a data center includes an automated monitoring architecture, a self-healing architecture, and a data operation fault prediction module; it should be explained that the automated monitoring architecture, the self-healing architecture, and the data operation fault prediction module cooperate with each other to perform intelligent monitoring of the data center, so as to improve the efficiency of operation and maintenance management; the thresholds used in this application are obtained by personnel in this field collecting corresponding data and averaging the historical operation and maintenance process of the data center, and are artificially set parameter values; the automated monitoring architecture is constructed by four underlying logics of monitoring, discovery, deployment, and troubleshooting, and is used to monitor the hardware environment of the data center in real time. The hardware environment of the current data center is detected through the automated monitoring architecture to ensure the operational safety of the data center and avoid the operating environment affecting the actual data processing efficiency when the data center is in operation. The influence of the characteristics of the data center itself is not considered, and only external influences are controlled, and timely discovery and redeployment are performed to troubleshoot faults; please refer to Figure 2As shown, the operating environment of the data center is monitored, and the operating environment is the hardware environment; the environmental parameters of the hardware environment are divided into operation-affecting parameters and operation-independent parameters, wherein the operation-affecting parameters are parameters that will cause the environmental parameters to rise when the data center hardware is operating, such as temperature; the operation-independent parameters are parameters that will not cause the environmental parameters to rise when the data center hardware is operating, such as dust concentration; the real-time environmental parameters are monitored and compared with the set parameter thresholds, and if the real-time monitored environmental parameters exceed the set parameter thresholds, the current period is marked as a parameter growth period; otherwise, if the real-time monitored environmental parameters do not exceed the set parameter thresholds, the current period is marked as a parameter stability period; at the same time, the hardware operation-affecting parameters are obtained. The control buffer time of the corresponding environmental parameters before and after the fluctuation occurs, where the control buffer time is expressed as the buffer time for the environmental parameters to be controlled within the set threshold range after the fluctuation occurs; before the fluctuation occurs, if the control buffer time exceeds the set time threshold, the environmental parameter control is abnormal, and the environmental control system of the data center is redeployed, where the environmental control system is represented by the environmental control system, such as the refrigeration system, ventilation system, etc.; the specific deployment is to control the environmental parameters in the operating environment around the hardware to ensure that the environmental parameters of the operating environment around the hardware are within the set range; after the fluctuation occurs, if the control buffer time exceeds the set time threshold, the control of the hardware's own operating parameters is abnormal, and the hardware operation operations that cause the operating parameters to fluctuate are deployed and controlled Control, reduce the real-time workload of the corresponding operation, and give priority to the current operation when the same type of operation-affecting parameters in the environmental parameters are at the valley value, and control the amount of hardware operation-affecting parameters generated, that is, the deviation of the amount of operation-affecting parameters generated in the same workload operation period during different commissioning stages of the same type of hardware is within the set deviation range; after parameter discovery and deployment, conduct an overall analysis of the environmental parameters, collect the corresponding environmental parameters of the parameter growth period and the parameter stability period, and construct the growth period parameter curve and the stability period parameter curve respectively; conduct curve analysis on the growth period parameter curve and the stability period parameter curve, collect the angle formed by the curve of the growth period parameter curve exceeding the risk level line, where the risk level line is represented by the environmental parameter setting The horizontal line of the corresponding value of the fixed threshold is in the same coordinate system as the curve; the angle formed by the curve is expressed as the part of the curve above the horizontal line of the red line value, corresponding to the angle formed by the different trends of the curve when the curve shows a downward trend after reaching the peak value of the curve; if the corresponding angle formed by the curve exceeds the set angle threshold, it is inferred that the environmental parameter control is abnormal during the parameter growth period, that is, the environmental parameter cannot be reduced quickly after exceeding the threshold and the parameter control speed is slow when the environmental parameter shows a downward trend, then the control performance of the environmental control system cannot match the current hardware operation scenario, and the operation process of the environmental control system in the corresponding operation period is detected and debugged to troubleshoot; if the corresponding angle formed by the curve does not exceed the set angle threshold, it is inferred that the environmental parameter control is normal during the parameter growth period;Collect the critical values corresponding to the growth trend and the decrease trend of the curve in the stable period parameter curve. When the critical value is continuously updated in the operation stage, the reciprocating floating frequency of the curve corresponding to the critical value is collected in real time; if the reciprocating floating frequency of the curve corresponding to the critical value exceeds the floating frequency threshold, it is inferred that the environmental parameter control in the parameter stable power transmission is inefficient, that is, the floating of environmental parameters cannot be effectively controlled, but the control is carried out after the environmental parameters show an increasing trend. Although it does not exceed the risk level line, the floating of environmental parameters has an impact on the operation of the hardware. Therefore, the control logic of the environmental control system is rectified. In addition to the control logic of the floating control of the environmental parameter values, a control logic is added, that is, the speed threshold of the environmental parameter growth rate is set. After exceeding the speed threshold, the increase in the environmental parameter value under two unit time of running at the current speed is limited. If the limited increase is exceeded, the environmental control system will be executed when the environmental parameter growth rate exceeds the threshold. Otherwise, the environmental control system will not run temporarily to continuously monitor the growth rate of the environmental parameters. When the automated monitoring architecture is running, the self-healing architecture runs synchronously. Among them, the self-healing architecture is divided into three underlying logics: resource distribution, parameter collection, and fault detection. The function of the self-healing architecture is to monitor the software environment of the data center. Through software environment monitoring combined with the analysis of the operating parameters of the data center, the risk faults of the data center itself are detected and controlled in time to reduce the operation of the data center. Fault transfer efficiency: the current operation fault is different from the troubleshooting of the automated monitoring architecture. The current operation fault is the network data operation fault of the data center, while the troubleshooting of the automated monitoring architecture is the hardware environment fault of the data center. The resources are distributed according to the processing resources of the data center, where the processing resources are represented by data processing memory, algorithm computing power pool and other resources. Parameters are collected according to the operation process of the network terminals covered by the data center. The network terminals with the same distributed resources are monitored, and the distribution resource occupancy rate of the same type of network terminals when the demand data volume is met and the distribution resource occupancy rate when the demand data volume is not met are collected, and the occupancy rate is analyzed: when the demand data volume is met, if the same type of If the distribution resource occupancy rate deviation of the same type of network terminal does not exceed the occupancy rate deviation threshold, it is inferred that the distribution of the distribution resources of the same type of network terminal is reasonable, and the resource distribution is performed according to the current resource distribution logic; on the contrary, if the distribution resource occupancy rate deviation of the same type of network terminal exceeds the occupancy rate deviation threshold, it is inferred that the distribution of the distribution resources of the same type of network terminal is unreasonable, and the current resource distribution logic is used as the adjustment starting point to adjust the resource distribution, reduce the occupancy rate deviation, and prevent the occurrence of excess data processing resources; if the distribution resource occupancy rates of the same type of network terminals are all within the resource occupancy rate range, it is inferred that the distribution resource utilization of the same type of network terminals is qualified, and the current distributed resources are used for data processing;On the contrary, if the occupancy rates of the distributed resources of the same type of network terminals are not all within the resource occupancy rate range, it is inferred that the utilization of the distributed resources of the same type of network terminals is unqualified, and the distributed resources are redistributed, and the resource distribution is carried out in combination with the data complexity of the network terminals, where the data complexity is expressed as the number of non-same type data, the frequency of continuous data processing, etc.; when the required data volume is not met, if the occupancy rates of the distributed resources of the same type of network terminals are all lower than the set occupancy rate threshold, it is inferred that the distributed resources are not suitable for the data processing of the current network terminals, such as inconsistent data transmission protocols, incompatible network hardware equipment, and inconsistent data transmission formats; for the current The same type of network terminals are divided into detailed categories, that is, the data processing requirements are synchronized and the processing requirements are used as the type classification standard to ensure the executability of data processing; if the distributed resource occupancy rates of the same type of network terminals are all higher than the set occupancy rate threshold, it is inferred that the distributed resources are suitable for the data processing of the current network terminal; the same type of network terminals are represented by the deviation of the required data volume within the set range, and the resource demand deviation is not large; the required data volume is represented by the amount of data that the network terminal can produce, and the data needs to be processed through distributed resources; the hardware environment and software environment are controlled and automatically monitored through the automated monitoring architecture and self-healing architecture to reduce manual work. The risk is checked to minimize the impact of external interference on the location where the data center hardware is stored; the automated monitoring architecture and the self-healing architecture complete the underlying logic operation, and after the fault is detected and eliminated, a completion signal is generated and sent to the data operation fault prediction module; the data operation fault prediction module is used to compare the historical operation of the data center with the real-time operation, and to predict the operation performance fault by comparing the operation data, so as to complete the fault prediction of the data center, and to improve the accuracy of fault prediction based on the operation data, and to improve the operation and maintenance efficiency of the data center; according to the historical operation process, the fault period of the data center is obtained, where the fault period includes all fault periods, such as equipment failure caused by hardware environment abnormalities and equipment processing failure caused by software environment abnormalities; according to the fault period, the hardware operation parameters corresponding to the distributed resources of the data center are obtained, such as hard disk read and write speed, temperature and other parameters; according to the fault period, the floating trend of the hardware operation parameters when the fault occurs is obtained and marked as a fault trend, and the excess increase span of the parameter growth rate of the hardware operation parameters in the current operation process that is in the fault trend relative to the peak value of the parameter growth rate in the historical period is collected, and at the same time, the increase span value of the average value of the parameter growth span of the hardware operation parameters that are not in the fault trend in the current operation period relative to the historical period is obtained;And compare the excess span of the parameter growth rate of the hardware operating parameters in the fault trend during the current operation process relative to the peak value of the parameter growth rate in the historical period, and the increased span value of the mean parameter growth span of the hardware operating parameters not in the fault trend during the current operation period relative to the historical period with the increased span threshold and the increased span threshold respectively: if the excess span of the parameter growth rate of the hardware operating parameters in the fault trend during the current operation process relative to the peak value of the parameter growth rate in the historical period exceeds the increased span threshold, or the increased span value of the mean parameter growth span of the hardware operating parameters not in the fault trend during the current operation period relative to the historical period exceeds the increased span threshold , then the fault prediction of the operation process is inferred to be high risk, a high risk signal is generated and sent to the administrator terminal, and the administrator terminal monitors the operation parameters of the data center after receiving it and controls the trend growth in time; if the hardware operation parameters in the current operation process are in the fault trend, the parameter growth rate relative to the peak value of the parameter growth rate in the historical period does not exceed the increase span threshold, and the hardware operation parameters in the current operation period are not in the fault trend, and the parameter growth span mean relative to the increase span value of the historical period does not exceed the increase span threshold, then the fault prediction of the operation process is inferred to be low risk, a low risk signal is generated and sent to the administrator terminal. ;
[0024] The intelligent operation and maintenance management method based on the data center is as follows: automated monitoring, based on the four underlying logics of monitoring, discovery, deployment and troubleshooting, to monitor the hardware environment of the data center in real time; self-healing fault detection, based on the three underlying logics of resource distribution, parameter collection and fault detection, to monitor the software environment of the data center; automated monitoring and self-healing fault detection complete the underlying logic operation, and after the fault is detected and eliminated, compare the historical operation of the data center with the real-time operation, and predict the operation performance fault through operation data comparison.
[0025] Additionally, a storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned intelligent operation and maintenance management method based on a data center.
[0026] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods.
[0027] Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory, the non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory, and the volatile memory may include random access memory (RAM) or external cache memory.
[0028] By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0029] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0030] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. Intelligent operation and maintenance management system based on data center, characterized by: It includes automated monitoring architecture, self-healing architecture and data operation fault prediction module; the automated monitoring architecture is constructed by four underlying logics: monitoring, discovery, deployment and troubleshooting, and is used to monitor the hardware environment of the data center in real time; When the automated monitoring architecture is operating, the self-healing architecture operates synchronously. The self-healing architecture is divided into three underlying logics: resource distribution, parameter collection, and fault detection. The function of the self-healing architecture is to monitor the software environment of the data center. The hardware environment and software environment are controlled and automatically monitored through the automated monitoring architecture and the self-healing architecture. After the automated monitoring architecture and the self-healing architecture complete the underlying logic operation and the fault is detected and eliminated, the data operation fault prediction module compares the historical operation of the data center with the real-time operation, and predicts operation performance faults through operation data comparison.
2. The intelligent operation and maintenance management system based on data center according to claim 1 is characterized in that: The underlying logical operation process of the automated monitoring architecture is as follows: monitor the operating environment of the data center, and the operating environment is the hardware environment; divide the environmental parameters of the hardware environment into operating-affecting parameters and operating-independent parameters; monitor the real-time environmental parameters and compare them with the set parameter thresholds, and if the real-time monitored environmental parameters exceed the set parameter thresholds, the current period is marked as a parameter growth period; conversely, if the real-time monitored environmental parameters do not exceed the set parameter thresholds, the current period is marked as a parameter stability period.
3. The intelligent operation and maintenance management system based on data center according to claim 2 is characterized in that: The control buffer time of the corresponding environmental parameters before and after the fluctuation of the hardware's own operation influencing parameters is obtained; before the fluctuation occurs, if the control buffer time exceeds the set time threshold, the environmental parameter control is abnormal, and the environmental control system of the data center is redeployed, specifically deployed to control the environmental parameters in the operating environment around the hardware to ensure that the environmental parameters of the operating environment around the hardware are within the set range; after the fluctuation occurs, if the control buffer time exceeds the set time threshold, the control of the hardware's own operation influencing parameters is abnormal, and the hardware operation operation that causes the fluctuation of the operation influencing parameters is deployed and controlled to reduce the real-time workload of the corresponding operation operation, and give priority to the current operation when the same type of operation influencing parameters in the environmental parameters are at the valley value, and the amount of hardware operation influencing parameters generated is controlled, that is, the deviation of the amount of operation influencing parameters generated in the same workload operation period in different commissioning stages of the same type of hardware is within the set deviation range.
4. The intelligent operation and maintenance management system based on data center according to claim 3 is characterized in that: After the parameters are discovered and deployed, the environmental parameters are analyzed as a whole, the environmental parameters corresponding to the parameter growth period and parameter stability period are collected, and the parameter curves for the growth period and the parameter curves for the stability period are constructed respectively; Perform curve analysis on the parameter curve of the growth period and the parameter curve of the stable period, and collect the corresponding angle formed by the curve of the parameter curve of the growth period exceeding the risk level line, where the risk level line is represented by the horizontal line corresponding to the value of the set threshold of the environmental parameter, and is in the same coordinate system as the curve; if the corresponding angle formed by the curve exceeds the set angle threshold, it is inferred that the environmental parameter control in the parameter growth period is abnormal, that is, the environmental parameter cannot be quickly reduced after exceeding the threshold and the parameter control speed is slow when the environmental parameter is on a downward trend, then the control performance of the environmental control system cannot match the current hardware operation scenario, and the operation process of the environmental control system in the corresponding operation period is detected and debugged to troubleshoot; If the angle formed by the corresponding curve does not exceed the set angle threshold, it is inferred that the environmental parameter control is normal during the parameter growth period.
5. The intelligent operation and maintenance management system based on data center according to claim 4 is characterized in that: Collect the critical values corresponding to the curve growth trend and the decrease trend in the parameter curve during the stable period, and collect the reciprocating floating frequency of the curve corresponding to the critical value in real time when the critical value is continuously updated during the operation stage; if the reciprocating floating frequency of the curve corresponding to the critical value exceeds the floating frequency threshold, it is inferred that the fluctuation of the environmental parameters cannot be effectively controlled, so the control logic of the environmental control system is rectified. In addition to the control logic of the floating control of the environmental parameter values, an additional control logic is added, that is, the speed threshold is set for the growth rate of the environmental parameters, and after exceeding the speed threshold, the increase in the environmental parameter value under two unit times of running at the current speed is limited. If the increase exceeds the limited increase, the environmental control system is executed when the growth rate of the environmental parameters exceeds the threshold. Otherwise, the environmental control system will not run for the time being to continuously monitor the growth rate of the environmental parameters.
6. The intelligent operation and maintenance management system based on data center according to claim 1 is characterized in that: The underlying logic operation process of the self-healing architecture is as follows: distribute resources according to the processing resources of the data center, monitor the network terminals with the same distributed resources, collect the distributed resource occupancy rate when the demand data volume of the same type of network terminals is met and the distributed resource occupancy rate when the demand data volume is not met, and analyze the occupancy rate: when the demand data volume is met, if the distribution resource occupancy rate deviation of the same type of network terminals does not exceed the occupancy rate deviation threshold, it is inferred that the distribution of the distributed resources of the same type of network terminals is reasonable, and the resources are distributed according to the current resource distribution logic; otherwise, if the distribution resource occupancy rate deviation of the same type of network terminals exceeds the occupancy rate deviation threshold, it is inferred that the distribution of the distributed resources of the same type of network terminals is unreasonable, and the current resource distribution logic is used as the adjustment starting point to adjust the resource distribution; if the distribution resource occupancy rates of the same type of network terminals are all within the resource occupancy rate range, it is inferred that the distribution resource utilization of the same type of network terminals is qualified, and data processing is performed with the current distributed resources; otherwise, if the distribution resource occupancy rates of the same type of network terminals are not all within the resource occupancy rate range, it is inferred that the distribution resource utilization of the same type of network terminals is unqualified, and the distributed resources are redistributed.
7. The intelligent operation and maintenance management system based on data center according to claim 6 is characterized in that: When the required data volume is not met, if the distributed resource occupancy rates of the same type of network terminals are all lower than the set occupancy rate threshold, it is inferred that the distributed resources are not suitable for the data processing of the current network terminal; the current network terminals of the same type are finely divided, that is, the data processing requirements are synchronized and the processing requirements are used as the type division standard to ensure the executability of data processing; if the distributed resource occupancy rates of the same type of network terminals are all higher than the set occupancy rate threshold, it is inferred that the distributed resources are suitable for the data processing of the current network terminal.
8. The intelligent operation and maintenance management system based on data center according to claim 1 is characterized in that: The process of the data operation fault prediction module is as follows: obtain the fault period of the data center; obtain the hardware operation parameters corresponding to the distributed resources of the data center according to the fault period; obtain the floating trend of the hardware operation parameters when the fault occurs according to the fault period and mark it as a fault trend, collect the excess increase span of the parameter growth rate of the hardware operation parameters in the fault trend during the current operation process relative to the peak value of the parameter growth rate during the historical period, and at the same time obtain the increase span value of the parameter growth span of the hardware operation parameters that are not in the fault trend in the current operation period relative to the historical period; if the excess increase span exceeds the increase span threshold, or the increase span value exceeds the increase span threshold, then it is inferred that the fault prediction of the operation process is high risk; if the excess increase span does not exceed the increase span threshold, and the increase span value does not exceed the increase span threshold, then it is inferred that the fault prediction of the operation process is low risk.
9. An intelligent operation and maintenance management method based on a data center, characterized in that: An intelligent operation and maintenance management system based on a data center as described in any one of claims 1 to 8 above.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to claim 9 is implemented.
Citation Information
Patent Citations
Container data center operation and maintenance system
CN112818403A
Data center environment monitoring method and system, electronic equipment and storage medium
CN113065293A
Intelligent operation and maintenance management method and system
CN115208742A
Gastrodia elata processing raw material product storage environment supervision system based on Internet of Things
CN116862731A
Network line intelligent operation and maintenance monitoring management system and method
CN119835143A