An operation and maintenance data management method and system based on big data analysis
By dividing the data center into multiple operation and maintenance sub-regions, calculating the operation and maintenance warning coefficients and conducting risk warnings, the problem of lack of automated monitoring and real-time operation and maintenance data management in the existing technology is solved, and efficient operation and maintenance data management and business continuity guarantee are achieved.
Patent Information
- Application Number
- CN202411414129.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-10-11
AI Technical Summary
The existing operation and maintenance data management methods lack an automated monitoring mechanism, which leads to untimely warning of emergencies, affecting the work efficiency of the operation and maintenance host, and lacks real-time and global fault monitoring.
By dividing the target data center into multiple operation and maintenance sub-regions, the equipment failure coefficient and task completion index of each operation and maintenance sub-regions are obtained, the regional operation and maintenance warning coefficient is calculated, the operation and maintenance sub-regions are divided into types, and the operation and risk warning is performed based on the operation and maintenance area division data.
It realizes automated dynamic monitoring of the target data center, improves the timeliness of emergency warnings, shortens response time, ensures business continuity, and effectively guarantees the real-time and global nature of data center fault monitoring.
Smart Images

Figure CN118916248B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computers and relates to operation and maintenance technology, specifically a method and system for operation and maintenance data management based on big data analysis. Background Art
[0002] The existing operation and maintenance data management methods have the following defects during the operation and maintenance process: 1. There is a lack of an automated monitoring mechanism for operation and maintenance hosts with failures or performance degradation, which easily leads to untimely early warnings of emergencies and thus affects the work efficiency of operation and maintenance hosts; 2. The process of fault monitoring for operation and maintenance hosts lacks real-time and overall perspectives. After a fault occurs, the monitoring system cannot issue an alarm in time, resulting in the operation and maintenance team being unable to respond quickly. For this reason, we propose a method and system for operation and maintenance data management based on big data analysis. Summary of the Invention
[0003] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a method and system for operation and maintenance data management based on big data analysis. To achieve the above purpose, the present invention adopts the following technical solutions: A method for operation and maintenance data management based on big data analysis specifically includes the following steps: Step S1: Divide the target data center into multiple operation and maintenance sub-regions respectively, obtain the equipment failure coefficient corresponding to each operation and maintenance sub-region respectively, obtain the task detection coefficient corresponding to each operation and maintenance sub-region respectively, calculate the mean value of the task detection coefficient to obtain the task completion index, and obtain the equipment operation and maintenance data; Calculate the equipment failure coefficient corresponding to the first operation and maintenance sub-region, and the specific formula configuration is as follows:
[0004] ; where Sbg is the equipment failure coefficient corresponding to the first operation and maintenance sub-region, Tb1 to Tbn are the nth reported failure values of the first reported failure value respectively, and Gzj is the failure monitoring time length value; Calculate the first task detection coefficient, and the specific formula configuration is as follows:
[0005] ; where Rwj1 is the first task detection coefficient, Cpu is the average CPU usage rate, Ncl is the average memory utilization rate, and Rzx is the task execution duration; Step S2: By analyzing the equipment operation and maintenance data, calculate the regional operation and maintenance warning coefficient corresponding to each operation and maintenance sub-region, obtain the threshold of the regional operation and maintenance warning coefficient and compare it with the regional operation and maintenance warning coefficient, and divide the operation and maintenance sub-region into the first type of operation and maintenance sub-region and the second type of operation and maintenance sub-region according to the numerical comparison result to obtain the operation and maintenance region division data; Calculate the regional operation and maintenance warning coefficient, and the specific formula configuration is as follows:
[0006] ; where, Qyy is the regional operation and maintenance warning coefficient, Sbg is the equipment failure coefficient, and Rwz is the task completion index; Step S3: Perform operation risk warnings on the operation and maintenance sub-regions and the target data center respectively according to the operation and maintenance area division data.
[0007] Furthermore, in the step S1, the following specific steps are further included: Step S11: Divide the target data center into m operation and maintenance equipment sub-regions with the same number of servers, and name them the first operation and maintenance sub-region to the mth operation and maintenance sub-region respectively; Step S12: Monitor the operation and maintenance equipment failures in the first operation and maintenance sub-region to obtain the equipment failure coefficient corresponding to the first operation and maintenance sub-region; Step S13: Obtain the task completion index corresponding to the first operation and maintenance sub-region; Step S14: Define the equipment failure coefficient and the task completion index corresponding to the first operation and maintenance sub-region as the first regional operation and maintenance data; Step S15: Obtain the regional operation and maintenance data corresponding to the second to the mth operation and maintenance sub-regions respectively to obtain the second regional operation and maintenance data to the mth regional operation and maintenance data; Step S16: Define the first regional operation and maintenance data to the mth regional operation and maintenance data as the equipment operation and maintenance data.
[0008] Furthermore, in the step S12, the following specific steps are further included: Step S121: During the process of detecting failures in the first operation and maintenance sub-region, obtain the time value corresponding to the current moment as the first reference time point, take the time point corresponding to one characteristic monitoring duration before the first reference time as the second reference time point, and define the time period between the first reference time point and the second reference time point as the failure monitoring period; Step S122: Obtain the time length value corresponding to the failure monitoring period to obtain the failure monitoring time length value; Step S123: Select n operation and maintenance hosts in the first operation and maintenance sub-region as test hosts respectively, and name them the first test host to the nth test host respectively; Step S124: Obtain the number of reported failure times of the first test host to the nth test host during the failure monitoring period to obtain the first reported failure value to the nth reported failure value; Step S125: Calculate the first reported failure value to the nth reported failure value and the failure monitoring time length value to obtain the equipment failure coefficient corresponding to the first operation and maintenance sub-region.
[0009] Further, in the step S13, the following specific steps are further included: Step S131: Obtain the work task list corresponding to the first operation and maintenance sub-region, select i work tasks from the work task list as sample work tasks respectively, and name them as the first sample work task to the i-th sample work task; Step S132: Conduct work detection on the first sample work task to obtain the first task monitoring coefficient; Step S133: Conduct work monitoring on the second sample work task to the i-th sample work task respectively to obtain the second task monitoring coefficient to the i-th task monitoring coefficient; Step S134: Calculate the average of the first task monitoring coefficient to the i-th task monitoring coefficient to obtain the task completion index corresponding to the first operation and maintenance sub-region.
[0010] Further, in the step S132, the following specific steps are further included: Step S1321: During the completion process of the first sample work task, obtain the average CPU usage rate of the operation and maintenance host executing the first sample work task; Step S1322: During the completion process of the first sample work task, obtain the average memory utilization rate of the operation and maintenance host executing the first sample work task; Step S1323: Take the time point when the first operation and maintenance sub-region starts to execute the first sample work task as the first execution monitoring time point; Step S1324: Take the time point when the first operation and maintenance sub-region ends to execute the first sample work task as the second execution monitoring time point; Step S1325: Calculate the numerical difference between the first execution monitoring time point and the second execution monitoring time point to obtain the task execution duration; Step S1326: Calculate the average CPU usage rate, the average memory utilization rate, and the task execution duration to obtain the first task monitoring coefficient.
[0011] Further, in the step S2, the following specific steps are further included: Step S21: Obtain the device operation and maintenance data, and obtain the first area operation and maintenance data to the m-th area operation and maintenance data respectively according to the device operation and maintenance data; Step S22: Conduct early warning analysis on the first operation and maintenance sub-region according to the first area operation and maintenance data to obtain the area operation and maintenance early warning coefficient corresponding to the first operation and maintenance sub-region; Step S23: Obtain the area operation and maintenance early warning coefficients corresponding to the second area operation and maintenance data to the m-th area operation and maintenance data respectively; Step S24: Obtain the numerical comparison between the area operation and maintenance early warning coefficient threshold and the area operation and maintenance early warning coefficient, and divide the operation and maintenance sub-region into the first type of operation and maintenance sub-region and the second type of operation and maintenance sub-region according to the numerical comparison result to obtain the operation and maintenance area division data.
[0012] Further, in the step S22, the following specific steps are further included: Step S221: Obtain the device failure coefficient and the task completion index respectively according to the first area operation and maintenance data; Step S222: Calculate the device failure coefficient and the task completion index to obtain the area operation and maintenance early warning coefficient.
[0013] Further, in step S24, the following specific steps are further included: Step S241: Obtain the equipment failure coefficient threshold and the task completion index threshold respectively, and calculate the regional operation and maintenance warning coefficient threshold by calculating the equipment failure coefficient threshold and the task completion index threshold; calculate the regional operation and maintenance warning coefficient threshold, and the specific formula configuration is as follows:
[0014] ; where Qyy1 is the regional operation and maintenance warning coefficient threshold, Sbg1 is the equipment failure coefficient threshold, and Rwz1 is the task completion index threshold; Step S242: When the regional operation and maintenance warning coefficient is greater than or equal to the regional operation and maintenance warning coefficient threshold, determine that the corresponding operation and maintenance sub-region is the first type of operation and maintenance sub-region; Step S243: When the regional operation and maintenance warning coefficient is less than the regional operation and maintenance warning coefficient threshold, determine that the corresponding operation and maintenance sub-region is the second type of operation and maintenance sub-region.
[0015] Further, in step S3, the following specific steps are further included: Step S31: Obtain the operation and maintenance area division data; Step S32: When the operation and maintenance sub-region is the first type of operation and maintenance sub-region, perform normal risk monitoring on the operation and maintenance sub-region; Step S33: When the operation and maintenance sub-region is the second type of operation and maintenance sub-region, issue an operation and maintenance risk warning for the operation and maintenance sub-region; Step S34: Obtain the numerical value corresponding to the second type of operation and maintenance sub-region in the target data center to obtain the abnormal operation and maintenance area numerical value, calculate the ratio of the abnormal operation and maintenance area numerical value to m to obtain the abnormal area ratio; Step S35: Obtain the abnormal area ratio threshold, and compare the abnormal area ratio with the abnormal area ratio threshold; specifically as follows: Step S351: When the abnormal area ratio is greater than or equal to the abnormal area ratio threshold, issue an operation risk warning for the target data center; Step S352: When the abnormal area ratio is less than the abnormal area ratio threshold, perform normal operation risk monitoring on the target data center.
[0016] An operation and maintenance data management system based on big data analysis, and the specific working processes of each module are as follows: Data acquisition module: used to divide the target data center into multiple operation and maintenance sub-regions respectively, and obtain the equipment failure coefficient and task completion index corresponding to each operation and maintenance sub-region to obtain equipment operation and maintenance data; Data analysis module: used to divide the operation and maintenance sub-regions into the first type of operation and maintenance sub-regions and the second type of operation and maintenance sub-regions by analyzing the equipment operation and maintenance data to obtain operation and maintenance area division data; Risk warning module: used to perform operation risk warnings on the operation and maintenance sub-regions and the target data center respectively according to the operation and maintenance area division data.
[0017] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows: 1. The present invention divides the target data center into multiple operation and maintenance sub-regions respectively, and obtains the equipment failure coefficient and task completion index corresponding to each operation and maintenance sub-region respectively, realizing automatic dynamic monitoring of the target data center, thereby improving the timeliness of emergency warning, shortening the response time, and ensuring business continuity. 2. The present invention conducts operation risk early warning on the operation and maintenance sub-regions and the target data center respectively according to the operation and maintenance area division data, which can effectively ensure the real-time and overall nature of the data center fault monitoring process. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the accompanying drawings.
[0019] Figure 1 It is a flowchart of the implementation steps of the present invention.
[0020] Figure 2 It is a block diagram of the overall system of the present invention.
[0021] Figure 3 It is a data interaction flowchart of the present invention. SPECIFIC EMBODIMENTS
[0022] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] Embodiment 1
[0024] Please refer to Figure 1 , the present invention provides a technical solution: a method for managing operation and maintenance data based on big data analysis, including the following specific steps:
[0025] Step S1: Obtain device operation and maintenance data; Step S11: Divide the target data center into m operation and maintenance device sub - regions with the same number of servers, and name them the first operation and maintenance sub - region to the mth operation and maintenance sub - region respectively; Step S12: Conduct operation and maintenance device fault monitoring on the first operation and maintenance sub - region to obtain the device fault coefficient corresponding to the first operation and maintenance sub - region; In the said Step S12, it further includes the following specific steps: Step S121: During the process of fault detection on the first operation and maintenance sub - region, obtain the time value corresponding to the current moment as the first reference time point, take the time point corresponding to one characteristic monitoring duration before the first reference time as the second reference time point, and define the time period between the first reference time point and the second reference time point as the fault monitoring period; Step S122: Obtain the time length value corresponding to the fault monitoring period to get the fault monitoring time length value; Step S123: Select n operation and maintenance hosts as test hosts in the first operation and maintenance sub - region respectively, and name them the first test host to the nth test host respectively; Step S124: Obtain the number of fault times reported by the first test host to the nth test host during the fault monitoring period respectively to get the first reported fault value and the nth reported fault value; Step S125: Calculate the device fault coefficient corresponding to the first operation and maintenance sub - region by calculating the first reported fault value and the nth reported fault value with the fault monitoring time length value; Calculate the device fault coefficient corresponding to the first operation and maintenance sub - region, and the specific formula configuration is as follows:
[0026] ; where Sbg is the equipment failure coefficient corresponding to the first operation and maintenance sub-region, Tb1 to Tbn are the nth reported fault values of the first reported fault values respectively, and Gzj is the fault monitoring time length value; Step S13: Obtain the task completion index corresponding to the first operation and maintenance sub-region; In the said Step S13, it further includes the following specific steps: Step S131: Obtain the work task list corresponding to the first operation and maintenance sub-region, select i work tasks from the work task list as sample work tasks respectively, and name them the first sample work task to the i-th sample work task; Step S132: Conduct work detection on the first sample work task to obtain the first task monitoring coefficient; In the said Step S132, it further includes the following specific steps: Step S1321: During the completion process of the first sample work task, obtain the average CPU usage rate of the operation and maintenance host executing the first sample work task; Step S1322: During the completion process of the first sample work task, obtain the average memory utilization rate of the operation and maintenance host executing the first sample work task; Step S1323: Use the time point when the first operation and maintenance sub-region starts to execute the first sample work task as the first execution monitoring time point; Step S1324: Use the time point when the first operation and maintenance sub-region ends to execute the first sample work task as the second execution monitoring time point; Step S1325: Calculate the numerical difference between the first execution monitoring time point and the second execution monitoring time point to obtain the task execution duration; Step S1326: Through calculation of the average CPU usage rate, average memory utilization rate, and task execution duration, obtain the first task monitoring coefficient; The calculation of the first task monitoring coefficient is specifically configured with the following formula:
[0027] ; where Rwj1 is the first task monitoring coefficient, Cpu is the average CPU usage rate, Ncl is the average memory utilization rate, and Rzx is the task execution duration; Step S133: Conduct work detection on the second sample work task to the i-th sample work task respectively to obtain the second task monitoring coefficient to the i-th task monitoring coefficient; Step S134: Calculate the average of the first task monitoring coefficient to the i-th task monitoring coefficient to obtain the task completion index corresponding to the first operation and maintenance sub-region; Step S14: Define the equipment failure coefficient and task completion index corresponding to the first operation and maintenance sub-region as the first region operation and maintenance data; Step S15: Obtain the region operation and maintenance data corresponding to the second to the mth operation and maintenance sub-regions respectively to obtain the second region operation and maintenance data to the mth region operation and maintenance data; Step S16: Define the first region operation and maintenance data to the mth region operation and maintenance data as the equipment operation and maintenance data.
[0028] Step S2: By analyzing the equipment operation and maintenance data, divide the operation and maintenance sub-areas into the first type of operation and maintenance sub-areas and the second type of operation and maintenance sub-areas to obtain the operation and maintenance area division data; Step S21: Obtain the equipment operation and maintenance data, and respectively obtain the operation and maintenance data of the first area to the mth area according to the equipment operation and maintenance data; Step S22: Conduct a warning analysis on the first operation and maintenance sub-area according to the operation and maintenance data of the first area to obtain the area operation and maintenance warning coefficient corresponding to the first operation and maintenance sub-area; In the said step S22, the following specific steps are further included: Step S221: Respectively obtain the equipment failure coefficient and the task completion index according to the operation and maintenance data of the first area; Step S222: Calculate the area operation and maintenance warning coefficient by calculating the equipment failure coefficient and the task completion index; Calculate the area operation and maintenance warning coefficient, and the specific formula configuration is as follows:
[0029] ; where, Qyy is the area operation and maintenance warning coefficient, Sbg is the equipment failure coefficient, and Rwz is the task completion index; Step S23: Respectively obtain the area operation and maintenance warning coefficients corresponding to the operation and maintenance data of the second area to the mth area; Step S24: Obtain the area operation and maintenance warning coefficient threshold and compare the numerical values with the area operation and maintenance warning coefficient. According to the numerical comparison result, divide the operation and maintenance sub-areas into the first type of operation and maintenance sub-areas and the second type of operation and maintenance sub-areas to obtain the operation and maintenance area division data; In the said step S24, the following specific steps are further included: Step S241: Respectively obtain the equipment failure coefficient threshold and the task completion index threshold, and calculate the area operation and maintenance warning coefficient threshold by calculating the equipment failure coefficient threshold and the task completion index threshold; Calculate the area operation and maintenance warning coefficient threshold, and the specific formula configuration is as follows:
[0030] ; where, Qyy1 is the area operation and maintenance warning coefficient threshold, Sbg1 is the equipment failure coefficient threshold, and Rwz1 is the task completion index threshold; Step S242: When the area operation and maintenance warning coefficient is greater than or equal to the area operation and maintenance warning coefficient threshold, determine that the corresponding operation and maintenance sub-area is the first type of operation and maintenance sub-area; Step S243: When the area operation and maintenance warning coefficient is less than the area operation and maintenance warning coefficient threshold, determine that the corresponding operation and maintenance sub-area is the second type of operation and maintenance sub-area.
[0031] Step S3: Please refer to Figure 3, perform operation risk early warning on the target data center according to the operation and maintenance area division data; Step S31: Obtain the operation and maintenance area division data; Step S32: When the operation and maintenance sub - area is the first - type operation and maintenance sub - area, perform risk monitoring on the operation and maintenance sub - area normally; Step S33: When the operation and maintenance sub - area is the second - type operation and maintenance sub - area, issue an operation and maintenance risk early warning for the operation and maintenance sub - area; Step S34: Obtain the quantity value corresponding to the second - type operation and maintenance sub - area in the target data center to get the abnormal operation and maintenance area quantity value, calculate the ratio of the abnormal operation and maintenance area quantity value to m to get the abnormal area quantity ratio; Step S35: Obtain the abnormal area quantity ratio threshold, and compare the abnormal area quantity ratio with the abnormal area quantity ratio threshold; specifically as follows: Step S351: When the abnormal area quantity ratio is greater than or equal to the abnormal area quantity ratio threshold, issue an operation risk early warning for the target data center; Step S352: When the abnormal area quantity ratio is less than the abnormal area quantity ratio threshold, perform operation risk monitoring on the target data center normally.
[0032] In this application, if there are corresponding calculation formulas, the above - mentioned calculation formulas are all dimensionless and take their numerical values for calculation. For coefficients such as weight coefficients and proportionality coefficients in the formulas, the magnitudes set are for obtaining a result value by quantifying each parameter. Regarding the magnitudes of the weight coefficients and proportionality coefficients, as long as the proportional relationship between the parameters and the result value is not affected.
[0033] Embodiment Two
[0034] Please refer to Figure 2 , based on another concept of the same invention, a new operation and maintenance data management system based on big - data analysis is proposed, including a data acquisition module, a data analysis module, a risk early - warning module, and a server. The data acquisition module, the data analysis module, and the risk early - warning module are respectively connected to the server, and the server controls the data acquisition module, the data analysis module, and the risk early - warning module respectively; The data acquisition module acquires equipment operation and maintenance data; The target data center is divided into m operation and maintenance equipment sub - areas with the same number of servers, and they are respectively named the first operation and maintenance sub - area to the m - th operation and maintenance sub - area.
[0035] It should be noted here that: In this application, m is the quantity value corresponding to the operation and maintenance equipment sub - area of the target data center, and m is an integer greater than 0; Perform operation and maintenance equipment fault monitoring on the first operation and maintenance sub - area to obtain the equipment fault coefficient corresponding to the first operation and maintenance sub - area; specifically as follows: During the process of fault detection on the first operation and maintenance sub - area, obtain the time value corresponding to the current moment as the first reference time point, take the time point corresponding to one characteristic monitoring duration before the first reference time as the second reference time point, and define the time period between the first reference time point and the second reference time point as the fault monitoring period; Obtain the time - length value corresponding to the fault monitoring period to get the fault monitoring time - length value.
[0036] It should be noted here that: in this application, the feature monitoring duration is specifically limited to 30 minutes; the fault monitoring period involved here is dynamically updated in real time as the time value corresponding to the current moment changes; n operation and maintenance hosts are respectively selected as test hosts in the first operation and maintenance sub-region, and they are respectively named the first test host to the nth test host.
[0037] It should be noted here that: in this application, n is the numerical value corresponding to the number of test hosts, and n is an integer greater than 0; the number of reported fault times of the first test host to the nth test host during the fault monitoring period is respectively obtained to get the first reported fault value to the nth reported fault value; the first reported fault value to the nth reported fault value and the fault monitoring time length value are calculated to obtain the equipment fault coefficient corresponding to the first operation and maintenance sub-region; the calculation of the equipment fault coefficient corresponding to the first operation and maintenance sub-region is specifically configured with the following formula:
[0038] ; where Sbg is the equipment fault coefficient corresponding to the first operation and maintenance sub-region, Tb1 to Tbn are respectively the first reported fault value to the nth reported fault value, and Gzj is the fault monitoring time length value; obtain the task completion index corresponding to the first operation and maintenance sub-region; specifically as follows: obtain the work task list corresponding to the first operation and maintenance sub-region, and respectively select i work tasks as sample work tasks in the work task list, and they are respectively named the first sample work task to the ith sample work task.
[0039] It should be noted here that: in this application, i is the numerical value corresponding to the number of sample work tasks, and i is an integer greater than 0; perform work detection on the first sample work task to obtain the first task monitoring coefficient; specifically as follows: during the completion process of the first sample work task, obtain the average CPU usage rate of the operation and maintenance host executing the first sample work task; during the completion process of the first sample work task, obtain the average memory utilization rate of the operation and maintenance host executing the first sample work task; use the time point when the first operation and maintenance sub-region starts to execute the first sample work task as the first execution monitoring time point; use the time point when the first operation and maintenance sub-region ends to execute the first sample work task as the second execution monitoring time point; calculate the numerical difference between the first execution monitoring time point and the second execution monitoring time point to obtain the task execution duration; calculate the average CPU usage rate, the average memory utilization rate, and the task execution duration to obtain the first task monitoring coefficient; the calculation of the first task monitoring coefficient is specifically configured with the following formula:
[0040] ; where, Rwj1 is the first task monitoring coefficient, Cpu is the average CPU usage rate, Ncl is the average memory utilization rate, and Rzx is the task execution duration; the second sample work task to the i-th sample work task are respectively monitored for work, and the second task monitoring coefficient to the i-th task monitoring coefficient are obtained; the first task monitoring coefficient to the i-th task monitoring coefficient are averaged to obtain the task completion index corresponding to the first operation and maintenance sub-region; the equipment failure coefficient and the task completion index corresponding to the first operation and maintenance sub-region are defined as the first region operation and maintenance data; the second to the m-th operation and maintenance sub-region corresponding region operation and maintenance data are respectively obtained to obtain the second region operation and maintenance data to the m-th region operation and maintenance data; the first region operation and maintenance data to the m-th region operation and maintenance data are defined as the equipment operation and maintenance data; the data acquisition module acquires the equipment operation and maintenance data and transports it to the data analysis module; the data analysis module analyzes the equipment operation and maintenance data, calculates the region operation and maintenance warning coefficient corresponding to each operation and maintenance sub-region, obtains the region operation and maintenance warning coefficient threshold, compares the numerical value with the region operation and maintenance warning coefficient, and divides the operation and maintenance sub-region into the first type of operation and maintenance sub-region and the second type of operation and maintenance sub-region according to the numerical comparison result to obtain the operation and maintenance region division data; acquire the equipment operation and maintenance data, and respectively obtain the first region operation and maintenance data to the m-th region operation and maintenance data according to the equipment operation and maintenance data; perform warning analysis on the first operation and maintenance sub-region according to the first region operation and maintenance data to obtain the region operation and maintenance warning coefficient corresponding to the first operation and maintenance sub-region; specifically as follows: respectively obtain the equipment failure coefficient and the task completion index according to the first region operation and maintenance data; calculate the region operation and maintenance warning coefficient from the equipment failure coefficient and the task completion index; calculate the region operation and maintenance warning coefficient, and the specific formula configuration is as follows:
[0041] ; where, Qyy is the region operation and maintenance warning coefficient, Sbg is the equipment failure coefficient, and Rwz is the task completion index; respectively obtain the region operation and maintenance warning coefficients corresponding to the second region operation and maintenance data to the m-th region operation and maintenance data; obtain the region operation and maintenance warning coefficient threshold, compare the numerical value with the region operation and maintenance warning coefficient, and divide the operation and maintenance sub-region into the first type of operation and maintenance sub-region and the second type of operation and maintenance sub-region according to the numerical comparison result to obtain the operation and maintenance region division data; specifically as follows: respectively obtain the equipment failure coefficient threshold and the task completion index threshold, and calculate the region operation and maintenance warning coefficient threshold from the equipment failure coefficient threshold and the task completion index threshold.
[0042] It should be noted here that: in this application, the equipment failure coefficient threshold is the maximum equipment failure coefficient corresponding to the operation and maintenance sub-region in the normal operation state, and the task completion index threshold is the minimum task completion index corresponding to the operation and maintenance sub-region in the normal operation state; calculate the region operation and maintenance warning coefficient threshold, and the specific formula configuration is as follows:
[0043] ; where Qyy1 is the threshold of the regional operation and maintenance warning coefficient, Sbg1 is the threshold of the equipment failure coefficient, and Rwz1 is the threshold of the task completion index; when the regional operation and maintenance warning coefficient is greater than or equal to the threshold of the regional operation and maintenance warning coefficient, it is determined that the corresponding operation and maintenance sub-region is the first type of operation and maintenance sub-region; when the regional operation and maintenance warning coefficient is less than the threshold of the regional operation and maintenance warning coefficient, it is determined that the corresponding operation and maintenance sub-region is the second type of operation and maintenance sub-region.
[0044] It should be noted here that: in this application, the first type of operation and maintenance sub-region is in a normal operation state, and the second type of operation and maintenance sub-region is in an abnormal operation state; the data analysis module obtains the operation and maintenance area division data and transports it to the risk warning module; the risk warning module performs operation risk warning on the target data center according to the operation and maintenance area division data; obtain the operation and maintenance area division data; when the operation and maintenance sub-region is the first type of operation and maintenance sub-region, risk monitoring is normally carried out on the operation and maintenance sub-region; when the operation and maintenance sub-region is the second type of operation and maintenance sub-region, an operation risk warning is issued for the operation and maintenance sub-region; obtain the numerical value corresponding to the second type of operation and maintenance sub-region in the target data center to obtain the number value of the abnormal operation and maintenance area, calculate the ratio of the number value of the abnormal operation and maintenance area to m to obtain the abnormal area ratio; obtain the abnormal area ratio threshold, and compare the abnormal area ratio with the abnormal area ratio threshold; specifically as follows: when the abnormal area ratio is greater than or equal to the abnormal area ratio threshold, an operation risk warning is issued for the target data center; when the abnormal area ratio is less than the abnormal area ratio threshold, operation risk monitoring is normally carried out on the target data center.
[0045] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to only the specific implementation manners. Obviously, according to the content of this specification, many modifications and changes can be made. These embodiments are selected and specifically described in this specification to better explain the principle and practical application of the present invention, so that those skilled in the art in the technical field can well understand and utilize the present invention.
Claims
1. An operation and maintenance data management method based on big data analysis, characterized in that: include: Step S1: Divide the target data center into multiple operation and maintenance sub-areas, obtain the equipment failure coefficient corresponding to each operation and maintenance sub-area, obtain the task detection coefficient corresponding to each operation and maintenance sub-area, calculate the mean of the task detection coefficients to obtain the task completion index, and obtain the equipment operation and maintenance data; Calculate the equipment failure coefficient corresponding to the first operation and maintenance sub-area. The specific formula configuration is as follows: Among them, Sbg is the equipment failure coefficient corresponding to the first operation and maintenance sub-area, Tb1 to Tbn are the first reported fault value and the nth reported fault value respectively, and Gzj is the fault monitoring time length value; Calculate the first task monitoring coefficient, the specific formula configuration is as follows: Among them, Rwj1 is the first task monitoring coefficient, Cpu is the average CPU usage, Ncl is the average memory utilization, and Rzx is the task execution time; Step S2: Analyze the equipment operation and maintenance data, calculate the regional operation and maintenance warning coefficient corresponding to each operation and maintenance sub-area, obtain the regional operation and maintenance warning coefficient threshold and compare the values with the regional operation and maintenance warning coefficient, and divide the operation and maintenance sub-area into a first type of operation and maintenance sub-area and a second type of operation and maintenance sub-area according to the numerical comparison result, and obtain the operation and maintenance area division data; Calculate the regional operation and maintenance warning coefficient. The specific formula configuration is as follows: Among them, Qyy is the regional operation and maintenance warning coefficient, Sbg is the equipment failure coefficient, and Rwz is the task completion index; Step S3: providing operation risk warnings for the operation and maintenance sub-areas and the target data center respectively according to the operation and maintenance area division data; The step S3 further includes the following specific steps: Step S31: Acquire operation and maintenance area division data; Step S32: When the operation and maintenance sub-area is a first type of operation and maintenance sub-area, risk monitoring is normally performed on the operation and maintenance sub-area; Step S33: When the operation and maintenance sub-area is a second type of operation and maintenance sub-area, an operation and maintenance risk warning is issued for the operation and maintenance sub-area; Step S34: obtaining the number value corresponding to the second type of operation and maintenance sub-area in the target data center, obtaining the number value of abnormal operation and maintenance areas, calculating the ratio of the number value of abnormal operation and maintenance areas to m, and obtaining the number ratio of abnormal areas; Step S35: obtaining an abnormal region quantity ratio threshold, and comparing the abnormal region quantity ratio with the abnormal region quantity ratio threshold; The details are as follows: Step S351: when the abnormal area quantity ratio is greater than or equal to the abnormal area quantity ratio threshold, an operation risk warning is issued to the target data center; Step S352: When the abnormal area quantity ratio is less than the abnormal area quantity ratio threshold, the target data center is normally monitored for operation risks.
2. The operation and maintenance data management method based on big data analysis according to claim 1 is characterized in that: The step S1 further includes the following specific steps: Step S11: Divide the target data center into m operation and maintenance equipment sub-areas with the same number of servers, and name them the first operation and maintenance sub-area to the mth operation and maintenance sub-area respectively; Step S12: monitoring the operation and maintenance equipment failures in the first operation and maintenance sub-area to obtain the equipment failure coefficient corresponding to the first operation and maintenance sub-area; Step S13: Acquire the task completion index corresponding to the first operation and maintenance sub-area; Step S14: defining the equipment failure coefficient and task completion index corresponding to the first operation and maintenance sub-area as the first area operation and maintenance data; Step S15: acquiring the regional operation and maintenance data corresponding to the second to m-th operation and maintenance sub-regions respectively, to obtain the second regional operation and maintenance data to the m-th regional operation and maintenance data; Step S16: define the first region operation and maintenance data to the mth region operation and maintenance data as equipment operation and maintenance data.
3. The operation and maintenance data management method based on big data analysis according to claim 2 is characterized in that: The step S12 further includes the following specific steps: Step S121: in the process of fault detection for the first operation and maintenance sub-area, a time value corresponding to the current time is obtained as a first reference time point, a time point corresponding to a characteristic monitoring duration before the first reference time is taken as a second reference time point, and a period between the first reference time point and the second reference time point is defined as a fault monitoring period; Step S122: acquiring the time length value corresponding to the fault monitoring period to obtain the fault monitoring time length value; Step S123: Select n operation and maintenance hosts in the first operation and maintenance sub-area as test hosts, and name them as the first test host to the nth test host respectively; Step S124: respectively obtaining the number of fault times reported by the first test host to the nth test host during the fault monitoring period, and obtaining a first reported fault value and an nth reported fault value; Step S125: The first reported fault value, the nth reported fault value and the fault monitoring time length value are calculated to obtain the equipment fault coefficient corresponding to the first operation and maintenance sub-area.
4. The operation and maintenance data management method based on big data analysis according to claim 2 is characterized in that: The step S13 further includes the following specific steps: Step S131: obtaining a work task list corresponding to the first operation and maintenance sub-area, selecting i work tasks from the work task list as sample work tasks, and naming them as the first sample work task to the i-th sample work task respectively; Step S132: Perform work detection on the first sample work task to obtain a first task monitoring coefficient; Step S133: performing work monitoring on the second sample work task to the i-th sample work task respectively, and obtaining the second task monitoring coefficient to the i-th task monitoring coefficient; Step S134: Calculate the average of the first task monitoring coefficient to the i-th task monitoring coefficient to obtain the task completion index corresponding to the first operation and maintenance sub-area.
5. The operation and maintenance data management method based on big data analysis according to claim 4 is characterized in that: The step S132 further includes the following specific steps: Step S1321: during the process of obtaining the first sample work task, obtaining the average CPU usage rate corresponding to the operation and maintenance host executing the first sample work task; Step S1322: during the process of obtaining the first sample work task, obtaining the average memory utilization rate corresponding to the operation and maintenance host executing the first sample work task; Step S1323: taking the time point when the first operation and maintenance sub-area starts to execute the first sample work task as the first execution monitoring time point; Step S1324: taking the time point when the first operation and maintenance sub-area finishes executing the first sample work task as the second execution monitoring time point; Step S1325: Calculate the numerical difference between the first execution monitoring time point and the second execution monitoring time point to obtain the task execution duration; Step S1326: Calculate the average CPU usage, average memory utilization, and task execution time to obtain a first task monitoring coefficient.
6. The operation and maintenance data management method based on big data analysis according to claim 1 is characterized in that: The step S2 further comprises the following specific steps: Step S21: Acquire equipment operation and maintenance data, and acquire first area operation and maintenance data to mth area operation and maintenance data respectively according to the equipment operation and maintenance data; Step S22: performing early warning analysis on the first operation and maintenance sub-region according to the first regional operation and maintenance data, and obtaining a regional operation and maintenance early warning coefficient corresponding to the first operation and maintenance sub-region; Step S23: respectively obtaining the regional operation and maintenance warning coefficients corresponding to the second regional operation and maintenance data to the mth regional operation and maintenance data; Step S24: obtaining a regional operation and maintenance warning coefficient threshold and performing numerical comparison with the regional operation and maintenance warning coefficient, dividing the operation and maintenance sub-region into a first type of operation and maintenance sub-region and a second type of operation and maintenance sub-region according to the numerical comparison result, and obtaining operation and maintenance region division data.
7. The operation and maintenance data management method based on big data analysis according to claim 6 is characterized in that: The step S22 further includes the following specific steps: Step S221: acquiring the equipment failure coefficient and the task completion index respectively according to the first area operation and maintenance data; Step S222: Calculate the equipment failure coefficient and the task completion index to obtain the regional operation and maintenance warning coefficient.
8. The operation and maintenance data management method based on big data analysis according to claim 6 is characterized in that: The step S24 further includes the following specific steps: Step S241: respectively obtaining a device failure coefficient threshold and a task completion index threshold, and calculating the device failure coefficient threshold and the task completion index threshold to obtain a regional operation and maintenance warning coefficient threshold; Calculate the regional operation and maintenance warning coefficient threshold. The specific formula configuration is as follows: Among them, Qyy1 is the regional operation and maintenance warning coefficient threshold, Sbg1 is the equipment failure coefficient threshold, and Rwz1 is the task completion index threshold; Step S242: when the regional operation and maintenance warning coefficient is greater than or equal to the regional operation and maintenance warning coefficient threshold, the corresponding operation and maintenance sub-region is determined to be a first type of operation and maintenance sub-region; Step S243: When the regional operation and maintenance warning coefficient is less than the regional operation and maintenance warning coefficient threshold, it is determined that the corresponding operation and maintenance sub-region is a second type of operation and maintenance sub-region.
9. An operation and maintenance data management system based on big data analysis, applicable to an operation and maintenance data management method based on big data analysis as claimed in any one of claims 1 to 8, characterized in that: The specific working process of each module of the management system is as follows: Data acquisition module: used to divide the target data center into multiple operation and maintenance sub-areas, obtain the equipment failure coefficient and task completion index corresponding to each operation and maintenance sub-area, and obtain equipment operation and maintenance data; Data analysis module: used to divide the operation and maintenance sub-areas into first-type operation and maintenance sub-areas and second-type operation and maintenance sub-areas by analyzing the equipment operation and maintenance data, and obtain operation and maintenance area division data; Risk warning module: used to issue operational risk warnings to the operation and maintenance sub-areas and target data centers based on the operation and maintenance area division data.
Citation Information
Patent Citations
Aircraft operation and maintenance management system based on data analysis
CN116664103A
Public energy consumption equipment operation and maintenance management system suitable for smart park
CN118096132A