An IT operation and maintenance management method based on Internet of Things
By acquiring alarm signals and log datasets through the Internet of Things, and combining the trends of normal and fault data, the correlation between devices can be determined, maintenance reports can be generated, and maintenance personnel can be screened. This solves the problem of frequent alarms in existing technologies, improves the efficiency and accuracy of IT operation and maintenance management, and reduces costs.
Patent Information
- Application Number
- CN202310756743.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-25
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2043-06-25
AI Technical Summary
In existing technologies, equipment malfunctions are determined by monitoring a single data point, which leads to frequent alarms, increases manpower and material costs, and fails to address the problem at its root.
By acquiring alarm signals and log datasets through the Internet of Things, fault analysis is performed. By combining the changing trends of normal and fault data, the correlation between devices is determined, maintenance reports are generated, and maintenance personnel are selected.
It improves the efficiency of IT operations and maintenance management, accurately identifies the cause of failures, reduces maintenance frequency and costs, and reduces unnecessary maintenance work.
Smart Images

Figure CN116708156B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of IT operation and maintenance management, and in particular to an IT operation and maintenance management method based on the Internet of Things. BACKGROUND
[0002] The Internet of Things is to connect various devices with each other through the Internet, so that the devices can better communicate and work together. Through the Internet of Things technology, the device information collected by various devices is transmitted to the IT device, and the received data information is processed by the IT device to determine whether the device has a fault and to allocate personnel to the operation and maintenance personnel.
[0003] In the related art, the sensors monitor the data of each device, and compare the monitored data with the standard data. When the monitored data exceeds the standard data, the system generates an alarm signal, and notifies the relevant operation and maintenance personnel to operate and maintain the corresponding fault device according to the alarm signal.
[0004] In the related art, when the device is alarmed, only the single data is monitored to determine the fault, and the joint analysis of multiple data is not performed to determine the alarm problem caused by the mutual influence between multiple devices. When maintaining, although the alarm is solved, the problem is not solved from the root, which leads to frequent alarms, increases the cost of manpower and material resources, and there is room for improvement. SUMMARY
[0005] In order to improve the IT operation and maintenance management efficiency, the present application provides an IT operation and maintenance management method based on the Internet of Things.
[0006] The present application provides an IT operation and maintenance management method based on the Internet of Things, which adopts the following technical scheme:
[0007] An IT operation and maintenance management method based on the Internet of Things, comprising:
[0008] An alarm signal is obtained and transmitted to an IT operation and maintenance platform through the Internet of Things, and the IT operation and maintenance platform receives the alarm signal and reads the corresponding log data set; the log data set includes normal data change information and fault data change information;
[0009] Based on the fault data change information, the fault device is analyzed to obtain a first analysis result; the first analysis result is a judgment result of the fault reason of the fault device according to the data change trend contained in the fault data change information;
[0010] The normal data change information and the fault data change information are processed to obtain the corresponding normal data change trend and fault data change trend, and the normal data change trend and the fault data change trend are matched in similarity.
[0011] If the similarity is greater than the preset similarity, it is determined that there is a correlation between the data changes of the normal device and the faulty device, and a second analysis result is output; the correlation is the influence of the normal device on the data changes of the faulty device during operation;
[0012] The first analysis result and the second analysis result are comprehensively judged to determine the influence factors of the first analysis result and the second analysis result on the faulty device under different conditions; the influence factors include external factors and internal factors, the external factors are factors leading to failure due to non-faulty device itself wear and tear, and the internal factors are factors leading to failure due to the faulty device itself wear and tear;
[0013] Based on the external factors or the internal factors, a report is generated to obtain a maintenance report under different influence factors, and maintenance personnel are screened and dispatched according to the maintenance report.
[0014] Preferably, the fault data change information is analyzed to obtain sorted fault change statistical data, and the first analysis result is obtained by analysis; the first analysis result includes external influence factors and internal influence factors;
[0015] If the fault change statistical data contains sudden change data, it is determined that the failure is caused by non-self reasons, and the external influence factors are output;
[0016] If the data in the fault change statistical data gradually changes over time and the change amplitude is small until the data exceeds the preset operation range, it is determined that the failure is caused by the device itself wear and tear, and the internal influence factors are output.
[0017] Preferably, the normal data change information is sorted and analyzed to obtain sorted normal change statistical data;
[0018] The normal change statistical data and the fault change statistical data are analyzed for data change trend, when the data change trend in the normal change statistical data changes, the data change trend in the fault change statistical data also changes, and the changed data size is regular, it is determined that there is a correlation between the normal change statistical data and the fault change statistical data;
[0019] The correlation also includes dividing the normal change statistical data according to the time information contained in the alarm signal to obtain a normal device whose data changes dramatically before the failure occurs and a normal device whose data changes dramatically after the failure occurs, and corresponding output of pre-fault statistical data and post-fault statistical data;
[0020] The pre-failure statistical data is that the normal device contained in the pre-failure statistical data has an influence on the running of the fault device; and the post-failure statistical data is that the fault device has an influence on the normal device in the post-failure statistical data.
[0021] Preferably, the first analysis result and the second analysis result are comprehensively judged, if the first analysis result is an internal influence factor and the second analysis result is that there is no correlation between the fault device and the normal device, the judgment result is consistent, and it is determined that the influence factor of the device is an internal factor;
[0022] If the first analysis result is an internal influence factor and the second analysis result is that there is a correlation between the fault device and the normal device, the judgment result is inconsistent, and data matching is performed;
[0023] If the matching result of the data matching is matching success, it indicates that the influence factor of the fault device is the influence of the normal device on the fault factor; if the matching result of the data matching is matching failure, it indicates that it is an internal factor; if the result of the data matching is matching success, the normal device which produces the influence is re-analyzed until the final root cause of the stability of the device in the Internet of Things is determined;
[0024] If the first analysis result is an external influence factor and the second analysis result is that there is no correlation between the fault device and the normal device, the judgment result is inconsistent, and it is determined that the influence factor of the fault device is the mutual influence between the devices, i.e. an external environmental influence factor;
[0025] If the first analysis result is an external influence factor and the second analysis result is that there is a correlation between the fault device and the normal device, the judgment result is consistent, and the normal device with the correlation is subjected to data change analysis and screening to screen out the normal device which produces sudden change data to the fault device.
[0026] Preferably, the data matching is to judge the data influence of the data change of the normal device on the fault device when the device does not fail according to the correlation, so as to obtain the change relationship between the normal device and the fault device, and calculate the change of the fault data in the first analysis result according to the change relationship, so as to determine whether the normal device with the correlation can be consistent with the fault data of the fault device, if yes, the matching is successful, otherwise the matching is failed.
[0027] Preferably, the judgment of the change relationship between the normal device and the faulty device specifically comprises obtaining data changes of the normal device and the faulty device and drawing corresponding data change trend charts with time lines as axes, comparing the data change trend charts to determine the similarity of the data change trends, if the similarity is high, analyzing the data change trends, analyzing the data change of the faulty device when the data change of the normal device, determining the data influence relationship of the normal device on the faulty device, and outputting the change relationship between the normal device and the faulty device.
[0028] The similarity judgment comprises comparison of data change trend charts with the same time starting point and comparison of data change trend charts with different time starting points.
[0029] Preferably, an alarm signal is obtained, and current fault data is determined according to the alarm signal.
[0030] A maintenance report is read, current fault data is matched with maintenance data in the maintenance report, and a matching result is output.
[0031] If the matching is successful, time information of the maintenance data is read according to the matching result, and a maintenance time interval of the maintenance data is obtained.
[0032] Time information of the current fault data and time information of adjacent maintenance data are read according to the matching result to determine a to-be-maintained time interval.
[0033] The maintenance time interval is matched with the to-be-maintained time interval, if the to-be-maintained time interval is greater than the maintenance time interval, it is determined that the normal equipment is faulty, and maintenance personnel are dispatched according to the fault information.
[0034] If the to-be-maintained time interval is less than the maintenance time interval, it is determined that the abnormal equipment is faulty, and fault analysis is performed on the abnormal equipment fault.
[0035] Preferably, different maintenance reports are generated according to different influence factors.
[0036] The maintenance report is judged to obtain a maintenance type of the maintenance fault, and keywords of the maintenance type are extracted.
[0037] The extracted keywords are used to screen maintenance personnel who have no current maintenance task to obtain a first screening result.
[0038] The first screening result is sorted according to work experience, maintenance accuracy, maintenance time consumption, and maintenance use time length respectively to obtain sorting results of each aspect of each maintenance personnel.
[0039] The ranking results of the same maintenance personnel are counted and score calculation is performed; the score calculation is to take the total number of personnel in the first screening result as the maximum value of the score, and to perform reverse distribution according to the ranking;
[0040] The scores of the same maintenance personnel are counted to obtain the personal total score, and the personal total scores are compared to screen out the best maintenance personnel.
[0041] Preferably, after the maintenance is performed, the method further comprises:
[0042] The device after the operation and maintenance is subjected to data memory and active debugging; the active debugging is used to verify the fault result of the judgment;
[0043] The data of the normal device are adjusted, and the data of the original fault device and other devices are monitored; the original fault device is the normal device after the maintenance of the fault device;
[0044] If the data of other devices are changed when the data of the normal device are adjusted, the modified data of other devices are corrected;
[0045] According to the data changes of the normal device and the original fault device, the change relationship between the normal device and the original fault device is determined, and the change relationship is compared with the previous change relationship to determine whether the change relationship is consistent; if the change relationship is consistent, it indicates that the judgment is accurate; if the change relationship is inconsistent, it indicates that the fault reasons of the original fault device have mutual influence, and the mutual influence is analyzed next time when a fault occurs.
[0046] Preferably, when a device fault occurs each time, the correlation between the normal device and the fault device is judged, and the normal device is divided into pre-fault statistical data and post-fault statistical data to determine the mutual influence relationship between each device in the IT operation and maintenance equipment, and gradually build and improve the influence relationship network.
[0047] In summary, the present application has at least one of the following beneficial technical effects:
[0048] 1. By acquiring alarm signals and reading the log dataset at the corresponding time point based on the alarm signals, the log dataset includes normal data change information for normal devices and fault data change information for faulty devices. Fault data change information is then analyzed to determine whether the fault originates from the device itself, and a first analysis result is output. Next, the fault data change information and normal data change information are processed for data change trends to obtain corresponding normal data change trends and fault data change trends. Similarity matching is then performed on the normal data change trends and fault data change trends. If the matching result shows a high similarity, a correlation in data changes is determined, and a second analysis result is output. The first and second analysis results are then combined. If the cause of the failure is determined, the influencing factors under different conditions, including internal and external factors, can be obtained. Different maintenance reports can be generated based on different influencing factors, and different maintenance personnel can be matched and screened according to different maintenance reports. A preliminary judgment can be made by analyzing the failure data change information of the failure device itself to determine whether the failure is caused by the device itself or by external influences. Then, a second analysis is carried out in conjunction with the normal data change trend to further determine the specific influencing factors of the failure. This allows for a better determination of the cause of the failure device, which provides convenience for maintenance, optimizes maintenance results, reduces the occurrence of failures, reduces maintenance costs, and improves operation and maintenance management efficiency.
[0049] 2. Based on the different states of the first and second analysis results, the influencing factors of the fault are judged. When the first analysis result is an internal influencing factor, if the second analysis result is not related, it is judged that the faulty equipment operates independently and the fault factor is an internal factor. If the second analysis result is related, it is judged that the faulty device is affected by other normal devices. The correlation between the normal devices is then used to determine whether they will cause the faulty device to malfunction, and the analysis results determine whether the fault is caused by a normal device or is an internal factor of the faulty equipment. When the first analysis result is an external influencing factor, if the second analysis result is not related, it is judged that the faulty device operates independently and is affected by external factors. The impact caused a sudden change in operating data, so it was determined to be an external environmental factor. If the second analysis result is correlated, it is determined that the influencing factors of the faulty device include external environmental factors and sudden changes in the normal device during operation. The normal devices with correlation are analyzed to determine whether the faulty device is caused by the influence of the normal device. If it is caused by the normal device, the sudden change in the normal device is re-analyzed to determine the final influencing factor. Through the comprehensive judgment of multiple influencing factors, the accuracy of the judgment is improved, the cause of the fault is found in the root and it is maintained, the fundamental problem of the fault is solved, the maintenance frequency of the load is reduced, and the maintenance cost is reduced.
[0050] 3. By comprehensively utilizing trend processing of normal and fault change information, corresponding normal and fault data change trend charts are obtained. Then, the similarity of these two trend charts is compared, including comparisons at the same time starting point and comparisons at different time starting points, to determine the corresponding similarity and comparison time starting point. Based on the similarity and comparison time starting point, the data change relationship is judged to determine the corresponding data change amount of the faulty device when the data of the normal device changes. This provides a basis for judging the impact of the normal device on the faulty device, making the judgment of the influencing factors of the faulty device more accurate and improving the accuracy of the IT operation and maintenance platform in judging the influencing factors of faults. Attached Figure Description
[0051] Figure 1 This embodiment mainly illustrates the steps of an IT operations and maintenance management method based on the Internet of Things;
[0052] Figure 2 This embodiment mainly illustrates the sub-step flowchart of S200 in the IT operation and maintenance management method based on the Internet of Things;
[0053] Figure 3 This embodiment mainly illustrates the sub-step flowchart of S300 in the IT operation and maintenance management method based on the Internet of Things;
[0054] Figure 4 This embodiment mainly illustrates the sub-step flowchart of S500 in the IT operation and maintenance management method based on the Internet of Things;
[0055] Figure 5 This embodiment mainly illustrates the sub-step flowchart of S100 in the IT operation and maintenance management method based on the Internet of Things;
[0056] Figure 6 This embodiment mainly illustrates the sub-step flowchart of S600 in the IT operation and maintenance management method based on the Internet of Things. Detailed Implementation
[0057] The following is in conjunction with the appendix Figures 1-6 This application will be described in further detail.
[0058] This application discloses an IT operation and maintenance management method based on the Internet of Things.
[0059] Example: Figure 1 As shown, the present invention provides an IT operations and maintenance management method based on the Internet of Things, comprising:
[0060] S100: An alarm signal is acquired and transmitted to the IT operations and maintenance platform via the Internet of Things. The IT operations and maintenance platform receives the alarm signal and reads the corresponding log dataset. The log dataset includes normal data change information and fault data change information. The IT operations and maintenance platform receives the alarm signal and acquires the time information of the alarm signal. Based on the time information, it acquires the log dataset of the device from the last device maintenance to the current time. The devices managed by the IT operations and maintenance platform include faulty devices that have been detected and normal devices that are in normal operating condition.
[0061] S200, perform fault analysis on the faulty device based on fault data change information to obtain a first analysis result; the first analysis result is a judgment result on the cause of the faulty device based solely on the data change trend contained in the fault data change information;
[0062] S300 performs data change trend processing on normal data change information and fault data change information to obtain the corresponding normal data change trend and fault data change trend, and performs similarity matching on the normal data change trend and fault data change trend.
[0063] S400, if the similarity is greater than the preset similarity, it is determined that there is a correlation between the data changes of the normal device and the faulty device, and the second analysis result is output; the correlation is the influence of the normal device on the data changes of the faulty device during operation.
[0064] S500, a comprehensive judgment is made on the first analysis result and the second analysis result to determine the influencing factors of the first analysis result and the second analysis result on the faulty device under different conditions; the influencing factors include external factors and internal factors, the external factors are factors that cause the fault due to wear and tear of the faulty device itself, and the internal factors are factors that cause the fault due to wear and tear of the faulty device itself.
[0065] S600 generates reports based on external or internal factors to obtain maintenance reports under different influencing factors, and selects and dispatches maintenance personnel based on the maintenance reports.
[0066] refer to Figure 2 In step S200, fault analysis is performed on the faulty device based on the fault data change information to obtain the first analysis result, including the following steps:
[0067] S210, organize and analyze the fault data change information to obtain organized fault change statistical data, and analyze to obtain the first analysis result; the first analysis result includes external influencing factors and internal influencing factors; by organizing the fault data change information, the various data in the fault data change information are arranged in a regular manner according to the time sequence, so that the organized fault change statistical data shows the trend and regulation of data change, which is more conducive to data analysis and judgment.
[0068] S220, if the fault change statistics include sudden changes in data, it is determined that the fault is not caused by its own reasons, and the external influencing factors are output. During the normal operation of the device, the operating data of the device tends to be stable and unchanged without the intervention of external influencing factors. Only when affected by factors other than its own reasons will the data of the device in stable operation change suddenly, thus causing the device to malfunction. Therefore, when performing maintenance, it is necessary not only to maintain the faulty device itself, but also to eliminate the external influencing factors that affect the stable operation of the fault.
[0069] S230, if the data in the fault change statistics gradually changes over time with small fluctuations, until the data exceeds the preset operating range, it is determined that the fault is caused by the device's own wear and tear, and the aforementioned internal influencing factors are output. During the stable operation of the device, with prolonged use, the device itself experiences wear and tear, causing it to reach its lifespan during operation, thus requiring replacement and maintenance.
[0070] In this embodiment, a preliminary judgment is made by observing the changes in the operating data of the faulty device itself. When the changes in the operating data are relatively gradual, it is determined that the device fault is caused by its own wear and tear, and maintenance can be performed on the faulty device. When the changes in the operating data are sudden, it is determined that the device fault is caused by external influencing factors. Therefore, maintenance is required not only to maintain the faulty device itself, but also to eliminate the influencing factors of the faulty device. Thus, further judgment on the influencing factors of the fault is required.
[0071] refer to Figure 3 In step S300, data change trend processing is performed on the normal data change information and the fault data change information to obtain the corresponding normal data change trend and fault data change trend, and similarity matching is performed on the normal data change trend and the fault data change trend, including the following steps:
[0072] S310, Organize and analyze the normal data change information to obtain the organized normal change statistics;
[0073] S320: Perform data trend analysis on the normal change statistics and fault change statistics. When the data trend in the normal change statistics changes, the data trend in the fault change statistics also changes accordingly, and the magnitude of the change shows a regularity. Then, it is determined that there is a correlation between the normal change statistics and the fault change statistics. By trend-oriented processing of both the normal change statistics and the fault change statistics, corresponding data trend graphs are generated. By comparing the data in the data trend graphs, it is determined whether there is a similarity between them. If there is a similarity, the data is compared to determine whether there is a certain regularity in the magnitude of the changes in the data in the normal change statistics and the fault change statistics. This indicates that there is a correlation between the normal change statistics and the fault change statistics.
[0074] S330, the correlation also includes dividing the normal change statistics data into devices based on the time information contained in the alarm signal, so as to obtain the normal devices whose data changed drastically before the fault occurred and the normal devices whose data changed drastically after the fault occurred, and output the pre-fault statistics data and post-fault statistics data accordingly; by dividing the normal devices before the fault and the normal devices after the fault with the fault time as the time node, the mutual influence between normal devices and faulty devices can be further judged. By recording the mutual influence between the devices included in the equipment managed by the IT operation and maintenance management platform, the influence network between devices can be gradually built, so that when analyzing the influencing factors of the fault, the influencing factors of the faulty device can be analyzed more quickly and conveniently, and whether there is mutual influence can be determined.
[0075] S340, the pre-fault statistical data includes the normal devices included in the pre-fault statistical data that will affect the operation of the faulty device; the post-fault statistical data includes the faulty device that will affect the normal devices included in the post-fault statistical data.
[0076] In this embodiment, trend analysis is performed on normal change statistics and fault change statistics. When the normal change statistics change, the fault change statistics also change at the same time or after a certain time. If the change trends of the two are similar and there is a certain regularity, it is determined that they are related. At the same time, the relationship between normal devices and faulty devices is divided to determine the mutual influence between normal devices and faulty devices, so as to construct an influence network. This provides a fast and convenient method for subsequent fault analysis, speeds up the analysis results, and improves the efficiency of IT operation and maintenance management.
[0077] refer to Figure 4In step S500, a comprehensive judgment is made on the first analysis result and the second analysis result to determine the influencing factors of the first analysis result and the second analysis result on the faulty device under different conditions, including the following steps:
[0078] S510, the first analysis result and the second analysis result are comprehensively judged. If the first analysis result is an internal influencing factor and the second analysis result is that there is no correlation between the faulty device and the normal device, the judgment result is consistent, and the influencing factor of the device is determined to be an internal factor. When the first analysis result is an internal influencing factor, the second analysis result needs to be judged again to avoid the occurrence of special situations such as the normal device accelerating the wear and tear of the faulty device during operation. If the second analysis result is that there is no correlation between the faulty device and the normal device, it indicates that the faulty device is completely independent during operation and is not affected by other normal devices. Therefore, it is judged that the faulty device is due to normal damage during daily use. For example, there are two devices A and B, which are related to each other. The lifespan of both A and B is 0 to 9, where 9 represents brand new and 0 represents faulty. During operation, the lifespan of A gradually decreases by 1 each time. Under the influence of A, the lifespan of B decreases by 2 each time. Therefore, when A is changed, B is also gradually changed. So, if only B is maintained when B fails, the lifespan of B will not be changed. Therefore, it is necessary to eliminate the influence between A and B to extend the lifespan of B and reduce the cost of operation and maintenance.
[0079] S520, if the first analysis result is an internal influencing factor and the second analysis result is that there is a correlation between the faulty device and the normal device, then the judgment result is inconsistent and data matching is performed; if the second analysis result is that there is a correlation between the faulty device and the normal device based on step S510, then further judgment needs to be made on the normal device to determine whether the operation of the normal device caused the faulty device to fail.
[0080] S530, if the data matching result is a successful match, it indicates that the influencing factor of the faulty device is the influence of the normal device on the faulty factor; if the data matching result is a failed match, it indicates that it is an internal factor; if the data matching result is a successful match, the normal device that has an impact is re-analyzed until the root cause affecting the operational stability of the device in the Internet of Things is determined; data analysis is performed on the normal device and the faulty device through the judgment of the correlation to determine the corresponding data change trend map, and the similarity of the data change trend map of the faulty device and the data change trend map of the normal device is compared to determine the time starting point of the similarity of the data change trend maps, and the data is analyzed according to the time starting point to determine the data change relationship between the normal device and the faulty device, and the theoretical data of the faulty device is predicted according to the data information of the normal device and the data change relationship when the fault occurs, and the theoretical data is compared with the actual data. If the theoretical data is less than the actual data, it is judged as a failed match, that is, the influence of the normal device on the faulty device is insufficient to cause the faulty device to fail; if the theoretical data is greater than the actual data, it is judged as a successful match, that is, the influence of the normal device on the faulty device will cause the faulty device to fail.
[0081] S540, if the first analysis result is an external influencing factor and the second analysis result is that there is no correlation between the faulty device and the normal device, then the judgment result is inconsistent, and the influencing factor of the faulty device is determined to be the non-interaction between devices, i.e., external environmental factors; when the first analysis result is an external influencing factor, it indicates that the data in the faulty device has changed drastically, and the second analysis result needs to be judged to determine whether the fault is caused by the influence of the normal device on the faulty device. If the second analysis result is that there is no correlation between the faulty device and the normal device, then the cause of the distance change is attributed to external causes, such as collisions, dust, etc., which affect the faulty device.
[0082] S550, if the first analysis result indicates an external influencing factor, and the second analysis result indicates a correlation between the faulty device and the normal device, then the judgment result is consistent. Data change analysis is then performed on the normal devices with the correlation to filter out those that caused the sudden data change in the faulty device. By analyzing the data changes of the normal devices correlated with the faulty device in the second analysis result, it is determined which normal device's data change caused the sudden data change in the faulty device. Since the fault data change information of the faulty device includes sudden data change, the normal data change information of the normal device also includes sudden data change due to the data change relationship between the normal device and the faulty device. Based on the time relationship of the sudden data change, it is determined whether the fault was caused by that normal device.
[0083] In this embodiment, different influencing factors of the faulty device are determined by analyzing the first and second analysis results under different conditions. When the first analysis result indicates an internal factor and the second analysis result indicates no correlation, it means that the faulty device operates independently and is not affected by external environmental factors or the operation of other devices. Therefore, maintenance only needs to be performed on the faulty device itself. When the first analysis result indicates an internal factor and the second analysis result indicates a correlation, it means that the faulty device does not experience drastic fluctuations during daily operation and is due to wear and tear on the faulty device itself. However, since there is a correlation between the normal device and the faulty device, it is necessary to analyze the normal device with the correlation to determine whether the influence of the normal device on the faulty device will cause the faulty device to malfunction. If it can cause a malfunction, it is determined to be an external influencing factor; if it cannot cause a malfunction, it is determined to be an internal influencing factor. When the first analysis result indicates an external factor and the second analysis result indicates no correlation, it means that the faulty device is caused by an external factor. However, since the second analysis result indicates no correlation, it means that the faulty device operates independently, thus excluding the influence of other normal devices on the faulty device. Therefore, the influencing factor of the faulty device is determined to be an external environmental factor. If the first analysis result indicates an external factor and the second analysis result indicates a correlation, then data analysis needs to be performed on the normal devices with the correlation to determine whether there is a sudden change in data in the normal devices. If so, the sudden change in data is analyzed to determine whether the sudden change in data in the normal devices will cause the faulty device to malfunction. If it can cause a malfunction, the normal device is re-analyzed until the cause of the sudden change in data is finally determined. If it cannot cause a malfunction, it is determined to be an external environment.
[0084] refer to Figure 5 In step S100, the alarm signal is acquired and transmitted to the IT operations and maintenance platform via the Internet of Things. The IT operations and maintenance platform receives the alarm signal and reads the corresponding log dataset. The steps also include:
[0085] S110: Acquire alarm signal and determine current fault data based on alarm signal; IT operation and maintenance management platform receives alarm signal and begins to analyze alarm signal for further operation and maintenance.
[0086] S120: Read the maintenance report, match the current fault data with the maintenance data in the maintenance report, and output the matching result; by matching the determined fault data with the maintenance data in the historical maintenance report, the maintenance record of the same fault type is identified.
[0087] S130, if the match is successful, read the time information of the maintenance data according to the matching result and obtain the maintenance time interval of the maintenance data; determine the daily maintenance cycle according to the historically matched maintenance data, and use this maintenance cycle as the standard maintenance time interval for equipment failure.
[0088] S140: Based on the matching results, read the time information of the current fault data and the time information of adjacent maintenance data to determine the maintenance time interval; and calculate the current maintenance time interval by using the time of the current fault and the maintenance time point of the maintenance data that is closest to the current time.
[0089] S150: Match the maintenance interval with the time interval to be maintained. If the time interval to be maintained is greater than the maintenance interval, it is determined to be a normal equipment fault, and maintenance personnel are dispatched based on the fault information. The comparison of time intervals determines whether the current fault is a normal fault. A normal equipment fault is a fault caused by normal wear and tear of the device without external factors. A time interval to be maintained greater than the maintenance interval indicates that the service life of the faulty device is greater than its average service life, i.e., a fault caused by normal wear and tear.
[0090] S160. If the maintenance interval is less than the maintenance interval, it is determined to be an abnormal equipment failure, and a fault analysis is performed on the abnormal equipment failure.
[0091] In this embodiment, the current fault data is determined by analyzing the alarm signal, and the current fault data is matched with the maintenance data in the maintenance report to determine whether there are historical maintenance records in the maintenance report. If there are historical maintenance records, the maintenance time interval is determined based on the maintenance data and used as the standard maintenance time interval. The maintenance time interval to be maintained for the current fault data is calculated and compared with the maintenance time interval. If the maintenance time interval to be maintained is greater than the maintenance time interval, it indicates that the fault of the faulty device is a normal equipment fault; otherwise, it indicates that the fault of the faulty device is an abnormal equipment fault.
[0092] refer to Figure 6 In step S600, a report is generated based on external or internal factors to obtain maintenance reports under different influencing factors. Maintenance personnel are then screened and dispatched based on these reports, including the following steps:
[0093] S610 generates different maintenance reports based on different influencing factors; since the influencing factors are different, the generated maintenance reports are also different. The maintenance report includes the original alarm signal of the faulty device and the fault influencing factors obtained from the analysis of the alarm signal.
[0094] S620: The maintenance report is evaluated to determine the maintenance type of the maintenance fault, and keywords are extracted from the maintenance type.
[0095] S630: The extracted keywords are used to filter maintenance personnel who do not currently have maintenance tasks to obtain the first screening result;
[0096] S640 sorts the first screening results by work experience, maintenance accuracy, maintenance time and maintenance usage time to obtain the ranking results of each maintenance personnel in various aspects.
[0097] S650, the ranking results of the same maintenance personnel are statistically analyzed and scores are calculated; the score calculation is to take the total number of personnel in the first screening result as the maximum score, and then distribute it in reverse according to the ranking; for example, the first screening result has 10 people, namely A, B, C, etc., and the ranking of A is first, second, and third respectively, then the corresponding scores are 10 points, 9 points, and 8 points, thus obtaining the total score of A.
[0098] S660 calculates the scores of the same maintenance personnel to obtain an individual total score, and compares these scores to select the best maintenance personnel. After maintenance is completed, the IT operations and maintenance platform stores the data of the devices after maintenance and performs proactive debugging to verify the fault diagnosis results in the maintenance report. By changing the operating data of normal devices that affected the faulty device and monitoring the data of the faulty device, the platform determines the data change relationship between the normal devices and the faulty device, obtaining a new relationship. This new relationship is compared with the original relationship. If the relationship matches, the IT operations and maintenance platform's judgment is accurate; otherwise, it indicates mutual influence between multiple devices.
[0099] In this embodiment, different maintenance reports are generated under different influencing factors, and then the maintenance reports are analyzed to extract keywords. Maintenance personnel are then screened based on the keywords. The screened maintenance personnel are then ranked according to their work experience, maintenance accuracy, maintenance time, and maintenance usage duration. Based on the ranking results, a score is calculated to obtain the maximum individual total score. The best maintenance personnel are then selected based on the total score.
[0100] Compared to existing IoT-based IT operations and maintenance management methods, this invention improves the efficiency of IT operations and maintenance management.
[0101] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. An IT operations and maintenance management method based on the Internet of Things, characterized in that, include: The alarm signal is acquired and transmitted to the IT operations and maintenance platform via the Internet of Things. The IT operations and maintenance platform receives the alarm signal and reads the corresponding log dataset. The log dataset includes normal data change information and fault data change information; Fault analysis is performed on the faulty device based on the fault data change information to obtain the first analysis result; The first analysis result is a judgment on the cause of the fault in the faulty device based solely on the data change trend contained in the fault data change information; Data change trend processing is performed on normal data change information and fault data change information to obtain the corresponding normal data change trend and fault data change trend, and similarity matching is performed on the normal data change trend and fault data change trend. The information on normal data changes is organized and analyzed to obtain the organized statistical data on normal changes; By performing data change trend analysis on normal change statistics and fault change statistics, when the data change trend in normal change statistics changes, the data change trend in fault change statistics also changes accordingly, and the magnitude of the change shows a regularity, it is determined that there is a correlation between normal change statistics and fault change statistics. The correlation also includes dividing the normal change statistics data into devices based on the time information contained in the alarm signal, so as to obtain normal devices whose data changed drastically before the fault occurred and normal devices whose data changed drastically after the fault occurred, and outputting the pre-fault statistics data and post-fault statistics data accordingly. The pre-fault statistics include the impact of normal devices on the operation of faulty devices. The post-fault statistics show the impact of the faulty device on the normal devices in the post-fault statistics. If the similarity is greater than the preset similarity, it is determined that there is a correlation between the data changes between the normal device and the faulty device, and the second analysis result is output. The correlation refers to the impact of normal device operation on the data changes of faulty device during operation. A comprehensive judgment is made on the first and second analysis results to determine the influencing factors of the first and second analysis results on the faulty device under different conditions; The influencing factors include external factors and internal factors. The external factors are those caused by wear and tear of the faulty device itself, while the internal factors are those caused by wear and tear of the faulty device itself. Reports are generated based on external or internal factors to obtain maintenance reports under different influencing factors, and maintenance personnel are screened and dispatched based on the maintenance reports.
2. The IT operation and maintenance management method based on the Internet of Things according to claim 1, characterized in that: The step of performing fault analysis on the faulty device based on fault data change information to obtain a first analysis result specifically includes: The fault data change information is organized and analyzed to obtain organized fault change statistics, and the first analysis result is obtained; the first analysis result includes external influencing factors and internal influencing factors; If the fault change statistics include sudden changes, it is determined that the fault is not caused by its own reasons, and the external influencing factors are output. If the data in the fault change statistics gradually changes over time and the change is small, until the data exceeds the preset operating range, it is determined that the fault is caused by the wear and tear of the device itself, and the internal influencing factors are output.
3. The IT operation and maintenance management method based on the Internet of Things according to claim 2, characterized in that: The step of comprehensively judging the first and second analysis results to determine the influencing factors of the first and second analysis results on the faulty device under different conditions specifically includes: The results of the first and second analyses are combined for judgment. If the first analysis result is an internal influencing factor and the second analysis result is that there is no correlation between the faulty device and the normal device, then the judgment result is consistent, and the influencing factor of the device is determined to be an internal factor. If the first analysis result indicates an internal influencing factor, and the second analysis result indicates a correlation between the faulty device and the normal device, then the judgment result is inconsistent, and data matching is performed. If the data matching result is a successful match, it indicates that the influencing factor of the faulty device is the influence of the normal device on the faulty factor; if the data matching result is a failed match, it indicates that it is an internal factor; if the data matching result is a successful match, the normal device that has an impact will be re-analyzed until the root cause affecting the stability of the device operation in the Internet of Things is determined. If the first analysis result is an external influencing factor, and the second analysis result is that there is no correlation between the faulty device and the normal device, then the judgment result is inconsistent, and the influencing factor of the faulty device is determined to be the non-interaction between devices, that is, the external environmental influencing factor. If the first analysis result is an external influencing factor, and the second analysis result is that there is a correlation between the faulty device and the normal device, then the judgment result is consistent. Data change analysis is performed on the normal devices with correlation to screen out the normal devices that cause sudden changes in data to the faulty device.
4. The IT operation and maintenance management method based on the Internet of Things according to claim 3, characterized in that: The data matching process involves determining the impact of data changes in normal devices on the data of faulty devices when no equipment failure occurs, based on the correlation relationship, to obtain the relationship between the changes in normal devices and faulty devices. The changes in fault data in the first analysis result are then calculated based on the relationship to determine whether the normal devices with the correlation relationship can match the fault data of the faulty devices. If they match, the matching is successful; otherwise, the matching fails.
5. The IT operation and maintenance management method based on the Internet of Things according to claim 4, characterized in that: The determination of the relationship between changes in normal and faulty devices involves acquiring the data changes of normal and faulty devices and plotting the corresponding data change trend graphs with the time line as the axis. The data change trend graphs are then compared to determine the similarity of the data change trends. If the similarity is high, the data change trends are analyzed to analyze the data changes of the faulty device when the data of the normal device changes, and to determine the data influence relationship between the normal device and the faulty device. Finally, the relationship between changes in normal and faulty devices is output. The similarity assessment includes comparing data trend charts from the same starting point in time and comparing data trend charts from different starting points in time.
6. The IT operation and maintenance management method based on the Internet of Things according to claim 2, characterized in that: The steps of acquiring alarm signals and transmitting them to the IT operations and maintenance platform via the Internet of Things, and the IT operations and maintenance platform receiving alarm signals and reading the corresponding log dataset, further include: Acquire alarm signals and determine current fault data based on alarm signals; Read the maintenance report, match the current fault data with the maintenance data in the maintenance report, and output the matching results; If a match is successful, the time information of the maintenance data is read based on the matching result, and the maintenance time interval of the maintenance data is obtained; Based on the matching results, the time information of the current fault data and the time information of adjacent maintenance data are read to determine the maintenance interval. The maintenance interval is matched with the time interval to be maintained. If the time interval to be maintained is greater than the maintenance interval, it is determined to be a normal equipment failure, and maintenance personnel are dispatched according to the failure information. If the maintenance interval is less than the maintenance interval, it is determined to be an abnormal equipment failure, and a fault analysis is performed on the abnormal equipment failure.
7. The IT operation and maintenance management method based on the Internet of Things according to claim 1, characterized in that: The steps of generating reports based on external or internal factors to obtain maintenance reports under different influencing factors, and then screening and dispatching maintenance personnel based on the maintenance reports, specifically include: Different maintenance reports are generated based on different influencing factors; The maintenance report is evaluated to determine the maintenance type of the maintenance failure, and keywords are extracted from the maintenance type. The extracted keywords are used to filter maintenance personnel who do not currently have maintenance tasks to obtain the first screening results; The first screening results are sorted by work experience, maintenance accuracy, maintenance time, and maintenance usage duration to obtain the ranking results of each maintenance personnel in various aspects. The ranking results of the same maintenance personnel are statistically analyzed and scores are calculated; the score calculation is to take the total number of personnel in the first screening result as the maximum score, and then distribute the scores in reverse according to the ranking. The scores of the same maintenance personnel are tallied to obtain an individual total score, and the individual total scores are compared to select the best maintenance personnel.
8. The IT operation and maintenance management method based on the Internet of Things according to claim 7, characterized in that: After maintenance is completed, it also includes: After maintenance is completed, the device is stored in data and actively debugged; the active debugging is used to verify the fault results. The data of the identified normal devices is adjusted and changed, and the data of the original faulty device and the data change of other devices are monitored; the original faulty device is the normal device after the faulty device has been maintained. If the data of other devices also changes when the data of a normal device is adjusted, the modified data of the other devices shall be corrected. Based on the data changes of the normal device and the original faulty device, determine the relationship between the changes between the normal device and the original faulty device, and compare it with the previous relationship to determine whether the relationship is consistent. If the relationship is consistent, it indicates that the judgment is accurate. If the relationship is inconsistent, it indicates that the cause of the original faulty device has mutual influence, and the mutual influence will be analyzed in the next fault.
9. The IT operation and maintenance management method based on the Internet of Things according to claim 2, characterized in that: Each time a device failure occurs, the correlation between normal devices and failed devices is determined. By dividing normal devices into pre-failure statistical data and post-failure statistical data, the mutual influence relationships among the various devices contained in the IT operation and maintenance equipment are determined, and the influence relationship network is gradually built and improved.
Citation Information
Patent Citations
A method and system for early warning and optimization of equipment failure based on similarity curves.
CN102270271A
Fault positioning method, device and system and storage medium
CN114781510A