Method and device for measuring event processing efficiency, electronic equipment and storage medium

By obtaining the critical periods of failure events and performing weighted averages, the problem of lack of quantitative indicators in operation and maintenance management is solved, and a comprehensive and objective evaluation of the performance of operation and maintenance personnel and teams is achieved, which improves the scientificity and accuracy of operation and maintenance management.

CN120355279APending Publication Date: 2025-07-22BEIJING FOLLOW YOU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510262149.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the existing technology, the performance of operation and maintenance personnel and teams cannot be evaluated objectively and comprehensively in operation and maintenance management, especially in terms of technical familiarity, work efficiency and sense of responsibility.

Method used

By obtaining the critical period of the fault event, including fault response time, fault processing time, fault complete processing time and business recovery time, calculate its average value, obtain event processing efficiency indicators, and weighted averages based on factors such as the severity of the event, resource consumption cost and business priority, to achieve a quantitative evaluation of operation and maintenance performance.

Benefits of technology

A comprehensive and objective quantitative evaluation of the performance of operation and maintenance personnel and teams has been achieved, which can reflect technical proficiency and sense of responsibility, and improve the scientificity and accuracy of operation and maintenance management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355279A_ABST
    Figure CN120355279A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and device for measuring event processing efficiency, electronic equipment and a storage medium, and the method comprises the steps: obtaining a key period of at least one target event, the key period comprising at least one of fault response time, fault processing time, fault complete processing time and service recovery time; determining an average value of the key period of the at least one target event to obtain an event processing efficiency index, the event processing efficiency index comprising at least one of average fault response time, average fault processing time, average fault complete processing time and average service recovery time; through the four indexes, the event processing efficiency of different personnel, teams, cross-team collaborative organization and the like and the proficiency and responsibility of related personnel for mastering the business technology can be seen, and the operation and maintenance management performance can be quantitatively and accurately evaluated comprehensively and objectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular, to a method, device, electronic device, and storage medium for measuring the processing efficiency of events. Background Art

[0002] In current operation and maintenance management, when a back-end system fails, the assessment of operation and maintenance personnel and teams mainly relies on subjective evaluation or simple failure recovery results, and it is impossible to comprehensively and objectively assess the performance of operation and maintenance personnel and teams, which is mainly reflected in the following aspects:

[0003] 1. Inaccurate assessment of technical familiarity: It is difficult to quantitatively evaluate the familiarity and mastery of operation and maintenance personnel and teams with technical architectures, businesses, and systems, and it is impossible to accurately determine whether the extended processing time is due to lack of proficiency in technology or other factors.

[0004] 2. Lack of basis for work efficiency assessment: There are no clear quantitative indicators to measure the work efficiency of operation and maintenance personnel and teams during the process of handling failures. It is impossible to accurately judge whether the team can respond, handle, and recover the business efficiently in different failure scenarios, and it is also difficult to evaluate its collaborative efficiency with other teams.

[0005] 3. Difficulty in quantifying the sense of responsibility assessment: The assessment of the sense of responsibility of operation and maintenance personnel lacks quantifiable data support. It is impossible to reflect the initiative and sense of responsibility of personnel during the process of handling failures through specific time data, such as whether they respond in a timely manner and cooperate actively. Summary of the Invention

[0006] The present invention provides a method, device, electronic device, and storage medium for measuring the processing efficiency of events to solve the technical problem of being unable to objectively and comprehensively assess the performance of operation and maintenance management.

[0007] In a first aspect, the present invention provides a method for measuring the processing efficiency of events, including: obtaining the critical time periods of at least one target event, where the critical time periods include at least one of a fault response time, a fault handling time, a complete fault handling time, and a service recovery time, where the fault response time refers to the time period between the moment when a fault occurs and the moment when the fault starts to be handled, the fault handling time refers to the time period between the moment when the fault starts to be handled and the moment when the fault is completely recovered, the complete fault handling time refers to the time period between the moment when the fault starts to be handled and the moment when a report and a rectification plan form are output, and the service recovery time refers to the time period between the moment when a fault occurs and the moment when a report and a rectification plan form are output; determining the average value of the critical time periods of the at least one target event to obtain an event processing efficiency index, where the event processing efficiency index includes at least one of an average fault response time, an average fault handling time, an average complete fault handling time, and an average service recovery time.

[0008] In some embodiments, obtaining the critical time period of at least one target event includes: obtaining the critical moment of at least one target event, where the critical moment includes the fault occurrence moment, the alarm receiving moment, the fault start processing moment, the fault emergency recovery completion moment, the fault complete recovery moment, the output report and rectification plan form moment; determining a first time period corresponding to the time from the fault occurrence moment to the alarm receiving moment, a second time period corresponding to the time from the alarm receiving moment to the fault start processing moment, a third time period corresponding to the time from the fault start processing moment to the fault emergency recovery completion moment, a fourth time period corresponding to the time from the fault emergency recovery completion moment to the fault complete recovery moment, and a fifth time period corresponding to the time from the fault complete recovery moment to the output report and rectification plan form moment; determining the fault response time according to the sum of the first time period and the second time period; determining the fault processing time according to the sum of the third time period and the fourth time period; determining the fault complete processing time according to the sum of the fault processing time and the fifth time period; and determining the service restoration time according to the sum of the fault response time and the fault complete processing time.

[0009] In some embodiments, determining the average value of the critical time period of the at least one target event to obtain an event processing efficiency index includes: determining the weight of the corresponding critical time period of each target event; and obtaining the corresponding event processing efficiency index according to the weighted average value of the corresponding critical time period of each target event.

[0010] In some embodiments, determining the weight of the corresponding critical time period of each target event includes at least one of the following: determining the weight corresponding to the fault response time of each target event according to the severity of the target event; determining the weight corresponding to the fault processing time of each target event according to the severity of the target event and the service processing time period; determining the weight corresponding to the fault complete processing time of each target event according to the resource consumption cost of the target event; determining the weight corresponding to the service restoration time of each target event according to the service priority of the target event; and obtaining the corresponding event processing efficiency index according to the weighted average value of the corresponding critical time period of each target event includes at least one of the following: obtaining the average fault response time according to the weighted average value of the fault response time of each target event; obtaining the average fault processing time according to the weighted average value of the fault processing time of each target event; obtaining the average fault complete processing time according to the weighted average value of the fault complete processing time of each target event; and obtaining the average service restoration time according to the weighted average value of the service restoration time of each target event.

[0011] In some embodiments, the resource consumption cost includes the cost of manual troubleshooting time and the analysis time cost of the report and rectification plan; then determining the weight corresponding to the complete troubleshooting time of each target event according to the resource consumption cost of the target event includes: determining the first weight corresponding to the troubleshooting time of each target event according to the manual troubleshooting time cost of the target event; determining the second weight corresponding to the fifth time period of each target event according to the analysis time cost of the report and rectification plan of the target event; obtaining the average complete troubleshooting time according to the weighted average of the complete troubleshooting times of each target event includes: adding the product of the troubleshooting time of each target event and the first weight and the product of the fifth time period and the second weight, and then dividing by the cumulative number of target events to obtain the average complete troubleshooting time.

[0012] In some embodiments, the method further includes: using at least one event processing efficiency indicator as an expected value; monitoring the current critical time period corresponding to the current target event in real time, and outputting an optimization prompt message when the current critical time period exceeds the corresponding expected value.

[0013] In some embodiments, the method further includes at least one of the following: periodically repeating the step of obtaining the critical time period of at least one target event to update the expected value; outputting an alarm message when it is detected in real time that the current critical time period of the current target event is greater than a preset threshold; receiving a modification to the preset threshold.

[0014] In a second aspect, the present invention provides a device for measuring event processing efficiency, including: an acquisition module for acquiring the critical time period of at least one target event, where the critical time period includes at least one of a fault response time, a fault processing time, a complete fault processing time, and a service restoration time, the fault response time is the time period from the moment when the fault occurs to the moment when the fault starts to be processed, the fault processing time is the time period from the moment when the fault starts to be processed to the moment when the fault is completely restored, the complete fault processing time is the time period from the moment when the fault starts to be processed to the moment when the report and rectification plan is output, and the service restoration time is the time period from the moment when the fault occurs to the moment when the report and rectification plan is output; a measurement module for determining the average value of the critical time periods of the at least one target event to obtain an event processing efficiency indicator, where the event processing efficiency indicator includes at least one of an average fault response time, an average fault processing time, an average complete fault processing time, and an average service restoration time.

[0015] In a third aspect, the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used for storing a computer program; when the processor executes the program stored on the memory, the steps of the method for measuring the event processing efficiency described in any item of the first aspect are implemented.

[0016] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for measuring the event processing efficiency described in any item of the first aspect are implemented.

[0017] A method, device, electronic device, and storage medium for measuring event processing efficiency provided by an embodiment of the present invention obtain the processing times of each stage of multiple failure events, and obtain four different-dimensional indicators: average failure response time, average failure processing time, average complete failure processing time, and average service recovery time by calculating the mean. Through these four indicators, the efficiency of different personnel, teams, and cross-team collaborative organizations in processing events, as well as the proficiency and sense of responsibility of relevant personnel in mastering business technologies, can be seen, realizing a comprehensive and objective quantitative and accurate evaluation of the operation and maintenance management performance. Description of the Drawings

[0018] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present invention, and are used together with the description to explain the principles of the present invention.

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic flowchart of a method for measuring event processing efficiency provided by an embodiment of the present invention;

[0021] Figure 2 It is a schematic diagram of each critical moment and critical period of a target event provided by an embodiment of the present invention;

[0022] Figure 3 For Figure 1 A detailed flowchart of step S102 in the shown embodiment;

[0023] Figure 4 It is a schematic structural diagram of a device for measuring event processing efficiency provided by an embodiment of the present invention;

[0024] Figure 5 A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0026] Figure 1 A schematic flowchart of a method for measuring the processing efficiency of metric events provided by an embodiment of the present invention.

[0027] As Figure 1 shown, the method includes the following steps:

[0028] Step S101: Obtain the key time periods of at least one target event, where the key time periods include at least one of a fault response time, a fault handling time, a complete fault handling time, and a service recovery time.

[0029] Among them, the fault response time refers to the time period from the moment when a fault occurs to the moment when the fault starts to be processed; the fault handling time refers to the time period from the moment when the fault starts to be processed to the moment when the fault is completely recovered; the complete fault handling time refers to the time period from the moment when the fault starts to be processed to the moment when a report and a rectification plan form are output; and the service recovery time refers to the time period from the moment when a fault occurs to the moment when a report and a rectification plan form are output.

[0030] Specifically, the target event is a fault event. An automated monitoring tool can be deployed in the background system to monitor and record the key time periods of the fault event in real time, such as the fault response time (Responsible Time, abbreviated as RT), the fault handling time (Event Holding Time, abbreviated as EHT), the complete fault handling time (Problem Full Holding Time, abbreviated as PFHT), and the service recovery time (Recover Business Time, abbreviated as RBT).

[0031] Among them, RT refers to the time period from the moment when a fault occurs to the moment when the fault starts to be processed, reflecting the time from the occurrence of the fault to the response of the operation and maintenance personnel; EHT refers to the time period from the moment when the fault starts to be processed to the moment when the fault is completely recovered, reflecting the time from the response of the operation and maintenance personnel to the completion of the initial processing of the fault; PFHT refers to the time period from the moment when the fault starts to be processed to the moment when the report and the rectification plan are output, reflecting the time from the start of the fault processing to the complete resolution of the fault; RBT refers to the time period from the moment when the fault occurs to the moment when the report and the rectification plan are output, reflecting the time from the occurrence of the fault to the restoration of normal business.

[0032] It should be noted that in this embodiment, with the help of the automatic monitoring tool in the background system, the actual processing time of each stage during the period from the response of the fault event to the restoration of normal business can be accurately obtained. Most traditional fault handling record methods rely on manual records, which are prone to inaccurate time records and may also miss key time nodes, thus making it impossible to truly present the time consumption situation during the fault handling process.

[0033] In some embodiments, the step S101 includes the following steps:

[0034] Obtain the critical moments of at least one target event, where the critical moments include the moment when the fault occurs, the moment when the alarm is received, the moment when the fault starts to be processed, the moment when the emergency recovery of the fault is completed, the moment when the fault is completely recovered, and the moment when the report and the rectification plan are output;

[0035] Determine a first time period corresponding to the time period from the moment when the fault occurs to the moment when the alarm is received, a second time period corresponding to the time period from the moment when the alarm is received to the moment when the fault starts to be processed, a third time period corresponding to the time period from the moment when the fault starts to be processed to the moment when the emergency recovery of the fault is completed, a fourth time period corresponding to the time period from the moment when the emergency recovery of the fault is completed to the moment when the fault is completely recovered, and a fifth time period corresponding to the time period from the moment when the fault is completely recovered to the moment when the report and the rectification plan are output;

[0036] Determine the fault response time according to the sum of the first time period and the second time period;

[0037] Determine the fault handling time according to the sum of the third time period and the fourth time period;

[0038] Determine the complete fault handling time according to the sum of the fault handling time and the fifth time period;

[0039] Determine the business recovery time according to the sum of the fault response time and the complete fault handling time.

[0040] Specifically, as Figure 2A schematic diagram of each critical moment and critical period of a target event provided by an embodiment of the present invention. An automated monitoring tool deployed in the background system can record the critical time nodes (i.e., critical moments) of a fault event in real time, such as the fault occurrence moment, the alarm receipt moment, the fault start processing moment, the fault emergency recovery completion moment, the fault complete recovery moment, and the output report and rectification plan form moment. These collected time data can be stored in a dedicated database and classified and managed according to dimensions such as fault event type and occurrence time, facilitating subsequent query and analysis.

[0041] Then, the critical periods are calculated based on these specific moments collected by the database, and the process is as follows:

[0042] Determine the first time period T1 from the fault occurrence moment to the alarm receipt moment, the second time period T2 from the alarm receipt moment to the fault start processing moment, determine the third time period T3 from the fault start processing moment to the fault emergency recovery completion moment, determine the fourth time period T4 corresponding to the fault emergency recovery completion moment to the fault complete recovery moment, determine the fifth time period T5 corresponding to the fault complete recovery moment to the output report and rectification plan form moment. T5 reflects the time from the complete resolution of the fault to the completion of the analysis report (Analysis Time, abbreviated as AT); RT, EHT, PFHT, and RBT are calculated respectively according to formulas (1), (2), (3), and (4):

[0043] RT = T1 + T2 (1)

[0044] EHT = T3 + T4 (2)

[0045] PFHT = EHT + T5 = EHT + AT = T3 + T4 + T5 (3)

[0046] RBT = RT + PFHT = T1 + T2 + T3 + T4 + T5 (4)

[0047] In addition, the automated monitoring tool of the background system will also record and store the context-related information of the fault event, such as the fault type (such as network fault, hardware fault, software fault, etc.), the fault level (such as level 1 fault, level 2 fault, etc.), the impact scope (such as the number of affected users, business modules, etc.), and the system state snapshot (such as server performance indicators, log information, etc.), which can facilitate the staff to perform some extended statistical analysis work based on this information.

[0048] Step S102: Determine the average value of the critical periods of the at least one target event to obtain an event processing efficiency index, where the event processing efficiency index includes at least one of the average fault response time, the average fault handling time, the average fault complete handling time, and the average business recovery time.

[0049] Specifically, when the target event has occurred n times in total, the fault response times of these n target events can be collected, and the mean value is calculated to obtain the mean time to hold (MTTH), as shown in formula (5): If the fault handling times of n target events are collected and the mean value is calculated, the mean time to normal (MTTN) is obtained, as shown in formula (6); the complete fault handling times of these n target events are collected, and the mean value is calculated to obtain the mean time to work (MTTW), as shown in formula (7); the service restoration times of these n target events are collected, and the mean value is calculated to obtain the mean time to Business (MTTB), as shown in formula (8). The metrics calculated based on multiple event occurrences can better reflect the real situation.

[0050] MTTH = ∑(T1 + T2) / n (5)

[0051] MTTN = ∑(T3 + T4) / n (6)

[0052] MTTW = ∑(EHT + AT) / n (7)

[0053] MTTB = ∑(RT + PFHT) / n (8)

[0054] After calculating the above four metrics, the performance of operation and maintenance management in different dimensions can be evaluated respectively. For example, the key points of MTTH assessment are the monitoring system efficiency, the status of personnel on duty, and the collaboration efficiency between operation and maintenance and R & D. If MTTH is relatively high, it indicates that there may be warning delays in the monitoring system, or the response of operation and maintenance personnel on duty is not timely, and it is necessary to optimize the monitoring configuration, strengthen personnel duty management, and improve the collaboration process. For example, the key points of MTTN assessment are the collaboration efficiency within the technical center, the familiarity with the system, and the mastery of key technologies. When MTTN is relatively high, it shows that the internal collaboration within the technical center is not smooth, and the technical proficiency of team members needs to be improved. It is necessary to strengthen technical training and optimize the collaboration mechanism. For example, the key points of MTTW assessment are the collaboration efficiency between the product center and the technical center, the service desk organization efficiency, the sense of responsibility, and the familiarity with the business line. Through MTTW, the cross-departmental collaboration effect and personnel responsibility can be evaluated. If the performance is not good, it is necessary to strengthen communication, clarify the division of labor, and improve the business understanding. For example, the key points of MTTB assessment are the organization collaboration efficiency, the cross-team organization efficiency, the familiarity with the technical architecture, and the business mastery. MTTB reflects the comprehensive performance of the whole organization during the business restoration process. Based on this metric, the organizational structure can be optimized, the technical architecture training can be strengthened, and the business awareness can be improved. It can be seen that the operation and maintenance assessment metric system of this embodiment is very comprehensive and can cover the key links and important factors in the fault handling process.

[0055] Generate a detailed report on the above four indicators and feedback it to relevant teams and individuals to provide a clear direction and basis for improving work. Of course, in addition to the above four indicators, other dimensions of assessment indicators can be introduced, such as fault resolution rate, customer satisfaction, etc., to achieve multi-dimensional assessment; personalized assessment indicators can also be formulated according to the responsibilities and work priorities of different departments and personnel; the assessment results can be linked to performance appraisal, promotion, etc. to motivate teams and individuals to actively improve work.

[0056] In some embodiments, the method further includes: using at least one event processing efficiency indicator as an expected value; monitoring in real time the current critical period corresponding to the current target event, and outputting an optimization prompt message when the current critical period exceeds the corresponding expected value.

[0057] Specifically, take the four calculated indicators as target values to provide a reference standard for subsequent fault event processing. In subsequent event processing, monitor the critical period of the event in real time, compare the actual critical period with the target value. If the actual critical period exceeds the target value, the reason can be analyzed in depth and improvement measures can be taken.

[0058] In some embodiments, the method further includes at least one of the following: periodically repeating the step of obtaining the critical period of at least one target event to update the expected value; outputting an alarm message when it is monitored in real time that the current critical period of the current target event is greater than a preset threshold; receiving a modification to the preset threshold.

[0059] Specifically, it is necessary to recalculate the event processing efficiency regularly to update the target value, form a virtuous cycle, and promote the operation and maintenance team to continuously improve technical level, work efficiency and collaboration ability, ensuring the stable operation of the background system and the continuous development of the business; in addition, this embodiment also adds an abnormal monitoring function, that is, when monitoring the current target event in real time, when a certain critical period deviates from the normal range, an alarm message is sent in time to remind the operation and maintenance personnel to pay attention; a preset threshold setting function is also provided, allowing users to adjust the warning threshold according to the actual situation.

[0060] The method for measuring the efficiency of event processing provided in this embodiment monitors and records the key time nodes of fault events in real time by deploying an automated monitoring tool in the background system, so as to calculate data such as response time (RT), event handling time (EHT), complete event handling time (PFHT), analysis report processing time (AT), and business recovery time (RBT) for each time, and save them in the database for subsequent query and calculation. After multiple (n) events have occurred cumulatively, the average response time (MTTH), average time to failure (MTTN), average complete event handling time (MTTW), and average business recovery time (MTTB) are calculated by taking the mean. Based on these four metrics, the familiarity and mastery of the technical architecture, business, and system by personnel and teams, as well as the work efficiency, organizational collaboration efficiency, and sense of responsibility of personnel and teams for their work, are respectively evaluated with emphasis, achieving comprehensive and objective quantitative and accurate assessment. In addition, the target values of the four metrics calculated in this embodiment can be updated regularly. When events occur again later, the target values can be referred to for further improvement, playing a role in a virtuous cycle.

[0061] On the basis of the foregoing embodiment, Figure 3 For Figure 1 A detailed flowchart of step S102 in the illustrated embodiment is shown in Figure 3 As shown, this step S102 includes the following steps:

[0062] Step S301: Determine the weights of the corresponding key time periods of each target event.

[0063] Step S302: Obtain the corresponding event processing efficiency metrics according to the weighted average values of the corresponding key time periods of each target event.

[0064] Specifically, by assigning weights to the key time periods of each target event according to the severity, business priority, business time period, etc. of each target event, and calculating the metrics by means of weighted averaging, interference from low-priority or less influential events can be avoided, and the actual situation can be more accurately reflected.

[0065] In some embodiments, step S301 includes at least one of the following: determining the weights corresponding to the fault response times of each target event according to the severity of the target event; determining the weights corresponding to the fault handling times of each target event according to the severity of the target event and the business processing period; determining the weights corresponding to the complete fault handling times of each target event according to the resource consumption cost of the target event; determining the weights corresponding to the service recovery times of each target event according to the business priority of the target event; step S302 includes at least one of the following: obtaining the average fault response time according to the weighted average of the fault response times of each target event; obtaining the average fault handling time according to the weighted average of the fault handling times of each target event; obtaining the average complete fault handling time according to the weighted average of the complete fault handling times of each target event; obtaining the average service recovery time according to the weighted average of the service recovery times of each target event.

[0066] Specifically, for the MTTH metric, the weights can be dynamically adjusted according to the severity of each target event (such as P0, P1, P2 levels). The higher the severity of the event, the higher the priority and the greater the corresponding weight. The lower the severity of the event, the lower the priority and the smaller the corresponding weight. The following is an example:

[0067] P0 event (severe fault): weight = 3;

[0068] P1 event (medium fault): weight = 2;

[0069] P2 event (minor fault): weight = 1;

[0070] Then the weighted calculation formula for MTTH is as follows:

[0071]

[0072] where n represents the cumulative number of occurrences of the target event, T1 i represents the value of the first time period corresponding to the i-th target event, T2 i represents the value of the second time period corresponding to the i-th target event, w i represents the RT weight value determined according to the severity of the i-th target event.

[0073] It should be noted that events with high severity often have a greater impact on the continuity and stability of the business. By weighting, it can be ensured that high-priority events have a greater impact on the overall metric, avoiding low-priority events from diluting key data and more truly reflecting the operation and maintenance capabilities of the system under high load or severe faults.

[0074] For the MTTN metric, affected by the characteristics of the business cycle and the severity of events, the weights can be dynamically adjusted according to different time periods of the business (such as levels T0, T1, T2, etc.). During peak business hours, which belong to high-priority events (such as T0), the weights are larger, and during idle business hours, which belong to low-priority events (such as T2), the weights are smaller. The weights are combined with the severity of the events. The calculation formula for MTTN is as follows:

[0075]

[0076] Among them, T3 i represents the value of the third time period corresponding to the i-th target event, and T4 i represents the value of the fourth time period corresponding to the i-th target event, and s i represents the weight coefficient of the business processing time period. Through the dual-weight mechanism, the dynamic perception ability of the fault handling time is improved, and the current true fault recovery ability can be more accurately reflected.

[0077] For the MTTW metric, considering that in the actual processing efficiency of target events, resource consumption and labor cost factors may also be involved, different weights can be assigned to the complete fault handling time according to the resource consumption cost.

[0078] In some embodiments, obtaining the average complete fault handling time according to the weighted average of the complete fault handling times of each target event includes: adding the product of the fault handling time of each target event and the first weight and the product of the fifth time period and the second weight, and then dividing by the cumulative number of target events to obtain the average complete fault handling time.

[0079] Specifically, as can be seen from formula (3), the complete fault handling time can be divided into the fault handling time EHT and the analysis report time AT. Different weights can be assigned to EHT (engineer handling time) and AT (analysis report handling time) to improve the sensitivity to resource costs. The calculation formula for MTTW is as follows:

[0080]

[0081] Among them, EHT i represents the fault handling time corresponding to the i-th target event, and AT i represents the analysis report time corresponding to the i-th target event, and C h represents the weight value of EHT of the corresponding target event, and C a represents the weight value of AT of the corresponding target event.

[0082] Regarding the MTTB metric, as can be seen from formula (4), the recovery business time RBT includes RT and PFHT, and both of these values are affected by the primary and secondary nature of the business. The weights can be dynamically adjusted according to the primary and secondary nature of the business (such as levels E0, E1, E2, etc.). For core businesses, the weight for high-priority events (such as E0) is larger, and for non-core businesses, the weight for low-priority events (such as E2) is smaller. Specific weight examples:

[0083] Event E0 (core business interruption event): weight = 3;

[0084] Event E1 (secondary business impact event): weight = 2;

[0085] Event E2 (non-business critical event): weight = 1;

[0086] The calculation formula for MTTB is as follows:

[0087]

[0088] where RT i represents the fault response time corresponding to the i-th target event, and PFHT i represents the fault complete handling time corresponding to the i-th target event, and u i represents the weight value of the RBT corresponding to the i-th target event.

[0089] Based on the foregoing embodiments, by determining the weights of the corresponding key time periods of each target event, and according to the weighted average value of the corresponding key time periods of each target event, the corresponding event handling efficiency metric is obtained. By the method of weighted averaging to obtain the metric, various factors can be comprehensively considered, making the metric more in line with the actual business requirements and more objectively and accurately measuring the operation and maintenance efficiency.

[0090] Figure 4 This is a schematic structural diagram of a device for measuring event handling efficiency provided by an embodiment of the present invention. As Figure 4 shown, the device includes:

[0091] A data acquisition module 401, configured to obtain the key time periods of at least one target event, where the key time periods include at least one of the fault response time, the fault handling time, the fault complete handling time, and the recovery business time. The fault response time refers to the time period between the moment when the fault occurs and the moment when the fault starts to be handled. The fault handling time refers to the time period between the moment when the fault starts to be handled and the moment when the fault is completely recovered. The fault complete handling time refers to the time period between the moment when the fault starts to be handled and the moment when the report and the rectification plan form are output. The recovery business time refers to the time period between the moment when the fault occurs and the moment when the report and the rectification plan form are output;

[0092] An index calculation module 402 is configured to determine an average value of critical periods of the at least one target event, so as to obtain an event processing efficiency index, where the event processing efficiency index includes at least one of an average fault response time, an average fault handling time, an average complete fault handling time, and an average service restoration time.

[0093] In some embodiments, the data collection module 401 is specifically configured to:

[0094] Obtain critical moments of at least one target event, where the critical moments include a fault occurrence moment, an alarm receiving moment, a fault start handling moment, a moment when a fault is urgently restored, a moment when a fault is completely restored, and a moment when a report and a rectification plan form are output;

[0095] Determine a first time period corresponding to the time period from the fault occurrence moment to the alarm receiving moment, a second time period corresponding to the time period from the alarm receiving moment to the fault start handling moment, a third time period corresponding to the time period from the fault start handling moment to the moment when the fault is urgently restored, a fourth time period corresponding to the time period from the moment when the fault is urgently restored to the moment when the fault is completely restored, and a fifth time period corresponding to the time period from the moment when the fault is completely restored to the moment when the report and the rectification plan form are output;

[0096] Determine the fault response time according to the sum of the first time period and the second time period;

[0097] Determine the fault handling time according to the sum of the third time period and the fourth time period;

[0098] Determine the complete fault handling time according to the sum of the fault handling time and the fifth time period;

[0099] Determine the service restoration time according to the sum of the fault response time and the complete fault handling time.

[0100] In some embodiments, the index calculation module 402 is specifically configured to:

[0101] Determine weights of corresponding critical periods of each target event;

[0102] Obtain corresponding event processing efficiency indexes according to weighted average values of corresponding critical periods of each target event.

[0103] In some embodiments, the index calculation module 402 is specifically configured to perform at least one of the following:

[0104] Determine weights corresponding to the fault response times of each target event according to the severity of the target event;

[0105] Determine the weight corresponding to the fault handling time of each target event according to the severity of the target event and the business processing period;

[0106] Determine the weight corresponding to the complete fault handling time of each target event according to the resource consumption cost of the target event;

[0107] Determine the weight corresponding to the service restoration time of each target event according to the business priority of the target event;

[0108] Obtain the average fault response time according to the weighted average of the fault response times of each target event;

[0109] Obtain the average fault handling time according to the weighted average of the fault handling times of each target event;

[0110] Obtain the average complete fault handling time according to the weighted average of the complete fault handling times of each target event;

[0111] Obtain the average service restoration time according to the weighted average of the service restoration times of each target event.

[0112] In some embodiments, the resource consumption cost includes the labor cost for handling faults and the analysis time cost of the report and rectification plan; then the index calculation module 402 is specifically configured to:

[0113] Determine the first weight corresponding to the fault handling time of each target event according to the labor cost for handling faults of the target event;

[0114] Determine the second weight corresponding to the fifth time period of each target event according to the analysis time cost of the report and rectification plan of the target event;

[0115] Add the product of the fault handling time of each target event and the first weight and the product of the fifth time period and the second weight, and then divide by the cumulative number of target events to obtain the average complete fault handling time.

[0116] In some embodiments, the device further includes a processing module 403, and the processing module 403 is configured to:

[0117] Use at least one event processing efficiency index as the expected value;

[0118] Real-time monitor the current critical period corresponding to the current target event, and output an optimization prompt message when the current critical time period exceeds the corresponding expected value.

[0119] In some embodiments, the processing module 403 is further configured to perform at least one of the following:

[0120] Periodically repeat the step of obtaining the critical time period of at least one target event to update the expected value;

[0121] When it is detected in real time that the current critical time period of the current target event is greater than a preset threshold, an alarm message is output;

[0122] Receive the modification of the preset threshold.

[0123] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process and corresponding beneficial effects of the device for measuring event processing efficiency described above can refer to the corresponding process in the foregoing method examples, and will not be elaborated here.

[0124] Figure 5 The following is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. As Figure 5 shown, the electronic device includes: a processor 501, a communication interface 502, a memory 503, and a communication bus 504. Among them, the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.

[0125] The memory 503 is used to store a computer program;

[0126] In an embodiment of the present application, when the processor 501 is used to execute the program stored on the memory 503, it implements the steps of the method for measuring event processing efficiency provided by any one of the foregoing method embodiments.

[0127] The electronic device provided by the embodiment of the present application has the same implementation principle and technical effects as the above embodiment, and will not be elaborated here.

[0128] The above-mentioned memory 503 may be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. The memory 503 has a storage space for program codes for executing any of the method steps in the above-mentioned method. For example, the storage space for program codes may include respective program codes for implementing the respective steps in the above method. These program codes may be read from or written into one or more computer program products. These computer program products include program code carriers such as hard disks, optical discs (CDs), memory cards, or floppy disks. Such computer program products are generally portable or fixed storage units. The storage unit may have a storage segment or a storage space, etc., arranged similarly to the memory 503 in the above-mentioned electronic device. The program codes may be compressed in an appropriate form, for example. Generally, the storage unit includes a program for executing the method steps according to the embodiments of the present application, that is, codes that can be read by a processor such as 501, and when these codes are run by the electronic device, they cause the electronic device to execute the respective steps in the method described above.

[0129] Embodiments of the present application also provide a computer-readable storage medium. A computer program is stored on the above-mentioned computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps of the method for measuring the efficiency of event processing as described above.

[0130] The computer-readable storage medium may be included in the device / apparatus described in the above embodiments; or it may exist separately and not be assembled into the device / apparatus. The above-mentioned computer-readable storage medium carries one or more programs, and when the one or more programs are executed, they implement the method according to the embodiments of the present application.

[0131] According to the embodiments of the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, device, or apparatus.

[0132] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0133] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for measuring the efficiency of event processing, characterized in that, Including: Obtain the critical periods of at least one target event, where the critical periods include at least one of the fault response time, the fault handling time, the complete fault handling time, and the business recovery time. Among them, the fault response time refers to the time period from the moment when the fault occurs to the moment when the fault starts to be handled; the fault handling time refers to the time period from the moment when the fault starts to be handled to the moment when the fault is completely recovered; the complete fault handling time refers to the time period from the moment when the fault starts to be handled to the moment when the report and the rectification plan form are output; the business recovery time refers to the time period from the moment when the fault occurs to the moment when the report and the rectification plan form are output; Determine the average value of the critical periods of the at least one target event to obtain an event handling efficiency index, where the event handling efficiency index includes at least one of the average fault response time, the average fault handling time, the average complete fault handling time, and the average business recovery time.

2. The method according to claim 1, characterized in that, The obtaining of the critical periods of at least one target event includes: Obtain the critical moments of at least one target event, where the critical moments include the moment when the fault occurs, the moment when the alarm is received, the moment when the fault starts to be handled, the moment when the emergency fault recovery is completed, the moment when the fault is completely recovered, and the moment when the report and the rectification plan form are output; Determine the first time period corresponding to the moment when the fault occurs to the moment when the alarm is received, the second time period corresponding to the moment when the alarm is received to the moment when the fault starts to be handled, the third time period corresponding to the moment when the fault starts to be handled to the moment when the emergency fault recovery is completed, the fourth time period corresponding to the moment when the emergency fault recovery is completed to the moment when the fault is completely recovered, and the fifth time period corresponding to the moment when the fault is completely recovered to the moment when the report and the rectification plan form are output; Determine the fault response time according to the sum of the first time period and the second time period; Determine the fault handling time according to the sum of the third time period and the fourth time period; Determine the complete fault handling time according to the sum of the fault handling time and the fifth time period; Determine the business recovery time according to the sum of the fault response time and the complete fault handling time.

3. The method according to claim 2, wherein The determining of the average value of the critical periods of the at least one target event to obtain an event handling efficiency index includes: Determine the weights of the corresponding critical periods of each target event; Obtain the corresponding event handling efficiency index according to the weighted average value of the corresponding critical periods of each target event.

4. The method according to claim 3, wherein The determining of the weights of the corresponding critical periods of each target event includes at least one of the following: Determine the weights corresponding to the fault response time of each target event according to the severity of the target event; Determine the weights corresponding to the fault handling time of each target event according to the severity of the target event and the business processing period; Determine the weights corresponding to the complete fault handling time of each target event according to the resource consumption cost of the target event; Determine the weights corresponding to the business recovery time of each target event according to the business priority of the target event; The obtaining of the corresponding event handling efficiency index according to the weighted average value of the corresponding critical periods of each target event includes at least one of the following: Obtain the average fault response time according to the weighted average of the fault response times of each target event; Obtain the average fault handling time according to the weighted average of the fault handling times of each target event; Obtain the average complete fault handling time according to the weighted average of the complete fault handling times of each target event; Obtain the average service restoration time according to the weighted average of the service restoration times of each target event.

5. The method according to claim 4, wherein The resource consumption cost includes the time cost of manual fault handling and the time cost of analyzing the report and rectification plan; then determining the weight corresponding to the complete fault handling time of each target event according to the resource consumption cost of the target event includes: Determine the first weight corresponding to the fault handling time of each target event according to the time cost of manual fault handling of the target event; Determine the second weight corresponding to the fifth time period of each target event according to the time cost of analyzing the report and rectification plan of the target event; The obtaining the average complete fault handling time according to the weighted average of the complete fault handling times of each target event includes: Add the product of the fault handling time of each target event and the first weight and the product of the fifth time period and the second weight, and then divide by the cumulative number of target events to obtain the average complete fault handling time.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Taking at least one event processing efficiency index as the expected value; Real-time monitor the current critical period corresponding to the current target event, and output an optimization prompt message when the current critical period exceeds the corresponding expected value.

7. The method according to claim 6, characterized in that, The method further includes at least one of the following: Periodically repeat the step of obtaining the critical period of at least one target event to update the expected value; Output an alarm message when it is detected in real time that the current critical period of the current target event is greater than the preset threshold; Receive the modification of the preset threshold.

8. A device for measuring the efficiency of event processing, characterized in that, It includes: A data acquisition module for obtaining the critical period of at least one target event, where the critical period includes at least one of the fault response time, fault handling time, complete fault handling time, and service restoration time. The fault response time refers to the time period from the moment of fault occurrence to the moment when fault handling starts. The fault handling time refers to the time period from the moment when fault handling starts to the moment of complete fault recovery. The complete fault handling moment refers to the time period from the moment when fault handling starts to the moment when the report and rectification plan are output. The service restoration time refers to the time period from the moment of fault occurrence to the moment when the report and rectification plan are output; An index calculation module for determining the average value of the critical periods of the at least one target event to obtain an event processing efficiency index, where the event processing efficiency index includes at least one of the average fault response time, average fault handling time, average complete fault handling time, and average service restoration time.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; A processor, when executing a program stored in a memory, implements the steps of the method for measuring event processing efficiency according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for measuring event processing efficiency according to any one of claims 1-7.