A method and system for device failure handling
Patent Information
- Application Number
- CN202511076957.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-08-01
AI Technical Summary
[0002]在现有的机房设备管理中,往往依赖人工巡检发现智联设备的故障并处置,效率低下且容易出错,由于缺乏自动化处理机制,智联设备的故障处置过程中存在信息延迟、信息不准确以及资源配置不合理等问题
(1)通过自动化告警采集与派单处置流程,减少了人工干预,缩短了智联设备故障的处理时间;
Smart Images

Figure CN121032050B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of operation and maintenance technology, and in particular to a method and system for handling equipment failures. Background Technology
[0002] In the existing management of data center equipment, manual inspections are often used to discover and handle faults in intelligent connected devices, which is inefficient and prone to errors. Due to the lack of automated processing mechanisms, there are problems such as information delays, inaccurate information, and unreasonable resource allocation in the process of handling faults in intelligent connected devices.
[0003] For example, manual data entry and work order dispatching processes: Fault entry and work order dispatching for intelligent connected equipment in the data center often require manual work by staff, communicating via telephone, SMS or email, which is inefficient and prone to errors.
[0004] The process based on manual review: In the stages of fault identification, fault handling scheduling, and fault acceptance, relevant personnel usually need to conduct manual review based on their experience. The process is tedious and highly subjective.
[0005] Database Management System: A database management system is used to record information about faults in smart connected devices, but it lacks automated processing capabilities and cannot achieve efficient work order flow and management.
[0006] Work order tracking systems can record work order status, but they often lack integrated process management and cannot achieve automated work order processing and resource allocation.
[0007] Because troubleshooting of smart connected devices relies on manual inspections, the response speed is slow, resulting in a longer processing cycle and low efficiency. Furthermore, the numerous manual steps lead to incomplete maintenance data, high error rates, and long processing times, affecting the quality of fault handling and customer satisfaction. In addition, the unreasonable resource allocation mechanism results in resource waste or uneven distribution.
[0008] Therefore, an automated solution for handling faults in intelligent connected devices is needed to optimize resource allocation and improve the efficiency and accuracy of fault handling. Summary of the Invention
[0009] The purpose of this invention is to provide a method and system for handling equipment failures, thereby optimizing resource allocation and improving failure handling efficiency.
[0010] To achieve the above objectives, the present invention provides the following technical solution: Firstly, a method for handling equipment malfunctions is provided, including: The alarm messages from multiple devices are acquired in real time. The alarm messages include the alarm device, alarm name, alarm type, alarm level and timestamp. The alarm association rules identify related alarm messages caused by the same fault. The related alarm messages are compressed with the existing fault group to obtain newly issued alarm messages. The work order type and workflow of the new work order are determined based on the alarm type of the newly issued alarm message. Based on the alarm level of the newly issued alarm message and the pre-configured permission information of each process step, the priority of dispatching is configured for each process step. Based on the work order type, the corresponding work order template is invoked, the work order template is configured, the new work order is generated, the entire lifecycle of the new work order is monitored, and the processing status of the new work order is updated synchronously.
[0011] Furthermore, after acquiring alarm messages from multiple devices in real time, it also includes: The alarm message is preprocessed; The preprocessing includes: Filtering and cleaning are used to remove redundant and invalid alarm messages; Standardization processing is used to process the alarm message into a preset standard format.
[0012] Furthermore, the alarm association rules include: spatiotemporal association rules, semantic association rules, and statistical association rules; The spatiotemporal association rules include: time window rules based on a set dynamic time threshold, topology association rules based on device dependencies, and propagation path rules based on fault propagation modes. The semantic association rules include: keyword matching rules based on alarm type mapping table, alarm combination patterns and fault mode library based on typical fault scenarios; The statistical association rules include: device association rules based on alarm frequency thresholds and region association rules based on alarm density thresholds.
[0013] Furthermore, the step of identifying related alarm messages caused by the same fault according to alarm association rules, compressing the related alarm messages with existing fault groups, and obtaining newly issued alarm messages includes: Extract the feature parameters of the alarm message, including alarm name, alarm type and timestamp; Based on the alarm association rules, the alarm messages are analyzed for correlation, and alarm messages with the same root cause of failure are identified as associated alarm messages. Query the current active alarm set, which contains fault groups that have not yet been cleared; The associated alarm messages are matched to existing fault groups and the alarm timestamps in the groups are updated. At the same time, an alarm suppression instruction is generated to block the subsequent work order dispatch process. For new alarms that do not match an existing fault group, create a new fault group record and trigger a standard alarm. Police handling procedures.
[0014] Furthermore, determining the work order type and workflow of the new work order based on the alarm type of the newly issued alarm message includes: Configure the new work order type according to the alarm type of the newly issued alarm message, including: work order type name, business category, fault level, fault type, creator, and creation time; According to the alarm type of the newly issued alarm message, the workflow template of the new work order is obtained from the unified workflow platform. The workflow template has multiple workflow steps, including creation, assignment, processing, review, closure, as well as timeout, rollback and escalation steps triggered in abnormal scenarios. The workflow template of the new work order is parsed to obtain the execution order and work permission information of each process step.
[0015] Furthermore, the step of configuring task dispatch priority for each process step based on the alarm level of the newly issued alarm message and the pre-configured permission information of each process step includes: Obtain the pre-configured execution permission information for each process step; Obtain multiple dynamic parameters for each permission role, including online status, real-time location, business skills matrix, and certificate / insurance validity status; Establish a multi-dimensional evaluation model, adjust the weight coefficients of each dynamic parameter according to the alarm level, and use the dynamic parameters and their weight coefficients to calculate the weighted comprehensive score of each permission role; Based on the weighted comprehensive score from high to low, the priority of dispatching is configured for the execution role of each process step.
[0016] Further, the step of calling the corresponding work order template based on the work order type, configuring the work order template, and generating the new work order includes: Obtain the corresponding work order template from the unified process platform according to the work order type; The work order template is configured with attributes, wherein the attribute information of the work order template includes: template name, work order type, business category, template ID, version number, and the process to which the work order belongs; Call the dispatch service interface to generate the new work order.
[0017] Furthermore, monitoring the entire lifecycle of the new work order and synchronously updating its processing status includes: The new work order is dispatched to the corresponding business system, and the entire lifecycle of the new work order is tracked and recorded. Establish a data synchronization mechanism so that when the processing status of a new work order changes, the status update information is promptly synchronized to the unified process platform.
[0018] Secondly, a device fault handling system is provided, comprising: The alarm compression module is used to acquire alarm messages from multiple devices in real time. The alarm messages include the alarm device, alarm name, alarm type, alarm level and timestamp. Based on the alarm association rules, it identifies related alarm messages caused by the same fault, compresses the related alarm messages with the existing fault group, and acquires newly issued alarm messages. The work order configuration module is used to determine the work order type and flow process of the new work order based on the alarm type of the newly issued alarm message, and to configure the dispatch priority for each process step based on the alarm level of the newly issued alarm message and the pre-configured permission information of each process step. The work order generation and monitoring module is used to call the corresponding work order template based on the work order type, configure the work order template, generate the new work order, monitor the entire life cycle of the new work order, and synchronously update the processing status of the new work order.
[0019] Furthermore, the work order configuration module specifically includes: The work order type and process acquisition unit is used to determine the work order type and flow process of the new work order based on the alarm type of the newly issued alarm message. The execution permission acquisition unit is used to acquire the execution permission information pre-configured for each process step; The parameter acquisition unit is used to acquire multiple dynamic parameters for each permission role, including online status, real-time location, business skills matrix, and certificate / insurance validity status. The scoring unit is used to establish a multi-dimensional evaluation model, adjust the weight coefficients of each dynamic parameter according to the alarm level, and calculate the weighted comprehensive score of each permission role using the dynamic parameters and their weight coefficients. The order dispatch priority configuration unit is used to configure the order dispatch priority for the execution role of each process step according to the weighted comprehensive score from high to low.
[0020] The technical effects and advantages of this invention are as follows: (1) By automating the alarm collection and dispatch process, manual intervention is reduced and the processing time for smart device failures is shortened; (2) By collecting detailed information on fault alarms, the work capabilities and location information of the personnel, etc., it provides support for data analysis and decision-making of the dispatch system, ensuring that tasks can be completed with high quality in the shortest time, realizing the rational allocation and effective use of resources, and reducing resource waste; (3) It provides an intelligent dispatch mechanism, which improves the efficiency and accuracy of fault services; (4) By managing permissions and assigning roles, the security of data and the standardization of operations are ensured, the risk of data leakage and improper operation is reduced, and the dispatch system is made more flexible and scalable.
[0021] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of the equipment fault handling method in an embodiment of the present invention; Figure 2 This is a schematic diagram of the work order docking and workflow in an embodiment of the present invention; Figure 3 This is a schematic diagram of the architecture of the equipment fault handling system according to an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] The first embodiment of the present invention provides a method for handling equipment malfunctions, such as... Figure 1 As shown, the method includes: S1. Real-time acquisition of alarm messages from multiple devices, wherein the alarm message includes alarm device, alarm name, alarm type, alarm level and timestamp, and identification of related alarm messages caused by the same fault according to alarm association rules, compression processing of the related alarm messages with existing fault groups, and acquisition of newly issued alarm messages. In this step, alarm messages reported by various intelligent connected devices in the computer room are received in real time through the alarm machine. The alarm messages are preprocessed into standard format alarm messages, and then compressed to obtain new alarm messages caused by new faults.
[0026] S2. Determine the work order type and workflow of the new work order based on the alarm type of the newly issued alarm message. Configure the dispatch priority for each workflow step based on the alarm level of the newly issued alarm message and the pre-configured permission information of each workflow step. S3. Based on the work order type, call the corresponding work order template, configure the work order template, generate the new work order, monitor the entire lifecycle of the new work order, and synchronously update the processing status of the new work order.
[0027] In this embodiment of the invention, the automated alarm collection and dispatch process reduces manual intervention and shortens the processing time for smart device failures. Through permission management and role allocation, it ensures data security and operational standardization, reduces the risk of data leakage and improper operation, and also makes the dispatch system more flexible and scalable.
[0028] In this embodiment of the invention, after acquiring alarm messages from multiple devices in real time, the method further includes: preprocessing the alarm messages. The preprocessing includes: filtering and cleaning to remove redundant and invalid alarm messages; standardization processing, which involves adding necessary basic information to the alarm message body and formatting the alarm messages into a preset standard format. For example, unifying the time base (time zone synchronization), formatting the alarm messages to the ISO8601 standard, and performing natural language processing on text-based alarms to extract standardized keywords.
[0029] The alarm association rules described in this embodiment of the invention include: spatiotemporal association rules, semantic association rules, and statistical association rules.
[0030] The spatiotemporal association rules include: time window rules based on set dynamic time thresholds, topology association rules based on device dependencies, and propagation path rules based on fault propagation patterns. For example, the time window rule sets a dynamic time threshold (e.g., continuous alarms within 5 minutes), the topology association rule establishes device dependencies based on the network topology map, and the propagation path rule defines the fault propagation pattern (e.g., core switch failure → cascading alarms in downstream devices). The semantic association rules include: keyword matching rules based on an alarm type mapping table, alarm combination patterns based on typical fault scenarios, and a fault mode library. For example, the keyword matching rules establish an alarm type mapping table (such as "link interruption" → "port down" → "BGP session termination"). The fault mode library predefines alarm combination patterns for typical fault scenarios.
[0031] Example: IF (interface traffic exceeds threshold AND CPU utilization > 90% AND memory leak alert) THEN is categorized as "Equipment Overload Failure Group". The statistical association rules include: device association rules based on alarm frequency thresholds (the same device alarms more than N times within a specific time period) and regional association rules based on alarm density thresholds (the number of regional alarms increases suddenly within a unit of time).
[0032] According to an embodiment of the present invention, the step of identifying related alarm messages caused by the same fault according to alarm association rules, compressing the related alarm messages with existing fault groups, and obtaining newly issued alarm messages includes: S11. Extract the feature parameters of the alarm message, including alarm name, alarm type and timestamp; S12. Based on the alarm association rules, perform correlation analysis on the alarm messages and identify alarm messages with the same root cause as associated alarm messages. S13. Query the current active alarm set, which contains fault groups that have not yet been cleared; S14. Match the associated alarm message to the existing fault group and update the alarm timestamp in the group. At the same time, generate an alarm suppression instruction to block the subsequent work order dispatch process. S15. For new alarms that do not match an existing fault group, create a new fault group record and trigger the standard alarm handling process.
[0033] For example, if two alarm messages point to the same network device and their fault type codes fall within a predefined range of strongly correlated fault types (e.g., a power module fault and a fan module fault are related due to power supply issues), it can be preliminarily determined that they may be caused by the same root cause. Furthermore, the time sequence of alarm occurrences can be considered. If a new alarm occurs within a preset time threshold after an active alarm (e.g., within 5 minutes), the likelihood of it being identified as associated with the same fault increases. After comprehensively considering multi-dimensional feature data and preset alarm association rules, it is finally determined which newly received alarm messages should be grouped with certain alarms in the active alarm set into a fault group.
[0034] For alarm messages already identified as belonging to the same fault group, compression processing is performed according to preset alarm convergence rules. For example, one time-interval-based convergence rule is that if devices at both ends of the same link generate link failure alarms and link recovery alarms respectively within 10 consecutive minutes, and the link recovery alarm occurs after the link failure alarm, then these two can be compressed and merged into a single link fluctuation fault group, retaining only key time nodes and fault status change information, without processing each individual alarm separately. Additionally, a convergence rule based on fault type correlation is used. When a server hardware failure alarm (such as a memory failure) and a corresponding operating system service anomaly alarm caused by the hardware failure occur simultaneously, they are merged into a server hardware-related fault group. Only hardware-related issues are followed up with subsequent dispatching, omitting separate dispatching for operating system service anomalies. This effectively reduces unnecessary dispatching caused by repetitive and highly correlated alarms, allowing maintenance personnel to focus on handling truly independent and critical faults, improving the overall focus and efficiency of maintenance work, and reducing resource waste and manpower costs caused by processing too many redundant dispatchings. These convergence rules can be flexibly configured and adjusted by system administrators based on the actual network environment and operational experience to adapt to optimization needs in different scenarios.
[0035] New alarms that do not match an existing fault group can be understood as alarms triggered by new faults. The work order type and workflow for the new work order are determined based on the alarm type of the new alarm message. (See [link / reference]). Figure 2 This includes the following steps: S21. Configure the work order type of the new work order according to the alarm type of the newly issued alarm message, including: work order type name, business category, fault level, fault type, creator, and creation time; S22. Obtain the workflow template of the new work order from the unified workflow platform according to the alarm type of the newly issued alarm message. The workflow template has multiple workflow steps, including creation, assignment, processing, review, closure, as well as timeout, rollback and escalation steps triggered in abnormal scenarios. S23. Parse the workflow template of the new work order to obtain the execution order and work permission information of each process step.
[0036] A unified workflow platform was used to design a workflow template for fault work orders, defining each business processing node of the fault work order as a workflow step and clarifying the flow relationship between nodes. The work order platform interfaces with the unified workflow platform to decouple business processes and achieve rapid support for business processes through the connection of business nodes. The workflow engine provides management of collaborative scenarios such as pending, completed, collaborative, reassigned, expedited, and handover of work orders, meeting the complex and diverse scenarios of various workflows such as branches, aggregations, multiple sub-workflows, and flexible workflows. It also provides an event-driven mechanism to trigger automated operations (escalation) through event listening (such as work order timeout events), records the entire lifecycle of work orders, supports backtracking and compliance checks, and provides visualization operations.
[0037] Furthermore, the step of configuring task dispatch priority for each process step based on the alarm level of the newly issued alarm message and the pre-configured permission information of each process step includes: S24. Obtain the pre-configured execution permission information for each process step; The unified workflow platform configures operation permissions (such as editing, viewing, and transferring) and execution roles / user groups for each workflow step. For example, work order creation (customer service), work order processing (technician), reassignment (supervisor), and closure (administrator).
[0038] S25. Obtain multiple dynamic parameters for each authorized role, including skill matching degree, real-time location, workload, and certificate insurance validity status; S26. Establish a multi-dimensional evaluation model, adjust the weight coefficients of each dynamic parameter according to the alarm level, and calculate the weighted comprehensive score of each permission role using the dynamic parameters and their weight coefficients. S27. Based on the weighted comprehensive score from high to low, assign priority to the execution role of each process step.
[0039] The multidimensional evaluation model calculates a weighted comprehensive score. S core as follows:
[0040] Skill matching S skill The score is based on the matching degree between the skills required for the task and the task (e.g., perfect match 1.0, secondary match 0.7), and the experience value is added as a weight (e.g., senior engineer experience bonus 20%).
[0041] straight-line distance S geo The Haversine formula is used to calculate the straight-line distance between the fault point and the real-time location of the personnel, which is then converted into an exponential decay function.
[0042] Where k is the attenuation coefficient, the closer the distance, the higher the score.
[0043] Workload Index S workload The score is based on the percentage of the current workload relative to the individual's maximum workload (e.g., 50% workload gets 0.8, 90% workload gets 0.2).
[0044] urgency S emergency The higher the alarm level (level 1 to level 4), the more urgent the alarm.
[0045] , , , These are weighting coefficients, which are dynamically adjusted based on fault handling requirements. For example, when handling a level 1 alarm, the weighting coefficient is increased. ,reduce .
[0046] In actual order dispatching, order recipient information is pushed to the person with the highest weighted composite score in descending order. If the person with the highest weighted composite score fails to accept the order within the time limit, the order is automatically passed to the second highest scorer and marked as "requiring manual intervention". In case of emergencies, such as when a person encounters traffic jams (GPS movement speed remains at 0), the system automatically triggers a re-dispatch. When a person receives a high load of tasks consecutively, the system automatically lowers the priority of their subsequent tasks to avoid overload.
[0047] In this embodiment of the invention, the step of calling the corresponding work order template based on the work order type, configuring the work order template, and generating the new work order includes: S31. Obtain the corresponding work order template from the unified process platform according to the work order type; S32. Configure the attributes of the work order template, wherein the attribute information of the work order template includes: template name, work order type, business category, template ID, version number, and work order process; S33. Call the dispatch service interface to generate the new work order.
[0048] After the work order center retrieves the corresponding work order template from the unified process platform based on the work order type, the form engine module allows users to personalize the work order details, such as whether the main and all fields are required, and the field types. The form engine module also supports the development of various business components, including enumerations, images, and tables, and supports external data sources. After configuration, a "template ID" is output for use by professional network administrators in dispatching work orders.
[0049] After a new work order is generated, it supports comprehensive querying, monitoring, and operational statistics data analysis in the work order center, and provides related services such as supervision, transfer, and assignment for work order scheduling, thereby improving work order processing efficiency and work order completion quality.
[0050] Furthermore, monitoring the entire lifecycle of the new work order and synchronously updating its processing status includes: S34. Dispatch the new work order to the corresponding business system, and track and record the entire lifecycle of the new work order; S35. Establish a data synchronization mechanism so that when the processing status of the new work order changes, the status update information is promptly synchronized to the unified process platform.
[0051] The work order center is responsible for the entire lifecycle management of work orders, from receiving information, creating, reviewing, scheduling, responding, completing, evaluating, and archiving. A data synchronization mechanism has been established to ensure that data between the work order center and the unified workflow platform can be updated in real time.
[0052] The second embodiment of the present invention provides a device fault handling system, the architecture of which is as follows: Figure 3 As shown, it includes: Data layer: Responsible for storing and managing various data during the fault handling process of intelligent connected devices, including alarm data, work order data, configuration information, etc. It adopts a combination of relational databases and non-relational databases to ensure efficient data storage and retrieval.
[0053] Business Logic Layer: The core layer of the system, which implements various business logics and algorithms, including alarm collection, alarm preprocessing, alarm compression, alarm dispatch, station staff (order takers) configuration, work order evaluation, data synchronization, etc., realizing automated processing and optimized management of intelligent connected device faults.
[0054] Presentation layer: Provides users with a user-friendly interface, facilitating alarm queries, work order queries, and fault handling personnel configuration. It adopts Web and mobile application technologies and supports access from multiple terminals.
[0055] In this embodiment of the invention, the equipment fault handling system includes: The alarm compression module is used to acquire alarm messages from multiple devices in real time. The alarm messages include the alarm device, alarm name, alarm type, alarm level and timestamp. Based on the alarm association rules, it identifies related alarm messages caused by the same fault, compresses the related alarm messages with the existing fault group, and acquires newly issued alarm messages. The work order configuration module is used to determine the work order type and flow process of the new work order based on the alarm type of the newly issued alarm message, and to configure the dispatch priority for each process step based on the alarm level of the newly issued alarm message and the pre-configured permission information of each process step. The work order generation and monitoring module is used to call the corresponding work order template based on the work order type, configure the work order template, generate the new work order, monitor the entire life cycle of the new work order, and synchronously update the processing status of the new work order.
[0056] According to a specific embodiment, the alarm compression module specifically includes: The preprocessing module includes a filtering and cleaning unit and a standardization processing unit. The filtering and cleaning unit is used to remove redundant and invalid alarm messages, and the standardization processing unit is used to supplement the alarm message body with necessary basic information and process the alarm message into a preset standard format. The correlation analysis unit is used to extract the feature parameters of the alarm message, including alarm name, alarm type and timestamp, and perform correlation analysis on the alarm message based on alarm correlation rules to identify alarm messages with the same root cause as related alarm messages. The matching unit is used to query the current active alarm set, which contains fault groups that have not yet been cleared, match the associated alarm message to the existing fault group and update the alarm timestamp in the group, and at the same time generate an alarm suppression instruction to block the subsequent work order dispatch process. Create a unit to create a new fault group record and trigger the standard alarm handling process for new alarms that do not match an existing fault group.
[0057] According to a specific embodiment, the work order configuration module specifically includes: The work order type and process acquisition unit is used to determine the work order type of the new work order based on the alarm type of the newly issued alarm message, including: work order type name, business category, fault level, fault type, creator, and creation time. It also obtains the workflow template of the new work order from the unified process platform based on the alarm type of the newly issued alarm message. The workflow template has multiple process steps, including creation, assignment, processing, review, closure, and timeout, rollback, and escalation steps triggered in abnormal scenarios. The execution permission acquisition unit is used to acquire the pre-configured execution permission information of each process step, and to acquire the execution order and work permission information of each process step by parsing the workflow template of the new work order. The parameter acquisition unit is used to acquire multiple dynamic parameters for each permission role, including online status, real-time location, business skills matrix, and certificate / insurance validity status. The scoring unit is used to establish a multi-dimensional evaluation model, adjust the weight coefficients of each dynamic parameter according to the alarm level, and calculate the weighted comprehensive score of each permission role using the dynamic parameters and their weight coefficients. The order dispatch priority configuration unit is used to configure the order dispatch priority for the execution role of each process step according to the weighted comprehensive score from high to low.
[0058] According to a specific embodiment, the work order generation and monitoring module specifically includes: The work order template configuration unit is used to obtain the corresponding work order template from the unified process platform according to the work order type, and configure the attributes of the work order template. The attribute information of the work order template includes: template name, work order type, business category, template ID, version number, and the process to which the work order belongs. The work order generation unit is used to call the dispatch service interface to generate the new work order. The monitoring unit is used to dispatch the new work order to the corresponding business system, and to track and record the entire lifecycle of the new work order. The data synchronization unit is used to establish a data synchronization mechanism, which promptly synchronizes the status update information to the unified process platform when the processing status of the new work order changes.
[0059] After a new work order is generated, the work order center supports comprehensive querying, monitoring, operational statistics, and other data analysis. It also provides related services such as supervision, transfer, and assignment for work order scheduling, improving work order processing efficiency and completion quality. The work order center is responsible for the entire lifecycle management of work orders, from receiving information, creation, review, scheduling, response, completion, evaluation, and archiving. A data synchronization mechanism is established to ensure real-time data updates between the work order center and the unified workflow platform.
[0060] In this embodiment of the invention, multivariate time series analysis and vector autoregression model can also be used to establish an intelligent alarm and early warning model, including the algorithm implementation details of real-time data stream analysis, the selection and training process of machine learning model, the calculation method of dynamic early warning threshold, and the specific implementation of the early warning triggering mechanism.
[0061] By deeply analyzing the complexity and resource requirements of alarm processing for intelligent connected devices, the system provides scientific decision support for the response team. Through correlation analysis of received activity alarm information, and based on preset alarm association rules, it identifies multiple alarms caused by the same fault or event and groups them into alarm groups, helping to reduce redundant alarm information and improve processing efficiency. Simultaneously, the system prioritizes alarms or alarm groups based on factors such as alarm severity, impact scope, and urgency, ensuring that high-priority alarms are processed first. The intelligent connected device alarm early warning model can predict potential resource shortages or mishandling during the processing based on real-time data and historical trends, issuing early warnings to help the fault handling team adjust resource allocation and processing strategies in a timely manner, ensuring efficient and accurate fault handling.
[0062] Furthermore, the work order status tracking algorithm can be optimized based on a state machine model. This includes optimizing the implementation of the event triggering mechanism, the real-time recording and updating mechanism of work order status information, and the optimization strategies for data storage and retrieval. By monitoring the lifecycle of new work orders in real time, the algorithm enables automatic recording and updating of work order status, providing data support for process management and problem tracking. The intelligence of the work order status tracking algorithm is reflected in its ability to automatically switch work order statuses through an event-driven approach and record detailed information and timestamps for each status. This automated status tracking mechanism not only improves the accuracy of data recording but also provides a basis for process optimization and user experience improvement.
[0063] Regarding the system in the above embodiments, the specific manner in which each unit module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0064] Based on the same inventive concept, embodiments of the present invention also provide an electronic device, including: a memory and a processor, wherein the processor is used to read and execute a computer program stored in the memory to implement the aforementioned device fault handling method.
[0065] Based on the same inventive concept, embodiments of the present invention also provide a computer storage medium storing computer-executable instructions, which, when executed, implement the aforementioned device fault handling method.
[0066] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0067] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional modules in the various embodiments of this invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0068] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0069] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0070] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not described in detail in a particular embodiment can be found in the relevant descriptions of other embodiments. Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for handling equipment malfunctions, characterized in that, The method includes: The alarm messages from multiple devices are acquired in real time. The alarm messages include the alarm device, alarm name, alarm type, alarm level and timestamp. The alarm association rules identify related alarm messages caused by the same fault. The related alarm messages are compressed with the existing fault group to obtain newly issued alarm messages. The work order type and workflow of the new work order are determined based on the alarm type of the newly issued alarm message. Based on the alarm level of the newly issued alarm message and the pre-configured permission information of each process step, the priority of dispatching is configured for each process step. This process involves: acquiring pre-configured execution permission information for each process step; acquiring multiple dynamic parameters for each permission role, including online status, real-time location, business skills matrix, and certificate / insurance validity status; establishing a multi-dimensional evaluation model, adjusting the weight coefficients of each dynamic parameter according to the alarm level, and calculating a weighted comprehensive score for each permission role using the dynamic parameters and their weight coefficients; and configuring task assignment priority for each process step's execution role according to the weighted comprehensive score from high to low. The multidimensional evaluation model calculates a weighted comprehensive score. S core as follows: S skill For skill matching, S geo The straight-line distance between the fault location and the real-time location of the personnel. S workload This is the workload index. S emergency Due to the level of urgency, , , , These are weighting coefficients, which are dynamically adjusted according to fault handling requirements. Based on the work order type, the corresponding work order template is invoked, the work order template is configured, the new work order is generated, the entire lifecycle of the new work order is monitored, and the processing status of the new work order is updated synchronously.
2. The method according to claim 1, characterized in that, After acquiring alarm messages from multiple devices in real time, it also includes: The alarm message is preprocessed; The preprocessing includes: Filtering and cleaning are used to remove redundant and invalid alarm messages; Standardization processing is used to process the alarm message into a preset standard format.
3. The method according to claim 1, characterized in that, The alarm association rules include: spatiotemporal association rules, semantic association rules, and statistical association rules; The spatiotemporal association rules include: time window rules based on a set dynamic time threshold, topology association rules based on device dependencies, and propagation path rules based on fault propagation modes. The semantic association rules include: keyword matching rules based on alarm type mapping table, alarm combination patterns and fault mode library based on typical fault scenarios; The statistical association rules include: device association rules based on alarm frequency thresholds and region association rules based on alarm density thresholds.
4. The method according to any one of claims 1 to 3, characterized in that, The step of identifying associated alarm messages caused by the same fault according to alarm association rules, compressing the associated alarm messages with existing fault groups, and obtaining newly issued alarm messages includes: Extract the feature parameters of the alarm message, including alarm name, alarm type and timestamp; Based on the alarm association rules, the alarm messages are analyzed for correlation, and alarm messages with the same root cause of failure are identified as associated alarm messages. Query the current active alarm set, which contains fault groups that have not yet been cleared; The associated alarm messages are matched to existing fault groups and the alarm timestamps in the groups are updated. At the same time, an alarm suppression instruction is generated to block the subsequent work order dispatch process. For new alarms that do not match an existing fault group, create a new fault group record and trigger a standard alarm. Police handling procedures.
5. The method according to claim 1, characterized in that, The step of determining the work order type and workflow of the new work order based on the alarm type of the newly issued alarm message includes: Configure the new work order type according to the alarm type of the newly issued alarm message, including: work order type name, business category, fault level, fault type, creator, and creation time; According to the alarm type of the newly issued alarm message, the workflow template of the new work order is obtained from the unified workflow platform. The workflow template has multiple workflow steps, including creation, assignment, processing, review, closure, as well as timeout, rollback and escalation steps triggered in abnormal scenarios. The workflow template of the new work order is parsed to obtain the execution order and work permission information of each process step.
6. The method according to claim 1, characterized in that, The process of calling the corresponding work order template based on the work order type, configuring the work order template, and generating the new work order includes: Obtain the corresponding work order template from the unified process platform according to the work order type; The work order template is configured with attributes, wherein the attribute information of the work order template includes: template name, work order type, business category, template ID, version number, and the process to which the work order belongs; Call the dispatch service interface to generate the new work order.
7. The method according to claim 1, characterized in that, The monitoring of the entire lifecycle of the new work order and the synchronous updating of the processing status of the new work order include: The new work order is dispatched to the corresponding business system, and the entire lifecycle of the new work order is tracked and recorded. Establish a data synchronization mechanism so that when the processing status of a new work order changes, the status update information is promptly synchronized to the unified process platform.
8. A device fault handling system, characterized in that, The system includes: The alarm compression module is used to acquire alarm messages from multiple devices in real time. The alarm messages include the alarm device, alarm name, alarm type, alarm level and timestamp. Based on the alarm association rules, it identifies related alarm messages caused by the same fault, compresses the related alarm messages with the existing fault group, and acquires newly issued alarm messages. The work order configuration module is used to determine the work order type and flow process of the new work order based on the alarm type of the newly issued alarm message, and to configure the dispatch priority for each process step based on the alarm level of the newly issued alarm message and the pre-configured permission information of each process step. This process involves: acquiring pre-configured execution permission information for each process step; acquiring multiple dynamic parameters for each permission role, including online status, real-time location, business skills matrix, and certificate / insurance validity status; establishing a multi-dimensional evaluation model, adjusting the weight coefficients of each dynamic parameter according to the alarm level, and calculating a weighted comprehensive score for each permission role using the dynamic parameters and their weight coefficients; and configuring task assignment priority for each process step's execution role according to the weighted comprehensive score from high to low. The multidimensional evaluation model calculates a weighted comprehensive score. S core as follows: S skill For skill matching, S geo The straight-line distance between the fault location and the real-time location of the personnel. S workload This is the workload index. S emergency Due to the level of urgency, , , , These are weighting coefficients, which are dynamically adjusted according to fault handling requirements. The work order generation and monitoring module is used to call the corresponding work order template based on the work order type, configure the work order template, generate the new work order, monitor the entire life cycle of the new work order, and synchronously update the processing status of the new work order.
9. The system according to claim 8, characterized in that, The work order configuration module specifically includes: The work order type and process acquisition unit is used to determine the work order type and flow process of the new work order based on the alarm type of the newly issued alarm message. The execution permission acquisition unit is used to acquire the execution permission information pre-configured for each process step; The parameter acquisition unit is used to acquire multiple dynamic parameters for each permission role, including online status, real-time location, business skills matrix, and certificate / insurance validity status. The scoring unit is used to establish a multi-dimensional evaluation model, adjust the weight coefficients of each dynamic parameter according to the alarm level, and calculate the weighted comprehensive score of each permission role using the dynamic parameters and their weight coefficients. The order dispatch priority configuration unit is used to configure the order dispatch priority for the execution role of each process step according to the weighted comprehensive score from high to low.
Citation Information
Patent Citations
Communication network fault positioning system and method based on big data analysis
CN118677759A