Operation and maintenance method and system based on big data automatic operation and maintenance platform

By building a big data automated operation and maintenance platform, using data integration and knowledge graph technology, the problem of untimely response of existing operation and maintenance platforms is solved, and the rapid and accurate handling of potential faults is achieved, and the operation and maintenance efficiency and system stability are improved.

CN120336065AActive Publication Date: 2025-07-18GUANGZHOU XIAOYUN NETWORK TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510489370.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing operation and maintenance platform is difficult to respond quickly and accurately to failures or changes that have not been foreseeable in modern IT systems, resulting in low operation and maintenance efficiency and instability in the system, affecting the normal operation of the business.

Method used

Build a big data automation operation and maintenance platform, and build an operation and maintenance knowledge graph through unified standardization and integration of multi-source heterogeneous data, use the fault feature library and historical operation and maintenance cases to conduct fault abnormality analysis and associated search, and generate automated operation and maintenance decisions.

Benefits of technology

It realizes the rapid and accurate identification and handling of potential faults, improves operation and maintenance efficiency and system stability, and ensures the normal operation of the business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336065A_ABST
    Figure CN120336065A_ABST
Patent Text Reader

Abstract

The invention provides an operation and maintenance method and system based on a big data automatic operation and maintenance platform, and the method comprises the steps: carrying out the unified standardization and integration processing of multi-source heterogeneous data collected from the operation and maintenance platform, and obtaining integrated multi-source data; constructing an operation and maintenance knowledge graph based on entities in the integrated multi-source data and entity relationships among the entities; performing fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform, and determining a potential fault event; performing association search in the operation and maintenance knowledge graph based on the potential fault event to obtain an association analysis result, and performing matching analysis on the association analysis result based on historical operation and maintenance cases in the big data storage platform to determine a target operation and maintenance case; and generating an automatic operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case. According to the invention, the operation and maintenance efficiency and the system stability are improved, and the normal operation of services is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to an operation and maintenance method and system based on a big data automated operation and maintenance platform. Background Art

[0002] In the current operation and maintenance method of the operation and maintenance platform, the widely adopted is the traditional operation and maintenance method based on rules and experience. This method relies on the rules preset by operation and maintenance personnel and the experience accumulated in the past to handle the problems that occur in the operation and maintenance process. When a specific event or abnormal index appears in the system, it is processed according to the established rules. For example, in terms of server performance monitoring, an alarm is issued when the CPU usage rate exceeds 80%. The operation and maintenance personnel judge whether to increase server resources or optimize relevant processes based on experience.

[0003] However, in the face of the increasingly complex and dynamically changing modern IT systems, it is often difficult to comprehensively cover various potential operation and maintenance scenarios in the formulation of the current method rules. With the expansion of the system scale, the complexity of business logic, and the introduction of new technologies, new fault modes and operation and maintenance requirements emerge continuously. The preset rules cannot adapt to these changes in a timely manner. When encountering unforeseen faults or system changes, it is difficult to respond quickly and accurately, resulting in potential problems that cannot be discovered and processed in a timely manner, greatly affecting the operation and maintenance efficiency and system stability, and even affecting the normal operation of the business. Summary of the Invention

[0004] The present invention provides an operation and maintenance method and system based on a big data automated operation and maintenance platform to achieve a quick and accurate response, timely discover and process potential problems, improve the operation and maintenance efficiency and system stability, and ensure the normal operation of the business.

[0005] In a first aspect, the present invention provides an operation and maintenance method based on a big data automated operation and maintenance platform, including:

[0006] Unify, standardize and integrate the multi-source heterogeneous data collected from the operation and maintenance platform based on multiple data collection interfaces to obtain the integrated multi-source data;

[0007] Construct an operation and maintenance knowledge graph based on each entity and the entity relationship between each entity in the integrated multi-source data;

[0008] Perform fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events;

[0009] Perform associated search in the operation and maintenance knowledge graph based on the potential fault events to obtain an associated analysis result, and perform matching analysis on the associated analysis result based on historical operation and maintenance cases in the big data storage platform to determine a target operation and maintenance case;

[0010] Generate an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case.

[0011] In a second aspect, the present invention further provides an operation and maintenance system based on a big data automated operation and maintenance platform, which is applied to the operation and maintenance method based on a big data automated operation and maintenance platform as described in the first aspect; the operation and maintenance system based on a big data automated operation and maintenance platform includes:

[0012] A data processing module, configured to perform unified standardization and integration processing on multi-source heterogeneous data collected from an operation and maintenance platform through multiple data collection interfaces to obtain integrated multi-source data;

[0013] A knowledge graph construction module, configured to construct an operation and maintenance knowledge graph based on each entity in the integrated multi-source data and the entity relationships between each entity;

[0014] A big data analysis module, configured to perform fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in a big data storage platform to determine potential fault events;

[0015] An operation and maintenance case matching module, configured to perform associated search in the operation and maintenance knowledge graph based on the potential fault event to obtain an associated analysis result, and perform matching analysis on the associated analysis result based on historical operation and maintenance cases in the big data storage platform to determine a target operation and maintenance case;

[0016] An operation and maintenance decision generation module, configured to generate an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case.

[0017] In a third aspect, the present invention further provides an electronic device, including: a memory, configured to store a computer software program; a processor, configured to read and execute the computer software program, thereby implementing the operation and maintenance method based on a big data automated operation and maintenance platform as described in any one of the above.

[0018] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, in which a computer software program is stored, and when the computer software program is executed by a processor, the operation and maintenance method based on a big data automated operation and maintenance platform as described in any one of the above is implemented.

[0019] In a fifth aspect, the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the operation and maintenance method based on a big data automated operation and maintenance platform as described above is implemented.

[0020] The operation and maintenance method based on the big data automated operation and maintenance platform provided by the embodiments of the present invention constructs a knowledge graph to graphically store entities and relationships in operation and maintenance data, enabling operation and maintenance personnel to intuitively understand the overall picture of the operation and maintenance system and facilitating correlation analysis. The fault event recognition is based on real-time data and a pre-set fault feature library, and can quickly and accurately detect potential fault events. Compared with the rule- and experience-based methods, it can detect problems more timely. The knowledge graph correlation analysis uses the structure of the knowledge graph to deeply mine the complex causal relationships behind fault events, rather than being limited to pre-set rules, improving the accuracy and comprehensiveness of fault cause analysis. The automated operation and maintenance decision generation combines the correlation analysis results and historical operation and maintenance cases to generate automated operation and maintenance decisions for fault events, avoiding the limitations of relying on the limited experience of operation and maintenance personnel to formulate solutions in traditional operation and maintenance. Therefore, through the coordination of multiple steps, the embodiments of the present invention can respond quickly and accurately, discover and handle potential problems in a timely manner, improve the operation and maintenance efficiency and system stability, and ensure the normal operation of the business. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic flowchart of the operation and maintenance method based on the big data automated operation and maintenance platform provided by the embodiments of the present invention;

[0022] Figure 2 is a schematic structural diagram of the operation and maintenance system based on the big data automated operation and maintenance platform provided by the embodiments of the present invention;

[0023] Figure 3 is an embodiment diagram of the electronic device provided by the embodiments of the present invention;

[0024] Figure 4 is an embodiment diagram of the computer-readable storage medium provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0026] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.

[0027] In the description of the present invention, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present invention is not necessarily construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to implement and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0028] Refer to Figure 1 , Figure 1 FIG.

[0029] Step 10: Perform unified standardization and integration processing on multi-source heterogeneous data collected from the operation and maintenance platform based on multiple data collection interfaces to obtain integrated multi-source data.

[0030] Optionally, the operation and maintenance platform includes a large number of servers, network devices, and business systems. The data collection interfaces of the servers, network devices, and business systems may be different. Therefore, the operation and maintenance platform provides data collection interfaces for the servers, network devices, and business systems. Thus, the operation and maintenance system collects multi-source heterogeneous data from the operation and maintenance platform through the multiple data collection interfaces provided by the operation and maintenance platform, such as obtaining hardware status data from the server monitoring system, obtaining business operation data from the application system logs, and other multi-source heterogeneous data.

[0031] Furthermore, since the data formats, structures, and semantics of different data in the multi-source heterogeneous data may be different, the operation and maintenance system performs unified standardization and integration processing on the collected multi-source heterogeneous data to obtain integrated multi-source data.

[0032] Among them, the standardization processing includes unifying the data format (such as unifying the time format to ISO8601), data encoding (unifying to UTF-8), and using a unified naming for data with the same meaning. The integration processing is to associate data from different sources but related to form a complete data set.

[0033] In one embodiment, the operation and maintenance system collects CPU usage rate data from the server monitoring interface in the format of "2023-10-10 10:00:00, 70%", and collects user request processing time data from the application system log interface in the format of "10 / 10 / 2023 10:00:00AM, 500ms".

[0034] Standardize the time format, all converted to "2023-10-10T10:00:00Z". Further, use the server identification information as the association key to integrate these two types of data together. The integrated multi-source data obtained is in the format of {"server_id": "srv001", "time": "2023-10-10T10:00:00Z", "CPU_usage": "70%", "request_process_time": "500ms"}.

[0035] Step 20: Based on each entity and the entity relationships among each entity in the integrated multi-source data, construct an operation and maintenance knowledge graph.

[0036] Furthermore, the operation and maintenance system performs entity extraction and relationship extraction on the integrated multi-source data to obtain each entity and the entity relationships among each entity. Among them, the entities in the embodiments of the present invention include device entities, fault entities, and business process entities, and the entity relationships include connection relationships, association relationships, reference relationships, service support relationships, dependency relationships, and causal relationships.

[0037] For example, the operation and maintenance system identifies device entities such as server A and server B, fault entities such as network failures and server crashes, and service support relationships between server A and business system 1, and causal relationships such as high CPU usage rate of server A may cause server crashes.

[0038] Furthermore, the operation and maintenance system constructs a knowledge graph using the Neo4j graph database according to each entity and the entity relationships among each entity.

[0039] In one embodiment, the server "srv001", application "app001", and database "db001" are identified as entities from the integrated multi-source data. It is found that "app001" runs on "srv001", and "srv001" connects to "db001" to obtain data. Therefore, nodes are created in Neo4j to represent these three entities respectively, and entity relationship edges of "runs on" and "connects to" are created to obtain the operation and maintenance knowledge graph.

[0040] Step 30: Perform fault anomaly analysis on the integrated multi-source data based on the preset fault feature library in the big data storage platform to determine potential fault events.

[0041] Furthermore, the preset fault feature library in the big data storage platform contains the feature patterns of various known faults. For example, in the fault feature library, it is defined that the CPU usage rate continuously exceeding 80% and the memory usage rate continuously exceeding 90% are the fault features of server performance bottleneck. Therefore, the operation and maintenance system performs matching analysis on the integrated multi-source data and the feature patterns, traverses the server performance data in the integrated multi-source data, and determines whether it conforms to these features, so as to determine potential fault events. In one embodiment, in the integrated multi-source data, for the server "srv001", the CPU usage rates are 85%, 88%, and 90% respectively within 10 consecutive minutes, and the memory usage rates are 92%, 93%, and 95% respectively. According to the performance bottleneck fault features in the preset fault feature library, the operation and maintenance system determines that there is a potential server performance bottleneck fault event for "srv001".

[0042] Step 40: Perform associated search in the operation and maintenance knowledge graph based on the potential fault event to obtain the associated analysis result, and perform matching analysis on the associated analysis result based on the historical operation and maintenance cases in the big data storage platform to determine the target operation and maintenance case.

[0043] Furthermore, the operation and maintenance system performs associated search according to the potential fault event in combination with the connection relationship, association relationship, reference relationship, service support relationship, dependency relationship, and causal relationship in the operation and maintenance knowledge graph to obtain the associated analysis result, as specifically described in steps 301 to 304.

[0044] Furthermore, the big data storage platform stores historical operation and maintenance cases, and the historical operation and maintenance cases may contain information such as fault phenomena, handling processes, and results. Therefore, the operation and maintenance system matches the associated analysis result with the historical operation and maintenance cases in the big data storage platform, and searches for similar cases by comparing information such as fault phenomena, handling processes, and results to determine the target operation and maintenance case. In one embodiment, for the potential fault event of the performance bottleneck of "srv001", it is found in the operation and maintenance knowledge graph that "app001" runs on it and the error rate of user requests of "app001" has recently increased. Among the historical operation and maintenance cases, a case is found where the server performance bottleneck causes a high error rate of the application running on it, and the problem is solved by adjusting the server resource allocation. This case is determined as the target operation and maintenance case.

[0045] Step 50: Generate an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case.

[0046] Furthermore, the event information includes fault event features, fault development logic, involved system environment information, acquisition time of operation resources, dependency relationship between operation steps, and real-time status change information of the potential fault event.

[0047] Therefore, the operation and maintenance system generates an automated operation and maintenance decision for the potential fault event according to the event information of the potential fault event and the target operation and maintenance case. The automated operation and maintenance decision at least includes operation steps, the operation sequence of the operation steps, and the operation resources required to execute the operation steps, as specifically described in steps 501 to 503.

[0048] The knowledge graph constructed in the embodiment of the present invention graphically stores entities and relationships in operation and maintenance data, enabling operation and maintenance personnel to intuitively understand the overall picture of the operation and maintenance system and facilitating correlation analysis. Fault event recognition is based on real-time data and a pre-set fault feature library, which can quickly and accurately detect potential fault events and detect problems more timely. The knowledge graph correlation analysis utilizes the structure of the knowledge graph to deeply explore the complex causal relationships behind fault events, rather than being limited to pre-set rules, improving the accuracy and comprehensiveness of fault cause analysis. The generation of automated operation and maintenance decisions generates automated operation and maintenance decisions for fault events by combining the correlation analysis results and historical operation and maintenance cases. Therefore, through the coordination of multiple steps, a quick and accurate response is made, potential problems are detected and processed in a timely manner, the operation and maintenance efficiency and the stability of the system are improved, and the normal operation of the business is guaranteed.

[0049] In one embodiment, the descriptions of steps 301 to 304 are as follows:

[0050] Step 301, perform an associated search in the operation and maintenance knowledge graph based on the potential fault event to determine the initial entities that have a connection relationship, an association relationship, or a reference relationship with the potential fault event.

[0051] Optionally, the operation and maintenance system performs an associated search in the operation and maintenance knowledge graph according to the potential fault event to determine the entities that are connected to the potential fault event through a connection relationship (such as a physical connection), an association relationship (such as a business association), or a reference relationship (such as a configuration reference), and obtains the initial entities. The initial entities in the embodiment of the present invention can be understood as an initial entity set. In one embodiment, the potential fault event is that the CPU usage rate of the server "srv001" is too high. In the operation and maintenance knowledge graph, "srv001" is connected to the network switch "switch001" through the "connected to" relationship, is associated with the application program "app001" through the "running" association relationship, and its configuration file references the path of the storage device "storage001", then "switch001", "app001", and "storage001" are the initial entities.

[0052] Step 302, screen out the first target entities that have a service support relationship with the potential fault event from the initial entities, and perform a path exploration along the service support relationship based on the first target entities to obtain the service support paths of the first target entities.

[0053] Further, in the initial entity set, the operation and maintenance system filters out the first target entities that have a service support relationship with potential fault events. The first target entities are those that provide key service support to the subjects involved in potential fault events.

[0054] Further, the operation and maintenance system starts from the first target entities along the service support relationship and gradually explores the upstream and downstream related entities to obtain the service support paths of the first target entities, as specifically described in steps 3021 to 3024. Continuing the above embodiment, among the initial entities, "app001" plays a key service support role in the business operation of "srv001", so "app001" is determined as the first target entity. Exploring along the service support relationship, it is found that "app001" obtains data from the database "db001" and is connected to the user request end through the load balancer "lb001", obtaining the service support paths of "app001→db001" and "app001→lb001".

[0055] Step 303: Based on each second target entity in the service support path and the dependency relationships between each second target entity, construct a dependency relationship subgraph.

[0056] Further, based on the service support paths obtained in step 302, the operation and maintenance system extracts each second target entity in the service support paths and the dependency relationships between each second target entity. Among them, the dependency relationships in the embodiments of the present invention can be data dependencies, resource dependencies, etc.

[0057] Further, the operation and maintenance system constructs a dependency relationship subgraph according to each second target entity and the dependency relationships between each second target entity, as specifically described in steps 3031 to 3034.

[0058] Continuing the above embodiment, in the service support paths of "app001→db001" and "app001→lb001", "db001" and "lb001" are the second target entities. "app001" depends on "db001" to provide data, "app001" depends on "lb001" for user request distribution, and "lb001" may depend on "db001" for user authentication data query. In the dependency relationship subgraph constructed by the operation and maintenance system, "app001", "db001" and "lb001" are nodes, and the dependency relationships between them are represented by directed edges, such as "app001→db001 (data dependency)", "app001→lb001 (request distribution dependency)", "lb001→db001 (authentication data query dependency)".

[0059] Step 304: Starting from the potential failure event, determine the event backtracking time window based on the occurrence time of the potential failure event, and perform root cause analysis based on the dependency subgraph and the event backtracking time window to determine the correlation analysis result of the cause chain that led to the occurrence of the potential failure event.

[0060] Further, starting from the potential failure event, the operation and maintenance system determines an event backtracking time window according to the occurrence time of the potential failure event, for example, 1 hour before the failure occurs. In the dependency subgraph, within this time window, the operation and maintenance system analyzes the states and events of each entity, traces back along the dependencies from the potential failure event in the reverse direction, and searches for a series of events and factors that may have caused the failure to obtain the correlation analysis result of the cause chain, as specifically described in Steps 3041 to 3044.

[0061] In one embodiment, the potential failure event "srv001" has an excessive CPU usage rate at 10:00. The determined backtracking time window is 9:00 - 10:00. It is found in the dependency subgraph that at 9:30, "db001" performed a large-scale data update operation, resulting in an instantaneous increase in the data request processing volume of "app001", which in turn increased the request load of "srv001" for processing "app001" requests, ultimately leading to an excessive CPU usage rate. The correlation analysis result of the cause chain "db001 large-scale data update → app001 request processing volume increase → srv001 excessive CPU usage rate" is determined.

[0062] The embodiment of the present invention utilizes the structure of the knowledge graph to be able to start from the potential failure event, deeply explore the related entities and relationships, deeply explore the complex causal relationships behind the failure event, construct the cause chain of the failure, enabling the operation and maintenance personnel to clearly understand the root cause of the failure, no longer being limited to the pre-set rules, and improving the accuracy and comprehensiveness of the failure cause analysis.

[0063] In one embodiment, the descriptions of Steps 3021 to 3024 are as follows:

[0064] Step 3021: Based on the first target entity, perform the first round of extended exploration in the operation and maintenance knowledge graph according to the service support relationship to obtain the first adjacent entity that has a direct service support relationship with it.

[0065] Optionally, after the operation and maintenance system determines the first target entity, according to the service support relationship in the operation and maintenance knowledge graph, it performs the first round of extended exploration starting from the first target entity, finds the entity that directly has a service support relationship with the first target entity, and obtains the first adjacent entity.

[0066] In one embodiment, the first target entity is "app001". In the operation and maintenance knowledge graph, "app001" is connected to the database "db001" through a service support relationship ("app001" obtains data from "db001"), and is also connected to the load balancer "lb001" ("lb001" distributes user requests for "app001"). Then, in the first round of extended exploration, "db001" and "lb001" are the first adjacent entities of "app001".

[0067] Step 3022: Using the first target entity as the root node and the first adjacent entities as the first-layer nodes, for each first-layer node, based on the first-layer node, conduct a second-round extended exploration in the operation and maintenance knowledge graph according to the service support relationship to obtain the second adjacent entities that have a direct service support relationship with it. Loop in turn. If after n rounds of extended exploration, there are no entities with a direct service support relationship for the (n - 1)-th layer nodes, then generate an extended exploration path of the first target entity based on the first target entity and the entities and the directions of the service support relationships obtained in each round of extended exploration.

[0068] Furthermore, the operation and maintenance system uses the first target entity as the root node and takes the first adjacent entities obtained from the first-round extended exploration as the first-layer nodes. For each first-layer node, conduct an extended exploration in the knowledge graph again according to the service support relationship to find the entities that have a direct service support relationship with them, that is, the second adjacent entities. This process will loop continuously, and each round of extended exploration is based on the previous layer of nodes and continues to expand outward. When the n-th round of extended exploration is carried out, if there are no new entities connected through the service support relationship for the (n - 1)-th layer nodes, the extended exploration process stops.

[0069] Furthermore, the operation and maintenance system constructs an extended exploration path starting from the first target entity according to the first target entity and the entities and the directions of the service support relationships found in each round of extended exploration.

[0070] In one embodiment, using "app001" as the root node, in the second-round extended exploration of its first-layer node "db001", it is connected to the storage server "storage001" through a service support relationship (data storage dependency), and in the second-round extended exploration of "lb001", it is connected to the network switch "switch001" through a service support relationship (network connection dependency). Continuing the exploration, there are no new adjacent entities for "storage001" based on the service support relationship, nor for "switch001". At this time, two extended exploration paths are generated: one is "app001 → db001 → storage001", and the other is "app001 → lb001 → switch001".

[0071] Step 3023: For any two first extended exploration paths and second extended exploration paths, if there are identical entities in the first extended exploration path and the second extended exploration path, then merge the identical entities and keep the different entities unchanged to obtain the fused exploration paths.

[0072] Furthermore, among the multiple generated extended exploration paths, the operation and maintenance system will conduct a comparative analysis on any two paths. If there are identical entities in the two paths, these identical entities will be merged into one node, while the different entities will remain as they are. Therefore, redundant information in the paths can be eliminated, making the path structure more concise and clear, and at the same time highlighting the differences and connections between different paths. Continuing with the above embodiment, there are two extended exploration paths. Path one is "app001 → db001 → storage001", and path two is "app001 → lb001 → db001". There is the entity "db001" in both paths. After fusion, the obtained fused exploration paths are "app001 → (db001) → storage001" and "app001 → lb001 → (db001)", where "db001" in the parentheses represents the merged node.

[0073] Step 3024: Integrate all the fused exploration paths to obtain the service support path of the first target entity.

[0074] Furthermore, the operation and maintenance system integrates all the exploration paths after the fusion process to obtain the service support path of the first target entity. Continuing with the above embodiment, after fusion, there are three exploration paths: Path one is "app001 → (db001) → storage001", path two is "app001 → lb001 → (db001)", and path three is "app001 → monitor001" ("monitor001" is used to monitor the running status of "app001"). The service support path of the first target entity "app001" obtained after integration is a comprehensive structure, showing the service support relationships between "app001" and "db001" (associated in different ways), "storage001", "lb001", and "monitor001", such as "app001 → db001 (data acquisition) → storage001 (data storage)", "app001 → lb001 (request distribution) → db001 (data query)", and "app001 → monitor001 (status monitoring)".

[0075] In the embodiment of the present invention, starting from the initial first-round exploration to determine directly related entities, through multiple rounds of extended exploration to construct a complete path, and then to fusion and integration operations, the finally formed service support path provides a clear visual view for the operation and maintenance personnel, showing the position of the first target entity in the entire operation and maintenance ecosystem and its close connection with other entities. Therefore, in the fault troubleshooting scenario, when a problem occurs with the first target entity, the operation and maintenance personnel can quickly locate each link that may affect it based on this complete service support path. Whether it is the upstream data provider, the downstream service recipient, or the intermediate transmission and monitoring links, they can be clearly seen at a glance. Therefore, the operation and maintenance personnel can respond quickly and accurately, timely discover and handle potential problems, improve the operation and maintenance efficiency and the stability of the system, and ensure the normal operation of the business.

[0076] In one embodiment, the descriptions of steps 3031 to 3034 are as follows:

[0077] Step 3031, for each second target entity in the service support path, traverse in the service support path the third target entities that have a direct dependency relationship with it, and based on the second target entity, the third target entity, and the nodes, with the direct dependency relationship as the node edge, construct an initial relationship subgraph.

[0078] Optionally, the operation and maintenance system traverses within the service support path for each second target entity in the path to find the third target entities that have a direct dependency relationship with the second target entity. After determining the third target entity and the direct dependency relationship, the operation and maintenance system uses the second target entity and the third target entity as nodes, and the direct dependency relationship between the second target entity and the third target entity as the edge connecting the nodes to construct a preliminary graph structure, that is, the initial relationship subgraph, where the initial relationship subgraph shows the most direct and basic dependency relationships in the service support path.

[0079] In one embodiment, the service support paths are "app001 → db001 → storage001" and "app001 → lb001 → switch001". For the second target entity "db001", it has a direct dependency on data provision with "app001" (i.e., "app001" depends on "db001" to provide data), and a direct dependency on data storage with "storage001" (i.e., "db001" depends on "storage001" to store data). Then, when constructing the initial relationship sub-graph, "app001", "db001", and "storage001" become nodes, and "app001 → db001 (data provision)", "db001 → storage001 (data storage)" become node edges. Similarly, for "lb001", it has a direct dependency on request distribution with "app001" and a direct dependency on network connection with "switch001", and the corresponding nodes and edges will also be added to the initial relationship sub-graph.

[0080] Step 3032: Starting from the second target entity in the service support path, traverse to determine the fourth target entity reachable through indirect dependencies within a preset step size, as well as the intermediate entities between the second target entity and the fourth target entity, and add corresponding node edges to the initial relationship sub-graph based on the fourth target entity and the intermediate entities to obtain the first updated relationship sub-graph.

[0081] Furthermore, the operation and maintenance system starts from each second target entity in the service support path and traverses according to a preset step size (for example, the step size is set to 2, indicating to find indirect dependencies within a certain range). The operation and maintenance system will determine the fourth target entity reachable through indirect dependencies within this preset step size, and at the same time find the intermediate entities between the second target entity and the fourth target entity. Further, the operation and maintenance system adds the newly discovered fourth target entity and intermediate entities to the initial relationship sub-graph with corresponding node edges to update the initial relationship sub-graph and obtain the first updated relationship sub-graph, so that the relationship sub-graph not only contains direct dependencies but also covers indirect dependencies within a certain range.

[0082] Continuing with the above service support path as an example, assume the preset step size is 2. For the second target entity "db001", through the indirect dependency relationship (where "app001" depends on "db001" to obtain data, and "db001" depends on "storage001" to store data, so "app001" indirectly depends on "storage001"), within the step size of 2, the fourth target entity "storage001" can be reached, and the intermediate entity is "db001". In the initial relationship subgraph, there are already edges of "app001→db001" and "db001→storage001". Now, to represent the indirect dependency relationship from "app001" to "storage001", the system can add a dotted line edge or use a certain identifier to indicate this indirect dependency relationship (such as adding a note to explain the indirect connection through "db001"). Similarly, for "lb001", if within the step size of 2, it is found that "app001" indirectly depends on the DNS server of the network provider (for example, as the fourth target entity) through "lb001" and "switch001", and the intermediate entities are "lb001" and "switch001", then corresponding nodes (DNS server) and edges representing the indirect dependency relationship will also be added to the initial relationship subgraph.

[0083] Step 3033: Detect the cyclic dependency relationship in the first updated relationship subgraph to obtain a detection result, and based on the detection result, process the first updated relationship subgraph to remove the cyclic dependency relationship, obtaining a second updated relationship subgraph.

[0084] Furthermore, a cyclic dependency relationship refers to the existence of a path in the graph structure where starting from a certain entity, after a series of dependency relationships, it returns to the entity itself. This situation brings complexity to the analysis and processing of dependency relationships. Therefore, the operation and maintenance system traverses the first updated relationship subgraph through a specific algorithm (such as a variant of the depth-first search algorithm) to detect whether there is a cyclic dependency relationship in the first updated relationship subgraph. If a cyclic dependency relationship is detected, the operation and maintenance system will process the first updated relationship subgraph based on the detection result to remove these cyclic dependency relationships, obtaining a second updated relationship subgraph. Common removal methods can be to cut off one of the dependency edges or adjust the representation of the dependency relationship.

[0085] In one embodiment, in the first updated relationship subgraph, due to some complex business logic, a cyclic dependency relationship of "app001 → db001 → app001" appears (it may be that some data updates of "app001" require processing by "db001", and after "db001" processes, it needs to feedback specific information to "app001"). To remove the cycle, the system can decide to cut the edge of "db001 → app001" according to the actual business importance (for example, in business, the feedback of "db001" to "app001" can be achieved through other means), and obtain a second updated relationship subgraph without cyclic dependency relationships.

[0086] Step 3034, hierarchically partition the entities in the second updated relationship subgraph according to the direct dependency relationship and the indirect dependency relationship with a preset step size, to obtain a dependency relationship subgraph. The dependency relationship subgraph is hierarchically partitioned downward starting from the entities without incoming edges as the top layer, so that the entities at each layer depend on the entities in its upper layer.

[0087] Furthermore, the operation and maintenance system hierarchically partitions the entities in the second updated relationship subgraph according to the direct dependency relationship and the indirect dependency relationship with a preset step size. Among them, the principle of partitioning is to use the entities without incoming edges (that is, not depending on other entities) as the top layer, and then hierarchically partition downward, so that the entities at each layer depend on the entities in its upper layer, to obtain a dependency relationship subgraph.

[0088] In one embodiment, in the second updated relationship subgraph, "storage001", "switch001", and the DNS server of the network provider have no incoming edges (they do not depend on other entities in the service support path), so they are partitioned as the top layer. "db001" depends on "storage001", and "lb001" depends on "switch001", so "db001" and "lb001" are partitioned as the second layer. "app001" depends on "db001" and "lb001", so "app001" is partitioned as the third layer. The finally obtained dependency relationship subgraph presents a clear hierarchical structure, such as the top layer being "storage001", "switch001", and the DNS server; the second layer being "db001", "lb001"; and the third layer being "app001".

[0089] The embodiments of the present invention can construct a dependency subgraph with distinct levels, no cyclic dependencies, and comprehensively reflecting entity dependencies from complex service support paths. This subgraph systematically integrates and sorts out the dependency information originally scattered in the service support paths. Therefore, in the scenarios of fault troubleshooting and system optimization, operation and maintenance personnel can rely on this dependency subgraph to quickly locate the fault propagation path and potential impact scope, and thus can respond quickly and accurately, discover and handle potential problems in a timely manner, improve the operation and maintenance efficiency and system stability, and ensure the normal operation of the business.

[0090] In one embodiment, the descriptions of steps 3041 to 3044 are as follows:

[0091] Step 3041: Using the entity corresponding to the potential fault event as the central entity, trace back reversely along the causal relationship in the dependency subgraph to determine the predecessor entities that caused the potential fault event to occur.

[0092] Optionally, the operation and maintenance system sets the entity corresponding to the potential fault event as the central entity, and in the already constructed dependency subgraph, traces back reversely from the central entity along the causal relationship represented in the dependency subgraph. Among them, the causal relationship is reflected as the dependency direction between entities in the dependency subgraph, such as data flow direction, service call direction, etc. Therefore, by searching reversely along these causal relationships, determine the predecessor entities that logically preceded the potential fault event and had an impact on it, and obtain the predecessor entities that caused the potential fault event to occur.

[0093] In one embodiment, for example, the potential fault event is that "app001" responds slowly, and "app001" is the central entity. In the dependency subgraph, "app001" depends on "db001" to obtain data and depends on "lb001" for request distribution. By tracing back reversely along the causal relationship, it is found that "db001" and "lb001" are the predecessor entities of "app001". Because if the data query of "db001" is slow or the request distribution of "lb001" is abnormal, it may cause the problem of slow response of "app001".

[0094] Step 3042: Generate a causal relationship path based on the central entity, the predecessor entity, and the reverse causal relationship direction from the central entity to the predecessor entity.

[0095] Furthermore, based on the determined central entity and predecessor entities, the operation and maintenance system generates a causal relationship path from the central entity to the predecessor entities. The causal relationship path clarifies the associations between entities and also identifies the direction of the causal relationship, that is, it points from the entity where the potential failure event occurs to the predecessor entities that may cause the failure, constructing a preliminary fault propagation logic framework, which helps to intuitively understand how the fault affects the central entity (potential failure event entity) from upstream entities.

[0096] Continuing with the above embodiment, taking "app001" as the central entity and "db001" and "lb001" as the predecessor entities, two causal relationship paths are generated. One is "app001←db001 (data acquisition dependency)", indicating that "app001" depends on "db001" for data acquisition, and problems with "db001" may affect "app001"; the other is "app001←lb001 (request distribution dependency)", indicating that the request distribution of "app001" depends on "lb001", and abnormalities in "lb001" may cause "app001" to respond slowly.

[0097] Step 3043, based on the business data characteristics of the predecessor entities and the changes in the business data of potential failure events within the event backtracking time window, perform correlation analysis to determine the business data correlation result.

[0098] Furthermore, the operation and maintenance system analyzes the business data characteristics of each predecessor entity in terms of business data. The business data characteristics may include data volume, data update frequency, request processing duration, etc. At the same time, the operation and maintenance system compares the changes in the business data of potential failure events (central entities) within the event backtracking time window. By performing correlation analysis on the business data characteristics of the predecessor entities and the changes in the business data of potential failure events, find the internal connection between the two, and judge whether the business data status of the predecessor entity is causally related to the occurrence of potential failure events to determine the business data correlation result.

[0099] In one embodiment, for example, the event backtracking time window is 30 minutes before the fault occurs. For the precursor entity "db001", it is found that the data update frequency has suddenly increased significantly within these 30 minutes, changing from the original 10 updates per hour to 1 update per 10 minutes. And for "app001" within the same time window, the average response time has extended from the original 200 ms to 800 ms. Through comparative analysis, it is found that there is an association between the abnormal increase in the data update frequency of "db001" and the sluggish response of "app001", that is, the business data association result shows that the change in the data update frequency of "db001" may be a factor leading to the sluggish response of "app001". For "lb001", within the event backtracking time window, its request distribution success rate has dropped from 99% to 90%, and the number of valid requests received by "app001" has also decreased accordingly, indicating that there is also an association between the drop in the request distribution success rate of "lb001" and the sluggish response of "app001".

[0100] Step 3044, integrate the causal relationship path and the business data association result to obtain the association analysis result of the cause chain leading to the occurrence of the potential fault event.

[0101] Furthermore, the operation and maintenance system integrates the previously generated causal relationship path and the determined business data association result to form a complete association analysis result of the cause chain leading to the occurrence of the potential fault event. Among them, the association analysis result includes both the logical dependency relationship (causal relationship path) between entities and the association information at the business data level.

[0102] In one embodiment, integrate the causal relationship path "app001 ← db001 (data acquisition dependency)" with the business data association result "there is an association between the abnormal increase in the data update frequency of db001 and the sluggish response of app001", and integrate the causal relationship path "app001 ← lb001 (request distribution dependency)" with the business data association result "there is an association between the drop in the request distribution success rate of lb001 and the sluggish response of app001". The final association analysis result of the cause chain leading to the occurrence of the potential fault event of the sluggish response of "app001" is: "app001" depends on "db001" for data acquisition, and the abnormal increase in the recent data update frequency of "db001" affects the response of "app001"; at the same time, "app001" depends on "lb001" for request distribution, and the drop in the request distribution success rate of "lb001" also has a negative impact on the response of "app001".

[0103] The embodiments of the present invention can start from two key elements, namely the dependency subgraph and the event backtracking time window, to comprehensively and deeply explore the cause chain leading to potential failure events, providing a structured and data-driven fault diagnosis perspective for operation and maintenance personnel. Therefore, in actual operation and maintenance scenarios, operation and maintenance personnel can quickly and accurately respond based on detailed correlation analysis results, timely discover and handle potential problems, improve operation and maintenance efficiency and system stability, and ensure the normal operation of the business.

[0104] In one embodiment, the descriptions of steps 501 to 503 are as follows:

[0105] Step 501: Split the target operation and maintenance case into multiple initial case segments, and map the operation and maintenance case features of each initial case segment to the fault event features in the event information to obtain a mapping matching result.

[0106] Optionally, the operation and maintenance system splits the target operation and maintenance case into multiple initial case segments according to a certain logic or operation process. Among them, these initial case segments can be a series of continuous operation and maintenance operations or the processing links for specific fault phenomena.

[0107] Furthermore, for each initial case segment, the operation and maintenance system extracts its operation and maintenance case features. Among them, the operation and maintenance case features include but are not limited to processing operation types, system components involved, processing time, etc. At the same time, fault event features are extracted from the event information of potential failure events, such as fault types, locations where faults occur, fault impact ranges, etc.

[0108] Furthermore, the operation and maintenance system maps and compares the operation and maintenance case features of each initial case segment with the fault event features one by one to judge the matching degree between them and obtain a mapping matching result. The obtained mapping matching result is used to measure the similarity between the initial case segment and the potential failure event at the feature level.

[0109] In one embodiment, for example, the target operation and maintenance case is the historical record of solving the performance bottleneck of server "srv001", which includes operations such as shutting down non-critical processes, increasing server memory, and optimizing database queries. It is split into three initial case fragments: fragment one is "shut down non-critical processes", fragment two is "increase server memory", and fragment three is "optimize database queries". The potential failure event is that the performance bottleneck of "srv001" appears again currently. The characteristics of the failure event include too high CPU usage (above 85%), memory usage approaching saturation (above 90%), and slow response of the involved application "app001". For fragment one, the operation and maintenance case characteristic is "operation type - process management, involved component - server process". When mapping with the failure event characteristics, it is found that the "server process" is associated with "srv001", and the high CPU usage may be related to the running of too many non-critical processes, with a certain degree of match. The operation and maintenance case characteristic of fragment two is "operation type - hardware resource adjustment, involved component - server memory", which highly matches the characteristic of approaching memory saturation in the failure event. The operation and maintenance case characteristic of fragment three is "operation type - database optimization, involved component - database", which has a certain association with the slow response of "app001" (possibly related to slow database queries) in the failure event, and the corresponding mapping and matching results are obtained.

[0110] Step 502: Based on the mapping and matching results of each initial case fragment, perform case fragment screening to determine the target case fragment, and construct a case association path based on the mapping and matching results of the target case fragment. The case association path represents the connection between different characteristics of the potential failure event and the target operation and maintenance case.

[0111] Furthermore, based on the mapping and matching results of each initial case fragment obtained in step 501, the operation and maintenance system selects those initial case fragments with a relatively high degree of match with the characteristics of the potential failure event as the target case fragments according to the preset matching degree threshold or comprehensive evaluation rules.

[0112] Further, the operation and maintenance system constructs a case association path based on the mapping and matching results of these target case segments. Among them, the case association path graphically or structurally shows how different characteristics of potential fault events are related to each processing link in the target operation and maintenance case, presenting the logical relationship from the fault phenomenon to the corresponding processing solution. In an embodiment, the matching degree threshold is 60%. According to the mapping and matching results in step 501, the matching degree of segment one is evaluated as 50%, the matching degree of segment two is 80%, and the matching degree of segment three is 70%. Therefore, segment two and segment three are determined as target case segments. When constructing the case association path, for segment two, since the fault event feature that the memory usage rate is close to saturation is closely related to the operation of "increasing the server memory", an association edge is established from "memory usage rate close to saturation" to "increasing the server memory". For segment three, because the slow response of "app001" is related to "optimizing database queries", an association edge is established from "slow response of app001" to "optimizing database queries", and the case association path is obtained.

[0113] Step 503, generate an automated operation and maintenance decision based on the case association path.

[0114] Further, the operation and maintenance system generates an automated operation and maintenance decision according to the association path, as specifically described in steps 5031 to 5034.

[0115] The embodiment of the present invention can closely combine historical target operation and maintenance cases with current potential fault events to generate targeted and operable automated operation and maintenance decisions. Therefore, in actual operation and maintenance, it can respond quickly and accurately, discover and handle potential problems in a timely manner, improve the operation and maintenance efficiency and system stability, and ensure the normal operation of the business.

[0116] In an embodiment, the descriptions of steps 5031 to 5034 are as follows:

[0117] Step 5031, screen out the target operation steps of potential fault events from the operation steps of the target case segment based on the case association path, and deduce the initial operation sequence of the target operation steps based on the fault development logic in the event information and the logical relationship between the operation steps in the target case segment.

[0118] Optionally, based on the established case association path, the operation and maintenance system filters out the target operation steps directly related to the potential fault event from among the numerous operation steps included in the target case segment. Here, the target operation steps are the core actions for resolving the current potential fault. At the same time, the operation and maintenance system analyzes the fault development logic in the event information, such as how the fault gradually affects other components from the anomaly of one component, and the internal logical relationship between the operation steps in the target case segment, such as that certain operations must be executed after other operations are completed. Therefore, the operation and maintenance system derives the preliminary execution order of the target operation steps, i.e., the initial operation order, based on the fault development logic and the logical relationship between the operation steps in the target case segment.

[0119] Continuing with the example of the performance bottleneck problem of server "srv001", the target case segments are "increasing server memory" and "optimizing database queries". In the "increasing server memory" segment, the operation steps include checking memory slots, selecting a suitable memory module, installing memory, etc.; the "optimizing database queries" segment includes steps such as analyzing query statements, creating indexes, and adjusting query parameters. According to the case association path, the target operation steps directly related to the performance bottleneck of "srv001" are "selecting a suitable memory module", "installing memory", "analyzing query statements", and "creating indexes". From the perspective of the fault development logic, insufficient server memory and slow database queries jointly cause the performance bottleneck. And from the operation logic of the target case segment, a suitable module needs to be selected before increasing the memory, and the query statement needs to be analyzed before optimizing the database query. So the derived initial operation order is: first, "select a suitable memory module", then "install memory", then "analyze query statements", and finally "create indexes".

[0120] Step 5032, determine the operation resources required to execute the target operation steps based on the system environment information involved in the event information.

[0121] Furthermore, the operation and maintenance system comprehensively evaluates various types of operation resources required to execute each target operation step according to the system environment information involved in the event information. Here, the operation resources cover hardware resources, such as additional memory modules and free server slots; software resources, such as database management tools and memory detection software; and human resources, including operation and maintenance personnel and database administrators with corresponding skills.

[0122] Continuing with the above embodiment, for the target operation step of "selecting a suitable memory module", the required operation resources include a memory specification manual (software resource, used to select a suitable memory according to the server configuration), and an operation and maintenance personnel familiar with the server memory configuration (human resource). The step of "installing the memory" requires a memory module (hardware resource), installation tools such as a screwdriver (hardware resource), and the above-mentioned operation and maintenance personnel. The step of "analyzing query statements" requires database management software (software resource), and a database administrator with database knowledge (human resource). The step of "creating an index" also depends on the database management software and the database administrator.

[0123] Step 5033: Adjust the initial operation sequence based on the acquisition time of the operation resources in the event information, the dependency relationship between the target operation steps, and the real-time status change information of the potential failure event to obtain the target operation sequence.

[0124] Furthermore, the operation and maintenance system comprehensively considers the acquisition time of the operation resources in the event information, that is, the time required to obtain the required hardware, software, or allocate human resources. At the same time, it analyzes the dependency relationship between the target operation steps. For example, some operations must wait for specific resources to be ready before they can start, or the completion of one operation is a prerequisite for the start of another operation. In addition, it also monitors the status change information of potential failure events in real time, such as whether the severity of the failure increases or the affected range expands. By comprehensively weighing these three factors, the operation and maintenance system adjusts the initial operation sequence to obtain a target operation sequence that is more in line with the actual situation and the urgency of fault handling.

[0125] In one embodiment, for example, it takes 2 hours to obtain a suitable memory module, while the database administrator can arrive within 1 hour. At the same time, it is found that as time goes by, the CPU usage rate of "srv001" continues to rise, and the performance bottleneck problem intensifies. Considering that database query optimization can perform some operations first with the existing memory, and the database administrator arrives earlier than the memory module acquisition time. Therefore, the initial operation sequence is adjusted, and the target operation sequence becomes: first "analyze query statements", then "create an index" (these two steps can start immediately after the database administrator arrives), and then after the memory module is obtained, perform the operations of "selecting a suitable memory module" and "installing the memory".

[0126] Step 5034: Integrate the target operation steps, operation sequence, and operation resources to generate an automated operation and maintenance decision for the potential failure event.

[0127] Furthermore, the operation and maintenance system integrates the determined target operation steps, the optimized target operation sequence, and the operation resources required to execute these operations to generate an automated operation and maintenance decision for potential fault events. Among them, the automated operation and maintenance decision is a detailed task list that clearly lists the specific content of each operation step, the execution sequence, and the detailed information of the required resources.

[0128] In one embodiment, the automated operation and maintenance decision generated after integration is as follows:

[0129] The first step: The database administrator uses database management software to analyze the database query statements related to "app001" within 1 hour.

[0130] The second step: After the analysis is completed, the database administrator immediately uses database management software to create an index.

[0131] The third step: After obtaining a suitable memory module (expected after 2 hours), the operation and maintenance personnel select a suitable memory module with reference to the memory specification manual.

[0132] The fourth step: The operation and maintenance personnel use tools such as screwdrivers to install the memory module into the "srv001" server.

[0133] The embodiment of the present invention can generate highly customized, practical, and highly operable automated operation and maintenance decisions. Therefore, when facing potential fault events, instead of mechanically applying historical cases, it fully combines the development logic of the current fault, the system environment, and the resource acquisition situation, dynamically adjusts the operation steps and sequence, enabling the operation and maintenance personnel to respond quickly and accurately, discover and handle potential problems in a timely manner, improving the operation and maintenance efficiency and system stability, and ensuring the normal operation of the business.

[0134] Furthermore, the operation and maintenance system provided by the present invention based on the big data automated operation and maintenance platform is described below. The operation and maintenance system based on the big data automated operation and maintenance platform described below can be mutually referred to and corresponding to the operation and maintenance method based on the big data automated operation and maintenance platform described above.

[0135] Optionally, referring to Figure 2 , Figure 2 is a schematic structural diagram of the operation and maintenance system provided by the present invention based on the big data automated operation and maintenance platform. The operation and maintenance system based on the big data automated operation and maintenance platform includes.

[0136] The data processing module 210 is used to uniformly standardize and integrally process the multi-source heterogeneous data collected from the operation and maintenance platform through multiple data collection interfaces to obtain the integrated multi-source data;

[0137] The knowledge graph construction module 220 is used to construct an operation and maintenance knowledge graph based on each entity and the entity relationships between entities in the integrated multi-source data;

[0138] The big data analysis module 230 is used to perform fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events;

[0139] The operation and maintenance case matching module 240 is used to perform associated search in the operation and maintenance knowledge graph based on potential fault events to obtain an associated analysis result, and perform matching analysis on the associated analysis result based on historical operation and maintenance cases in the big data storage platform to determine a target operation and maintenance case;

[0140] The operation and maintenance decision generation module 250 is used to generate an automated operation and maintenance decision for potential fault events based on the event information of potential fault events and the target operation and maintenance case.

[0141] The embodiment of the present invention realizes a fast and accurate response, timely discovers and processes potential problems, improves the operation and maintenance efficiency and the stability of the system, and ensures the normal operation of the business.

[0142] Please refer to Figure 3 , Figure 3 which is the embodiment diagram of the electronic device provided by the embodiment of the present invention. As Figure 3 shown, the embodiment of the present invention provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored on the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, the following steps are implemented:

[0143] Perform unified standardization and integration processing on multi-source heterogeneous data collected from the operation and maintenance platform based on multiple data collection interfaces to obtain integrated multi-source data;

[0144] Construct an operation and maintenance knowledge graph based on each entity and the entity relationships between entities in the integrated multi-source data;

[0145] Perform fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events;

[0146] Perform associated search in the operation and maintenance knowledge graph based on potential fault events to obtain an associated analysis result, and perform matching analysis on the associated analysis result based on historical operation and maintenance cases in the big data storage platform to determine a target operation and maintenance case;

[0147] Generate an automated operation and maintenance decision for potential fault events based on the event information of potential fault events and the target operation and maintenance case.

[0148] Please refer toFigure 4 , Figure 4 This is an embodiment diagram of the computer-readable storage medium provided by the embodiments of the present invention. As Figure 4 shown, this embodiment provides a computer-readable storage medium 400, on which a computer program 311 is stored. When the computer program 311 is executed by a processor, the following steps are implemented:

[0149] Unify and standardize and integrate multi-source heterogeneous data collected from the operation and maintenance platform based on multiple data collection interfaces to obtain integrated multi-source data;

[0150] Construct an operation and maintenance knowledge graph based on each entity in the integrated multi-source data and the entity relationships between each entity;

[0151] Perform fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events;

[0152] Perform associated search in the operation and maintenance knowledge graph based on the potential fault event to obtain an associated analysis result, and perform matching analysis on the associated analysis result based on historical operation and maintenance cases in the big data storage platform to determine the target operation and maintenance case;

[0153] Generate an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case.

[0154] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the operation and maintenance method based on the big data automated operation and maintenance platform provided by the above various methods. The method includes:

[0155] Unify and standardize and integrate multi-source heterogeneous data collected from the operation and maintenance platform based on multiple data collection interfaces to obtain integrated multi-source data;

[0156] Construct an operation and maintenance knowledge graph based on each entity in the integrated multi-source data and the entity relationships between each entity;

[0157] Perform fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events;

[0158] Perform associated search in the operation and maintenance knowledge graph based on the potential fault event to obtain an associated analysis result, and perform matching analysis on the associated analysis result based on historical operation and maintenance cases in the big data storage platform to determine the target operation and maintenance case;

[0159] Generate an automated operation and maintenance decision for a potential fault event based on the event information of the potential fault event and the target operation and maintenance case.

[0160] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0161] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, and of course also by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An operation and maintenance method based on a big data automated operation and maintenance platform, characterized in that, Including: Unify, standardize and integrate multi-source heterogeneous data collected from the operation and maintenance platform based on multiple data collection interfaces to obtain integrated multi-source data; Construct an operation and maintenance knowledge graph based on each entity in the integrated multi-source data and the entity relationships between each entity; Perform fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events; Perform associated search in the operation and maintenance knowledge graph based on the potential fault events to obtain an associated analysis result, and perform matching analysis on the associated analysis result based on historical operation and maintenance cases in the big data storage platform to determine a target operation and maintenance case; Generate an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case.

2. The operation and maintenance method of the operation and maintenance platform based on big data automation according to claim 1, wherein, The performing associated search in the operation and maintenance knowledge graph based on the potential fault events to obtain an associated analysis result includes: Perform associated search in the operation and maintenance knowledge graph based on the potential fault events to determine initial entities that have a connection relationship, an association relationship, or a reference relationship with the potential fault events; Screen out first target entities that have a service support relationship with the potential fault events from the initial entities, and explore paths along the service support relationship based on the first target entities to obtain the service support paths of the first target entities; Construct a dependency relationship subgraph based on each second target entity in the service support path and the dependency relationships between each second target entity; Taking the potential fault event as the starting point, determine an event backtracking time window based on the event occurrence time of the potential fault event, and perform source analysis based on the dependency relationship subgraph and the event backtracking time window to determine the associated analysis result of the cause chain that leads to the occurrence of the potential fault event.

3. The operation and maintenance method of the operation and maintenance platform based on big data automation according to claim 2, characterized in that, The performing source analysis based on the dependency relationship subgraph and the event backtracking time window to determine the associated analysis result of the cause chain that leads to the occurrence of the potential fault event includes: Taking the entity corresponding to the potential fault event as the central point entity, perform reverse tracing along the causal relationship in the dependency relationship subgraph to determine the precursor entities that lead to the occurrence of the potential fault event; Generate a causal relationship path based on the central point entity, the precursor entities, and the reverse causal relationship direction from the central point entity to the precursor entities; Perform associated analysis on the business data characteristics of the precursor entities and the changes in the business data of the potential fault event within the event backtracking time window to determine a business data association result; Integrate the causal relationship path and the business data association result to obtain the associated analysis result of the cause chain that leads to the occurrence of the potential fault event.

4. The operation and maintenance method of the operation and maintenance platform based on big data automation according to claim 2, characterized in that, The exploring paths along the service support relationship based on the first target entities to obtain the service support paths of the first target entities includes: Perform a first-round expansion exploration in the operation and maintenance knowledge graph based on the first target entities according to the service support relationship to obtain first adjacent entities that have a direct service support relationship with them; Taking the first target entity as the root node and the first adjacent entity as the first-layer node, for each first-layer node, based on the first-layer node, perform a second-round expansion exploration in the operation and maintenance knowledge graph according to the service support relationship to obtain the second adjacent entity that has a direct service support relationship with it. Loop in turn. If after n rounds of expansion exploration, there is no entity with a direct service support relationship among the (n - 1)-layer nodes, then based on the first target entity and the direction of the entity and service support relationship obtained in each round of expansion exploration, generate the expansion exploration path of the first target entity; For any two first expansion exploration paths and second expansion exploration paths, if there are the same entities in the first expansion exploration path and the second expansion exploration path, then merge the same entities and keep the different entities unchanged to obtain the merged exploration path; Integrate all the merged exploration paths to obtain the service support path of the first target entity.

5. The operation and maintenance method of the big data-based automated operation and maintenance platform according to claim 2, characterized in that, Constructing a dependency relationship subgraph based on each second target entity in the service support path and the dependency relationship between each second target entity, including: For each second target entity in the service support path, traverse in the service support path the third target entity that has a direct dependency relationship with it, and based on the second target entity, the third target entity and the nodes, with the direct dependency relationship as the node edge, construct an initial relationship subgraph; Starting from the second target entity, traverse in the service support path to determine the fourth target entity that can be reached by the indirect dependency relationship within the preset step length, and the intermediate entity between the second target entity and the fourth target entity, and add the corresponding node edges in the initial relationship subgraph based on the fourth target entity and the intermediate entity to obtain the first updated relationship subgraph; Perform cyclic dependency relationship detection on the first updated relationship subgraph to obtain a detection result, and based on the detection result, perform cyclic dependency relationship removal processing on the first updated relationship subgraph to obtain the second updated relationship subgraph; Divide the entities in the second updated relationship subgraph according to the direct dependency relationship and the indirect dependency relationship of the preset step length to obtain the dependency relationship subgraph; the dependency relationship subgraph is divided step by step from the entity without incoming edges as the top layer, so that the entities in each layer depend on the entities in its upper layer.

6. The operation and maintenance method based on the big data automated operation and maintenance platform according to any one of claims 1 to 5, characterized in that, Generating an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case, including: Split the target operation and maintenance case into multiple initial case segments, and map the operation and maintenance case features of each initial case segment to the fault event features in the event information to obtain a mapping matching result; Based on the mapping matching results of each initial case segment, perform case segment screening to determine the target case segment, and construct a case association path based on the mapping matching results of the target case segment; the case association path represents the connection between the different features of the potential fault event and the target operation and maintenance case; Generate the automated operation and maintenance decision based on the case association path.

7. The operation and maintenance method of the operation and maintenance platform based on big data automation according to claim 6, characterized in that, Generating the automated operation and maintenance decision based on the case association path includes: Screening out the target operation steps of the potential fault event from the operation steps of the target case segment based on the case association path, and deriving the initial operation sequence of the target operation steps based on the fault development logic in the event information and the logical relationship between the operation steps in the target case segment; Determining the operation resources required to execute the target operation steps based on the system environment information involved in the event information; Adjusting the initial operation sequence based on the acquisition time of the operation resources in the event information, the dependency relationship between the target operation steps, and the real-time status change information of the potential fault event to obtain the target operation sequence; Integrating the target operation steps, the operation sequence, and the operation resources to generate an automated operation and maintenance decision for the potential fault event.

8. An operation and maintenance system based on a big data automated operation and maintenance platform, characterized in that, Applied to the operation and maintenance method based on the big data automated operation and maintenance platform as described in any one of claims 1 to 7; The operation and maintenance system based on the big data automated operation and maintenance platform includes: A data processing module for uniformly standardizing and integrating multi-source heterogeneous data collected from the operation and maintenance platform based on multiple data acquisition interfaces to obtain integrated multi-source data; A knowledge graph construction module for constructing an operation and maintenance knowledge graph based on each entity in the integrated multi-source data and the entity relationship between each entity; A big data analysis module for performing fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events; An operation and maintenance case matching module for performing associated search in the operation and maintenance knowledge graph based on the potential fault event to obtain an associated analysis result, and performing matching analysis on the associated analysis result based on historical operation and maintenance cases in the big data storage platform to determine a target operation and maintenance case; An operation and maintenance decision generation module for generating an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case.

9. An electronic device, comprising: A memory for storing computer software programs; A processor for reading and executing the computer software program, characterized in that when the processor executes the computer software program, it implements the operation and maintenance method based on the big data automated operation and maintenance platform as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing a computer software program, characterized in that, When the computer software program is executed by the processor, it implements the operation and maintenance method based on the big data automated operation and maintenance platform as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Operation and maintenance management system and method based on knowledge base

    CN114706994A

  • Fault root cause analysis method and device of distributed software system and storage medium

    CN119621396A

  • Operation and maintenance management system and method based on artificial intelligence

    CN119676055A

  • Automatic root cause analysis and prediction for a large dynamic process execution system

    US20220066852A1

Cited By

  • Intelligent operation and maintenance knowledge base construction method

    CN120979957A

  • Fault case intelligent matching method and system applied to operation and maintenance technology service

    CN121614892A

  • Intelligent matching method and system for fault cases applied to operation and maintenance technical services

    CN121614892B