Operation and maintenance method and system based on big data automated operation and maintenance platform
By building a big data automated operation and maintenance platform and utilizing multi-source data integration and knowledge graph analysis, the problem of delayed response in existing operation and maintenance methods has been solved, achieving fast and accurate fault handling and improving system stability.
Patent Information
- Application Number
- CN202510489370.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Existing rule- and experience-based operation and maintenance methods are unable to quickly and accurately respond to unforeseen failures or system changes in modern IT systems, resulting in the inability to timely discover and handle potential problems, affecting operation and maintenance efficiency and system stability.
Build an automated operation and maintenance platform based on big data, generate automated operation and maintenance decisions through multi-source data integration, knowledge graph construction, fault anomaly analysis and association search, use knowledge graphs to deeply explore the complex cause and effect relationships behind fault events, and respond quickly based on historical operation and maintenance cases.
It enables rapid and accurate identification and handling of potential faults, improves operation and maintenance efficiency and system stability, and ensures the normal operation of the business.
Smart Images

Figure CN120336065B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to an operation and maintenance method and system based on a big data automated operation and maintenance platform. Background Art
[0002] Current platform operation and maintenance methods widely utilize traditional rules-based and experience-based approaches. This approach relies on pre-defined rules and accumulated experience to address issues that arise during the operation and maintenance process. When specific system events or abnormal indicators occur, they are handled according to established rules. For example, in server performance monitoring, an alarm is set to be issued when CPU utilization exceeds 80%. Based on experience, the operator then determines whether to increase server resources or optimize related processes.
[0003] However, faced with the increasingly complex and dynamically changing nature of modern IT systems, current methodologies and rules often fail to fully cover all potential O&M scenarios. With the expansion of system scale, the complexity of business logic, and the introduction of new technologies, new failure modes and O&M requirements continue to emerge. Pre-defined rules are unable to adapt to these changes in a timely manner. When encountering unforeseen failures or system changes, it is difficult to respond quickly and accurately, resulting in potential issues not being discovered and addressed in a timely manner. This significantly impacts O&M efficiency and system stability, and can even affect the normal operation of the business. Summary of the Invention
[0004] The present invention provides an operation and maintenance method and system based on a big data automated operation and maintenance platform, which are used to achieve rapid and accurate response, timely discover and handle potential problems, improve operation and maintenance efficiency and system stability, and ensure the normal operation of the business.
[0005] In a first aspect, the present invention provides an operation and maintenance method based on a big data automated operation and maintenance platform, comprising:
[0006] Standardize and integrate the multi-source heterogeneous data collected from the operation and maintenance platform based on various data collection interfaces to obtain integrated multi-source data;
[0007] Building an operation and maintenance knowledge graph based on the entities in the integrated multi-source data and the entity relationships between the entities;
[0008] Performing fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events;
[0009] Performing an association search in the operation and maintenance knowledge graph based on the potential fault event to obtain an association analysis result, and performing a matching analysis on the association analysis result based on historical operation and maintenance cases in the big data storage platform to determine a target operation and maintenance case;
[0010] Based on the event information of the potential fault event and the target operation and maintenance case, an automated operation and maintenance decision for the potential fault event is generated.
[0011] In a second aspect, the present invention further provides an operation and maintenance system based on a big data automated operation and maintenance platform, which is applied to the operation and maintenance method based on a big data automated operation and maintenance platform as described in the first aspect; the operation and maintenance system based on a big data automated operation and maintenance platform comprises:
[0012] The data processing module is used to standardize and integrate the multi-source heterogeneous data collected from the operation and maintenance platform based on various data collection interfaces to obtain integrated multi-source data;
[0013] A knowledge graph construction module is used to construct an operation and maintenance knowledge graph based on the entities in the integrated multi-source data and the entity relationships between the entities;
[0014] A big data analysis module is used to perform fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events;
[0015] An operation and maintenance case matching module is used to perform an association search in the operation and maintenance knowledge graph based on the potential fault event to obtain an association analysis result, and perform a matching analysis on the association analysis result based on the historical operation and maintenance cases in the big data storage platform to determine a target operation and maintenance case;
[0016] The operation and maintenance decision generating module is used to generate an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case.
[0017] In a third aspect, the present invention also provides an electronic device comprising: a memory for storing a computer software program; and a processor for reading and executing the computer software program, thereby implementing an operation and maintenance method based on a big data automated operation and maintenance platform as described above.
[0018] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium, in which a computer software program is stored. When the computer software program is executed by a processor, it implements an operation and maintenance method based on a big data automated operation and maintenance platform as described above.
[0019] In a fifth aspect, the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above operation and maintenance methods based on a big data automated operation and maintenance platform.
[0020] The operation and maintenance method based on a big data automated operation and maintenance platform provided by an embodiment of the present invention constructs a knowledge graph that graphically stores entities and relationships in the operation and maintenance data, allowing operation and maintenance personnel to intuitively understand the overall picture of the operation and maintenance system and facilitate association analysis. Fault event identification is based on real-time data and a pre-set fault feature library, which can quickly and accurately discover potential fault events and detect problems more promptly than rule-based and experience-based methods. Knowledge graph association analysis utilizes the structure of the knowledge graph to deeply explore the complex causal relationships behind fault events, rather than being limited to pre-set rules, thereby improving the accuracy and comprehensiveness of fault cause analysis. Automated operation and maintenance decision generation generates automated operation and maintenance decisions for fault events by combining association analysis results with historical operation and maintenance cases, avoiding the limitations of traditional operation and maintenance that rely on the limited experience of operation and maintenance personnel to formulate solutions. Therefore, the embodiment of the present invention responds quickly and accurately through the coordination of multiple steps, promptly discovers and handles potential problems, improves operation and maintenance efficiency and system stability, and ensures the normal operation of the business. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flowchart of an operation and maintenance method based on a big data automated operation and maintenance platform provided by an embodiment of the present invention;
[0022] Figure 2 Schematic diagram of the structure of an operation and maintenance system based on a big data automated operation and maintenance platform provided by an embodiment of the present invention;
[0023] Figure 3 An embodiment diagram of an electronic device provided by an embodiment of the present invention;
[0024] Figure 4 An embodiment diagram of a computer-readable storage medium provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0027] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.
[0028] See Figure 1 , Figure 1 This is a flow chart of the operation and maintenance method based on the big data automated operation and maintenance platform provided by the present invention. In the embodiment of the present invention, the execution subject of the operation and maintenance method based on the big data automated operation and maintenance platform is the operation and maintenance system. Therefore, the operation and maintenance method based on the big data automated operation and maintenance platform includes:
[0029] Step 10: standardize and integrate the multi-source heterogeneous data collected from the operation and maintenance platform based on multiple data collection interfaces to obtain integrated multi-source data.
[0030] Optionally, the operation and maintenance platform includes a large number of servers, network devices, and business systems. The data collection interfaces for these servers, network devices, and business systems may differ. Therefore, the operation and maintenance platform provides data collection interfaces for these servers, network devices, and business systems. Therefore, the operation and maintenance system collects heterogeneous data from multiple sources through the various data collection interfaces provided by the operation and maintenance platform, such as hardware status data from server monitoring systems and business operation data from application system logs.
[0031] Furthermore, since the data formats, structures and semantics of different data in multi-source heterogeneous data may be different, the operation and maintenance system will standardize and integrate the collected multi-source heterogeneous data to obtain integrated multi-source data.
[0032] Standardization involves unifying data formats (e.g., ISO 8601 for time formats), data encoding (UTF-8), and naming for data with the same meaning. Integration involves linking related data from different sources to form a complete dataset.
[0033] In one embodiment, the operation and maintenance system collects CPU usage data from the server monitoring interface in the format of "2023-10-10 10:00:00, 70%", and collects user request processing time data from the application system log interface in the format of "10 / 10 / 2023 10:00:00AM, 500ms".
[0034] The time format is standardized and converted to "2023-10-10T10:00:00Z". The server identification information is further used as the association key to integrate the two types of data together. The resulting integrated multi-source data is in the format of {"server_id":"srv001","time":"2023-10-10T10:00:00Z","CPU_usage":"70%","request_process_time":"500ms"}.
[0035] Step 20: construct an operation and maintenance knowledge graph based on the entities in the integrated multi-source data and the entity relationships between the entities.
[0036] Furthermore, the operation and maintenance system performs entity extraction and relationship extraction on the integrated multi-source data to obtain the entity relationships between each entity and each entity. Among them, the entities in the embodiment of the present invention include equipment entities, fault entities and business process entities, and the entity relationships include connection relationships, association relationships, reference relationships, service support relationships, dependency relationships and causal relationships.
[0037] For example, the operation and maintenance system identifies device entities such as server A and server B, fault entities such as network failure and server crash, as well as the service support relationship between server A and business system 1, and the causal relationship that the high CPU usage of server A may cause the server crash.
[0038] Furthermore, the operation and maintenance system uses the Neo4j graph database to build a knowledge graph based on each entity and the entity relationships between entities.
[0039] In one embodiment, the server "srv001," the application "app001," and the database "db001" are identified as entities from the integrated multi-source data. It is discovered that "app001" runs on "srv001," and "srv001" connects to "db001" to retrieve data. Therefore, nodes are created in Neo4j to represent these three entities, and entity relationship edges are created for "running on" and "connected to," resulting in an operations and maintenance knowledge graph.
[0040] Step 30: Perform fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events.
[0041] Furthermore, the preset fault feature library in the big data storage platform contains characteristic patterns of various known faults. For example, the fault feature library defines that the CPU usage rate continuously exceeds 80% and the memory usage rate continuously exceeds 90% as server performance bottleneck fault characteristics. Therefore, the operation and maintenance system matches and analyzes the integrated multi-source data with the characteristic pattern, traverses the server performance data in the integrated multi-source data, and determines whether it meets these characteristics, thereby determining potential fault events. In one embodiment, in the integrated multi-source data, the CPU usage of the server "srv001" is 85%, 88%, and 90% respectively, and the memory usage is 92%, 93%, and 95% respectively within 10 consecutive minutes. Based on the performance bottleneck fault characteristics in the preset fault feature library, the operation and maintenance system determines that "srv001" has a potential server performance bottleneck fault event.
[0042] Step 40: perform an association search in the operation and maintenance knowledge graph based on potential fault events to obtain association analysis results, and perform matching analysis on the association analysis results based on historical operation and maintenance cases in the big data storage platform to determine the target operation and maintenance case.
[0043] Furthermore, the operation and maintenance system performs an association search based on the potential fault events in combination with the connection relationships, association relationships, reference relationships, service support relationships, dependency relationships and causal relationships in the operation and maintenance knowledge graph to obtain association analysis results, as specifically described in steps 301 to 304.
[0044] Furthermore, the big data storage platform stores historical operation and maintenance cases, which may contain information such as failure phenomena, processing procedures, and results. Therefore, the operation and maintenance system matches the correlation analysis results with the historical operation and maintenance cases in the big data storage platform, and searches for similar cases by comparing information such as failure phenomena, processing procedures, and results, and determines the target operation and maintenance case. In one embodiment, for the performance bottleneck potential failure event of "srv001", it is found in the operation and maintenance knowledge graph that "app001" is running on it and the recent user request error rate of "app001" has increased. In the historical operation and maintenance cases, a case is found in which the server performance bottleneck causes a high error rate in the application running on it, and the problem is solved by adjusting the server resource allocation. This case is then determined as the target operation and maintenance case.
[0045] Step 50: Generate an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case.
[0046] Furthermore, event information includes fault event characteristics, fault development logic, involved system environment information, acquisition time of operation resources, dependencies between operation steps, and real-time status change information of potential fault events.
[0047] Therefore, the operation and maintenance system generates an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case, wherein the automated operation and maintenance decision at least includes the operation steps, the operation sequence of the operation steps, and the operation resources required to execute the operation steps, as specifically described in steps 501 to 503.
[0048] The knowledge graph constructed by the embodiment of the present invention graphically stores the entities and relationships in the operation and maintenance data, so that the operation and maintenance personnel can intuitively understand the overall picture of the operation and maintenance system, and facilitate association analysis. Fault event identification is based on real-time data and a pre-set fault feature library, which can quickly and accurately discover potential fault events and detect problems more promptly. Knowledge graph association analysis utilizes the structure of the knowledge graph to deeply explore the complex cause-and-effect relationships behind fault events, rather than being limited to pre-set rules, thereby improving the accuracy and comprehensiveness of fault cause analysis. Automated operation and maintenance decision generation generates automated operation and maintenance decisions for fault events by combining association analysis results and historical operation and maintenance cases. Therefore, through the collaboration of multiple steps, responses can be made quickly and accurately, potential problems can be discovered and handled in a timely manner, operation and maintenance efficiency and system stability can be improved, and the normal operation of the business can be guaranteed.
[0049] In one embodiment, steps 301 to 304 are described as follows:
[0050] Step 301: perform an association search in the operation and maintenance knowledge graph based on the potential fault event to determine an initial entity that has a connection relationship, an association relationship, or a reference relationship with the potential fault event.
[0051] Optionally, the operation and maintenance system performs an association search in the operation and maintenance knowledge graph based on the potential fault event to determine the entity connected to the potential fault event through a connection relationship (such as physical connection), an association relationship (such as business association) or a reference relationship (such as configuration reference), and obtains the initial entity, wherein the initial entity in the embodiment of the present invention can be understood as an initial entity set. In one embodiment, the potential fault event is that the CPU usage of the server "srv001" is too high. In the operation and maintenance knowledge graph, "srv001" is connected to the network switch "switch001" through the "connected to" relationship, and is related to the application "app001" through the "running" association relationship, and its configuration file references the path of the storage device "storage001", then "switch001", "app001" and "storage001" are the initial entities.
[0052] Step 302 : Screen out a first target entity that has a service support relationship with the potential fault event from the initial entities, and perform path exploration along the service support relationship based on the first target entity to obtain a service support path of the first target entity.
[0053] Furthermore, in the initial entity set, the operation and maintenance system will screen out first target entities that have a service support relationship with the potential fault event. The first target entities are entities that provide key service support to the subjects involved in the potential fault event.
[0054] Furthermore, the operation and maintenance system will start from the first target entity along the service support relationship, and gradually explore the paths of its upstream and downstream related entities to obtain the service support path of the first target entity, as described in steps 3021 to 3024. Continuing with the above embodiment, in the initial entity, "app001" plays a key service support role in the business operation of "srv001", so "app001" is determined as the first target entity. Exploring along the service support relationship, it was found that "app001" obtains data from the database "db001" and is connected to the user request end through the load balancer "lb001", thus obtaining the service support paths of "app001→db001" and "app001→lb001".
[0055] Step 303: construct a dependency subgraph based on the second target entities in the service support path and the dependency relationships between the second target entities.
[0056] Furthermore, based on the service support path obtained in step 302, the operation and maintenance system extracts the second target entities in the service support path and the dependency relationships between the second target entities, wherein the dependency relationships in the embodiment of the present invention may be data dependency, resource dependency, etc.
[0057] Furthermore, the operation and maintenance system constructs a dependency subgraph according to each second target entity and the dependency relationship between each second target entity, as specifically described in steps 3031 to 3034 .
[0058] Continuing with the above embodiment, in the "app001→db001" and "app001→lb001" service support paths, "db001" and "lb001" are the second target entities. "app001" depends on "db001" to provide data, "app001" depends on "lb001" to distribute user requests, and "lb001" may depend on "db001" to query user authentication data. In the dependency subgraph constructed by the operation and maintenance system, "app001", "db001" and "lb001" are nodes, and the dependency relationships between them are represented by directed edges, such as "app001→db001 (data dependency)", "app001→lb001 (request distribution dependency)", and "lb001→db001 (authentication data query dependency)".
[0059] Step 304, starting from the potential fault event, determines the event backtracking time window based on the event occurrence time of the potential fault event, and performs source analysis based on the dependency subgraph and the event backtracking time window to determine the correlation analysis results of the cause chain that led to the potential fault event.
[0060] Furthermore, starting with the potential failure event, the O&M system determines an event backtracking time window based on the time of the potential failure event, for example, one hour before the failure. Within this time window, the O&M system analyzes the states and events of each entity in the dependency subgraph, tracing back from the potential failure event along the dependency relationships to identify the series of events and factors that may have led to the failure, thereby obtaining a correlation analysis result for the cause chain, as described in steps 3041 through 3044.
[0061] In one embodiment, the potential fault event "srv001" (excessive CPU usage) occurred at 10:00. The backtracking time window was determined to be 9:00-10:00. The dependency subgraph revealed that at 9:30, "db001" performed a large-scale data update, which caused a sudden increase in the data request processing volume of "app001." This in turn increased the request load on "srv001" to process "app001," ultimately leading to excessive CPU usage. This determined the cause chain association analysis result: "db001 large-scale data update → increased request processing volume on app001 → excessive CPU usage on srv001."
[0062] The embodiment of the present invention utilizes the structure of the knowledge graph, which can start from potential fault events, deeply explore the entities and relationships related to them, deeply explore the complex cause-effect relationship behind the fault events, and construct a chain of causes of the fault, so that operation and maintenance personnel can clearly understand the root cause of the fault and are no longer limited to pre-set rules, thereby improving the accuracy and comprehensiveness of fault cause analysis.
[0063] In one embodiment, steps 3021 to 3024 are described as follows:
[0064] Step 3021: Based on the first target entity, a first round of expansion exploration is performed in the operation and maintenance knowledge graph according to the service support relationship to obtain the first adjacent entity that has a direct service support relationship with it.
[0065] Optionally, after the operation and maintenance system determines the first target entity, it performs a first round of expansion exploration starting from the first target entity based on the service support relationship in the operation and maintenance knowledge graph, finds entities that have a direct service support relationship with the first target entity, and obtains the first adjacent entity.
[0066] In one embodiment, the first target entity is "app001." In the operation and maintenance knowledge graph, "app001" is connected to the database "db001" through a service support relationship ("app001" obtains data from "db001") and is also connected to the load balancer "lb001" ("lb001" distributes user requests to "app001"). Therefore, in the first round of expansion exploration, "db001" and "lb001" are the first adjacent entities of "app001."
[0067] Step 3022: take the first target entity as the root node and the first adjacent entity as the first-layer node. For each first-layer node, perform a second round of extended exploration in the operation and maintenance knowledge graph based on the service support relationship of the first-layer node to obtain the second adjacent entity with a direct service support relationship therewith, and repeat in sequence. If after n rounds of extended exploration, there is no entity with a direct service support relationship at the n-1th layer node, then generate an extended exploration path for the first target entity based on the first target entity and the direction of the entity and service support relationship obtained in each round of extended exploration.
[0068] Furthermore, the O&M system uses the first target entity as the root node and the first adjacent entity obtained from the first round of extended exploration as the first-level node. For each first-level node, extended exploration is again performed within the knowledge graph based on the service support relationship to identify entities with which these entities have direct service support relationships, i.e., the second-level adjacent entities. This process continues in a continuous loop, with each round of extended exploration continuing to expand outward based on the previous-level nodes. When the nth round of extended exploration is reached, if no new entities connected to the n-1th-level nodes through service support relationships exist, the extended exploration process ceases.
[0069] Furthermore, the operation and maintenance system constructs an extended exploration path starting from the first target entity based on the first target entity and the direction of the entity and service support relationship discovered during each round of extended exploration.
[0070] In one embodiment, with "app001" as the root node, its first-level node "db001" is connected to the storage server "storage001" via a service support relationship (data storage dependency) during the second round of expanded exploration. In the second round of expanded exploration, "lb001" is connected to the network switch "switch001" via a service support relationship (network connection dependency). Continuing the exploration, "storage001" has no new adjacent entities based on the service support relationship, nor does "switch001." At this point, two expanded exploration paths are generated: one is "app001 → db001 → storage001," and the other is "app001 → lb001 → switch001."
[0071] Step 3023: For any two first extended exploration paths and second extended exploration paths, if there are identical entities in the first extended exploration path and the second extended exploration path, the identical entities are merged and the different entities remain unchanged to obtain a fused exploration path.
[0072] Furthermore, among the multiple extended exploration paths generated, the operation and maintenance system will perform a comparative analysis on any two paths. If two paths contain identical entities, these identical entities are merged into a single node, while different entities remain unchanged. This eliminates redundant information in the paths, making the path structure more concise and clear, while also highlighting the differences and connections between different paths. Continuing with the above example, there are two extended exploration paths: Path 1 is "app001 → db001 → storage001," and Path 2 is "app001 → lb001 → db001." Both paths contain the entity "db001." After fusion, the resulting fused exploration paths are "app001 → (db001) → storage001" and "app001 → lb001 → (db001)," where the "db001" in parentheses represents the merged node.
[0073] Step 3024: Integrate all fused exploration paths to obtain the service support path of the first target entity.
[0074] Furthermore, the operation and maintenance system integrates all the exploration paths after the fusion process to obtain the service support path of the first target entity. Continuing with the above embodiment, there are three exploration paths after fusion: Path one is "app001→(db001)→storage001", Path two is "app001→lb001→(db001)", and Path three is "app001→monitor001" ("monitor001" is used to monitor the running status of "app001"). The service support path of the first target entity "app001" obtained after integration is a comprehensive structure, showing the service support relationship between "app001" and "db001" (associated in different ways), "storage001", "lb001", and "monitor001", such as "app001→db001 (data acquisition)→storage001 (data storage)", "app001→lb001 (request distribution)→db001 (data query)", and "app001→monitor001 (status monitoring)".
[0075] The embodiment of the present invention goes from the initial first round of exploration to determine directly related entities, to multiple rounds of expanded exploration to build a complete path, and then to fusion and integration operations. The service support path finally formed provides operation and maintenance personnel with a clear visual view, showing the position of the first target entity in the entire operation and maintenance ecosystem and its close connection with other entities. Therefore, in the troubleshooting scenario, when a problem occurs with the first target entity, the operation and maintenance personnel can quickly locate the various links that may affect it based on this complete service support path, whether it is the upstream data provider, the downstream service contractor, or the intermediate transmission and monitoring links. They can all be seen at a glance, so the operation and maintenance personnel can respond quickly and accurately, discover and handle potential problems in a timely manner, improve operation and maintenance efficiency and system stability, and ensure the normal operation of the business.
[0076] In one embodiment, steps 3031 to 3034 are described as follows:
[0077] Step 3031: For each second target entity in the service support path, traverse the third target entities that have a direct dependency relationship with it in the service support path, and build an initial relationship subgraph based on the second target entity, the third target entity and the node, with the direct dependency relationship being the node edge.
[0078] Optionally, the operation and maintenance system traverses each second target entity in the service support path to identify third target entities that have a direct dependency relationship with the second target entity. After determining the third target entity and the direct dependency relationship, the operation and maintenance system constructs a preliminary graph structure, namely the initial relationship subgraph, using the second target entity and the third target entity as nodes and the direct dependency relationship between the second target entity and the third target entity as the edge connecting the nodes. The initial relationship subgraph shows the most direct and basic dependency relationships in the service support path.
[0079] In one embodiment, the service support paths are "app001→db001→storage001" and "app001→lb001→switch001". For the second target entity "db001", it has a direct dependency relationship with "app001" for data provision ("app001" depends on "db001" to provide data), and a direct dependency relationship with "storage001" for data storage ("db001" depends on "storage001" to store data). Then, when constructing the initial relationship subgraph, "app001", "db001", and "storage001" become nodes, and "app001→db001 (data provision)" and "db001→storage001 (data storage)" become node edges. Similarly, for "lb001", it has a direct dependency relationship with "app001" for request distribution, and a direct dependency relationship with "switch001" for network connection, and the corresponding nodes and edges will also be added to the initial relationship subgraph.
[0080] Step 3032: Start traversing from the second target entity in the service support path, determine the fourth target entity that is reachable by the indirect dependency relationship within the preset step size, and the intermediate entity between the second target entity and the fourth target entity, and add corresponding node edges to the initial relationship subgraph based on the fourth target entity and the intermediate entity to obtain the first updated relationship subgraph.
[0081] Furthermore, the operation and maintenance system begins with each second target entity in the service support path and traverses it according to a preset step size (for example, a step size of 2 indicates searching for indirect dependencies within a certain range). The operation and maintenance system determines the fourth target entity that can be reached through indirect dependencies within this preset step size, and simultaneously finds the intermediate entities between the second and fourth target entities. Furthermore, the operation and maintenance system adds corresponding node edges to the initial relationship subgraph for these newly discovered fourth target entities and intermediate entities, and updates the initial relationship subgraph to obtain a first updated relationship subgraph, so that the relationship subgraph not only includes direct dependencies but also covers indirect dependencies within a certain range.
[0082] Continuing with the above service support path as an example, let's assume a preset step size of 2. For the second target entity, "db001," through indirect dependencies ("app001" depends on "db001" to obtain data, "db001" depends on "storage001" to store data, so "app001" indirectly depends on "storage001"), the fourth target entity, "storage001," can be reached within step size 2, with "db001" being the intermediate entity. In the initial relationship subgraph, there are already edges "app001→db001" and "db001→storage001." To reflect the indirect dependency from "app001" to "storage001," the system can add a dashed edge or use some other identifier to indicate this indirect dependency (for example, by adding a comment indicating the indirect connection through "db001"). Similarly, for "lb001", if in step 2, it is found that "app001" indirectly depends on the network provider's DNS server (for example, the fourth target entity) through "lb001" and "switch001", and the intermediate entities are "lb001" and "switch001", then the corresponding node (DNS server) and the edge representing the indirect dependency relationship will also be added to the initial relationship subgraph.
[0083] Step 3033: perform a circular dependency check on the first updated relationship subgraph to obtain a detection result, and remove the circular dependency from the first updated relationship subgraph based on the detection result to obtain a second updated relationship subgraph.
[0084] Furthermore, a circular dependency refers to a path in the graph structure that starts from a certain entity, passes through a series of dependencies, and then returns to the entity itself. This situation complicates the analysis and processing of dependencies. Therefore, the operation and maintenance system traverses the first updated relationship subgraph using a specific algorithm (such as a variant of the depth-first search algorithm) to detect whether there are circular dependencies in the first updated relationship subgraph. If circular dependencies are detected, the operation and maintenance system processes the first updated relationship subgraph based on the detection results, removing these circular dependencies and obtaining the second updated relationship subgraph. Common removal methods include severing one of the dependency edges or adjusting the representation of the dependencies.
[0085] In one embodiment, a circular dependency relationship "app001 → db001 → app001" appears in the first updated relationship subgraph due to complex business logic (perhaps because some data updates in "app001" require processing by "db001," and after processing, "db001" needs to provide specific information back to "app001"). To eliminate the loop, the system can decide to cut the "db001 → app001" edge based on actual business importance (for example, feedback from "db001" to "app001" can be implemented through other business methods), resulting in a second updated relationship subgraph without the circular dependency relationship.
[0086] Step 3034: The entities in the second updated relationship subgraph are hierarchically partitioned according to direct dependencies and indirect dependencies with a preset step size, thereby obtaining a dependency subgraph. The dependency subgraph is partitioned downwards, starting with entities without inbound edges at the top level, so that entities at each level are dependent on entities at the upper level.
[0087] Furthermore, the operation and maintenance system divides the entities in the second updated relationship subgraph into relationship hierarchies according to direct dependencies and indirect dependencies with a preset step size. The principle of division is to use entities without incoming edges (i.e., not dependent on other entities) as the top layer, and then gradually divide downwards so that entities at each level depend on their upper-level entities, thereby obtaining a dependency subgraph.
[0088] In one embodiment, in the second updated dependency subgraph, "storage001," "switch001," and the network provider's DNS server have no incoming edges (they do not depend on other entities in the service support path), so they are grouped as the top layer. "db001" depends on "storage001," and "lb001" depends on "switch001," grouping "db001" and "lb001" as the second layer. "app001" depends on "db001" and "lb001," grouping "app001" as the third layer. The resulting dependency subgraph exhibits a clear hierarchical structure, with "storage001," "switch001," and the DNS server as the top layer; "db001," "lb001" as the second layer; and "app001" as the third layer.
[0089] The embodiment of the present invention can construct a dependency subgraph with clear hierarchy, no circular dependencies and comprehensive reflection of entity dependencies from a complex service support path. This subgraph systematically integrates and sorts the dependency information originally scattered in the service support path. Therefore, in troubleshooting and system optimization scenarios, operation and maintenance personnel can use this dependency subgraph to quickly locate the fault propagation path and potential impact range, so that they can respond quickly and accurately, discover and handle potential problems in a timely manner, improve operation and maintenance efficiency and system stability, and ensure the normal operation of the business.
[0090] In one embodiment, steps 3041 to 3044 are described as follows:
[0091] Step 3041 , taking the entity corresponding to the potential fault event as the central entity, reverse tracing is performed along the causal relationship in the dependency subgraph to determine the predecessor entity that causes the potential fault event to occur.
[0092] Optionally, the operation and maintenance system sets the entity corresponding to the potential fault event as the central point entity, and in the constructed dependency subgraph, performs reverse tracing starting from the central point entity along the causal relationship represented in the dependency subgraph, wherein the causal relationship is reflected in the dependency subgraph as the dependency direction between entities, such as data flow direction, service call direction, etc. Therefore, by searching along these causal relationships in reverse, the predecessor entity that logically precedes the potential fault event and affects it is determined, and the predecessor entity that causes the potential fault event is obtained.
[0093] In one embodiment, for example, a potential failure event is a slow response from "app001." "app001" is the central entity. In the dependency subgraph, "app001" relies on "db001" for data acquisition and "lb001" for request distribution. Tracing back the causal relationships, it is found that "db001" and "lb001" are predecessor entities of "app001." This is because slow data queries from "db001" or abnormal request distribution from "lb001" can both lead to slow responses from "app001."
[0094] Step 3042: Generate a causal relationship path based on the center point entity, the predecessor entity, and the reverse causal relationship direction from the center point entity to the predecessor entity.
[0095] Furthermore, based on the identified central point entity and predecessor entity, the operation and maintenance system generates a causal path from the central point entity to the predecessor entity. The causal path clarifies the relationship between entities and also identifies the direction of the causal relationship, that is, from the entity where the potential fault event occurs to the predecessor entity that may cause the fault. This constructs a preliminary fault propagation logic framework, which helps to intuitively understand how the fault affects the central point entity (potential fault event entity) from the upstream entity.
[0096] Continuing with the above example, with "app001" as the central entity and "db001" and "lb001" as predecessor entities, two causal paths are generated. One is "app001←db001 (data acquisition dependency)", indicating that "app001" depends on "db001" for data acquisition. If a problem with "db001" occurs, "app001" may be affected. The other is "app001←lb001 (request dispatch dependency)", indicating that "app001" depends on "lb001" for request dispatch. An exception in "lb001" may cause "app001" to respond slowly.
[0097] Step 3043 , performing correlation analysis based on the business data features of the predecessor entity and the changes in the business data of the potential fault event within the event backtracking time window, and determining a business data correlation result.
[0098] Furthermore, the operation and maintenance system analyzes the business data characteristics of each predecessor entity, including data volume, data update frequency, request processing time, etc. At the same time, the operation and maintenance system compares the business data changes of the potential failure event (central point entity) within the event backtracking time window. By correlating the business data characteristics of the predecessor entity with the business data changes of the potential failure event, the system identifies the inherent connection between the two, determines whether the business data status of the predecessor entity is causally related to the occurrence of the potential failure event, and determines the business data correlation result.
[0099] In one embodiment, for example, the event backtracking window is 30 minutes before the failure occurred. For the predecessor entity "db001," it was found that its data update frequency suddenly increased significantly during these 30 minutes, from 10 updates per hour to 1 update every 10 minutes. Meanwhile, the average response time of "app001" increased from 200ms to 800ms during the same time window. Comparative analysis revealed a correlation between the abnormal increase in "db001's" data update frequency and the delayed response of "app001." In other words, the business data correlation results indicate that the change in "db001's" data update frequency may be a factor contributing to the delayed response of "app001." For "lb001," its request dispatch success rate dropped from 99% to 90% during the event backtracking window, while the number of valid requests received by "app001" also decreased. This suggests that the decreased request dispatch success rate of "lb001" is also related to the delayed response of "app001."
[0100] In step 3044, the causal relationship path and the business data association result are integrated to obtain the association analysis result of the cause chain leading to the potential failure event.
[0101] Furthermore, the operation and maintenance system integrates the previously generated causal path and the determined business data association results to form a complete association analysis result of the cause chain leading to the potential failure event. The association analysis result includes both the logical dependency relationship between entities (causal path) and the association information at the business data level.
[0102] In one embodiment, the causal relationship path "app001←db001 (data acquisition dependency)" is integrated with the business data association result "an abnormal increase in db001 data update frequency is associated with app001's delayed response." The causal relationship path "app001←lb001 (request dispatch dependency)" is also integrated with the business data association result "a decrease in lb001's request dispatch success rate is associated with app001's delayed response." The resulting correlation analysis results for the causal chain leading to the potential fault event of "app001's delayed response" are as follows: "app001" depends on "db001" for data acquisition; the recent abnormal increase in "db001's" data update frequency affects "app001's" response; and "app001" depends on "lb001" for request dispatch. The decrease in "lb001's" request dispatch success rate also negatively impacts "app001's" response.
[0103] The embodiments of the present invention can comprehensively and deeply explore the cause chain leading to potential failure events based on the two key elements of dependency subgraph and event backtracking time window, providing operation and maintenance personnel with a structured, data-driven fault diagnosis perspective. Therefore, in actual operation and maintenance scenarios, operation and maintenance personnel can respond quickly and accurately based on detailed correlation analysis results, discover and handle potential problems in a timely manner, improve operation and maintenance efficiency and system stability, and ensure the normal operation of the business.
[0104] In one embodiment, steps 501 to 503 are described as follows:
[0105] Step 501: split the target operation and maintenance case into multiple initial case segments, map the operation and maintenance case features of each initial case segment with the fault event features in the event information, and obtain a mapping matching result.
[0106] Optionally, the operation and maintenance system splits the target operation and maintenance case into multiple initial case segments according to a certain logic or operation process, where these initial case segments can be a series of continuous operation and maintenance operations, or processing links for specific fault phenomena.
[0107] Furthermore, for each initial case segment, the operation and maintenance system extracts its operation and maintenance case features, where these include but are not limited to the processing operation type, the system components involved, the processing time, etc. At the same time, fault event features are extracted from the event information of potential fault events, such as the fault type, the location of the fault, and the scope of the fault impact.
[0108] Furthermore, the operation and maintenance system maps and compares the operation and maintenance case features of each initial case fragment with the fault event features one by one, determines the degree of matching between them, and obtains the mapping matching results, where the mapping matching results are used to measure the similarity between the initial case fragment and the potential fault event at the feature level.
[0109] In one embodiment, for example, the target operation and maintenance case is a historical record of resolving a performance bottleneck on server "srv001," which includes operations such as shutting down non-critical processes, increasing server memory, and optimizing database queries. This case is broken down into three initial case segments: Segment 1: "Shutting down non-critical processes," Segment 2: "Increasing server memory," and Segment 3: "Optimizing database queries." The potential failure event is the recurrence of a performance bottleneck on "srv001." Failure event characteristics include excessive CPU usage (above 85%), near-saturation memory usage (above 90%), and a slow response from the application "app001" involved. For Segment 1, the operation and maintenance case characteristics are "Operation Type - Process Management, Involved Component - Server Process." When mapped against the failure event characteristics, "Server Process" is associated with "srv001," and high CPU usage is likely related to excessive non-critical processes running, indicating a certain degree of match. Segment 2's operation and maintenance case characteristics are "Operation Type - Hardware Resource Adjustment, Involved Component - Server Memory," which closely matches the near-saturation memory usage characteristic of the failure event. The operation and maintenance case feature of fragment 3 is "Operation Type - Database Optimization, Involved Component - Database." This feature is related to the slow response of "app001" in the failure event (possibly related to slow database queries), and a corresponding mapping matching result is obtained.
[0110] In step 502, case segments are screened based on the mapping and matching results of each initial case segment to determine the target case segment. A case association path is then constructed based on the mapping and matching results of the target case segment. The case association path represents the connection between different characteristics of the potential fault event and the target operation and maintenance case.
[0111] Furthermore, based on the mapping matching results of each initial case segment obtained in step 501, the operation and maintenance system selects those initial case segments that have a higher matching degree with the potential fault event characteristics as target case segments according to a preset matching threshold or comprehensive evaluation rule.
[0112] Furthermore, the operation and maintenance system constructs a case association path based on the mapping and matching results of these target case segments. This case association path graphically or structuredly illustrates how the different characteristics of potential fault events relate to the various processing steps in the target operation and maintenance case, presenting a logical relationship from the fault phenomenon to the corresponding processing solution. In one embodiment, the matching degree threshold is 60%. Based on the mapping and matching results in step 501, the matching degree of segment 1 is assessed as 50%, that of segment 2 as 80%, and that of segment 3 as 70%. Therefore, segments 2 and 3 are determined as target case segments. When constructing the case association path, for segment 2, since the fault event characteristic of near-saturation memory usage is closely related to the "increase server memory" operation, an edge is established from "near-saturation memory usage" to "increase server memory." For segment 3, since the slow response of "app001" is related to "optimize database queries," an edge is established from "slow response of app001" to "optimize database queries," resulting in the case association path.
[0113] Step 503: Generate an automated operation and maintenance decision based on the case association path.
[0114] Furthermore, the operation and maintenance system generates an automated operation and maintenance decision based on the associated path, as specifically described in steps 5031 to 5034 .
[0115] The embodiments of the present invention can closely combine historical target operation and maintenance cases with current potential failure events to generate targeted and operational automated operation and maintenance decisions. Therefore, in actual operation and maintenance, it can respond quickly and accurately, discover and handle potential problems in a timely manner, improve operation and maintenance efficiency and system stability, and ensure the normal operation of the business.
[0116] In one embodiment, steps 5031 to 5034 are described as follows:
[0117] Step 5031, based on the case association path, the target operation steps of the potential fault event are screened out from the operation steps of the target case segment, and based on the fault development logic in the event information and the logical relationship between the operation steps in the target case segment, the initial operation sequence of the target operation steps is derived.
[0118] Optionally, the operation and maintenance system, based on the established case association path, filters out the target operation steps that are directly related to the potential fault event from the numerous operation steps contained in the target case fragment, where the target operation steps are the core actions to resolve the current potential fault. At the same time, the operation and maintenance system analyzes the fault development logic in the event information, such as how the fault gradually affects other components from the abnormality of one component, and the inherent logical relationship between the operation steps in the target case fragment, such as some operations must be performed after other operations are completed. Therefore, the operation and maintenance system derives the preliminary execution order of the target operation steps, that is, the initial operation order, based on the fault development logic and the logical relationship between the operation steps in the target case fragment.
[0119] Continuing with the performance bottleneck issue on server "srv001," the target case segments are "Increase Server Memory" and "Optimize Database Queries." In the "Increase Server Memory" segment, the steps include checking the memory slots, selecting the appropriate memory module, and installing memory. The "Optimize Database Queries" segment includes steps like analyzing query statements, creating indexes, and adjusting query parameters. Based on the case association path, the target steps directly related to the "srv001" performance bottleneck are "Select an Appropriate Memory Module," "Install Memory," "Analyze Query Statements," and "Create Indexes." From the perspective of the fault's development logic, insufficient server memory and slow database queries jointly contribute to the performance bottleneck. Furthermore, based on the target case segment's operational logic, selecting the appropriate module is necessary before increasing memory, and analyzing query statements is necessary before optimizing database queries. Therefore, the deduced initial operation sequence is: "Select an Appropriate Memory Module," followed by "Install Memory," then "Analyze Query Statements," and finally, "Create Indexes."
[0120] Step 5032: Determine the operation resources required to execute the target operation step based on the system environment information involved in the event information.
[0121] Furthermore, the operation and maintenance system comprehensively evaluates the various types of operation resources required to execute each target operation step based on the system environment information involved in the event information. Among them, operation resources include hardware resources, such as additional memory modules and free server slots; software resources, such as database management tools and memory detection software; human resources, including operation and maintenance personnel with corresponding skills, database administrators, etc.
[0122] Continuing with the above example, for the target operation step of "selecting a suitable memory module," the required operational resources include memory specifications (software resources, used to select appropriate memory based on server configuration) and operations and maintenance personnel familiar with server memory configuration (human resources). The "installing memory" step requires memory modules (hardware resources), installation tools such as screwdrivers (hardware resources), and the aforementioned operations and maintenance personnel. "Analyzing query statements" requires database management software (software resources) and a database administrator with database knowledge (human resources). "Creating an index" also relies on database management software and a database administrator.
[0123] Step 5033: Based on the acquisition time of the operation resources in the event information, the dependency relationship between the target operation steps, and the real-time status change information of the potential fault event, the initial operation sequence is adjusted to obtain the target operation sequence.
[0124] Furthermore, the O&M system comprehensively considers the acquisition time of operational resources in the event information, namely the length of time required to acquire the required hardware, software, or deploy human resources. Simultaneously, it analyzes the dependencies between target operational steps. For example, some operations must wait for specific resources to be ready before they can begin, or the completion of one operation is a prerequisite for the initiation of another. Furthermore, it monitors status changes of potential fault events in real time, such as whether the severity of the fault is increasing or the scope of impact is expanding. By comprehensively weighing these three factors, the O&M system adjusts the initial operation sequence to obtain a target operation sequence that better reflects the actual situation and the urgency of fault handling.
[0125] In one embodiment, for example, obtaining a suitable memory module requires two hours, while a database administrator can be in place within one hour. At the same time, it was found that the CPU usage of "srv001" continued to rise over time, exacerbating the performance bottleneck. Considering that database query optimization can initially perform some operations within the available memory, and that the database administrator's arrival time precedes the memory module acquisition time, the initial operation sequence is adjusted, and the target operation sequence becomes: first "analyze the query statement," then "create the index" (these two steps can be started immediately after the database administrator is in place), and then, after the memory module is acquired, proceed to "select the appropriate memory module" and "install the memory."
[0126] Step 5034: Integrate the target operation steps, operation sequence, and operation resources to generate automated operation and maintenance decisions for potential failure events.
[0127] Furthermore, the operation and maintenance system integrates the determined target operation steps, the optimized target operation sequence, and the operation resources required to perform these operations to generate automated operation and maintenance decisions for potential failure events. The automated operation and maintenance decision is a detailed task list that clearly lists the specific content of each operation step, the execution sequence, and detailed information on the required resources.
[0128] In one embodiment, the automated operation and maintenance decisions generated after integration are as follows:
[0129] Step 1: The database administrator uses database management software to analyze the database query statements related to "app001" within 1 hour.
[0130] Step 2: After the analysis is completed, the database administrator immediately uses the database management software to create the index.
[0131] Step 3: After obtaining the appropriate memory modules (estimated to take 2 hours), the operation and maintenance personnel refer to the memory specifications and select the appropriate memory modules.
[0132] Step 4: The operation and maintenance personnel use tools such as screwdrivers to install the memory module into the "srv001" server.
[0133] The embodiments of the present invention can generate highly customized, practical, and highly operational automated operation and maintenance decisions. Therefore, when faced with potential failure events, historical cases are no longer mechanically applied. Instead, the development logic of the current failure, the system environment, and resource acquisition conditions are fully combined to dynamically adjust the operation steps and sequence, so that operation and maintenance personnel can respond quickly and accurately, discover and handle potential problems in a timely manner, improve operation and maintenance efficiency and system stability, and ensure the normal operation of the business.
[0134] Furthermore, the operation and maintenance system based on the big data automated operation and maintenance platform provided by the present invention is described below. The operation and maintenance system based on the big data automated operation and maintenance platform described below and the operation and maintenance method based on the big data automated operation and maintenance platform described above can be referenced to each other.
[0135] Optional, see Figure 2 , Figure 2 It is a structural diagram of the operation and maintenance system based on the big data automated operation and maintenance platform provided by the present invention. The operation and maintenance system based on the big data automated operation and maintenance platform includes.
[0136] The data processing module 210 is used to standardize and integrate the multi-source heterogeneous data collected from the operation and maintenance platform based on various data collection interfaces to obtain integrated multi-source data;
[0137] The knowledge graph construction module 220 is used to construct an operation and maintenance knowledge graph based on the entities and the entity relationships between the entities in the integrated multi-source data;
[0138] The big data analysis module 230 is used to perform fault anomaly analysis on the integrated multi-source data based on the preset fault feature library in the big data storage platform to determine potential fault events;
[0139] The operation and maintenance case matching module 240 is used to perform an association search in the operation and maintenance knowledge graph based on potential fault events, obtain association analysis results, and perform matching analysis on the association analysis results based on historical operation and maintenance cases in the big data storage platform to determine the target operation and maintenance case;
[0140] The operation and maintenance decision generating module 250 is used to generate an automated operation and maintenance decision for a potential fault event based on the event information of the potential fault event and a target operation and maintenance case.
[0141] The embodiments of the present invention achieve rapid and accurate response, timely discovery and processing of potential problems, improve operation and maintenance efficiency and system stability, and ensure the normal operation of services.
[0142] See also Figure 3 , Figure 3 This is a diagram of an embodiment of an electronic device provided by an embodiment of the present invention. Figure 3 As shown, an embodiment of the present invention provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, the following steps are implemented:
[0143] Standardize and integrate the multi-source heterogeneous data collected from the operation and maintenance platform based on various data collection interfaces to obtain integrated multi-source data;
[0144] Build an operation and maintenance knowledge graph based on the entities and the entity relationships between entities in the integrated multi-source data;
[0145] Perform fault anomaly analysis on the integrated multi-source data based on the preset fault feature library in the big data storage platform to identify potential fault events;
[0146] Based on potential failure events, an association search is performed in the operation and maintenance knowledge graph to obtain association analysis results. The association analysis results are then matched with historical operation and maintenance cases in the big data storage platform to determine the target operation and maintenance cases.
[0147] Generate automated operation and maintenance decisions for potential failure events based on event information and target operation and maintenance cases.
[0148] See also Figure 4 , Figure 4 Detailed description of an embodiment of a computer-readable storage medium provided by an embodiment of the present invention. Figure 4 As shown, this embodiment provides a computer-readable storage medium 400 on which a computer program 311 is stored. When the computer program 311 is executed by a processor, the following steps are implemented:
[0149] Standardize and integrate the multi-source heterogeneous data collected from the operation and maintenance platform based on various data collection interfaces to obtain integrated multi-source data;
[0150] Build an operation and maintenance knowledge graph based on the entities and the entity relationships between entities in the integrated multi-source data;
[0151] Perform fault anomaly analysis on the integrated multi-source data based on the preset fault feature library in the big data storage platform to identify potential fault events;
[0152] Based on potential failure events, an association search is performed in the operation and maintenance knowledge graph to obtain association analysis results. The association analysis results are then matched with historical operation and maintenance cases in the big data storage platform to determine the target operation and maintenance cases.
[0153] Generate automated operation and maintenance decisions for potential failure events based on event information and target operation and maintenance cases.
[0154] On the other hand, the present invention further provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the operation and maintenance method based on the big data automated operation and maintenance platform provided by the above methods, which includes:
[0155] Standardize and integrate the multi-source heterogeneous data collected from the operation and maintenance platform based on various data collection interfaces to obtain integrated multi-source data;
[0156] Build an operation and maintenance knowledge graph based on the entities and the entity relationships between entities in the integrated multi-source data;
[0157] Perform fault anomaly analysis on the integrated multi-source data based on the preset fault feature library in the big data storage platform to identify potential fault events;
[0158] Based on potential failure events, an association search is performed in the operation and maintenance knowledge graph to obtain association analysis results. The association analysis results are then matched with historical operation and maintenance cases in the big data storage platform to determine the target operation and maintenance cases.
[0159] Generate automated operation and maintenance decisions for potential failure events based on event information and target operation and maintenance cases.
[0160] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0161] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An operation and maintenance method based on a big data automated operation and maintenance platform, characterized in that: include: Standardize and integrate the multi-source heterogeneous data collected from the operation and maintenance platform based on various data collection interfaces to obtain integrated multi-source data; Building an operation and maintenance knowledge graph based on the entities in the integrated multi-source data and the entity relationships between the entities; Performing fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events; Performing an association search in the operation and maintenance knowledge graph based on the potential fault event to obtain an association analysis result, and performing a matching analysis on the association analysis result based on historical operation and maintenance cases in the big data storage platform to determine a target operation and maintenance case; generating an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case; The associated search based on the potential fault event in the operation and maintenance knowledge graph to obtain the associated analysis results includes: Performing an association search on the operation and maintenance knowledge graph based on the potential fault event to determine an initial entity that has a connection relationship, an association relationship, or a reference relationship with the potential fault event; Screening out a first target entity that has a service support relationship with the potential fault event from the initial entities, and performing path exploration along the service support relationship based on the first target entity to obtain a service support path for the first target entity; Constructing a dependency subgraph based on the second target entities in the service support path and the dependency relationships between the second target entities; Taking the potential fault event as a starting point, determining an event backtracking time window based on the event occurrence time of the potential fault event, and performing a source analysis based on the dependency subgraph and the event backtracking time window to determine a correlation analysis result of the cause chain leading to the occurrence of the potential fault event; The generating of an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case includes: Splitting the target operation and maintenance case into multiple initial case segments, and mapping the operation and maintenance case features of each initial case segment with the fault event features in the event information to obtain a mapping matching result; Based on the mapping and matching results of each initial case fragment, case fragments are screened to determine the target case fragment, and a case association path is constructed based on the mapping and matching results of the target case fragment; the case association path represents the connection between different characteristics of the potential fault event and the target operation and maintenance case; The automated operation and maintenance decision is generated based on the case association path.
2. The operation and maintenance method based on the big data automated operation and maintenance platform according to claim 1, characterized in that: The tracing back the source analysis based on the dependency subgraph and the event backtracking time window to determine the correlation analysis result of the cause chain leading to the occurrence of the potential fault event includes: Taking the entity corresponding to the potential fault event as the central entity, tracing back along the causal relationship in the dependency subgraph to determine the predecessor entity that caused the potential fault event to occur; Generate a causal relationship path based on the central point entity, the predecessor entity, and the reverse causal relationship direction from the central point entity to the predecessor entity; Performing correlation analysis based on the business data characteristics of the predecessor entity and changes in the business data of the potential fault event within the event backtracking time window to determine a business data correlation result; The causal relationship path and the business data association result are integrated to obtain an association analysis result of the cause chain leading to the occurrence of the potential fault event.
3. The operation and maintenance method based on the big data automated operation and maintenance platform according to claim 1, characterized in that: The performing path exploration along the service support relationship based on the first target entity to obtain the service support path of the first target entity includes: Based on the first target entity, a first round of expansion exploration is performed in the operation and maintenance knowledge graph according to the service support relationship to obtain a first adjacent entity that has a direct service support relationship with the first target entity; Taking the first target entity as the root node and the first adjacent entity as the first-layer node, for each first-layer node, a second round of extended exploration is performed in the operation and maintenance knowledge graph based on the service support relationship of the first-layer node to obtain a second adjacent entity with a direct service support relationship therewith, and the process is repeated in sequence. If, after n rounds of extended exploration, there is no entity with a direct service support relationship at the n-1th layer node, then based on the first target entity and the direction of the entity and service support relationship obtained in each round of extended exploration, an extended exploration path for the first target entity is generated; For any two first extended exploration paths and second extended exploration paths, if there are identical entities in the first extended exploration path and the second extended exploration path, the identical entities are merged, and the different entities remain unchanged to obtain a fused exploration path; All fused exploration paths are integrated to obtain the service support path of the first target entity.
4. The operation and maintenance method based on the big data automated operation and maintenance platform according to claim 1, characterized in that: The step of constructing a dependency subgraph based on the second target entities in the service support path and the dependency relationships between the second target entities includes: For each second target entity in the service support path, traverse the third target entities that have a direct dependency relationship with the second target entity in the service support path, and construct an initial relationship subgraph based on the second target entity, the third target entity and the node, with the direct dependency relationship being a node edge; Starting from the second target entity, traversing the service support path, determining a fourth target entity that is reachable through an indirect dependency relationship within a preset step size, and an intermediate entity between the second target entity and the fourth target entity, and adding corresponding node edges to the initial relationship subgraph based on the fourth target entity and the intermediate entity to obtain a first updated relationship subgraph; Performing a circular dependency detection on the first updated relationship subgraph to obtain a detection result, and removing the circular dependency from the first updated relationship subgraph based on the detection result to obtain a second updated relationship subgraph; The entities in the second updated relationship subgraph are divided into relationship hierarchies according to direct dependencies and indirect dependencies with a preset step length to obtain the dependency subgraph; the dependency subgraph is divided gradually downward with entities without input edges as the top layer, so that entities at each layer are dependent on entities at the upper layer.
5. The operation and maintenance method based on the big data automated operation and maintenance platform according to claim 1, characterized in that: Generating the automated operation and maintenance decision based on the case association path includes: Filtering target operation steps of the potential fault event from the operation steps of the target case segment based on the case association path, and deducing the initial operation sequence of the target operation steps based on the fault development logic in the event information and the logical relationship between the operation steps in the target case segment; Determining the operation resources required to execute the target operation step based on the system environment information involved in the event information; Adjusting the initial operation sequence based on the acquisition time of the operation resources in the event information, the dependency relationship between the target operation steps, and the real-time status change information of the potential fault event to obtain a target operation sequence; The target operation steps, the operation sequence, and the operation resources are integrated to generate an automated operation and maintenance decision for the potential failure event.
6. An operation and maintenance system based on a big data automated operation and maintenance platform, characterized in that: Applicable to the operation and maintenance method based on the big data automated operation and maintenance platform as described in any one of claims 1 to 5; The operation and maintenance system based on the big data automated operation and maintenance platform includes: The data processing module is used to standardize and integrate the multi-source heterogeneous data collected from the operation and maintenance platform based on various data collection interfaces to obtain integrated multi-source data; A knowledge graph construction module is used to construct an operation and maintenance knowledge graph based on the entities in the integrated multi-source data and the entity relationships between the entities; A big data analysis module is used to perform fault anomaly analysis on the integrated multi-source data based on a preset fault feature library in the big data storage platform to determine potential fault events; An operation and maintenance case matching module is used to perform an association search in the operation and maintenance knowledge graph based on the potential fault event to obtain an association analysis result, and perform a matching analysis on the association analysis result based on the historical operation and maintenance cases in the big data storage platform to determine a target operation and maintenance case; An operation and maintenance decision generating module, configured to generate an automated operation and maintenance decision for the potential failure event based on the event information of the potential failure event and the target operation and maintenance case; The associated search based on the potential fault event in the operation and maintenance knowledge graph to obtain the associated analysis results includes: Performing an association search on the operation and maintenance knowledge graph based on the potential fault event to determine an initial entity that has a connection relationship, an association relationship, or a reference relationship with the potential fault event; Screening out a first target entity that has a service support relationship with the potential fault event from the initial entities, and performing path exploration along the service support relationship based on the first target entity to obtain a service support path for the first target entity; Constructing a dependency subgraph based on the second target entities in the service support path and the dependency relationships between the second target entities; Taking the potential fault event as a starting point, determining an event backtracking time window based on the event occurrence time of the potential fault event, and performing a source analysis based on the dependency subgraph and the event backtracking time window to determine a correlation analysis result of the cause chain leading to the occurrence of the potential fault event; The generating of an automated operation and maintenance decision for the potential fault event based on the event information of the potential fault event and the target operation and maintenance case includes: Splitting the target operation and maintenance case into multiple initial case segments, and mapping the operation and maintenance case features of each initial case segment with the fault event features in the event information to obtain a mapping matching result; Based on the mapping and matching results of each initial case fragment, case fragments are screened to determine the target case fragment, and a case association path is constructed based on the mapping and matching results of the target case fragment; the case association path represents the connection between different characteristics of the potential fault event and the target operation and maintenance case; The automated operation and maintenance decision is generated based on the case association path.
7. An electronic device comprising: Memory for storing computer software programs; A processor for reading and executing the computer software program, characterized in that when the processor executes the computer software program, it implements the operation and maintenance method based on the big data automated operation and maintenance platform as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium storing a computer software program, wherein: When the computer software program is executed by a processor, the operation and maintenance method based on the big data automated operation and maintenance platform as claimed in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Operation and maintenance management system and method based on knowledge base
CN114706994A
Fault root cause analysis method and device of distributed software system and storage medium
CN119621396A
Cited By
Depth evaluation method and device based on multi-dimensional knowledge association graph
CN121707306A