System correlation analysis and fault range positioning method, device, medium and equipment
Patent Information
- Application Number
- CN202610867047.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]有鉴于此,本申请实施例提供了一种系统关联分析及故障范围定位方法、装置、介质及设备,主要目的在于解决现有的系统关联关系维护方式在系统发生故障时,无法快速准确定位根源系统的支撑体系及其影响范围的技术问题
[0009]By employing the above technical solutions, this application provides a system correlation analysis and fault range localization method, apparatus, medium, and equipment. It constructs an inter-system correlation graph using system entities as nodes and inter-system relationships as edges. It dynamically optimizes the correlation relationships and strength between systems using user access information and system tag overlap. Simultaneously, real-time monitoring alarm information is overlaid on the graph, and a fault impact range is analyzed starting from the root system based on a graph traversal algorithm, thereby generating a fault analysis report. This transforms the originally static inter-system correlation relationships into dynamic, quantifiable correlation relationships and strengths, thus solving the problems of information lag and ambiguous impact boundary localization caused by traditional reliance on manual maintenance. Furthermore, by adding real-time alarm information to the inter-system correlation graph and marking the correlation strength between systems, it helps maintenance personnel quickly identify the core of the fault and the affected systems, replacing the manual troubleshooting process. This significantly shortens the system's emergency response time and improves the accuracy of secondary disaster assessment.
Smart Images

Figure CN122802338A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of system operation and maintenance monitoring technology, and in particular to a system correlation analysis and fault range location method, device, medium and equipment. Background Technology
[0002] In the field of system operation and maintenance, system architectures across various industries are becoming increasingly complex, with dense inter-system dependencies and data interactions. For such systems, accurately understanding the interrelationships between them is crucial for locating the root cause of system failures and assessing their impact.
[0003] Traditional methods of maintaining system relationships typically rely on manually compiled static documents or simple configuration management databases, which are difficult to cope with dynamic and complex environments. Especially in the event of sudden failures, maintenance personnel often cannot quickly obtain the complete upstream and downstream dependencies and support systems of the failed system, resulting in low emergency response efficiency and unclear positioning of the boundaries of secondary disaster impacts. Summary of the Invention
[0004] In view of this, the embodiments of this application provide a system correlation analysis and fault range location method, apparatus, medium and equipment, the main purpose of which is to solve the technical problem that the existing system correlation maintenance methods cannot quickly and accurately locate the root cause system's support system and its scope of influence when a system failure occurs.
[0005] According to one aspect of this application, a method for system correlation analysis and fault range localization is provided, the method comprising: Using user-created system entities as nodes and inter-system relationships as edges, an inter-system relationship graph is constructed, wherein the system entities include basic system information and system tags; Real-time acquisition of user access information, and updating of the inter-system relationships based on the user access information and the overlap of system tags between systems, and calculation of the degree of inter-system association; Based on the updated inter-system relationships, the edges of the inter-system relationship graph are adjusted, and the corresponding edges in the inter-system relationship graph are marked according to the degree of association. The system acquires alarm information from each system in real time through a monitoring system, and marks the corresponding nodes in the inter-system relationship graph according to the source and level of the alarm information. In response to a fault analysis request, the startup graph traversal algorithm starts from the root system selected by the user and traverses along the edges with relationships. Based on the degree of association of each node and alarm information, a fault analysis report is generated.
[0006] According to another aspect of this application, a system correlation analysis and fault range location device is provided, the device comprising: The relationship graph creation module is used to construct an inter-system relationship graph using user-created system entities as nodes and inter-system relationships as edges. The system entities include basic system information and system tags. The association strength calculation module is used to acquire user access information in real time, update the association relationship between systems based on the user access information and the overlap of system tags between systems, and calculate the degree of association between systems. The association adjustment module is used to adjust the edges of the inter-system association graph according to the updated inter-system association relationship, and to mark the corresponding edges in the inter-system association graph according to the degree of association. The alarm information marking module is used to obtain alarm information from each system in real time through the monitoring system, and mark the corresponding nodes in the inter-system relationship graph according to the source and level of the alarm information; The analysis report generation module is used to respond to fault analysis requests. It starts the graph traversal algorithm from the root cause system selected by the user and traverses the edges with relationships. Based on the degree of association of each node and alarm information, it generates a fault analysis report.
[0007] According to another aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described system correlation analysis and fault range location method.
[0008] According to another aspect of this application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described system correlation analysis and fault range localization method.
[0009] By employing the above technical solutions, this application provides a system correlation analysis and fault range localization method, apparatus, medium, and equipment. It constructs an inter-system correlation graph using system entities as nodes and inter-system relationships as edges. It dynamically optimizes the correlation relationships and strength between systems using user access information and system tag overlap. Simultaneously, real-time monitoring alarm information is overlaid on the graph, and a fault impact range is analyzed starting from the root system based on a graph traversal algorithm, thereby generating a fault analysis report. This transforms the originally static inter-system correlation relationships into dynamic, quantifiable correlation relationships and strengths, thus solving the problems of information lag and ambiguous impact boundary localization caused by traditional reliance on manual maintenance. Furthermore, by adding real-time alarm information to the inter-system correlation graph and marking the correlation strength between systems, it helps maintenance personnel quickly identify the core of the fault and the affected systems, replacing the manual troubleshooting process. This significantly shortens the system's emergency response time and improves the accuracy of secondary disaster assessment.
[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a system correlation analysis and fault range localization method provided in an embodiment of this application is shown. Figure 2 A flowchart illustrating another system correlation analysis and fault range localization method provided in an embodiment of this application is shown; Figure 3 A schematic diagram of a system correlation analysis and fault range location device provided in an embodiment of this application is shown. Detailed Implementation
[0012] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.
[0013] In one embodiment, refer to Figure 1 and Figure 2 This paper provides a method for system correlation analysis and fault range localization. Taking the application of this method to computer equipment as an example, the method includes the following steps: Step 101: Construct a system relationship graph using user-created system entities as nodes and inter-system relationships as edges. System entities include basic system information and system tags.
[0014] In this context, a system entity refers to an independent software or hardware unit within an information system that carries business functions, such as the core accounting system, user authentication system, and database system of a bank's data center. Basic system information includes descriptive data such as system name, system description, and importance level. System tags are keywords used to identify the functional attributes or categories of a system. A system can be associated with multiple tags, such as core accounting or transaction processing. Inter-system relationships refer to the logical connections between systems, such as call relationships, dependency relationships, and authentication relationships. These relationships, as edges in a graph, can be accompanied by relationship descriptions and tags, such as upstream, downstream, strong association, weak association, dependency, call, and authentication.
[0015] Specifically, the computer device can provide a user interface to receive system entities manually created by the user, including basic system information entered and selected or customized system labels. Simultaneously, users can draw preliminary relationships between systems on the interface through drag-and-drop or connecting lines, and define relationship labels for each relationship. Subsequently, the computer device can store this data in a graph database, forming an initial inter-system relationship graph. Based on this, the computer device can also integrate network traffic analysis systems, log analysis centers, and other operation and maintenance monitoring tools through automatic learning methods to automatically discover communication calls and other relationships between systems, generate a potential candidate list of relationships, and dynamically add them to the inter-system relationship graph after user confirmation, thereby continuously enriching the graph's completeness and accuracy.
[0016] Step 102: Obtain user access information in real time, update the relationship between systems based on the user access information and the overlap of system tags between systems, and calculate the degree of relationship between systems.
[0017] User access information refers to behavioral data generated when operations and maintenance personnel or system administrators view or operate the graph through the platform. This includes the number of times a user performs a relational search through a certain system, the number of times they are redirected to the related system page after the search, and the duration of their visit to the related system page. The overlap of system tags refers to the number of tags shared by two systems or the semantic similarity between them. The more identical tags, the more related the functions may be. The degree of association is a quantitative indicator used to measure the closeness of the relationship between two systems. The higher the score, the stronger the degree of association. It can also be quantified into tags such as strong association and weak association.
[0018] Specifically, the computer device can periodically collect user access information for each target system. For example, it can daily count the number of times users perform relationship searches via the graph on the target system's interface, the number of times they successfully access other system pages after the search, and the number of times the access duration after the jump exceeds 60 seconds. This information is then used to calculate a first association score between each system and the target system. Next, the computer device can count the number of identical tags between other systems and the target system, and calculate a second association score based on this number. Finally, these two association scores are added together to obtain a cumulative total score between the two systems. If this total score exceeds a preset threshold, such as 10 points, a relationship is determined to exist between the two systems, with higher scores indicating stronger associations. Based on these calculations, the computer device can dynamically update the association strength of corresponding edges in the graph.
[0019] Step 103: Adjust the edges of the inter-system relationship graph according to the updated inter-system relationship, and mark the corresponding edges in the inter-system relationship graph according to the degree of relationship.
[0020] Adjusting the edges of the inter-system relationship graph includes adding new relationship edges, deleting edges that are no longer valid, or modifying the attributes of existing edges; marking the edges refers to adjusting the visual style of the edges according to the degree of association, such as the thickness of the edges, the depth of the colors, or dynamic effects, to intuitively reflect the strength of the relationship.
[0021] Specifically, the computer device can add or modify corresponding relationship edges in the graph database based on the calculated associations and their strength, and simultaneously update the edge weight attributes and relationship labels. Furthermore, in the visualization interface, the computer device can render edges with stronger associations as thicker lines or more prominent colors, and edges with weaker associations as thinner lines or lighter colors, allowing users to easily identify core dependency paths. In addition, the computer device can filter the display based on relationship labels, for example, displaying only strong associations or specific types of relationships.
[0022] Step 104: Obtain alarm information from each system in real time through the monitoring system, and mark the corresponding nodes in the inter-system relationship graph according to the source and level of the alarm information.
[0023] Among them, the monitoring system refers to various monitoring tools deployed in the operation and maintenance environment, such as monitoring and alarm platforms, intelligent analysis centers, network monitoring systems, etc.; alarm information includes alarm source system identifier, alarm level, such as severe, warning, prompt, alarm content description, etc.; marking nodes refers to adjusting the visualization style of nodes according to the alarm status, such as node color, flashing effect or alarm icon, etc.
[0024] Specifically, computer equipment can interface with existing monitoring and alarm platforms via application programming interfaces (APIs) to receive alarm data from various systems in real time or periodically retrieve it. When an alarm message is received from a system, the computer equipment can locate the corresponding system node in the data map based on the source of the alarm and set the display style of the node according to the alarm level; for example, critical alarms are marked in red, warnings in orange, and alerts in yellow. Simultaneously, the node can display an alarm count or a pop-up window displaying detailed alarm content. Furthermore, the computer equipment can record the alarm occurrence time, which serves as an important reference during fault analysis.
[0025] Step 105: In response to the fault analysis request, the graph traversal algorithm is started from the root system selected by the user and traversed along the edges with relationships. Based on the degree of association of each node and alarm information, a fault analysis report is generated.
[0026] Among them, the fault analysis request refers to the operation of the user to initiate a request for fault analysis on the platform interface. This operation requires specifying a system as the starting point of the analysis; the root cause system refers to the system selected as the source of the fault or the assessment center; the graph traversal algorithm screenshot includes graph theory algorithms such as breadth-first search or depth-first search, which are used to explore all reachable nodes starting from the starting point; the fault analysis report can include a list of affected systems, the impact type of each system, the business flows involved, the responsible person information, and optimization suggestions, etc.
[0027] Specifically, when a user selects a root system on the platform, such as the core accounting system, and initiates a fault analysis request, the computer device can initiate a breadth-first search algorithm. Starting from the core accounting system node, it traverses along all outgoing and incoming edges to access all system nodes directly or indirectly related to it. During the traversal, the computer device can classify the traversed systems into different impact types based on the degree of correlation and the current alarm status of each node. For example, systems with a strong correlation and close dependency with the root system are classified as directly impacted, meaning these systems will be immediately blocked due to the root system's failure; systems with a weak correlation or partial call relationship are classified as indirectly impacted, meaning only some functions or transactions are restricted; and systems without a direct correlation but whose business indicators, such as transaction volume and success rate, fluctuate significantly during the root system's failure are classified as potentially impacted. By aggregating this information, a structured fault analysis report can be generated, including the name of the affected system, the type of impact, alarm details, relevant business lines and responsible persons, etc., and optimization suggestions can be given according to preset rules, such as suggesting the implementation of downgrade or isolation measures.
[0028] In a specific example, assume a bank's data center has deployed the method described in this embodiment. Initially, the system administrator can manually create and input system entities such as the core accounting system, user authentication system, transaction risk control system, SMS notification system, and database system into the platform. Each system is assigned a corresponding system label, and relationships are established based on known associations. For example, the core accounting system is labeled "Accounting Core" and "Transaction Processing," while the user authentication system is labeled "Authentication Service" and "Transaction Processing." Dependencies are then established between the core accounting system and the user authentication system. Subsequently, the platform automatically integrates network traffic analysis tools and discovers significant communication between the core accounting system and the external payment gateway. After administrator confirmation, the payment gateway is added as a new node to the graph, and a call relationship is established. During routine maintenance, the platform detects that maintenance personnel frequently search for the transaction risk control system on the core accounting system page and remain there for extended periods, accumulating points exceeding a threshold. Therefore, the association strength between the two systems is marked as strong, and the transaction risk control system node is enlarged and displayed in the graph. One day, the monitoring system detected a large number of timeout alarms in the core accounting system. The platform immediately marked the core accounting system node in red. Simultaneously, the transaction risk control system also generated secondary alarms due to call failures and was marked in orange. The operations and maintenance personnel selected the core accounting system as the root cause system and clicked the impact analysis button. The platform initiated a breadth-first search, traversing the user authentication system, transaction risk control system, database system, and SMS notification system. Combining the correlation strength and alarm information, it determined that the user authentication system and transaction risk control system were directly affected, the database system was indirectly affected, and the SMS notification system was not included in the impact scope due to no direct relationship and no fluctuation in business volume. Subsequently, the platform can generate a fault analysis report, displaying a list of affected systems, the impact type of each affected system, related alarms, and recommended measures, and push this information to the operations and maintenance team. The team can then quickly initiate an emergency response to avoid a wider business interruption.
[0029] The above embodiments construct an inter-system relationship graph using system entities as nodes and inter-system relationships as edges. They dynamically optimize the relationships and strengths between systems using user access information and system tag overlap. Real-time monitoring and alarm information is overlaid on the graph, and a graph traversal algorithm is used to analyze the scope of the fault's impact starting from the root system, generating a fault analysis report. This transforms the originally static inter-system relationships into dynamic, quantifiable relationships and strengths, thus solving the problems of information lag and ambiguous impact boundary positioning caused by traditional manual maintenance. Furthermore, by adding real-time alarm information to the inter-system relationship graph and marking the strength of inter-system relationships, operations and maintenance personnel can quickly identify the core of the fault and the affected systems, replacing manual troubleshooting. This significantly shortens the system's emergency response time and improves the accuracy of secondary disaster assessment.
[0030] In one embodiment, step 101 can be implemented by the following method: receiving system entities created by the user, system tags corresponding to the system entities, and relationships between systems; then building logical relationships based on the system entities, system tags, and relationships between systems to generate an initial relationship graph; receiving and calculating and identifying potential relationships between systems based on the overlap of system tags between systems, as well as collected network traffic data and system application logs; finally, in response to the user's confirmation operation on potential relationships, adjusting the relationships in the initial relationship graph to obtain an inter-system relationship graph.
[0031] In this embodiment, when creating an inter-system relationship graph, the system first receives user-created system entities, corresponding system tags, and inter-system relationships. System entities may include basic system information such as system name, system description, and importance level, as well as system tags identifying system functional attributes. Inter-system relationships may include call relationships, dependency relationships, authentication relationships, etc., and may be accompanied by relationship tags such as upstream system, downstream system, weak association, strong association, dependency, call, authentication, etc. Further, the computer device can construct logical relationships based on system entities, system tags, and inter-system relationships to generate an initial relationship graph. Subsequently, semantic analysis can be performed based on the overlap of system tags, and network traffic data and system application logs can be collected. By analyzing communication patterns, potential inter-system relationships can be identified; for example, two systems with frequent data interactions but not yet established relationships in the graph can be identified as potential relationships. Subsequently, the potential relationships between systems can be pushed to users, and in response to users' confirmation of the potential relationships, the relationships in the initial relationship graph can be adjusted, such as adding new relationship edges or adding new relationship labels to the relationships, thereby obtaining a complete relationship graph between systems.
[0032] The above embodiments combine manual creation with automated learning to construct inter-system relationship graphs, which can effectively improve the accuracy and reliability of the construction of inter-system relationship graphs.
[0033] In one embodiment, step 102 can be implemented by the following method: For any target system, firstly, periodically count the number of searches, the number of redirects after searches, and the number of visits exceeding a preset threshold for accessing each system through the inter-system relationship graph under the target system; then, calculate a first association score between each system and the target system based on the number of searches, the number of redirects, the number of visits, and the weight coefficients corresponding to the number of searches, the number of redirects, and the number of visits; next, calculate a second association score between each system and the target system based on the overlap between the system labels of each system and the system labels of the target system; finally, determine the association relationship and degree of association between each system and the target system based on the sum of the first association score and the second association score and the preset association score threshold.
[0034] In this embodiment, when calculating the relationships and degree of association between systems, for any target system, the number of searches performed on the target system interface through the inter-system relationship graph to retrieve relationships with other systems, the number of redirects to other systems after the search, and the number of accesses exceeding a preset threshold can be periodically counted. The preset threshold can be set according to actual operational needs, such as 60 seconds. Then, the computer device can calculate a first association score for each other system and the target system based on the number of searches, redirects, and accesses, as well as the weighting coefficients pre-configured for these metrics. For example, 0.1 points are awarded for each search, 0.2 points for a successful redirect after the search, and 0.5 points for access timeouts. These scores are then summed to obtain the first association score. Next, the computer device can calculate a second association score based on the overlap between the system tags of each other system and the system tags of the target system, such as 1 point for each shared tag. Finally, the computer device can add the first association score and the second association score to obtain a total score, and compare this total score with a preset association score threshold. If the total score exceeds the threshold, for example, if the total score exceeds 10 points, it is determined that there is an association between the system and the target system, and the higher the total score, the stronger the association. In this embodiment, user behavior information can be collected periodically and the overlap of system tags between systems can be calculated. The association relationship and association strength between systems can be adjusted based on the updated user behavior information and the overlap of system tags.
[0035] In this embodiment, in addition to calculating the correlation and degree of correlation between systems through the overlap of user behavior information and system tags, the correlation and strength between systems can also be adjusted based on the potential correlation between systems, and a potential correlation weight that can be dynamically adjusted during system operation can be preset. For example, when it is subsequently determined that there is a potential correlation between two systems based on actual operation and maintenance information, the potential correlation weight between the two systems can be increased to enhance the correlation strength between the two systems, or a new correlation can be established between the two systems; when it is determined that the potential correlation between the two systems is weak based on actual operation and maintenance information, the potential correlation weight between the two systems can be appropriately reduced to weaken the correlation strength between the two systems, or even delete the correlation between the two systems.
[0036] The above embodiments, by statistically analyzing the number of user searches, jumps, and access durations in the graph, and combining this with weighted integration based on the overlap of system tags, can dynamically quantify and automatically update the relationships between systems. This allows the relationships to no longer rely on static configurations, but rather reflect the relationships and functional similarities between systems in actual operation and maintenance, thereby improving the accuracy and timeliness of the relationship graph.
[0037] In one embodiment, step 105 can be implemented as follows: First, in response to the fault analysis request, a graph traversal algorithm is initiated, starting from the root system selected by the user, and traversing along the edges with relationships to identify at least one affected system; then, based on the relationship and degree of association between the affected system and the root system, the impact type of the affected system is determined, wherein the impact type includes direct impact, indirect impact, and potential impact; finally, a fault analysis report is generated based on the affected system, the impact type of the affected system, and the alarm information of each system.
[0038] In this embodiment, when a user initiates a fault analysis request on the platform interface, the computer device can parse the root cause system selected by the user in the request and initiate a graph traversal algorithm. This algorithm can be, for example, a graph theory algorithm such as breadth-first search or depth-first search. Then, the computer device can start from the node corresponding to the root cause system, traverse along the edges connected to that node in the inter-system relationship graph, and visit all system nodes that are directly or indirectly related to the root cause system, identifying these nodes as affected systems. Then, based on the relationship and degree of association between each affected system and the root cause system, the impact type of each affected system can be determined.
[0039] In this embodiment, the impact types can include direct impact, indirect impact, and potential impact. Direct impact refers to situations where the affected system is heavily dependent on the root system, and will immediately suffer a blocking effect when the root system fails. Indirect impact refers to situations where the affected system has a partial call relationship with the root system, and only some functions or transactions are restricted during a failure. Potential impact refers to situations where the affected system has no direct relationship with the root system, but its business indicators fluctuate significantly during the root system's failure. Finally, the computer equipment can generate a fault analysis report based on all affected systems and their impact types, combined with the alarm information obtained from each system. The fault analysis report may include a list of affected systems, the impact type of each system, relevant alarm information details, involved business flows, and optimization suggestions, among other things.
[0040] The above embodiments use a graph traversal algorithm to search for affected systems starting from the root system, and combine the degree of correlation between systems and real-time alarm information to classify the impact types of different affected systems. This can automatically analyze and accurately locate the scope of fault impact, thereby significantly shortening the fault emergency response time and helping operation and maintenance personnel to accurately assess the boundaries of secondary disasters, thus improving the timeliness of fault handling.
[0041] In one embodiment, the impact type of the affected system can be determined by the following methods: when the relationship between the affected system and the root system is a functional dependency or a strong correlation, the impact type of the affected system is determined to be a direct impact; when the relationship between the affected system and the root system is a data call relationship or a weak correlation, the impact type of the affected system is determined to be an indirect impact; during the failure of the root system, if the fluctuation range of a preset indicator of a system with no correlation is found to be greater than a preset threshold, that system is identified as an affected system, and the impact type of the affected system is determined to be a potential impact.
[0042] In this embodiment, after identifying the affected systems using a graph traversal algorithm, the impact type of each affected system can be further determined. Specifically, the determination method is as follows: When the relationship between the affected system and the root system is a functional dependency relationship—for example, the affected system provides core business logic support to the root system, or the relationship between the two systems is strong—the impact type of the affected system can be determined as a direct impact. Secondly, when the relationship between the affected system and the root system is a data call relationship—for example, the affected system only obtains non-critical data or periodically synchronizes data from the root system, or the relationship between the two systems is weak—the impact type of the affected system can be determined as an indirect impact. Furthermore, during the failure of the root system, if any one or more preset business indicators of other systems that are not directly related to the root system, such as transaction volume, transaction success rate, and average response time, are monitored to fluctuate above a preset threshold—for example, a decrease in transaction success rate exceeding 20%, or an increase in average response time exceeding 100%—this system can be identified as an affected system, and its impact type can be determined as a potential impact.
[0043] In a specific example, after the core accounting system, as the root cause, fails, a graph traversal algorithm identifies four affected systems: the user authentication system, the transaction risk control system, the database system, and the SMS notification system. The user authentication system has a functional dependency on the core accounting system (user authentication is a prerequisite for transaction initiation), and the correlation is strong; therefore, it is considered a direct impact. The transaction risk control system has a real-time data exchange dependency on the core accounting system, also a strong correlation, and is also considered a direct impact. The database system has a data call relationship with the core accounting system (the core accounting system reads and writes to the database), but the database itself does not depend on the core accounting system's business logic, and the correlation is weak; therefore, it is considered an indirect impact. The SMS notification system only receives notification messages from the core accounting system unidirectionally, has a weak correlation, and is a data call relationship; therefore, it is also considered an indirect impact. Meanwhile, computer equipment detected that during the period when the core accounting system malfunctioned, the transaction success rate of the external payment gateway dropped from 99% to 60%, with the fluctuation range far exceeding the preset threshold of 10%. Since the external payment gateway and the core accounting system are not directly related in the graph, the external payment gateway can be identified as a potential impact and included in the list of affected systems.
[0044] The above embodiments categorize affected systems into direct, indirect, and potential impacts based on three dimensions: the type and degree of association, and fluctuations in business metrics. This not only allows for the analysis of fault propagation paths within known dependencies but also enables the discovery of potential unknown related impacts through changes in business metrics. Consequently, the content of fault analysis reports becomes more comprehensive and accurate. Furthermore, it helps operations and maintenance personnel quickly grasp systems with varying degrees of impact and take differentiated emergency measures.
[0045] In one embodiment, refer to Figure 2 In the process of fault range analysis through the inter-system correlation graph, the display mode of the interface can be switched according to the user's operation. The display modes include impact range analysis mode and correlation overview mode. Based on this, the above method also includes the following steps: Step 201: In response to the impact range analysis request, mark and display the node corresponding to the root system selected by the user, and mark and display the affected systems associated with the root system, the coverage of the affected systems, and the alarm information of each system.
[0046] In this embodiment, when a user initiates an impact scope analysis request and selects a root cause system on the platform interface, the computer device can respond to the request by highlighting the node corresponding to the selected root cause system in the system relationship graph. Simultaneously, the graph displays the various affected systems associated with the root cause system and the coverage area of each affected system. Then, the nodes of the affected systems can be differentiated based on their impact type. For example, direct impact types are marked in dark red, indirect impact types in orange, and potential impact types in yellow, etc. Furthermore, the system can display full and partial alarm information near each affected system node, including alarm level, alarm content, and occurrence time, allowing users to promptly observe the impact scope and severity related to the fault, thus enabling timely handling of the affected systems.
[0047] Step 202: In response to the relationship overview request, display the relationship and degree of relationship between each system, or display the system associated with the root system selected by the user and the corresponding relationship, and mark and display the nodes and edges corresponding to each system according to the degree of relationship between the systems.
[0048] In this embodiment, when a user initiates a request for an overview of relationships on the platform interface, the computer device can fully display the relationships and strengths of connections between systems in the graph, and mark and display the nodes and edges corresponding to each system based on the strength of the connections. In this mode, if a root system is selected, the nodes corresponding to the selected root system can be highlighted in the graph, and the nodes of each system that are related to the root system, as well as the connections between each system and the root system, can be marked. Furthermore, the nodes and edges corresponding to each system can be visually marked according to the degree of connection between each system and the root system. For example, nodes with strong connections can be enlarged, nodes with moderate connections can be kept at a normal size, and nodes with weak connections can be reduced in size. Simultaneously, the thickness of the edges related to the connections can also be adjusted accordingly based on the degree of connection; this embodiment does not impose specific limitations on these adjustments.
[0049] The above embodiments provide two visualization modes: fault range viewing and correlation viewing. These modes allow operations and maintenance personnel to switch between views as needed. The fault range viewing mode focuses on real-time alarms and impact types, helping operations and maintenance personnel quickly locate disaster boundaries. The correlation viewing mode focuses on the normal dependency structure and correlation strength between systems, facilitating routine maintenance and change impact assessment. These two modes complement each other, effectively improving the efficiency and usability of the visualization.
[0050] In one embodiment, refer to Figure 2 After generating the fault analysis report, the above-mentioned system correlation analysis and fault range localization method may further include the following methods: First, receive user feedback information on the fault analysis report, wherein the feedback information includes user annotation information on the affected systems missing or falsely reported in the fault analysis report; then, based on the feedback information, determine the affected systems to be adjusted and their corresponding weight adjustment values; finally, based on the affected systems to be adjusted and their corresponding weight adjustment values, update the correlation and degree of correlation between the affected systems to be adjusted and the root source system.
[0051] In this embodiment, after generating the fault analysis report, user feedback can be received. This feedback may include user annotations regarding missing or falsely reported affected systems. For example, a user might annotate a system that should be affected but is not listed in the report, or a system listed in the report that is not actually affected. Based on this feedback, the affected systems requiring adjustment in the system graph and their corresponding weight adjustments can be determined. For instance, for missing systems, the potential association weight between that system and the root system can be increased; for falsely reported systems, the potential association weight between that system and the root system can be decreased. Then, based on the affected systems to be adjusted and their corresponding weight adjustments, the association relationship and degree of association between the affected system and the root system in the system association graph can be updated. The adjusted weight values are then rewritten into the graph database, and the updated association degree is used in subsequent fault analyses.
[0052] For example, after a core accounting system failure, the failure analysis report listed the user authentication system, transaction risk control system, and database system as affected systems. During actual emergency response, operations personnel discovered that although the SMS notification system did not have a direct correlation with the core accounting system in the graph, its SMS sending volume decreased significantly during the failure period, indicating it should be within the potential impact range. Meanwhile, the database system, listed in the report, was actually operating normally and unaffected. Based on this, operations personnel can provide feedback on the failure analysis report on the platform, marking the SMS notification system as a missing system and the database system as a false alarm in the feedback information. Upon receiving the feedback, the computer equipment can increase the potential correlation weight between the SMS notification system and the core accounting system by 5 points and decrease the potential correlation weight between the database system and the core accounting system by 3 points. In subsequent correlation calculations, these adjustments will affect the total correlation score, thereby changing the correlation and degree between the two systems. This makes the SMS notification system more likely to be identified as an affected system in future failure analyses, while the correlation degree of the database system will decrease accordingly.
[0053] The above embodiments receive user feedback on fault analysis reports and dynamically adjust the relationships and degrees of correlation between systems based on the feedback. This allows for continuous optimization of the inter-system relationship map, enabling the map to continuously absorb actual operation and maintenance experience and gradually correct deviations in the initial construction and automated learning map. This effectively improves the accuracy and reliability of fault analysis.
[0054] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. In addition, the labels corresponding to each step in the above embodiments are only for identification purposes and are not intended to limit the execution order of the steps. The execution order of the steps in each embodiment can be set according to the actual situation.
[0055] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this application provides a system correlation analysis and fault range location device, such as... Figure 3 As shown, the device includes: The relationship graph creation module 31 can be used to construct an inter-system relationship graph using user-created system entities as nodes and inter-system relationships as edges, wherein the system entities include basic system information and system tags; The association strength calculation module 32 can be used to obtain user access information in real time, update the association relationship between systems based on the user access information and the overlap of system tags between systems, and calculate the degree of association between systems. The association adjustment module 33 can be used to adjust the edges of the inter-system association graph according to the updated inter-system association relationship, and mark the corresponding edges in the inter-system association graph according to the degree of association; The alarm information marking module 34 can be used to obtain alarm information of each system in real time through the monitoring system, and mark the corresponding nodes in the inter-system relationship graph according to the source and level of the alarm information; The analysis report generation module 35 can be used to respond to fault analysis requests by initiating a graph traversal algorithm that starts from the root cause system selected by the user and traverses along the edges with relationships. Based on the degree of association of each node and alarm information, a fault analysis report is generated.
[0056] In a specific application scenario, the relationship graph creation module 31 can be used to receive system entities created by the user, system tags corresponding to the system entities, and relationships between systems; build logical relationships based on the system entities, system tags, and relationships between systems to generate an initial relationship graph; calculate and identify potential relationships between systems based on the overlap of system tags between systems, as well as collected network traffic data and system application logs; and adjust the relationships in the initial relationship graph in response to the user's confirmation operation of the potential relationships to obtain the relationship graph between systems.
[0057] In specific application scenarios, the association strength calculation module 32 can be used to periodically count, for any target system, the number of searches performed on each system through the inter-system association graph under the target system, the number of redirects after the search, and the number of visits whose access time exceeds a preset threshold; calculate a first association score for each system and the target system based on the number of searches, redirects, and visits, and the weight coefficients corresponding to the number of searches, redirects, and visits; calculate a second association score for each system and the target system based on the overlap between the system tags of each system and the system tags of the target system; and determine the association relationship and degree of association between each system and the target system based on the sum of the first association score and the second association score and a preset association score threshold.
[0058] In a specific application scenario, the analysis report generation module 35 can be used to respond to a fault analysis request by initiating a graph traversal algorithm starting from the root system selected by the user, traversing along the edges with relationships, and identifying at least one affected system; determining the impact type of the affected system based on the relationship and degree of association between the affected system and the root system, wherein the impact type includes direct impact, indirect impact, and potential impact; and generating a fault analysis report based on the affected system, the impact type of the affected system, and the alarm information of each system.
[0059] In specific application scenarios, the analysis report generation module 35 can also be used to determine the impact type of the affected system as direct impact when the relationship between the affected system and the root cause system is a functional dependency relationship or a strong correlation; to determine the impact type of the affected system as indirect impact when the relationship between the affected system and the root cause system is a data call relationship or a weak correlation; and to determine the affected system as a potential impact when, during the failure of the root cause system, the fluctuation range of a preset indicator of a system with no correlation is found to be greater than a preset threshold.
[0060] In specific application scenarios, the device further includes a visualization module 36. The visualization module 36 can be used to respond to an impact range analysis request, mark and display the nodes corresponding to the root system selected by the user, and mark and display the affected systems associated with the root system, the coverage of the affected systems, and the alarm information of each system; respond to an association overview request, display the association relationships and degree of association between each system, or display the systems associated with the root system selected by the user and their corresponding association relationships, and mark and display the nodes and edges corresponding to each system according to the degree of association between each system.
[0061] In specific application scenarios, the correlation adjustment module 33 can also be used to receive user feedback information on the fault analysis report, wherein the feedback information includes user annotation information on the affected systems missing or falsely reported in the fault analysis report; based on the feedback information, determine the affected systems to be adjusted and their corresponding weight adjustment values; and update the correlation relationship and correlation degree between the affected systems to be adjusted and the root cause system based on the affected systems to be adjusted and their corresponding weight adjustment values.
[0062] It should be noted that other corresponding descriptions of the functional units involved in the system correlation analysis and fault range location device provided in this application embodiment can be found in the following references. Figures 1 to 2 The corresponding descriptions in the methods will not be repeated here.
[0063] This application also provides a computer device, specifically a personal computer, server, network device, etc. The computer device includes a bus, processor, memory, and communication interface, and may also include input / output interfaces and a display device. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores location information. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.
[0064] Those skilled in the art will understand that the structure of the computer device described above is only a partial structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. A specific computer device may include more or fewer components, or combine certain components, or have different component arrangements.
[0065] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, having stored thereon a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0066] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0067] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, graphics processors, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0068] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0069] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for system correlation analysis and fault range localization, characterized in that, The method includes: Using user-created system entities as nodes and inter-system relationships as edges, an inter-system relationship graph is constructed, wherein the system entities include basic system information and system tags; Real-time acquisition of user access information, and updating of the inter-system relationships based on the user access information and the overlap of system tags between systems, and calculation of the degree of inter-system association; Based on the updated inter-system relationships, the edges of the inter-system relationship graph are adjusted, and the corresponding edges in the inter-system relationship graph are marked according to the degree of association. The system acquires alarm information from each system in real time through a monitoring system, and marks the corresponding nodes in the inter-system relationship graph according to the source and level of the alarm information. In response to a fault analysis request, the startup graph traversal algorithm starts from the root system selected by the user and traverses along the edges with relationships. Based on the degree of association of each node and alarm information, a fault analysis report is generated.
2. The method according to claim 1, characterized in that, The construction of an inter-system relationship graph, using user-created system entities as nodes and inter-system relationships as edges, includes: Receive user-created system entities, system tags corresponding to system entities, and relationships between systems; Logical relationships are constructed based on the system entities, system tags, and relationships between the systems to generate an initial relationship graph. Based on the overlap of system labels between systems, as well as the collected network traffic data and system application logs, potential relationships between systems are calculated and identified. In response to the user's confirmation of the potential relationship, the relationships in the initial relationship graph are adjusted to obtain the inter-system relationship graph.
3. The method according to claim 1, characterized in that, The step of updating the inter-system relationships based on the user access information and the overlap of system tags between systems, and calculating the degree of inter-system association, includes: For any target system, periodically count the number of searches performed on each system through the inter-system relationship graph under the target system, the number of redirects after the search, and the number of visits whose access time exceeds a preset threshold. Based on the number of searches, the number of redirects, the number of visits, and the weighting coefficients corresponding to the number of searches, the number of redirects, and the number of visits, a first association score is calculated for each system and the target system; Based on the overlap between the system labels of each system and the system labels of the target system, a second association score is calculated for each system and the target system. Based on the sum of the first association score and the second association score, as well as a preset association score threshold, the association relationship and degree of association between each system and the target system are determined.
4. The method according to claim 1, characterized in that, In response to the fault analysis request, the startup graph traversal algorithm starts from the root cause system selected by the user and traverses along the edges with relationships. Based on the degree of association of each node and alarm information, a fault analysis report is generated, including: In response to a fault analysis request, the startup graph traversal algorithm starts from the root system selected by the user and traverses along the edges with relationships to identify at least one affected system. Based on the correlation and degree of correlation between the affected system and the root cause system, the impact type of the affected system is determined, wherein the impact type includes direct impact, indirect impact, and potential impact; A fault analysis report is generated based on the affected systems, the impact types of the affected systems, and the alarm information of each system.
5. The method according to claim 4, characterized in that, The step of determining the impact type of the affected system based on the correlation and degree of correlation between the affected system and the root cause system includes: When the relationship between the affected system and the root system is a functionally dependent relationship or a strong correlation, the impact type of the affected system is determined to be a direct impact. When the relationship between the affected system and the root system is a data call relationship or a weak correlation, the impact type of the affected system is determined to be an indirect impact. During the root cause system failure, if the fluctuation range of a preset indicator of a system with no correlation is found to be greater than a preset threshold, the system is identified as an affected system, and the impact type of the affected system is determined to be a potential impact.
6. The method according to claim 4 or 5, characterized in that, The method includes: In response to the impact range analysis request, the node corresponding to the root system selected by the user is marked and displayed, and the affected systems associated with the root system, the coverage of the affected systems, and the alarm information of each system are marked and displayed. In response to a request for an overview of relationships, the system displays the relationships and degree of association between systems, or displays the systems associated with the root system selected by the user and their corresponding relationships. Based on the degree of association between systems, the system marks and displays the nodes and edges corresponding to each system.
7. The method according to claim 4 or 5, characterized in that, The method includes: Receive user feedback on the fault analysis report, wherein the feedback includes user annotations of affected systems that are missing or falsely reported in the fault analysis report; Based on the feedback information, determine the affected systems to be adjusted and their corresponding weight adjustment values; Based on the affected systems to be adjusted and their corresponding weight adjustment values, the correlation and degree of correlation between the affected systems to be adjusted and the root system are updated.
8. A system correlation analysis and fault range location device, characterized in that, The device includes: The relationship graph creation module is used to construct an inter-system relationship graph using user-created system entities as nodes and inter-system relationships as edges. The system entities include basic system information and system tags. The association strength calculation module is used to acquire user access information in real time, update the association relationship between systems based on the user access information and the overlap of system tags between systems, and calculate the degree of association between systems. The association adjustment module is used to adjust the edges of the inter-system association graph according to the updated inter-system association relationship, and to mark the corresponding edges in the inter-system association graph according to the degree of association. The alarm information marking module is used to obtain alarm information from each system in real time through the monitoring system, and mark the corresponding nodes in the inter-system relationship graph according to the source and level of the alarm information; The analysis report generation module is used to respond to fault analysis requests. It starts the graph traversal algorithm from the root cause system selected by the user and traverses the edges with relationships. Based on the degree of association of each node and alarm information, it generates a fault analysis report.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.
10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.