Fault positioning method and device based on decision tree, medium and product
By obtaining alarm information from the monitoring system of the financial business system, extracting key elements using the rule base, and associating data across multiple platforms, combined with a decision tree orchestration engine for fault location, the problem of low efficiency in fault location in distributed systems is solved, and efficient and accurate fault root cause analysis is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-04-10
AI Technical Summary
In distributed and microservice-based financial business systems, fault location is inefficient. Existing technologies lack a systematic location mechanism, resulting in low fault recovery efficiency. Existing methods rely on manual judgment and have poor accuracy.
By obtaining alarm information from the target monitoring system, extracting key element fields using a pre-set rule base, performing correlation queries on multiple observation platforms, aggregating the data, and inputting it into a decision tree orchestration engine for fault location, a fault location result is generated.
It has achieved automation and intelligence in fault location, significantly improving the accuracy and efficiency of location, and solving the problems of difficulty in integrating multi-source operation and maintenance data and lack of automated analysis methods.
Smart Images

Figure CN121841955A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis, and in particular to a fault location method, device, medium, and product based on decision trees. Background Technology
[0002] As financial business system architectures evolve towards distributed and microservice architectures, the complexity and coupling of core business processes have increased significantly. A single business failure often involves multiple technology platforms and components. Against this backdrop, operations and maintenance personnel face significant challenges when troubleshooting: on the one hand, the long system call chains and dynamically changing dependencies make it difficult to quickly construct an accurate business chain topology view reflecting real-time dependencies, leading to difficulties in analyzing the scope of the failure's impact; on the other hand, failures may originate from different levels, and existing technologies lack a systematic localization mechanism, making it difficult to quickly pinpoint the root cause of the problem and severely impacting failure recovery efficiency.
[0003] Current solutions to these problems primarily rely on the experience of operations and maintenance personnel, who manually analyze system metrics, logs, and link data using various independent tools. Some systems attempt to assist in fault localization by setting threshold alarm rules or displaying static topology diagrams, but these methods have significant limitations. Fault localization decisions still heavily depend on manual judgment and cannot visually demonstrate the fault propagation path. These shortcomings lead to inefficient fault localization and have become a major bottleneck in improving system availability. Summary of the Invention
[0004] This invention provides a fault location method, device, medium, and product based on decision trees to solve the problems of low automation, poor accuracy, and low efficiency in fault location.
[0005] According to one aspect of the present invention, a fault location method based on a decision tree is provided, the method comprising: Obtain target alarm information that matches the target fault scenario from the target monitoring system; The target alarm information is matched with a pre-set rule base, and key element fields are extracted from the target alarm information based on the rule matching results. The key element fields are correlated and queried across multiple target observation platforms, and the resulting query results are aggregated to obtain multidimensional descriptive data that matches the target alarm information. The target fault scenario and multidimensional description data are input into the decision tree orchestration engine. The engine then uses a target decision tree that matches the target fault scenario to generate fault location results that match the target alarm information based on the multidimensional description data.
[0006] According to another aspect of the present invention, a fault location device based on a decision tree is provided, the device comprising: The alarm information acquisition module is used to acquire target alarm information that matches the target fault scenario from the target monitoring system. The feature field extraction module is used to match target alarm information with a pre-set rule base and extract key feature fields from the target alarm information based on the rule matching results. The descriptive data acquisition module is used to perform correlation queries on key element fields across multiple target observation platforms, and aggregate the multiple query results to obtain multidimensional descriptive data that matches the target alarm information. The location result acquisition module is used to input the target fault scenario and multi-dimensional description data into the decision tree orchestration engine. Through the target decision tree in the decision tree orchestration engine that matches the target fault scenario, the fault location result that matches the target alarm information is generated based on the multi-dimensional description data.
[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the decision tree-based fault location method according to any embodiment of the present invention.
[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the decision tree-based fault location method according to any embodiment of the present invention.
[0009] According to another aspect of the present invention, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps of the method as described in any embodiment of the present invention.
[0010] The technical solution of this invention obtains target alarm information matching the target fault scenario from the target monitoring system, performs rule matching between the target alarm information and a pre-set rule base, extracts key element fields based on the matching results, performs correlation queries on multiple target observation platforms for the key element fields, aggregates the multiple query results into multi-dimensional descriptive data, inputs the target fault scenario and multi-dimensional descriptive data into a decision tree orchestration engine, and generates fault location results through a target decision tree matching the target fault scenario. This solves the problems of low efficiency and insufficient accuracy in fault location caused by difficulties in integrating multi-source operation and maintenance data, poor correlation of fault scenarios, and lack of automated analysis methods in the prior art, and achieves the beneficial effects of automating and intelligentizing the fault location process and significantly improving the accuracy and efficiency of location.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a fault location method based on a decision tree according to Embodiment 1 of the present invention; Figure 2 This is a flowchart of another fault location method based on a decision tree provided according to Embodiment 2 of the present invention; Figure 3 This is a flowchart of a fault location method based on a decision tree in a specific scenario applicable to an embodiment of the present invention. Figure 4 This is a schematic diagram of a fault location device based on a decision tree according to Embodiment 3 of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device that implements a fault location method based on a decision tree according to an embodiment of the present invention. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] Example 1 Figure 1 The flowchart below shows a fault location method based on a decision tree according to Embodiment 1 of the present invention. This embodiment is applicable to the situation of rapid and accurate automated fault location of complex systems. The method can be executed by a fault location device based on a decision tree. The device can be implemented in hardware and / or software and is generally configured in electronic devices.
[0017] Among them, the decision tree is a tree-structured analysis model that simulates the reasoning logic of human experts, decomposing the complex fault diagnosis process into a series of hierarchical judgment nodes. Each node sets judgment conditions for specific dimensions of operation and maintenance data, selects different analysis branches based on whether the conditions are met, and gradually converges from the global scope to the specific root cause of the fault, finally forming a clear localization conclusion at the leaf node.
[0018] Correspondingly, such as Figure 1 As shown, the method includes: S110. Obtain target alarm information that matches the target fault scenario from the target monitoring system.
[0019] In this context, the target monitoring system can be understood as a collection of various professional monitoring tools deployed in an enterprise environment. These tools are responsible for real-time monitoring of infrastructure, cloud platforms, application performance, and business transactions. They continuously report abnormal events and performance indicators generated during system operation through standardized interfaces, forming the data source for fault localization. Target alarm information can be understood as standardized event reports generated by the monitoring system that characterize an abnormal state of a specific resource or service. Target fault scenarios can be understood as systemic fault states with clear business impact, indicated by relevant target alarm information.
[0020] In this embodiment, it is first necessary to receive target alarm information in real time from the target monitoring system. This target monitoring system covers multiple aspects such as transaction monitoring and performance / capacity monitoring, and data access is achieved through standardized interfaces to ensure comprehensive acquisition of original alarms or query requests related to potential faults. This target alarm information includes specific target fault scenarios.
[0021] Optionally, based on the above embodiments, obtaining target alarm information matching the target fault scenario from the target monitoring system may include: The system obtains target alarm information generated and reported in real time by the target monitoring system through a preset alarm information reporting interface; or The system sends an alarm information query request to the target monitoring system through a preset alarm information query interface, and receives target alarm information that matches the alarm information query request through the same alarm information query interface.
[0022] The reporting interface can be understood as a dedicated data receiving channel with unidirectional data input. Its core function is to receive alarm information proactively pushed by the target monitoring system when an anomaly is detected. The alarm information query interface can be understood as a data service channel provided externally. Its data flow is on-demand output. It responds to requests from external systems, passively queries and filters its own alarm information, and returns target alarm information that meets the criteria.
[0023] Generally, the fault localization process first requires acquiring target alarm information. Typically, two parallel target alarm information acquisition methods are designed to adapt to different operational scenarios: First, the target monitoring system continuously monitors various metrics. Once a metric exceeds a preset threshold or an abnormal state occurs, the target monitoring system generates a target alarm in real time and reports it through a preset alarm information reporting interface. Second, when operations personnel need to check a specific range based on experience or preliminary signs, they send a specific alarm information query request to the target monitoring system through a preset alarm information query interface. This query request typically includes information such as the time range, host IP, and alarm level. Upon receiving the query request, the target monitoring system matches the alarm information and receives the successfully matched target alarms through the preset alarm information reporting interface.
[0024] In a specific example, when the database response time of a core payment application fluctuates, the target monitoring system generates a target alarm message in real time (containing key fields such as the database node IP, response time threshold exceeding the limit, and the current timestamp), and reports the target alarm message through a preset alarm message reporting interface. Additionally, if operations and maintenance personnel need to investigate potential problems in batch transaction processing within the previous hour, they can send an alarm message query request to the target monitoring system through a preset alarm message query interface, and receive target alarm messages matching the query request through the same interface. These two information acquisition mechanisms together ensure the real-time nature and flexibility of fault data collection.
[0025] S120. Match the target alarm information with the preset rule base, and extract the key element fields from the target alarm information based on the rule matching results.
[0026] The pre-built rule base can be understood as a database storing expert operational knowledge. It systematically defines the identification patterns and processing logic for different alarm types. Each rule includes specific matching conditions and definitions of the data fields to be extracted. Key element fields are the core data units extracted from the target alarm information, typically including the name of the affected business application, the specific server node identifier where the failure occurred, and the specific type classification of the alarm event. These fields serve as index keys for subsequent cross-platform data association, acting as the link to achieve multi-source data aggregation and correlation analysis.
[0027] In this embodiment, after obtaining the target alarm information, a pre-set expert rule base is invoked for analysis. The rule base contains pre-defined processing logic for different alarm types. By parsing the metadata and specific content of the target alarm information, applicable rules are matched, and based on the mapping relationship defined by the rules, key element fields such as application group, service node, and alarm type are accurately extracted, laying a data foundation for subsequent in-depth analysis.
[0028] Optionally, based on the above embodiments, the target alarm information is matched with a preset rule base, and key element fields are extracted from the target alarm information according to the rule matching results. This may include: Metadata and content fields are parsed and obtained from the target alarm information; The metadata and the content fields are matched against the rules included in the rule base, and the target rules that are successfully matched are obtained. Based on the field mapping relationship defined in the target rule, extract the key element fields from the content fields.
[0029] Generally, once target alarm information is obtained, the first task is to perform structured parsing. This process will parse out two different types of data: one is metadata, which describes the attributes of the target alarm information itself, such as which monitoring tool generated the target alarm information, the specific time of its generation, and the severity level; the other is content fields, which contain the specific content carried by the target alarm information, recording in detail which component triggered what anomaly under what conditions.
[0030] Generally, after structured parsing is completed, these two types of information need to be matched with the pre-defined rules in the knowledge base. Each rule clearly defines its applicable scope, such as specific alarm types, source systems, resource tags, or time periods. The matching process involves logically judging the parsed data against the conditions set by the rules. If all conditions are met, the rule is identified as a target rule that successfully matches the current target alarm information, thereby triggering the subsequent analysis and handling process defined by the rule.
[0031] Generally, the target rule predefines a set of field mapping relationships, which clearly indicates which specific data should be found from the content fields of the target alarm information, and assigns them a unified, system-recognizable key element field name.
[0032] S130. Perform correlation queries on key element fields across multiple target observation platforms, and aggregate the resulting query data to obtain multidimensional descriptive data that matches the target alarm information.
[0033] The target observation platform can be understood as a specialized storage and query center for operational data, typically including a log platform, a metrics database, and a link tracing system. These observation platforms record detailed system operation logs, real-time performance metrics, and service call dependencies, collectively forming the core source of observable data. Multidimensional descriptive data can be understood as a unified data view formed by integrating and processing various operational information associated with target alarm information.
[0034] In this embodiment, key element fields extracted from target alarm information are used to initiate correlation queries to multiple target observation platforms (including the log center, monitoring center, and link service center). By constructing standardized query requests, heterogeneous data such as log records, performance indicator sets, and service call chains are obtained in parallel from each target observation platform. Subsequently, using the key element fields as correlation keys, these query result data are time-series aligned and context-correlated. Finally, through aggregation processing, multi-dimensional descriptive data with a unified time-series format that matches the target alarm information is generated, thereby completing the construction of a complete data profile of the fault scenario.
[0035] S140. Input the target fault scenario and multidimensional description data into the decision tree orchestration engine. Through the target decision tree in the decision tree orchestration engine that matches the target fault scenario, generate fault location results that match the target alarm information based on the multidimensional description data.
[0036] The decision tree orchestration engine can be understood as an automated reasoning core capable of executing predefined fault analysis logic. It loads decision trees pre-orchestrated by operations experts, performs clear logical judgments and path selections on the input multi-dimensional data, and ultimately outputs structured fault location conclusions. Specifically, the decision tree orchestration engine pre-configures multiple decision trees, each used for fault location in different fault scenarios.
[0037] In this embodiment, the fused multidimensional data and the target fault scenario are input into the decision tree orchestration engine, triggering the execution of the localization logic of a specific decision tree (i.e., the target decision tree) pre-configured for the target fault scenario. This target decision tree is constructed based on expert knowledge, and by judging the health of the cloud platform, the status of middleware, and the performance of applications layer by layer, the analysis scope gradually converges, and finally generates a clear conclusion at the leaf node that includes the root cause of the fault, a list of affected components, and the severity level, thus completing automated fault localization.
[0038] The technical solution of this invention obtains target alarm information matching the target fault scenario from the target monitoring system, performs rule matching between the target alarm information and a pre-set rule base, extracts key element fields based on the matching results, performs correlation queries on multiple target observation platforms for the key element fields, aggregates the multiple query results into multi-dimensional descriptive data, inputs the target fault scenario and multi-dimensional descriptive data into a decision tree orchestration engine, and generates fault location results through a target decision tree matching the target fault scenario. This solves the problems of low efficiency and insufficient accuracy in fault location caused by difficulties in integrating multi-source operation and maintenance data, poor correlation of fault scenarios, and lack of automated analysis methods in the prior art, and achieves the beneficial effects of automating and intelligentizing the fault location process and significantly improving the accuracy and efficiency of location.
[0039] Example 2 Figure 2 This is a flowchart of another fault location method based on decision trees provided in Embodiment 2 of the present invention. This embodiment is based on the above embodiments and optimized. Specifically, the operation of "inputting the target fault scenario and multi-dimensional description data into the decision tree orchestration engine, and generating a fault location result that matches the target alarm information based on the multi-dimensional description data through the target decision tree in the decision tree orchestration engine that matches the target fault scenario" has been refined.
[0040] Correspondingly, such as Figure 2 As shown, the method includes: S210. Obtain target alarm information that matches the target fault scenario from the target monitoring system.
[0041] S220. Match the target alarm information with the preset rule base, and extract the key element fields from the target alarm information based on the rule matching results.
[0042] S230. Perform correlation queries on key element fields across multiple target observation platforms, and aggregate the resulting query data to obtain multidimensional descriptive data that matches the target alarm information.
[0043] Optionally, based on the above embodiments, the key element fields are correlated and queried across multiple target observation platforms, and the resulting query results are aggregated to obtain multi-dimensional descriptive data matching the target alarm information. This may include: Based on the key element fields and the information query rules corresponding to each target observation platform, standardized query requests corresponding to each target observation platform are constructed respectively. Based on each standardized query request, the data query results of each observable platform are called in parallel to obtain the query result data returned by each target observation platform; The target observation platform includes a log center, a monitoring center, and a link service center. The query result data corresponding to the log center is log records, the query result data corresponding to the monitoring center is a set of performance indicators, and the query result data corresponding to the link service center is service call chain information. Using the key element fields as association keys, time series alignment and context association processing are performed on each query result data, and the processed result data are merged to generate the multidimensional description data in a unified time series format.
[0044] In this context, the association key can be understood as one or more common data fields used to connect and match relevant information from different data sources. Time series alignment refers to the process of arranging and synchronizing multiple sets of data with timestamps according to a unified timeline. Fusion processing refers to the process of integrating multi-source data, after association and time alignment, into a composite data view with a unified structure and inherent connections. Information query rules can be understood as a pre-defined set of standardized instructions used to define how to accurately obtain data related to fault analysis from a specified target observation platform.
[0045] Generally, the processing first requires generating standardized query requests that conform to the query syntax of each platform, based on the extracted key element fields and the information query rules of different observation platforms. These standardized query requests will accurately include key information such as the time range to be queried, the specific application or node identifier, ensuring that relevant data records can be retrieved from the corresponding platform.
[0046] Generally, after generating a standardized query request, data queries are simultaneously initiated to multiple observable platforms, such as the log center, monitoring center, and link service center. This process retrieves three types of data results in parallel (i.e., query result data returned by each target observation platform): detailed text log records from the log center, a set of performance indicator values from the monitoring center, and information on the call relationship chain between services from the link service center, thus forming a multi-faceted data perspective.
[0047] Generally, after obtaining various types of data, the initially extracted key element fields are used as the connecting links to arrange and align the data from different sources in chronological order and establish contextual relationships between the data. Finally, all the linked data is integrated into a unified format to form a multi-dimensional descriptive data containing logs, metrics, and link information, providing a comprehensive data foundation for subsequent analysis.
[0048] S240. Input the target fault scenario into the decision tree orchestration engine and activate the target decision tree in the decision tree orchestration engine that matches the target fault scenario.
[0049] In this embodiment, once the target fault scenario to be analyzed is determined, it is input into the decision tree orchestration engine. The engine retrieves and activates a specific decision tree corresponding to the target fault scenario from a pre-set set of decision trees. For example, for a scenario of abnormal cloud platform resources, a decision tree focused on analyzing infrastructure health will be activated; while for a scenario of slow application service response, another decision tree focused on analyzing middleware dependencies and application performance will be activated, thus ensuring a highly targeted analysis process.
[0050] S250. Provide the multidimensional description data to the root node of the target decision tree, and through the information flow logic defined by each node in the target decision tree, transfer the multidimensional description data level by level until it reaches the target leaf node of the target decision tree.
[0051] The information flow logic can be understood as the data processing rules preset in each node of the decision tree. After verifying the input data according to the judgment conditions of the node, it automatically selects the corresponding branch path that meets the conditions and guides the analysis process to the next level node.
[0052] In this embodiment, after the target decision tree is activated, the aggregated multidimensional descriptive data is injected from the starting node of the decision tree. The data flows from top to bottom along the tree structure. At each judgment node, the data is examined according to the preset analysis logic of that node. For example, a node may determine whether the CPU utilization of a specific cluster of the cloud platform exceeds a threshold, or whether the number of active connections in a certain database connection pool is abnormal. Based on the judgment result, the node selects the corresponding branch path and guides the data to the next level node, thereby achieving gradual convergence and refinement of the analysis scope.
[0053] S260. The list of fault-affected components, the type of fault component, and the fault severity level defined in the target leaf node are determined as the fault location result that matches the target alarm information.
[0054] The list of affected components can be understood as the collection of all system components that experience performance degradation or functional abnormalities due to a root cause failure. This list includes not only components directly connected to the root cause but also upstream and downstream services that depend on these components, clearly outlining the actual impact of the failure across the entire business chain and providing a direct basis for assessing the impact and prioritizing recovery. A faulty component can be understood as the specific system unit ultimately identified as the root cause of the problem through analysis. The severity level of the failure can be understood as a quantitative rating of the degree of impact based on preset standards. This level comprehensively considers factors such as the core nature of the affected business, the scope of user impact, and the duration of service interruption, and is typically divided into different levels such as urgent, major, and minor, to guide the operations team in making decisions regarding the appropriate emergency response speed and resource allocation.
[0055] In this embodiment, when the fault analysis process executes along the branch logic of the decision tree to the target leaf node, the predefined output information in that node—namely, the list of affected components, the type of the faulty components, and the severity level of the fault—is determined as the fault location result matching the target alarm information. This result clearly provides the conclusion of fault root cause location (e.g., a host machine on a cloud platform), the list of affected related components, and the severity rating of the fault.
[0056] Optionally, based on the above embodiments, the target monitoring system includes multiple monitoring platforms, and each monitoring platform includes multiple monitoring nodes; Correspondingly, after generating fault location results that match target alarm information based on multi-dimensional descriptive data through the target decision tree matched with the target fault scenario in the decision tree orchestration engine, it may also include: Obtain a topology view of monitoring nodes that match the target monitoring platform; The fault location results are mapped to the topology view of the monitoring nodes, and the mapping results are visualized.
[0057] In this context, a topology view can be understood as a structural diagram that graphically presents the logical connections and dependencies between various components of a system.
[0058] Generally, a target monitoring system consists of multiple monitoring platforms with different functions. For example, it may include a platform dedicated to monitoring system performance metrics, a platform responsible for log collection, and a platform for tracing service chains. The monitoring capabilities of each monitoring platform are further implemented by distributed monitoring nodes. Each monitoring node represents the entity being monitored, such as a physical server, a virtual machine instance, a container, or a microservice instance.
[0059] Generally, after arriving at a fault location conclusion through decision tree analysis, it is necessary to combine the abstract conclusion with the specific system architecture. This process first requires obtaining a topology diagram reflecting the dependencies between nodes in the monitoring environment. This diagram is usually obtained from a graph database that stores entity relationship data and can clearly show the connections and dependencies between all monitoring nodes.
[0060] Generally, after obtaining the topology diagram, it is necessary to associate the faulty components and their affected areas identified in the location results with the nodes in the diagram. By comparing component identifiers, the root cause node, affected nodes, and their fault attributes are accurately mapped to the corresponding positions in the topology diagram, establishing a visual association between logical conclusions and the physical architecture.
[0061] Generally, after mapping is completed, different visual elements are used to distinguish and display the status of various nodes. For example, the root cause node of the fault is marked in red, the affected nodes are marked in yellow, and the potential propagation path of the fault is shown through highlighted connecting lines. The final visualization view can intuitively present the location of the fault point, the scope of impact, and the propagation relationship, providing maintenance personnel with a clear display of problem location.
[0062] Optionally, based on the above embodiments, mapping the fault location results to the monitoring node topology view and visually displaying the mapping results may include: The topology data of the monitoring node topology view is obtained from the graph database, wherein the topology data includes the plurality of monitoring nodes and the dependencies between the nodes; The list of fault-affected components in the fault location results is matched with the monitoring nodes in the topology data to determine the set of abnormal monitoring nodes; Based on the fault component type and fault severity level in the fault location results, configure corresponding visualization attributes for each abnormal monitoring node in the abnormal monitoring node set. In the monitoring node topology view, the abnormal monitoring node is rendered and displayed based on the visualization attributes, and the fault propagation path is displayed based on the dependency relationship between the nodes.
[0063] Graph databases can be understood as databases specifically designed for storing and querying complex relationships between entities. They represent data as nodes and edges connecting those nodes, making them ideal for expressing and quickly retrieving interconnected data such as network topology, social relationships, or dependencies. In this embodiment, they are used to efficiently store and query dependencies between monitoring nodes. Visual attributes can be understood as the visual characteristics given to graphical elements to intuitively convey information on a graphical interface. These attributes include color, shape, size, transparency, and animation effects. By configuring different visual attributes for objects of different states or types, observers can quickly understand the state, type, or severity level they represent; for example, red indicates a serious fault, and yellow indicates a warning.
[0064] Generally, the first step is to query and retrieve the topology data corresponding to the monitoring node topology view describing the entire system architecture from a graph database that specializes in storing relational data. This topology data is essentially structured relational data, primarily consisting of two elements: first, information on multiple monitoring nodes representing all monitored entities (such as servers, service instances, etc.) in the system; and second, clearly defined dependencies between these monitoring nodes that define their mutual calls and dependencies. This complete set of topology data forms the foundation for subsequent fault impact analysis, path display, and visualization rendering. Typically, precise identification matching is performed by comparing the list of fault-affected components in the fault location results with the unique identifiers (such as IP addresses, service names, etc.) of all monitoring nodes in the topology data obtained from the graph database. This matching operation aims to identify nodes that exist in the topology structure and also appear in the fault impact list, thereby accurately filtering and determining the set of abnormal monitoring nodes that need to be highlighted, providing precise target objects for subsequent visualization marking and impact analysis.
[0065] Generally, two core attributes are extracted based on the fault location results: the type of the faulty component (such as database, middleware, application service, etc.) and the severity level of the fault (such as fatal, severe, warning, etc.). Based on predefined mapping rules, corresponding visual attributes are dynamically assigned to each abnormal monitoring node in the identified set of abnormal monitoring nodes. These attributes typically include visual elements such as the node's display color, icon shape, and flashing frequency, thereby intuitively transforming abstract conclusions such as the type and severity of the fault into easily understandable graphical information in the topology view.
[0066] Generally, in the constructed monitoring node topology view, the nodes are first rendered differently based on the visualization attributes (such as color, icon, etc.) configured for each abnormal monitoring node to intuitively reflect their fault status. At the same time, based on the dependencies between nodes defined in the topology data, the fault propagation path from the fault source to the affected nodes is automatically analyzed and drawn, usually displayed in the form of highlighted lines or arrow animations, thus fully presenting the scope of the fault's impact and the propagation chain.
[0067] The technical solution of this invention obtains target alarm information matching the target fault scenario from the target monitoring system, performs rule matching between the target alarm information and a pre-set rule base to extract key element fields, uses the key element fields to perform correlation queries and aggregation processing on multiple target observation platforms to obtain multi-dimensional descriptive data, and then inputs the target fault scenario and multi-dimensional descriptive data into a decision tree orchestration engine to activate the matching target decision tree. Through the information flow logic defined by the nodes, the multi-dimensional descriptive data is processed step by step, and finally, a fault location result containing a list of fault-affected components, fault component types, and fault severity levels is generated at the target leaf node. This solves the problems of low efficiency and insufficient accuracy in fault location caused by reliance on human experience in the prior art. In particular, through the automated analysis and reasoning of the decision tree engine, it achieves accurate location of complex fault root causes and rapid determination of the scope of influence, and achieves the beneficial effect of significantly improving the automation level of fault handling and operation and maintenance efficiency.
[0068] To facilitate understanding, the specific operation and maintenance scenarios to which this technical solution is applicable are described. In this specific embodiment, in order to solve the problem of difficulty in locating faults in complex financial systems due to multi-source alarms and cross-platform dependencies, this embodiment of the invention designs a complete fault location scheme based on decision trees.
[0069] Specifically, Figure 3 A flowchart illustrating a decision tree-based fault location solution is shown below. Figure 3As shown, this solution begins by collecting and aggregating multi-source alarm information, including faults A, B, C, and D, from various alarm platforms. Next, it performs element analysis on the alarms, extracting key element fields such as application groups and service nodes. Based on these elements, it conducts horizontal and vertical strong dependency analysis, outlining and presenting the complete architecture from the underlying cloud platform infrastructure (including campus, cluster, and host machine levels), upper-layer application services, to various middleware components, and their inter-call and inter-dependency relationships. Then, by aggregating anomaly logs, performance metrics, and link data from multiple levels such as applications, middleware, and infrastructure, a unified multi-dimensional indicator view is formed. Finally, relying on a decision tree analysis engine, the aggregated panoramic data is comprehensively analyzed to achieve precise fault location, clearly distinguishing and determining whether the root cause of the fault belongs to the cloud platform infrastructure, a specific application service, or a specific middleware component, thus completing the closed-loop process from alarm perception to root cause location.
[0070] Example 3 Figure 4 This is a schematic diagram of a fault location device based on a decision tree, provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes: The alarm information acquisition module 410 is used to acquire target alarm information that matches the target fault scenario from the target monitoring system. The feature field extraction module 420 is used to match the target alarm information with a preset rule base and extract key feature fields from the target alarm information based on the rule matching results. The descriptive data acquisition module 430 is used to perform correlation queries on key element fields across multiple target observation platforms, and to aggregate the multiple query result data to obtain multi-dimensional descriptive data that matches the target alarm information. The location result acquisition module 440 is used to input the target fault scene and multi-dimensional description data into the decision tree orchestration engine. Through the target decision tree in the decision tree orchestration engine that matches the target fault scene, the fault location result that matches the target alarm information is generated based on the multi-dimensional description data.
[0071] The technical solution of this invention obtains target alarm information matching the target fault scenario from the target monitoring system, performs rule matching between the target alarm information and a pre-set rule base, extracts key element fields based on the matching results, performs correlation queries on multiple target observation platforms for the key element fields, aggregates the multiple query results into multi-dimensional descriptive data, inputs the target fault scenario and multi-dimensional descriptive data into a decision tree orchestration engine, and generates fault location results through a target decision tree matching the target fault scenario. This solves the problems of low efficiency and insufficient accuracy in fault location caused by difficulties in integrating multi-source operation and maintenance data, poor correlation of fault scenarios, and lack of automated analysis methods in the prior art, and achieves the beneficial effects of automating and intelligentizing the fault location process and significantly improving the accuracy and efficiency of location.
[0072] Based on the above embodiments, the alarm information acquisition module 410 can specifically be used for: The system obtains target alarm information generated and reported in real time by the target monitoring system through a preset alarm information reporting interface; or The system sends an alarm information query request to the target monitoring system through a preset alarm information query interface, and receives target alarm information that matches the alarm information query request through the same alarm information query interface.
[0073] Based on the above embodiments, the feature field extraction module 420 can specifically be used for: Metadata and content fields are parsed and obtained from the target alarm information; The metadata and the content fields are matched against the rules included in the rule base, and the target rules that are successfully matched are obtained. Based on the field mapping relationship defined in the target rule, extract the key element fields from the content fields.
[0074] Based on the above embodiments, the data acquisition module 430 can be specifically used for: Based on the key element fields and the information query rules corresponding to each target observation platform, standardized query requests corresponding to each target observation platform are constructed respectively. Based on each standardized query request, the data query results of each observable platform are called in parallel to obtain the query result data returned by each target observation platform; The target observation platform includes a log center, a monitoring center, and a link service center. The query result data corresponding to the log center is log records, the query result data corresponding to the monitoring center is a set of performance indicators, and the query result data corresponding to the link service center is service call chain information. Using the key element fields as association keys, time series alignment and context association processing are performed on each query result data, and the processed result data are merged to generate the multidimensional description data in a unified time series format.
[0075] Based on the above embodiments, the positioning result acquisition module 440 can specifically be used for: The target failure scenario is input into the decision tree orchestration engine, which then activates the target decision tree that matches the target failure scenario in the decision tree orchestration engine. The multidimensional description data is provided to the root node of the target decision tree, and the multidimensional description data is transferred level by level through the information flow logic defined by each node in the target decision tree until it reaches the target leaf node of the target decision tree. The list of affected components, the type of the faulty component, and the severity level of the fault, defined in the leaf node of the target, are used as the fault location results that match the target alarm information.
[0076] Based on the above embodiments, the target monitoring system includes multiple monitoring platforms, and each monitoring platform includes multiple monitoring nodes; Accordingly, based on the above embodiments, the fault location device based on decision trees may further include: The topology view acquisition module is used to acquire a monitoring node topology view that matches the target monitoring platform after generating a fault location result that matches the target alarm information based on the multi-dimensional description data through the target decision tree matching the target fault scenario in the decision tree orchestration engine. The visualization module is used to map fault location results to the topology view of monitoring nodes and to visualize the mapping results.
[0077] Based on the above embodiments, the visualization module can specifically be used for: The topology data of the monitoring node topology view is obtained from the graph database, wherein the topology data includes the plurality of monitoring nodes and the dependencies between the nodes; The list of fault-affected components in the fault location results is matched with the monitoring nodes in the topology data to determine the set of abnormal monitoring nodes; Based on the fault component type and fault severity level in the fault location results, configure corresponding visualization attributes for each abnormal monitoring node in the abnormal monitoring node set. In the monitoring node topology view, the abnormal monitoring node is rendered and displayed based on the visualization attributes, and the fault propagation path is displayed based on the dependency relationship between the nodes.
[0078] The fault location device based on decision tree provided in the embodiments of the present invention can execute the fault location method based on decision tree provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0079] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0080] Example 4 Figure 5 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0081] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0082] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0083] Processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as performing a decision tree-based fault location method as described in any embodiment of the present invention, i.e.: Obtain target alarm information that matches the target fault scenario from the target monitoring system; The target alarm information is matched with a pre-set rule base, and key element fields are extracted from the target alarm information based on the rule matching results. The key element fields are correlated and queried across multiple target observation platforms, and the resulting query results are aggregated to obtain multidimensional descriptive data that matches the target alarm information. The target fault scenario and multidimensional description data are input into the decision tree orchestration engine. The engine then uses a target decision tree that matches the target fault scenario to generate fault location results that match the target alarm information based on the multidimensional description data.
[0084] In some embodiments, a decision tree-based fault location method as described in any one of the embodiments of the present invention can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the decision tree-based fault location method described above as described in any one of the embodiments of the present invention can be performed. Alternatively, in other embodiments, processor 11 can be configured by any other suitable means (e.g., by means of firmware) to perform the decision tree-based fault location method as described in any one of the embodiments of the present invention.
[0085] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0086] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0087] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0088] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0089] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0090] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0091] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0092] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A decision tree based fault location method, characterized in that, The method comprises: obtaining target alarm information matched with a target fault scenario from a target monitoring system; performing rule matching on the target alarm information and a preset rule library, and extracting a key element field from the target alarm information according to a rule matching result; performing associated query on the key element field in multiple target observation platforms, and performing aggregation processing on multiple query result data obtained, to obtain multi-dimensional description data matched with the target alarm information; inputting the target fault scenario and the multi-dimensional description data into a decision tree arrangement engine, generating a fault positioning result matched with the target alarm information for the multi-dimensional description data via a target decision tree matched with the target fault scenario in the decision tree arrangement engine.
2. The method of claim 1, wherein, The method comprises: obtaining target alarm information matched with a target fault scenario from a target monitoring system, comprising: obtaining target alarm information generated and reported in real time in the target monitoring system through a preset alarm information reporting interface; or 3. The method of claim 1, wherein, sending an alarm information query request to the target monitoring system through a preset alarm information query interface, and receiving target alarm information matched with the alarm information query request via the alarm information query interface. Performing rule matching on the target alarm information and a preset rule library, and extracting a key element field from the target alarm information according to a rule matching result, comprising: parsing metadata and content fields in the target alarm information; matching the metadata and the content fields with each rule included in the rule library respectively, and obtaining a target rule matched successfully; 4. The method of claim 1, wherein, extracting a key element field from the content field according to a field mapping relationship defined in the target rule. Performing associated query on the key element field in multiple target observation platforms, and performing aggregation processing on multiple query result data obtained, to obtain multi-dimensional description data matched with the target alarm information, comprising: constructing a standardized query request corresponding to each target observation platform respectively according to the key element field and an information query rule corresponding to each target observation platform respectively; calling data query results of each observable platform in parallel according to each standardized query request, and obtaining query result data fed back by each target observation platform respectively; wherein the target observation platform comprises a log center, a monitoring center and a link service center; the query result data corresponding to the log center is log record, the query result data corresponding to the monitoring center is performance index set, and the query result data corresponding to the link service center is service call chain information; 5. The method of claim 1, wherein, performing time series alignment and context association processing on each query result data by taking the key element field as an association key, and performing fusion processing on each processing result data to generate the multi-dimensional description data in a unified time sequence format. Inputting the target fault scenario and the multi-dimensional description data into a decision tree arrangement engine, generating a fault positioning result matched with the target alarm information for the multi-dimensional description data via a target decision tree matched with the target fault scenario in the decision tree arrangement engine, comprising: inputting the target fault scenario into the decision tree arrangement engine to activate the target decision tree matched with the target fault scenario in the decision tree arrangement engine; The multi-dimensional description data is provided to a root node of the target decision tree, and the multi-dimensional description data is sequentially transferred through information flow transfer logic defined by each node in the target decision tree until a target leaf node of the target decision tree is reached. The fault impact component list, the type of the fault component, and the fault severity level defined in the target leaf node are determined as the fault locating result matched with the target alarm information.
6. The method according to any one of claims 1 to 5, characterized in that, The target monitoring system includes a plurality of monitoring platforms, and each monitoring platform includes a plurality of monitoring nodes. Correspondingly, after the target decision tree matched with the target fault scenario in the decision tree arrangement engine generates the fault locating result matched with the target alarm information for the multi-dimensional description data, the method further includes: Obtaining a monitoring node topology view matched with the target monitoring platform; Mapping the fault locating result to the monitoring node topology view, and visualizing and displaying the mapping result.
7. The method of claim 6, wherein, Mapping the fault locating result to the monitoring node topology view, and visualizing and displaying the mapping result, includes: Obtaining topology data of the monitoring node topology view from a graph database, wherein the topology data includes the plurality of monitoring nodes and the dependency relationship between the nodes; Identifying and matching the fault impact component list in the fault locating result with the monitoring nodes in the topology data to determine a set of abnormal monitoring nodes; According to the type of the fault component and the fault severity level in the fault locating result, configuring a corresponding visual attribute for each abnormal monitoring node in the set of abnormal monitoring nodes; In the monitoring node topology view, rendering and displaying the abnormal monitoring nodes based on the visual attribute, and displaying the fault propagation path based on the dependency relationship between the nodes.
8. An electronic device, comprising: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the decision tree-based fault locating method of any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to execute the decision tree-based fault locating method of any one of claims 1-7 when executed by the processor.
10. A computer program product, characterised in that, The computer program product includes a computer program that, when executed by a processor, implements the decision tree-based fault locating method according to any one of claims 1-7.