Method for fault localization and apparatus
By generating business topology models and determining exception subgraphs, the problem of over-reliance on manual experience in the prior art is solved, and efficient and accurate problem delimitation is achieved.
Patent Information
- Application Number
- PCT/CN2024/133627
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-11-21
- Publication Date
- 2025-06-19
AI Technical Summary
Existing problem-delimiting techniques rely too much on manual experience, resulting in high delivery costs and low efficiency.
By obtaining business data, generating business topology models, determining exception subgraphs, and determining the root cause node based on the attribute information of nodes and edges, reducing dependence on manual experts.
It has achieved rapid narrowing of the analysis scope, accurately identifying the root cause nodes, reducing delivery costs, and improving problem delimiting efficiency.
Smart Images

Figure CN2024133627_19062025_PF_FP_ABST
Abstract
Description
A method and device for problem demarcation
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 15, 2023, with application number 202311734804.X, and priority to the Chinese patent application entitled “A method and device for problem delimiting”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present invention relates to the field of network technology, and in particular to a method and device for problem demarcation. Background Art
[0003] The field of network technology refers to the broad area of using various technologies and equipment to create, manage, and operate networks for data, voice, and video communications. Network technology includes, but is not limited to, routers, switches, bridges, and various types of communications hardware, as well as the software and protocols that support this hardware. These technologies allow different computers, servers, and other devices to connect and exchange data with each other, forming the foundation of the modern Internet and various private networks.
[0004] Currently, a method has been proposed to determine the root cause of network failures by manually setting up fault trees. Specifically, this method mainly builds design-state rule orchestration capabilities and an operational-state rule execution engine. Based on the design-state rule orchestration capabilities, expert experience is manually configured to form executable fault tree rules. The fault tree rules are then injected into the operational-state rule engine. The operational-state engine parses and executes the fault tree rules to provide demarcation results, thereby forming an automatic problem analysis capability.
[0005] However, existing problem demarcation technologies all rely on design-state tools and manually configured fault tree rules to demarcate root cause nodes. This approach relies on expert rule configuration for each scenario. Furthermore, the complexity and dynamic nature of network environments require that fault tree rules be constantly updated to adapt to new situations. This process is not only time-consuming and costly, but also overly reliant on manual experience, resulting in high delivery costs and low efficiency. Therefore, addressing the overreliance on manual experience, resulting in high delivery costs and low efficiency, is a pressing technical issue. Summary of the Invention
[0006] The embodiments of the present application provide a method and apparatus for problem demarcation, which solve the problem that existing problem demarcation technologies rely too much on manual experience, resulting in high delivery costs and low efficiency.
[0007] In a first aspect, an embodiment of the present application provides a method for problem demarcation, which may include:
[0008] Obtain business data, wherein the business data includes M network entities, N business paths and business statistical information, where M and N are both positive integers; generate a first business topology model based on the business data, wherein the first business topology model includes M nodes, N edges, first attribute information of the M nodes, and second attribute information of the N edges, wherein one of the M nodes corresponds to a network entity in the M network entities, and one of the N edges corresponds to a business path in the N business paths; determine J abnormal subgraphs in the first business topology model, wherein each of the abnormal subgraphs includes one or more abnormal nodes and abnormal edges linking the one or more abnormal nodes, the abnormal node is a node in the first business topology model that meets a preset abnormal condition, and J is a positive integer; based on the first attribute information corresponding to each node and the second attribute information corresponding to each edge in the first business topology model, and the J abnormal subgraphs, determine J first root cause nodes, where one abnormal subgraph corresponds to one first root cause node.
[0009] In the prior art, in order to improve the user's network experience, rules are usually formulated based on expert experience, and then based on the rules, demarcation analysis is performed on the problems that occur in the network. Since each network service has its own uniqueness, such as different equipment, software and usage scenarios may all be different, experts may need to set different fault tree rules according to different network services, and may require continuous investment and continuous optimization by experts. This will lead to low efficiency and a large amount of human resources and time resources will be wasted. In response to this technical problem, in the embodiment of the present application, different business topology models can be generated based on different business data, and then each business topology model can be analyzed. The scope of analysis can be quickly narrowed, and then a step-by-step analysis can be performed based on network data and performance indicators to identify the root cause node that is truly problematic. There is no need to rely on the experience and knowledge of human experts to set rules for different network environments, and the problem demarcation results can be obtained, which greatly reduces the reliance on human experience in the prior art, reduces delivery costs, and improves system efficiency. Specifically, in an embodiment of the present application, during the problem demarcation process, the intelligent demarcation platform creates a business topology model for the business data. The business topology model may include nodes and edges (wherein one of the M nodes corresponds to one of the M network entities, and one of the N edges corresponds to one of the N business paths), as well as attribute information corresponding to the nodes and edges (for example, statistics of poor quality events and business volume of network elements and business paths in the business data). Furthermore, based on the statistical values of the business data, abnormal subgraphs are segmented to determine J abnormal subgraphs, which can quickly narrow the analysis scope and find the minimum business topology range where the problem exists. Based on the attribute information of the nodes and edges in the business topology model, as well as the attribute information of the abnormal nodes and the attribute information of the abnormal edges in each abnormal subgraph, the root cause node corresponding to each abnormal subgraph can be finally determined. In summary, in the embodiments of the present application, step-by-step analysis can be performed based on different business data to narrow the scope, and finally the root cause node can be determined based on attribute information and abnormal subgraphs. Problem demarcation can be performed without relying on the experience configuration rules of human experts. Therefore, the dependence on human experts in the problem demarcation process is greatly reduced, the delivery cost is reduced, and the efficiency of problem demarcation is improved.
[0010] In one possible implementation, generating the first business topology model includes: constructing a second business topology model based on the business data, the second business topology model including M nodes and N edges; determining the first attribute information of each of the network entities and the second attribute information of each of the business paths based on the business statistical information; generating the first business topology model based on the M nodes, the N edges, the first attribute information of the M nodes and the second attribute information of the N edges. In an embodiment of the present application, a business topology model is first constructed using business data, and then attribute information is determined based on the business data, so that the complete link of the business can be portrayed, and the specific node or path where the problem occurs can be accurately displayed, so as to support subsequent problem demarcation based on the business topology model; specifically, in an embodiment of the present application, a second business topology model is first constructed using business data, and the second business topology model includes nodes and edges. The attribute information of each node and edge (for example, poor quality events and business volume statistics in network elements and business paths) can be determined based on the business statistical information in the business data, and then a first business topology model is generated based on the attribute information and the nodes and edges in the second business topology model. Through the above-mentioned node attributes, poor and good quality of experience can be recorded, and the complete link of the business can be portrayed; in the prior art, simple business data may not be able to accurately point out the specific location of the problem, especially in a complex network environment, but in an embodiment of the present application, a business topology model with business data can be generated, and a complete business link can be restored based on the business topology model. Since the connection and dependency relationship of each network entity is displayed in the business topology model, the source of the problem can be located more accurately and quickly. In summary, in the embodiment of the present application, a second business topology model with nodes and edges and the attribute information of each node and edge therein are determined through business data, and then a first business topology model is generated based on the above-mentioned second business topology model and the corresponding attribute information to form a business topology model with business data. Based on the business data, a complete business link can be portrayed, and the specific node or path where the problem occurs can be accurately displayed, so as to support subsequent problem demarcation based on the business topology model.
[0011] In one possible implementation, the service statistics include abnormal events and service volume of each of the M network entities, and abnormal events and service volume of each of the N service paths. In the embodiment of the present application, abnormal events (e.g., "poor quality events") generally refer to events whose quality does not meet standards or expectations, while "service volume" refers to the number or volume of services completed within a certain period of time. By obtaining poor quality events and service volume data on network nodes and inter-node links through service data, and by observing the service volume and poor quality events of each node or edge in the service, the location of the node with the actual problem can be quickly pointed out, which helps to shorten the time for problem demarcation and improve efficiency.
[0012] In a possible implementation, the first attribute information of each of the M network entities includes the statistical value of the abnormal event and the statistical value of the business volume; the second attribute information of each business path in the N business paths includes the statistical value of the abnormal event and the statistical value of the business volume. In an embodiment of the present application, by calculating the statistical value of abnormal events (such as poor quality events) and business volume for each network entity and each business path, further analysis can be performed in the subsequent problem demarcation process based on these attribute information to confirm the various abnormal subgraphs in the business topology model, so as to more accurately identify the root cause node. For example, the weight ratio of each node in the business topology model can be confirmed by the statistical value of the poor quality event and the statistical value of the business volume, and then the failure probability of each abnormal node in the abnormal subgraph can be confirmed based on the weight ratio, and finally the corresponding root cause node can be determined in the abnormal subgraph. In an embodiment of the present application, the data of poor quality events and business volume can be used to help determine the final root cause node to improve the accuracy of the problem demarcation solution.
[0013] In a possible implementation, the determining of J first root cause nodes based on the first attribute information corresponding to each node in the first business topology model and the second attribute information corresponding to each edge, as well as the J abnormal subgraphs, includes: determining the weight ratio of the business volume of each node in the M nodes based on the statistical value of the abnormal event and the statistical value of the business volume in the first attribute information of each node in the first business topology model, and the statistical value of the abnormal event and the statistical value of the business volume in the second attribute information of each edge; determining the weight ratio of the business volume of each abnormal node in each abnormal subgraph based on the weight ratio of each node in the M nodes; determining the failure probability of each abnormal node in each abnormal subgraph based on the weight ratio of the business volume of each abnormal node, as well as the statistical value of the abnormal event in the first attribute information and the statistical value of the business volume; and determining the first root cause node in each abnormal subgraph based on the failure probability of each abnormal node. In existing technologies, problem demarcation technology can often only identify the problematic node, but cannot determine whether the node is the main cause or affected by the transmission relationship. If the real root cause node needs to be found, experts are required to make manual judgments. Due to manual misjudgment, the node found may not be the main cause of the problem, but an affected node, which affects the accuracy of problem demarcation and causes a large amount of resource waste. To address this technical problem, in an embodiment of the present application, the weight of each node in the business is first determined through an intelligent delimitation platform (for example, the weight can be the weight of the business volume of each node in the business, or the weight ratio of the importance of the quality-poor events and business volume of each node to the total quality-poor events and business volume in the business), and then the failure probability of each node is determined in each abnormal subgraph based on the weight and attribute information of each node. The root cause node can be determined based on the failure probability. For example, the root cause node can be the node with the highest failure probability. Therefore, determining the root cause node based on the failure probability can reduce the probability of misjudgment and greatly improve the accuracy of problem delimitation. In summary, in an embodiment of the present application, the weight ratio of each node can be analyzed based on the intelligent delimitation platform, and then the failure probability of each abnormal node in each abnormal subgraph is determined. Finally, the root cause node is determined based on the failure probability. No manual intervention is required in the middle. The process of finding the root cause node can be independently completed by the intelligent delimitation platform, reducing the possibility of manual misjudgment, greatly improving the accuracy of problem delimitation, and at the same time, reducing delivery costs.
[0014] In one possible implementation, the method is applied to an intelligent delimitation platform, wherein the intelligent delimitation platform includes a fault probability model; determining J first root cause nodes based on the first attribute information corresponding to each node and the second attribute information corresponding to each edge in the first business topology model, as well as the J abnormal subgraphs, includes: inputting the first attribute information of each node and the second attribute information of each edge in the first business topology model into the fault probability model, outputting the fault probability of each abnormal node in each abnormal subgraph; and determining the first root cause node in each abnormal subgraph based on the failure probability of each abnormal node. In an embodiment of the present application, the attribute information in the first business topology model and the abnormal subgraph can be input into the fault probability model to determine the failure probability of each abnormal node in each abnormal subgraph (for example, the attribute information of the node and the edge can be input into the fault probability model to determine the weight ratio of the node, and then the corresponding failure probability is determined based on the attribute information and weight ratio of the node in the abnormal subgraph), and then the corresponding root cause node is found according to the failure probability of the node in each abnormal subgraph; the fault probability model can be a pre-trained classification machine learning algorithm model, and the fault probability is determined based on the fault probability model, so that rapid fault location can be performed in subsequent delimitation, the corresponding root cause node can be found, and the accuracy of the delimitation result can be increased; in summary, in an embodiment of the present application, the attribute information of the node and the edge can be input into the fault probability model to determine the failure probability of each abnormal node in the abnormal subgraph and the corresponding root cause node, so as to improve the accuracy and efficiency of the delimitation result.
[0015] In a possible implementation, the fault probability model includes a weight proportion sub-model and a fault probability sub-model; the first attribute information of each node and the second attribute information of each edge in the first business topology model are input into the fault probability model, and the failure probability of each abnormal node in each abnormal sub-graph is output, including: the statistical value of the abnormal event in the first attribute information of each node and the statistical value of the business volume, the statistical value of the abnormal event in the second attribute information of each edge and the statistical value of the business volume are input into the weight proportion sub-model, and the weight proportion of the business volume of each node in the M nodes is output; the weight proportion of the business volume of each abnormal node, and the statistical value of the abnormal event in the first attribute information and the statistical value of the business volume are input into the fault probability sub-model, and the failure probability of each abnormal node in each abnormal sub-graph is output. In an embodiment of the present application, the fault probability model cooperates with the nested sub-models therein to jointly realize the function of determining the fault probability. Specifically, the fault probability model includes two nested sub-models, and the weight ratio of each node in the first business topology model can be determined by the weight ratio sub-model (for example, the weight ratio of the node can be determined by inputting the statistical values of the number of poor quality events and the business volume of the node and the edge into the weight ratio sub-model), and the failure probability of each abnormal node in each abnormal sub-graph can be determined in the fault probability sub-model through the weight ratio of each node mentioned above (for example, the corresponding failure probability can be determined by inputting the statistical values of the number of poor quality events and the business volume of the abnormal nodes in each abnormal sub-graph, and the weight ratio of the abnormal nodes into the fault probability sub-model); in an embodiment of the present application, the fault probability sub-model can be based on the weight ratio determined by the weight ratio sub-model to determine the failure probability of each abnormal node in the abnormal sub-graph,
[0016] By implementing the functionality of the fault probability model through a nested sub-model division of labor and cooperation approach, the accuracy of the intelligent delimitation platform in determining fault probabilities can be significantly improved, facilitating more accurate root cause node identification and problem delimitation.
[0017] In one possible implementation, the weight ratio sub-model includes a first sub-model in an unsupervised stage; in the first sub-model in the unsupervised stage, the weight ratio of the business volume of each of the M nodes output is preset based on experience, or, in the first sub-model in the unsupervised stage, the weight ratio of the business volume of each of the M nodes output is obtained by training the first sub-model based on historical sample documents, wherein the historical sample documents include abnormal events and business volume of each node and abnormal events and business volume of each edge in the first business topology model in which the problem has been identified. In an embodiment of the present application, the weight ratio sub-model may include a first sub-model in an unsupervised stage, and the weight ratio corresponding to each node may be determined by the first sub-model in the unsupervised stage; specifically, the weight ratio of each node in the first business topology model may be obtained in a preset manner (for example, the business volume weight ratio of the preset node may be preset based on expert experience), or, the weight ratio sub-model nested in the fault probability model may be trained in an unsupervised stage based on historical sample documents, and the historical sample documents include abnormal events and business volume of each node and abnormal events and business volume of each edge in the first business topology model where the problem has been determined (for example, abnormal events may be poor quality events, and the first business topology model may include poor quality events and business volume of nodes and edges for unsupervised training). By analyzing these historical data, the above-mentioned weight ratio model can learn and train from past events to more accurately predict possible future failures, so that the weight ratio of each node in the first business topology model can be obtained in the subsequent problem demarcation process. This training strategy based on historical data enables the model to not only cope with the current network conditions, but also predict and adapt to future changes in network behavior. In summary, the embodiment of the present application can train the first sub-model of the unsupervised stage in the weight ratio sub-model through preset weight ratios or historical event analysis, so that the model can effectively adapt to changes in network behavior and improve the ability to respond to new network configurations or unknown faults, significantly improving the accuracy and efficiency of the problem demarcation solution.
[0018] In one possible implementation, the weighted proportion sub-model also includes a second sub-model in a supervised phase; the method also includes: receiving user feedback results, the feedback results including the K second root cause nodes of the real problem determined by the user for the J first root cause nodes; using the first business topology model and the abnormal subgraph as sample data, and the K second root cause nodes as labels, to tune the second sub-model. In the prior art, the system or platform for problem demarcation generally lacks self-tuning capabilities, and cannot automatically adapt to changes in network behavior or learn from the demarcation process to improve future performance, which may cause the system to perform poorly when dealing with unknown faults or new network configurations; to address this problem, an embodiment of the present application provides a solution for tuning the second sub-model in the supervised phase of the weighted proportion sub-model, which can be trained in a supervised phase each time the problem demarcation is performed. Specifically, the user's feedback results can be received, and tuning can be performed based on the root cause nodes of the real problem in the user's feedback results (for example, the K second root cause nodes can be some or all of the root cause nodes in the J first root cause nodes; or the above-mentioned K second root cause nodes The model may include some or all of the J first root cause nodes, as well as other root cause nodes other than the above-mentioned J first root cause nodes) to improve the efficiency of network problem demarcation; specifically, when the user's feedback result shows that the demarcation result is correct, the model will strengthen the current learning path; if the feedback result is wrong, the model will be retrained based on the new second root cause node (as a label) and sample data (first business topology model and abnormal subgraph) provided by the user, so as to continuously optimize and adjust its parameters, and calculate the corresponding fault probability according to the fault probability submodel until the abnormal node with the highest fault probability is more consistent with the above-mentioned label information (such as the second root cause node of user feedback) to obtain a trained weight proportion submodel. In this way, the weight proportion submodel can continuously improve the accuracy and efficiency of problem demarcation not only in theory but also in practical applications. In summary, in the embodiment of the present application, different measures can be taken based on user feedback, and the second submodel of the supervised phase of the weight proportion submodel can be tuned and trained based on the user's feedback results.
[0019] In one possible implementation, determining the first root cause node in each abnormal subgraph based on the failure probability of each abnormal node includes: adding the failure probability of each abnormal node to the first attribute information of the corresponding abnormal node; and for each first abnormal subgraph, finding the abnormal node with the highest failure probability as the first root cause node. In this embodiment of the present application, the corresponding root cause node can be determined based on the failure probability in each abnormal subgraph. Since each abnormal subgraph only outputs the abnormal node with the highest failure probability as the root cause node, the possibility of incorrect demarcation results is reduced, and a clear problem object is locked in to support subsequent problem demarcation analysis and optimization.
[0020] In one possible implementation, the node that meets the preset abnormal condition is a node whose statistical value of the abnormal event is higher than the preset normal statistical value. In an embodiment of the present application, abnormal nodes are defined as nodes whose statistical values of abnormal events are significantly higher than normal levels (for example, showing obvious quality differences in poor quality events). Abnormal nodes can be screened out from the first business topology model by comparing the preset normal statistical values to confirm an abnormal subgraph, which can be a minimum connected graph; further, the abnormal subgraph is determined by finding abnormal nodes with obvious poor quality, and then including the nodes connected to these nodes in the subgraph (when searching for connected nodes, it is not considered whether these nodes are of good quality or poor quality) to form a complete business topology graph; through this method, the search scope of the problem can be effectively narrowed, the efficiency of problem demarcation can be improved, and resource overhead can be reduced, and the problem areas in the network business can be more accurately identified and analyzed, thereby promoting more effective problem solving.
[0021] In one possible implementation, the first business topology model, the second topology model and the abnormal subgraph are graph models. In an embodiment of the present application, the business topology model and the abnormal subgraph can be graph models. In an embodiment of the present application, a graph model is used to construct the first business topology model, the second topology model and the abnormal subgraph, which provides an efficient and intuitive method for business analysis and anomaly detection. First, the graph model can clearly describe complex business relationships and interactions through the structure of nodes and edges, making the structure and relationship of the business easy to understand and analyze; secondly, the graph model constructed by the attribute information of nodes and edges (for example, poor quality events and business volume in the attribute information of nodes and edges) can show the complete process of a business, and can effectively identify and locate abnormal behaviors or root cause nodes in the business process. In summary, in an embodiment of the present application, by using a graph model to construct a business topology and abnormal subgraph for the analysis of business processes and the definition of root cause nodes, the business topology model and the abnormal subgraph can be made more flexible and adaptable, thereby enhancing the applicability and effectiveness of the problem delimitation solution when dealing with more diverse and complex business scenarios.
[0022] In one possible implementation, the network entity is one or more of a network element, a terminal device, and a support device. In embodiments of the present application, the network entity can be diverse, such as one or more of a network element, a terminal device, and a support device. The supported service topology model can present more service scenarios to implement a problem-bounded solution. Furthermore, different types of devices can be flexibly adjusted and expanded based on network requirements to adapt to ever-changing needs and challenges.
[0023] In a second aspect, an embodiment of the present application provides a problem delimiting apparatus, which may include:
[0024] A data acquisition unit, configured to acquire service data, wherein the service data includes M network entities, N service paths, and service statistics, where M and N are both positive integers;
[0025] A first model generating unit is configured to generate a first service topology model based on the service data, where the first service topology model includes M nodes, N edges, first attribute information of the M nodes, and second attribute information of the N edges, wherein one of the M nodes corresponds to one of the M network entities, and one of the N edges corresponds to one of the N service paths;
[0026] a subgraph generation unit, configured to determine J abnormal subgraphs in the first service topology model, wherein each abnormal subgraph includes one or more abnormal nodes and abnormal edges linking the one or more abnormal nodes, the abnormal nodes being nodes in the first service topology model that meet preset abnormal conditions, and J being a positive integer;
[0027] The root cause determination unit determines J first root cause nodes based on the first attribute information corresponding to each node and the second attribute information corresponding to each edge in the first business topology model, as well as the J abnormal subgraphs, where one abnormal subgraph corresponds to one first root cause node.
[0028] In a possible implementation, the first model generating unit is specifically configured to:
[0029] Building a second service topology model based on the service data, wherein the second service topology model includes M nodes and N edges;
[0030] Determining first attribute information of each of the network entities and second attribute information of each of the service paths based on the service statistical information;
[0031] The first service topology model is generated based on the M nodes, the N edges, the first attribute information of the M nodes, and the second attribute information of the N edges.
[0032] In a possible implementation, the service statistical information includes abnormal events and service volume of each of the M network entities, and abnormal events and service volume of each of the N service paths.
[0033] In one possible implementation, the first attribute information of each of the M network entities includes the statistical value of the abnormal event and the statistical value of the business volume; the second attribute information of each of the N business paths includes the statistical value of the abnormal event and the statistical value of the business volume.
[0034] In a possible implementation, the root cause determination unit is specifically configured to:
[0035] Determine a weighted proportion of the traffic volume of each of the M nodes based on the statistical value of the abnormal events and the statistical value of the traffic volume in the first attribute information of each node in the first traffic topology model, and the statistical value of the abnormal events and the statistical value of the traffic volume in the second attribute information of each edge;
[0036] Determine the weight ratio of the traffic volume of each abnormal node in each abnormal subgraph based on the weight ratio of each node in the M nodes;
[0037] Determine the failure probability of each abnormal node in each abnormal subgraph based on the weighted proportion of the business volume of each abnormal node, the statistical value of the abnormal events in the first attribute information, and the statistical value of the business volume;
[0038] Based on the failure probability of each abnormal node, a first root cause node in each abnormal subgraph is determined.
[0039] In one possible implementation, the method is applied to an intelligent delimitation platform, wherein the intelligent delimitation platform includes a fault probability model;
[0040] The root cause determination unit is specifically configured to:
[0041] Inputting the first attribute information of each node and the second attribute information of each edge in the first service topology model into the fault probability model, and outputting the fault probability of each abnormal node in each abnormal subgraph;
[0042] Based on the failure probability of each abnormal node, a first root cause node in each abnormal subgraph is determined.
[0043] In a possible implementation, the fault probability model includes a weight proportion sub-model and a fault probability sub-model; and the root cause determination unit is specifically configured to:
[0044] Input the statistical value of the abnormal event and the statistical value of the business volume in the first attribute information of each node, and the statistical value of the abnormal event and the statistical value of the business volume in the second attribute information of each edge into the weight proportion sub-model, and output the weight proportion of the business volume of each node in the M nodes;
[0045] The weighted proportion of the business volume of each abnormal node, the statistical value of the abnormal event in the first attribute information and the statistical value of the business volume are input into the fault probability sub-model, and the failure probability of each abnormal node in each abnormal sub-graph is output.
[0046] In a possible implementation, the weight proportion sub-model includes a first sub-model in an unsupervised stage;
[0047] In the first sub-model in the unsupervised stage, the weight ratio of the business volume of each of the M nodes output is preset based on experience, or, in the first sub-model in the unsupervised stage, the weight ratio of the business volume of each of the M nodes output is obtained by training the first sub-model based on historical sample documents, wherein the historical sample documents include abnormal events and business volume of each node and abnormal events and business volume of each edge in the first business topology model in which the problem has been identified.
[0048] In a possible implementation, the weight proportion sub-model further includes a second sub-model in a supervised phase; and the apparatus further includes:
[0049] a feedback receiving unit, receiving a user's feedback result, wherein the feedback result includes K second root cause nodes of a real problem determined by the user for the J first root cause nodes;
[0050] A model tuning unit is used to use the first business topology model and the abnormal subgraph as sample data, and the K second root cause nodes as labels to tune the second submodel.
[0051] In a possible implementation, the root cause determination unit is specifically configured to:
[0052] Adding the failure probability of each abnormal node to the first attribute information of the corresponding abnormal node;
[0053] For each of the first abnormal subgraphs, an abnormal node with the highest failure probability is found as the first root cause node.
[0054] In a possible implementation, the node that meets the preset abnormal condition is a node whose statistical value of the abnormal event is higher than a preset normal statistical value.
[0055] In one possible implementation, the first topology model, the second topology model, and the abnormal subgraph are graph models.
[0056] In a possible implementation, the network entity is one or more of a network element, a terminal device, and a support device.
[0057] In a third aspect, an embodiment of the present application provides a computer storage medium for storing computer software instructions used in a problem delimiting device provided in the second aspect, which includes a program designed for executing the above aspect.
[0058] In a fourth aspect, an embodiment of the present application provides a computer program, which includes instructions. When the computer program is executed by a computer, the computer can execute the process executed in the problem delimiting device in the second aspect above.
[0059] In a fifth aspect, embodiments of the present invention provide an intelligent delimiting platform configured to support the corresponding functions of the problem delimiting method provided in the first aspect. The intelligent delimiting platform may further include a memory coupled to the processor and storing program instructions and data necessary for the intelligent delimiting platform. The intelligent delimiting platform may further include a communication interface for communicating between the problem delimiting device and other devices or a communication network.
[0060] In a sixth aspect, the present application provides a chip system comprising a processor for supporting a problem delimiting device in implementing the functions described in the first aspect, such as generating or processing information used in the problem delimiting method. In one possible design, the chip system further comprises a memory for storing program instructions and data necessary for the data sending device. The chip system may consist of a chip or may include a chip and other discrete components. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background technology, the drawings required for use in the embodiments of the present application or the background technology will be described below.
[0062] FIG1A is a schematic diagram of a system architecture provided in an embodiment of the present application.
[0063] FIG1B is a schematic diagram of another system architecture provided in an embodiment of the present application.
[0064] FIG2 is a module diagram of an intelligent delimitation platform provided in an embodiment of the present application.
[0065] FIG3 is a flow chart of a method for problem demarcation provided in an embodiment of the present application.
[0066] FIG4A is a schematic diagram of a service topology model provided in an embodiment of the present application.
[0067] FIG4B is a schematic diagram of another service topology model provided in an embodiment of the present application.
[0068] FIG4C is a schematic diagram of another service topology model provided in an embodiment of the present application.
[0069] FIG4D is a schematic diagram of an abnormal subgraph provided in an embodiment of the present application.
[0070] FIG4E is a schematic diagram of determining the fault probability in an abnormal subgraph provided by an embodiment of the present application.
[0071] FIG4F is a schematic diagram of determining a root cause node provided in an embodiment of the present application.
[0072] FIG4G is a schematic diagram of another method for determining a root cause node according to an embodiment of the present application.
[0073] FIG5 is a schematic structural diagram of a problem delimiting device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0074] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.
[0075] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The terms "first," "second," etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0076] The terms "first," "second," "third," and "fourth," etc. in the specification and claims of this application and the accompanying drawings are used to distinguish different objects, not to describe a specific order. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. A process, apparatus, system, product, or device that illustratively includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to the process, apparatus, product, or device.
[0077] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0078] First, some terms in this application are explained to facilitate understanding by those skilled in the art.
[0079] (1) Web Entities: In the fields of computer science and network technology, web entities refer to any object or concept on the Internet that has a unique identity and attributes. These entities can include user terminals (such as personal computers and mobile devices), wireless cells (such as the coverage area of a mobile communication network), sites (such as network nodes or access points), core network elements (equipment and technology at the center of the network), and service providers (SPs). The application areas of web entities include network management, performance optimization, fault diagnosis, and network security. Through in-depth identification and analysis of these key web entities, network performance can be improved, resource allocation can be optimized, and the overall efficiency and security of the network system can be enhanced.
[0080] (2) Quality Deterioration Events: Quality Deterioration Events refer to various events that occur in telecommunications networks and affect service quality. These events can be caused by various factors, such as network congestion, equipment failure, and signal interference, leading to communication service interruptions, decreased data transmission speeds, and reduced call quality. Monitoring and analyzing quality deterioration events is a key part of telecommunications network management. The goal is to quickly identify and resolve these issues to maintain high network service standards and user satisfaction.
[0081] (3) Network Management / Probes: Network management refers to the maintenance, monitoring, and administration of telecommunications or computer networks, involving multiple aspects such as ensuring reliable network operation, configuring network equipment, and monitoring network performance and security. Probes are monitoring devices deployed in the network to collect network performance data, monitor traffic, and identify potential network problems. Network management combined with the use of probes can effectively monitor network status, promptly identify and resolve problems, and ensure stable network operation and service quality.
[0082] (4) Network Issue Delimitation: Network issue delimitation refers to the process of identifying and locating the specific areas or elements that cause performance problems or failures in complex telecommunications or computer networks. This process is crucial because it directly affects the speed and efficiency of problem resolution. Network problem delimitation typically involves a series of systematic steps, including initial problem identification, data collection, analysis, and precise fault location. This process may utilize various tools and techniques, such as traffic analysis, log file review, and fault tree analysis.
[0083] (5) Fault Tree: A fault tree is a graphical tool used to analyze and identify potential causes of failures in a system. It is primarily used in engineering and safety engineering to help engineers systematically identify the various possible paths and conditions that lead to a specific failure event. A fault tree uses logical symbols to represent a failure event and its possible causes, forming a tree-like structure with the failure event at the top and the causes leading to the event at the branches.
[0084] (6) Runtime Rule Engine: A runtime rule engine is a technical mechanism that processes and executes predefined rules in real time within a software system or telecommunications network. The primary function of this engine is to dynamically make decisions or trigger specific actions based on a set of rules while the system is running. These rules are typically defined based on business logic, network policies, or performance parameters and can handle a variety of runtime scenarios. The application of runtime rule engines is particularly important in telecommunications networks and IT systems. By executing predefined demarcation rules, they can provide demarcation analysis results when the problem is demarcated.
[0085] First, we analyze and propose the technical problems that this application aims to solve. In the prior art, the technology for problem demarcation (taking network problem demarcation as an example) includes the following solutions:
[0086] Solution: Problem demarcation is performed based on the rule orchestration capability in the design state and the rule execution engine in the runtime state. Specifically, the following steps 1 and 2 can be included.
[0087] Step 1: Manually configure expert experience to form fault tree rules.
[0088] Step 2: The fault tree rules are injected into the runtime engine to parse the execution and give the delimited results.
[0089] Disadvantage 1: Manual configuration of fault tree rules is scenario-dependent and relies on expert experience, resulting in high delivery costs and low efficiency. Specifically, in existing autonomous intelligent networks (ANs), fault tree rule configuration can be performed manually by experts based on their experience. This process involves in-depth analysis of network data, identification of fault modes, and formulation of fault handling strategies. Experts need to continuously analyze emerging network fault cases and update the fault tree database to ensure it can address the latest network issues. However, the complexity and dynamic nature of network environments require fault tree rules to be constantly updated to adapt to new situations, a process that is both time-consuming and costly. Furthermore, with the rapid development of network technologies and the diversification of service types, network environments are becoming increasingly complex, further increasing the difficulty of expert fault tree rule configuration and placing extremely high demands on human resources. Because this configuration method relies heavily on the individual experience of experts, it limits the universality and reusability of the rules. For example, different network environments may require different fault tree rules, and these rules may not be directly transferable to new environments. Due to the uncertainty of the rules, continuous expert input and optimization may be required. Furthermore, whenever the network topology, configuration, or services change, a large amount of manual intervention may be required to reconfigure the fault tree rules, significantly increasing labor and time costs. In addition, the manual configuration process may inevitably lead to human errors, which may cause inaccurate fault tree rules, thus affecting the efficiency and accuracy of fault demarcation.
[0090] Disadvantage 2: Fault tree rules are rigid and lack flexibility, making them difficult to adapt to the rapidly changing network environment. Fixed fault tree rules are often designed based on static network environments and known failure modes. They perform well for simple services (for example, when network status and service models do not change significantly). However, modern networks are characterized by highly dynamic and evolving service requirements. These static rules struggle to capture real-time changes in the network environment and emerging failure modes. Furthermore, the rapid development of network technology has introduced new devices and protocols, as well as new services and applications. These changes increase network complexity, making fault tree rules less applicable over time. For example, with the introduction of virtualization and cloud services, traditional network problem demarcation rules may no longer be applicable because they cannot handle fault propagation and isolation in virtual environments. Furthermore, fixed rules lack adaptability and cannot automatically update to address changes in network policies and configurations. In large-scale networks, even small configuration changes can trigger complex chain reactions, requiring fault tree rules to be updated in real time to reflect these changes. In terms of automated fault recovery, rigid rules may hinder the implementation of more flexible and adaptive recovery strategies, as these strategies need to be able to dynamically adjust according to the current network status.
[0091] Disadvantage 3: Inability to automatically tune and adapt to network changes, impacting the accuracy and responsiveness of demarcation analysis. Existing problem demarcation platforms lack automatic tuning capabilities. This means that once fault tree rules are developed and put into use, they cannot automatically adapt to changes in network behavior or self-optimize to improve the efficiency and accuracy of fault handling. This can lead to significant shortcomings in rapidly changing network environments. For example, as network traffic grows and new applications are deployed, the original rules may not correctly identify faults or predict their potential impact. Furthermore, the lack of tuning capabilities means that the rules cannot automatically learn from their mistakes to reduce future errors. Another issue is that due to the lack of an automatic tuning mechanism, network operators must manually adjust rules when they discover that they are no longer applicable. This is not only time-consuming and error-prone, but can also lead to service interruptions and customer dissatisfaction. Without automatic tuning, even small network updates can be labor-intensive, as each change requires manual evaluation and reconfiguration of the relevant fault tree rules. In addition, the lack of automatic tuning capabilities also limits the portability of fault tree rules in different network environments. For example, when rules need to be migrated from one network service to another, the above fault tree rules usually need to be extensively modified to adapt to the new network service, which increases the complexity of rule management and limits the reusability of the rules.
[0092] In order to solve the problem that current network problem demarcation technology does not meet actual business needs and achieve the goal of accurately identifying network problem demarcation capabilities, taking into account the shortcomings of existing technologies, the technical problems actually to be solved by this application include one or more of the following three aspects:
[0093] 1. Reduce reliance on expert experience and lower delivery costs (Disadvantage 1). The current network problem demarcation process is highly dependent on the experience of network experts. Manual rule configuration can require significant time for data analysis, rule design, and testing, resulting in excessively high human resource and time costs. Therefore, a technical solution for automated network problem demarcation is needed to reduce reliance on expert experience and lower delivery costs.
[0094] 2. Increase the flexibility of fault tree rules (disadvantage 2). In the existing technology, fault tree rules do not have the ability to adapt to multiple different businesses. In order to improve the flexibility of fault tree rules, the design and implementation mechanism of the rules must be fundamentally reconsidered. The rules should no longer be statically coded, but should be dynamically generated and updated. This means that it is necessary to develop an intelligent system that can monitor the network status in real time and automatically analyze the impact and propagation path of faults. Such a system can dynamically generate and adjust fault tree rules to match the current state of the network and predict future state changes. Therefore, it is necessary to propose a technical solution that can increase the flexibility of fault tree rules, so that the rule system can be adaptive and dynamic.
[0095] 3. Introducing automatic tuning capabilities (disadvantage 3). Existing problem demarcation platforms lack self-tuning capabilities and are unable to automatically adapt to changes in network behavior or learn from the demarcation process to improve future performance. This may result in the system performing poorly when dealing with unknown faults or new network configurations. The problem demarcation platform needs to be able to learn from past faults and network behavior, predict possible problems, and automatically adjust rules to prevent faults before problems occur. It should be able to respond not only to known fault modes, but also to new or unknown fault modes. Therefore, an intelligent demarcation platform with self-tuning capabilities is needed that can update its fault tree rules and processing strategies through training to improve the efficiency of network problem demarcation.
[0096] In summary, the existing network problem demarcation technology cannot meet the actual business needs, seriously affecting its demarcation analysis function. Therefore, the network problem demarcation technology provided in this application is used to solve some or all of the above technical problems.
[0097] Figure 1A is a schematic diagram of a system architecture provided by an embodiment of the present application. The system may include: a network management / probe 10, an intelligent delimitation platform 20, a network management operation platform 30, and a network optimization operation platform 40.
[0098] The network management / probe (Set Top Box, STB) 10 is used to collect business log data. It can be a combination of one or more network management and monitoring systems based on different technologies, such as the Simple Network Management Protocol (SNMP), Remote Monitoring (RMON) or customized probe technology. The network management / probe 10 can be deployed in the core network to monitor key indicators such as network traffic, device status, and service quality. The network management / probe 10 can also be applied to a variety of scenarios, such as network performance monitoring, fault diagnosis, security analysis, traffic management, etc., and collect operating data of network devices such as bandwidth utilization, error rate, device operating time, etc. through protocols such as SNMP. Further, in an embodiment of the present application, please refer to Figure 1B, which is another system architecture diagram provided in an embodiment of the present application, including the interaction relationship between various components. For example, business log data can be obtained based on the network management / probe 10. The business data log mainly includes two types of data, one is the experience quality poor event document, and the other is the business volume statistics of the network node. By analyzing the data collected from the network management / probe 10, the network operation and maintenance team or the intelligent demarcation platform 20 can promptly discover and resolve network bottlenecks to ensure the continuity and reliability of network services.
[0099] The intelligent demarcation platform 20 is a network fault location and analysis system designed to improve the efficiency and accuracy of network fault handling. It is one possible implementation of a problem demarcation device. For example, it can employ a range of technologies and algorithms to analyze network data and quickly locate the source of network problems. The platform can process large amounts of network traffic data and log information, leveraging data mining and machine learning techniques to identify fault patterns and abnormal behavior. The intelligent demarcation platform 20 typically integrates multiple functions, such as automatic fault detection, root cause analysis, network performance monitoring, and predictive maintenance; it may include multiple modules and components. For example, in an embodiment of the present application, the intelligent delimitation platform 20 can obtain service log data from the network management / probe 10. The service log data may include two categories, one of which is a poor experience quality event document containing the service topology, and the other is the service volume statistics of the network nodes. The intelligent delimitation platform 20 can extract the data required to build the service topology model from these two categories of data, for example, the network nodes through which the service passes, the link relationships between network nodes, and the poor quality events and service volume data on each network node and each link relationship between nodes. The intelligent delimitation platform 20 first builds a service topology model based on the above data, and then performs abnormal subgraph segmentation based on the service topology model, finds the node with the highest failure probability in each abnormal subgraph as the root cause node, and outputs the above root cause node as the delimitation result to the human-computer interaction page in the intelligent delimitation platform 20, supporting customers to confirm whether the delimitation result is accurate and input information about the network nodes that actually have problems. The model is retrained based on the feedback results as labels. By integrating network data analysis, machine learning, and human-computer interaction interfaces, the Intelligent Delimitation Platform 20 can not only quickly locate and diagnose network problems, but also help network operations teams optimize network configurations and policies, improving network stability and performance.
[0100] The network management operation platform 30 is a platform for problem handling and maintenance of the core network. It provides a comprehensive solution for monitoring and managing core network facilities, such as data centers, switches, routers and other key network equipment. The platform can monitor the health of the network in real time, quickly identify and diagnose problems, thereby ensuring the high availability and performance of the network. It includes modules such as fault management, configuration management, account management and security management, which support network administrators to effectively handle various network events and faults. For example, in an embodiment of the present application, the root cause node can be determined based on the intelligent delimitation platform 20, and then the problem type can be determined based on the root cause node. If the root cause node is caused by a core network problem, it can be uploaded to the network management operation platform 30 for closed-loop processing. In addition, the network management operation platform 30 can also be integrated with other systems to provide a more comprehensive network view and more efficient problem-solving strategies. It is an indispensable tool for operators in core network management.
[0101] The network optimization operation platform 40 is a platform for performance optimization and management of wireless networks. The network optimization operation platform 40 can integrate a variety of tools and technologies to monitor key indicators such as coverage, signal quality, and data transmission rate of the wireless network. By collecting and analyzing wireless network data, the platform can identify network coverage blind spots, signal interference sources, and other factors that affect network performance. The platform can also support network planning, capacity analysis, spectrum management and other functions to help operators optimize wireless network layout and resource allocation. For example, in an embodiment of the present application, the root cause node can be determined based on the intelligent delimitation platform 20, and then the problem type can be determined based on the root cause node. If the root cause node is caused by a wireless problem, it can be uploaded to the network optimization operation platform 40 for closed-loop processing. The goal of the network optimization operation platform 40 is to improve the stability and user experience of the wireless network and ensure the high quality and efficiency of wireless services. For operators, this platform is an important tool for optimizing wireless network performance and improving service quality.
[0102] Based on the above system architecture, an embodiment of the present application provides an intelligent delimiting platform 20 applied to the above system architecture. Please refer to Figure 2, which is a schematic diagram of the modules of the intelligent delimiting platform provided by the embodiment of the present application. The intelligent delimiting platform 20 may include a composition module 100, a delimiting module 200, and a training module 300. Among them:
[0103] Mapping module 100: Responsible for extracting the required topological data from service data and constructing a service topology model based on this topological data. Mapping module 100 may include a topology node extraction unit 101 and a service topology mapping unit 102. Topology node extraction unit 101 extracts topological nodes based on service data acquired by network management / probes. Service raw data primarily consists of two types: poor experience quality event records containing service topology information, and traffic statistics on network nodes. From these two types of raw data, the data required for service topology construction is extracted. Specifically, the data includes the network nodes that services traverse and the inter-node links, as well as poor experience quality events and traffic volume data on network nodes and inter-node links. After topology node extraction, service topology mapping unit 102 is executed to dynamically generate a visual representation of the network, including all nodes, connections, and their attributes. This unit can track network changes in real time, such as device additions and removals, or configuration changes, and update the topology map accordingly. Furthermore, the mapping module 100 integrates multiple data sources, including but not limited to routing information, traffic statistics, and device status, to ensure the accuracy and practicality of the topology map. This dynamic and comprehensive service topology model allows for the depiction of a complete service chain, providing the necessary foundation for the subsequent delimitation functions of the delimitation module 200.
[0104] Delimitation Module 200: Responsible for analyzing the network topology provided by Mapping Module 100, it locates the root cause of network issues based on fault probabilities without manually configuring a fault rule tree. Delimitation Module 200 may include an abnormal subgraph separation unit 201, a node fault probability update unit 202, a root cause object search unit based on fault probability 203, and a delimitation result unit 204. When a potential user experience issue is detected, Delimitation Module 200 can quickly locate the fault point, identify the impact range, and propose possible solutions. Specifically, Delimitation Module 200 first divides the abnormal subgraph into smaller subgraphs to gradually narrow the search scope. Node fault probability update unit 202 then updates the fault probabilities to the abnormal nodes in the abnormal subgraph. After the root cause node is found using the root cause object search unit based on fault probability 203, Delimitation Result Unit 204 outputs the root cause node. Furthermore, Delimitation Module 200 features a user-friendly interface, enabling network administrators to easily monitor network status, understand fault analysis results, and implement appropriate solutions.
[0105] Training module 300: By inputting the attribute information of nodes and edges into the fault probability model, the fault probability of each abnormal node in the abnormal subgraph can be determined, and the fault probability model can be trained to achieve a self-tuning function. The training module 300 may include a fault probability model periodic training unit 301, a fault probability model generation unit 302, and a fault probability self-optimization unit 303. For example, the fault probability model may include a nested weight ratio sub-model and a fault probability sub-model. For example, the fault probability model periodic training unit 301 includes the first sub-model of the unsupervised stage in the weight ratio sub-model. The first sub-model can be trained by presetting the weight ratio through expert experience or analyzing a large amount of historical data to output the weight ratio of each node; the fault probability model generation unit 302 may include a fault probability sub-model, which is obtained by inputting the weight ratio obtained by the above-mentioned fault probability model periodic training unit 301, as well as the attribute information of each abnormal node in the abnormal subgraph. , information outputs the failure probability of the abnormal node; the failure probability self-optimization unit 303 may include a second sub-model in the supervised stage in the weight proportion sub-model, and the second sub-model is trained by taking the first business topology model and the abnormal sub-graph as sample data and the real and effective second root cause node as a label, that is, the second sub-model can be tuned based on the user's feedback results to continuously optimize and update its recognition ability. The failure probability model jointly realizes the function of determining the failure probability through the nested weight proportion sub-model and the failure probability sub-model, ensuring that it can still maintain high efficiency and accuracy in the rapidly developing network technology and the ever-changing network usage mode, which helps to improve the overall performance and accuracy of the platform in the intelligent delimitation platform 20 in problem delimitation.
[0106] It is understandable that the structure of the intelligent delimiting platform 20 in FIG. 2 is only an exemplary implementation in the embodiment of the present application. The intelligent delimiting platform 20 in the embodiment of the present application includes but is not limited to the above structure.
[0107] Based on the system architecture diagram provided in Figures 1A-1B and the module diagram of the intelligent delimitation platform provided in Figure 2, combined with the problem delimitation method provided in this application, the technical problems raised in this application are specifically analyzed and solved.
[0108] Referring to Figure 3 , which is a schematic flow diagram of a problem delimiting method provided in an embodiment of the present application, this method can be applied to the system architecture described in Figures 1A and 1B . The intelligent delimiting platform 20 can be used to support and execute steps S300 through S303 of the method flow shown in Figure 3 . The following description will be provided from the perspective of the intelligent delimiting platform 20 in conjunction with Figure 3 . The method can include the following steps S300 through S303.
[0109] Step S300: Acquire business data.
[0110] Specifically, service data includes M network entities, N service paths, and service statistics, where M and N are both positive integers. For example, the intelligent delimitation platform 20 can collect service log data through network management / probes. There are two main types of service raw data: poor experience quality event records containing service topology, and traffic volume statistics of network nodes. From these two types of raw data, the data required for service topology construction is extracted. For example, this data may include the network nodes that the service passes through, the links between network nodes, poor experience quality events on network nodes and their links, and traffic volume data.
[0111] In one possible implementation, service data can be the core data of a service. The network entities included in the service data can be the various components that make up the network, such as one or more network elements, routers, switches, servers, terminal devices, support devices, and the connection points between them. Each of the M network entities is a node in network communication and can be a physical device, such as a hardware router, or a virtual node, such as a virtual switch in a software-defined network. A service path refers to the specific route along which data and communication services are transmitted in a telecommunications network. These paths determine the routing of data packets from source to destination. N service paths connect the M network entities and represent the path along which data flows in the network. Each service path may have different attributes, such as bandwidth, latency, and packet loss rate. These attributes may directly impact service performance and reliability. Optimizing and managing service paths can improve data transmission efficiency, reduce latency, and enhance user experience.
[0112] In one possible implementation, service statistics provide quantitative data on a network service, including but not limited to poor quality events, service volume, traffic statistics, service quality indicators, failure rates, etc. Based on these service statistics, detailed information on the service can be quickly understood, such as which paths may face congestion and which devices may have performance issues. For example, assume that a network service includes 50 network entities (M=50), which may include one or more network entities in various network elements, switches, routers, and firewalls. There are 100 service paths (N=100) in this network service, each path representing a data exchange route between different network entities. Service statistics can display the poor quality events and service volume of each of the 50 network entities and 100 service paths, indicating potential network problems. Through comprehensive analysis of these service data, the intelligent demarcation platform can have a comprehensive understanding of the network's operating status, so as to more accurately find the cause of the fault in the subsequent demarcation process.
[0113] Step S301: Generate a first service topology model based on service data.
[0114] Specifically, the first service topology model includes M nodes, N edges, first attribute information of the M nodes, and second attribute information of the N edges, wherein one of the M nodes corresponds to one of the M network entities, and one of the N edges corresponds to one of the N service paths. For example, please refer to Figure 4A, which is a schematic diagram of a service topology model provided in an embodiment of the present application. The first service topology model in the figure may include nodes 1-node 15, and edges linking these nodes, wherein the white normal nodes represent nodes that have not sent abnormal events, the black second abnormal nodes represent nodes that actually have problems, and the black and white first abnormal nodes represent affected nodes. Each node here corresponds to a network entity in the service data, such as a network element, server, router or other key network device, and each edge corresponds to a service path, indicating the communication path or connection relationship between these nodes. Furthermore, in the above-mentioned first business topology model, each node carries key first attribute information, such as one or more of the network element's configuration details, operating status, performance indicators, poor quality event statistics, and business volume statistics; similarly, each edge contains second attribute information, such as connection speed, reliability, the type of protocol used, one or more of the poor quality event statistics, and business volume statistics. Further, please refer to Figure 4B, which is a schematic diagram of another business topology model provided in an embodiment of the present application. The nodes in the business topology model in this figure may include user 1, local network devices, local switch devices, other local devices, the enterprise's main router, external network node 1, external network node 2, Server 1, and edges linking these nodes. In a complete link of a network service, for example, user 1 attempts to access Server 1. The network communication involved in this process can be represented by the nodes and edges in the first business topology model. User 1's request first reaches their local network device, which can be a wireless router or access point. The device appears as the first node in the business topology model. Subsequently, the user's request is transmitted through the network, passing through multiple nodes, each of which can represent a network element in the network; these nodes are connected by edges, which represent actual network connections, which can include wired connections or wireless connections. The user's request may first pass through a local network switch (the second node), then through the enterprise's main router (the third node), and then may pass through several external network nodes, and finally reach the network device of the data center hosting Server 1 (the last node). In the above example, this topological network service model representation method can clearly understand the complete path from the user to the server. Based on the attributes of the nodes and edges, a service link can be depicted, which helps the intelligent delimitation platform 20 better identify the cause of the fault in the subsequent delimitation process.
[0115] In one possible implementation, a second business topology model is constructed based on business data, and the second business topology model includes M nodes and N edges; based on business statistical information, the first attribute information of each network entity and the second attribute information of each business path are determined; based on M nodes, N edges, the first attribute information of M nodes and the second attribute information of N edges, the first business topology model is generated. Specifically, please refer to another business topology model schematic diagram provided in Figure 4C. The second business topology model in the figure may include nodes 1-node 15, and edges linking these nodes. In the second business topology model, a second business topology model is first generated based on the M network entities and N business paths obtained from the business data. The second business topology model may include corresponding M nodes and N edges, representing the link relationship in the network business structure; further, the first attribute information of each node (for example, the statistical value of the quality difference events and the statistical value of the business volume of each network entity) and the second attribute information of each edge (for example, the statistical value of the quality difference events and the statistical value of the business volume of each business path) can be calculated based on the business data, and then the first business topology model is generated based on the M nodes, N edges, and the corresponding first attribute information and second attribute information in the second business topology model. The attributes of the nodes / edges include the statistical value of the number of quality differences obtained based on the business data statistics, which can support the subsequent problem demarcation based on the business topology model.
[0116] Step S302: Determine J abnormal subgraphs in the first service topology model.
[0117] Specifically, each of the abnormal subgraphs includes one or more abnormal nodes and abnormal edges linking the one or more abnormal nodes, and the abnormal nodes are nodes that meet the preset abnormal conditions in the first business topology model, and J is a positive integer; exemplarily, firstly, based on the first attribute information of each node in the first business topology model, the statistical value of the poor quality events of each node (i.e., the statistical value of the abnormal events) is found, and compared with the preset abnormal conditions, the nodes of "obviously good quality" (for example, the nodes of obviously good quality can be nodes with a statistical value of poor quality events of 0) and "obviously poor quality" (for example, the nodes of obviously poor quality can be nodes with a statistical value of poor quality events of not less than 1) in the first business topology model are screened out, and then the nodes with obviously good quality are found. Nodes connected to nodes with poor quality and their corresponding edges are identified as abnormal edges to determine abnormal subgraphs. An abnormal subgraph is composed of several closely connected nodes connected by edges, forming a subnetwork. Each subnetwork is an abnormal subgraph. (In the graph representation of a telecommunications network, an abnormal subgraph refers to a network portion that exhibits abnormal behavior or performance indicators. These anomalies may indicate potential network problems or failures. By monitoring and analyzing abnormal subgraphs, potential network failures can be identified and prevented in advance, thereby improving network reliability and service quality.) Each abnormal subgraph represents a specific problem area in the network service. In complex network environments, multiple independent abnormal subgraphs may exist simultaneously. These subgraphs may be caused by different reasons. For example, one subgraph may be caused by a hardware failure, while another may be caused by a configuration error.Further, referring to FIG4A, for example, in a network service (i.e., the first service topology model), it may include normal nodes (i.e., nodes that meet the "obviously good quality" condition), first abnormal nodes, and second abnormal nodes (i.e., nodes that meet the "obviously good quality" condition, wherein the second abnormal node is a node that actually has a problem, and the first abnormal node is an affected node). Referring to FIG4D, FIG4D is a schematic diagram of an abnormal subgraph provided in an embodiment of the present application, which may include a first abnormal subgraph, a second abnormal subgraph, and a third abnormal subgraph, wherein each abnormal subgraph may include a second abnormal node that actually has a problem, a first abnormal node that is affected by the second abnormal node, and edges linking these abnormal nodes. In the first service topology model, the intelligent delimitation platform may identify three abnormal nodes based on the first service topology model. Each abnormal subgraph (i.e., J = 3) includes a second abnormal node with significantly poor quality and a first abnormal node connecting to it. The first abnormal subgraph may include several closely connected network elements or routers, which may be exhibiting abnormal behavior due to overload, resulting in a poor quality event that meets the significantly poor quality criteria. The second abnormal subgraph may consist of several remote network elements or servers, which frequently experience service interruptions due to improper security configuration, resulting in a poor quality event that meets the significantly poor quality criteria. The third abnormal subgraph may include three abnormal nodes, which may exhibit abnormal behavior due to hardware failure, configuration errors, or security attacks, resulting in a poor quality event that meets the significantly poor quality criteria. The edges connecting these nodes (i.e., abnormal edges) also exhibit abnormal behavior, for example, they may have very high data packet loss rates or abnormal traffic spikes. By identifying and analyzing these abnormal subgraphs, the intelligent delimitation platform can quickly narrow the analysis scope and find the minimum service topology range where the problem exists, thereby more accurately defining the root cause of the problem.
[0118] Step S303: Determine J first root cause nodes.
[0119] Specifically, based on the first attribute information corresponding to each node in the first business topology model and the second attribute information corresponding to each edge, as well as the J abnormal subgraphs, J first root cause nodes are determined, and one abnormal subgraph corresponds to one first root cause node. Exemplarily, the intelligent delimitation platform can determine a corresponding first root cause node in each abnormal subgraph based on the statistical value of the poor quality events and business volume of each node in the first business topology model (i.e., the first attribute information), as well as the statistical value of the poor quality events and business volume of each edge (i.e., the second attribute information). For example, in each abnormal subgraph, the node that is most likely to be the main cause of the network abnormality corresponds to a root cause network element (Root Cause Network Elements) in the business data. By accurately identifying the root cause network element, network operation and maintenance personnel can solve problems more effectively, reduce the impact of network failures, and improve recovery speed. The intelligent delimitation platform can use preset processing strategies to analyze data patterns and network behaviors in abnormal subgraphs to determine the main influencing factors of each abnormal subgraph, which may involve complex data association analysis;
[0120] Furthermore, on one hand, for example, the first root cause node in each abnormal subgraph can be determined through a preset processing strategy. The intelligent delimitation platform first builds a library containing multiple algorithms and rules based on the first attribute information of each node and the second attribute information of each edge. Based on the analysis of historical fault data, these preset processing strategies can effectively identify the primary fault cause node in each abnormal subgraph that causes network anomalies. On the other hand, for example, the first root cause node in each abnormal subgraph can be determined through data pattern analysis. The intelligent delimitation platform can configure an automated data collection system to collect key data from network entities (such as routers, switches, network elements, terminal devices, etc.) and network management systems, including traffic statistics, device logs, performance indicators, poor quality events, and business volume. Data mining algorithms such as cluster analysis and anomaly detection can be used to identify unusual patterns or trends in the data to determine the primary fault cause node in each abnormal subgraph that causes network anomalies. On another hand, for example, the first root cause node in each abnormal subgraph can be determined through data association analysis. Statistical analysis and machine learning models are used to analyze the correlation between different data sources and establish causal relationships. For example, time series analysis is used to correlate performance changes of different nodes and edges to determine the primary fault cause node that causes network anomalies. In the embodiment of the present application, step-by-step analysis can be performed based on different business data to narrow the scope, and finally the root cause node can be determined based on the attribute information and the abnormal subgraph. The delimitation of the experience problem can be performed without relying on the experience configuration rules of human experts. Therefore, the dependence on human experts in the problem delimitation process is greatly reduced, the delivery cost is reduced, and the efficiency of problem delimitation is improved.
[0121] In one possible implementation, based on the statistical values of abnormal events and business volume in the first attribute information of each node in the first business topology model, and the statistical values of abnormal events and business volume in the second attribute information of each edge, the weight ratio of the business volume of each node in the M nodes is determined; based on the weight ratio of each node in the M nodes, the weight ratio of the business volume of each abnormal node in each abnormal subgraph is determined; based on the weight ratio of the business volume of each abnormal node, and the statistical values of abnormal events and business volume in the first attribute information, the failure probability of each abnormal node in each abnormal subgraph is determined; based on the failure probability of each abnormal node, the first root cause node in each abnormal subgraph is determined. Exemplarily, in an embodiment of the present application, first, based on the first business topology graph model, A period of time (A is a number not less than 0) is accumulated. After A period of time, the first business topology graph model is generated based on the business data, and the business volume statistics and quality difference event statistics of the nodes therein, and the business volume statistics and quality difference event statistics on the edges are used as parameters to obtain the weight ratio of each node in the entire business (i.e., the first business topology model) (for example, it can be the weight ratio of the business volume of each node in the business, or it can be the weight ratio of the importance of the quality difference events and business volume of each node in the total quality difference events and business volume in the business). For example, it can be obtained by calculating the weight of the network element based on manual experience, or by learning based on the topological data of the historical problem and the determined root cause node as a historical sample document; after the first business topology model determines J abnormal subgraphs and the weight ratio of each node therein, the fault probability is determined based on the business volume statistics, quality difference event statistics and weight ratio of each abnormal node.
[0122] Specifically, the failure probability of an abnormal node = f(quality difference event statistics, business volume statistics, weight), see Figure 4E, Figure 4E is a schematic diagram of a method for determining the failure probability in an abnormal subgraph provided by an embodiment of the present application, the figure may include a first abnormal subgraph, a second abnormal subgraph and a third abnormal subgraph, wherein each abnormal subgraph may include a second abnormal node that is the main cause of the failure, a first abnormal node affected by the second abnormal node and edges linking these abnormal nodes, wherein each abnormal node in each abnormal subgraph in the figure has a corresponding failure probability, and the corresponding first root cause node can be determined based on the failure probability of each abnormal subgraph, further, see Figure 4F, Figure 4F is a schematic diagram of a method for determining the root cause node provided by an embodiment of the present application, the figure may include a first abnormal subgraph, a second abnormal subgraph and The third abnormal subgraph, wherein each abnormal subgraph may include a second abnormal node, which is the primary cause of the fault, a first abnormal node affected by the second abnormal node, edges linking these abnormal nodes, the fault probability of each abnormal node, and a root cause node determined based on the fault probability. In this graph, it is assumed that the rule for determining the root cause node is that the abnormal node with the highest fault probability is the root cause node. For example, in the first abnormal subgraph, there may be three abnormal nodes: abnormal node 9 has a fault probability of 70%, abnormal node 8 has a fault probability of 40%, and abnormal node 10 has a fault probability of 50%. The intelligent demarcation platform can find the abnormal node 9 with the highest fault probability based on the magnitude of the fault probability and define this abnormal node 9 as the root cause node. Other abnormal nodes with a lower fault probability than abnormal node 9 are defined as affected nodes. In addition, defining the root cause node based on the fault probability can also mean determining the abnormal node with the second highest fault probability as the root cause node, or determining the abnormal node with the lowest fault probability as the root cause node. The above is only one possible implementation method for defining the root cause node based on the fault probability and is not specifically limited in the embodiments of this application. In existing technologies, problem demarcation technology can often only identify the node with the problem, but cannot determine whether the node is the main cause or affected by the transitive relationship. If the real root cause node needs to be found, experts are required to make manual judgments. Due to manual misjudgment, the node found may not be the main cause of the problem, but an affected node, which affects the accuracy of problem demarcation and causes a large amount of resource waste.To address this technical problem, in an embodiment of the present application, the weight ratio of each node in the business is first determined through the intelligent delimitation platform, and then the failure probability of each abnormal node is determined in each abnormal subgraph based on the weight and attribute information. Finally, the root cause node can be determined based on the failure probability. For example, the root cause node can be the node with the highest failure probability. Therefore, determining the root cause node based on the failure probability can reduce the probability of misjudgment and greatly improve the accuracy of problem delimitation. In summary, the weight ratio of each node is analyzed based on the intelligent delimitation platform, and then the failure probability of each abnormal node is determined in each abnormal subgraph. The root cause node is determined based on the failure probability. No human intervention is required in the middle. The process of finding the root cause node can be independently completed by the intelligent delimitation platform, which reduces the possibility of human misjudgment, greatly improves the accuracy of problem delimitation, and at the same time, reduces delivery costs.
[0123] In one possible implementation, the fault probability training model is used to output the probability, which includes a nested weight proportion sub-model and a fault probability sub-model. The weight proportion sub-model is used to calculate the weight. The weight proportion sub-model includes a supervised stage and an unsupervised stage. The fault probability sub-model is used to calculate the fault probability. (1) Unsupervised stage of the weight proportion sub-model: In this stage, the weight proportion sub-model can determine the weight based on manual experience or historical sample documents. (2) Fault probability calculation of the fault probability sub-model: The fault probability sub-model can receive the weight calculated by the weight proportion sub-model and the attribute information of the node (for example, the statistical value of the node's poor quality events and the statistical value of the business volume) to calculate and generate the fault probability A. (3) Supervised stage of the weight proportion sub-model: In the supervised learning stage, the user can evaluate the accuracy of the fault probability A output by the fault probability sub-model. For example, if the fault probability A is considered inaccurate, the user will provide a fault probability B as a label; then, the weight proportion sub-model will use these user-provided fault probabilities B for training, so that the output weights are more accurate, so that the fault probability A generated by the final fault probability training model is closer to the fault probability B.
[0124] In one possible implementation, the method is applied to an intelligent delimitation platform, wherein the intelligent delimitation platform includes a fault probability model; determining J first root cause nodes based on the first attribute information corresponding to each node and the second attribute information corresponding to each edge in the first business topology model, as well as the J abnormal subgraphs, includes: inputting the first attribute information of each node and the second attribute information of each edge in the first business topology model into the fault probability model, outputting the fault probability of each abnormal node in each abnormal subgraph; and determining the first root cause node in each abnormal subgraph based on the failure probability of each abnormal node. Exemplarily, the first business topology model and the attribute information in the abnormal subgraph can be input into the fault probability model to determine the failure probability of each abnormal node in each abnormal subgraph (for example, the attribute information of the node and the edge can be input into the fault probability model to determine the weight ratio of the node, and then the corresponding failure probability is determined based on the attribute information and weight ratio of the node in the abnormal subgraph), and then the corresponding root cause node is found according to the failure probability of the node in each abnormal subgraph; the fault probability model can be a pre-trained classification machine learning algorithm model, and the failure probability determined based on the fault probability model can be used to quickly locate the fault in the subsequent problem delimitation, find the corresponding root cause node, and increase the accuracy of the delimitation result; in the embodiment of the present application, the attribute information of the node and the edge can be input into the fault probability model to determine the failure probability of each abnormal node in the abnormal subgraph and the corresponding root cause node, so as to improve the accuracy and efficiency of the delimitation result.
[0125] In a possible implementation, the fault probability model includes a weight proportion sub-model and a fault probability sub-model; the first attribute information of each node and the second attribute information of each edge in the first business topology model are input into the fault probability model, and the failure probability of each abnormal node in each abnormal sub-graph is output, including: the statistical value of the abnormal event in the first attribute information of each node and the statistical value of the business volume, the statistical value of the abnormal event in the second attribute information of each edge and the statistical value of the business volume are input into the weight proportion sub-model, and the weight proportion of the business volume of each node in the M nodes is output; the weight proportion of the business volume of each abnormal node, and the statistical value of the abnormal event in the first attribute information and the statistical value of the business volume are input into the fault probability sub-model, and the failure probability of each abnormal node in each abnormal sub-graph is output. In an embodiment of the present application, the fault probability model cooperates with the nested sub-models therein to jointly realize the function of determining the fault probability. Specifically, the fault probability model includes two nested sub-models, and the weight ratio of each node in the first business topology model can be determined by the weight ratio sub-model (for example, the weight ratio of the node can be determined by inputting the statistical values of the number of poor quality events and the business volume of the node and the edge into the weight ratio sub-model), and the failure probability of each abnormal node in each abnormal sub-graph can be determined in the fault probability sub-model through the weight ratio of each node above (for example, the corresponding failure probability can be determined by inputting the statistical values of the number of poor quality events and the business volume of the abnormal nodes in each abnormal sub-graph, as well as the weight ratio of the abnormal nodes into the fault probability sub-model); in an embodiment of the present application, the fault probability sub-model can be based on the weight ratio determined by the weight ratio sub-model to determine the failure probability of each abnormal node in the abnormal sub-graph. The function of the fault probability model is realized by the division of labor and cooperation of nested sub-models, which can significantly improve the accuracy of the intelligent delimitation platform in determining the fault probability, so as to facilitate more accurate determination of the root cause node and problem delimitation in the subsequent process.
[0126] In one possible implementation, the weight ratio sub-model includes a first sub-model in an unsupervised stage; in the first sub-model in the unsupervised stage, the weight ratio of the business volume of each of the M nodes output is preset based on experience, or, in the first sub-model in the unsupervised stage, the weight ratio of the business volume of each of the M nodes output is obtained by training the first sub-model based on historical sample documents, wherein the historical sample documents include abnormal events and business volume of each node and abnormal events and business volume of each edge in the first business topology model in which the problem has been identified. Specifically, the weight proportion sub-model may include the first sub-model of the unsupervised stage, and the weight proportion corresponding to each node may be determined by the first sub-model of the unsupervised stage; illustratively, the weight proportion of the business volume of each node in the M nodes may be preset based on experience (for example, the weight proportion is preset based on expert experience); or, the weight proportion sub-model nested in the fault probability model may be trained in an unsupervised stage based on historical sample documents, and the historical sample documents include abnormal events and business volume of each node and abnormal events and business volume of each edge in the first business topology model in which the problem has been identified (for example, abnormal events may be poor quality events, and the poor quality events and business volume of nodes and edges may be included in the first business topology model for unsupervised training). By analyzing these historical data, the model can learn from past events to more accurately predict possible future failures. This training strategy based on historical data enables the model to not only cope with current network conditions, but also predict and adapt to future changes in network behavior. Furthermore, the fault probability model can be obtained by training a classification machine learning algorithm model multiple times. The intelligent delimitation platform can collect a large amount of data related to network services during each experience problem definition process, such as the operating status and fault history in the service data, or the J abnormal subgraphs in the above-mentioned first service topology model as sample data input into the classification machine learning algorithm model; the classification machine learning algorithm model can include one or more models such as decision trees, random forests, K-nearest neighbors (KNN) and neural networks.
[0127] In a possible implementation, the weighted proportion sub-model also includes a second sub-model in a supervised phase; receiving user feedback results, the feedback results include the K second root cause nodes of real problems determined by the user for the J first root cause nodes; using the first business topology model and the abnormal subgraph as sample data, and the K second root cause nodes as labels, to tune the second sub-model. Exemplarily, an embodiment of the present application provides a solution for tuning the second sub-model in the supervised phase of the weighted proportion sub-model, which can be trained in a supervised phase each time the problem is delimited. Specifically, assuming that the initial classification machine learning algorithm model is a neural network model, when the intelligent delimitation platform executes the problem delimitation process to obtain the first root cause node, the intelligent delimitation platform can provide a human-computer interaction page to support the user to confirm whether the delimitation result is accurate, and to input the information of the second root cause node that actually has problems (for example, the K second root cause nodes can be some or all of the root cause nodes in the J first root cause nodes; or the above-mentioned K second root cause nodes can include some or all of the root cause nodes in the J first root cause nodes). point, and other root cause nodes other than the above-mentioned J first root cause nodes), the second sub-model is input into the neural network model as label information based on the feedback result (i.e., the second root cause node); and the data used in this delimitation process (for example, the first business topology model and the abnormal subgraph) is input into the neural network model as sample data, and the neural network model is trained with the label information as the training target to obtain the weight ratio of each node obtained based on the training of the neural network model, and the corresponding fault probability is calculated according to the fault probability sub-model until the abnormal node with the highest fault probability is more consistent with the above-mentioned label information (for example, the second root cause node of user feedback), so as to obtain a trained weight ratio sub-model. In an embodiment of the present application, a solution for tuning the fault probability model is provided. In each process of problem delimitation, in an embodiment of the present application, different measures can be taken based on user feedback, and the second sub-model in the supervised stage of the weight ratio sub-model can be trained based on the user feedback results, thereby achieving self-tuning.
[0128] In one possible implementation, the failure probability of each abnormal node is added to the first attribute information of the corresponding abnormal node; and for each first abnormal subgraph, the abnormal node with the highest failure probability is found as the first root cause node. Specifically, in this embodiment of the present application, the failure probability corresponding to each abnormal node can first be added to the first attribute information of the corresponding abnormal node, and then the corresponding root cause node can be determined by comparing the failure probabilities in each abnormal subgraph. For example, please refer to Figure 4G, which is another schematic diagram of determining the root cause node provided by an embodiment of the present application. The first business topology model in the figure may include nodes 1-node 15, and edges linking these nodes. Three abnormal subgraphs can be determined based on the abnormal nodes in nodes 1-node 15, which may include a first abnormal subgraph, a second abnormal subgraph, and a third abnormal subgraph. Each abnormal subgraph may include multiple abnormal nodes and abnormal edges linking abnormal nodes. The embodiment of the present application can add the failure probability corresponding to each abnormal node to the attributes of the corresponding abnormal node to identify the root cause node; further, taking the root cause node in the first abnormal subgraph as an example, the first abnormal subgraph may include nodes 8, nodes 9, and nodes 10. After adding the failure probability corresponding to each abnormal node in the first abnormal subgraph to the attribute information of the corresponding abnormal node, it can be known that the failure probability of node 8 is 40%, the failure probability of node 9 is 70%, and the failure probability of node 10 is 50%. By comparing the sizes of the failure probabilities, node 9 with a failure probability of 70% is determined to be the root cause node in the first abnormal subgraph. In summary, in the embodiment of the present application, the fault probability is first updated to the abnormal subgraph, and then the root cause node in each abnormal subgraph is determined based on the fault probability. Since each abnormal subgraph only outputs the abnormal node with the highest fault probability as the root cause node, the possibility of incorrect delimitation results is reduced, and a clear problem object is locked to support the analysis and optimization of subsequent problem delimitation.
[0129] In one possible implementation, after obtaining the root cause node, a specific analysis of the root cause node can be performed to determine the true cause of the fault, locate the main cause of the fault, and propose targeted solutions. Furthermore, for example, for core network problems, it can be uploaded to the network management operation platform for closed-loop processing. Assuming that the fault type of the root cause node is a core network problem (for example, the problem is one or more of hardware failure, software defects, configuration errors, and network congestion), for hardware failure, the network management operation platform can solve the core network problem by replacing the faulty network equipment or components; if the problem is caused by a software defect, the network management operation platform may need to update or patch the software of the network equipment to solve the core network problem. In summary, in the embodiment of the present application, the true cause of the fault can be determined based on the root cause node, and the cause of the fault can be classified, and the problem can be quickly and accurately identified and targeted solutions can be taken, rather than relying on guesswork and trial and error, thereby minimizing the uncertainty and risk of system maintenance and repair; it helps to improve the stability and security of the system, reduce maintenance costs, thereby improving user satisfaction, and optimizing resource utilization.
[0130] The above describes in detail the method of the embodiment of the present application, and the following provides the relevant device of the embodiment of the present application.
[0131] Please refer to Figure 5, which is a structural diagram of a problem delimiting device provided in an embodiment of the present application. The problem delimiting device 60 may include a data acquisition unit 601, a first model generation unit 602, a subgraph generation unit 603, a root cause determination unit 604, a feedback receiving unit 605, and a model tuning unit 606, wherein each unit is described in detail as follows.
[0132] A data acquisition unit 601 is configured to acquire service data, where the service data includes M network entities, N service paths, and service statistics, where M and N are both positive integers.
[0133] A first model generating unit 602 generates a first service topology model based on the service data, where the first service topology model includes M nodes, N edges, first attribute information of the M nodes, and second attribute information of the N edges, wherein one of the M nodes corresponds to one of the M network entities, and one of the N edges corresponds to one of the N service paths;
[0134] In a possible implementation, the first model generating unit 602 is specifically configured to:
[0135] Building a second service topology model based on the service data, wherein the second service topology model includes M nodes and N edges;
[0136] Determining first attribute information of each of the network entities and second attribute information of each of the service paths based on the service statistical information;
[0137] The first service topology model is generated based on the M nodes, the N edges, the first attribute information of the M nodes, and the second attribute information of the N edges.
[0138] In a possible implementation, the service statistical information includes abnormal events and service volume of each of the M network entities, and abnormal events and service volume of each of the N service paths.
[0139] In one possible implementation, the first attribute information of each of the M network entities includes the statistical value of the abnormal event and the statistical value of the business volume; the second attribute information of each of the N business paths includes the statistical value of the abnormal event and the statistical value of the business volume.
[0140] A subgraph generation unit 603 is configured to determine J abnormal subgraphs in the first service topology model, wherein each abnormal subgraph includes one or more abnormal nodes and abnormal edges connecting the one or more abnormal nodes, wherein the abnormal nodes are nodes in the first service topology model that meet a preset abnormal condition, and J is a positive integer.
[0141] The root cause determination unit 604 determines J first root cause nodes based on the first attribute information corresponding to each node and the second attribute information corresponding to each edge in the first business topology model, as well as the J abnormal subgraphs, where one abnormal subgraph corresponds to one first root cause node.
[0142] In a possible implementation, the root cause determination unit is specifically configured to:
[0143] Determine a weighted proportion of the traffic volume of each of the M nodes based on the statistical value of the abnormal events and the statistical value of the traffic volume in the first attribute information of each node in the first traffic topology model, and the statistical value of the abnormal events and the statistical value of the traffic volume in the second attribute information of each edge;
[0144] Determine the weight ratio of the traffic volume of each abnormal node in each abnormal subgraph based on the weight ratio of each node in the M nodes;
[0145] Determine the failure probability of each abnormal node in each abnormal subgraph based on the weighted proportion of the business volume of each abnormal node, the statistical value of the abnormal events in the first attribute information, and the statistical value of the business volume;
[0146] Based on the failure probability of each abnormal node, a first root cause node in each abnormal subgraph is determined.
[0147] In one possible implementation, the method is applied to an intelligent delimitation platform, wherein the intelligent delimitation platform includes a fault probability model;
[0148] The root cause determination unit 604 is specifically configured to:
[0149] Inputting the first attribute information of each node and the second attribute information of each edge in the first service topology model into the fault probability model, and outputting the fault probability of each abnormal node in each abnormal subgraph;
[0150] Based on the failure probability of each abnormal node, a first root cause node in each abnormal subgraph is determined.
[0151] In a possible implementation, the fault probability model includes a weight proportion sub-model and a fault probability sub-model; and the root cause determination unit is specifically configured to:
[0152] Input the statistical value of the abnormal event and the statistical value of the business volume in the first attribute information of each node, and the statistical value of the abnormal event and the statistical value of the business volume in the second attribute information of each edge into the weight proportion sub-model, and output the weight proportion of the business volume of each node in the M nodes;
[0153] The weighted proportion of the business volume of each abnormal node, the statistical value of the abnormal event in the first attribute information and the statistical value of the business volume are input into the fault probability sub-model, and the failure probability of each abnormal node in each abnormal sub-graph is output.
[0154] In a possible implementation, the weight proportion sub-model includes a first sub-model in an unsupervised stage;
[0155] In the first sub-model in the unsupervised stage, the weight ratio of the business volume of each of the M nodes output is preset based on experience, or, in the first sub-model in the unsupervised stage, the weight ratio of the business volume of each of the M nodes output is obtained by training the first sub-model based on historical sample documents, wherein the historical sample documents include abnormal events and business volume of each node and abnormal events and business volume of each edge in the first business topology model in which the problem has been identified.
[0156] In a possible implementation, the root cause determination unit 604 is specifically configured to:
[0157] Adding the failure probability of each abnormal node to the first attribute information of the corresponding abnormal node;
[0158] For each of the first abnormal subgraphs, an abnormal node with the highest failure probability is found as the first root cause node.
[0159] In a possible implementation, the node that meets the preset abnormal condition is a node whose statistical value of the abnormal event is higher than a preset normal statistical value.
[0160] In one possible implementation, the first topology model, the second topology model, and the abnormal subgraph are graph models.
[0161] In a possible implementation, the weight proportion sub-model further includes a second sub-model in a supervised stage;
[0162] The feedback receiving unit 605 receives a user's feedback result, where the feedback result includes K second root cause nodes of a real problem determined by the user for the J first root cause nodes.
[0163] The model tuning unit 606 is configured to use the first service topology model and the abnormal subgraph as sample data, and the K second root cause nodes as labels to tune the second sub-model.
[0164] In a possible implementation, the network entity is one or more of a network element, a terminal device, and a support device.
[0165] It should be noted that the functions of each functional unit in the problem delimiting device described in the embodiments of the present application can be found in the relevant descriptions of steps S300 to S303 in the method embodiment described in Figure 3 above, and will not be repeated here. In the above embodiments, the descriptions of each embodiment have their own emphasis. If a part is not detailed in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0166] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0167] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative, and the aforementioned unit division is merely a logical functional division. In actual implementation, other division methods may be used. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not implemented.
[0168] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0169] In addition, the functional units in the embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0170] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc., specifically a processor in a computer device) to execute all or part of the steps of the above-mentioned methods of each embodiment of the present application. Among them, the aforementioned storage medium may include: U disk, mobile hard disk, magnetic disk, optical disk, read-only memory (Read-Only Memory, abbreviated: ROM) or random access memory (Random Access Memory, abbreviated: RAM) and other media that can store program codes.
[0171] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for problem delimitation, characterized in that: The method comprises: Acquire service data, where the service data includes M network entities, N service paths, and service statistics information, where M and N are both positive integers; Based on the service data, a first service topology model is generated, wherein the first service topology model includes M nodes, N edges, first attribute information of the M nodes, and second attribute information of the N edges, wherein one of the M nodes corresponds to one of the M network entities, and one of the N edges corresponds to one of the N service paths; Determine J abnormal subgraphs in the first service topology model, wherein each of the abnormal subgraphs includes one or more abnormal nodes and abnormal edges linking the one or more abnormal nodes, the abnormal nodes are nodes in the first service topology model that meet preset abnormal conditions, and J is a positive integer; Based on the first attribute information corresponding to each node and the second attribute information corresponding to each edge in the first business topology model, and the J abnormal subgraphs, J first root cause nodes are determined, and one abnormal subgraph corresponds to one first root cause node.
2. The method according to claim 1, characterized in that The generating of the first service topology model comprises: Building a second service topology model based on the service data, wherein the second service topology model includes M nodes and N edges; Based on the service statistical information, determine first attribute information of each of the network entities and second attribute information of each of the service paths; The first service topology model is generated based on the M nodes, the N edges, the first attribute information of the M nodes, and the second attribute information of the N edges.
3. The method according to claim 2, characterized in that The service statistics information includes abnormal events and service volume of each of the M network entities, and abnormal events and service volume of each of the N service paths.
4. The method according to claim 3, characterized in that The first attribute information of each of the M network entities includes the statistical value of the abnormal event and the statistical value of the business volume; the second attribute information of each of the N business paths includes the statistical value of the abnormal event and the statistical value of the business volume.
5. The method according to claim 4, characterized in that The determining J first root cause nodes based on the first attribute information corresponding to each node and the second attribute information corresponding to each edge in the first service topology model, and the J abnormal subgraphs, includes: Determine a weight proportion of the business volume of each of the M nodes based on the statistical value of the abnormal events and the statistical value of the business volume in the first attribute information of each of the nodes in the first business topology model, and the statistical value of the abnormal events and the statistical value of the business volume in the second attribute information of each of the edges; Determine the failure probability of each abnormal node in each abnormal subgraph based on the weighted proportion of the business volume of each abnormal node, and the statistical value of the abnormal event in the first attribute information and the statistical value of the business volume; Based on the failure probability of each abnormal node, a first root cause node in each abnormal subgraph is determined.
6. The method according to claim 4, characterized in that Applied to an intelligent delimitation platform, the intelligent delimitation platform includes a fault probability model; the first attribute information corresponding to each node and the second attribute information corresponding to each edge in the first service topology model, and the J abnormal subgraphs, determining J first root cause nodes, including: Inputting the first attribute information of each node and the second attribute information of each edge in the first service topology model into the fault probability model, and outputting the fault probability of each abnormal node in each abnormal subgraph; Based on the failure probability of each abnormal node, a first root cause node in each abnormal subgraph is determined.
7. The method according to claim 6, characterized in that The fault probability model includes a weight proportion sub-model and a fault probability sub-model; the first attribute information of each node and the second attribute information of each edge in the first service topology model are input into the fault probability model, and the fault probability of each abnormal node in each abnormal sub-graph is output, including: Input the statistical value of the abnormal event and the statistical value of the business volume in the first attribute information of each node, and the statistical value of the abnormal event and the statistical value of the business volume in the second attribute information of each edge into the weight proportion sub-model, and output the weight proportion of the business volume of each node in the M nodes; The weighted proportion of the business volume of each abnormal node, the statistical value of the abnormal events in the first attribute information and the statistical value of the business volume are input into the fault probability sub-model, and the fault probability of each abnormal node in each abnormal sub-graph is output.
8. The method according to claim 7, characterized in that The weight proportion sub-model includes a first sub-model in the unsupervised stage; In the first sub-model in the unsupervised stage, the weighted proportion of the business volume of each of the M nodes output is preset based on experience, or, in the first sub-model in the unsupervised stage, the weighted proportion of the business volume of each of the M nodes output is obtained by training the first sub-model based on historical sample documents, wherein the historical sample documents include abnormal events and business volume of each node and abnormal events and business volume of each edge in the first business topology model in which the problem has been identified.
9. The method according to claim 8, characterized in that The weight proportion sub-model also includes a second sub-model in the supervised stage; the method also includes: Receive a user's feedback result, where the feedback result includes K second root cause nodes of real problems determined by the user for the J first root cause nodes, where K is an integer greater than 0; The first business topology model and the abnormal subgraph are used as sample data, and the K second root cause nodes are used as labels to tune the second submodel.
10. The method according to any one of claims 5 to 9, characterized in that: The determining, based on the failure probability of each abnormal node, the first root cause node in each abnormal subgraph comprises: Adding the failure probability of each abnormal node to the first attribute information of the corresponding abnormal node; For each of the first abnormal subgraphs, an abnormal node with the highest failure probability is found as the first root cause node.
11. The method according to any one of claims 4 to 10, characterized in that: The node satisfying the preset abnormal condition is a node whose statistical value of the abnormal event is higher than a preset normal statistical value.
12. The method according to any one of claims 2 to 11, characterized in that: The first service topology model, the second service topology model and the abnormal subgraph are graph models.
13. The method according to any one of claims 1 to 12, characterized in that The network entity is one or more of a network element, a terminal device, and a supporting device.
14. A device for problem delimitation, characterized in that: The device comprises: A data acquisition unit, used to acquire service data, wherein the service data includes M network entities, N service paths and service statistics information, where M and N are both positive integers; A first model generating unit generates a first service topology model based on the service data, wherein the first service topology model includes M nodes, N edges, first attribute information of the M nodes, and second attribute information of the N edges, wherein one of the M nodes corresponds to one of the M network entities, and one of the N edges corresponds to one of the N service paths; A subgraph generation unit is configured to determine J abnormal subgraphs in the first service topology model, wherein each of the abnormal subgraphs includes one or more abnormal nodes and abnormal edges linking the one or more abnormal nodes, the abnormal nodes are nodes in the first service topology model that meet preset abnormal conditions, and J is a positive integer; The root cause determination unit determines J first root cause nodes based on the first attribute information corresponding to each node and the second attribute information corresponding to each edge in the first business topology model, and the J abnormal subgraphs, where one abnormal subgraph corresponds to one first root cause node.
15. The device according to claim 14, characterized in that The first model generating unit is specifically used for: Building a second service topology model based on the service data, wherein the second service topology model includes M nodes and N edges; Based on the service statistical information, determine first attribute information of each of the network entities and second attribute information of each of the service paths; The first service topology model is generated based on the M nodes, the N edges, the first attribute information of the M nodes, and the second attribute information of the N edges.
16. The device according to claim 15, characterized in that The service statistics information includes abnormal events and service volume of each of the M network entities, and abnormal events and service volume of each of the N service paths.
17. The device according to claim 16, characterized in that The first attribute information of each of the M network entities includes the statistical value of the abnormal event and the statistical value of the business volume; the second attribute information of each of the N business paths includes the statistical value of the abnormal event and the statistical value of the business volume.
18. The device according to claim 17, characterized in that The root cause determination unit is specifically used for: Determine a weight proportion of the business volume of each of the M nodes based on the statistical value of the abnormal events and the statistical value of the business volume in the first attribute information of each of the nodes in the first business topology model, and the statistical value of the abnormal events and the statistical value of the business volume in the second attribute information of each of the edges; Determine the failure probability of each abnormal node in each abnormal subgraph based on the weight proportion of the traffic of each abnormal node, and the statistical value of the abnormal event in the first attribute information and the statistical value of the traffic; Based on the failure probability of each abnormal node, a first root cause node in each abnormal subgraph is determined.
19. The device according to claim 18, characterized in that Applied to an intelligent delimitation platform, the intelligent delimitation platform including a fault probability model; The root cause determination unit is specifically used for: Inputting the first attribute information of each node and the second attribute information of each edge in the first service topology model into the fault probability model, and outputting the fault probability of each abnormal node in each abnormal subgraph; Based on the failure probability of each abnormal node, a first root cause node in each abnormal subgraph is determined.
20. The device according to claim 19, characterized in that The fault probability model includes a weight proportion sub-model and a fault probability sub-model; the root cause determination unit is specifically used to: Input the statistical value of the abnormal event and the statistical value of the business volume in the first attribute information of each node, and the statistical value of the abnormal event and the statistical value of the business volume in the second attribute information of each edge into the weight proportion sub-model, and output the weight proportion of the business volume of each node in the M nodes; The weighted proportion of the business volume of each abnormal node, the statistical value of the abnormal events in the first attribute information and the statistical value of the business volume are input into the fault probability sub-model, and the fault probability of each abnormal node in each abnormal sub-graph is output.
21. The device according to claim 20, characterized in that The weight proportion sub-model includes a first sub-model in the unsupervised stage; In the first sub-model in the unsupervised stage, the weighted proportion of the business volume of each of the M nodes output is preset based on experience, or, in the first sub-model in the unsupervised stage, the weighted proportion of the business volume of each of the M nodes output is obtained by training the first sub-model based on historical sample documents, wherein the historical sample documents include abnormal events and business volume of each node and abnormal events and business volume of each edge in the first business topology model in which the problem has been identified.
22. The device according to claim 21, characterized in that The weight proportion sub-model also includes a second sub-model in the supervised stage; the device also includes: A feedback receiving unit receives a user's feedback result, wherein the feedback result includes K second root cause nodes of a real problem determined by the user for the J first root cause nodes, where K is an integer greater than 0; A model tuning unit is used to use the first business topology model and the abnormal subgraph as sample data, and the K second root cause nodes as labels to tune the second submodel.
23. The device according to any one of claims 18 to 22, characterized in that The root cause determination unit is specifically used for: Adding the failure probability of each abnormal node to the first attribute information of the corresponding abnormal node; For each of the first abnormal subgraphs, an abnormal node with the highest failure probability is found as the first root cause node.
24. The device according to any one of claims 17 to 23, characterized in that The node satisfying the preset abnormal condition is a node whose statistical value of the abnormal event is higher than a preset normal statistical value.
25. The device according to any one of claims 15 to 24, characterized in that The first service topology model, the second service topology model and the abnormal subgraph are graph models.
26. The device according to any one of claims 14 to 25, characterized in that The network entity is one or more of a network element, a terminal device, and a supporting device.
27. A computer storage medium, characterized in that The computer storage medium stores a computer program, which implements the method of any one of claims 1 to 13 when executed by a processor.
28. A computer program, characterized in that The computer program comprises instructions, and when the computer program is executed by a computer, the computer is caused to perform the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Information analysis method and device and electronic equipment
CN114844768A
Business exception root cause determination method, device and system
CN115348146A
Fault positioning method and device based on multi-index root cause positioning algorithm
CN116418653A
Fault root cause determination method and device, and storage medium and electronic device
WO2023029654A1
Cited By
Abnormal event detection method and system based on multi-source operation and maintenance data fusion
CN120803804A
Power distribution network topology correction method fusing topology anomaly recognition and structure credible reasoning
CN120833077A