Network operation and fault elimination method and system and related device
Patent Information
- Application Number
- CN202311871378.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2043-12-29
AI Technical Summary
面对可能发生的大量排障任务,如此运维方式,需要依赖更多的专业运维人员,这对于运营主体而言显然成本较高,影响用户体验,因此有必要提供有效的解决方案
[0046] Fundamentally, the network health status of a dimensional object can be obtained step by step based on the aggregated service dataset, the service datasets that participate in the aggregation of the aggregated service dataset, the service data subsets of each service stage in the service dataset, and the service inspection items of each service sub-stage in the service data subset. Therefore, when the network health status of the network topology is abnormal, the cause of the abnormality can be traced step by step from coarse to fine by using the network health status of each dimensional object under any troubleshooting dimension. This improves the accuracy and efficiency of fault diagnosis and resolution during network operation and maintenance, and reduces operation and maintenance costs.
Smart Images

Figure CN117768299B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to network operation and maintenance troubleshooting methods, systems and related equipment. Background Technology
[0002] In this era of rapid development in information technology, users can connect to the network via physical or wireless access methods using terminal devices to successfully access web pages or conduct business activities. However, in reality, at any stage of the entire network lifecycle (from network access to network exit), a failure point may occur. This could be due to authentication errors or malfunctions in the access equipment in the area, leading to failure to access web pages or interruption of business activities.
[0003] The current solution involves operations and maintenance personnel systematically checking several nodes in the network topology until the fault is located and resolved. Given the potentially large number of troubleshooting tasks, this approach requires a significant number of specialized operations and maintenance personnel, which is obviously costly for the operating entity and negatively impacts user experience. Therefore, it is necessary to provide an effective solution. Summary of the Invention
[0004] This application provides network operation and maintenance troubleshooting methods, systems, and related equipment for efficiently identifying abnormal problems during network operation and maintenance.
[0005] The first aspect of this application provides a network operation and maintenance troubleshooting method, including:
[0006] Obtain the service dataset of at least one network service within the network topology; wherein, one network service includes multiple service phases of the network lifecycle, each service phase has a service data subset, and the service data subset is obtained by combining service inspection items of multiple service sub-phases;
[0007] According to multiple preset troubleshooting dimensions, the service dataset of the at least one network service is summarized to obtain multiple dimension objects under each troubleshooting dimension and the summarized service dataset of each dimension object.
[0008] For each dimension object under each troubleshooting dimension, the network health status of the dimension object is obtained based on the aggregated service dataset of the dimension object;
[0009] Based on the network health status of the dimension objects under each of the troubleshooting dimensions, the network health status of the network topology under each of the troubleshooting dimensions is obtained;
[0010] If the network health status of the network topology is abnormal, troubleshooting is performed on the network topology based on the network health status.
[0011] Optionally, troubleshooting the network topology based on its health status includes:
[0012] If the network health status of the network topology is abnormal, search for the dimension object with abnormal network health status under the target troubleshooting dimension; the target troubleshooting dimension can be any of the troubleshooting dimensions.
[0013] In the service dataset of the abnormal dimension object, find the abnormal service data subset corresponding to the abnormal service stage, and then find the abnormal service inspection item corresponding to the abnormal service sub-stage in the abnormal service data subset.
[0014] Optionally, after searching for the abnormal service inspection item corresponding to the abnormal service sub-stage in the abnormal service data subset, the method further includes:
[0015] If the network health status of the network topology is abnormal, network maintenance is performed on the network topology according to the abnormal solution pre-configured for the abnormal service inspection item.
[0016] Optionally, the network health status of the dimension object includes: the health status score of the dimension object, and / or the proportion of abnormal service inspection items associated with the dimension object.
[0017] Optionally, obtaining the network health status of the dimension object based on the aggregated service dataset of the dimension object includes:
[0018] Determine the network health status of the service dataset for each of the network services described.
[0019] For the dimension object, the network health status of the dimension object is determined based on the network health status of the service dataset associated with the aggregated service dataset of the dimension object.
[0020] Optionally, after obtaining the service dataset of at least one network service within the network topology, the method further includes:
[0021] The service dataset for a single network service is displayed in a tree structure.
[0022] Wherein, the root node of the tree structure is the identification information of the network service, the child nodes of the tree structure are the service stages within the network lifecycle, the child nodes of each service stage are the service sub-stages included in the service stage, and the service sub-stages include service inspection items.
[0023] Optionally, after obtaining the summary service dataset for each of the dimension objects, the method further includes:
[0024] The summary service dataset of the dimensional objects is displayed in the form of a tree structure;
[0025] The root node of the tree structure is the dimension object, and the child nodes of the dimension object are the service datasets belonging to the dimension object.
[0026] Optionally, the step of summarizing the service dataset of the at least one network service according to multiple preset troubleshooting dimensions to obtain multiple dimension objects under each troubleshooting dimension and a summarized service dataset of each dimension object includes:
[0027] Collect service datasets of the network service at least once within a preset time span, and obtain multi-dimensional objects based on the unique object identifiers carried by the service datasets;
[0028] The service datasets belonging to the same dimension object are aggregated to obtain the aggregated service dataset of the dimension object.
[0029] Optionally, after obtaining the network health status of the network topology under each of the troubleshooting dimensions based on the network health status of the dimension objects under each of the troubleshooting dimensions, the method further includes:
[0030] The network health status of the network topology in each of the troubleshooting dimensions is displayed respectively.
[0031] Optionally, after obtaining the service dataset of at least one network service within the network topology, the method further includes:
[0032] Regularly monitor each of the service inspection items to identify potential problems in the network services.
[0033] In practice, the method described in the first aspect of this application may be implemented using the content described in the second aspect of this application.
[0034] A second aspect of this application provides a network operation and maintenance troubleshooting system, including:
[0035] The acquisition unit is used to acquire the service dataset of at least one network service within the network topology, wherein one network service includes multiple service phases of the network lifecycle, each service phase has a service data subset, and the service data subset is obtained by combining service inspection items of multiple service sub-phases;
[0036] The processing unit is used to summarize the service dataset of the at least one network service according to multiple preset troubleshooting dimensions, so as to obtain multiple dimension objects under each troubleshooting dimension and the summarized service dataset of each dimension object.
[0037] The processing unit is further configured to obtain the network health status of the dimension object for each dimension object under each troubleshooting dimension, based on the aggregated service dataset of the dimension object;
[0038] The processing unit is further configured to obtain the network health status of the network topology under each of the troubleshooting dimensions based on the network health status of the dimension objects under each of the troubleshooting dimensions.
[0039] A third aspect of this application provides an electronic device, including:
[0040] Central processing unit, memory, and input / output interfaces;
[0041] The memory is either a short-term storage memory or a persistent storage memory;
[0042] The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method described in the first aspect of the embodiments of this application or any specific implementation thereof.
[0043] A fourth aspect of this application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect of this application or any specific implementation thereof.
[0044] The fifth aspect of this application provides a computer program product comprising instructions or a computer program, which, when run on a computer, causes the computer to perform the method described in the first aspect of this application or any specific implementation thereof.
[0045] As can be seen from the above technical solutions, the embodiments of this application have at least the following advantages:
[0046] Fundamentally, the network health status of a dimensional object can be obtained step by step based on the aggregated service dataset, the service datasets that participate in the aggregation of the aggregated service dataset, the service data subsets of each service stage in the service dataset, and the service inspection items of each service sub-stage in the service data subset. Therefore, when the network health status of the network topology is abnormal, the cause of the abnormality can be traced step by step from coarse to fine by using the network health status of each dimensional object under any troubleshooting dimension. This improves the accuracy and efficiency of fault diagnosis and resolution during network operation and maintenance, and reduces operation and maintenance costs. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0048] It should be noted that although the steps in the flowcharts (if any) involved in the embodiments are drawn sequentially according to the arrows, unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts involved in the embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0049] Figure 1 This is a schematic diagram of a system architecture for a network operation and maintenance troubleshooting method according to an embodiment of this application;
[0050] Figure 2 This is a flowchart illustrating a network operation and maintenance troubleshooting method according to an embodiment of this application;
[0051] Figure 3 This is a schematic diagram of the service chain structure in the network operation and maintenance troubleshooting method of this application embodiment;
[0052] Figure 4 This is a schematic diagram of the structure of the precise troubleshooting tree at each level in the network operation and maintenance troubleshooting method of this application embodiment;
[0053] Figure 5 This is a schematic diagram of the structure of the precise troubleshooting tree for a unit in the network operation and maintenance troubleshooting method of this application embodiment;
[0054] Figure 6 This is a schematic diagram of the precise troubleshooting tree structure for a single-dimensional object in the network operation and maintenance troubleshooting method of this application embodiment;
[0055] Figure 7 This is an interface diagram showing the number of users in the network operation and maintenance troubleshooting method of this application embodiment;
[0056] Figure 8 This is a UI troubleshooting guide page diagram in the network operation and maintenance troubleshooting method of this application embodiment;
[0057] Figure 9 This is a schematic diagram of the network operation and maintenance troubleshooting system according to an embodiment of this application;
[0058] Figure 10 This is another structural diagram of the network operation and maintenance troubleshooting system according to an embodiment of this application;
[0059] Figure 11 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0060] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0061] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0062] In the following description, expressions such as "one specific implementation" or "one specific example" describe a subset of all possible embodiments. However, it is understood that "one specific implementation" or "one specific example" can be the same or different subset of all possible embodiments and can be combined with each other without conflict. In the following description, the term "multiple" means at least two. When a certain value mentioned in this application reaches a threshold (if it exists), in some specific examples, it may include the former being greater than the latter. When "any" or "at least one" or similar expressions are mentioned, it specifically refers to any one of the listed examples or any combination of these examples.
[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0064] The method provided in this application embodiment can be applied to, for example, Figure 1In the system architecture environment shown, terminal devices can access the target network (such as an SXF-5G network) through at least one network device among the following: wireless access point (AP), network controller (NAC), network management center (NMC), switch (SW), and gateway (GW) to perform a network service. This network service can refer to a network activity such as web page access or online gaming during a network lifecycle. The terminal devices within a certain range and their connected network devices can be used to form the network topology. For the server (or central end), it can receive service data generated by network devices during a network service, such as login authentication results, radio frequency transmission power, or message response information. This data can be stored in a data storage system (located on a server or in the cloud). This data can be used to deploy service inspection items to evaluate the status of network services. Subsequently, the server can gradually obtain service datasets from the service inspection items, and statistically analyze the service datasets to obtain a summary service dataset for an object in any troubleshooting dimension. The existence of each summary service dataset in any troubleshooting dimension can be used to reflect the network health status of the aforementioned network topology in this troubleshooting dimension. When the network health status of the network topology is abnormal, the server can investigate the cause of the fault step by step (i.e., locate the problem) based on the network health status of the dimension object under any troubleshooting dimension. For example, if the fault is in a certain service inspection item under a certain service sub-stage, the server can then perform targeted troubleshooting on the related items of that service inspection item to ensure the normal operation of network services.
[0065] The aforementioned AP (full name: Access Point): In Chinese, it is called a wireless access point. It is an essential communication device in a basic wireless network, used to transmit and receive wireless signals to realize wireless communication functions.
[0066] NAC (Network Access Controller) is a diversified, high-performance professional network control and management platform independently developed by Sinrui Technology. The controller supports centralized, visual management of devices such as wireless access points and switches, integrating authentication services, internet access behavior management, internet access behavior auditing, internal network terminal status visualization analysis, and east-west traffic security systems, significantly improving your network operation and maintenance experience.
[0067] NMC (Network Management Center) is a centralized management and intelligent operation and maintenance center for network devices. It can be flexibly deployed on integrated hardware and software devices or enterprise private clouds. As the intelligent brain of network management, the platform supports centralized management of devices such as service gateways, wireless APs, and switches. It can realize rapid activation, policy configuration, fault location, security alarms, and intelligent operation and maintenance of all network devices. At the same time, it supports unified identity authentication and role authorization for all network users, reducing your network management difficulty in one stop.
[0068] SW (full name: Switch): In Chinese, it means switch. It is a device that forwards and exchanges data using a type of (optical) electrical signal. It can provide an (optical) electrical signal path for any two network nodes connected to the switch and is an indispensable device for network connection.
[0069] GW (Gateway): A gateway is a device or system that transmits data between different networks. A gateway connects two or more different networks and forwards data packets, enabling these networks to communicate with each other. Gateways are typically located at the edge of a network and can be either hardware or software, with routing and forwarding capabilities. They can also perform tasks such as protocol conversion, data format conversion, and the implementation of security policies to ensure the secure and efficient transmission of data.
[0070] Network lifecycle: refers to the entire process from a terminal accessing the network to exiting the network. It can be understood as the process of a terminal using the network, mainly divided into the following five service stages (each stage can be further subdivided into sub-stages):
[0071] (1) User Onboarding stage: The terminal device connects to the network through physical connection or wireless access. In this stage, the terminal needs to establish a physical or logical link with the network and obtain network access permissions.
[0072] (2) IP Address Acquisition Stage: After a user accesses the network, the terminal device needs to obtain a valid IP address to communicate on the network. This can be achieved through Dynamic Host Configuration Protocol (DHCP) or static IP allocation. (3) User Authentication Stage: After accessing the network, the terminal device needs to authenticate itself to confirm its legitimacy and authorized access rights. Common authentication methods include username and password, certificates, two-factor authentication, etc., to ensure that only legitimate users can use network resources.
[0073] (4) Network Access Phase: After successful authentication, the terminal device obtains authorized access to network resources. In this phase, the terminal can use network services, access the Internet, and perform data transmission and other operations.
[0074] (5) Network Logout Phase: When the terminal device no longer needs to access the network, or when the network service provider requests a disconnection, the terminal device will perform a network logout operation. This may involve steps such as clearing network configuration, releasing IP addresses, and disconnecting physical or logical connections.
[0075] Inspection items: Specific projects or tasks that detect and evaluate certain indicators that may affect the network's lifecycle. These inspection items aim to provide real-time monitoring and analysis of key indicators such as network performance, security, and availability, enabling timely identification of potential problems and corresponding optimization measures. Common inspection items include checking whether a terminal responds to authentication messages, server connectivity testing, and RF transmission power testing. Specifically, the authentication message response inspection item can have two results: normal or abnormal. It describes whether the server responds to the authentication message within a time limit; failure to respond is considered abnormal. Server connectivity testing can also have two results: normal or abnormal. It is used to periodically send packets to the service chain for testing; failure to respond for multiple consecutive rounds is considered abnormal. RF transmission power testing is used to detect the specific value of RF transmission power in dBm. If the current dBm test result is within the normal range, it is considered normal. Through continuous monitoring and management of these inspection items, the stability, performance, and security of the network can be improved, ensuring the network operates normally and meets user needs. Correspondingly, abnormal inspection items refer to inspection items whose results are not within the normal range.
[0076] It should be noted that the methods provided in the embodiments of this application can be implemented jointly by the terminal device and the server as described above, or they can be implemented entirely on the server side, or they can be implemented entirely on the terminal device side. The specific implementation can be determined according to the actual application scenario, and no restrictions are imposed here.
[0077] The method described in this application will be further explained in detail below.
[0078] Please see Figure 1 The first aspect of this application provides a specific embodiment of a network operation and maintenance troubleshooting method, which includes the following operation steps:
[0079] Step 21: Obtain the service dataset of at least one network service within the network topology.
[0080] A single network service comprises multiple service phases throughout the network lifecycle. Each service phase has a subset of service data, which is obtained by combining service inspection items from multiple service sub-phases.
[0081] like Figure 3 As shown, this can be understood as a hierarchical or progressive relationship of "inspection item - service sub-stage - sub-chain - service stage - service dataset (i.e., service chain)" (the latter is composed of the former); each business stage (i.e., service stage) in the network lifecycle can be further subdivided into multiple business sub-stages (i.e., service sub-stages). For example, the user access stage can include service sub-stages such as the association stage, the layer 2 authentication stage, and the terminal verification stage. Each service sub-stage can contain or be configured with inspection items; if an inspection item in a sub-stage is abnormal, then that sub-stage is also considered abnormal, and so on, until the state of the service chain can be considered abnormal.
[0082] The aforementioned service dataset (which can be called a service chain) can be viewed as a collection of data aggregated within a single network lifecycle. A service chain is a concept that visualizes the various business stages (i.e., service stages), related business sub-stages (i.e., service sub-stages), and inspection items of a terminal device within a single network lifecycle. Graphically, a service chain can present the entire process of a terminal device from network access to network exit (i.e., presenting the network lifecycle), identifying key services, traffic paths, or data flows involved in different business stages. In short, a service chain visualizes a terminal's single network lifecycle, including definitions of stages, sub-stages, and inspection items. From the perspective of data flow status or direction, a service chain is obtained by collecting various data statistics during user network usage and can be categorized as raw data. For users, a service chain can be visualized as a tree structure (which can be called a unit-based precise troubleshooting tree) or a chart displayed on the UI for viewing. Generally, the data from one service chain, displayed through the UI, can be considered the simplest troubleshooting tree.
[0083] Taking the pre-authentication traffic during the authentication phase as an example, the traffic path described above includes processes such as existing user verification, pop-up portal page, portal page display, password modification, and authentication submission (or sub-phases). Users can intuitively see how the data is processed throughout the entire authentication process and whether there are any anomalies. Key services: Each of the above sub-phases has corresponding service checks (which can be considered inspection items), such as checking whether the user's password is correct, whether the user exists, whether the access is denied by permission settings, and whether DNS traffic is allowed.
[0084] Correspondingly, the aforementioned service data subset (which can be called a service subchain) refers to a situation where a certain stage of the service chain can occur multiple times within a single network lifecycle. Specifically, a certain service stage in the service chain (such as the user authentication stage) may be executed repeatedly, with each execution representing an independent behavior or attempt within that stage. By defining and describing service subchains, we can more accurately reflect the repetitive behaviors (such as multiple login attempts) at a certain stage of the network lifecycle. This helps in better analyzing and understanding user behavior, performance issues, or security incidents, as well as conducting network optimization and troubleshooting.
[0085] Step 22: Summarize the service datasets to obtain the summary service dataset for each dimension object under the troubleshooting dimension.
[0086] Specifically, according to multiple preset troubleshooting dimensions, the service datasets of at least one network service are summarized to obtain multiple dimension objects under each troubleshooting dimension and the summarized service dataset of each dimension object. Taking a tree structure as an example, the tree structure of the summarized service dataset of a certain dimension object (which can be referred to as the troubleshooting tree) is obtained by summarizing the service chains of that dimension object (such as the region dimension or the user dimension) (modeling result).
[0087] For example, troubleshooting dimensions can include four dimensions: region, device, network, and user. Each dimension can be further subdivided into multiple dimension objects. For instance, given a network topology within a certain range, the region dimension can be subdivided into two dimension objects: Region 1 and Region 2. Each region has its own summary service dataset (which can be called a precise troubleshooting tree under a single dimension object). Other dimensions are similar. For example, the device dimension can be subdivided into network devices such as Device 1 (e.g., AP_1), Device 2 (e.g., GW_1), Device 2 (e.g., GW_2)... Device 10 (e.g., SW_n). The network dimension can be subdivided into topology network 1... topology network 5. The user dimension (specifically, it can be a unique identifier for the terminal device or a user identifier) can be subdivided into user 1... user 20. It should be noted that the number of dimension objects subdivided into each of the four dimensions can be the same or different, depending on the composition of the network topology (e.g., the delineation of regions or the number of selected users).
[0088] Step 23: Obtain the network health status of the dimension object.
[0089] Specifically, for each dimension object under each troubleshooting dimension, the network health status of the dimension object is obtained based on the aggregated service dataset of the dimension object.
[0090] For example, the network health status of a dimension object can be calculated by directly analyzing the aggregated service dataset of that dimension object, or indirectly by calculating the network health status of each service dataset that participates in assembling the aggregated service dataset. In this embodiment, the network health status can be presented in the form of specific numerical values (such as health scores) and / or levels, and can be used to reflect the network operation and maintenance status or whether a link in the network service is functioning normally.
[0091] Step 24: Based on the network health status of the dimension objects under each troubleshooting dimension, obtain the network health status of the network topology under each troubleshooting dimension.
[0092] Specifically, the network health status of a network topology under any troubleshooting dimension can be obtained by summarizing or centrally reflecting the network health status of objects in each dimension under that troubleshooting dimension. In other words, the overall service dataset (global level) obtained by summarizing the various summary service datasets under a certain troubleshooting dimension can be regarded as an expression of the overall network topology. The network health status of the network topology reflected under different troubleshooting dimensions may be the same or different. The specific differences may lie in the slight differences in scores and / or levels, but in general they are the same, because regardless of which troubleshooting dimension is used, it is an assessment of the status of the same batch of network topologies.
[0093] Step 25: If the network health status of the network topology is abnormal, troubleshoot the network topology based on the network health status.
[0094] The troubleshooting process can generate corresponding troubleshooting suggestions based on the network health status of the network topology (specifically, based on the network health status of objects in each dimension under a certain troubleshooting dimension).
[0095] In summary, fundamentally, the network health status of a dimensional object is obtained by progressively tracing the data from the aggregated service dataset, the service datasets that participate in the aggregation of the aggregated service dataset, the service data subsets of each service stage within the service dataset, and the service inspection items of each service sub-stage within the service data subset. Therefore, when the network health status of the network topology is abnormal, the cause of the anomaly can be traced step by step from coarse to fine by examining the network health status of each dimensional object under any troubleshooting dimension. This improves the accuracy and efficiency of fault diagnosis and resolution during network operation and maintenance, and reduces operation and maintenance costs.
[0096] Based on the examples above, some specific possible implementation examples will be provided below. In practical applications, the implementation content of these examples can be combined or implemented separately as needed according to the corresponding functional principles and application logic, depending on the actual scenario.
[0097] based on Figure 2In some specific examples, step 25 may include: finding dimension objects with abnormal network health status under the target troubleshooting dimension; the target troubleshooting dimension is any troubleshooting dimension; in the service dataset of the abnormal dimension objects, finding the abnormal service data subset corresponding to the abnormal service stage, and finding the abnormal service inspection item corresponding to the abnormal service sub-stage in the abnormal service data subset.
[0098] like Figure 4 As shown, (1) Unit-based precise troubleshooting tree (which can be classified as unit level): refers to the visualization result of a single service chain (i.e., a service dataset of a single network service). Based on this, one can intuitively understand the relationship between related functions, operations, or events in each service stage of network operation and maintenance, and can assist in troubleshooting faults in a certain network lifecycle. It should be noted that each terminal device can generate at least one service dataset of a network service per day (i.e., a terminal device can have multiple service chains per day). For example, the process from the terminal device accessing the network to exiting the network is considered to generate a service dataset. In fact, each service chain starts with the user access stage and ends with the network exit stage. The examples in this application are for effect illustration and may contain errors or omissions.
[0099] Based on the above description, in some specific examples, after step 21, the embodiments of this application may further include the following operation: displaying the service dataset of a single network service in the form of a tree structure; wherein, the root node of the tree structure is the identification information of the network service (such as the timing information of the first service phase), the child nodes of the tree structure are the service phases within the network lifecycle, and the child nodes of each service phase are the service sub-phases included in the service phase, and the service sub-phases include service inspection items. Generally, operating the root node can trigger the display of child node information. For example, the content of the service phase is displayed first in the UI interface, and the operation of the service phase can trigger the display of the content of the service sub-phases.
[0100] (2) Precise Troubleshooting Tree (can be categorized at the dimensional object level): This involves statistically analyzing the service chains occurring within a single dimensional object (such as a single region, single device, single network, or single user) to obtain a precise troubleshooting tree for that dimensional object (e.g., the summary service dataset for device 1). This tree assists in quickly locating the cause of a fault within the dimensional object (or dimensional unit). In other words, a precise troubleshooting tree is a collection of multiple service chains. For example, poor network access for wireless users in a certain area might be due to low AP RF power in that area. Specifically, the abnormal AP can be quickly identified through precise troubleshooting trees at the device, region, or network dimensions. Furthermore, as... Figure 4The "pop-up portal page" sub-stage shown can be shared by multiple service chains. Each sub-stage can have its own inspection items. Therefore, the inspection items and their reports in this sub-stage may contain information from multiple service chains. Each inspection item in the inspection item report can specifically point to at least one service chain.
[0101] based on Figure 2 In some specific examples, after the above-mentioned "finding the abnormal service inspection item corresponding to the abnormal service sub-stage in the abnormal service data subset", the embodiments of this application may further include the following operation: performing network maintenance on the network topology according to the abnormal solutions pre-configured for the abnormal service inspection items. For example, solutions for abnormal situations can be configured for each service inspection item through experience (such as an experience handling database), so that the precise troubleshooting tree at each level can not only assist in discovering problems, but also provide solutions, making network operation and maintenance simpler; for example, after discovering that the network health status of the network topology is abnormal, the abnormal service inspection item can be quickly found, and the corresponding abnormal solution can be taken to restore the network to normal as soon as possible and improve the user experience.
[0102] Based on the above description, in some specific examples, after step 22, the embodiments of this application may further include the following operation: the summary service dataset of the dimension object can be displayed in the form of a tree structure; wherein, the root node of the tree structure is the dimension object (such as device 1), and the child nodes of the dimension object are the service datasets belonging to the dimension object, such as the service datasets in the form of a tree structure under device 1. Compared with statistical charts such as lists, pie charts, or bar charts, tree structures are more hierarchical and concise, which helps to reduce data storage and troubleshooting workload.
[0103] (3) Generalized Precision Troubleshooting Tree, or Multi-Dimensional Generalized Precision Troubleshooting Tree (which can be classified as overall or at the troubleshooting dimension level): refers to the visualization result of the overall network topology. Based on this, one can intuitively understand the network health status reflected by the network topology under any troubleshooting dimension (such as region, device, network, or user). It can assist maintenance personnel in troubleshooting the cause of the fault step by step from coarse to fine. For example, when the network health status of the generalized precision troubleshooting tree (with the device dimension as the target troubleshooting dimension) is abnormal → troubleshooting down to find that the precision troubleshooting tree of device 1 (i.e., the abnormal dimension object) is the cause → troubleshooting down to find that the nth unit precision troubleshooting tree contained in device 1 has the cause (i.e., there is an abnormal service data subset), and then troubleshooting down to find the abnormal service sub-stage in the abnormal service data subset, until the abnormal service inspection item corresponding to the abnormal service sub-stage is traced. This abnormal service inspection item can be used to finally determine the cause of the fault that caused the generalized precision troubleshooting tree (i.e., the network topology) to be abnormal, and to take corresponding solutions.
[0104] In summary, as an example, today the operations and maintenance administrator logs into the system to check the overall status (i.e., network topology). If an anomaly is found in the network health status, such as a low score, the administrator can then investigate which dimensions (e.g., Region 1, Network 2, Network 3, Device 1 or Device 2) are causing the anomaly. Next, the administrator checks the precise troubleshooting tree for that region / network / device to determine which service stage is causing the low score. This investigation continues downwards to find the abnormal sub-stage (e.g., a high number of service interruptions and / or abnormal inspection items). For example, if the user authentication stage score for network SSID1 is found to be very low, with a particularly high number of abnormal inspection items in the "submit verification" sub-stage, further investigation of the inspection item reports for that sub-stage reveals a particularly high number of anomalies in the "user expiration" inspection item. This reasonably suggests that the configured authentication validity period has expired, and the administrator can be guided to extend the validity period (i.e., provide a solution) to overcome this anomaly.
[0105] As can be seen, the embodiments of this application can locate which sub-stage is abnormal by analyzing the network health status of the dimension object, from coarse to fine, thereby ultimately identifying which inspection items are causing the abnormality. This makes troubleshooting work traceable and reduces the workload and time required in traditional operation and maintenance troubleshooting. In some specific examples, when the number of objects (such as the number of users) contained in a certain troubleshooting dimension exceeds a certain number, it will be difficult and time-consuming to check the summary service datasets and their service data subsets under the user dimension. Therefore, other troubleshooting dimensions with a relatively smaller order of magnitude (such as the region dimension) can be used as the target troubleshooting dimension. For example, the summary datasets and their service data subsets of region 1 to region n under the region dimension can be checked step by step to find the abnormal inspection items.
[0106] based on Figure 2 In some specific examples, step 22 may include: collecting service datasets of at least one network service within a preset time span, and obtaining multiple dimension objects based on the unique object identifier carried by the service datasets; summarizing the service datasets belonging to the same dimension objects to obtain a summary service dataset of dimension objects.
[0107] The aforementioned preset time span can be a period of one day, 7 days (i.e., a week), or 30 days (i.e., a month). For example, data (i.e., service data) that occurs in each network lifecycle within the preset time span can be collected from various network devices (such as AP, SW, NMC, etc.) connected to the terminal device. This includes data on service phases, sub-chains, sub-phases, inspection items, etc. This data can carry a unique service chain identifier (such as circle_id) – used to mark which service chain it belongs to (service chains may have temporal differences). Through the service chain aggregation module, data belonging to the same circle_id can be combined to obtain a complete service chain. This service chain (i.e., service dataset) can be used to aggregate and obtain a higher-level accurate troubleshooting tree.
[0108] Each service chain carries a unique object identifier pointing to a dimension object, such as region ID, device ID, network ID, user terminal MAC address, etc. Based on the unique object identifiers carried in the service dataset, we can identify the dimension objects within this data, such as Device 1…Device 15, Region 1…Region 3, Topology 1…Topology 5, User 1…User 20. By aggregating service chains pointing to the same dimension object (e.g., all carrying mac_1), we can obtain a precise troubleshooting tree for that dimension object (i.e., aggregating the service dataset). Furthermore, the health status of the corresponding precise troubleshooting tree can be obtained based on the service chains involved in the aggregation.
[0109] For example, service data over a 7-day or 30-day time span can be collected and summarized to obtain various levels of troubleshooting trees, such as long-term precise troubleshooting trees and (multi-dimensional) general precise troubleshooting trees. These troubleshooting trees composed of data from historical time spans can be called historical data precise troubleshooting trees. Each level of troubleshooting tree can assist in analyzing recent network conditions, helping to independently discover problems and investigate historical issues. The collected data may include: whether the username exists, whether the login password is correct, whether the IP address has been obtained, the specific value of the radio frequency transmission power in dBm, whether user groups are denied access, multicast optimization rate in Mbps, etc. The service chain can be regarded as the most basic (raw) information, the basic unit of summarization (the summarized object), and various precise troubleshooting trees can be summarized based on the service chain. The service chain will not be further subdivided; it is simply composed of multiple stages, each stage is composed of sub-chains, each sub-chain is composed of sub-stages, and each sub-stage contains multiple inspection items. Sub-stages and inspection items can be further improved and expanded in the future.
[0110] based on Figure 2In some specific examples, the network health status of a dimension object includes: the dimension object's health status score, and / or, the percentage of abnormal service inspection items associated with the dimension object (e.g., number of abnormal inspection items: total number of inspection items), and other health indicators. In practical applications, the network health status can be the aforementioned health indicators themselves, or a comprehensive result calculated based on these indicators.
[0111] based on Figure 2 In some specific examples, step 23 may include: determining the network health status of the service dataset for each network service; and, for a dimension object, determining the network health status of the dimension object based on the network health status of the service dataset associated with the aggregated service dataset of the dimension object.
[0112] For example, service datasets (which can be viewed as unit-level precise troubleshooting trees), aggregated service datasets (which can be viewed as precise troubleshooting trees), and aggregated service datasets under a certain troubleshooting dimension (which can be viewed as generalized precise troubleshooting trees), etc., can each have a set of health status evaluation standards to assist users in quickly troubleshooting or locating anomalies layer by layer. Taking the health status score (which can be abbreviated as health score) as the network health status as an example, the health score is calculated level by level. For example, the level order can be inspection item - service sub-stage - sub-chain - service stage - service dataset (i.e., service chain), where the sub-stage score is calculated by combining the health scores of its various inspection items, the sub-chain score is calculated by the health scores of the sub-stage, the service stage score is calculated by the health scores of the sub-chain, and the service chain score is calculated by the health scores of the service stage. Furthermore, within a preset time span, the health score of a certain dimension object is calculated by the health scores of multiple service chains during the period (which can be specified under a certain dimension object, such as user 1 dimension). Similarly, the health scores of a single device, a single region, a single network, and other dimension objects are calculated by the health scores (i.e., status) of the service chains included during that period. Furthermore, the overall health score (i.e., the overall network topology) can be obtained or centrally reflected by summing the health scores of objects in each dimension under any troubleshooting dimension (such as all devices under the device dimension).
[0113] Specifically, (1) the unit-level precise troubleshooting tree (i.e., the service dataset, which can be classified as a unit level): contains or expresses the health status of service chains, service stages, service sub-chains, and sub-stages; among them, at least one dimension, such as health score or the proportion of abnormal inspection items (i.e., the number of abnormal inspection items / the total number of inspection items), can be used as the evaluation dimension for the unit-level precise troubleshooting tree and its subordinate elements. For example, the subordinate elements (or constituent elements) of a service sub-chain are sub-stages, and the subordinate elements of a sub-stage are inspection items. In some specific examples, such as Figure 5As shown, the collected data from each individual service chain can be visualized according to the time sequence, that is, the centralized display unit's precise troubleshooting tree, which makes it convenient for users to locate abnormal inspection items step by step according to the network health status.
[0114] The aforementioned time sequence refers to the order in which stages, sub-chains, and sub-stages occur. These can be displayed according to the actual order of terminal operations, allowing users to intuitively see the health status of each stage during network usage, facilitating troubleshooting and resolution. The aforementioned network health status (hereinafter referred to as "status") may include at least one of the following: service stage score, service stage status (whether interrupted or normal), sub-chain status (whether successful), number or percentage of abnormal inspection items in the service stage, and number or percentage of abnormal inspection items in the sub-stage. Alternatively, the status may be a comprehensive result of at least one of the aforementioned items. Because the stage status and score are calculated by the lower level (sub-chain), the problem can be identified level by level based on the sub-chain's status (whether it is normal, high or low), number of abnormal inspection items, etc., ultimately pinpointing which sub-stage is abnormal and identifying which inspection items are causing the abnormality. This allows for a gradual narrowing of the investigation from a broad scope to a narrower scope, efficiently identifying abnormal inspection items and reducing the workload and time spent investigating massive amounts of data during troubleshooting.
[0115] (2) Precise Troubleshooting Tree (i.e., a summary service dataset, which can be categorized as a dimension object level): This includes or expresses the health status of each service stage and sub-stage under a dimension unit (such as a single region, a single device, a single network, or a single user). At least one of the following dimensions can be used as the evaluation dimension for the precise troubleshooting tree: health score, number of service interruptions / total number of services, percentage of abnormal inspection items (i.e., number of abnormal inspection items / total number of inspection items), and number of abnormal users / total number of users. Service interruption: refers to the abnormal termination of a service chain, resulting in an interruption of the network lifecycle. For example, failing to successfully transition from one service stage to the next in a network service session reflects an abnormal situation in the previous service stage. Abnormal user: a terminal device or user whose network service status evaluation is below the normal threshold within a certain time span.
[0116] like Figure 6 As shown, a precise troubleshooting tree (i.e., a summary service dataset) can be displayed for each dimension object, allowing users to intuitively understand the network health status under that dimension object and quickly locate abnormal inspection items level by level. For example, an abnormal inspection item may be in the "submit authentication" sub-stage. Furthermore, the cause of the anomaly may be an anomaly in inspection items such as air interface packet loss rate and / or bit error rate. It should be noted that... Figure 6 The “SXF” and “SXF-5G” shown can be two access networks in the network topology, and users can complete network activities such as web page access through either of them.
[0117] (3) (Multi-dimensional) Generalized Precise Troubleshooting Tree (i.e., the overall network topology, which can be categorized as either the overall structure or a troubleshooting dimension): This includes or expresses the overall network topology, the health status of each dimension (region, device, network, user, etc.), and the health status of each dimension's objects. Similarly, with a precise troubleshooting tree, at least one dimension can be used as the evaluation dimension for the generalized precise troubleshooting tree: health score, number of service interruptions / total number of services, percentage of abnormal inspection items, and number of abnormal users / total number of users. In some examples, the number of users and their corresponding levels can be as follows: Figure 7 As shown, pie charts or bar charts can be used to reflect whether a user is abnormal. For example, users who are below a predetermined percentage or a preset star rating can be considered abnormal users.
[0118] As can be seen, by aggregating the service chains, we can obtain the precise troubleshooting tree and health status of objects in each dimension. By aggregating the precise troubleshooting tree of objects in each (all) dimension under a certain troubleshooting dimension, we can obtain the health status of that troubleshooting dimension to express the overall health status of the entire network topology environment. This allows operations and maintenance personnel to understand the health status of the network from the overall perspective or from the dimensions of interest, and quickly locate the affected areas, devices, networks, users, etc., providing an entry point for the next step of accurately locating the problem.
[0119] (4) Historical data accurate troubleshooting tree: such as summarizing service data within a 7-day or 30-day time span to obtain a long-term accurate troubleshooting tree or a (multi-dimensional) general accurate troubleshooting tree; the health status evaluation criteria of the historical data accurate troubleshooting tree can be similar to the evaluation dimensions used in the above-mentioned accurate troubleshooting tree and (multi-dimensional) general accurate troubleshooting tree.
[0120] In this embodiment, referencing and displaying historical data for precise troubleshooting is based on the consideration that operations and maintenance personnel often need to not only focus on the current network status and troubleshooting, but also evaluate the overall network condition in the recent period in order to independently identify problems and prevent them in advance. The advantages of aggregating operations and maintenance data from the past 7 or 30 days include a larger data sample and a longer time span, making it easier to highlight network problems. In some specific examples, to prevent performance issues caused by aggregating large amounts of data, this embodiment can also adopt measures such as aggregating during idle periods or temporarily caching the aggregated data, thereby providing customers with a better operations and maintenance experience.
[0121] Since the network health status of a network topology under any troubleshooting dimension can be obtained by summarizing or centrally reflecting the network health status of objects in each dimension under that troubleshooting dimension, therefore, based on Figure 2In some specific examples, after step 24, the embodiments of this application may also include the following operations: displaying the network health status of the network topology in each troubleshooting dimension, so that users or maintenance personnel can intuitively know which troubleshooting dimension is appropriate as the target troubleshooting dimension, so as to reduce troubleshooting time.
[0122] based on Figure 2 In order to effectively prevent network failures, in some specific examples, after step 21, the embodiments of this application may also include the following operation: periodically monitoring each service inspection item to identify potential problems in the network service.
[0123] Compared to Figure 2 The examples shown illustrate that the additional or detailed examples or possible implementation methods described above may not necessarily be executed in actual implementation. If more than two examples or possible implementation methods are added or detailed, these examples or possible implementation methods can be implemented in combination or individually, depending on the actual scenario.
[0124] In summary, the embodiments of this application aim to solve the technical problems of service interruption and efficiency degradation caused by anomalies during network operation and maintenance. Specifically, the following technical problems can be solved:
[0125] (1) Troubleshooting efficiency: Current network operation and maintenance methods are inefficient in troubleshooting, which affects the normal operation of services. The embodiments of this application construct a precise troubleshooting tree and combine it with data collection and analysis tools to quickly and accurately locate and resolve faults in the network operation and maintenance process, thereby improving the efficiency of troubleshooting.
[0126] (2) System optimization and preventive measures: In addition to troubleshooting, this application also focuses on system optimization and preventive measures. By analyzing and mining data during network operation and maintenance, potential problems and risks are identified, and corresponding improvement measures are provided to improve network stability and security.
[0127] By addressing the above technical issues, the embodiments of this application aim to provide an efficient and accurate network troubleshooting solution to help enterprises and organizations better cope with faults and problems in network operation and maintenance.
[0128] Accordingly, the technical advantages or effects of the embodiments of this application are as follows:
[0129] 1. Efficient Troubleshooting: The purpose of this application is to provide an efficient network troubleshooting solution. By collecting historical data, a precise troubleshooting tree is used to quickly and accurately locate fault points during network operation and maintenance. Specifically, by using a solution based on a precise troubleshooting tree, each step can be checked progressively downwards (from coarse to fine), accurately locating fault points during network operation and maintenance. This method greatly improves the accuracy and efficiency of fault location, enabling maintenance personnel to find the problem more quickly.
[0130] 2. Providing Solutions and Troubleshooting Suggestions: This application aims not only to locate the fault point but also to provide corresponding solutions and troubleshooting suggestions, helping enterprises and organizations resolve network operation and maintenance issues more quickly. Specifically, in addition to locating the fault point, the solutions in this application can also provide corresponding solutions and troubleshooting suggestions based on a comparison of expected results and actual data. This helps operation and maintenance personnel better understand and solve problems, shortens the fault handling cycle, and reduces the complexity of fault diagnosis.
[0131] For example, the above actual data may refer to the inspection results of the inspection items. If the authentication server is not connected, the following two suggestions will be given: 1. Please check whether the configured server IP or port is normal. 2. Please check whether the relevant services of the server are enabled and whether the server is working properly.
[0132] For example, if the channel utilization rate is 93%, exceeding the normal range of 0% to 75%, the following four suggestions would be given: 1. Adjust the channel: High channel utilization may be due to multiple wireless networks operating on the same channel. Adjusting the channel of the wireless network can reduce interference and cross-interference, thereby reducing channel utilization. 2. Optimize the wireless network layout: Plan the layout of the wireless network reasonably to avoid wireless signal overlap and interference. Ensure that the distance between wireless access points is moderate, avoiding being too dense or too sparse. Adjust the location and number of access points according to actual needs and network load. 3. Increase the number of access points: If a large number of terminal devices are connected to the same access point, it may lead to excessively high channel utilization. Consider increasing the number of access points to reduce the load on each access point, thereby reducing channel utilization. 4. Use the 5GHz band: The 5GHz band has more channels, which can reduce interference and cross-interference, thereby reducing channel utilization.
[0133] 3. System Optimization and Prevention: In addition to troubleshooting, this application also provides system optimization and prevention measures. By analyzing and mining network operation and maintenance data, potential problems and risks can be identified, and corresponding improvement measures can be provided to enhance network stability and security.
[0134] For example, the aforementioned network operation and maintenance data mainly refers to the service chain data described above, including the completion status of each stage during the internet access process and the collection of various key information, such as the channel utilization rate of access devices and the signal strength of terminals. For example, excessively high channel utilization may lead to frequent internet disconnections for users; whether multiple DHCP server responses are received may result in the terminal failing to obtain an IP address, leading to the risk of network outages.
[0135] 4. User-friendly UI: Faced with complex network topologies, the technical solution in this application features a user-friendly UI, enabling even non-professionals to troubleshoot through simple operations and guidance. This reduces reliance on technical personnel and lowers training and learning costs. Figure 8 The troubleshooting guide page shown allows users to navigate to a specific page by clicking on the content requiring troubleshooting. The page can be something like... Figure 6 , Figure 5 The tree-based obstacle-solving UI diagram shown is used to guide users through the obstacle-solving process step by step.
[0136] 5. Visualization of the Troubleshooting Process: By using a UI page and a tree-structured, precise troubleshooting tree, users can intuitively understand the various stages of network operations and maintenance and the relationships between related functions, operations, or events. This visualization helps non-professionals quickly understand and locate problems, reducing the complexity of the troubleshooting process.
[0137] 6. Guided Troubleshooting Guide: This application provides a UI-based guided troubleshooting guide that leads users step-by-step to potentially problematic sub-stages and provides corresponding checkpoints and suggestions. This guided troubleshooting helps non-professionals follow certain procedures and steps during troubleshooting, improving the accuracy and efficiency of the process.
[0138] 7. Reduce maintenance costs: The network operation status is displayed through a simple UI page. When a network failure occurs, it guides maintenance personnel to troubleshoot and locate the problem step by step and provides effective solutions. Even non-professionals can quickly find and solve problems through the guidance of the UI page, which greatly reduces the reliance on professional maintenance personnel and the cost investment.
[0139] Please see Figure 9 The second aspect of this application provides a specific embodiment of a network operation and maintenance troubleshooting system, the system comprising:
[0140] The acquisition unit 901 is used to acquire the service dataset of at least one network service within the network topology, wherein a network service includes multiple service phases of the network lifecycle, each service phase has a service data subset, and the service data subset is obtained by combining service inspection items of multiple service sub-phases.
[0141] Processing unit 902 is used to summarize the service dataset of at least one network service according to multiple preset troubleshooting dimensions, so as to obtain multiple dimension objects under each troubleshooting dimension and the summarized service dataset of each dimension object.
[0142] The processing unit 902 is also used to obtain the network health status of the dimension object for each dimension object under each troubleshooting dimension, based on the aggregated service dataset of the dimension object;
[0143] The processing unit 902 is also used to obtain the network health status of the network topology under each troubleshooting dimension based on the network health status of the dimension objects under each troubleshooting dimension.
[0144] The processing unit 902 is also used to troubleshoot the network topology based on the network topology health status if the network health status of the network topology is abnormal.
[0145] In some specific examples, processing unit 902 is specifically used for:
[0146] If the network health status of the network topology is abnormal, find the dimension object with abnormal network health status under the target troubleshooting dimension; the target troubleshooting dimension can be any troubleshooting dimension.
[0147] In the service dataset of the abnormal dimension object, find the abnormal service data subset corresponding to the abnormal service stage, and then find the abnormal service inspection item corresponding to the abnormal service sub-stage in the abnormal service data subset.
[0148] In some specific examples, processing unit 902 is also used for:
[0149] If the network health status of the network topology is abnormal, network maintenance is performed on the network topology according to the abnormality solutions pre-configured for the abnormal service inspection items.
[0150] In some specific examples, the network health status of a dimension object includes: the dimension object's health status score, and / or, the percentage of abnormal service inspection items associated with the dimension object.
[0151] In some specific examples, processing unit 902 is specifically used for:
[0152] Determine the network health status of the service dataset for each network service;
[0153] For a dimension object, determine its network health status based on the network health status of the service dataset associated with the dimension object's aggregated service dataset.
[0154] In some specific examples, processing unit 902 is also used for:
[0155] The service dataset for a single network service is displayed in a tree structure.
[0156] In this tree structure, the root node is the identification information of the network service, the child nodes are the service stages within the network lifecycle, the child nodes of each service stage are the service sub-stages included in the service stage, and the service sub-stages include service inspection items.
[0157] In some specific examples, processing unit 902 is also used for:
[0158] A summary service dataset that displays dimensional objects in a tree structure;
[0159] In this tree structure, the root node is the dimension object, and the child nodes of the dimension object are the service datasets belonging to the dimension object.
[0160] In some specific examples, processing unit 902 is specifically used for:
[0161] Collect service datasets of at least one network service within a preset time span, and obtain multi-dimensional objects based on the unique object identifiers carried by the service datasets;
[0162] By aggregating service datasets belonging to the same dimension object, we obtain the aggregated service dataset of the dimension object.
[0163] In some specific examples, processing unit 902 is also used for:
[0164] The network health status of the network topology is displayed in each troubleshooting dimension.
[0165] In some specific examples, processing unit 902 is also used for:
[0166] Regularly monitor each service inspection item to identify potential problems in network services.
[0167] In this embodiment, the operations performed by each unit of the network operation and maintenance troubleshooting system are similar to those described in the first aspect or any specific method embodiment of the first aspect, and will not be repeated here. Of course, the specific implementation process of each operation in the first aspect of this application can also be found in the relevant description of the second aspect.
[0168] The network operation and maintenance troubleshooting system of this application embodiment can be specifically described as follows: Figure 10 As shown in the table below, the functions of each component module can be integrated into the acquisition unit 901 or the processing unit 902.
[0169]
[0170] In summary, the network operation and maintenance troubleshooting system of this application embodiment involves:
[0171] 1. Construction of a Precise Troubleshooting Tree: This involves a method for constructing a precise troubleshooting tree. It models the relationships between various stages, functions, and operations during network operation and maintenance, displaying these relationships in a tree structure on the UI page. By using algorithms to obtain and model the relationships between various terminal operations, users can intuitively see the actual status of each step of the terminal's operation. This construction method is innovative and practical, and can improve the accuracy of fault location.
[0172] 2. Design of Guided Troubleshooting Guide: This application provides a design method for a UI-based guided troubleshooting guide. It guides users step-by-step through potentially problematic sub-stages using a precise troubleshooting tree, providing corresponding checkpoints and suggestions. This design method enables non-professionals to troubleshoot by following the process and steps, improving the efficiency and accuracy of the troubleshooting process.
[0173] 3. Generation of Fault Solutions and Troubleshooting Suggestions: This embodiment of the application uses data collection and analysis tools to automatically generate fault solutions and troubleshooting suggestions by comparing expected results with actual data. This method can quickly identify problems and provide targeted solutions, reducing fault handling time and costs.
[0174] 4. System Optimization and Preventive Measures: This application focuses on system optimization and preventive measures. By analyzing and mining network operation and maintenance data, potential problems and risks are identified, and corresponding improvement measures are provided. This technology can improve network stability and security, and reduce the occurrence of future failures.
[0175] 5. Service chain related definitions: Visualize the network lifecycle of a terminal, including the definitions of stages, sub-stages, inspection items, etc. Users can intuitively understand the relationship between each stage of network operation and maintenance and related functions, operations or events, which greatly improves the maintainability of the network system.
[0176] Please see Figure 11 The electronic device 110 in this application embodiment may include one or more central processing units (CPUs) 111 and a memory 115, wherein the memory 115 stores one or more applications or data.
[0177] The memory 115 can be volatile or persistent storage. The program stored in the memory 115 can include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the central processing unit 111 can be configured to communicate with the memory 115 and execute the series of instruction operations stored in the memory 115 on the electronic device 110.
[0178] Electronic device 110 may also include one or more power supplies 112, one or more wired or wireless network interfaces 113, one or more input / output interfaces 114, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0179] The central processing unit 111 can perform the operations performed by the first aspect or any specific method embodiment of the first aspect, which will not be described in detail here.
[0180] This application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method as described in the first aspect or any specific implementation thereof.
[0181] This application provides a computer program product containing instructions or computer programs, which, when run on a computer, causes the computer to perform the method described in the first aspect or any specific implementation thereof.
[0182] It is understood that, in the various embodiments of this application, the sequence number of each step does not imply the order of execution. The execution order of each step should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The operations added or refined to the above-described methods, systems, or devices (if any) may not necessarily be executed in specific implementations. If more than two operations are added, these operations can be implemented in combination or individually, depending on the actual scenario.
[0183] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system (if it exists) and device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0184] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system or apparatus, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0185] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0186] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0187] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product (computer program product) is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a business server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A network operation and maintenance troubleshooting method, characterized in that, include: Obtain the service dataset of at least one network service within the network topology; wherein, one network service includes multiple service phases of the network lifecycle, each service phase has a service data subset, and the service data subset is obtained by combining service inspection items of multiple service sub-phases; According to multiple preset troubleshooting dimensions, the service dataset of the at least one network service is summarized to obtain multiple dimension objects under each troubleshooting dimension and the summarized service dataset of each dimension object. For each dimension object under each troubleshooting dimension, the network health status of the dimension object is obtained based on the aggregated service dataset of the dimension object; Based on the network health status of the dimension objects under each of the troubleshooting dimensions, the network health status of the network topology under each of the troubleshooting dimensions is obtained; If the network health status of the network topology is abnormal, troubleshooting is performed on the network topology based on the network health status.
2. The method according to claim 1, characterized in that, The troubleshooting of the network topology based on the network health status includes: Find the dimension objects with abnormal network health status under the target troubleshooting dimension; the target troubleshooting dimension can be any of the aforementioned troubleshooting dimensions. In the service dataset of the abnormal dimension object, find the abnormal service data subset corresponding to the abnormal service stage, and then find the abnormal service inspection item corresponding to the abnormal service sub-stage in the abnormal service data subset.
3. The method according to claim 2, characterized in that, After searching for the abnormal service inspection item corresponding to the abnormal service sub-stage in the abnormal service data subset, the method further includes: Network maintenance is performed on the network topology based on the anomaly solutions pre-configured for the anomaly service inspection items.
4. The method according to claim 1, characterized in that, The network health status of the dimension object includes: the health status score of the dimension object, and / or the percentage of abnormal service inspection items associated with the dimension object.
5. The method according to claim 1, characterized in that, The step of obtaining the network health status of the dimension object based on the aggregated service dataset of the dimension object includes: Determine the network health status of the service dataset for each of the network services described. For the dimension object, the network health status of the dimension object is determined based on the network health status of the service dataset associated with the aggregated service dataset of the dimension object.
6. The method according to claim 1, characterized in that, After obtaining the service dataset of at least one network service within the network topology, the method further includes: The service dataset for a single network service is displayed in a tree structure. Wherein, the root node of the tree structure is the identification information of the network service, the child nodes of the tree structure are the service stages within the network lifecycle, the child nodes of each service stage are the service sub-stages included in the service stage, and the service sub-stages include service inspection items.
7. The method according to claim 1 or 6, characterized in that, After obtaining the summary service dataset for each of the said dimension objects, the method further includes: The summary service dataset of the dimensional objects is displayed in the form of a tree structure; The root node of the tree structure is the dimension object, and the child nodes of the dimension object are the service datasets belonging to the dimension object.
8. The method according to any one of claims 1 to 6, characterized in that, The step of summarizing the service dataset of the at least one network service according to multiple preset troubleshooting dimensions to obtain multiple dimension objects under each troubleshooting dimension and a summarized service dataset of each dimension object includes: Collect service datasets of the network service at least once within a preset time span, and obtain multi-dimensional objects based on the unique object identifiers carried by the service datasets; The service datasets belonging to the same dimension object are aggregated to obtain the aggregated service dataset of the dimension object.
9. The method according to any one of claims 1 to 6, characterized in that, After obtaining the network health status of the network topology under each of the troubleshooting dimensions based on the network health status of the dimension objects under each of the troubleshooting dimensions, the method further includes: The network health status of the network topology in each of the troubleshooting dimensions is displayed respectively.
10. The method according to any one of claims 1 to 6, characterized in that, After obtaining the service dataset of at least one network service within the network topology, the method further includes: Regularly monitor each of the service inspection items to identify potential problems in the network services.
11. A network operation and maintenance troubleshooting system, characterized in that, include: The acquisition unit is used to acquire the service dataset of at least one network service within the network topology, wherein one network service includes multiple service phases of the network lifecycle, each service phase has a service data subset, and the service data subset is obtained by combining service inspection items of multiple service sub-phases; The processing unit is used to summarize the service dataset of the at least one network service according to multiple preset troubleshooting dimensions, so as to obtain multiple dimension objects under each troubleshooting dimension and the summarized service dataset of each dimension object. The processing unit is further configured to obtain the network health status of the dimension object for each dimension object under each troubleshooting dimension, based on the aggregated service dataset of the dimension object; The processing unit is further configured to obtain the network health status of the network topology under each of the troubleshooting dimensions based on the network health status of the dimension objects under each of the troubleshooting dimensions. The processing unit is also configured to troubleshoot the network topology based on the network health status if the network health status of the network topology is abnormal.
12. An electronic device, characterized in that, include: Central processing unit, memory, and input / output interfaces; The memory is either a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The system stores instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Microservice link anomaly positioning method and device, equipment and storage medium
CN115309578A
Abnormal positioning method and device, equipment and storage medium
CN115774648A