Fault determination method and device, nonvolatile storage medium and electronic equipment
By simulating service traffic data and clustering algorithms to process abnormal indicators, and combining service logs to match fault points, the problem of untimely detection of network equipment and network links in the service issuance platform is solved, and the reliability of network services is improved.
Patent Information
- Application Number
- CN202510378225.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-17
AI Technical Summary
The fault detection of network equipment and network links in the service issuing platform is not timely, resulting in low reliability of network services.
By sending multiple simulated service traffic data to network element devices, receiving response dial-test results, determining abnormal indicators, and using clustering algorithms to process abnormal indicators, obtaining business logs of front-end business systems, matching exception clusters and logs, and determining faulty network element devices and/or network links.
It realizes timely detection of faults in network equipment and network links in the service issuance platform, and improves the reliability of network services.
Smart Images

Figure CN120166052A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data analysis, and in particular, to a fault determination method and apparatus, a non-volatile storage medium, and an electronic device. Background Art
[0002] In the modern communication field, the service provisioning platform plays a key role in connecting front-end service operations (such as the business hall issuing service requests through the BOSS system) and back-end network configuration (automatically executing scripts for network element configuration). The success rate of this platform directly affects the efficiency of front-end service handling and service quality. However, in actual applications, the service provisioning platform faces multiple challenges. Especially when building a new network, the old network architecture is often used, resulting in equipment aging and increased maintenance complexity. The main disadvantages of related fault determination methods are as follows: 1. Abnormal detection is not timely: The abnormal phenomena of the platform lack regularity, and the process of discovering faults takes a long time. For example, a fault that occurs at night may not be detected until the front-end service is affected during the day, with an average time-consuming of 1-2 hours. 2. Decentralized monitoring means, difficult to correlate faults: The decentralized monitoring method makes the alarm information of each network element reported independently, lacking unified management. This not only increases the difficulty of analyzing the cause of the fault but also easily causes misjudgment because not all alarms directly reflect the actual problems of the platform.
[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0004] The present application provides a fault determination method and apparatus, a non-volatile storage medium, and an electronic device, so as to at least solve the technical problem that the reliability of network services is relatively low due to the untimely detection of faults in network devices and network links in the service provisioning platform.
[0005] According to one aspect of the present application, a fault determination method is provided, including: sending multiple simulated service traffic data to a network element device, and receiving the test results generated by the network element device in response to the multiple simulated service traffic data, where the network element device is communicatively connected to a front-end service system through a service provisioning platform; determining abnormal indicators in the test results, and performing clustering processing on the abnormal indicators by using a clustering algorithm to obtain multiple abnormal clusters; obtaining abnormal logs in the service logs of the front-end service system, matching the multiple abnormal clusters with the abnormal logs, and determining the network element device and / or network link with faults according to the matching results.
[0006] Optionally, match multiple exception clusters with the exception logs, and determine the network element devices and / or network links with faults according to the matching results, including: matching the time window corresponding to the exception cluster with the first field in the exception logs to obtain the first log data in the exception logs, where the first field is used to represent timestamp information; filtering out the second log data in the first log data that is consistent with the service type of the exception cluster according to the second field in the first log data, where the second field is used to identify the service affected by the event or operation; when the frequency of the preset error code in the second log data is greater than the first preset threshold, determine the root cause of the fault corresponding to the preset error code according to the preset mapping relationship, where the preset mapping relationship includes: the mapping relationship between the preset error code and the root cause of the fault, and the root cause of the fault includes: the fault subject and the fault content; determine the correlation index between the root cause of the fault and the exception index in the exception cluster, and when the correlation index is greater than the second preset threshold, determine the fault subject in the root cause of the fault as the network element device and / or network link with faults.
[0007] Optionally, after filtering out the second log data in the first log data that is consistent with the service type of the exception cluster according to the second field in the first log data, the method further includes: extracting a third field from the second log data and determining the service traffic record corresponding to the third field, where the third field is used to represent the session identifier, and the service traffic record includes five-tuple data, and the five-tuple data includes: source IP, destination IP, source port, destination port, and protocol; determining the network element devices and network links in the transmission path of the service traffic through the five-tuple data; if there is a target subject in the network element devices and network links in the transmission path of the service traffic that is consistent with the exception subject in the exception index of the exception cluster, determine the target subject as the network element device and / or network link with faults.
[0008] Optionally, the method further includes: performing thresholding discretization on the dial test results to obtain discrete indicators, aggregating the discrete indicators according to the time window to obtain a transaction database, where the transaction data in the transaction database includes the fault status labels triggered within the same time window; extracting frequent itemsets from the transaction database and determining association rules according to the frequent itemsets, where the association rules include: the association relationship between the connectivity index between the front-end business system and the network element device, the interface status data, the network element status data, and the probability of the target subject having a fault, and the target subject includes network element devices and network links; determining a target rule in the association rules whose support degree is greater than the third preset threshold and confidence degree is greater than the fourth preset threshold; obtaining the service traffic data between the front-end business system and the network element device, and if it is determined that there is target service traffic data that satisfies the target rule in the service traffic data, generating a target message, where the target message includes the identification information of the target subject.
[0009] Optionally, the clustering algorithm includes: a density clustering algorithm; using the clustering algorithm to cluster the abnormal metrics to obtain multiple abnormal clusters, including: obtaining a preset transmission rate of the simulated service traffic data within different time windows; in the case where the preset transmission rate is less than the first threshold, determining the first time window during which the preset transmission rate is less than the first threshold, and narrowing the ranges of the neighborhood radius and the minimum sample number in the density clustering algorithm within the first time window, and clustering the abnormal metrics according to the neighborhood radius and the minimum sample number after the range narrowing to obtain multiple abnormal clusters; in the case where the preset transmission rate is greater than the second threshold, determining the second time window during which the preset transmission rate is greater than the second threshold, and expanding the ranges of the neighborhood radius and the minimum sample number in the density clustering algorithm within the second time window, and clustering the abnormal metrics according to the neighborhood radius and the minimum sample number after the range expansion to obtain multiple abnormal clusters, where the second threshold is greater than the first threshold.
[0010] Optionally, sending multiple simulated service traffic data to a network element device and receiving the test results generated by the network element device in response to the multiple simulated service traffic data, including: generating multiple simulated service traffic data according to a preset service type; sending the multiple simulated service traffic data to the network element device through a service deployment platform, dynamically adjusting the traffic rate, the number of concurrent connections, and the packet size during the sending process, and injecting controllable abnormal metrics into the transmission path, where the controllable abnormal metrics include at least one of the following: network jitter and packet loss rate; receiving the test results generated by the network element device in response to the multiple simulated service traffic data, where the test results include: connectivity metrics between the front-end service system and the network element device, interface status data, and network element status data, where the connectivity metrics include at least one of the following: delay, packet loss rate, and path hop information, the interface status data includes at least one of the following: input / output error count, cyclic redundancy check error rate, and port negotiation status, and the network element status data includes at least one of the following: bandwidth utilization rate, processor utilization rate, and memory occupancy rate.
[0011] Optionally, after determining the network element device and / or network link with a fault according to the matching result, the method further includes: generating a root cause analysis report for the fault, where the root cause analysis report for the fault includes: abnormal metrics of the abnormal cluster, abnormal logs matching the abnormal metrics, and repair suggestions for the network element device and / or network link with the fault; sending the root cause analysis report to the target device.
[0012] According to another aspect of the present application, a fault determination device is further provided, including: a sending module, configured to send multiple simulated service traffic data to a network element device and receive a test result generated by the network element device in response to the multiple simulated service traffic data, wherein the network element device is communicatively connected to a front-end service system through a service deployment platform; a determination module, configured to determine abnormal indicators in the test result and perform clustering processing on the abnormal indicators by using a clustering algorithm to obtain multiple abnormal clusters; a matching module, configured to obtain abnormal logs in the service logs of the front-end service system, match the multiple abnormal clusters with the abnormal logs, and determine the network element device and / or network link with a fault according to the matching result.
[0013] According to another aspect of the present application, a non-volatile storage medium is further provided. The storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the above-mentioned fault determination method.
[0014] According to another aspect of the present application, an electronic device is further provided, including: a memory and a processor, where the processor is configured to run a program stored in the memory, and when the program runs, it executes the above-mentioned fault determination method.
[0015] According to another aspect of the present application, a computer program is further provided, wherein when the computer program is executed by a processor, it implements the above-mentioned fault determination method.
[0016] According to another aspect of the present application, a computer program product is further provided. The computer program product includes a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned fault determination method.
[0017] In the present application, by sending multiple simulated service traffic data to a network element device and receiving a test result generated by the network element device in response to the multiple simulated service traffic data, wherein the network element device is communicatively connected to a front-end service system through a service deployment platform; determining abnormal indicators in the test result and performing clustering processing on the abnormal indicators by using a clustering algorithm to obtain multiple abnormal clusters; obtaining abnormal logs in the service logs of the front-end service system, matching the multiple abnormal clusters with the abnormal logs, and determining the network element device and / or network link with a fault according to the matching result, the purpose of timely detecting faults of network devices and network links in the service deployment platform is achieved, thereby realizing the technical effect of improving the reliability of network services, and further solving the technical problem that the reliability of network services is relatively low due to the untimely detection of faults of network devices and network links in the service deployment platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0019] Figure 1 is a flowchart of a fault determination method according to an embodiment of the present application;
[0020] Figure 2 is a network topology structure diagram of a service delivery platform according to an embodiment of the present application;
[0021] Figure 3 is a flowchart of another fault determination method according to an embodiment of the present application;
[0022] Figure 4 is a structure diagram of a fault determination device according to an embodiment of the present application;
[0023] Figure 5 is a hardware structure block diagram of a computer terminal of a fault determination method according to an embodiment of the present application. Detailed implementation manners
[0024] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned accompanying drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] According to an embodiment of the present application, a method embodiment of a fault determination method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0027] Figure 1 is a flowchart of a fault determination method according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0028] Step S102, send multiple simulated service traffic data to the network element device, and receive the test results generated by the network element device in response to the multiple simulated service traffic data. Among them, the network element device is communicatively connected to the front-end service system through the service distribution platform.
[0029] The service distribution platform is located between the front-end service system and the back-end network device. The service distribution platform can receive service instructions sent by the front-end service system and convert them into a format suitable for processing by the back-end network device.
[0030] Figure 2 is a network topology structure diagram of a service distribution platform according to an embodiment of the present application. As Figure 2 shown, the service distribution platform is used for service requests in the operator BOSS system, generates specific network configuration instructions, and then sends the instructions to relevant network element devices through the Service Provider Network (SPN), such as the Access Gateway Control Function Entity (AGCF), Home Subscriber Server (HSS), Multimedia Session Control / Telephone Service Function (SSS / MM TEL), Electronic Number Mapping (ENUM), etc., to complete the opening or change of services. Among them, the AGCF is used for access control and manages the communication between users and the IMS network. The HSS is used to store user subscription data and is the user database in the IMS network. The SSS / MM TEL is responsible for processing multimedia session control and telephone services, including call establishment, routing selection, etc. The ENUM is used to convert telephone numbers into SIP URIs for addressing in the network.
[0031] The network element device receives service instructions from the distribution platform, performs corresponding operations, such as updating user configurations, configuring call routing, etc., and returns the execution results to the distribution platform.
[0032] In step S102, according to the preset business scenarios, the simulator generates multiple simulated service traffic data. These service traffic data cover different service types and cities, so as to comprehensively detect the health status of the network. For example, simulated traffic related to IMS services can be sent to network elements such as AGCF, HSS, SSS / MM TEL, and ENUM to simulate the service provisioning process. The generated simulated service traffic data is sent to the target network element device through the service provider network, and at the same time, the response data of the network element device is received. These response data are the test results for subsequent analysis.
[0033] Step S104, determine the abnormal indicators in the test results, and use the clustering algorithm to cluster the abnormal indicators to obtain multiple abnormal clusters.
[0034] Perform anomaly detection on each received test result. This can be based on preset health metric thresholds, such as response time, packet loss rate, retransmission count, etc., to determine whether there are abnormal indicators. The detected abnormal indicators are classified through a clustering algorithm. The purpose is to group similar abnormal indicators into a cluster, that is, an abnormal cluster. Through clustering analysis, the system will generate multiple abnormal clusters, and each cluster contains the same type or related abnormal indicators, which helps in subsequent fault location.
[0035] Step S106, obtain the abnormal logs in the business logs of the front-end business system, match the multiple abnormal clusters with the abnormal logs, and determine the network element devices and / or network links with faults according to the matching results.
[0036] Collect business logs from the front-end business system, especially the abnormal logs, which record any abnormal or error information during the business acceptance process. Match the collected abnormal logs with the abnormal clusters generated in step S104. By analyzing information such as timestamps, operation types, and network element names in the abnormal logs, corresponding to the abnormal indicators in the abnormal clusters, to determine the specific location of the fault. After the matching process is completed, the system will be able to determine whether the fault is caused by a specific network element device or a network link problem, and may even locate multiple fault sources. This information will be quickly pushed to the maintenance team through an instant warning notification for quick response and fault repair.
[0037] For example, assume that after a certain simulated service traffic is sent, the AGCF network element returns an abnormal response with an abnormally extended response time. When analyzing this abnormal response, the system identifies that the AGCF network element has a problem of too long response time and classifies it into an abnormal cluster. At the same time, the front-end service logs collected by the system also record that during the same time period, timeout errors occurred in the service processing related to the AGCF. By matching the abnormal cluster with the abnormal logs, the system determines that the AGCF network element has a fault, specifically manifested as an abnormal response time, and immediate inspection and repair are required. In this way, the accuracy of fault location and the response speed have been significantly improved.
[0038] According to the above steps, multiple simulated service traffic data are sent to the network element device, and the test results generated by the network element device in response to the multiple simulated service traffic data are received, where the network element device is communicatively connected to the front-end service system through the service deployment platform; abnormal indicators are determined from the test results, and a clustering algorithm is used to cluster the abnormal indicators to obtain multiple abnormal clusters; abnormal logs in the service logs of the front-end service system are obtained, the multiple abnormal clusters are matched with the abnormal logs, and according to the matching results, the network element device and / or network link with faults are determined, achieving the purpose of timely detecting faults in the network devices and network links in the service deployment platform, thereby realizing the technical effect of improving the reliability of network services.
[0039] The following Figure 1 exemplarily illustrates and explains the steps shown.
[0040] According to some optional embodiments of the present application, matching the multiple abnormal clusters with the abnormal logs and determining the network element device and / or network link with faults according to the matching results can be achieved by the following method: matching the time window corresponding to the abnormal cluster with the first field in the abnormal log to obtain the first log data in the abnormal log, where the first field is used to represent timestamp information; according to the second field in the first log data, filtering out the second log data in the first log data that is consistent with the service type of the abnormal cluster, where the second field is used to identify the service affected by the event or operation; in the case where the frequency of the preset error code appearing in the second log data is greater than the first preset threshold, according to the preset mapping relationship, determining the fault root cause corresponding to the preset error code, where the preset mapping relationship includes: the mapping relationship between the preset error code and the fault root cause, and the fault root cause includes: the fault subject and the fault content; determining the correlation index between the fault root cause and the abnormal indicators in the abnormal cluster, and in the case where the correlation index is greater than the second preset threshold, determining the fault subject in the fault root cause as the network element device and / or network link with faults.
[0041] In the above embodiments, first, according to the start time and end time of the abnormal cluster, a time window is defined, for example, [start time, end time]. Business logs are collected in real time from the log system of the distribution platform, and these logs record all events during the business processing. The collected logs are converted into structured data to ensure that each log contains keyword fields such as timestamp and affected_service. Determine the SQL query statement for filtering log records within a specific time window. For example: SELECT * FROM logs_table WHERE timestamp BETWEEN'start time' AND 'end time'; Execute the SQL query to extract all log entries within the abnormal cluster time window from the log database. From the above-filtered log records, extract all the first log data containing the timestamp field timestamp.
[0042] Secondly, determine the business types involved in the abnormal cluster, such as "video stream service" or "voice service". Identify the affected_service field in the logs to identify the business types involved in the log records. By comparing affected_service with the business types identified by the abnormal cluster, further filter the second log data consistent with the abnormal cluster business types. Use SQL queries or data processing scripts to extract all the second log data from the first log data where the affected_service field matches the abnormal cluster business type.
[0043] Thirdly, according to the error codes that appear in the second log data, such as HTTP status codes 500, 503 or SIP error code 500, etc., identify the preset error codes. Implement a statistical function to calculate the occurrence times of each error code in the second log data. Determine the abnormal threshold, for example, the first preset threshold = 10 times / 5 minutes, for detecting the abnormal frequency. Maintain a preset mapping relationship table from error codes to the root causes of faults, such as "HTTP 500 → database connection pool exhaustion". For error codes whose occurrence frequency exceeds the first preset threshold, determine the corresponding root cause of the fault according to the mapping relationship.
[0044] Finally, extract all abnormal metrics from the abnormal cluster, such as packet loss rate > 3% or CPU utilization rate > 90%. For each determined root cause of the fault, calculate its correlation with various abnormal metrics in the abnormal cluster. For example, when the database connection pool is exhausted, the CPU utilization rate usually increases significantly. Set the second preset threshold, for example, 0.8, for judging the association strength between the root cause of the fault and the abnormal metrics. When the correlation index between the root cause of the fault and the abnormal metrics in the abnormal cluster exceeds the second preset threshold, determine the faulty entity in the root cause of the fault as the faulty network element device and / or network link.
[0045] For example, assume that the time window of the anomaly cluster is from 2024-09-04 11:45:00 to 2024-09-04 11:50:00, and the service type is "video subscription service". Extract all log data with the timestamp field within this time window from the log database as the first log data. Further filter the records in the first log data where the affected_service field is "video subscription service" to obtain the second log data. Count the occurrence frequency of error codes in the second log data. For example, HTTP 503 appears 15 times, exceeding the first preset threshold (assumed to be 10 times / 5 minutes), then determine the root cause of the failure as CPU overload according to the preset mapping relationship (such as "HTTP 503 → NE device CPU overload"). Calculate the correlation index between the CPU overload and the anomaly metrics in the anomaly cluster (such as CPU utilization > 95%). If the result exceeds the second preset threshold (assumed to be 0.8), then determine the NE device with CPU overload as the faulty entity.
[0046] Further, according to the second field in the first log data, after filtering out the second log data that is consistent with the service type of the anomaly cluster from the first log data, the following steps can also be performed: Extract the third field from the second log data and determine the service traffic record corresponding to the third field, where the third field is used to represent the session identifier, and the service traffic record includes five-tuple data, and the five-tuple data includes: source IP, destination IP, source port, destination port, and protocol; Determine the NE devices and network links in the transmission path of the service traffic through the five-tuple data; If there is a target entity in the NE devices and network links in the transmission path of the service traffic that is consistent with the anomaly entity in the anomaly metrics of the anomaly cluster, determine the target entity as the faulty NE device and / or network link.
[0047] This embodiment details how to associate service traffic records using the Session ID, determine the service traffic transmission path through the five-tuple data (source IP, destination IP, source port, destination port, and protocol), and combine the anomaly metrics in the anomaly cluster to accurately locate the faulty NE device and / or network link. The specific steps are as follows:
[0048] 1. Extract the third field containing the session identifier from the second log data to ensure that each exception log has a unique Session ID. Integrate flow data analysis tools such as Telegraf or NetFlow to collect business traffic information in real time. Establish a mapping relationship between the Session ID and the five-tuple data to ensure that each session identifier can be associated with its corresponding network flow record. Use SQL or data processing scripts to extract all five-tuple data that matches the Session ID in the second log data from the network traffic database to generate the third log data (business traffic records).
[0049] 2. Compare the five-tuple data in the third log data with the network topology diagram to determine the complete transmission path of each session. For example, the path: user terminal → front-end server → firewall → switch → target network element. Decompose the complete transmission path into a series of network element devices and network links, such as {source: 'user terminal', destination: 'target network element', path: ['front-end server', 'firewall','switch']}. For each link, mark its source IP, destination IP, source port, destination port, and protocol type to provide link-level fault location information. Compare the exception entity of the exception metric in the exception cluster (such as "AGCF device with CPU overload") with the network element devices or network links in the business traffic path. If a network element device or network link that is the same as the exception entity is found in the business traffic path, it is determined as the fault point.
[0050] For example, assume that the exception Session ID extracted from the second log data is SID5678, and the corresponding five-tuple data is {sourceIP: '10.1.1.100', destIP: '10.1.1.200','sourcePort': 4000, 'destPort': 5060, 'protocol': 'SIP'}. Combining with the network topology diagram, determine the transmission path of this session as: user terminal → front-end server (10.1.1.100) → firewall → core switch → target SIP network element (10.1.1.200). If the exception metric in the exception cluster contains "SIP network element CPU overload" and the CPU utilization rate of the target SIP network element > 95%, then the target SIP network element is determined as the network element device with a fault, realizing accurate fault location.
[0051] According to some other alternative embodiments of the present application, the above-mentioned fault determination further includes the following steps: thresholding and discretizing the probing results to obtain discrete metrics, aggregating the discrete metrics according to a time window to obtain a transaction database, where the transaction data in the transaction database includes fault status tags triggered within the same time window; extracting frequent item sets from the transaction database and determining association rules based on the frequent item sets, where the association rules include: the association relationship between the connectivity metrics, interface status data, network element status data of the front-end business system and the network element device and the probability of a fault existing in the target entity, and the target entity includes network element devices and network links; determining target rules in the association rules whose support degree is greater than a third preset threshold and confidence degree is greater than a fourth preset threshold; obtaining the service traffic data between the front-end business system and the network element device, and if it is determined that there is target service traffic data satisfying the target rules in the service traffic data, generating a target message, where the target message includes the identification information of the target entity.
[0052] In the above embodiment, raw monitoring data such as connectivity metrics, interface status data, and network element status data are obtained from the probing system. According to historical data statistics, abnormal thresholds for each metric are set, such as "packet loss rate > 2%" or "CPU utilization rate > 85%". The raw monitoring data is converted into boolean tags, such as discrete metrics like "packet loss rate abnormal" and "CPU too high". For key metrics, such as CPU utilization rate, it can be further divided into tags such as "normal", "slightly abnormal", and "severely abnormal" to enhance the description ability of the rules. The length of the time window is defined, such as 5 minutes, for aggregating the discrete metrics. The discrete metrics are aggregated according to the time window to generate a transaction database, and each record contains a set of triggered fault status tags.
[0053] The Apriori algorithm is used to extract frequent item sets from the transaction database, including combinations of multiple simultaneously triggered fault status tags. The minimum support degree (the third preset threshold) is set, such as 0.1, to filter out item sets with insufficient support degree. Association rules are generated based on the frequent item sets, such as "packet loss rate abnormal ∧ CPU too high → high probability of fault". The confidence degree of each rule is calculated, and only the rules with a confidence degree greater than the fourth preset threshold (assumed to be 0.7) are retained to form a target rule library. The third preset threshold (support degree) and the fourth preset threshold (confidence degree) are set for screening the target rules. The rule library is dynamically updated to eliminate rules with low confidence or low support degree to ensure the timeliness and accuracy of the rules.
[0054] In the network monitoring system, the service traffic data between the front-end service system and network elements is obtained in real time, including five-tuple information and protocol type. The service traffic data is preprocessed to ensure consistency with the data formats in the discrete metrics and transaction database. For the service traffic data within each time window, the rules in the target rule library are used for matching. For example, if the rule "A: Packet loss rate is abnormal ∧ B: CPU is too high → Network element device failure" is satisfied, then mark the traffic data as abnormal. Identify the service traffic data that meets the target rules and mark it as "target service traffic data". Extract the identification information of network element devices or network links from the target service traffic data, such as IP addresses, port numbers, and device IDs. Construct a target message containing the identification information of the fault entity (such as "device ID: 1000, IP: 10.1.1.10, fault type: CPU is too high") and push an alarm.
[0055] For example, assume that the discrete metrics extracted from the dial test results are "packet loss rate is abnormal" and "CPU is too high". A transaction record {timestamp: '2024-09-04T11:30:00', tags: ['packet loss rate is abnormal', 'CPU is too high']} is formed in the transaction database. Through Apriori algorithm mining, the rule {packet loss rate is abnormal ∧ CPU is too high → high failure probability} is obtained, with a support degree of 0.15 and a confidence degree of 0.80, meeting the screening conditions. Traffic data that meets this rule is detected in the real-time service traffic data, and a target message {deviceID: '2001', IP: '10.1.1.20', port: 5060, protocol: 'SIP', status: 'fault', reason: 'CPU is too high'} is generated and an alarm is pushed in a timely manner to achieve rapid prediction and location of faults.
[0056] In some alternative embodiments of the present application, a clustering algorithm is used to cluster abnormal metrics to obtain multiple abnormal clusters, which can be achieved by the following method: Obtain the preset transmission rate of the simulated service traffic data in different time windows; when the preset transmission rate is less than the first threshold, determine the first time window during which the preset transmission rate is less than the first threshold, and within the first time window, narrow the ranges of the neighborhood radius and the minimum number of samples in the density clustering algorithm, and cluster the abnormal metrics according to the narrowed neighborhood radius and minimum number of samples to obtain multiple abnormal clusters; when the preset transmission rate is greater than the second threshold, determine the second time window during which the preset transmission rate is greater than the second threshold, and within the second time window, expand the ranges of the neighborhood radius and the minimum number of samples in the density clustering algorithm, and cluster the abnormal metrics according to the expanded neighborhood radius and minimum number of samples to obtain multiple abnormal clusters, where the second threshold is greater than the first threshold.
[0057] In the above embodiments, traffic data is obtained in real time from the simulated service traffic system, and the average transmission rate within different time windows is recorded. According to historical data and service requirements, a first threshold (low traffic threshold, e.g., 1 Mbps) and a second threshold (high traffic threshold, e.g., 10 Mbps, ensuring that the second threshold is greater than the first threshold) are set. When the preset transmission rate is less than the first threshold (e.g., 0.5 Mbps), it is identified as a low traffic scenario. When the preset transmission rate is greater than the second threshold (e.g., 15 Mbps), it is identified as a high traffic scenario. The time windows corresponding to the low traffic or high traffic scenarios are marked. For example, the first time window of the low traffic scenario is from 2024-09-04T11:20:00 to 2024-09-04T11:30:00.
[0058] Within the first time window, the neighborhood radius eps parameter of the DBSCAN algorithm is reduced to 0.5 times the standard value to capture sparse anomalies. For example, the standard eps value is 0.02 s, and after adjustment, it is 0.01 s. The min_samples parameter is reduced to 70% of the standard value, reducing the density requirement for anomaly detection. For example, the standard min_samples value is 5, and after adjustment, it is 3.
[0059] Within the second time window, the eps parameter of DBSCAN is expanded to 1.5 times the standard value to filter high-density noise. For example, the standard eps value is 0.02 s, and after adjustment, it is 0.03 s. The min_samples is increased to 120% of the standard value to ensure the aggregation strength of the anomaly clusters. For example, the standard min_samples value is 5, and after adjustment, it is 6.
[0060] According to the traffic scenario of the current time window, the adjusted neighborhood radius and minimum number of samples are dynamically selected. DBSCAN clustering is performed on the anomaly metrics, and multiple anomaly clusters are identified according to the adjusted parameters. Each cluster represents a set of possible fault points.
[0061] For example, assume that within the first time window (low traffic scenario, average transmission rate 0.8 Mbps), according to the preset low traffic parameter adjustment, the eps of the DBSCAN algorithm is adjusted from 0.02 s to 0.01 s, and the min_samples is reduced from 5 to 3. Clustering is performed on the anomaly metrics (such as CPU utilization, latency), and 3 anomaly clusters are generated, corresponding to "too high CPU", "abnormal latency", and "abnormal packet loss rate" respectively. Within the second time window (high traffic scenario, average transmission rate 20 Mbps), according to the preset high traffic parameter adjustment, the eps of DBSCAN is expanded from 0.02 s to 0.03 s, and the min_samples is increased from 5 to 6. Clustering is performed on the anomaly metrics under the expanded parameters, and 2 anomaly clusters are identified, effectively filtering out false anomalies under high-density noise.
[0062] Through the above process, an adaptive density clustering fault detection method based on simulated service traffic data is implemented, ensuring that the system can accurately and timely detect faults under different traffic scenarios, providing strong technical support for improving service quality and reducing operation and maintenance costs.
[0063] As some optional embodiments of the present application, sending multiple simulated service traffic data to a network element device and receiving the test results generated by the network element device in response to the multiple simulated service traffic data can be achieved by the following method: generating multiple simulated service traffic data according to a preset service type; sending the multiple simulated service traffic data to the network element device through a service distribution platform, dynamically adjusting the traffic rate, concurrent connection number, and packet size during the sending process, and injecting controllable exception metrics into the transmission path, where the controllable exception metrics include at least one of the following: network jitter and packet loss rate; receiving the test results generated by the network element device in response to the multiple simulated service traffic data, where the test results include: connectivity metrics between the front-end service system and the network element device, interface status data, and network element status data, where the connectivity metrics include at least one of the following: delay, packet loss rate, and path hop information, the interface status data includes at least one of the following: input / output error count, cyclic redundancy check error rate, and port negotiation status, and the network element status data includes at least one of the following: bandwidth utilization rate, processor utilization rate, and memory occupancy rate.
[0064] In the above embodiments, according to the service types involved in the distribution platform, such as video streams, voice services, data transmission, etc., set the types and characteristics of the simulated services. Design traffic templates for various service types, including packet size, protocol type, data structure, etc. Preset different traffic rates and concurrent connection numbers according to the processing capabilities of the network element device and the network bandwidth. Generate simulated service traffic data, and for each type of service, dynamically adjust the traffic rate, concurrent connection number, and packet size according to the network conditions.
[0065] Integrate the traffic generation tool with the service distribution platform to achieve the automatic sending of simulated service flows. Send the simulated service traffic to the target network element device according to the preset parameters (traffic rate, concurrent number, packet size). Simulate network jitter or packet loss situations on specific links or devices through network testing tools. Dynamically adjust the amplitude of network jitter (such as delay fluctuations) and the packet loss rate (such as 1% - 5%) to test the stability of the network element device. Collect and record the connectivity metrics, interface status data, and network element status data between the front-end service system and the network element device.
[0066] Furthermore, invalid or duplicate data can be filtered from the probing results to ensure the accuracy of the analyzed data. Classify and store connectivity metrics, interface status data, and network element status data for subsequent analysis. Set anomaly thresholds for each type of data based on historical data statistics. For example, packet loss rate > 3%, CPU utilization > 90%, etc. Use anomaly detection algorithms (such as IQR or Z-Score) to identify anomaly data points or anomaly clusters.
[0067] In some alternative embodiments of the present application, after determining the network element devices and / or network links with faults according to the matching results, the following steps can also be performed: Generate a root cause analysis report for the faults, where the root cause analysis report includes: anomaly metrics of the anomaly cluster, anomaly logs matching the anomaly metrics, and repair suggestions for the network element devices and / or network links with faults; Send the root cause analysis report to the target device.
[0068] From the real-time monitoring data, identify anomaly clusters through the DBSCAN density clustering algorithm, and each cluster represents a concentrated area of a set of anomaly metrics. Record the anomaly metrics in each anomaly cluster, including but not limited to CPU utilization, memory occupancy, packet loss rate, network latency, etc. Extract all anomaly logs with keyword fields such as timestamp, error code, session ID, affected service, etc. from the log database. Determine the time range of the anomaly cluster, compare it with the timestamp field in the log data, and filter out the log records within the same time window. Associate the anomaly metrics in the anomaly cluster with the relevant information in the log records. For example, if the log shows an "excessive CPU utilization" error code and the time window coincides, then this log is associated with the "high CPU" anomaly metric.
[0069] Design a template for the root cause analysis report for faults, including parts such as anomaly metrics, anomaly logs, faulty devices / links, possible causes, and repair suggestions. Insert the matching anomaly metrics and anomaly logs into the report template to form a fault description. Based on the analysis of the anomaly metrics and logs, locate the fault point to a specific network element device or network link, such as "abnormal packet loss rate of the core switch port Gi0 / 1". Retrieve repair suggestions from the mapping table of preset fault types to repair strategies according to the fault point. For example, "excessive memory occupancy" is mapped to "restart the device" or "adjust memory allocation". Customize repair scripts: For common faults, develop custom repair scripts and integrate them into the report for quick execution.
[0070] Integrate the root cause analysis report of faults with the alarm push system (such as enterprise WeChat robots). When the report is generated, it is immediately sent to the maintenance personnel or the alarm receiving end of the target device. According to the repair suggestions in the report, schedule the automated operation and maintenance tools to perform fault recovery operations. After the fault is processed, continuously monitor the status of the faulty device to verify the repair effect. Confirm whether the service has been restored through the dial test system or business traffic monitoring. According to the fault handling results, update the abnormal index thresholds, log correlation rules, and repair strategies to achieve the adaptive optimization of the system.
[0071] Figure 3 is a flowchart of another fault determination method according to an embodiment of the present application, as Figure 2 shown, the method includes the following steps:
[0072] Step S301, scenario analysis. Based on the current network status and historical data, the system dynamically adjusts the frequency and type of business flow monitoring to ensure that when the real business traffic changes, the simulated business flow can be effectively monitored, especially potential problems can still be discovered during the business traffic trough. According to factors such as network change operations (such as software upgrades, configuration modifications), intelligently adjust the business flow monitoring strategy to ensure that the monitoring covers key scenarios and improve the pertinence of fault detection.
[0073] Step S302, simulate business flow. The simulation process involves simulating the issuance of instructions in the distribution platforms of multiple cities (such as 5 cities) for specific business types (such as voice, data services) to detect the health status of the entire link from front-end acceptance to back-end network element configuration. Capture and parse the logs generated in the simulated business flow, and extract key information such as business execution status, network element response time, error codes, etc. for subsequent analysis and fault location.
[0074] Step S303, obtain the actual business flow. In parallel with the simulated business flow, the system continuously monitors the logs generated in the actual business flow to real-time monitor the health status of front-end business acceptance and the response of back-end network elements. It includes continuously monitoring the logs of business issuance to identify any anomalies inconsistent with the normal operation mode, such as delays, error responses, etc.
[0075] Step S304, clustering and location. Analyze the collected abnormal log data, and classify similar abnormal indicators into the same cluster through clustering algorithms to achieve rapid convergence and locate the problem points. Once the clustering analysis locates an anomaly, the system will immediately push the anomaly information to the operation and maintenance personnel through tools such as enterprise WeChat robots for quick response. The abnormal data and clustering results will be stored in the database as the basis for subsequent data analysis and model optimization.
[0076] Step S305, relevant network elements issue alarms. The system monitors the status of relevant network elements (such as AGCF, HSS, SSS / MM TEL, ENUM), collects their alarm information, and further confirms the fault point through correlation analysis with log data. Regularly check the connectivity and performance of the database to ensure that alarm information and other data can be recorded and analyzed in a timely and accurate manner.
[0077] Step S306, precise positioning. After the fault is identified, a warning notice will be immediately sent through communication tools such as enterprise WeChat, including fault details, impact scope, and preliminary positioning information, to help operation and maintenance personnel respond quickly. According to the abnormal data collected and the positioning results, continuously optimize the clustering algorithm, adaptive scheduling strategy, and fault positioning model to improve the accuracy and efficiency of automatic fault positioning.
[0078] Figure 4 is a structural diagram of a fault determination device according to an embodiment of the present application, as Figure 4 shown, the device includes:
[0079] A sending module 41, configured to send multiple simulated service traffic data to a network element device, and receive a test result generated by the network element device in response to the multiple simulated service traffic data, where the network element device is communicatively connected to a front-end service system through a service provider network.
[0080] A determining module 42, configured to determine abnormal indicators in the test result, and perform clustering processing on the abnormal indicators by using a clustering algorithm to obtain multiple abnormal clusters.
[0081] A matching module 43, configured to obtain abnormal logs in the service logs of the front-end service system, match the multiple abnormal clusters with the abnormal logs, and determine a network element device and / or a network link with a fault according to the matching result.
[0082] Optionally, the matching module 43 is further configured to perform the following steps: match the time window corresponding to the abnormal cluster with a first field in the abnormal log to obtain first log data in the abnormal log, where the first field is used to represent timestamp information; according to a second field in the first log data, filter out second log data in the first log data that is consistent with the service type of the abnormal cluster, where the second field is used to identify the service affected by an event or operation; in the case where the frequency of a preset error code appearing in the second log data is greater than a first preset threshold, determine a fault root cause corresponding to the preset error code according to a preset mapping relationship, where the preset mapping relationship includes: a mapping relationship between the preset error code and the fault root cause, and the fault root cause includes: a fault subject and a fault content; determine a correlation index between the fault root cause and the abnormal indicators in the abnormal cluster, and in the case where the correlation index is greater than a second preset threshold, determine the fault subject in the fault root cause as the network element device and / or the network link with a fault.
[0083] Optionally, the fault determination device is further configured to perform the following steps after filtering out second log data consistent with the service type of the abnormal cluster from the first log data according to the second field in the first log data: extract a third field from the second log data, and determine a service traffic record corresponding to the third field, where the third field is used to represent a session identifier, and the service traffic record includes five-tuple data, and the five-tuple data includes: source IP, destination IP, source port, destination port, and protocol; determine network element devices and network links in the transmission path of the service traffic through the five-tuple data; if there is a target entity in the network element devices and network links in the transmission path of the service traffic that is consistent with the abnormal entity in the abnormal metrics of the abnormal cluster, determine the target entity as the network element device and / or network link with a fault.
[0084] Optionally, the fault determination device is further configured to perform the following steps: perform threshold discretization on the probing result to obtain a discrete metric, aggregate the discrete metrics according to a time window to obtain a transaction database, where the transaction data in the transaction database includes fault status tags triggered within the same time window; extract frequent itemsets from the transaction database, and determine association rules according to the frequent itemsets, where the association rules include: the association relationship between the connectivity metric between the front-end service system and the network element device, the interface status data, the network element status data, and the probability of the target entity having a fault, and the target entity includes network element devices and network links; determine a target rule in the association rules whose support degree is greater than a third preset threshold and whose confidence degree is greater than a fourth preset threshold; obtain service traffic data between the front-end service system and the network element device, and if target service traffic data that meets the target rule is determined in the service traffic data, generate a target message, where the target message includes the identification information of the target entity.
[0085] Optionally, the clustering algorithm includes: a density clustering algorithm. The determination module 42 is further configured to perform the following steps: obtain the preset transmission rate of the simulated service traffic data in different time windows; in the case where the preset transmission rate is less than the first threshold, determine the first time window during which the preset transmission rate is less than the first threshold, and narrow the range of the neighborhood radius and the range of the minimum sample number in the density clustering algorithm within the first time window, and perform clustering processing on the abnormal metrics according to the neighborhood radius and the minimum sample number after the range is narrowed to obtain multiple abnormal clusters; in the case where the preset transmission rate is greater than the second threshold, determine the second time window during which the preset transmission rate is greater than the second threshold, and expand the range of the neighborhood radius and the range of the minimum sample number in the density clustering algorithm within the second time window, and perform clustering processing on the abnormal metrics according to the neighborhood radius and the minimum sample number after the range is expanded to obtain multiple abnormal clusters, where the second threshold is greater than the first threshold.
[0086] Optionally, the sending module 41 is further configured to perform the following steps: generate multiple simulated service traffic data according to a preset service type; send the multiple simulated service traffic data to the network element device through the service delivery platform, dynamically adjust the traffic rate, the concurrent connection number, and the packet size during the sending process, and inject controllable exception metrics into the transmission path, where the controllable exception metrics include at least one of the following: network jitter and packet loss rate; receive the test results generated by the network element device in response to the multiple simulated service traffic data, where the test results include: connectivity metrics between the front-end service system and the network element device, interface status data, and network element status data, where the connectivity metrics include at least one of the following: latency, packet loss rate, and path hop information, the interface status data includes at least one of the following: input / output error count, cyclic redundancy check error rate, and port negotiation status, and the network element status data includes at least one of the following: bandwidth utilization rate, processor utilization rate, and memory occupancy rate.
[0087] Optionally, after determining the network element device and / or network link with a fault according to the matching result, the fault determination device is further configured to perform the following steps: generate a fault root cause analysis report, where the fault root cause analysis report includes: exception metrics of the exception cluster, exception logs matching the exception metrics, and repair suggestions for the network element device and / or network link with the fault; send the root cause analysis report to the target device.
[0088] It should be noted that the above Figure 4 each module may be a program module (for example, a set of program instructions for implementing a specific function), or a hardware module. For the latter, it may be presented in the following forms, but not limited to: the manifestation form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.
[0089] It should be noted that Figure 4 The preferred implementation manners of the illustrated embodiments may be referred to Figure 1 the relevant descriptions of the illustrated embodiments, which will not be elaborated here.
[0090] Figure 5 shows a hardware structure block diagram of a computer terminal for implementing the fault determination method. As Figure 5As shown, the computer terminal 50 may include one or more processors 502 (illustrated as 502a, 502b, ……, 502n in the figure) (the processor 502 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 504 for storing data, and a transmission module 506 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 5 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 50 may further include more or fewer components than those Figure 5 shown in, or have a different configuration from that Figure 5 shown.
[0091] It should be noted that the above one or more processors 502 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 50. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0092] The memory 504 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the fault determination method in the embodiments of the present application. The processor 502 executes various functional applications and data processing by running the software programs and modules stored in the memory 504, that is, implements the above-mentioned fault determination method. The memory 504 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 504 may further include a memory remotely set relative to the processor 502, and these remote memories can be connected to the computer terminal 50 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0093] The transmission module 506 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the computer terminal 50. In one example, the transmission module 506 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 506 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0094] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 50.
[0095] It should be noted here that in some alternative embodiments, the above-mentioned Figure 5 shown computer terminal may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 5 is only an example of a specific specific instance and is intended to show the types of components that may exist in the above computer terminal.
[0096] It should be noted that Figure 5 the shown computer terminal is used to execute Figure 1 the shown fault determination method. Therefore, the relevant explanations in the execution method of the above commands also apply to this electronic device and will not be elaborated here.
[0097] The embodiment of the present application also provides a non-volatile storage medium. The non-volatile storage medium includes a stored program. Wherein, when the program runs, it controls the device where the storage medium is located to execute the above fault determination method.
[0098] A program for the non-volatile storage medium to execute the following functions: sending multiple simulated service traffic data to a network element device and receiving the dial test results generated by the network element device in response to the multiple simulated service traffic data, wherein the network element device is communicatively connected to a front-end service system through a service provisioning platform; determining abnormal indicators in the dial test results and performing clustering processing on the abnormal indicators by using a clustering algorithm to obtain multiple abnormal clusters; obtaining abnormal logs in the service logs of the front-end service system, matching the multiple abnormal clusters with the abnormal logs, and determining the network element device and / or network link with faults according to the matching results.
[0099] The embodiment of the present application also provides an electronic device, including: a memory and a processor. The processor is used to run the program stored in the memory. Wherein, when the program runs, it executes the above fault determination method.
[0100] The processor is used to run a program that performs the following functions: sending multiple simulated service traffic data to network element devices, and receiving the test results generated by the network element devices in response to the multiple simulated service traffic data. Among them, the network element devices are communicatively connected to the front-end service system through a service deployment platform; determining abnormal indicators in the test results, and performing clustering processing on the abnormal indicators using a clustering algorithm to obtain multiple abnormal clusters; obtaining abnormal logs in the service logs of the front-end service system, matching the multiple abnormal clusters with the abnormal logs, and determining the network element devices and / or network links with faults according to the matching results.
[0101] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0102] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0103] In the above embodiments of the present application, the information collected is information and data authorized by the user or fully authorized by all parties, and the processing of the relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with relevant laws, regulations, and standards, takes necessary protection measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.
[0104] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the units or modules can be in an electrical or other form.
[0105] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0106] In addition, the functional units in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0107] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0108] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A fault determination method, characterized in that: include: Sending a plurality of simulated service flow data to a network element device, and receiving a dialing test result generated by the network element device in response to the plurality of simulated service flow data, wherein the network element device is communicatively connected with a front-end service system through a service provisioning platform; Determining abnormal indicators in the dialing test results, and clustering the abnormal indicators using a clustering algorithm to obtain multiple abnormal clusters; Acquire an exception log in the service log of the front-end service system, match the plurality of exception clusters with the exception log, and determine a faulty network element device and / or network link according to the matching result.
2. The method according to claim 1, characterized in that Matching the plurality of abnormal clusters with the abnormal log, and determining the faulty network element device and / or network link according to the matching result, includes: Matching the time window corresponding to the abnormal cluster with the first field in the abnormal log to obtain first log data in the abnormal log, wherein the first field is used to represent timestamp information; According to a second field in the first log data, second log data consistent with the business type of the abnormal cluster is screened out from the first log data, wherein the second field is used to identify a service affected by an event or operation; In the case where the frequency of the preset error code appearing in the second log data is greater than the first preset threshold, determining the fault root cause corresponding to the preset error code according to a preset mapping relationship, wherein the preset mapping relationship includes: a mapping relationship between the preset error code and the fault root cause, and the fault root cause includes: a fault subject and a fault content; Determine a correlation index between the root cause of the fault and the abnormal index in the abnormal cluster, and when the correlation index is greater than a second preset threshold, determine the fault subject in the root cause of the fault as a faulty network element device and / or network link.
3. The method according to claim 2, characterized in that After filtering out second log data consistent with the business type of the abnormal cluster from the first log data according to the second field in the first log data, the method further includes: Extracting a third field from the second log data, and determining a service flow record corresponding to the third field, wherein the third field is used to indicate a session identifier, and the service flow record includes five-tuple data, and the five-tuple data includes: source IP, destination IP, source port, destination port, and protocol; Determine the network element devices and network links in the transmission path of the service flow through the five-tuple data; If there is a target subject in the network element equipment and network link in the transmission path of the business traffic that is consistent with the abnormal subject in the abnormal indicator in the abnormal cluster, the target subject is determined to be a faulty network element equipment and / or network link.
4. The method according to claim 1, characterized in that: The method further comprises: Discretizing the dial test result by thresholding to obtain discrete indicators, aggregating the discrete indicators according to time windows to obtain a transaction database, wherein the transaction data in the transaction database includes fault status tags triggered within the same time window; Extracting frequent item sets from the transaction database, and determining association rules according to the frequent item sets, wherein the association rules include: an association relationship between a connectivity indicator between the front-end service system and the network element device, interface status data, network element status data, and a target subject having a failure probability, wherein the target subject includes a network element device and a network link; Determine a target rule in the association rules whose support is greater than a third preset threshold and whose confidence is greater than a fourth preset threshold; The service flow data between the front-end service system and the network element device is obtained. If it is determined that there is target service flow data satisfying the target rule in the service flow data, a target message is generated, wherein the target message includes identification information of the target subject.
5. The method according to claim 1, characterized in that The clustering algorithm includes: density clustering algorithm; A clustering algorithm is used to cluster the abnormal indicators to obtain multiple abnormal clusters, including: Obtaining a preset transmission rate of the simulated service flow data in different time windows; In the case where the preset transmission rate is less than a first threshold, determining a first time window during which the preset transmission rate is less than the first threshold, and reducing the range of the neighborhood radius and the range of the minimum number of samples in the density clustering algorithm within the first time window, and clustering the abnormal indicators according to the reduced neighborhood radius and the minimum number of samples to obtain a plurality of abnormal clusters; In the case where the preset transmission rate is greater than a second threshold, a second time window during which the preset transmission rate is greater than the second threshold is determined, and the range of the neighborhood radius and the range of the minimum number of samples in the density clustering algorithm are expanded within the second time window, and the abnormal indicators are clustered according to the expanded neighborhood radius and the minimum number of samples to obtain a plurality of abnormal clusters, wherein the second threshold is greater than the first threshold.
6. The method according to claim 1, characterized in that Sending a plurality of simulated service flow data to a network element device, and receiving a dialing test result generated by the network element device in response to the plurality of simulated service flow data, comprises: Generate a plurality of simulated service flow data according to a preset service type; Sending a plurality of the simulated service flow data to the network element device through the service issuance platform, dynamically adjusting the flow rate, the number of concurrent connections and the message size during the sending process, and injecting a controllable abnormality indicator into the transmission path, wherein the controllable abnormality indicator includes at least one of the following: network jitter and packet loss rate; Receive the dialing test result generated by the network element device in response to the plurality of simulated service flow data, wherein the dialing test result includes: connectivity indicators, interface status data and network element status data between the front-end service system and the network element device, wherein the connectivity indicators include at least one of the following: delay, packet loss rate and path hopping information, the interface status data includes at least one of the following: input / output error count, cyclic redundancy check error rate and port negotiation status, and the network element status data includes at least one of the following: Bandwidth utilization, processor utilization, and memory utilization.
7. The method according to claim 1, characterized in that After determining the faulty network element device and / or network link according to the matching result, the method further includes: Generate a fault root cause analysis report, wherein the fault root cause analysis report includes: abnormal indicators of the abnormal cluster, abnormal logs matching the abnormal indicators, and repair suggestions for the faulty network element equipment and / or network links; send the root cause analysis report to the target device.
8. A fault determination device, characterized in that: include: A sending module, used for sending a plurality of simulated service flow data to a network element device, and receiving a dialing test result generated by the network element device in response to the plurality of simulated service flow data, wherein the network element device is connected to the front-end service system through a service issuance platform; A determination module, used to determine abnormal indicators in the dialing test result, and cluster the abnormal indicators using a clustering algorithm to obtain multiple abnormal clusters; The matching module is used to obtain an abnormal log in the service log of the front-end service system, match the multiple abnormal clusters with the abnormal log, and determine the faulty network element device and / or network link according to the matching result.
9. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored program, wherein when the program is executed, the device where the non-volatile storage medium is located is controlled to execute the fault determination method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the fault determination method according to any one of claims 1 to 7 is executed when the program is run.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the fault determination method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
System fault detection method, device, equipment, medium and program product
CN120582958A
Log collection method, device and system, server and storage medium
CN121125472A
Dial test data storage method, device and equipment, medium and program product
CN121255108A