Internet line fault prediction and health management method based on multi-protocol active probing
By using a multi-protocol proactive dialing method, a set of dialing protocol-Internet dialing target binary pairs is generated. By utilizing the Poisson time axis and cross-protocol linkage density, the problems of sampling distortion and insufficient early fault samples in Internet line quality monitoring are solved, and efficient line fault prediction and health management are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHUN ELECTRONIC TECH (SHANGHAI) CO LTD
- Filing Date
- 2026-06-17
- Publication Date
- 2026-07-21
AI Technical Summary
In existing Internet line quality monitoring solutions, fixed-interval polling is prone to phase synchronization with periodic network interference, resulting in sampling distortion; single-protocol or single-target detection is difficult to distinguish between local line faults and remote target maintenance fluctuations; fixed absolute delay thresholds cannot adapt to the natural path differences of different targets; and there is insufficient early-stage fault sample collection.
A multi-protocol active dialing test method is adopted. By generating a set of dialing test protocols and Internet dialing test targets, maintaining the exploration time axis and the dense time axis, using the Poisson time axis to drive dialing test, triggering cross-protocol linkage density, and combining the cumulative deviation counter for fault early warning, a multi-protocol and multi-target collaborative judgment is achieved.
It achieves an unbiased reflection of the true quality distribution of the line, effectively distinguishes between local line faults and remote target maintenance fluctuations, reduces false alarms and missed alarms, and has good deployability, real-time performance and interpretability, ensuring reliable reporting even when the line is interrupted.
Smart Images

Figure CN122437779A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of Internet technology, specifically relating to a method for predicting Internet line faults and managing health based on multi-protocol proactive dialing. Background Technology
[0002] Internet line quality monitoring is fundamental to ensuring the normal operation of various network services. Current mainstream line quality monitoring methods can be broadly categorized into two types. The first is passive monitoring, which assesses line health by collecting indicators such as traffic statistics, packet loss counts, and interface status from network devices. This method relies on the operational status of the monitored equipment; when the line is completely interrupted, the equipment often cannot continue reporting data, thus the monitoring system loses its ability to detect the most severe moment of the fault. The second type is active probing, where probing devices periodically initiate probe requests to the target, assessing line quality based on response latency and success rate. Existing active probing schemes mostly use fixed-interval polling methods, relying on a single protocol, a single target, or a fixed absolute latency threshold as the basis for judgment. Fixed-interval polling is prone to phase synchronization with periodic network interference such as operator policy updates, equipment batch processing, and session keep-alive, causing the sampling results to systematically deviate from the true quality distribution of the line. Probing with a single protocol or a single target is difficult to distinguish between local line faults and operational fluctuations of remote targets, often misjudging temporary maintenance of remote targets as local line anomalies. Fixed absolute latency thresholds cannot adapt to the inherent differences in physical distance and network path between different targets. They are too lenient for near-end targets and too strict for far-end targets, resulting in both false negatives and false positives. Furthermore, most existing solutions maintain low-frequency detection even after anomalies occur, making it difficult to quickly collect sufficient samples in the early stages of a fault to support reliable judgment. Therefore, how to simultaneously improve sampling unbiasedness, multi-protocol, multi-target collaborative judgment, and rapid sample collection in the early stages of a fault remains an unresolved issue. Summary of the Invention
[0003] The main objective of this invention is to provide a method for predicting and managing Internet line faults based on multi-protocol active probing, in order to solve problems such as fixed-interval polling being prone to phase synchronization with periodic network interference leading to sampling distortion, single-protocol or single-target detection being unable to distinguish between local line faults and remote target maintenance fluctuations, fixed absolute delay thresholds being unable to adapt to the natural path differences of different targets, and insufficient early-stage fault sample collection in existing line quality monitoring schemes.
[0004] To solve the above problems, the technical solution of the present invention is implemented as follows: A multi-protocol proactive testing-based method for predicting and managing internet line faults is implemented within a distributed monitoring system comprised of testing equipment, IoT channels, and a data collection platform. This includes: Step 1: Generate a set of tuples consisting of the dialing test protocol - Internet dialing test target tuple. The tuple is an abbreviation for the dialing test protocol - Internet dialing test target tuple. Set the initial baseline value for each tuple. The working mode identifier is initially set to exploration mode. The suspicious tuple set is initially empty. Step 2: Each tuple maintains two Poisson time axes: an exploration time axis and a dense time axis. The intensity parameters of the exploration time axis and the dense time axis are the exploration time base intensity and the dense time base intensity, respectively, with the dense time base intensity being greater than the exploration time base intensity. In exploration mode, only the exploration time axis triggers active probing, while in dense mode, both the exploration time axis and the dense time axis trigger active probing. Abnormal metrics are reported via the IoT channel. Step 3: When an abnormal measurement value is generated, the instantaneous deviation ratio is obtained based on the baseline reference value of the abnormal measurement value and the binary pair that generated the abnormal measurement value. In the exploration mode, when the instantaneous deviation ratio of any binary pair reaches the instantaneous deviation multiple, cross-protocol linkage intensification is triggered, the working mode identifier is switched to intensive mode, the intensive time axis of all binary pairs in the binary pair set is regenerated for the next batch of measurement time, and the intensive observation window is started. During the intensive observation window, binary pairs whose instantaneous deviation ratio reaches the instantaneous deviation multiple are registered to the suspicious binary pair set, and a cumulative deviation counter is assigned to the registered binary pairs. When the intensive observation window ends, if the value of any cumulative deviation counter reaches the alarm count threshold, a line fault warning event is reported, and a new intensive observation window is started. If the values of all cumulative deviation counters are less than the alarm count threshold, the system switches to exploration mode.
[0005] Furthermore, the targets of the internet dialing test cover at least two of the following: web servers on the internet, server TCP ports on the internet, devices on the internet that support the ICMP protocol, and public DNS servers; the dialing test protocols cover at least two of the following: ICMP-based dead / live detection, HTTP / HTTPS-based web application response speed detection, TCP connection-based network application response speed detection, and DNS-based domain name resolution response speed detection.
[0006] Furthermore, the testing device downloads the testing configuration from the system management platform. The testing configuration includes a set of testing targets, a set of testing protocols, and a protocol-target adaptation relationship. The set of testing targets contains more than two Internet testing targets, the set of testing protocols contains more than two testing protocols, and the protocol-target adaptation relationship defines the category of Internet testing targets adapted to each testing protocol. Based on the protocol-target adaptation relationship, the testing device generates a set of tuples from the set of testing targets and the set of testing protocols. Each tuple in the set of tuples consists of one testing protocol and one Internet testing target, and the combination of testing protocol and Internet testing target in each tuple conforms to the protocol-target adaptation relationship.
[0007] Furthermore, the abnormal measurement values are uniformly normalized to a form where the larger the value, the worse the line condition. The normalization rules are as follows: when the response time, connection establishment time, or parsing time is obtained from active dialing, the corresponding time value is used as the abnormal measurement value; when the packet loss result, timeout result, or error result is obtained from active dialing, the preset packet loss penalty value, preset timeout penalty value, or preset error penalty value is used as the abnormal measurement value, respectively; the preset packet loss penalty value, preset timeout penalty value, and preset error penalty value are all greater than 0.
[0008] Furthermore, the initial baseline value of each tuple is calibrated as follows: the testing device continuously performs calibration tests on each tuple an equal number of times as the calibration test count, obtaining a calibration anomaly metric sequence for each tuple, where the calibration test count is a positive integer greater than 1; elements corresponding to preset packet loss penalty values, preset timeout penalty values, and preset error penalty values are removed from the calibration anomaly metric sequence to obtain a valid calibration anomaly metric sequence; when the valid calibration anomaly metric sequence is not empty, the median of the valid calibration anomaly metric sequence is used as the initial baseline value of the corresponding tuple; when the valid calibration anomaly metric sequence is empty, the preset calibration failure baseline value is used as the initial baseline value of the corresponding tuple, where the preset calibration failure baseline value is greater than the preset minimum baseline value; when the initial baseline value is less than the preset minimum baseline value, the value of the preset minimum baseline value is used as the initial baseline value of the corresponding tuple, where the preset minimum baseline value is greater than 0.
[0009] Furthermore, for each pair, the baseline reference value is the larger of the initial baseline value and the preset minimum baseline value; the instantaneous deviation ratio is equal to the abnormal measurement value divided by the baseline reference value; the instantaneous deviation multiple is the threshold multiple for triggering cross-protocol linkage intensification, and the instantaneous deviation multiple is greater than 1.
[0010] Furthermore, the next measurement time on the Poisson timeline is generated using the inverse exponential distribution transformation method. The generation method is as follows: the measurement device independently extracts a random number from a uniform distribution in the interval 0 to 1. The value of the random number is greater than 0 and less than or equal to 1. The time interval between the next measurement time and the current time is equal to the negative of the natural logarithm of the random number divided by the intensity parameter. The unit of the time interval is seconds. The unit of the intensity parameter is times / second. On the exploration timeline, the intensity parameter is the value of the exploration time base intensity. On the dense timeline, the intensity parameter is the value of the dense time base intensity. After each active measurement is completed, the corresponding Poisson timeline immediately regenerates the next measurement time in the above manner.
[0011] Furthermore, in exploration mode, the dense time axis is in a standby state; when cross-protocol linkage density is triggered, the probe device uses the working mode identifier switching time as the current time, and regenerates the next probe time for the dense time axis of all tuples in the tuple set, using the dense time base intensity as the intensity parameter; the duration of the dense observation window is the dense observation window length. During the dense observation window, the exploration time axis and the dense time axis of all tuples in the tuple set trigger active probe in parallel, and the instantaneous probe frequency of each tuple is the sum of the exploration time base intensity and the dense time base intensity.
[0012] Furthermore, when cross-protocol linkage intensification is triggered, the tuple that triggers cross-protocol linkage intensification is registered in the suspicious tuple set, and a cumulative deviation counter is assigned to the registered tuple, which is initialized to 1. During the intensive observation window, when the instantaneous deviation ratio of a tuple reaches the instantaneous deviation multiple, registration and counting actions are performed: if the tuple that generates the instantaneous deviation ratio is already registered in the suspicious tuple set, the corresponding cumulative deviation counter is incremented by 1; if the tuple that generates the instantaneous deviation ratio is the first to reach the instantaneous deviation multiple during this intensive observation window, the tuple that generates the instantaneous deviation ratio is registered in the suspicious tuple set, and a cumulative deviation counter is assigned to the registered tuple, which is initialized to 1. The alarm counting threshold is a positive integer.
[0013] Furthermore, the line fault warning event includes the identifier of the binary pair that triggered the fault warning, all abnormal measurement values generated by the corresponding binary pair during the dense observation window, and all instantaneous deviation ratios. Starting a new dense observation window includes: maintaining the dense mode identifier, using the start time of the new dense observation window as the current time, regenerating the next test time for the dense time axis of all binary pairs in the binary pair set, clearing the suspicious binary pair set, and deleting each cumulative deviation counter. When switching to the exploration mode, the dense time axis of all binary pairs in the binary pair set is restored to the ready-to-use state, all binary pairs in the binary pair set are restored to active testing triggered only by the exploration time axis, the suspicious binary pair set is cleared, and each cumulative deviation counter is deleted.
[0014] This invention offers the following advantages: By maintaining two independent Poisson time axes—a probing time axis and a dense time axis—in parallel for each binary tuple, the initiation time of active probing lacks a predictable periodic structure. This prevents phase locking from various periodic interferences on the line, ensuring that the obtained observation samples unbiasedly reflect the true quality distribution of the line and avoiding the systematic distortion caused by sampling aliasing in fixed-interval polling. This invention uses a binary tuple composed of the probing protocol and the Internet probing target as the smallest observation and judgment unit. It calculates the relative deviation based on the baseline reference value determined by the typical normal level of each binary tuple. This unifies the judgment of targets at different physical distances and on different network paths under the criterion of how many times worse than their normal state. This avoids the imbalance of being too lenient for near-end targets and too strict for far-end targets, and effectively distinguishes between local line faults and operational fluctuations of far-end targets through cross-observation of multiple protocols and multiple targets, significantly reducing false alarms and missed alarms. This invention, in probe mode, triggers cross-protocol linkage intensive monitoring upon observing any transient anomalies, causing all binary pairs to simultaneously enter a high-frequency observation phase. This allows for rapid collection of sufficient samples in the early stages of a fault, and graded judgments are made based on whether the cumulative number of deviations reaches a threshold, balancing alarm sensitivity and reliability. The entire method does not rely on machine learning models or long-term historical data storage, and can be executed entirely locally on testing equipment with limited computing power. It possesses good deployability, real-time performance, and interpretability, and because the reporting channel is physically independent of the monitored line, it ensures reliable reporting even when the line is completely interrupted. Attached Figure Description
[0015] Figure 1 A timing diagram illustrating a dual-time-base Poisson time process for a single binary tuple provided in an embodiment of the present invention; Figure 2 A schematic diagram illustrating the collaborative response of a set of binary tuples during cross-protocol linkage intensification, provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the cumulative deviation counter value trajectory and alarm determination during the dense observation window provided in this embodiment of the invention. Detailed Implementation
[0016] A multi-protocol proactive testing-based method for predicting and managing internet line faults is implemented within a distributed monitoring system comprised of testing equipment, IoT channels, and a data collection platform. This includes: Step 1: Generate a set of tuples consisting of the dialing test protocol - Internet dialing test target tuple. The tuple is an abbreviation for the dialing test protocol - Internet dialing test target tuple. Set the initial baseline value for each tuple. The working mode identifier is initially set to exploration mode. The suspicious tuple set is initially empty. Step 2: Each tuple maintains two Poisson time axes: an exploration time axis and a dense time axis. The intensity parameters of the exploration time axis and the dense time axis are the exploration time base intensity and the dense time base intensity, respectively, with the dense time base intensity being greater than the exploration time base intensity. In exploration mode, only the exploration time axis triggers active probing, while in dense mode, both the exploration time axis and the dense time axis trigger active probing. Abnormal metrics are reported via the IoT channel. Step 3: When an abnormal measurement value is generated, the instantaneous deviation ratio is obtained based on the baseline reference value of the abnormal measurement value and the binary pair that generated the abnormal measurement value. In the exploration mode, when the instantaneous deviation ratio of any binary pair reaches the instantaneous deviation multiple, cross-protocol linkage intensification is triggered, the working mode identifier is switched to intensive mode, the intensive time axis of all binary pairs in the binary pair set is regenerated for the next batch of measurement time, and the intensive observation window is started. During the intensive observation window, binary pairs whose instantaneous deviation ratio reaches the instantaneous deviation multiple are registered to the suspicious binary pair set, and a cumulative deviation counter is assigned to the registered binary pairs. When the intensive observation window ends, if the value of any cumulative deviation counter reaches the alarm count threshold, a line fault warning event is reported, and a new intensive observation window is started. If the values of all cumulative deviation counters are less than the alarm count threshold, the system switches to exploration mode.
[0017] The entire method is implemented in a distributed monitoring system, consisting of three parts: testing devices, IoT channels, and a data collection platform. Each testing device connects to a nearby monitored internet line and actively observes its quality. The IoT channels and monitored lines are physically independent, ensuring that even if a monitored line completely fails, the testing devices can still report locally observed anomalies to the data collection platform. The data collection platform aggregates, stores, and visualizes anomaly metrics and line fault warning events from all testing devices.
[0018] When the testing device is powered on for the first time, redeployed, or reconfigured by maintenance personnel, it establishes a communication link with the system management platform and downloads the testing configuration required by the device. The testing configuration is organized in a structured format, and its format does not constitute a limitation on this method. After the testing configuration is loaded, the testing device persists it to local non-volatile storage media. Even if there is a brief loss of connection with the system management platform during subsequent operation, the testing device can still autonomously complete the entire process of active testing, anomaly measurement, cross-protocol linkage intensiveization, and fault early warning based on the locally retained testing configuration, ensuring that the monitoring function is not interrupted due to short-term unavailability of the control plane.
[0019] The configuration for dial-up testing includes: the set of dial-up targets, the set of dial-up protocols, the protocol-target adaptation relationship, the probing time base intensity, the dense time base intensity, the instantaneous deviation multiple, the dense observation window length, the alarm count threshold, the number of calibration dial-up tests, the preset minimum baseline value, the preset calibration failure baseline value, the preset packet loss penalty value, the preset timeout penalty value, and the preset error penalty value. The following is a set of example values, provided as an implementation method: probing time base intensity is 0.5, dense time base intensity is 2, instantaneous deviation multiple is 5, dense observation window length is 30, alarm count threshold is 3, the number of calibration dial-up tests is 20, the preset minimum baseline value is 10, the preset calibration failure baseline value is 1000, the preset packet loss penalty value is 3000, the preset timeout penalty value is 5000, and the preset error penalty value is 10000. The units for detection time base intensity and dense time base intensity are times / second, the unit for dense observation window length is seconds, the instantaneous deviation multiple and alarm count threshold are dimensionless, and the units for the remaining values involving time consumption and baseline are milliseconds.
[0020] Each item in the target set is an accessible object on the Internet, called an Internet target, which covers two or more of the following: web servers, server TCP ports, devices that support the ICMP protocol, and public DNS servers.
[0021] Web servers on the internet, typically including those provided by major portal websites, content service websites, e-commerce websites, and search engines. As an example, this could include... www.baidu.com , www.taobao.com , www.qq.com , www.sina.com.cn Choose publicly accessible web servers with high historical availability. Prioritize these targets to ensure that even if the target itself undergoes short-term maintenance, the assessment related to the local internet connection will not be affected by fluctuations in the maintenance of a single target.
[0022] On the Internet, server TCP ports are specified in the form of IP address plus port number. Examples include port 80 of 180.76.76.76, port 53 of 114.114.114.114, and port 443 of 180.184.1.1. These are often used to evaluate the end-to-end connection capability based on the TCP protocol.
[0023] Devices on the Internet that support the ICMP protocol are designated by a single IP address. Examples include publicly accessible IP addresses such as 8.8.8.8, 114.114.114.114, and 180.76.76.76. Because some ISPs or target peers may employ QoS rate limiting or dropping policies for ICMP packets, the results of ICMP testing should be corroborated with results from other protocols and should not be used as the sole basis for judging line quality.
[0024] The public DNS server is specified in the form of "DNS server IP address plus the domain name to be resolved". For example, it could include using 114.114.114.114 for resolution. www.baidu.com Use 8.8.8.8 for parsing. www.example.com Combinations such as... To avoid DNS resolution results being masked by caching, in one optional implementation, a random string prefix can be added to the domain name to be resolved (e.g., ... abcd1234.www.baidu.com This ensures that each query triggers a complete recursive parsing path.
[0025] Limiting the target set to at least two categories is to effectively distinguish local line faults from remote target faults through cross-type observation consensus, and to avoid misjudging maintenance fluctuations of a single type of target as local line faults. In one implementation, the total size of the target set can be between 8 and 12.
[0026] Each item in the probe test protocol set is a specific method of active probe test, called a probe test protocol, covering two or more of the following: ICMP-based dead / live detection, HTTP / HTTPS-based web application response speed detection, TCP connection-based network application response speed detection, and DNS-based domain name resolution response speed detection.
[0027] Based on ICMP liveness / death detection, the testing device sends ICMP Echo Request messages to the target internet connection and keeps track of the time. When an Echo Reply is received within the protocol layer's waiting limit, the round-trip time is recorded as the elapsed time. If no reply is received, it is recorded as a timeout. When an indicative message such as Destination Unreachable is received, it is recorded as an error.
[0028] For HTTP / HTTPS-based web application response speed testing, the testing device sends a GET request to the web server and records the time taken from sending the first byte of the request to receiving the first byte of the response. This time includes the cumulative impact of multiple stages such as connection establishment, TLS handshake, and remote server processing. When the response status code is in the 5xx range or an exception such as TCP reset occurs, it is recorded as an error. When no response is received, it is recorded as a timeout.
[0029] The network application response speed test based on TCP connection involves the testing device initiating a three-way handshake to the server's TCP port, recording the connection establishment time from the sending of SYN to the receipt of SYN+ACK, and immediately sending an RST packet to close the connection after the handshake is completed. If the three-way handshake is not completed within the protocol layer's waiting limit, it is recorded as a timeout; if an RST packet is received from the other end, it is recorded as an error.
[0030] DNS-based domain name resolution response speed testing involves sending a query message to a public DNS server and recording the resolution time from the time the query is sent to the time the response is received. If the RCODE field of the DNS response indicates an error such as SERVFAIL or REFUSED, it is recorded as an error; if no response is received, it is recorded as a timeout.
[0031] Using at least two different testing protocols is necessary because ICMP, HTTP / HTTPS, TCP, and DNS reflect different aspects such as IP layer reachability and latency, application layer service response capability, transport layer connection establishment capability, and domain name resolution service health. Cross-evaluation of multiple protocols can avoid blind spots or misjudgments due to rate limiting or abnormality of a single protocol.
[0032] Not every dialing protocol can be used for every type of internet dialing target. DNS-based domain name resolution response speed testing is only meaningful when the internet dialing target is a public DNS server. Forcing a DNS protocol to dial a regular web server will not yield meaningful information related to line quality. Similarly, HTTP / HTTPS-based testing requires the internet dialing target to be a web server, TCP connection-based testing requires an open TCP port, and ICMP-based dead / live testing requires ICMP protocol support.
[0033] If the protocol-target adaptation relationship is not introduced and the Cartesian product of the test protocol set and the test target set is directly used as the binary set, the test device will perform a large number of tests that are known to fail. This will not only consume local line bandwidth, but also inject a large number of failure results that are unrelated to line quality into the abnormal metric stream, thus increasing the probability of false positive alarms.
[0034] The protocol-target adaptation relationship defines the category of Internet probe targets that each probe testing protocol adapts to. One implementation may use the following adaptation relationship: ICMP-based live / dead detection adapts to devices supporting the ICMP protocol; HTTP / HTTPS-based web application response speed detection adapts to web server targets; TCP connection-based network application response speed detection adapts to server TCP port targets; and DNS-based domain name resolution response speed detection adapts to public DNS server targets.
[0035] The adaptation relationship can also be further expanded according to business needs. In one optional implementation, a web server that has both open TCP ports and HTTP services is allowed to be classified into both the web server class and the server TCP port class in the protocol-target adaptation relationship, so that the same web server can be observed independently from the application layer and the transport layer. In another optional implementation, the adaptation relationship can be extended to support custom protocols.
[0036] The testing device, based on the loaded protocol-target adaptation relationship, sequentially filters all combinations of testing protocols and internet testing targets that conform to the adaptation relationship in the testing protocol set and the testing target set. Each combination is denoted as a testing protocol-internet testing target tuple, or simply a tuple. All tuples form a tuple set. Each tuple consists of one testing protocol and one internet testing target, and their combination conforms to the protocol-target adaptation relationship.
[0037] For example, suppose a certain probe testing device's probe target set includes 3 web servers, 2 server TCP ports, 4 devices supporting the ICMP protocol, and 2 public DNS servers. The probe testing protocol set includes the above 4 probe testing protocols. Using the above protocol-target adaptation relationship, after filtering according to the adaptation relationship, we get 3 HTTP / HTTPS protocol tuples, 2 TCP protocol tuples, 4 ICMP protocol tuples, and 2 DNS protocol tuples, for a total of 11 tuples.
[0038] The binary tuple is the smallest scheduling object for all subsequent logic: each active dial test targets a certain binary tuple, initiates a dial test to its specified Internet dial test target according to its specified dial test protocol, and obtains an abnormal measurement value. Subsequent logic such as instantaneous deviation judgment, cross-protocol linkage intensification, cumulative deviation counting, and line fault early warning also uses the binary tuple as the basic object.
[0039] The raw results of proactive dial-up tests are divided into two categories: time-consuming and failure-related. Time-consuming results include: ICMP tests obtain round-trip time; HTTP / HTTPS tests obtain the time from sending the first byte of the request to receiving the first byte of the response; TCP tests obtain the connection establishment time from sending the SYN signal to receiving the SYN+ACK signal; and DNS tests obtain the resolution time from sending the query to receiving the response. These results are continuous positive real numbers in milliseconds; a larger value indicates a slower response, i.e., a worse line condition.
[0040] Failure results can be categorized into three types: First, packet loss, where the testing device sends a request but receives no response and a timeout cannot be determined at the protocol stack level; second, timeout, where the testing device sends a request but receives no response within the protocol layer's preset waiting limit; and third, error, where the testing device receives a response, but the status code or return code indicates a protocol layer error. Failure results themselves do not carry a timeout value; they are qualitative failure markers.
[0041] To ensure that subsequent decision-making logic can handle the two types of results in a uniform numerical manner, a normalization rule for anomaly metrics is introduced. This rule stipulates that anomaly metrics are positive real numbers, and the larger the value, the worse the line condition. The normalization rule is as follows: When the response time, connection establishment time, or resolution time is obtained during proactive dialing, the corresponding time value is used as the anomaly metric for that proactive dialing, in milliseconds.
[0042] When packet loss, timeout, or error results are obtained during active dial-up testing, the preset packet loss penalty value, preset timeout penalty value, or preset error penalty value are used as the anomaly metric for that active dial-up test, respectively. The three preset penalty values are manually preset, and for example, they can be set to 3000, 5000, and 10000, respectively, all in milliseconds.
[0043] The three preset penalty values satisfy the following two constraints: First, any preset penalty value is no less than 10 times the typical time consumption of the vast majority of tuples in the set, ensuring that the abnormal metric value generated by any packet loss, timeout, or error has an instantaneous deviation of at least 10 times from the baseline reference value, stably triggering subsequent judgments. Second, there is a strict order among the three preset penalty values: the preset error penalty value is greater than the preset timeout penalty value, and the preset timeout penalty value is greater than the preset packet loss penalty value. ,in This indicates the preset packet loss penalty value. This indicates the preset timeout penalty value. This indicates the preset error penalty value. This order ensures that after the three failure results are normalized to the same type of numerical variable, they still retain their respective severity information: error results usually correspond to the most severe failures with explicit error indications at the protocol layer, hence the highest penalty value; timeouts usually correspond to a complete lack of response for an entire waiting period, with the next highest severity; packet loss usually refers to a single transmission that did not meet expectations, with the lowest relative severity.
[0044] In one alternative implementation, the specific values of the three preset penalty values can also be dynamically determined in relation to the initial baseline value, but regardless of the specific values, they should all satisfy the above two constraints.
[0045] Each tuple corresponds to a specific local internet line-testing protocol-internet testing target observation path. Its typical anomaly metric during normal operation depends on multiple factors, including the local line, backbone network congestion level, the load of the internet testing target, and the protocol layer overhead of the selected testing protocol. These differences can cause significant variations in the typical latency of different tuples, ranging from orders of magnitude: a tuple accessing a local device supporting the ICMP protocol within the same city might have a typical latency of 10 to 30 milliseconds, while a tuple accessing an overseas web server across continents might have a typical latency of 150 to 400 milliseconds.
[0046] Applying the same absolute time threshold to all pairs presents a dilemma: a low threshold means that the normal time of cross-continental pairs already exceeds the threshold, resulting in continuous false alarms; a high threshold means that even if local pairs experience several times the degradation, they will not be triggered. Therefore, for each pair, an initial baseline value reflecting its typical normal time is calculated based on its own calibration data. This value is then used as a reference to determine instantaneous deviations, unifying the judgment logic to a degradation factor relative to its own normal value.
[0047] The calibration process is performed after the testing equipment completes the testing configuration loading and the generation of the binary set, but before entering the formal monitoring stage. The testing equipment sequentially performs a calibration test on each binary set an equal number of times as the calibration test count. Each test yields an anomaly metric value according to the aforementioned normalization rules, and these values are arranged in chronological order to form the calibration anomaly metric value sequence for that binary set. The number of calibration tests is preset manually and is a positive integer greater than 1; in one implementation, it can be 20.
[0048] In engineering practice, a single calibration test may result in packet loss, timeout, or erroneous results due to unforeseen circumstances. For example, a calibration anomaly metric sequence for a binary tuple might consist of 20 elements: 25, 28, 31, 27, 30, 26, 29, 28, 5000, 27, 31, 26, 30, 28, 27, 32, 28, 26, 31, 27. The 9th test might receive a preset timeout penalty of 5000 due to an unforeseen timeout. If the average of this sequence is directly taken, it would be approximately 275, significantly deviating from the true normal latency level of the binary tuple. Subsequent calibration tests using this excessively high average as a baseline would result in significant degradation without triggering false negatives.
[0049] To avoid the aforementioned distortion, the initial baseline value is calculated using the following two-step method.
[0050] First, remove all elements from the calibration anomaly metric sequence that are equal to the preset packet loss penalty, preset timeout penalty, and preset error penalty. These elements are generated by occasional packet loss, timeout, and errors during the calibration phase and should not be included in the estimation of the normal time consumption of the binary tuple. The sequence after removal is called the effective calibration anomaly metric sequence for the binary tuple. Continuing with the previous example, the effective calibration anomaly metric sequence after removal has a total of 19 elements.
[0051] Secondly, the median of the valid calibration anomaly measurement sequence is taken as the initial baseline value of the pair. Continuing with the previous example, the 19 elements are arranged in ascending order, and the median is 28. Therefore, the initial baseline value of the pair is 28.
[0052] The reason for choosing the median instead of the mean is that even after removing elements equal to the preset penalty value, there may still be some slightly oversized samples remaining in the effective calibrated abnormal measure value sequence. These samples make the mean significantly higher, while the impact on the median is negligible. Therefore, the median makes the estimate of the initial baseline value closer to the normal and typical level.
[0053] There are two boundary cases that require additional handling.
[0054] The first boundary case: The effective calibration anomaly metric sequence is empty, meaning that all tests on this tuple during the calibration phase resulted in packet loss, timeout, or error results, making it impossible to calculate the median. To address this, a preset calibration failure baseline value is introduced and used as the initial baseline value for this tuple. This preset calibration failure baseline value is manually preset; for example, 1000 milliseconds can be used. Its value must be greater than the preset minimum baseline value and less than any preset penalty value, ensuring that it is not truncated by the lower limit while still generating a sufficient instantaneous deviation proportion to trigger subsequent judgments when the tuple subsequently experiences another failure.
[0055] The second boundary case: The calculated initial baseline value is too low, falling into a very small range close to 0. For local area networks or data center scenarios, the normal round-trip time and its median may only be 0.3 to 1 millisecond. In this case, even if the subsequent response time is still within the normal range, the relatively low median is already several times lower, which may frequently trigger the judgment even before the line has actually deteriorated. Therefore, a preset minimum baseline value is introduced. If the calculated initial baseline value is less than the preset minimum baseline value, then the preset minimum baseline value is used as the initial baseline value for the binary tuple. The preset minimum baseline value is manually preset and greater than 0; for example, 10 milliseconds can be used.
[0056] Combining the handling of the two boundary cases described above, the calculation of the initial baseline value can be formally expressed as follows. Let the first... The sequence of calibrated outlier measures for each binary tuple is as follows: ,in Indicates the first The first of the two pairs The abnormal measurement value obtained from the second calibration test. This indicates the number of calibration tests. Let the set of preset penalty values be... The meanings of each symbol are the same as above. A sequence of valid calibrated outlier metrics for each binary tuple Defined as
[0057] in Indicates the first A sequence of valid calibrated outlier measures for a pair of tuples. It means "does not belong to".
[0058] No. The initial baseline value of each binary tuple Calculate using the following formula:
[0059] in Indicates the first The initial baseline values of the two tuples This operation represents the median of a sequence (if the number of elements in the sequence is even, the arithmetic mean of the two middle elements is taken). This means taking the larger of the two values. This indicates the preset minimum baseline value. This indicates the preset calibration failure baseline value. Represents the empty set. and These represent "not equal to" and "equal to", respectively. and Between satisfy This ensures that the initial baseline values under both boundary conditions are not lower than the preset minimum baseline value, thereby maintaining the physical meaning of subsequent baseline reference values across all pairs.
[0060] In one alternative implementation, after completing one calibration and entering the formal monitoring phase, the maintenance personnel can reissue the dial-up configuration and trigger one recalibration through the system management platform based on the long-term drift of the line quality, so that the initial baseline value is updated to follow the changes in the long-term quality level of the line.
[0061] After completing the loading of the dial test configuration, the generation of the binary set, and the calibration of the initial baseline value, the dial test device performs the following two initializations.
[0062] The operating mode identifier is initialized to probe mode. The operating mode identifier characterizes the current operating phase of the testing equipment, and its value is limited to two modes: probe mode and intensive mode. Probe mode corresponds to the normal operating phase, where the testing equipment actively probes each binary pair at a lower frequency determined by the probe time base intensity. Intensive mode corresponds to the high-frequency observation phase after a transient anomaly is observed on a particular binary pair. No transient anomalies have been observed at system startup, therefore the operating mode identifier is initialized to probe mode.
[0063] The suspicious pair set is initialized to empty. This set records all pairs that have triggered momentary deviations during the dense observation window. Each recorded pair maintains its own cumulative deviation counter, and subsequent determination of line fault warning events is based on whether the counter value reaches the alarm count threshold. No pairs trigger momentary deviations at system startup, therefore it is initialized to empty.
[0064] After completing the above two initialization steps, the testing equipment enters the formal monitoring phase and proceeds to the execution of subsequent steps.
[0065] After completing step 1, each tuple has an initial baseline value that reflects its normal typical time consumption. The testing equipment enters the formal monitoring stage. Each tuple is repeatedly and actively tested according to its own independent Poisson time process. The abnormal measurement values generated are continuously reported to the data collection platform through the Internet of Things channel.
[0066] During the formal monitoring phase, the testing and scheduling adopts a Poisson time process rather than fixed-interval polling. Various types of interference on internet lines have their own periodicity, such as operators periodically issuing QoS policy updates, network devices periodically batch processing logs, and session keep-alive mechanisms on cross-border paths periodically releasing idle connections. If the initiation time of fixed-interval polling happens to be synchronous or nearly synchronous with the frequency of a certain periodic interference, the interference will be continuously observed or continuously avoided. The resulting abnormal metrics will not accurately reflect the true quality distribution of the line over time, i.e., sampling aliasing.
[0067] The events in a Poisson time process occur independently and lack a predictable periodic structure, thus failing to induce phase locking for various periodic interferences on the line. Any segment of length is... The expected number of dial-up tests occurring within a given time period is [value]. ,in Indicates the length of the time period under consideration, in seconds; This represents the intensity parameter of the Poisson time process, expressed in times per second. Due to the memoryless nature of the Poisson process, the time interval between any two adjacent measurements follows a parameter... The exponential distribution of has a probability density function of .
[0068] in This indicates the time interval between two consecutive dialing tests, in seconds. This represents the natural constant (with a value of approximately 2.71828); Indicates time interval The probability density of the Poisson time process. The Poisson time process also possesses the PASTA property (Poisson Arrivals See Time Averages), meaning that the samples drawn from the time axis by the Poisson process unbiasedly reflect the time mean of the observed system in a long-term statistical sense. Therefore, the sample mean, sample quantiles, and sample median of the anomaly metrics obtained from active probing driven by the Poisson time axis can all serve as unbiased estimates of the true quality of the line within the corresponding time period. This is something that fixed-time-interval polling cannot guarantee when there is periodic interference with the line.
[0069] To achieve dialing and scheduling of the Poisson time process, the inverse transform method of the exponential distribution is employed. This is based on the following: If... It is obedience For a random variable that is uniformly distributed over an interval, then That is, it follows the parameter. The random variable is distributed exponentially, where Represents a uniformly distributed random variable. Represented by natural constant Logarithmic function with base 0. The meaning is the same as above. When generating the next test time, the testing equipment first calls the local uniformly distributed random number generator to extract a random number. , The range of values is limited to (in (This indicates the specific value drawn in this lottery; excluding 0 is to avoid...) Numerical anomalies caused by the absence of a given value; if the value returned by the random number generator falls within... If the interval (i.e., includes 0 but does not include 1) is such that when a 0 is drawn, either the interval is redrawn until a non-zero value is reached, or the 0 is replaced with a 1 to avoid this issue. These two approaches are statistically equivalent. Then proceed as follows...
[0070] Calculate the time interval between the next dialing time and the current time. The meanings of the symbols are the same as above. The next dialing time is set to the current time plus... When the real time advances to the next dialing time, the dialing device immediately performs an active dialing test, and after the test is completed, it re-extracts random numbers and calculates the new next dialing time using the above method, and so on.
[0071] Uniformly distributed random numbers The source can be a pseudo-random number generator with a sufficiently long period and good distribution quality, or a random number interface based on the operating system's entropy pool. Each pair of tuples maintains its own independent random number sequence and internal state for both the exploration time axis and the dense time axis, so that the probe times between different pairs and between the two Poisson time axes of the same pair are statistically uncorrelated, avoiding the concentrated initiation of probes by all pairs within the same short time window.
[0072] Each tuple has two Poisson timelines: an exploration timeline and a dense timeline. For example... Figure 1 As shown, Figure 1 Taking a binary pair as an example, the exploration time axis and the dense time axis of the binary pair are drawn side by side from top to bottom on the same time axis. The horizontal axis is the time in seconds, and each vertical short line on the axis represents one active dialing on that time axis. Figure 1 The time axis is divided into two segments: the exploration mode on the left and the dense mode on the right, with the trigger moment as the boundary. In the exploration mode segment, only the exploration time axis has probes, while the dense time axis is indicated by a dashed line as it is in a pending state. In the dense mode segment, probes are distributed on both the exploration and dense time axes, with the probes on the dense time axis being significantly more concentrated, indicating that the dense time base intensity is greater than the exploration time base intensity. Figure 1 At the bottom, a dense observation window, with a length equal to the length of the dense observation window, is marked with a double-headed arrow, starting from the moment of triggering. The intensity parameter of the exploration time axis is the exploration time base intensity, and the intensity parameter of the dense time axis is the dense time base intensity. The exploration time base intensity is denoted as... Dense time base intensity is denoted as The units are all times per second, where the subscript 'p' is taken from the first letter of 'probe' and the subscript 'd' from the first letter of 'dense'. The intensity of the dense time base is greater than that of the probe time base, i.e. Following the values used in one implementation method of step 1, , That is, on average, one active dialing test is triggered every 2 seconds on the exploration timeline, and on average, one active dialing test is triggered every 0.5 seconds on the dense timeline.
[0073] The reason for configuring two independent Poisson timelines for each pair, instead of switching between exploration and dense modes using the intensity parameters of a single Poisson timeline, is as follows.
[0074] Firstly, the overall measurement process formed when two time axes work in parallel can be directly characterized by the superposition of Poisson processes: two independent time axes with different intensities. and The superposition of Poisson processes still results in a single Poisson process, whose intensity parameter is equal to... Using the example parameters, the instantaneous dialing frequency for each tuple in dense mode naturally increases to [value missing]. Hours / second; within the dense observation window length of 30 seconds as specified in step 1, the average cumulative number of observations obtained is... The number of observation samples is sufficient to support subsequent statistical judgments and is significantly higher than that obtained in the same period of time under the exploration mode. One sample.
[0075] Secondly, the work of retaining the exploration timeline in the dense mode ensures that the sample stream on the exploration timeline remains continuous and comparable before and after the mode switch, which facilitates the comparison of the three time periods before, during and after the event with the same caliber.
[0076] Third, the design of two independent time axes makes the activation of the dense mode equivalent to superimposing a new high-frequency observation on the existing low-frequency observation, thereby avoiding the generation of transitional samples at the moment of switching.
[0077] When the working mode is identified as exploration mode, the exploration timeline continuously generates the next probe time using the inverse exponential distribution transform method and triggers an active probe upon arrival. The dense timeline is in a pending state, meaning it does not extract the next probe time, does not trigger active probes, and does not consume associated resources. The pending state is equivalent to being initialized but not yet started, rather than being uninitialized; that is, the intensity parameters of the dense timeline... The associated data, such as the identifier of the binary tuple to which it belongs and the internal state of the corresponding pseudo-random number generator, are always retained and are only actually started the moment the working mode identifier switches to dense mode.
[0078] When the working mode is set to Intensive Mode, both Poisson timelines are active simultaneously. The Exploration timeline continues from the next test time set before Intensive Mode was enabled, without being re-sampled due to mode switching. The Intensive timeline, on the other hand, uses the instant of mode switching as the reference time, immediately extracts random numbers using the inverse exponential distribution method, and calculates its first next test time. Thereafter, the next test time is regenerated using the same method after each test is completed, until the working mode is switched back to Exploration Mode.
[0079] When the next probe time arrives on any Poisson timeline, the probe device initiates an active probe to the internet probe target specified by the tuple to which that timeline belongs, according to the probe protocol specified by that tuple. The result is mapped to an anomaly metric value according to the normalization rule described in step 1. In subsequent processing, the anomaly metric value is only bound to the tuple that generated it, and not to the exploration timeline or dense timeline. Subsequent instantaneous deviation determination and cumulative deviation count are both based on the tuple as the aggregation unit.
[0080] The scheduling implementation can employ a priority queue based on a min-heap: The next probe time for each active Poisson timeline is inserted into the min-heap as a priority key. When the time at the top of the heap is reached, the scheduling thread pops the top element and triggers an active probe. After completion, a new next probe timeline is generated and re-inserted into the heap. Let the total number of active Poisson timelines be... ,in If the integer is positive, then the time complexity of each scheduling operation is O(n). Another alternative implementation involves creating a separate software timer for each activity's Poisson timeline, which is then centrally scheduled by the operating system. This approach is simple to implement, but when... When the number of timer entries is large (e.g., exceeding 1000), too many entries will increase kernel scheduling overhead, making it more suitable to... For smaller deployment scenarios, when hardware resources are sufficient, a separate user-space thread can be started for each Poisson timeline to wait until the next probing time.
[0081] One type of boundary case that needs to be handled during scheduling is when the actual time of a single active probe exceeds the interval between the probe time and the trigger time of the previous probe on the current Poisson timeline, and a very small value is drawn on that timeline. This can occur when the target internet connection deteriorates significantly. One implementation strategy is to continue testing based on the actual completion time, that is, after the current test ends, the next test result is obtained using the inverse exponential distribution transformation method. And newly drawn The timing is calculated from the end of the initial test. Under this strategy, at most one active test is in progress at any given instant on the same Poisson timeline, avoiding overload on the internet testing target. The trade-off is a slight downward bias in the instantaneous testing frequency compared to the theoretical intensity. At a typical value of times / second, the impact of this downward bias on line quality statistics is negligible in engineering. Another alternative implementation allows multiple incomplete active tests to exist simultaneously on the same Poisson timeline. Each time the next test is scheduled, a new active test is triggered independently, regardless of whether the previous test was completed. This approach is closer to the ideal definition of the Poisson process, but it places higher demands on the concurrent resources of the testing equipment.
[0082] Upon the generation of each abnormal metric, the testing device immediately or after a short buffer period reports the abnormal metric to the data collection platform via the IoT channel. The reported payload shall include at least the following fields: the identifier of the tuple that generated the abnormal metric, the identifier of the Poisson timeline to which it belongs, the initiation time of the active testing, the working mode identifier at the moment the active testing was initiated, and the normalized abnormal metric value.
[0083] The physical transmission path of the IoT channel is independent of the monitored internet line, which is a key guarantee that the distributed monitoring system can still report anomalies even in the most severe fault scenarios. In a typical implementation, the testing device integrates a cellular IoT module that accesses the data collection platform through the cellular operator's network. Even if the monitored line is completely interrupted due to any link such as the access gateway, local fiber optic cable, or in-home wiring, the cellular IoT channel can still report the fault as an anomaly metric, avoiding the inability to inform maintenance personnel when the fault is at its most severe.
[0084] The choice of application layer protocol for reporting does not constitute a limitation on this method. It can be implemented using a publish-subscribe model based on the MQTT protocol, a reporting method based on the HTTPS protocol, or a lightweight reporting method based on the CoAP protocol.
[0085] To balance real-time reporting and transmission efficiency, two timing strategies can be adopted: immediate reporting one by one or batch buffered reporting. The former has the lowest end-to-end latency, while the latter packages multiple abnormal metrics accumulated in the buffer into a single transmission to reduce protocol overhead and power consumption. Moreover, it does not affect the timeliness of alarm judgment when the buffering time is much shorter than the length of the dense observation window.
[0086] The IoT channel itself may also be temporarily unavailable. In one alternative implementation, the testing device is configured with a local circular buffer. During the period when the channel is unavailable, newly generated abnormal metric values are written cyclically, and after the channel is restored, they are resent in the original order of their generation time.
[0087] After anomaly metrics are reported to the data collection platform, the platform assigns them to the corresponding time series of the binary tuples based on fields such as the binary identifier and initiation time of each record. This data is then used by the cross-protocol linkage intensification and fault prediction logic in step 3. Anomaly metrics generated by the same binary tuple on both the exploration time axis and the intensification time axis are merged into a single time series for processing in subsequent instantaneous deviation determination.
[0088] All pairs in the set continuously generate abnormal metrics under the scheduling of the aforementioned dual-time-base Poisson time process and report them to the data collection platform via the IoT channel. However, a single isolated abnormal metric cannot be directly used to determine an anomaly. For example, a 100-millisecond timeframe might indicate significant degradation for a pair corresponding to a local device supporting the ICMP protocol, but it might still be within the normal range for a pair accessing an overseas web server across continents. Therefore, the determination of whether an anomaly exists should not be based on the absolute value of the abnormal metric, but rather on the deviation factor of the abnormal metric from the typical normal level of the pair itself. This relative factor is defined as the instantaneous deviation ratio.
[0089] Suppose the currently generated anomaly metric belongs to the first... A set of two tuples, the value of this outlier measure is denoted as . The baseline reference value for this binary tuple is denoted as The instantaneous deviation ratio of this active dialing test according to
[0090] Calculation. Among them, This represents the instantaneous deviation percentage of the active dialing test, and its value is a dimensionless positive real number. This indicates the abnormal measurement value of this proactive dialing test, in milliseconds; Indicates the first Baseline reference values for each pair of bytes, in milliseconds; subscript This represents the index of the tuple in the set of tuples, and its value is a positive integer.
[0091] Baseline reference value It is determined jointly by the initial baseline value and the preset minimum baseline value. Let the first... The initial baseline value set by each tuple during the initialization and configuration phase is The preset minimum baseline value is ,but
[0092] in, Indicates the first The initial baseline values for each pair of tuples are in milliseconds. This represents the preset minimum baseline value, in milliseconds, and has a value greater than 0. This indicates that the larger of the two values is taken. Although the calibration process has already applied a lower limit protection of no less than the preset minimum baseline value to the initial baseline value once, the same lower limit is applied again here as a unified constraint. This is because in some optional implementations, the initial baseline value may be directly injected by the system management platform without the lower limit protection being applied before injection. Applying it again can prevent the baseline reference value from taking a tiny value close to 0, and prevent the denominator of the instantaneous deviation from the ratio from approaching zero, which would lead to an excessively large value.
[0093] The reason why the instantaneous deviation ratio is based on and The ratio, rather than the difference, is used because differences are not comparable between different pairs: the same 80-millisecond degradation is 5 times that of a local pair with a normal value of 20 milliseconds, which is clearly abnormal, but less than 1.3 times that of a cross-border pair with a normal value of 350 milliseconds, which is considered normal fluctuation. By using a ratio, all pairs share a single dimensionless instantaneous deviation multiple as a unified threshold, achieving a semantic determination of how many times the degradation is relative to its normal value, without needing to set separate thresholds for each pair.
[0094] Numerical examples are as follows. A given pair corresponds to the ICMP testing protocol, and the internet testing target is a local device within the same city that supports the ICMP protocol. The initial baseline value obtained from calibration is 28 milliseconds, and the preset minimum baseline value is 10 milliseconds. Therefore, the baseline reference value for this pair is... Milliseconds. If the anomaly metric for a single active dial-up test is 168 milliseconds, then the instantaneous deviation percentage for that active dial-up test is... If the obtained abnormal measurement value is the preset timeout penalty value of 5000 (corresponding to a timeout in this dial-up test), then the instantaneous deviation ratio of this active dial-up test is: For example, if a given pair corresponds to a test protocol of HTTP / HTTPS, and the internet test target is a cross-border web server, with an initial baseline value of 380 milliseconds and a baseline reference value of... When an outlier of 800 milliseconds is obtained in a single instance, the instantaneous deviation ratio is... This is within the normal deviation range and does not constitute an abnormality.
[0095] Instantaneous deviation multiple The value is preset manually and is a dimensionless, strictly positive real number greater than 1, following the value used in one implementation method during the initialization and configuration phase. ,in This indicates the instantaneous deviation multiple, which is the threshold multiple for triggering cross-protocol linkage intensification. The following experience can be used as a reference for the value: Even when the local line is running normally without faults, it may experience slight fluctuations of 1 to 2 times due to factors such as background traffic, peer service load, and minor adjustments to cross-domain routing. Such fluctuations should not trigger subsequent judgments. The lower limit should be at least higher than this natural fluctuation range. Meanwhile, The value should not be too high; otherwise, the instantaneous deviation caused by moderate line degradation will not be sufficient to reach the threshold, resulting in missed detections. One engineering-effective method for determining the value is based on the statistical distribution of the effective calibration anomaly measurement sequence obtained from calibration. The ratio is 2 to 3 times the ratio of the 99th percentile to the median of the sequence. This method has a very low false trigger probability in the normal distribution and sufficient sensitivity in typical moderate degradation scenarios.
[0096] When the working mode is identified as exploration mode, each abnormal metric value originates from the completion of an active probe on the exploration time axis of a certain binary tuple. The probe device immediately calculates the instantaneous deviation ratio of that active probe according to the above formula: If This abnormal metric is only reported to the data collection platform via the IoT channel as a regular observation sample and does not trigger any additional actions; if The detection equipment determined that the active detection constituted a momentary anomaly with further observation value, triggering a cross-protocol linkage intensive operation.
[0097] Triggering cross-protocol linkage intensive monitoring indicates that transient anomalies have been observed and the system has entered a high-frequency observation phase to collect more samples, rather than indicating a line fault. This method uses two independent thresholds—the transient deviation multiple and the alarm count threshold—for two levels of judgment: the transient deviation multiple serves as the first-level criterion to filter out the vast majority of natural fluctuations, while the alarm count threshold serves as the second-level criterion, requiring a certain number of transient deviations to accumulate during the intensive observation window before a line fault warning event can be reported. This mechanism ensures that the method will not generate false alarms directly from a single random jitter, nor will it miss early stages of a fault due to overly stringent initial conditions.
[0098] The intensive collaborative effects of cross-protocol linkage, such as Figure 2 As shown, Figure 2 Taking a set of 5 tuples as an example, draw the event timeline of each of the 5 tuples from top to bottom. On the left side of each timeline, mark the corresponding dialing protocol and Internet dialing target of the tuple. The horizontal axis is the time in seconds. Figure 2 The system is divided into two segments: exploration mode and intensive mode, with the trigger moment as the boundary. The active probe that triggered this cross-protocol linkage intensification is marked with a circle and a pentagram on the timeline of the second binary tuple; that is, the moment when the instantaneous deviation ratio of this binary tuple first reaches the instantaneous deviation multiple. From the trigger moment onwards... Figure 2The simultaneous, densely distributed probing events across all five pairs of data points demonstrate the collaborative behavior of cross-protocol linkage intensification, which enables the dense timelines of all pairs in the set to synchronously regenerate the next probing event moment with the trigger instant as a reference. At the moment of triggering, the following actions are executed atomically. "Atomization" means that these actions are not interrupted by newly generated abnormal metrics on other Poisson timelines during execution, thus ensuring that the boundaries of the dense observation window remain consistent with the initial state of the suspicious pair set.
[0099] The first action is the switching of the working mode identifier: the working mode identifier is switched from exploration mode to intensive mode.
[0100] The second action is to register the tuple that triggers cross-protocol linkage intensive processing: register the tuple in the suspicious tuple set, assign it a cumulative deviation counter and initialize it to 1 instead of 0. This tuple is the first tuple whose instantaneous deviation ratio reaches the instantaneous deviation multiple during the intensive observation window, and its deviation count starts from 1; when the alarm count threshold is a small integer, setting the count starting point to 1 ensures that the first deviation is counted, guaranteeing the timeliness of alarm judgment.
[0101] The third action is a unified resampling of the dense timelines of all binary pairs: taking the instant when the cross-protocol linkage intensification is triggered as the reference time, a random number is immediately resampled using the inverse exponential distribution transformation method for the dense timelines of all binary pairs in the binary pair set. ,according to
[0102] Calculate the time interval between the next dialing time and the reference time, and set the next dialing time of this dense time axis to the reference time plus... .in, This represents a random number independently drawn from a uniform distribution between 0 and 1, with a value range of 1. ; Represented by natural constant A logarithmic function with base 0; Indicates the intensity of dense time base, measured in times per second; This indicates the time interval between one measurement time and the reference time under this dense time axis, in seconds; This represents the natural constant, with a value of approximately 2.71828. The random number drawn from the dense timeline for each pair of tuples. They are independent of each other, meaning that the dense timeline of the tuple that triggers cross-protocol linkage is also simultaneously re-extracted once, instead of using any "existing" next dialing time.
[0103] All pairs of tuples are uniformly resampled at the moment of triggering. This is because the dense timeline is in an unavailable state during the exploration mode, and its next test time does not exist. Furthermore, it ensures that the expected number of tests for each pair on the dense timeline within the dense observation window is [value missing]. Using the example parameters , The time is 60 times, of which This indicates the length of the dense observation window, in seconds. This uniform resampling ensures that the dense observation window itself has clear boundaries and repeatable statistical significance, preventing the "accidental sampling of a larger value" from occurring on the dense timeline of a particular pair. "And there is a random bias that there are no samples in the first half of the window."
[0104] The fourth action is the activation of the dense observation window: the dense observation window takes the moment when the cross-protocol linkage density is triggered as the starting point, and adds the length of the dense observation window to the starting point. The subsequent moment serves as the end point of the window. Specifically, the upstream testing device starts a software timer, and the timer's expiration time is the end point of the window. Before the expiration, the window is in a state of intensive observation, and abnormal measurement values are processed according to the intensive mode logic. When the window expires, the judgment and switching actions at the end point of the window are executed.
[0105] During periods when the working mode is identified as intensive mode, the exploration time axis of all pairs in the pair set is triggered in parallel with the intensive time axis, and the instantaneous detection frequency of each pair is [missing information]. Using the example parameters times / second, of which This indicates the detection time base intensity, measured in times per second. Each time an abnormal measurement value is generated during an active probe, the probe device... Immediately calculate the instantaneous deviation ratio of this active dialing test; the meanings of each symbol are the same as above. At this time, the abnormal metric does not trigger any additional actions; In such cases, the following two situations should be handled separately.
[0106] In the first scenario, the binary pair that generates the instantaneous deviation ratio has been registered in the suspicious binary pair set during the current intensive observation window. In this case, the testing device increments the cumulative deviation counter corresponding to the binary pair by 1.
[0107] In the second scenario, the pair that exhibits the instantaneous deviation ratio has not yet been registered in the suspicious pair set during the current intensive observation window; that is, this is the first time the pair has appeared during the current intensive observation window. At this point, the testing device registers the binary pair that generates the instantaneous deviation ratio into the suspicious binary pair set, assigns a new cumulative deviation counter to the binary pair, and initializes the new cumulative deviation counter to 1.
[0108] The distinction between the two scenarios above can be achieved in engineering implementation using a hash table with a binary identifier as the key and a cumulative deviation counter as the value: each time an instantaneous deviation ratio is generated and The hash table is queried using a tuple identifier. If a match is found, the corresponding value is incremented by 1; otherwise, a new entry is inserted and its value is set to 1. Hash table queries, insertions, and updates all have constant time complexity, providing good real-time performance in high-frequency testing scenarios. In optional implementations where memory usage is more critical, a bitwise tagging approach combined with a compact counting array can be used to directly locate the counter position using the tuple index.
[0109] The suspicious tuple set records all tuples that have triggered instantaneous deviations during the current window from the coverage dimension, while the cumulative deviation counter records the cumulative number of times each registered tuple has triggered instantaneous deviations from the frequency dimension. The two-dimensional information together support the reliability of subsequent alarm determination.
[0110] Numerical examples are shown below. The example parameters are used as well. , , Second, Alarm count threshold ,in This represents the alarm count threshold, and its value is a positive integer. The process of accumulating the cumulative deviation from the counter during this intensive observation window and the alarm determination at the end of the window are as follows: Figure 3 As shown, Figure 3 The horizontal axis represents the time in seconds since the start of the window, and the vertical axis represents the value of the cumulative deviation counter. Figure 3 The trajectory of the cumulative deviation counter of the binary α is drawn with a solid-line stepped curve as it increments by 1 with each instantaneous deviation. The trajectory of the binary β is drawn with a dashed-line stepped curve. The instantaneous deviation ratio that triggered the count is marked next to each step. Figure 3 The alarm count threshold is marked with a horizontal dashed line, and the position where the counter of the binary tuple α first reaches the threshold is marked with a circle, indicating that although the threshold has been reached at this position, it continues to accumulate in a dense mode until the end of the window for unified judgment. Figure 3 The right side displays the determination results for the tuples α and β at the end of the window. Assume a certain cross-protocol linkage intensification is achieved using tuples... Last time Triggered by instantaneous deviation, the testing device will Registered to the set of suspicious pairs, for Assign a cumulative deviation counter and initialize it to 1. In the subsequent 30-second intensive observation window, assume the events are in the following chronological order: 0.7 seconds after the start of the window, Regenerated above Instantaneous deviation, The counter increments from 1 to 2; 3.4 seconds after the start of the window, above Instantaneous deviation, The counter incremented from 2 to 3; 5.9 seconds after the start of the window, above Instantaneous deviation, The counter increments from 3 to 4; 11.2 seconds after the start of the window, another tuple... First generation Instantaneous deviation, Once registered in the suspicious pair set, its counter is initialized to 1; thereafter until the end of the window, The above will produce 1 more time. Instantaneous deviation, The counter incremented from 4 to 5; no other pairs were generated during this window. The sample. At the end of the window, the set of suspicious pairs is... There are 2 pairs in total. and The cumulative deviation counters for each have values of 5 and 1, respectively.
[0111] When the end time of the dense observation window is reached, the testing device traverses the cumulative deviation counter corresponding to each tuple in the suspicious tuple set, and compares its value with the alarm count threshold. For comparison, handle the following two situations separately.
[0112] In the first scenario, at least one cumulative deviation counter in the suspicious binary set reaches the alarm counting threshold. The suspicious binary set is denoted as... There exists at least one pair of tuples. Make ,in This represents the set of suspicious pairs at the end of this intensive observation window. express The index of a tuple in the middle, express The index is The value of the cumulative deviation counter for the binary tuple. Continuing with the numerical example above, , , Therefore, the condition is met. In this case, the testing equipment determines that at least one tuple corresponds to an Internet line fault warning during this intensive observation window. For each tuple whose cumulative deviation counter reaches the alarm count threshold, a line fault warning event is reported to the data collection platform via the Internet of Things channel.
[0113] The payload of a line fault warning event includes at least the following fields: the identifier of the tuple that triggered the fault warning; the start and end times of this intensive observation window; an ordered list of all abnormal metrics of the tuple during this window and their occurrence times; an ordered list of all instantaneous deviation ratios of the tuple during this window and their occurrence times; and the value of the cumulative deviation counter for the tuple at the end of this window. Optionally, the payload may further carry the same fields of other tuples in the suspicious tuple set that did not reach the alarm count threshold during this window for correlation analysis.
[0114] After reporting a line fault warning event, the working mode identifier remains in dense mode, and the testing equipment immediately starts a new dense observation window: using the start time of the new dense observation window as a reference time, the dense time axis of all tuples is regenerated to generate the next testing time, in the same way as when triggering cross-protocol linkage density; the suspicious tuple set is cleared, and each of the assigned cumulative deviation counters is deleted, so that the new dense observation window starts counting from a blank state.
[0115] The reason for maintaining intensive mode and starting a new window after an alarm, instead of directly reverting to exploration mode, is that the line fault may not have been restored when an alarm is triggered. Continuously maintaining high-frequency observation can collect complete samples during the fault period, which helps to locate the cause of the fault and continuously record the development process of the fault.
[0116] The reason for clearing the suspicious binary set and deleting all cumulative deviation counters when a new window starts, instead of carrying over the statistical results from the current window to the new window, is twofold. Firstly, the alarm count threshold is set based on a length of... The semantics of the cumulative number of deviations during the dense observation window, if carried over, will cause scattered deviations over a long period of time to be incorrectly superimposed into one alarm; secondly, clearing and deleting make each dense observation window completely independent and repeatable in alarm judgment, and also facilitates the retrospective tracing of the sample interval corresponding to each alarm.
[0117] In the second scenario, the cumulative deviation counter values in the entire suspicious binary set are all less than the alarm count threshold, i.e., for all... All The meanings of each symbol are the same as above. In this case, it is determined that no continuous instantaneous deviation was observed during the current intensive observation window, and the current trigger is an isolated instantaneous jitter rather than a line fault. Therefore, no line fault warning event is reported, and the working mode indicator is switched from intensive mode back to exploration mode and the following three actions are performed. First, all tuples of intensive time axes are restored to the ready-to-use state, that is, the next dialing time that has been set but not yet arrived on each intensive time axis is cleared, but the intensity parameter of that intensive time axis is retained. The system retrieves and assigns data such as the identifier of the assigned binary tuple and the internal state of the corresponding pseudo-random number generator, so that it can be directly activated the next time cross-protocol linkage intensification is triggered. Secondly, all binary tuples are restored and active probing is performed only according to their respective exploration timelines. The next probing time on the exploration timeline is not regenerated due to mode switching, so as to maintain the statistical continuity of the sample stream before and after mode switching. Thirdly, the suspicious binary tuple set is cleared, and each assigned cumulative deviation counter is deleted.
[0118] When no alarm is triggered, the probe mode is rolled back instead of continuing to maintain the intensive mode, due to considerations of both resource consumption and target load. In intensive mode, the instantaneous probe frequency for each tuple reaches [a certain value]. Using the example parameter of 2.5 times / second, the number of elements in the set of tuples is denoted as... , This represents the number of pairs in the set of pairs; the total instantaneous dialing frequency of the distributed monitoring system in dense mode is... times / second. Under the deployment, the total instantaneous dialing frequency in intensive mode is At a rate of [number] times per second, maintaining this level for an extended period will continuously consume the resources of both the local internet line and the internet-based target being probed. Switching to intensive mode only when momentary anomalies are observed, and immediately reverting to probing mode when no alarms occur, can keep the average probe overhead of the distributed monitoring system within an acceptable range.
[0119] Several common boundary situations in engineering are explained here, along with their handling methods.
[0120] The first type of boundary case involves deviations from the assigned position near the end time of the dense observation window. Let the end time of the dense observation window be... The time when a certain abnormal metric was generated was ,and and The difference is within a few milliseconds, of which Indicates the end time of the dense observation window. This indicates the time when the abnormal metric value was generated. The time taken for the testing device to complete the corresponding judgment and counting actions (typically several hundred microseconds) may include this deviation in this window or the next window. In one implementation, a strict comparison is made between the time when the abnormal metric value was generated (i.e., the time when the active testing request was initiated) and the endpoint time: Add them to this window. The instantaneous deviation trigger is categorized into the next window or the exploration mode. This categorization method, which is based on the event generation time rather than the event processing time, ensures that the categorization results are independent of the processing delay of the testing equipment itself and remain consistent even when the load of the testing equipment fluctuates.
[0121] The second type of boundary case concerns whether the counter continues to accumulate after reaching a threshold. Whether a single cumulative deviation counter continues to accumulate after reaching the alarm count threshold does not affect the alarm determination result. In one implementation, the cumulative deviation counter continues to accumulate until the end of the window; in another optional implementation, the cumulative deviation counter reaches... Then, the accumulation is stopped to slightly save memory and computational overhead. The two implementation methods are completely equivalent in terms of the correctness of alarm determination.
[0122] The third type of boundary case is that cross-protocol linkage intensification is not repeatedly triggered during the intensive mode. When the working mode identifier is already in intensive mode, if the instantaneous deviation ratio of a certain tuple reaches the instantaneous deviation multiple for the first time, the tuple is only registered in the suspicious tuple set and a cumulative deviation counter initialized to 1 is assigned. The mode switch and resampling of all tuple intensive timelines are not performed again.
[0123] The fourth type of boundary case involves independent determination of continuous alarms. During the newly started dense observation window, if a tuple of cumulative deviation counters reaching the alarm count threshold reappears, a line fault warning event is reported again and a new window is started; if all counters do not reach the threshold, the detection mode is rolled back. The determination of the end point of each dense observation window is executed independently.
[0124] Several alternative implementation methods can further enhance the adaptability to specific scenarios while keeping the core logic of this method unchanged.
[0125] The first optional implementation is to differentiate the alarm counting threshold. In a basic implementation, the alarm counting threshold... All pairs of data are assigned the same value; in more refined alternative implementations, different alarm counting thresholds can be set for different dialing protocols or different Internet dialing target categories, such as setting different alarm counting thresholds for ICMP-based live / death detection. 1. Detect the response speed of HTTP / HTTPS based web applications This is to reflect the differences in jitter characteristics of different dialing protocols. Specifically, dialing based on ICMP is more affected by rate limiting policies and requires a higher number of cumulative deviations, while dialing based on HTTP / HTTPS is less affected by protocol layer factors and a lower number of cumulative deviations is sufficient to make a reliable judgment.
[0126] The second optional implementation method is a two-way determination of the instantaneous deviation ratio. The default trigger condition for this method in the determination direction is... This means the abnormal metric value is significantly greater than the baseline reference value. In special scenarios where a response time significantly lower than the baseline reference value may indicate an anomaly, it may be necessary to issue an alert for a significant increase in speed. In this case, a symmetrical judgment condition can be added. In other words, cross-protocol linkage intensification is also triggered when the abnormal metric value is significantly less than the baseline reference value. The main logic of this method does not need to be changed under this optional implementation.
[0127] The third optional implementation method is a sliding update of the baseline reference value. In scenarios where the line quality drifts slowly over a long period of time, the data collection platform can periodically recalculate the baseline reference value based on a certain quantile of all abnormal measurement values over a past period. After being sent to the testing equipment by the system management platform, the instantaneous deviation judgment continues to be performed. The update period should avoid the period when line fault warning events have been reported.
[0128] The fourth optional implementation is to strengthen the triggering conditions for cross-protocol linkage density. In an optional implementation where a further reduction in false triggering rate is required, the triggering conditions can be strengthened to a certain length of... Within the small window, any one or more pairs of tuples cumulatively appear More than one instantaneous deviation, among which This indicates the length of the small window for enhanced triggering, in seconds, with a typical value of 5. This represents the cumulative threshold for enhanced triggering, and is a positive integer, typically 2. This enhancement condition can significantly reduce the false trigger rate of single-point transient jitter, at the cost of a slight delay in response time to genuine transient faults. The threshold can be adjusted based on the stability level of the deployed lines. and We need to compromise on the value of .
[0129] The fifth optional implementation is redundant transmission of line fault early warning events. The testing device can report via the IoT channel while simultaneously sending the same line fault early warning event redundantly to maintenance personnel or a second data collection platform via a pre-configured backup channel.
[0130] The above process completes the transformation from abnormal metric streams to line fault early warning events: filtering out most natural fluctuations with instantaneous deviation multiples, further filtering out scattered single-point deviations with alarm count thresholds, characterizing the coverage of deviations with a set of suspicious binary pairs, and characterizing the persistence of deviations with a cumulative deviation counter. The entire process does not incorporate machine learning models or rely on long-term historical data storage, and can be executed entirely locally on testing equipment with limited computing power, exhibiting excellent deployability and interpretability.
[0131] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting and managing Internet line faults based on multi-protocol proactive dialing, executed in a distributed monitoring system consisting of dialing equipment, IoT channels, and a data collection platform, including: Step 1: Generate a set of tuples consisting of the dialing test protocol - Internet dialing test target tuple. The tuple is an abbreviation for the dialing test protocol - Internet dialing test target tuple. Set the initial baseline value for each tuple. The working mode identifier is initially set to exploration mode. The suspicious tuple set is initially empty. Step 2: Each tuple maintains two Poisson time axes: an exploration time axis and a dense time axis. The intensity parameters of the exploration time axis and the dense time axis are the exploration time base intensity and the dense time base intensity, respectively, with the dense time base intensity being greater than the exploration time base intensity. In exploration mode, only the exploration time axis triggers active probing, while in dense mode, both the exploration time axis and the dense time axis trigger active probing. Abnormal metrics are reported via the IoT channel. Step 3: When an abnormal measurement value is generated, the instantaneous deviation ratio is obtained based on the abnormal measurement value and the baseline reference value of the binary pair that generated the abnormal measurement value. In the exploration mode, when the instantaneous deviation ratio of any binary pair reaches the instantaneous deviation multiple, cross-protocol linkage intensification is triggered, the working mode identifier is switched to intensive mode, the intensive time axis of all binary pairs in the binary pair set is regenerated for the next batch of measurement time, and the intensive observation window is started. During the intensive observation window, binary pairs whose instantaneous deviation ratio reaches the instantaneous deviation multiple are registered to the suspicious binary pair set, and a cumulative deviation counter is assigned to the registered binary pairs. When the intensive observation window ends, if the value of the cumulative deviation counter reaches the alarm count threshold, a line fault warning event is reported, and a new intensive observation window is started. If all cumulative deviation counter values are less than the alarm count threshold, switch to probe mode.
2. The method as described in claim 1, characterized in that, Internet dialing tests cover at least two of the following: web servers on the Internet, server TCP ports on the Internet, devices on the Internet that support the ICMP protocol, and public DNS servers. The testing protocols cover two or more of the following: ICMP-based live / dead detection, HTTP / HTTPS-based web application response speed detection, TCP connection-based network application response speed detection, and DNS-based domain name resolution response speed detection.
3. The method as described in claim 2, characterized in that, The testing device downloads the testing configuration from the system management platform. The testing configuration includes a set of testing targets, a set of testing protocols, and a protocol-target adaptation relationship. The set of testing targets contains more than two Internet testing targets, and the set of testing protocols contains more than two testing protocols. The protocol-target adaptation relationship defines the category of Internet testing targets that each testing protocol adapts to. Based on the protocol-target adaptation relationship, the testing device generates a set of tuples from the set of testing targets and the set of testing protocols. Each tuple in the set of tuples consists of one testing protocol and one Internet testing target, and the combination of testing protocol and Internet testing target in each tuple conforms to the protocol-target adaptation relationship.
4. The method as described in claim 1, characterized in that, Anomaly metrics are uniformly normalized to a form where the larger the value, the worse the line condition. The normalization rule is: when actively dialing to obtain response time, connection establishment time, or parsing time, the corresponding time value is used as the anomaly metric. When packet loss, timeout, or error results are obtained from active dialing, the preset packet loss penalty value, preset timeout penalty value, or preset error penalty value are used as the anomaly measurement value, respectively. The preset packet loss penalty value, preset timeout penalty value, and preset error penalty value are all greater than 0.
5. The method as described in claim 4, characterized in that, The initial baseline value of each tuple is calibrated as follows: the testing equipment continuously performs calibration tests on each tuple for the same number of calibration tests as the calibration test count, to obtain a calibration anomaly metric sequence for each tuple. The calibration test count is a positive integer greater than 1. Elements corresponding to the preset packet loss penalty value, preset timeout penalty value, and preset error penalty value in the calibration anomaly metric sequence are removed to obtain a valid calibration anomaly metric sequence. When the effective calibration outlier measure sequence is not empty, the median of the effective calibration outlier measure sequence is used as the initial baseline value of the corresponding binary tuple; When the valid calibration anomaly measurement sequence is empty, the preset calibration failure baseline value is used as the initial baseline value of the corresponding binary pair. The value of the preset calibration failure baseline value is greater than the preset minimum baseline value. When the value of the initial baseline value is less than the preset minimum baseline value, the value of the preset minimum baseline value is used as the initial baseline value of the corresponding binary pair. The value of the preset minimum baseline value is greater than 0.
6. The method as described in claim 5, characterized in that, For each pair, the baseline reference value is the larger of the initial baseline value and the preset minimum baseline value; the instantaneous deviation ratio is equal to the abnormal measurement value divided by the baseline reference value; the instantaneous deviation multiple is the threshold multiple for triggering cross-protocol linkage intensification, and the instantaneous deviation multiple is greater than 1.
7. The method as described in claim 1, characterized in that, The next measurement time on the Poisson timeline is generated using the inverse exponential distribution transform method. The generation method is as follows: the measurement device independently draws a random number from a uniform distribution in the interval 0 to 1. The value of the random number is greater than 0 and less than or equal to 1. The time interval between the next measurement time and the current time is equal to the negative of the natural logarithm of the random number divided by the intensity parameter. The unit of the time interval is seconds. The unit of the intensity parameter is times / second. On the exploration timeline, the intensity parameter is the value of the exploration time base intensity. On the dense timeline, the intensity parameter is the value of the dense time base intensity. After each active measurement is completed, the next measurement time on the corresponding Poisson timeline is immediately regenerated in the above manner.
8. The method as described in claim 1, characterized in that, In exploration mode, the dense time axis is in a standby state. When cross-protocol linkage density is triggered, the probe device uses the working mode identifier switching time as the current time and regenerates the next probe time for the dense time axis of all tuples in the tuple set. The intensity parameter used is the dense time base intensity. The duration of the dense observation window is the dense observation window length. During the dense observation window, the exploration time axis and the dense time axis of all tuples in the tuple set trigger active probe in parallel. The instantaneous probe frequency of each tuple is the sum of the exploration time base intensity and the dense time base intensity.
9. The method as described in claim 1, characterized in that, When cross-protocol linkage intensification is triggered, the tuple that triggered cross-protocol linkage intensification is registered in the suspicious tuple set, and a cumulative deviation counter is assigned to the registered tuple, which is initialized to 1; during the intensive observation window, when the instantaneous deviation ratio of the tuple reaches the instantaneous deviation multiple, the registration and counting actions are performed: if the tuple that generated the instantaneous deviation ratio has already been registered in the suspicious tuple set, the corresponding cumulative deviation counter is incremented by 1; If a tuple that generates an instantaneous deviation ratio is the first tuple to have an instantaneous deviation ratio that reaches the instantaneous deviation multiple during this intensive observation window, the tuple that generates the instantaneous deviation ratio will be registered in the suspicious tuple set. One cumulative deviation counter will be assigned to the registered tuple and initialized to 1. The alarm count threshold will be a positive integer.
10. The method as described in claim 1, characterized in that, Line fault warning events include the identifier of the binary tuple that triggered the fault warning, all abnormal measurement values generated by the corresponding binary tuple during the dense observation window, and all instantaneous deviation ratios. Starting a new dense observation window includes: maintaining the dense mode identifier, using the start time of the new dense observation window as the current time, regenerating the next test time for the dense time axis of all binary tuples in the binary tuple set, clearing the suspicious binary tuple set, and deleting each cumulative deviation counter. When switching to exploration mode, the dense time axis of all binary tuples in the binary tuple set is restored to the ready-to-use state, all binary tuples in the binary tuple set are restored to active testing triggered only by the exploration time axis, the suspicious binary tuple set is cleared, and each cumulative deviation counter is deleted.