Equipment dial testing method and device and related equipment

By constructing the device network topology and using the LSTM-Transformer hybrid model to predict the failure probability, the testing strategy is dynamically adjusted, solving the problems of resource waste and delayed fault detection in traditional testing, and achieving efficient network monitoring.

CN121750489APending Publication Date: 2026-03-27CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-05
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In traditional network monitoring, the fixed-period indiscriminate testing mechanism leads to resource waste and delayed fault detection, and cannot effectively respond to changes in network status.

Method used

By acquiring the configuration information and historical alarm data of communication devices, the device network topology is constructed, and the LSTM-Transformer hybrid model is used to predict the failure probability. The probe testing strategy is dynamically adjusted, and the priority and risk value drive the probe testing.

Benefits of technology

It improved testing efficiency, reduced resource waste, and enhanced the timeliness of fault detection and the reliability of monitoring critical business operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750489A_ABST
    Figure CN121750489A_ABST
Patent Text Reader

Abstract

The invention provides an equipment dial testing method and device and related equipment, and relates to the technical field of communication. The method comprises the following steps: acquiring configuration information and historical alarm data of various communication devices; constructing a device network topology based on the configuration information of various communication devices; starting a probe to carry out dial testing on various communication devices according to the initial strategy to obtain a current dial testing result; inputting the historical alarm data of various communication devices, the device network topology and the current dial test result into the hybrid model, and outputting the fault probability of each region in the device network topology; and based on the fault probability of each area, determining a dial testing strategy for each area, and starting a probe to carry out dial testing on communication equipment in each area according to the dial testing strategy of each area. By means of the technical means, the problem that in the related technology, equipment dial testing cannot differentially treat a stable network state scene and an abnormal scene is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of communication technology, and in particular to a device probing method and device and related equipment. BACKGROUND

[0002] In the current large-scale and high-dynamic communication and network environment, the traditional network monitoring generally relies on a fixed-period non-discriminatory probing mechanism, that is, performing a probing task at a uniform time interval for all nodes. Although this method is simple to implement, it has significant defects: on the one hand, it still continues to probe at a high frequency during a period of stable network state, causing a large amount of redundancy of computing, bandwidth and energy resources; on the other hand, due to the lack of the ability to predict the evolution of the network state, it is difficult to cope with abnormal scenarios such as sudden traffic surge or link degradation, resulting in a lag in fault discovery and affecting business continuity.

[0003] SUMMARY

[0004] The present disclosure provides a device probing method, device and related equipment, which improves the probing efficiency.

[0005] According to one aspect of the present disclosure, a device probing method is provided, comprising: obtaining configuration information and historical alarm data of various communication devices; constructing a device network topology based on the configuration information of the various communication devices; starting a probe to probe the various communication devices according to an initial strategy to obtain a current probing result; inputting the historical alarm data of the various communication devices, the device network topology and the current probing result into a hybrid model to output a fault probability of each region in the device network topology; determining a probing strategy for each region based on the fault probability of each region, and starting the probe to probe the communication devices in each region according to the probing strategy for each region.

[0006] In one embodiment of the present disclosure, before starting the probe to probe the communication devices in each region according to the probing strategy for each region, the method further comprises: determining the priority of the various communication devices; and determining the probing strategy for each region based on the fault probability of each region and the priority of the communication devices in the region.

[0007] In one embodiment of the present disclosure, determining the priority of the various communication devices comprises: obtaining the priority of the services carried by the various communication devices; and determining the priority of the various communication devices based on the priority of the services carried by the various communication devices.

[0008] In one embodiment of the present disclosure, determining the probing strategy for each region based on the fault probability of each region and the priority of the communication devices in the region comprises: determining a risk value of each region based on the fault probability of each region and the priority of the communication devices in the region; and determining the probing strategy for each region based on the risk value of each region.

[0009] In one embodiment of this disclosure, a testing strategy for each region is determined based on the risk value of each region, including: for any region: when the risk value of the region is greater than a first preset threshold, the testing interval of the region is shortened based on the initial strategy; when the risk value of the region is less than or equal to a second preset threshold, the testing interval of the region is increased based on the initial strategy, wherein the first preset threshold is greater than the second preset threshold; when the risk value of the region is less than or equal to the first preset threshold but greater than the second preset threshold, the testing interval of the region remains unchanged based on the initial strategy.

[0010] In one embodiment of this disclosure, a testing strategy for each region is determined based on the risk value of each region, including: for any region: when the risk value of the region is greater than a first preset threshold, a first interval is determined based on the difference between the risk value of the region and the first preset threshold, and the testing interval of the region is shortened according to the first interval based on the initial strategy; when the risk value of the region is less than or equal to a second preset threshold, a second interval is determined based on the difference between the risk value of the region and the second preset threshold, and the testing interval of the region is increased according to the second interval based on the initial strategy; when the risk value of the region is less than or equal to the first preset threshold but greater than the second preset threshold, the testing interval of the region is kept unchanged based on the initial strategy.

[0011] In one embodiment of this disclosure, after starting probes to test various communication devices according to an initial strategy and obtaining the current test results, the method further includes: constructing a historical failure rate matrix using historical alarm data of various communication devices; inputting the historical failure rate matrix, device network topology and current test results into a hybrid model, and outputting the failure probability of each region.

[0012] According to another aspect of this disclosure, a device testing apparatus is provided, comprising: an acquisition unit configured to acquire configuration information and historical alarm data of various communication devices; a construction unit configured to construct a device network topology based on the configuration information of various communication devices; a testing unit configured to initiate probes to test various communication devices according to an initial strategy and obtain current testing results; a prediction unit configured to input historical alarm data of various communication devices, device network topology and current testing results into a hybrid model and output the fault probability of each area in the device network topology; and a determination unit configured to determine a testing strategy for each area based on the fault probability of each area, and initiate probes to test communication devices in each area according to the testing strategy for each area.

[0013] According to another aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any of the methods described above by executing the executable instructions.

[0014] According to another aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements any of the methods described above.

[0015] In the embodiments of this disclosure, a device network topology is constructed based on the configuration information of various communication devices. Historical alarm data from various communication devices, the device network topology, and current test results are input into a hybrid model. The model outputs the fault probability of each region within the device network topology. Based on the fault probability of each region, a test strategy is determined for each region. Through the above technical means, the problem in related technologies where device test scenarios cannot differentiate between stable and abnormal network conditions is solved, thereby improving test efficiency.

[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0018] Figure 1 A schematic diagram of a device testing system according to an embodiment of this disclosure is shown.

[0019] Figure 2 A flowchart of a device testing method according to an embodiment of this disclosure is shown.

[0020] Figure 3 A flowchart of a dialing strategy determination method is shown in an embodiment of this disclosure.

[0021] Figure 4 A flowchart of another device testing method in an embodiment of this disclosure is shown.

[0022] Figure 5 A flowchart illustrating a method for predicting the probability of regional failures according to an embodiment of this disclosure is shown.

[0023] Figure 6 A schematic diagram of a device testing apparatus according to an embodiment of the present disclosure is shown.

[0024] Figure 7A schematic diagram of an electronic device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0025] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0026] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0027] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0029] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0030] It should be noted that, unless otherwise specified, the embodiments of this disclosure and the technical features thereof can be combined with each other.

[0031] To facilitate understanding, the following is an explanation of several terms used in this disclosure: Probe protocol: refers to the type of communication protocol used by the probe when performing active probe tasks, including ICMP, TCP, HTTP, etc., to simulate different levels of business access behavior to detect network or service status.

[0032] ICMP (Internet Control Message Protocol): A network layer protocol used to transmit control messages and error reports between network devices, often used to detect host or network connectivity.

[0033] TCP (Transmission Control Protocol): A connection-oriented, reliable transport layer protocol that ensures data is transmitted in order and without errors between communicating parties. It is often used for port reachability detection.

[0034] HTTP (Hypertext Transfer Protocol): An application layer protocol used for clients and servers to request and transfer resources such as web pages and APIs. It is often used to verify the availability and response performance of web services.

[0035] Test interval: The time interval between two consecutive test tasks, used to control the detection frequency of the target communication device or area, and is the core scheduling parameter in the test strategy.

[0036] IP address (Internet Protocol Address): A unique identifier assigned to each communication device in a network. It is used to locate and address devices in the network and is the basic information for building network topology and initiating dial-up tests.

[0037] LSTM-Transformer hybrid model: A time series prediction model that combines Long Short-Term Memory (LSTM) network and Transformer architecture. LSTM is used to capture long-term dependencies in time series, while Transformer is used to model spatial correlations across regions. Together, they are used to predict the failure probability of each region.

[0038] Long Short-Term Memory (LSTM) network: a special type of recurrent neural network structure that effectively alleviates the gradient vanishing problem by introducing a gating mechanism. It can learn and remember temporal dependencies over long periods of time and is suitable for modeling time series such as historical fault data.

[0039] Transformer: A deep neural network architecture based on self-attention mechanism. It can process sequential data in parallel without relying on recurrent structures. It is good at capturing global dependencies between elements within a sequence and has advantages in modeling the spatial correlation of multi-region network states.

[0040] 5G (Fifth Generation Mobile Networks): A new generation of cellular mobile communication technology with high bandwidth, low latency, and massive connectivity. Communication equipment in its core network and radio access network needs to ensure service continuity through intelligent dialing testing.

[0041] Figure 1 This diagram illustrates a device testing system according to an embodiment of the present disclosure. The device testing system includes: The communication system includes communication equipment 101, acquisition module 102, scheduling module 103, and network probe library 104.

[0042] Communication device 101 in the communication system: represents the set of communication devices to be monitored, serving as the target of the test and the source of result feedback.

[0043] Acquisition module 102: Responsible for obtaining configuration information and historical alarm data from communication device 101 and transmitting the data to dynamic scheduling module 103.

[0044] Network probe library 104: Stores probe templates at various protocol levels for the dynamic scheduling module to call to execute probe tasks.

[0045] Dynamic scheduling module 103: consists of multiple sub-modules, including: Network configuration module 1031: used to construct the device network topology and determine the priority of communication devices; Time-series prediction submodule 1032: Combines historical failure rate matrix, device network topology and current test results, and uses a hybrid model to predict the failure probability of each area; Strategy Generator 1033: Calculates regional risk values ​​based on failure probability and communication equipment priority, and generates dialing test strategies for each region accordingly; Policy executor 1034: Based on the generated probe testing policy, it retrieves the corresponding probes from the network probe library 106 and initiates the probe testing task. Specifically, policy executor 1034 also initiates probe testing on various communication devices according to the initial policy, obtaining the current probe testing results.

[0046] Perform a dial test: The policy executor 1034 initiates a dial test request to the communication device 101.

[0047] Return of test results: Communication device 101 returns the test response to strategy executor 1034, forming a closed-loop feedback.

[0048] The device testing system can be equipped with an application program to perform the following actions: acquiring configuration information and historical alarm data of various communication devices; constructing a device network topology based on the configuration information of various communication devices; launching probes to test various communication devices according to an initial strategy and obtaining the current test results; inputting the historical alarm data of various communication devices, the device network topology, and the current test results into a hybrid model and outputting the fault probability of each area in the device network topology; determining the testing strategy for each area based on the fault probability of each area, and launching probes to test the communication devices in each area according to the testing strategy of each area.

[0049] Figure 2 This diagram illustrates a flowchart of a device testing method according to an embodiment of the present disclosure. The method is as follows: Figure 2 As shown, it includes the following steps: S201, obtain configuration information and historical alarm data of various communication devices.

[0050] Exemplary communication equipment refers to physical or logical devices used to implement network communication functions, including routers, switches, servers, wireless access points, and various network elements.

[0051] As an example, configuration information is data that describes the attributes and connectivity of a communication device in a network. Attributes include IP address, port status, link information, and device type, while connectivity includes other communication devices to which the communication device is connected.

[0052] For example, historical alarm data is a time-series information that records faults or abnormal events that occurred in communication equipment over a past period of time, including alarm time, alarm type, and associated device identifiers.

[0053] S202 constructs the device network topology based on the configuration information of various communication devices.

[0054] As an example, a device network topology is a structured graph of the connections between devices. In a device network topology, each communication device is a node, and the information of a node is the attribute of the corresponding communication device.

[0055] S203, according to the initial strategy, the probe is started to perform dial tests on various communication devices and the current dial test results are obtained.

[0056] As an example, the initial strategy is to pre-set uniform dialing test rules for all communication devices, including fixed dialing test protocols, target addresses, and dialing test intervals.

[0057] As an example, a probe is an executable probe unit that encapsulates specific probe protocols and parameters to initiate active probes to communication devices and return probe results.

[0058] As an example, the current probe test results are real-time performance or status indicators returned by the probe after performing the probe test task, including latency, packet loss rate, connectivity status, and service response codes.

[0059] S204 inputs historical alarm data of various communication devices, device network topology and current dial-up test results into the hybrid model, and outputs the fault probability of each area in the device network topology.

[0060] The exemplary hybrid model is a time-series prediction model that combines LSTM and Transformer architectures to fuse historical alarm data, device network topology, and current probing results to predict the probability of failure in each area in the future.

[0061] As an example, the device network topology is divided into multiple regions according to a preset grid size, with each region containing one or more communication devices.

[0062] As an example, the failure probability is a numerical value output by the hybrid model, representing the likelihood of a network failure occurring in a certain area within a preset time window in the future.

[0063] S205, based on the failure probability of each area, determine the dialing test strategy for each area, and start the probe to dial the communication equipment in each area according to the dialing test strategy of each area.

[0064] As an example, a probe strategy for a region is a dynamic probe rule developed for that region, including the type of probe used and the probe interval.

[0065] In this embodiment, configuration information and historical alarm data of the communication equipment are first acquired, and a device network topology is constructed based on the configuration information. Then, probes are initiated to perform dial-up testing according to an initial strategy, and the current dial-up test results are obtained. The historical alarm data, device network topology, and current dial-up test results are input into a hybrid model to obtain the fault probability of each area in the device network topology. Finally, based on the fault probability of each area, a dial-up testing strategy for each area is determined, and probes are initiated to perform dial-up testing on the communication equipment in each area according to this strategy. Through the above technical means, the problem that related technologies cannot distinguish between stable and abnormal network conditions is solved, thereby improving dial-up testing efficiency.

[0066] Figure 3 This diagram illustrates a flowchart of a dialing strategy determination method according to an embodiment of the present disclosure. The method is as follows: Figure 3 As shown, it includes the following steps: S301 determines the priority of various communication devices; S302, based on the failure probability of each area and the priority of the communication equipment in that area, determine the dialing test strategy for each area.

[0067] As an example, the priority of communication equipment reflects the order of processing or priority of guarantees in resource scheduling or fault response.

[0068] In this embodiment, the priorities of various communication devices are determined, and these priorities, along with the region's failure probability, are used as decision factors to generate a testing strategy that better meets business needs. Through these technical means, high-priority communication devices receive higher-density testing support under the same risk level, improving the monitoring reliability of critical services.

[0069] In one optional embodiment, configuration information and historical alarm data of various communication devices are acquired; a device network topology is constructed based on the configuration information of various communication devices; probes are initiated to perform dial-up tests on various communication devices according to an initial strategy, and the current dial-up test results are obtained; the historical alarm data of various communication devices, the device network topology, and the current dial-up test results are input into a hybrid model, and the failure probability of various communication devices in the device network topology is output; for each region in the device network topology, the average failure probability of all communication devices in that region is calculated; based on the average failure probability and the service type weight of the communication devices in that region, a dial-up test strategy for that region is determined. Through the above technical means, the dial-up test strategy can reflect the comprehensive impact of the overall risk level and service attributes of the region, avoiding strategy oscillations caused by relying solely on the anomaly of a single device.

[0070] As an example, the service type weight is a pre-set numerical parameter based on the importance or sensitivity of the service category carried by the communication equipment in the overall service system. It is used to assign differentiated impact to different service types in risk assessment or resource scheduling.

[0071] In one embodiment of this disclosure, determining the priority of various communication devices includes: obtaining the priority of the services carried by various communication devices; and determining the priority of various communication devices based on the priority of the services carried by various communication devices.

[0072] As an example, the priority of services carried by communication equipment is a level identifier set according to the importance of the types of services supported by the communication equipment in the overall service system.

[0073] In this embodiment, when determining the priority of various communication devices, the priority of the services carried by each communication device is obtained and directly mapped to the priority of the communication device, ensuring that the device priority is consistent with the criticality of the service. Through the above technical means, it is ensured that testing resources are tilted towards communication devices carrying high-priority services, enhancing the monitoring coverage of core service links.

[0074] Figure 4 This invention discloses a flowchart of another device testing method according to an embodiment of the present disclosure. The method is as follows: Figure 4 As shown, it includes the following steps: S401, based on the failure probability of each area and the priority of the communication equipment in that area, determine the risk value of each area; S402, based on the risk value of each region, determines the testing strategy for each region.

[0075] As an example, the risk value of a region is a quantitative indicator calculated by combining the failure probability of the region with the priority of communication equipment in the region, and is used to characterize the level of attention that region receives in network monitoring.

[0076] In this embodiment, based on the failure probability of each region and the priority of the communication equipment in that region, the risk value of each region is first generated by fusion, and then the dialing test strategy for each region is determined based on the risk value, so that the strategy decision is based on a unified risk assessment.

[0077] By employing the aforementioned technical means, we can achieve a combined driving force between the probability of failure and the importance of business operations, thereby preventing high-priority areas from being overlooked due to low failure probability, or low-priority areas from excessively consuming resources due to occasional alarms.

[0078] In one optional embodiment, a regional complexity coefficient is calculated based on the number of communication devices and the topology depth in each region; the failure probability is then weighted and fused with the regional complexity coefficient to generate a risk value. Through these technical means, the probing strategy not only responds to the failure probability but also considers the impact of regional structural complexity on the difficulty of failure propagation and location, thereby improving the robustness of scheduling decisions.

[0079] As an example, topology depth is the number of shortest path hops from a region or node in a device's network topology to a reference root node (such as a core gateway or data center exit), used to measure the depth of that region's position in the network hierarchy.

[0080] In one embodiment of this disclosure, a testing strategy for each region is determined based on the risk value of each region, including: for any region: when the risk value of the region is greater than a first preset threshold, the testing interval of the region is shortened based on the initial strategy; when the risk value of the region is less than or equal to a second preset threshold, the testing interval of the region is increased based on the initial strategy, wherein the first preset threshold is greater than the second preset threshold; when the risk value of the region is less than or equal to the first preset threshold but greater than the second preset threshold, the testing interval of the region remains unchanged based on the initial strategy.

[0081] In this embodiment, the detection strategy for each region dynamically adjusts the detection interval by comparing the region's risk value with two preset thresholds: the detection interval is shortened when the risk value is higher than the first preset threshold, increased when the risk value is not higher than the second preset threshold, and kept constant when the risk value is between the two thresholds, thereby achieving adaptive detection control based on risk classification. Through these techniques, the detection frequency is strictly aligned with the actual risk level of the region, avoiding resource waste caused by continuously performing high-density detection in low- and medium-risk areas.

[0082] In one optional embodiment, for each region, a corresponding dialing interval is linearly mapped based on the position of its risk value in a continuous numerical space; probes are then activated to dial the communication devices in that region according to the mapped dialing interval. Through the above technical means, fine-grained continuous adjustment of the dialing interval is achieved, rather than being limited to discrete level switching, thereby improving the accuracy and smoothness of resource allocation.

[0083] In one embodiment of this disclosure, a testing strategy for each region is determined based on the risk value of each region, including: for any region: when the risk value of the region is greater than a first preset threshold, a first interval is determined based on the difference between the risk value of the region and the first preset threshold, and the testing interval of the region is shortened according to the first interval based on the initial strategy; when the risk value of the region is less than or equal to a second preset threshold, a second interval is determined based on the difference between the risk value of the region and the second preset threshold, and the testing interval of the region is increased according to the second interval based on the initial strategy; when the risk value of the region is less than or equal to the first preset threshold but greater than the second preset threshold, the testing interval of the region is kept unchanged based on the initial strategy.

[0084] In this embodiment, the detection strategy for each region dynamically calculates the adjustment amount based on the difference between the region's risk value and a preset threshold: when the risk value is higher than a first preset threshold, the size of the excess is used as the first interval, which shortens the detection interval; when the risk value is not higher than a second preset threshold, the size of the difference is used as the second interval, which increases the detection interval; when the risk value is between the two thresholds, the initial detection interval remains unchanged, thereby achieving quantitative adjustment based on the degree of risk deviation. Through the above technical means, the adjustment range of the detection interval is proportional to the severity of the risk, avoiding the use of the same intensity of detection response for minor risk fluctuations or extremely high-risk scenarios.

[0085] In one optional embodiment, a nonlinear mapping function is set from risk value to probe interval, which has a larger slope in high-risk areas. The risk value of each region is substituted into this mapping function to obtain the corresponding probe interval. Probes are then activated according to this probe interval to probe communication devices in that region. Through these techniques, the system becomes more sensitive to risk changes in high-risk areas, enabling rapid concentration of detection resources towards the most urgent areas.

[0086] Figure 5 This invention discloses a flowchart illustrating a method for predicting the probability of regional failures according to an embodiment of the present disclosure. The method is as follows: Figure 5 As shown, it includes the following steps: S501 uses historical alarm data from various communication devices to construct a historical failure rate matrix; S502 inputs the historical failure rate matrix, device network topology, and current test results into the hybrid model and outputs the failure probability of each region.

[0087] As an example, the historical failure rate matrix is ​​a numerical matrix formed by aggregating historical alarm data of various communication devices according to time and regional dimensions. It is used to characterize the frequency or probability distribution of failures in each region during historical periods.

[0088] In this embodiment, after the probe is initiated according to the initial strategy to complete the probe test and obtain the current test results, a historical failure rate matrix is ​​further constructed using historical alarm data. This matrix, along with the device network topology and the current test results, is then provided as input to the hybrid model to output the failure probability of each region. This enhances the predictive model's ability to jointly perceive long-term failure patterns and short-term state changes. Through these technical means, the accuracy of the hybrid model's prediction of regional failure trends is improved, providing a more reliable basis for subsequent risk value calculation and probe test strategy generation.

[0089] In one optional embodiment, a historical fault time-series sequence set is constructed using historical alarm data from various communication devices, where each sequence corresponds to a region. The historical fault time-series sequence set, the graph embedding representation of the device network topology, and the current test results are concatenated into a multimodal input feature. The multimodal input is then fed into a hybrid model, which outputs the fault probability for each region. By employing the above techniques, the original time-series dynamic details are preserved, rather than simply using the aggregated fault rate matrix. This enables the model to capture sudden, non-stationary fault modes, improving its predictive sensitivity to novel or rare anomalies.

[0090] For example, in a certain operator's backbone network, the configuration information and historical alarm data of all communication devices are first obtained. These devices include core routers, aggregation switches, access switches, firewalls, load balancers, DNS servers, and business application servers. Based on the configuration information of these communication devices, such as IP addresses, interface status, and link connection relationships, a nationwide device network topology is constructed. Simultaneously, device priorities are determined according to the priority of the services carried by each communication device. For example, core routers carrying 5G user authentication services are given high priority, while access switches used only for internal monitoring log forwarding are given low priority. Probes are initiated according to the initial strategy, performing ICMP connectivity probes and HTTP service availability probes on all communication devices every 5 minutes to obtain current test results, including latency, packet loss rate, and HTTP status codes. Historical alarm data is aggregated by region and time window to construct a historical failure rate matrix. This historical failure rate matrix, device network topology, and current test results are then input into a hybrid model. The system outputs the failure probability of each area in the network topology of the output device. For each area, a risk value is generated by weighting the failure probability with the priority of the communication devices in the area. For any area, if its risk value is greater than the first preset threshold of 0.8, the first interval is determined to be 3.8 minutes based on the excess risk value (e.g., 0.92-0.8=0.12) using a preset function, and the dialing interval is shortened from 5 minutes to 1.2 minutes. If its risk value is less than or equal to the second preset threshold of 0.3, the second interval is determined to be 7 minutes based on the difference (e.g., 0.3-0.15=0.15), and the dialing interval is extended to 12 minutes. If the risk value is between 0.3 and 0.8, the dialing interval remains unchanged at 5 minutes. Finally, according to the above dynamically generated dialing strategy, probe instances with corresponding protocol templates are called from the network probe library to perform differentiated dialing tasks on communication devices such as core routers and DNS servers in each area, thereby significantly reducing the overall detection resource consumption while ensuring the monitoring strength of high-priority service links.

[0091] Based on the same disclosed concept, this disclosure also provides a device testing apparatus, as shown in the following embodiments. Since the principle by which the device testing apparatus solves the problem is similar to that of the above method embodiments, the implementation of the device testing apparatus can refer to the implementation of the above method embodiments, and repeated details will not be elaborated further.

[0092] Figure 6 This diagram illustrates a device testing apparatus according to an embodiment of the present disclosure, such as... Figure 6 As shown, the testing device may include: The acquisition unit 601 is configured to acquire configuration information and historical alarm data of various communication devices; Construction unit 602 is configured to construct a device network topology based on the configuration information of various communication devices; The probe testing unit 603 is configured to initiate probe testing on various communication devices according to an initial strategy and obtain the current testing results. The prediction unit 604 is configured to input historical alarm data of various communication devices, device network topology and current dial-up test results into the hybrid model, and output the fault probability of each area in the device network topology. The determination unit 605 is configured to determine the dialing test strategy for each region based on the fault probability of each region, and start the probe to dial the communication equipment in each region according to the dialing test strategy of each region.

[0093] In some embodiments, the determining unit 605 is further configured to determine the priority of various communication devices; and to determine a dialing strategy for each region based on the failure probability of each region and the priority of the communication devices in that region.

[0094] In some embodiments, the determining unit 605 is further configured to acquire the priority of the services carried by various communication devices; and to determine the priority of various communication devices based on the priority of the services carried by various communication devices.

[0095] In some embodiments, the determining unit 605 is further configured to determine the risk value of each region based on the failure probability of each region and the priority of the communication equipment in that region; and to determine the dialing test strategy for each region based on the risk value of each region.

[0096] In some embodiments, the determining unit 605 is further configured to, for any region: if the risk value of the region is greater than a first preset threshold, shorten the dialing interval of the region based on the initial strategy; if the risk value of the region is less than or equal to a second preset threshold, increase the dialing interval of the region based on the initial strategy, wherein the first preset threshold is greater than the second preset threshold; if the risk value of the region is less than or equal to the first preset threshold but greater than the second preset threshold, keep the dialing interval of the region unchanged based on the initial strategy.

[0097] In some embodiments, the determining unit 605 is further configured to, for any region: when the risk value of the region is greater than a first preset threshold, determine a first interval based on the difference between the risk value of the region and the first preset threshold, and shorten the detection interval of the region according to the first interval based on the initial strategy; when the risk value of the region is less than or equal to a second preset threshold, determine a second interval based on the difference between the risk value of the region and the second preset threshold, and increase the detection interval of the region according to the second interval based on the initial strategy; when the risk value of the region is less than or equal to the first preset threshold but greater than the second preset threshold, keep the detection interval of the region unchanged based on the initial strategy.

[0098] In some embodiments, the prediction unit 604 is further configured to construct a historical failure rate matrix using historical alarm data from various communication devices; input the historical failure rate matrix, device network topology, and current test results into a hybrid model, and output the failure probability of each region.

[0099] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."

[0100] The following reference Figure 7 To describe an electronic device 700 according to such an embodiment of the present disclosure. Figure 7 The electronic device 700 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0101] like Figure 7 As shown, the electronic device 700 is presented in the form of a general-purpose computing device. The components of the electronic device 700 may include, but are not limited to: at least one processor 710, at least one memory 720, and a bus 730 connecting different system components (including memory 720 and processor 710).

[0102] The memory stores program code that can be executed by the processor 710, causing the processor 710 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processor 710 can perform the following steps of the above method embodiments: Acquire configuration information and historical alarm data of various communication devices; construct device network topology based on the configuration information of various communication devices; initiate probes to perform dial-up tests on various communication devices according to the initial strategy, and obtain the current dial-up test results; input the historical alarm data of various communication devices, device network topology and current dial-up test results into the hybrid model, and output the fault probability of each area in the device network topology; determine the dial-up test strategy for each area based on the fault probability of each area, and initiate probes to perform dial-up tests on communication devices in each area according to the dial-up test strategy of each area.

[0103] Determine the priority of various communication devices; based on the failure probability of each area and the priority of communication devices in that area, determine the dialing test strategy for each area.

[0104] Obtain the priority of the services carried by various communication devices; determine the priority of various communication devices based on the priority of the services carried by various communication devices.

[0105] Based on the failure probability of each region and the priority of the communication equipment in that region, the risk value of each region is determined; based on the risk value of each region, the dialing test strategy for each region is determined.

[0106] For any given region: if the risk value of the region is greater than the first preset threshold, then the call interval for that region is shortened based on the initial strategy; if the risk value of the region is less than or equal to the second preset threshold, then the call interval for that region is increased based on the initial strategy, wherein the first preset threshold is greater than the second preset threshold; if the risk value of the region is less than or equal to the first preset threshold but greater than the second preset threshold, then the call interval for that region remains unchanged based on the initial strategy.

[0107] For any given region: when the risk value of the region is greater than a first preset threshold, a first interval is determined based on the difference between the risk value of the region and the first preset threshold, and the detection interval of the region is shortened according to the first interval based on the initial strategy; when the risk value of the region is less than or equal to a second preset threshold, a second interval is determined based on the difference between the risk value of the region and the second preset threshold, and the detection interval of the region is increased according to the second interval based on the initial strategy; when the risk value of the region is less than or equal to the first preset threshold but greater than the second preset threshold, the detection interval of the region remains unchanged based on the initial strategy.

[0108] A historical failure rate matrix is ​​constructed using historical alarm data from various communication devices. The historical failure rate matrix, device network topology, and current test results are input into a hybrid model to output the failure probability of each region.

[0109] The memory 720 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 7201 and / or cache memory 7202, and may further include read-only memory (ROM) 7203.

[0110] The memory 720 may also include a program / utility 7204 having a set (at least one) of program modules 7205, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0111] Bus 730 can represent one or more of several types of bus structures, including a memory bus or memory controller, peripheral bus, graphics acceleration port, processor, or a local bus using any of the various bus structures.

[0112] Electronic device 700 can also communicate with one or more external devices 740 (e.g., keyboard, pointing device, Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with electronic device 700, and / or with any device that enables electronic device 700 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 750. Furthermore, electronic device 700 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 760. As shown, network adapter 760 communicates with other modules of electronic device 700 via bus 730. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 700, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0113] In the disclosed exemplary embodiments, a computer-readable storage medium is also provided, which may be a readable signal medium or a readable storage medium.

[0114] In some possible implementations, various aspects of this disclosure may also be implemented as a program product comprising program code that, when run on a terminal device, causes the terminal device to perform the steps described in the foregoing “Detailed Description” section of this specification according to various exemplary embodiments of this disclosure.

[0115] More specific examples of computer-readable storage media in this disclosure may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0116] In this disclosure, a computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device.

[0117] Optionally, the program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0118] In practical implementation, program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on a terminal device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0119] This disclosure provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a device testing method provided in various optional embodiments of this disclosure.

[0120] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0121] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0122] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0123] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope of this disclosure is indicated by the appended claims.

Claims

1. A method for testing equipment, characterized in that, include: Obtain configuration information and historical alarm data of various communication devices; Construct a device network topology based on the configuration information of various communication devices; The probes are activated according to the initial strategy to perform tests on various communication devices, and the current test results are obtained. The historical alarm data of various communication devices, the network topology of the devices, and the current test results are input into the hybrid model, and the failure probability of each area in the network topology of the devices is output. Based on the failure probability of each region, a testing strategy is determined for each region, and probes are started to test the communication equipment in each region according to the testing strategy.

2. The method according to claim 1, characterized in that, Before initiating probe testing of communication devices in each region according to the testing strategy for each region, the method further includes: Determine the priority of various communication devices; Based on the failure probability of each region and the priority of communication equipment in that region, a dialing test strategy is determined for each region.

3. The method according to claim 2, characterized in that, Determining the priority of various communication devices includes: Obtain the priority of services carried by various communication devices; The priority of each communication device is determined based on the priority of the services it carries.

4. The method according to claim 2, characterized in that, The determination of the dialing test strategy for each region based on the failure probability of each region and the priority of communication equipment in that region includes: Based on the failure probability of each region and the priority of the communication equipment in that region, the risk value of each region is determined. Based on the risk values ​​of each region, a testing strategy is determined for each region.

5. The method according to claim 4, characterized in that, The determination of the detection strategy for each region based on the risk value of each region includes: For any given region: If the risk value of the area is greater than the first preset threshold, the detection interval of the area will be shortened based on the initial strategy. When the risk value of the area is less than or equal to the second preset threshold, the detection interval of the area is increased based on the initial strategy, wherein the first preset threshold is greater than the second preset threshold; If the risk value of the area is less than or equal to the first preset threshold, but greater than the second preset threshold, the detection interval of the area will remain unchanged based on the initial strategy.

6. The method according to claim 4, characterized in that, The determination of the detection strategy for each region based on the risk value of each region includes: For any given region: When the risk value of the area is greater than the first preset threshold, a first interval is determined based on the difference between the risk value of the area and the first preset threshold, and the detection interval of the area is shortened according to the first interval based on the initial strategy. When the risk value of the area is less than or equal to the second preset threshold, a second interval is determined based on the difference between the risk value of the area and the second preset threshold, and the detection interval of the area is increased according to the second interval based on the initial strategy; If the risk value of the area is less than or equal to the first preset threshold, but greater than the second preset threshold, the detection interval of the area will remain unchanged based on the initial strategy.

7. The method according to claim 1, characterized in that, After the method involves initiating probe tests on various communication devices according to the initial strategy and obtaining the current test results, it further includes: Construct a historical failure rate matrix using historical alarm data from various communication devices; The historical failure rate matrix, the device network topology, and the current test results are input into the hybrid model to output the failure probability of each region.

8. A device for testing equipment, characterized in that, include: The acquisition unit is configured to acquire configuration information and historical alarm data of various communication devices; The building unit is configured to construct the device network topology based on the configuration information of various communication devices; The probe unit is configured to initiate probes to perform probe tests on various communication devices according to an initial strategy and obtain the current probe test results. The prediction unit is configured to input historical alarm data of various communication devices, the device network topology, and the current dial-up test results into a hybrid model, and output the fault probability of each area in the device network topology. The determination unit is configured to determine the dialing test strategy for each region based on the failure probability of each region, and start probes to dial test the communication equipment in each region according to the dialing test strategy of each region.

9. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the method of any one of claims 1-7 by executing the executable instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-7.