A network telemetry method and device for distinguishing scenes
By building a hybrid perception mechanism switching architecture under the SDN architecture, combining with P4 programmable switches, dynamically adjusting the detection frequency and paths, the problem of insufficient perceptual response capabilities of INT technology in complex network scenarios is solved, and efficient and flexible network telemetry strategies and information acquisition are achieved.
Patent Information
- Application Number
- CN202410827550.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2044-06-25
AI Technical Summary
In the face of complex network scenarios, existing INT technology is difficult to adapt to the measurement indicator requirements of multiple network business scenarios, resulting in difficulty in obtaining network status information, insufficient perceived response capabilities, and there are problems of perceived information lag and loss caused by fixed length of data packets.
A hybrid perception mechanism switching architecture based on in-band network telemetry is built, combined with the software-defined network SDN architecture, through the separation of the control plane and the data plane, the P4 programmable switch is used to realize the coordinated work of the network state collection unit, the working mode switching unit, the detection frequency adjustment unit and the path control unit, dynamically adjust the detection frequency and path, and improve the active perception function of nodes in the network.
It realizes efficient and flexible telemetry strategies in different network business scenarios, meets the requirements of anti-packet loss and high real-time, improves the accuracy and timeliness of network performance and information perception, and reduces network overhead.
Smart Images

Figure CN118740647B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and specifically to a method for real-time telemetry of network status information, which is applied to various network scenarios to perform full network monitoring, real-time network status assessment, and timely and accurate control of network anomalies. In particular, it relates to a network telemetry method and device that distinguishes scenarios. Background Art
[0002] The network is a highly complex distributed system. Therefore, network management tools are required to manage network devices, regulate and control the overall network's operational status, and promptly detect and report network faults to endpoints. Traditional network monitoring methods are typically based on a server-client model. For example, SNMP allows network control and management systems to extract statistics from their network elements for non-real-time data collection. sFlow provides a full network view by periodically sampling interface statistics and randomly sampling packet information. NetFlow collects IP flow information at a constant sampling rate and sends the aggregated results to a flow collector for analysis. While these methods have proven effective in the early days of the internet, their limitations in achieving more complex monitoring are also evident. Over the past decade, the internet has undergone revolutionary changes and dramatic expansion worldwide. As a result, networks have become increasingly flexible and programmable, at the cost of increased complexity. This means that troubleshooting and recovering from unclear faults (such as congestion, link failures, and black holes) will become more challenging. Therefore, network management solutions need to evolve to become more automated, capable of dynamically deepening and even automatically reacting to network events based on current conditions. This automated response requires network telemetry systems to provide high-precision, fine-grained events (e.g., instantaneous traffic bursts) on a fine timescale. The rise of OpenFlow-based SDN has provided new insights into network monitoring technology. Data plane switches need to regularly report their characteristic information to the SDN controller to achieve centralized network control and management. However, due to the relatively long latency of data acquisition, incomplete end-to-end single-flow information, and the non-scalable processing burden on the switch, problems still exist. Therefore, it is necessary to consider solving them through a more programmable data plane. Programming the protocol-independent packet processor P4 or Protocol Oblivious Forwarding (POF) can achieve a programmable and protocol-independent data plane. Therefore, the in-band network telemetry technology proposed based on the P4 language provides network operators with the ability to customize network monitoring solutions. INT greatly improves network visibility, realizes real-time, fine-grained network monitoring, makes it easier to diagnose network anomalies, and solves the performance and scalability issues existing in traditional monitoring methods. INT, with its real-time, accurate, and control plane-free features, has brought new directions to network measurement technology. INT-based measurement solutions and applications have become a research hotspot in current network operations, management, and maintenance.
[0003] Domestic and international scholars have produced a wealth of research results on in-band network telemetry technology. Currently, the INT specification, for the first time, describes the concept of INT and a complete INT system prototype, along with relevant application cases. To reduce the application overhead of INT technology, numerous studies have introduced new technologies. For example, source routing (SR), based on the INT specification, can purposefully guide the forwarding path of probe packets, thereby reducing flow table overhead. In most network monitoring scenarios, particularly when link capacity utilization is relatively high, highly accurate, real-time, per-packet INT monitoring strategies are not necessary. Kim et al. proposed a P4-based sampling sINT scheme that selectively inserts INT headers into terminal hosts. PINT divides telemetry data collection operations into three categories: packet aggregation, static flow aggregation, and dynamic flow aggregation. This reduces INT monitoring overhead to the bit level while ensuring monitoring accuracy. Another major technical improvement focus on INT is fault detection in the network. Tang et al. proposed PAINT, an intelligent SDN fault location system based on INT. Tan et al. also designed a packet loss monitoring system, LossSight, for INT, which includes packet loss detection, location, diagnosis, and recovery capabilities.
[0004] With the development of its technology, the current application research of INT mainly focuses on the optimization of a certain requirement for a single working mode. However, the existing different working modes of INT have shortcomings and lack application scalability in networks of various scales. There are two reasons: (1) Affected by traffic duration, traffic characteristics and traffic distribution, the INT method of on-path perception is difficult to achieve a high overall network link telemetry coverage and cannot achieve a high network fault detection accuracy; (2) The INT method of uploading perception information hop by hop requires each INT transmission node to collect perception data and construct a detection message for uploading, which will result in additional network bandwidth occupation. Therefore, how to take into account the advantages of the two working modes and achieve efficient telemetry task scheduling is a problem worth studying. In addition, different network scenarios have different requirements for network internal perception information. The traditional single network perception method is difficult to guarantee the efficiency and real-time performance of monitoring under various network scales. Research on dynamically adjusting perception measurement tasks and timely obtaining network status information is still weak. Therefore, considering the adaptability of network perception scenarios, taking the two scenarios of high real-time and anti-packet loss as examples, the application of INT is analyzed:
[0005] Because end-to-end telemetry mechanisms make INT unavoidable for packet loss, and various reasons can cause packet loss during the perception process, INT perception can become unreliable due to potential network packet loss. Incomplete telemetry data can severely impact the performance of upper-layer network telemetry applications. Current research on INT packet loss primarily focuses on analyzing the packet loss itself, such as locating the specific location of lost packets, recovering from lost packets, or performing detailed analysis of the loss situation. Various techniques, such as network coding and machine learning, are employed to address this problem. However, these approaches are difficult to quickly and easily scale in complex network scenarios, and the multiple optimization methods introduced by focusing solely on the causes of packet loss increase the complexity of network perception technology itself.
[0006] Due to the nature of INT's metadata collection along the way, excessive data transmission through too many routing devices can lead to excessive overhead and increased perception time. This can be minimized by averaging the number of nodes per perception path or the information acquisition latency. To reduce this overhead, current INT technology improvements focus on covering all links within the network and reducing duplicate perception nodes. However, while covering all links within the network, these improvements rarely consider the uneven update rates of perception information among nodes within the network, and they also fail to decouple perception operations for individual vulnerable nodes.
[0007] The technical solution of one of the prior arts retrieved by the inventor is described as follows:
[0008] The Shandong Provincial Computing Center has disclosed a gray fault detection and positioning method and system based on hybrid in-band network telemetry, which relates to the field of fault detection. It includes: the server collects hop-by-hop telemetry information of passive INT detection packets, performs a primary detection on whether a fault exists, and sends a secondary detection instruction for the faulty path to the controller of the virtual SDN network; the controller sends an active INT detection packet to the server, and performs a secondary detection on the path that has a fault in the primary detection; the source server reroutes the data traffic of the path information that actually has a fault; the controller sets a priority for all path information that actually has a fault, compares the paths based on the priority, and obtains the fault location; the controller feeds the fault location back to the server, and the server searches for all paths related to the fault location and ages them in advance. This invention integrates active in-band network telemetry and passive in-band network telemetry to make up for the shortcomings of a single telemetry method and improve the efficiency and reliability of network telemetry.
[0009] The relative shortcomings of the prior art include:
[0010] (1) This technology sends active INT detection packets twice to locate the fault point and then adjusts the path. It mainly focuses on locking the fault location and is achieved by integrating active and passive INT detection packets. However, it does not consider the changes in detection time cost under different network topology scales.
[0011] (2) Both the active and passive detection methods involved in this technology are INT flow-by-flow detection methods. There is a possibility that the fault point will be lost due to fault recovery caused by the long detection time during the detection process.
[0012] The technical solution of the second prior art retrieved by the inventor is described as follows:
[0013] The Guangdong New Generation Communications and Network Innovation Research Institute has disclosed a network status-based in-band network telemetry method. The method includes: obtaining network status information during packet telemetry, constructing a network status table based on this information; configuring a mapping relationship between the network status table and telemetry frequency, and determining the telemetry frequency based on this mapping relationship; and adding telemetry control information to packets that meet the telemetry conditions based on the telemetry frequency to implement in-band network telemetry. This invention can automatically adjust the frequency of network telemetry, reducing the additional network overhead incurred by network telemetry and ensuring the normal forwarding function of the network.
[0014] The relative shortcomings of the prior art include:
[0015] Technique 2 uses INT to collect network status as a benchmark to adjust the frequency of probes at the transmitting end, improving the real-time nature of network information acquisition. However, this technique focuses on comparing the number and arrival times of probe packets between the transmitting and receiving ends. Actual network conditions are complex and diverse, so comparing only one transmitting and receiving end can lead to problems such as misjudgment and inaccurate frequency adjustment.
[0016] In view of the above description of the prior art, the technical problems to be solved by the present invention include:
[0017] Due to the rapid development of network equipment and the enhancement of network programmability, the demand for network measurement technology is also increasing. More new applications are increasingly sensitive to changes in internal network indicators. In order to adapt to various types of monitoring needs, traditional INT technology collects massive amounts of network status information and uploads it to the control end, which also greatly increases its workload. In addition, there are risks such as delayed and missing perception information due to the limitations of fixed data packet length.
[0018] In addition, due to the complex and changeable internal conditions of the network and the frequent occurrence of network anomalies, the application adaptability of traditional INT technology is insufficient, making it difficult to flexibly change the measurement indicators required in various network business scenarios. The method of the present invention addresses this technical issue and, based on the network status characteristics in different scenarios, constructs a hybrid perception mechanism switching architecture based on in-band network telemetry. This architecture focuses on key issues such as implementing INT switching between different working modes, active perception and notification functions of nodes within the network, and perception attribute priority path planning. It addresses the difficulties in obtaining network status information in complex scenarios and the lack of in-network perception response capabilities, making network operation more efficient and stable, ensuring the timeliness and accuracy of perception information, and improving network performance.
[0019] In view of the above-mentioned deficiencies in the prior art, the present invention is proposed. Summary of the Invention
[0020] The purpose of the present invention is to address the problems of difficulty in obtaining network status information and insufficient network perception response capabilities in existing technologies in complex scenarios. According to the network status characteristics in different scenarios, a hybrid perception mechanism switching architecture based on in-band network telemetry is constructed. Focusing on key issues such as realizing INT multi-working mode combined switching, active perception notification function of in-network nodes and perception attribute priority path planning, the present invention considers the perception needs of different network business scenarios, gives in-network nodes active perception function, improves the in-network information perception and update capability, and realizes the hybrid technology implementation of INT working mode and in-network node perception.
[0021] In order to achieve the above object, the present invention adopts the following technical solutions:
[0022] A scenario-specific network telemetry method is based on a software-defined network (SDN) architecture. The control plane and data plane are separated. The control plane is managed by a central controller, while the data plane is handled by network devices that perform actual packet processing. The central controller is responsible for decision-making for the entire network. The data plane uses a P4 programmable switch to support network node functions. The method includes a network status collection unit, a working mode switching unit, a detection frequency adjustment unit, and a path control unit. The working mode switching unit and the detection frequency adjustment unit are both included in the telemetry control unit.
[0023] The network status collection unit is used to collect network scenario information and real-time status information. The network scenario information describes the business scenarios to which the technology is applicable, facilitating the selection of different working modes for path detection and adjustment of detection frequency. The real-time status information is the collected node metadata that accurately reflects the network situation, facilitating timely adjustment of telemetry strategies and ensuring flexible and reliable network detection.
[0024] The working mode switching unit selects the telemetry mode according to the data collected by the network status collection module based on different network scenario characteristics, path status and different working mode characteristics of INT;
[0025] The detection frequency adjustment unit is responsible for receiving active sensing notifications from nodes in the network, thereby adjusting the detection frequency;
[0026] The path control unit is responsible for generating multiple detection paths that evenly cover nodes in the network according to the network topology; and replanning the detection paths based on the detection frequency adjustment;
[0027] The specific workflow includes the following steps:
[0028] Step S1: Initial path planning is performed based on network scenario attributes, with the goal of covering all network links. The network collection unit obtains all network information.
[0029] Step S2, the path control unit generates multiple detection paths that evenly cover nodes in the network according to the network topology based on the information fed back by the network collection unit;
[0030] Step S3: Based on the information collected in step S1, select a working mode according to different network scales or different network scenarios;
[0031] Step S4: After the network information is collected, it is sent to the network status collection unit for storage in the network scenario information; using the network resource information, the detection method is further adjusted; services are provided for detection path re-planning, detection frequency optimization, and node detection action adjustment;
[0032] The network status collection unit, working mode switching unit, detection frequency adjustment unit and path control unit are functionally independent but mutually supportive. They continuously adjust the detection behavior and provide feedback through network status information to ensure the real-time and accuracy of perception.
[0033] Preferably, the nodes in the network are all provided with a function of autonomously detecting changes in network conditions between nodes, and the specific method is:
[0034] Step S1, receiving INT telemetry message;
[0035] Step SII, collecting node status information and comparing the collected information with a set threshold based on a preset network indicator;
[0036] Step SIII, determining whether the metadata collected by the INT detection packet is greater than a preset threshold; if not, proceeding to step SIV; if yes, proceeding to step SV;
[0037] Step SIV: Telemetry messages that do not exceed the threshold continue to collect telemetry information along the detection path;
[0038] In step SV, the telemetry message node that exceeds the threshold reports the telemetry to the central controller for active notification, and the telemetry message continues to be detected along the path;
[0039] In step SVI, the central controller controls the transmitting end to adjust the telemetry frequency according to the notification information.
[0040] Preferably, the working mode switching unit includes a hybrid perception mode switching method based on network scenario information, or a dynamic adjustment perception method based on network status changes,
[0041] The hybrid perception mode switching method based on network scenario information is specifically as follows:
[0042] First, we analyze the perception requirements of network scenarios and construct multi-objective optimization problems based on different scenario requirements. In each scenario, we can select reasonable information based on the perception priority.
[0043] Secondly, based on the hybrid telemetry switching model, a set of constraints is constructed around network resource availability constraints, dependency constraints on scenario telemetry requirements, detection path direction, and perception effect requirements.
[0044] The method for dynamically adjusting perception based on network status changes is specifically as follows:
[0045] In the adaptive switching framework, considering the complex internal situation of the network and the different needs of collecting telemetry data, the perception parameter thresholds are set based on the flexible switching of the perception working mode. The detection frequency is adjusted dynamically on demand by the nodes, so that dynamic interaction between nodes can realize more measurement possibilities and achieve flexible and efficient telemetry tasks.
[0046] Preferably, the detection frequency adjustment unit is connected to the threshold setting of the nodes in the network, and multiple adjustments are made to different detection frequencies according to whether the threshold is exceeded and the degree of exceeding the threshold.
[0047] Preferably, in step S4,
[0048] Replanning of the detection path includes, but is not limited to, utilizing network scenario information and real-time feedback on the link transmission status and node queue depth provided by real-time recovered network status information, or other real-time feedback on significant network conditions, to ensure high coverage of network information detection links and high information volume obtained from network resources in accordance with scenario requirements;
[0049] Alternatively, for abnormal nodes, a separate detection path is set up, and a hop-by-hop detection mode is adopted to ensure the accuracy and completeness of detection information collection. In addition, to prevent node deterioration from being caused by the detection behavior itself, path redundancy operations are performed for the node directly connected to the upstream node to avoid excessive detection traffic flowing through the abnormal node.
[0050] Preferably, in step S4,
[0051] Detection frequency optimization is based on the initial detection frequency and performs real-time updates on the detection frequencies of different paths according to the speed of network information update.
[0052] Preferably, in step S4,
[0053] The node detection action adjustment is mainly aimed at switching the detection method and selecting the detection action of abnormal nodes in the network. It can select and switch the appropriate working mode for the node according to the changes in the network scenario or the changes in the conditions of a certain link. In addition, because the nodes in the network have active perception functions, they can also detect anomalies in the network in time and take active information upload actions, thereby improving the whole system's ability to handle abnormalities.
[0054] The present invention also provides a network telemetry device for distinguishing scenarios using a network telemetry method, including:
[0055] The control plane is managed by a central controller, while the data plane, where network devices perform actual packet processing, uses a central controller responsible for decision-making for the entire network; the data plane uses P4 programmable switches to support network node functions; and
[0056] Network status collection module, working mode switching module, detection frequency adjustment module and path control module, wherein the working mode switching module and the detection frequency adjustment module are both included in the telemetry control module; wherein,
[0057] The network status collection module is used to collect network scenario information and real-time status information. The network scenario information describes the business scenarios to which the technology is applicable, facilitating the selection of different working modes for path detection and adjustment of detection frequency. The real-time status information is the collected node metadata that accurately reflects the network situation, facilitating timely adjustment of telemetry strategies and ensuring flexible and reliable network detection.
[0058] The working mode switching module selects the telemetry mode according to the data collected by the network status collection module based on the characteristics of different network scenarios, path status and different working modes of INT;
[0059] The detection frequency adjustment module is responsible for receiving active sensing notifications from nodes in the network and adjusting the detection frequency;
[0060] The path control module is responsible for generating multiple detection paths that evenly cover nodes in the network according to the network topology; and replanning the detection paths based on the detection frequency adjustment.
[0061] The network status collection module, working mode switching module, detection frequency adjustment module and path control module are functionally independent but mutually supportive. They continuously adjust detection behavior and provide feedback through network status information to ensure the real-time and accuracy of perception.
[0062] The present invention also provides an electronic device, comprising a processor and a memory; wherein, when the processor executes a computer program stored in the memory, the aforementioned network telemetry method for distinguishing scenarios is implemented.
[0063] The present invention also provides a method for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned network telemetry method for distinguishing scenarios is implemented.
[0064] The present invention can achieve the following beneficial effects:
[0065] With the rapid development of software-defined networks and programmable data planes, in-band network telemetry has received widespread attention as an advanced network monitoring method. INT uses data packets to carry network telemetry information and collects network status information hop by hop for congestion control, traffic scheduling, and anomaly detection. The method of the present invention proposes a network information hybrid perception method to combine the two working modes to achieve efficient telemetry task scheduling. Moreover, in different network service scenarios, such as anti-packet loss and high real-time requirement scenarios, it can meet various telemetry needs, implement flexible telemetry strategies, and obtain accurate telemetry information. The advantages of the present invention are
[0066] (1) The method of the present invention considers the problem of uneven update rates of node perception information within the network and decouples the perception operations of individual vulnerable nodes. On the basis of full network telemetry coverage, the method considers the perception requirements of different network service scenarios, increases the node actions within the network, and improves the network's information perception and update capabilities.
[0067] (2) With the development of INT technology, current INT application research mainly focuses on optimizing a specific requirement for a single working mode. However, the existing different INT working modes all have shortcomings and lack scalability in applications across networks of various scales. The method of the present invention combines two INT detection methods, taking into account the advantages of both working modes and switching mechanisms based on different network scenario requirements, such as network size, to achieve better perception performance.
[0068] (3) Since the switching nodes are configured using the P4 language, the method of the present invention has the following three characteristics: repeatability, platform independence, and protocol independence.
[0069] 1. Reconfigurability: This method allows users to flexibly define data plane processing behavior without having to replace hardware. Users can modify the P4 code based on their specific forwarding requirements for data plane nodes without having to replace equipment or wait for new equipment to be developed. This allows for dynamic modification of message processing after the forwarding logic code has been compiled and deployed on a specific platform.
[0070] ② Protocol independence: P4 code is not bound to any specific network protocol. This allows switching nodes to use network protocols as needed, fully utilizing device resources.
[0071] ③ Platform independence: Developers can write data packet processing logic independently of the specific underlying operating platform. Code can be quickly ported across different platforms, such as hardware switches, FPGAs, SmartNICs, and software switches, using device-specific backend compilers, reducing the burden on developers and improving development efficiency.
[0072] Explanation of terms:
[0073] ① Software Defined Network (SDN) is a new and innovative network architecture proposed by the Clean-Slate research group at Stanford University. It is a method for implementing network virtualization. Its core technology, OpenFlow, separates the control and data planes of network devices, enabling flexible control of network traffic. This makes the network, as a pipeline, more intelligent and provides an excellent platform for innovation in core networks and applications.
[0074] ②P4 programmable data plane: Programming Protocol-independent Packet Processors, Programming Protocol-independent Packet Processors (P4) is a domain-specific language for network devices that specifies how data plane devices (switches, NICs, routers, filters, etc.) process data packets.
[0075] ③INT Inband Network Telemetry: Inband Network Telemetry (INT) is a network measurement technology that fundamentally uses data plane services to collect, carry, organize, and report network status, without using separate control plane management traffic to collect the above information. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0077] Figure 1 It is the overall design framework diagram of the present invention;
[0078] Figure 2 This is a schematic diagram of the unit structure of a network telemetry method capable of distinguishing scenarios according to the present invention;
[0079] Figure 3 Schematic diagram of INT-MD working mode in the prior art;
[0080] Figure 4 It is a schematic diagram of the INT-MX working mode in the prior art;
[0081] Figure 5 This is a structural diagram of the node active perception working process;
[0082] Figure 6 Diagram of the node active perception workflow;
[0083] Figure 7 This is a schematic diagram of the business scenario perception work for high real-time requirements;
[0084] Figure 8 A flow chart for dynamically adjusting the detection path based on changes in network status;
[0085] Figure 9 A schematic diagram of an example of applying the present invention to a low-latency scenario. DETAILED DESCRIPTION
[0086] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0087] Figure 1 It is the overall design framework diagram of the present invention, such as Figure 1As shown, the present invention provides a scenario-differentiated network telemetry method that adopts a software-defined networking (SDN) architecture. Compared with the traditional network control plane and data plane integrated in the network device architecture, it introduces the separation of the control plane and the data plane. The control plane is managed by a central controller, while the data plane is performed by the network device to perform actual data packet processing. In addition, a central controller is used to be responsible for decision-making for the entire network. Through the programmatic flexibility of the controller, the SDN network is more programmable and can adjust network behavior as needed to support new protocols and services, providing greater flexibility and innovation space for the network, so that the network can adapt to new applications and services. P4 (Programming Protocol-Independent Packet Processors) programmable switches are used in the data plane to support network node functions, which improves the flexibility, customizability and programmability of the network, enables network devices to better adapt to the ever-changing network environment and needs, and provides more space for advanced optimization, innovation and management of the network.
[0088] Figure 2 This is a schematic diagram of the unit structure of a network telemetry method capable of distinguishing scenarios according to the present invention. Figure 2 As shown, the present invention includes a network status collection unit, a working mode switching unit, a detection frequency adjustment unit and a path control unit, wherein the working mode switching unit and the detection frequency adjustment unit are both included in the telemetry control unit; wherein,
[0089] The network status collection unit is used to collect network scenario information and real-time status information. The network scenario information describes the business scenarios to which the technology is applicable, facilitating the selection of different working modes for path detection and adjustment of detection frequency. The real-time status information is the collected node metadata that accurately reflects the network situation, facilitating timely adjustment of telemetry strategies and ensuring flexible and reliable network detection.
[0090] The working mode switching unit selects the telemetry mode according to the data collected by the network status collection module based on different network scenario characteristics, path status and different working mode characteristics of INT;
[0091] The detection frequency adjustment unit is responsible for receiving active sensing notifications from nodes in the network, thereby adjusting the detection frequency;
[0092] The path control unit is responsible for generating multiple detection paths that evenly cover nodes in the network according to the network topology; and replanning the detection paths based on the detection frequency adjustment;
[0093] The specific workflow includes the following steps:
[0094] Step S1: Initial path planning is performed based on network scenario attributes, with the goal of covering all network links. The network collection unit obtains all network information.
[0095] Network attributes are considered in multiple dimensions, and working modes are selected based on the size of the network. For example, for large-scale networks with a large average number of path hops, the hop-by-hop detection mode is selected to avoid incomplete information collection or excessive duplication of information collection due to data packet length limitations due to too many hops. For small-scale topologies or topologies with simple node connections, the flow-based detection mode is adopted, using a small number of detection packets to reclaim network resources. For different network scenarios, such as those with severe packet loss, a major difficulty in detecting network information is the loss of detection packets during transmission. Therefore, the hop-by-hop detection mode is adopted to cover as much network node information as possible. For network scenarios with high real-time requirements, the real-time acquisition of network resources is required to be high and the time difference of recovering information from multiple detection paths is small. Therefore, the path is reasonably planned based on the hop-by-hop detection mode.
[0096] Step S2, the path control unit generates multiple detection paths that evenly cover nodes in the network according to the network topology based on the information fed back by the network collection unit;
[0097] Step S3: Based on the information collected in step S1, select a working mode according to different network scales or different network scenarios;
[0098] In step S4, after the network information is collected, it is sent to the network status collection unit for storage in the network scenario information. This network resource information can be used to further adjust the detection method and improve the flexibility of the system. Specifically, the network status information unit can provide services for detection path replanning, detection frequency optimization, and node detection action adjustment.
[0099] The re-planning of the detection path has two functions. The first is to use the real-time transmission status of the link and the real-time feedback of the node queue depth and other landmark network conditions provided by the network scenario information and the real-time recovered network status information, and combine the scenario requirements to ensure the high coverage of the network information detection link and the high amount of information obtained by the network resources; the second is to set a separate detection path for abnormal nodes, such as queue depth and transmission delay exceeding the threshold setting, which may cause node transmission problems. The hop-by-hop detection mode is adopted to ensure the accuracy and completeness of the detection information collection. In addition, to prevent the deterioration of the node from being affected by the detection behavior itself, the path redundancy operation is performed for the node directly connected to the upstream node to avoid excessive detection traffic flowing through the abnormal node.
[0100] Probe frequency optimization involves updating the probe frequencies of different paths in real time based on the initial probe frequency and the speed of network information updates. For example, if collecting network-wide link information requires four probe paths, one of which includes multiple nodes with high exchange activity, more frequent probes between nodes are required. Therefore, a faster frequency is set to ensure the timeliness of the information.
[0101] The node detection action adjustment is mainly aimed at switching the detection method and selecting the detection action of abnormal nodes in the network. It can select and switch the appropriate working mode for the node according to the changes in the network scenario or the changes in the conditions of a certain link. In addition, because the nodes in the network have active perception functions, they can also detect anomalies in the network in time and take active information upload actions, thereby improving the whole system's ability to handle abnormalities.
[0102] Therefore, the network status collection unit, working mode switching unit, detection frequency adjustment unit and path control unit are functionally independent but mutually supportive, and continuously adjust the detection behavior and provide feedback through network status information to ensure the real-time and accuracy of perception.
[0103] The working mode switching unit includes a hybrid perception switching method based on network scenario information, or a dynamic perception adjustment method based on network status changes.
[0104] The hybrid perception mode switching method based on network scenario information is specifically described as follows:
[0105] To address the efficiency and accuracy issues of network information perception in complex scenarios, a hybrid perception switching method based on in-band network telemetry is proposed. As a hybrid measurement technology, in-band network telemetry technology is fundamentally a technology that uses data plane services to collect, carry, organize, and report network conditions, without using a separate control plane management traffic to collect the above information. INT technology involves three functional nodes, each of which provides different functions. For telemetry managers, service traffic that requires telemetry will have an INT header added to the source node, which contains an instruction set indicating the type of information to be collected, thus becoming an INT message. When it reaches the INT forwarding node, the collected information will be inserted into the INT message according to the instruction set. Finally, all INT information will pop up on the INT sink node and sent to the monitoring device.
[0106] a. Introduction to INT working mode:
[0107] The INT standard proposes three working modes: In the classic INT-MD, the INT source node embeds instructions into the forwarding service data packet, and then embeds metadata hop by hop as the data packet passes through the INT forwarding node. Finally, the INT sink node strips the instructions from the data packet and sends the accumulated telemetry data to the telemetry controller. The working mode is as follows: Figure 3 shown.
[0108] Based on the working mode of the accompanying path, the packet processing process of the INT mode (also known as INT-MX) that feeds back the metadata to the controller hop by hop is to add the INT instruction in the packet header at the source node, and then each forwarding node encapsulates the telemetry information into metadata according to the INT instruction in the packet and forwards the report message directly to the telemetry server. The INT host node finally strips off the INT header and forwards it to the receiver. In this mode, the modification of the data packet is limited to modifying the instruction header. The size of the data packet will not increase with the increase of the devices passed through during the monitoring process. The processing complexity of the forwarding device is reduced during the telemetry process, and the MTU limitation is avoided. However, the INT mode of packet-by-packet and hop-by-hop perception will introduce more out-of-band bandwidth overhead, and will also increase the pressure on data integration and processing of the data analysis system. The working method is as follows Figure 4 shown.
[0109] On the basis of INT-MX, the network switching equipment exports metadata to the telemetry monitoring system directly from the data plane according to the INT instructions of its own flow monitoring list configuration, without the need to modify the data message. Since the INT data packet only uses the flag bit, the INT-XD mode will cause this mode to be overly dependent on the monitoring list configured in advance by the network administrator, and cannot flexibly customize the deployment of monitoring solutions as needed. Therefore, the application of this working mode is not considered in the present invention.
[0110] The metadata categories that can be collected by INT technology are shown in Table 1 below:
[0111] Table 1 Metadata categories collected by INT technology
[0112]
[0113]
[0114] The use of basic metadata can complete the information perception within the switching node. On this basis, according to the secondary perception information processing, customized perception information can be used to achieve more detailed network information measurement, as shown in Table 2 below:
[0115] Table 2 Customized perception information table
[0116] Information Name Information Description node_pkt_drop_count Total number of node data packets lost node_byte_pro_time Node message processing time node_byte_pro_count Total number of node message processing route_hop_count Total number of nodes the path passes through route_trace_rtt Round trip time of the path route_hop_throughput Instantaneous throughput of nodes passing through the path send_probe_count Total number of probe packets sent receive_probe_count Total number of probe packets received
[0117] b. Introduction to node active perception function:
[0118] In view of the current situation that the application scope of single INT working mode perception is limited, the present invention combines the advantages of the two telemetry data collection working modes of INT along-the-path perception and hop-by-hop uploading of perception information, and switches the mechanism according to the requirements of different network scenarios, such as the size of the network, to achieve better perception performance. Considering the versatility of the switching mechanism in different network scenarios, with the network scale as a reference, the along-the-path perception information method is first selected to broadcast the unknown scenario to obtain the entire network topology information, and the information feedback is used to judge the switching according to the switching trigger conditions. The designed working mode switching module is used to switch the corresponding scenario applicable mode to achieve efficient telemetry tasks. On the basis of INT working mode switching, in order to address the timeliness of perception information, the function of autonomous detection of network status changes between nodes in the network is designed. The working method is as follows Figure 5 shown.
[0119] The node queue depth is used as a reference factor for triggering active perception, and the network congestion is measured according to the node queue depth, and the perception data is updated in a timely manner. Figure 6 As shown, all nodes in the network are equipped with an active notification function. When network anomalies occur, congestion is used as a criterion. When the node queue depth exceeds an initial threshold set based on the network scenario, congestion is determined to be present. The congested node then begins actively reporting the network's real-time status, sending INT detection packets to a remote telemetry controller. Furthermore, the node's active notification function offers high flexibility, with the frequency dynamically adjusted based on the percentage of nodes exceeding the threshold.
[0120] refer to Figure 6 , the specific method is:
[0121] Step S1, receiving INT telemetry message;
[0122] Step SII, collecting node status information and comparing the collected information with a set threshold based on a preset network indicator;
[0123] Step SIII, determining whether the metadata collected by the INT detection packet is greater than a preset threshold; if not, proceeding to step SIV; if yes, proceeding to step SV;
[0124] Step SIV: Telemetry messages that do not exceed the threshold continue to collect telemetry information along the detection path;
[0125] In step SV, the telemetry message node that exceeds the threshold reports the telemetry to the central controller for active notification, and the telemetry message continues to be detected along the path;
[0126] In step SVI, the central controller controls the transmitting end to adjust the telemetry frequency according to the notification information.
[0127] c. Hybrid perception mode switching method based on network scenario information:
[0128] Based on the two INT operating modes and the perception capabilities of network nodes, we first analyze the perception requirements of network scenarios. For example, high-real-time scenarios require high latency in acquiring perception information, while packet loss-resistant scenarios require more frequent acquisition of perception information, meaning more frequent exploration of the perception coverage range. Therefore, we formulate a multi-objective optimization problem based on the requirements of different scenarios, ensuring that appropriate information is selected based on the priority of perception requirements in each scenario.
[0129] Secondly, based on the hybrid telemetry switching model, a set of constraints is constructed around network resource availability constraints, the dependencies of scenario telemetry requirements, the detection path direction, and the perception effect requirements. Taking the packet loss resistance scenario as an example, due to suboptimal network conditions, there is a probability of loss of network status information during the detection process, which affects the overall network telemetry coverage. Therefore, the hop-by-hop upload method is preferred for the initial collection of network status information.
[0130] After collecting network-wide status information, topological regions can be divided based on the node's switching activity and link packet loss. The INT working mode of uploading information hop by hop is maintained for the detection path formed by nodes with frequent switching activity. The INT working mode of uploading information along the path is selected for nodes with good link packet loss and moderate switching status. This completes the in-network status refinement adjustment of the detection working mode, and introduces constraint methods such as path redundancy planning to avoid bandwidth loss and increased congestion caused by repeated detection on congested links.
[0131] During the periodic detection of network information, when a node anomaly occurs, the detection path covering the abnormal node will detect abnormal network resource information, such as a sharp increase in queue depth and an abnormal increase in transmission delay between nodes. Based on the indicators with high priority for network resources in different scenarios, the detection path is adjusted and the working mode is refined.
[0132] Specifically, based on the high priority of queue depth in anti-packet loss scenarios, when an abnormality occurs in a node in the network, that is, the queue depth exceeds the set initial threshold, the node actively perceives the function and uploads information. At the same time, the detection frequency of the abnormal node detection path is increased and the INT working mode of uploading information hop by hop is selected to better monitor the status changes of abnormal nodes and provide real-time information for subsequent adjustments.
[0133] When the detection method is adjusted through multiple feedback of abnormal node information in the network, the node abnormality cannot be alleviated, then only the node active perception function will be selected for the abnormal nodes, and the network status detection path will exclude the abnormal nodes for path planning to prevent the secondary impact of the detection path on the abnormal nodes and ensure the rapid update of node information.
[0134] The method for dynamically adjusting perception based on network status changes is described below in combination with two typical network scenarios:
[0135] Within the adaptive switching framework, considering the complexities of the network and the varying needs for collecting telemetry data, sensing parameter thresholds are set based on flexible switching of sensing modes. Nodes dynamically adjust detection frequencies on demand, enabling dynamic interaction between nodes to expand measurement possibilities and achieve flexible and efficient telemetry tasks. The following application examples illustrate two typical network scenarios with high real-time requirements and varying levels of packet loss.
[0136] a) For scenarios with high real-time requirements: For networks with high real-time requirements, the speed of network information collection is required to be higher. A value can be set as the transmission time between congested nodes as a reference time. The frequency of active node perception can be increased or decreased according to the degree of exceeding or falling below the reference time. In addition, while the nodes are fully covered, the path with the shortest perception time is selected according to the transmission delay between nodes and set as a high priority. The path is dynamically selected during the perception process. Taking the selection of working mode 1 as an example, the working process is described as follows Figure 7 shown.
[0137] Automatic perception switching strategy for scenarios with high real-time requirements:
[0138] In scenarios with high real-time requirements, the network's response speed is crucial, so perception switching should be fast and accurate.
[0139] Dynamic latency monitoring: Monitors latency indicators of all paths in real time and compares them against preset thresholds.
[0140] Dynamic sensing mode switching: When the latency of a specific path exceeds a threshold, the system automatically switches to a higher-frequency sensing mode. Conversely, if the latency falls below the threshold, the detection frequency is reduced to optimize resource usage.
[0141] Priority path selection: Dynamically selects based on real-time data, with latency as the core indicator, and selects multiple non-overlapping detection paths with low latency for detection packet detection path selection to ensure high-speed response for critical services.
[0142] Emergency response mechanism: For sudden high-latency events, the detection frequency is increased, network information is reported multiple times for abnormal nodes, and detection and service path switching are urgently initiated to quickly locate and resolve the problem.
[0143] b) Figure 8 Dynamically adjust the detection path flow chart based on network status changes, such as Figure 8 As shown, for different packet loss scenarios:
[0144] For scenarios with different packet loss severities, based on the INT working mode switching mechanism, congested nodes are selected based on queue depth and excluded from the plan through path redundancy control. INT selects paths passing through nodes with smaller queue depths globally for path planning perception. In addition, congested nodes actively send detection packets one or more times based on different excess ratios, promptly monitoring network information to resolve INT packet loss situations and achieve more accurate perception and measurement results.
[0145] Automatic perception switching strategy for scenarios with different packet loss levels:
[0146] When faced with different levels of packet loss severity, it is particularly important to ensure that the collected network information is not lost.
[0147] Real-time monitoring of queue depth: Continuously monitor the queue depth of each node in the network and adjust the detection strategy based on this.
[0148] Packet loss sensitivity adjustment: Dynamically adjusts the frequency of probe packets and path selection based on queue depth and packet loss rate. For nodes with high packet loss rates, the probe frequency is increased to closely monitor and adjust.
[0149] Path redundancy and selection: When a path with a high packet loss rate is detected, a path redundancy mechanism is automatically introduced to select an alternative path to maintain the comprehensiveness of network information collection.
[0150] Rapid response to packet loss events: Once a packet loss event is detected that exceeds the preset threshold, the node will immediately trigger active detection behavior and quickly adjust the route to reduce the impact.
[0151] Through these automatic sensing switching strategies, the network can dynamically adjust telemetry and path planning based on real-time requirements and packet loss conditions, effectively improving network stability and performance. This dynamic adjustment not only enhances the network's adaptability but also ensures the continuity and efficiency of critical applications.
[0152] c) Changes for different scenarios:
[0153] When the network scenario changes, for example, from a low-latency scenario to a high-packet-loss scenario, the system needs to be able to quickly identify this change and automatically adjust its perception strategy to adapt to the new network conditions, requiring the system to be highly flexible and adaptable. The following example uses the automatic perception strategy adjustment process for a low-latency scenario to change to a high-packet-loss scenario to cope with changes in different scenarios:
[0154] 1. Scene recognition and evaluation:
[0155] Data analysis: The system continuously analyzes collected telemetry data (such as latency, packet loss rate, bandwidth, etc.) and can use statistical methods to predict changing trends in network status.
[0156] Scenario determination: When data shows a significant increase in packet loss and latency is no longer a primary issue, the system automatically identifies it as a high packet loss scenario.
[0157] 2. Strategy switching and parameter adjustment:
[0158] Adjust the detection frequency: In high packet loss scenarios, increase the frequency of sending detection messages to more intensively monitor network status and quickly identify the source of problems.
[0159] Path optimization: Automatically adjust data transmission paths based on historical network status information, giving priority to paths with low packet loss rates. When abnormal nodes are found in the network, redundant paths are added to prevent data loss.
[0160] 3. Fault diagnosis and recovery:
[0161] Real-time diagnosis: The system uses multi-dimensional detection data information to locate the fault point, such as identifying whether the packet loss is caused by network congestion or equipment failure. Network congestion usually occurs when the network traffic is overloaded, and the data packets are discarded because the network path or node exceeds the processing capacity. Packet loss in this case is usually related to high traffic periods or traffic surges on specific paths. The situation can be alleviated by sending perception notifications multiple times by the node and collecting detection path indicators; Equipment failure: This may be due to hardware or software problems in a forwarding device in the network (such as routers, switches), resulting in the inability to process the passing data packets normally. Packet loss caused by equipment failure is usually random and cannot be associated with traffic changes, so it cannot be alleviated by node detection behavior.
[0162] Automatic recovery: After determining the cause of a failure, the system attempts to automatically resolve the problem, such as by adjusting routing to bypass the failed node or automatically triggering recovery measures such as device restarts.
[0163] 4. Performance monitoring and feedback adjustment:
[0164] Continuous monitoring: After adjustments are made, the system will continue to monitor network performance to ensure the effectiveness of the measures taken, such as determining whether information collection is comprehensive and whether the collection latency remains unchanged, while preparing to further fine-tune the strategy.
[0165] Feedback mechanism: The system collects performance data after executing the strategy, which will be fed back into the strategy adjustment algorithm to optimize future strategy decisions.
[0166] By automatically sensing and adjusting the policy process, the network management system can respond to changes in different network scenarios in real time, ensuring stable network operation and service continuity. This dynamic adjustment strategy not only improves the network's adaptability, but also greatly enhances network reliability and user service experience.
[0167] The present invention also provides a network telemetry device for distinguishing scenarios using a network telemetry method, including:
[0168] The control plane is managed by a central controller, while the data plane, where network devices perform actual packet processing, uses a central controller responsible for decision-making for the entire network; the data plane uses P4 programmable switches to support network node functions; and
[0169] Network status collection module, working mode switching module, detection frequency adjustment module and path control module, wherein the working mode switching module and the detection frequency adjustment module are both included in the telemetry control module; wherein,
[0170] The network status collection module is used to collect network scenario information and real-time status information. The network scenario information describes the business scenarios to which the technology is applicable, facilitating the selection of different working modes for path detection and adjustment of detection frequency. The real-time status information is the collected node metadata that accurately reflects the network situation, facilitating timely adjustment of telemetry strategies and ensuring flexible and reliable network detection.
[0171] The working mode switching module selects the telemetry mode according to the data collected by the network status collection module based on the characteristics of different network scenarios, path status and different working modes of INT;
[0172] The detection frequency adjustment module is responsible for receiving active sensing notifications from nodes in the network and adjusting the detection frequency;
[0173] The path control module is responsible for generating multiple detection paths that evenly cover nodes in the network according to the network topology; and replanning the detection paths based on the detection frequency adjustment.
[0174] The network status collection module, working mode switching module, detection frequency adjustment module and path control module are functionally independent but mutually supportive. They continuously adjust detection behavior and provide feedback through network status information to ensure the real-time and accuracy of perception.
[0175] The following combination Figure 9 The specific process of applying the present invention to low-latency scenarios is described as follows:
[0176] This example involves a network optimization framework for dynamically adjusting network paths and telemetry strategies in real time to cope with abnormal situations in the network and ensure key performance indicators such as low latency. The following is an extended description of a specific implementation process of the invention, involving steps such as message sending, receiving and processing. Figure 9 As shown, the present invention can be applied to a variety of scenarios, adjusting the dual perception of the originating device and the nodes within the network based on the characteristics of the scenario. The framework of the present invention includes functional descriptions of the terminal device and the node devices within the network. Taking the low-latency scenario as an example, the demand for real-time communication, autonomous driving, and other services in daily life is increasing. To ensure the reliability and real-time performance of information transmission, this invention can provide support for such scenarios.
[0177] exist Figure 9 The middle and lower-layer forwarding planes collect information and transmit data along the initial detection path. If an anomaly occurs at a node within the network, for example, metadata collected by the detection packet at node S4 indicates that the transmission delay is higher than the set threshold, continuing to select the path passing through this node may cause local congestion within the network.
[0178] Therefore, based on this anomaly, the abnormal node in the forwarding plane notifies the control plane to adjust the telemetry solution. The control plane receives the abnormal information and can perform two operations:
[0179] ① Adjust the path. Adjust the path selected by the previous hop node based on the abnormal node, bypass the abnormal node, and divert telemetry data through the upstream node to alleviate the abnormal situation. You can also design a dedicated detection path for the abnormal node to eliminate other interference and improve the accuracy of information collection at the node;
[0180] ② If the abnormality persists or the number of abnormalities exceeds the initial value, the controller will adjust the fault based on the degree of abnormality. For example, the important indicator in low-latency scenarios is transmission delay. Adjusting a single node for troubleshooting is not enough to solve the problem of normal transmission in the network.
[0181] Therefore, by resetting the detection path and detection frequency at the initiator, abnormal problems can be solved, and active actions can be used instead of passive actions for adjustment, which greatly improves the troubleshooting speed and information collection speed.
[0182] Specific implementation process:
[0183] 1. Initialization phase:
[0184] Transmitting device configuration: The transmitting device presets the detection path according to business requirements and sets the frequency of sending detection messages (for example, once per second).
[0185] Telemetry configuration: Each node in the network is configured with basic telemetry functions, which can intercept the passing detection messages and collect key parameters such as latency and packet loss rate. The telemetry information is inserted into the INT detection message according to the telemetry instructions.
[0186] 2. Sending and collecting detection messages:
[0187] Message generation and transmission: The originating device periodically sends probe messages, each of which contains a unique identifier and a sending timestamp.
[0188] Data collection: Nodes within the network capture passing probe packets, record their latency data and any abnormal indicators, and append this data to the packet metadata.
[0189] 3. Anomaly detection and response:
[0190] Anomaly detection: Each node analyzes the passing detection messages and identifies an anomaly if it finds that scenario-sensitive indicators, such as the latency indicator in a low-latency scenario, exceed the preset threshold.
[0191] Message reporting: The node adds metadata to the abnormal message and sends it to the control plane.
[0192] Control plane analysis and decision-making: After receiving abnormal information, the control plane quickly analyzes it and decides on the next action. The actions include:
[0193] ① Path adjustment: If the anomaly is local, the control plane will instruct the relevant upstream nodes to bypass the abnormal node and select an alternative path for data transmission.
[0194] ② Increase the detection frequency: For persistent or severe anomalies, you may need to increase the detection frequency of the node or path to monitor the problem more carefully.
[0195] ③ Exclusive detection path design: Design exclusive detection paths for abnormal nodes, eliminate other traffic interference, specifically monitor and adjust these nodes, perform redundant path planning for other nodes, and remove abnormal nodes to calculate the path using graph theory optimization path algorithm.
[0196] 4. Dynamic adjustment and optimization:
[0197] Real-time adjustment of detection paths: Based on real-time feedback from network conditions, the control plane can dynamically adjust the detection path.
[0198] Information feedback loop: The adjusted paths and strategies will be fed back to the originators so that they can update their probe settings according to the latest network status.
[0199] 5. Troubleshooting and reporting:
[0200] Implementation of troubleshooting measures: At the same time, paths and nodes are monitored in real time during system operation, and specific troubleshooting measures such as hardware replacement and software upgrades are implemented on nodes or paths where problems are found.
[0201] Performance Report: Generate network performance reports regularly, including overall and local performance indicators of the network, as well as evaluation of the effectiveness of measures taken.
[0202] This detailed implementation process ensures efficient network operation and smooth operation of critical applications (such as real-time communications and autonomous driving), especially in scenarios requiring low latency and high reliability.
[0203] The present invention also provides an electronic device, comprising a processor and a memory; wherein, when the processor executes a computer program stored in the memory, the aforementioned network telemetry method for distinguishing scenarios is implemented.
[0204] The present invention also provides a method for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned network telemetry method for distinguishing scenarios is implemented.
[0205] The proposed scenario-based network telemetry method and device addresses the increasingly complex and diverse development trend of network internal states and the challenges faced by INT solutions in different network scenarios, such as poor perception flexibility, low timeliness of perception information, and insufficient dynamic perception. Combining P4, INT, path planning, and other related technologies, the method and device selects a perception working mode based on the characteristic attributes of different business scenarios and prioritizes requirements to obtain network-wide status information. Through the proactive behavior of nodes within the network, abnormalities are sensitively captured, and real-time perception updates are completed. Ultimately, efficient monitoring of the entire network, real-time assessment of network status, and precise control of network anomalies are achieved. From this perspective, it is difficult to find other alternative solutions to achieve this invention's purpose. However, in the specific implementation process, other telemetry technologies can be selected, such as IOAM (In-band Operation, Administration, and Maintenance) technology, which also has two equivalent working modes to INT technology and is also in-band, and is expected to achieve the same results, demonstrating the scalability of the solution.
[0206] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0207] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0208] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0209] The above is a detailed introduction to a network telemetry method and device for distinguishing scenarios provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A network telemetry method for distinguishing scenarios, characterized in that: Based on the software-defined network (SDN) architecture, the control plane and the data plane are separated. The control plane is managed by a central controller, and the data plane is handled by network devices that perform actual data packet processing. The central controller is responsible for decision-making for the entire network. The data plane uses a P4 programmable switch to support network node functions. The method is executed by a network status collection unit, a working mode switching unit, a detection frequency adjustment unit, and a path control unit, wherein the working mode switching unit and the detection frequency adjustment unit are both included in the telemetry control unit. The network status collection unit is used to collect network scenario information and real-time status information. The network scenario information describes the business scenarios to which the technology is applicable, facilitating the selection of different working modes for path detection and adjustment of detection frequency. The real-time status information is the collected node metadata that accurately reflects the network situation, facilitating timely adjustment of telemetry strategies and ensuring flexible and reliable network detection. The working mode switching unit selects the telemetry mode according to the data collected by the network status collection module, based on the characteristics of different network scenarios, path status and different working modes of the in-band network telemetry technology INT; The detection frequency adjustment unit is responsible for receiving active sensing notifications from nodes in the network, thereby adjusting the detection frequency; The path control unit is responsible for generating multiple detection paths that evenly cover nodes in the network according to the network topology; and replanning the detection paths based on the detection frequency adjustment; The specific workflow includes the following steps: Step S1: Initial path planning is performed based on network scenario attributes, with the goal of covering all network links. The network collection unit obtains all network information. Step S2, the path control unit generates multiple detection paths that evenly cover nodes in the network according to the network topology based on the information fed back by the network collection unit; Step S3: Based on the information collected in step S1, select a working mode according to different network scales or different network scenarios; Step S4: After the network information is collected, it is sent to the network status collection unit for storage in the network scenario information; using the network resource information, the detection method is further adjusted; services are provided for detection path re-planning, detection frequency optimization, and node detection action adjustment; The network status collection unit, working mode switching unit, detection frequency adjustment unit and path control unit are functionally independent but mutually supportive. They continuously adjust the detection behavior and provide feedback through network status information to ensure the real-time and accuracy of perception.
2. A network telemetry method for distinguishing scenarios according to claim 1, characterized in that: The nodes in the network are all equipped with the function of autonomously detecting changes in network conditions between nodes. The specific method is as follows: Step S1, receiving INT telemetry message; Step SII, collecting node status information and comparing the collected information with a set threshold based on a preset network indicator; Step SIII, determining whether the metadata collected by the INT detection packet is greater than a preset threshold; if not, proceeding to step SIV; if yes, proceeding to step SV; Step SIV: Telemetry messages that do not exceed the threshold continue to collect telemetry information along the detection path; In step SV, the telemetry message node that exceeds the threshold reports the telemetry to the central controller for active notification, and the telemetry message continues to be detected along the path; In step SVI, the central controller controls the transmitting end to adjust the telemetry frequency according to the notification information.
3. The network telemetry method for distinguishing scenarios according to claim 1, characterized in that: The working mode switching unit executes a hybrid perception mode switching method based on network scenario information, or dynamically adjusts the perception method based on changes in network status. The hybrid perception mode switching method based on network scenario information is specifically as follows: First, we analyze the perception requirements of network scenarios and construct multi-objective optimization problems based on different scenario requirements. In each scenario, we can select reasonable information based on the perception priority. Secondly, based on the hybrid telemetry switching model, a set of constraints is constructed around network resource availability constraints, dependency constraints on scenario telemetry requirements, detection path direction, and perception effect requirements. The method for dynamically adjusting perception based on network status changes is specifically as follows: In the adaptive switching framework, considering the complex internal situation of the network and the different needs of collecting telemetry data, the perception parameter thresholds are set based on the flexible switching of the perception working mode. The detection frequency is adjusted dynamically on demand by the nodes, so that dynamic interaction between nodes can realize more measurement possibilities and achieve flexible and efficient telemetry tasks.
4. The network telemetry method for distinguishing scenarios according to claim 1, characterized in that: In the step S4, Replanning of the detection path includes, but is not limited to, utilizing network scenario information and real-time feedback on the link transmission status and node queue depth provided by real-time recovered network status information, or other real-time feedback on significant network conditions, to ensure high coverage of network information detection links and high information volume obtained from network resources in accordance with scenario requirements; Alternatively, for abnormal nodes, a separate detection path is set up, and a hop-by-hop detection mode is adopted to ensure the accuracy and completeness of detection information collection. In addition, to prevent node deterioration from being caused by the detection behavior itself, path redundancy operations are performed for the node directly connected to the upstream node to avoid excessive detection traffic flowing through the abnormal node.
5. The network telemetry method for distinguishing scenarios according to claim 1, characterized in that: In the step S4, Detection frequency optimization is based on the initial detection frequency and performs real-time updates on the detection frequencies of different paths according to the speed of network information update.
6. The network telemetry method for distinguishing scenarios according to claim 1, characterized in that: In the step S4, The node detection action adjustment is mainly aimed at switching the detection method and selecting the detection action of abnormal nodes in the network. It can select and switch the appropriate working mode for the node according to the changes in the network scenario or the changes in the conditions of a certain link. In addition, because the nodes in the network have active perception functions, they can also detect anomalies in the network in time and take active information upload actions, thereby improving the whole system's ability to handle abnormalities.
7. A network telemetry device using the network telemetry method for distinguishing scenarios according to any one of claims 1 to 6, characterized in that: include: The control plane is managed by a central controller, while the data plane, where network devices perform actual packet processing, uses a central controller responsible for decision-making for the entire network; The data plane uses P4 programmable switches to support network node functions; and, Network status collection unit, working mode switching unit, detection frequency adjustment unit and path control unit.
8. An electronic device, characterized in that: It comprises a processor and a memory; wherein, when the processor executes the computer program stored in the memory, it implements the network telemetry method for distinguishing scenarios as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that Used to store computer programs; wherein, when the computer program is executed by a processor, it implements the network telemetry method for distinguishing scenarios as described in any one of claims 1 to 6.