Network equipment edge detection method and system based on heartbeat detection mechanism
By employing a dynamic detection strategy based on a heartbeat detection mechanism and in-band telemetry data fusion, the problem of rigidity and single fault recovery in traditional network device detection in edge computing scenarios is solved. This enables real-time status awareness and efficient fault location of network devices, improving their adaptability and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BRINGSPRING SCIENCE & TECHNOLOGY CO LTD
- Filing Date
- 2026-04-03
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional network device detection technologies in edge computing scenarios suffer from rigid frequency settings and limited fault recovery strategies, failing to meet the demands for network reliability, real-time performance, and adaptability in scenarios such as medical IoT and industrial internet.
By generating dynamic detection strategies based on a heartbeat detection mechanism, fusing in-band telemetry data, and optimizing network behavior profiles in a closed loop, a dynamic detection strategy is generated that includes target device identification, detection signaling type, and detection mode. The detection frequency and mode are adjusted in real time, and combined with network context and behavior profiles, real-time perception of network status and fault location are achieved.
It achieves nanosecond-level network status awareness, improves the real-time performance, adaptability, and fault location accuracy of detection, optimizes resource utilization, forms a closed-loop mechanism of detection-analysis-optimization, and enhances the adaptability and long-term stability of network devices.
Smart Images

Figure CN121967281A_ABST
Abstract
Description
A method and system for edge detection of network devices based on heartbeat detection mechanism Technical Field
[0001] This invention relates to the field of network device detection technology, and in particular to a method and system for edge detection of network devices based on a heartbeat detection mechanism. Background Technology
[0002] In the field of traditional network device detection, mainstream technologies have significant limitations. Rule-based intrusion detection systems (IDS) rely on static policies, making it difficult to adapt to dynamic changes in network topology and new attack patterns, resulting in high false negative and false positive rates. While periodic polling mechanisms can monitor device status, they cannot capture sudden network performance changes in real time, such as microsecond-level latency fluctuations or sudden packet loss. In-band telemetry (INT) can provide fine-grained network status information, but traditional implementations lack dynamic adjustment capabilities. For example, fixed probe frequencies can lead to bandwidth waste or delayed status perception, and they do not form a closed-loop optimization mechanism with network behavior profiling.
[0003] In edge computing scenarios, devices are widely distributed and resources are limited. Traditional heartbeat detection mechanisms suffer from rigid frequency settings and simplistic fault recovery strategies. For example, KubeEdge edge node health checks only determine offline status based on timeout thresholds, without dynamically adjusting the detection mode based on network context. Scenarios such as medical IoT and industrial internet have extremely high requirements for network reliability. Traditional static detection strategies can no longer meet the demands for real-time performance, adaptability, and precise positioning. There is an urgent need for a technical solution that can dynamically generate detection strategies, fuse telemetry data in real time, and update network behavior profiles.
[0004] Therefore, it is necessary to provide a network device edge detection method and system based on a heartbeat detection mechanism to solve the above-mentioned technical problems. Summary of the Invention
[0005] To address the aforementioned technical issues, this invention provides a network device edge detection method and system based on a heartbeat detection mechanism. By generating dynamic detection strategies, fusing in-band telemetry data, and optimizing network behavior profiles in a closed loop, the real-time performance, adaptability, and fault location accuracy of network device detection are improved.
[0006] This invention provides a network device edge detection method based on a heartbeat detection mechanism. The detection method includes the following steps: generating a dynamic detection strategy based on a network behavior profile and network context, and allocating it to target detection nodes in an edge detection device cluster, wherein the dynamic detection strategy includes a target device identifier, a detection signaling type, and a detection mode; the target detection node constructs a detection packet corresponding to the detection mode according to the dynamic detection strategy, and encapsulates signaling data including the detection signaling type in the detection packet, and sends it to the target device; the network device carrying the forwarding of the detection packet identifies the detection signaling type, and when forwarding a response packet returned by the target device, adds the network device's in-band telemetry data to the response packet for reporting; receiving a response packet including the in-band telemetry data, associating the arrival information of the response packet with the in-band telemetry data by time and task, and generating fused network state information; updating the network behavior profile according to the network state information, and adjusting the detection mode and detection signaling type in the dynamic detection strategy for the next cycle based on the updated network behavior profile.
[0007] Preferably, the step of generating a dynamic detection strategy based on the network behavior profile and network context and allocating it to target detection nodes in the edge detection device cluster includes: acquiring the network behavior profile including historical detection data and network event logs, and the network context including network topology and device attribute information; calculating the detection requirements for different target devices based on the network behavior profile and the network context, and generating the dynamic detection strategy including target device identifier, detection signaling type and detection mode; evaluating the cost of each detection node executing the dynamic detection strategy based on the load status and network location of each detection node in the edge detection device cluster, and allocating the dynamic detection strategy to the target detection node with the lowest evaluated cost.
[0008] Preferably, the target detection node constructs a detection packet corresponding to the detection mode according to the dynamic detection strategy, and encapsulates signaling data including the detection signaling type in the detection packet, and sends it to the target device, including: selecting a corresponding detection protocol from a preset protocol library based on the detection mode; generating an initial detection packet carrying a task identifier according to the selected detection protocol; writing the signaling instruction corresponding to the detection signaling type into the initial detection packet to generate an in-band telemetry enabled detection packet; and tagging the detection packet with a corresponding target VLAN tag according to the target device identifier, and sending it by the target detection node.
[0009] Preferably, the network device carrying the probe message forwarding identifies the probe signaling type and, when forwarding the response message returned by the target device, adds the network device's in-band telemetry data to the response message for reporting. This includes: determining a collection template based on the probe signaling type, wherein the collection template defines data fields to be collected, the data fields including at least the port identifier, timestamp, and queue status involved when the network device forwards the response message; when forwarding the response message, collecting the real-time port identifier, timestamp, and queue status from the network device according to the definition of the collection template, and combining them to generate the in-band telemetry data; binding the task identifier in the response message with the in-band telemetry data and writing it into the selected in-band telemetry extension field in the response message, and reporting it through the network device.
[0010] Preferably, the step of receiving a response message including the in-band telemetry data, and associating the arrival information of the response message with the in-band telemetry data by time and task to generate fused network state information includes: matching the corresponding initial probe task record in a preset policy task library based on the task identifier to obtain the target device identifier, the probe signaling type, and the probe mode; normalizing each local timestamp in the sequence of in-band telemetry data formation based on the task generation timestamp in the initial probe task record, and aligning the local timestamp with the final timestamp of the response message arrival to reconstruct the timing state vector of the response message on the transmission path; and calculating and generating the fused network state information including connectivity state features, hop-by-hop performance features, and topology features based on the timing state vector, the probe mode, the port identifier and queue status in the in-band telemetry data.
[0011] Preferably, the step of updating the network behavior profile based on the network state information and adjusting the detection mode and detection signaling type in the dynamic detection strategy for the next cycle based on the updated network behavior profile includes: updating the state profile of the corresponding target device in the network behavior profile based on the connectivity state characteristics, and updating the performance baseline profile and structure profile of the corresponding network path in the network behavior profile based on the hop-by-hop performance characteristics and the topology characteristics; calculating the detection mode adjustment parameters and detection signaling type adjustment parameters for the target device in the next cycle based on the updated state profile, the performance baseline profile, and the structure profile through a preset detection strategy adjustment model; and inputting the detection mode adjustment parameters and the detection signaling type adjustment parameters as constraints into the dynamic detection strategy generation process for the next cycle to generate an optimized dynamic detection strategy adapted to the current network state.
[0012] This invention also provides a network device edge detection system based on a heartbeat detection mechanism, used to execute the aforementioned network device edge detection method based on a heartbeat detection mechanism. The detection system includes: a policy generation and allocation module, used to generate dynamic detection policies based on network behavior profiles and network context, and allocate them to target detection nodes in an edge detection device cluster, wherein the dynamic detection policy includes a target device identifier, a detection signaling type, and a detection mode; and a detection packet construction module, used by the target detection node to construct a detection packet corresponding to the detection mode according to the dynamic detection policy, and encapsulate signaling data including the detection signaling type in the detection packet, and send it to the target device. The telemetry data injection module is used to carry the network device that forwards the probe message, identify the probe signaling type, and add the network device's in-band telemetry data to the response message when forwarding the response message returned by the target device. The state information fusion module is used to receive the response message including the in-band telemetry data, associate the arrival information of the response message with the in-band telemetry data in terms of time and task, and generate fused network state information. The profile update and adjustment module is used to update the network behavior profile according to the network state information, and adjust the probe mode and probe signaling type in the dynamic probe strategy for the next cycle based on the updated network behavior profile.
[0013] Compared with related technologies, the network device edge detection method and system based on heartbeat detection mechanism provided by the present invention have the following beneficial effects: by dynamically adjusting the detection frequency through the heartbeat detection mechanism, and combining in-band telemetry data to capture network latency, packet loss, queue status and other indicators in real time, nanosecond-level network status perception is achieved.
[0014] Based on network behavior profiles and context, detection strategies are dynamically generated, detection signaling types are adjusted according to device load, and detection modes are optimized according to network topology changes, thereby improving detection efficiency and resource utilization.
[0015] By associating response messages with telemetry data in terms of time and task, the end-to-end timing state vector can be reconstructed to accurately locate network fault points, such as switch queue congestion or link jitter.
[0016] By updating network behavior profiles based on the fused network status information, a closed-loop mechanism of "detection-analysis-optimization" is formed, which improves the adaptive capabilities and long-term stability of network devices. Attached Figure Description
[0017] Figure 1 is a flowchart of a network device edge detection method based on a heartbeat detection mechanism provided by the present invention; Figure 2 is a module structure diagram of a network device edge detection system based on a heartbeat detection mechanism provided by the present invention. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the drawings, not all structures. Moreover, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0019] It should also be noted that, for ease of description, the accompanying drawings show only the parts relevant to the invention and not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but it may also have additional steps not included in the drawings. The process may correspond to a method, function, procedure, subroutine, subroutine, etc.
[0020] Example 1: This invention provides a network device edge detection method based on a heartbeat detection mechanism. Referring to Figure 1, the detection method includes the following steps: S1: Based on network behavior profile and network context, a dynamic detection strategy is generated and allocated to the target detection nodes in the edge detection device cluster. The dynamic detection strategy includes the target device identifier, detection signaling type, and detection mode.
[0021] Specifically, step S1 includes the following sub-steps: S11: Obtain the network behavior profile including historical probe data and network event logs, and the network context including network topology and device attribute information.
[0022] In this embodiment, this step is the data acquisition and fusion stage, designed to provide a comprehensive data foundation for policy generation. Specifically, the system acquires data from various components in the network through multiple interfaces and protocols.
[0023] First, network behavior profiles are obtained from the database of the central network data analysis module. These profiles primarily contain two types of time-series data: one is historical probe data, which records, in time series form, the response status (success / timeout), response time to various probe packets (such as ICMP, TCP), and packet loss rate of each target device within a specific past period (e.g., 24 hours). This data reveals the periodic activity patterns of the devices and the stability of the network path. The second is network event logs, which systematically record key events affecting network stability, such as device interface oscillations, frequent changes in MAC address entries, and BGP session interruptions, along with their timestamps.
[0024] Secondly, network context is obtained from the Network Management System (NMS) or SDN controller. This includes static and dynamic information: the network topology describes the physical and logical connections between network devices such as switches and routers, usually obtained through LLDP (Link Layer Discovery Protocol) or configuration information issued by the controller; device attribute information is extracted from the CMDB (Configuration Management Database) or the device's SNMPMIB (Management Information Base), containing metadata such as the device's IP address, MAC address, VLAN, device type (e.g., server, camera, IoT sensor), and service importance level.
[0025] Finally, the system aligns and correlates the aforementioned multi-source data on timestamps to form a unified network state view with a time dimension, preparing for subsequent intelligent analysis.
[0026] S12: Based on the network behavior profile and the network context, calculate the detection requirements for different target devices, and generate the dynamic detection strategy including target device identifier, detection signaling type and detection mode.
[0027] In this embodiment, this step is the core of the strategy decision-making process. Based on the fused data, a customized detection scheme for each device is calculated through a predefined strategy engine or a lightweight algorithm.
[0028] The specific process is as follows: The policy engine first assesses the urgency and granularity requirements of the target device's probe. For example, for a device with highly stable historical behavior and low business importance (such as an internal printer), the calculated requirement is "basic connectivity monitoring". Therefore, in the generated policy, the probe mode can be selected as low-frequency ICMP_PING, and the probe signaling type is set to BASIC_LIVENESS, which only needs to collect basic connectivity status.
[0029] Conversely, for a core server that is critical to business operations and whose historical logs show that its path has recently experienced microsecond-level latency fluctuations, the calculated requirement is "high-performance deep monitoring".
[0030] At this time, the probe mode will select high-frequency TCP_SYN probe that can simulate real business, and may be combined with UDP_JITTER probe to measure jitter; the probe signaling type will be set to ADVANCED_PERFORMANCE, which requires network devices to report detailed performance indicators. Its instructions will require the collection of more detailed data, such as egress queue depth, link utilization, etc.
[0031] Furthermore, for newly joined unknown devices (whose device attribute information indicates an "unknown" type), the engine initializes a "high-frequency, multi-protocol" discovery phase policy. Ultimately, a separate policy record is generated for each target device identifier (i.e., its IP address), and all these records together constitute a policy set, namely the dynamic detection policy.
[0032] S13: Based on the load status and network location of each detection node in the edge detection device cluster, evaluate the cost of each detection node executing the dynamic detection strategy, and allocate the dynamic detection strategy to the target detection node with the lowest evaluated cost.
[0033] In this embodiment, this step is the strategy optimization and scheduling phase, aiming to efficiently and rationally allocate probing tasks to the most suitable hardware probes. The system maintains a real-time status table of an edge probing device cluster. For each probe node in the cluster, the system continuously monitors its load status, including but not limited to: current CPU utilization, memory usage, and the number of probe tasks being executed concurrently.
[0034] At the same time, based on the network topology, the system calculates the network location relationship between each probe node and the IP address represented by each target device identifier, which is usually measured by logical hop count or predefined network areas (such as "same data center AZ" or "across backbone network").
[0035] Next, the system calculates the execution cost of a probe strategy to be assigned (e.g., strategy P for target device D) on each candidate probe node N. The cost function (C) can be simplified to: C = α * (current load rate of node N) + β * (network distance from node N to target D), where α and β are configurable weight coefficients.
[0036] The system calculates the cost C of all candidate nodes for policy P, and then selects the node with the lowest cost as the target probe node. For example, if a probe node is very close to the target device (small network distance), but its current CPU load exceeds 80%, its total cost may be higher than that of a node with a lighter load but a slightly farther distance. Through this policy allocation mechanism that combines dynamic load balancing with proximity service, the efficiency and reliability of the entire probe cluster are ensured, and single-point overload is avoided.
[0037] Ultimately, the system uses internal API calls or message queues (such as RabbitMQ and Kafka) to explicitly direct the issuance of policy P to the selected target probe node.
[0038] S2: The target detection node constructs a detection message corresponding to the detection mode according to the dynamic detection strategy, and encapsulates signaling data including the detection signaling type in the detection message and sends it to the target device.
[0039] Specifically, step S2 includes the following sub-steps: S21: Based on the detection mode, select the corresponding detection protocol from the preset protocol library.
[0040] In this embodiment, after receiving the dynamic detection strategy, the target detection node first parses the "detection mode" field in the strategy.
[0041] This node maintains a configurable protocol library, which is essentially a mapping table that maps different probe mode enumeration values (such as ICMP_PING, TCP_SYN_HEALTH, UDP_JITTER) to specific protocol stack construction parameters.
[0042] For example, when the probe mode is ICMP_PING, the protocol library instructs the node to use the network layer ICMP protocol and specifies the message type as EchoRequest with a code of 0. If the probe mode is TCP_SYN_HEALTH, the protocol library instructs the use of the transport layer TCP protocol, specifies the target port (such as port 80 commonly used for web services or a custom health check port), and sets the SYN bit in the TCP flags to 1 to simulate a TCP connection establishment request. For UDP_JITTER mode, the protocol library instructs the use of the UDP protocol and may associate a specific payload pattern or sequence number generation rule for subsequent jitter calculation. By querying this protocol library, the target probe node translates the abstract probe mode instructions into specific, executable network protocol stack operation instructions, laying the foundation for the precise construction of subsequent messages.
[0043] S22: Generate an initial probe message carrying a task identifier according to the selected probe protocol.
[0044] In this embodiment, after the detection protocol is determined, the target detection node calls its operating system kernel or built-in RawSocket function to begin constructing the initial detection message.
[0045] First, the node generates a globally unique task identifier. This identifier is typically composed of a random number (or a sequence number based on the node ID) and the high-order bits of the current system timestamp to ensure its uniqueness and traceability.
[0046] Subsequently, the node fills in the various fields of the message according to the protocol specifications determined in step S21: For the IP layer, the source IP is set to the interface IP of the probe node in the target VLAN, the destination IP is the target device identifier (IP address) specified in the policy, and an appropriate TTL value is set. For the transport layer, the port number, flag bits, etc., are filled in according to the protocol type. A crucial step is to embed this task identifier into the message.
[0047] The embedding location can be: 1) the Identifier and Sequence Number fields of the ICMP message; 2) the source port number of the TCP or UDP message (mapped to the task ID via a specific port segment) or a specific data structure defined at the beginning of its payload to carry the task ID. This task identifier serves as the "identity card" for this probe task and will play a core role in subsequent response correlation and data analysis.
[0048] S23: Write the signaling instruction corresponding to the probe signaling type into the initial probe message to generate an in-band telemetry enabled probe message.
[0049] In this embodiment, this step is crucial for implementing in-band network telemetry (INT). The target probe node parses the "probe signaling type" (such as BASIC_LIVENESS or ADVANCED_PERFORMANCE) in the strategy. The node internally pre-sets signaling instruction sets corresponding to different signaling types. These instructions are used to inform network devices along the path which data needs to be collected and reported.
[0050] For example, the instruction for BASIC_LIVENESS might be very simple, only requiring the network device to add its device ID and ingress port information when forwarding the response packet. The instruction for ADVANCED_PERFORMANCE, however, would be more complex, requiring the device to report detailed performance data such as timestamps, egress queue lengths, and link utilization.
[0051] These signaling instructions are encoded and written to specific locations in the initial probe message. In IPv4 networks, the Options field can be used; in IPv6 networks, the Hop-by-Hop OptionsHeader can be used; or tunnel encapsulation headers that support extensible metadata, such as Geneve or VXLAN-GPE, can be used.
[0052] By writing these instructions, the initial probe message is "enabled" to be an in-band telemetry enabled probe message, which not only performs connectivity checks but also issues a detailed "data acquisition task list" to the network device.
[0053] S24: Based on the target device identifier, mark the target VLAN tag corresponding to the probe message, and send it by the target probe node.
[0054] In this embodiment, the node queries its local VLAN routing information or determines the target VLAN ID to which the IP address belongs by querying the central controller based on the target device identifier (IP address) in the policy.
[0055] Subsequently, the node adds the corresponding IEEE 802.1Q VLAN tag to the constructed probe packet on the network interface on which it sends the packet.
[0056] Specifically, a 4-byte 802.1Q tag is inserted between the source MAC address and Ethernet type fields in the Ethernet frame header. The VLAN Identifier (VID) field in this tag is set to the queried target VLAN ID. This tagging operation is typically performed in kernel mode by a VLAN-enabled network interface card driver or a virtual switch (vSwitch) to ensure efficiency. Once a packet is tagged with a VLAN tag, it becomes a frame that can be broadcast or unicast within the specified VLAN.
[0057] Finally, the target probe node sends a complete probe message with a VLAN tag and containing the task ID and INT command to the network through the specified physical or virtual network interface, thereby initiating a probe task that combines connectivity checks and deep telemetry for the target device within a specific VLAN.
[0058] S3: The network device carrying the probe message forwarding identifies the probe signaling type and, when forwarding the response message returned by the target device, adds the network device's in-band telemetry data to the response message for reporting.
[0059] Specifically, step S3 includes the following sub-steps: S31: Determine the collection template according to the probe signaling type, wherein the collection template defines data fields to be collected, and the data fields include at least the port identifier, timestamp and queue status involved when the network device forwards the response message.
[0060] In this embodiment, when a network device that supports in-band network telemetry (INT) (such as a switch that supports the P4 programmable data plane or a traditional switch that supports specific INT functions) receives a response message returned by the target device, it first checks whether the message header contains a specific INT instruction.
[0061] The device has a series of pre-set acquisition templates or dynamically issued through the control plane. These templates correspond one-to-one with different probe signaling types (such as BASIC_LIVENESS or ADVANCED_PERFORMANCE). After parsing the probe signaling type carried in the response message, the device immediately indexes the corresponding acquisition template. This template is essentially a data structure or configuration instruction set, which explicitly defines the metadata fields to be acquired and their format.
[0062] For example, for the BASIC_LIVENESS signaling type, its collection template may only define two mandatory fields: ingress_port_id (the port identifier of the packet entering the device) and switch_id (the device's own identifier). However, for the ADVANCED_PERFORMANCE signaling type, its collection template defines a richer set of fields, including at least: ingress_port_id (ingress port identifier), egress_port_id (the egress port identifier from which the packet will be forwarded), timestamp (the precise timestamp of the device processing the packet, usually taken from the hardware clock), egress_port_tx_utilization (egress port transmission direction utilization), queue_id (the identifier of the queue in which the packet resides), queue_occupancy (the current queue occupancy depth), and congestion_status (congestion status flag), etc.
[0063] S32: When forwarding the response message, according to the definition of the collection template, collect the real-time port identifier, the timestamp and the queue status from the network device, and combine them to generate the in-band telemetry data.
[0064] In this embodiment, after the acquisition template is determined, the network device synchronously performs data acquisition actions in its data plane pipeline for forwarding and processing the response message. This process is real-time and inline, typically completed in hardware or a high-speed forwarding ASIC to minimize the impact on forwarding performance. Just before the message is about to be sent out from the egress port, the device's data plane reads the required information in real time from various hardware registers or status counters according to the template instructions.
[0065] Specifically, it collects: port identifier (obtained directly from the ingress and egress port information of the packet), timestamp (obtained from the current time from a high-precision hardware clock), and queue status (read from the management unit corresponding to the egress queue into which the packet is scheduled, such as queue ID and current cache depth).
[0066] The device then combines the raw data collected, representing the device's current state, into one or more in-band telemetry data tuples in memory, according to a predefined template order and format (e.g., using TLV (Type-Length-Value) format or some compact binary encoding). Each tuple corresponds to one acquisition action and contains the values of all specified fields acquired from the same device.
[0067] S33: Bind the task identifier in the response message to the in-band telemetry data and write it into the selected in-band telemetry extension field in the response message, and report it through the network device.
[0068] In this embodiment, after generating local in-band telemetry data tuples, the network device needs to associate them with specific probe tasks and report them. The device first extracts the key task identifier from the received response message. This identifier is usually located at a specific location in the message, such as in the IP options field, at a specific offset in the TCP / UDP payload, or in fields such as the ICMP sequence number mentioned above.
[0069] This identifier is extracted to ensure that subsequent analysis systems can accurately determine which specific probe mission these telemetry data belong to.
[0070] Next, the device performs data binding and writing operations. It associates the in-band telemetry data tuples generated by the device with the extracted task identifier (for example, by adding a task identifier field to the header of the tuple data block), and then appends this data block bound to the task identifier to a specific in-band telemetry extension field in the response message.
[0071] In IPv6 networks, this is typically achieved through hop-by-hop options extension headers; in IPv4, it might utilize IP options fields or encapsulation methods defined by the INT specification (e.g., adding an INT metadata tail to the original packet payload). If the packet length exceeds the MTU due to added data, the device may need to fragment the packet or employ other encapsulation strategies. Finally, after modifying the packet, the network device, following the normal routing and forwarding process, sends this response packet, now "attached" with its real-time status information and the original task identifier, to the next-hop device or ultimately returns it to the probe source (analysis system). In this way, each INT-enabled network device on the path "stamps" its status onto the packet, achieving distributed, in-band data collection and reporting.
[0072] S4: Receive a response message including the in-band telemetry data, associate the arrival information of the response message with the in-band telemetry data by time and task, and generate fused network status information.
[0073] Specifically, step S4 includes the following sub-steps: S41: Based on the task identifier, match the corresponding initial probe task record in the preset policy task library to obtain the target device identifier, the probe signaling type and the probe mode.
[0074] In this embodiment, when the central analysis platform (or the designated data collector) receives a response message from the network that embeds in-band telemetry data, it first performs a task association operation.
[0075] The platform extracts a unique task identifier from specific locations in the response message (such as the ICMP sequence number field, a specific structure in the TCP / UDP payload, or the IP options field).
[0076] Subsequently, the platform queries its pre-built strategy task library. This task library is a temporary or persistent database (e.g., using Redis or a relational database), whose records are created synchronously when the dynamic probing strategy is generated and sent to the target probing node in step S1.
[0077] Each record must contain at least the following key metadata: task identifier, target device identifier (IP address), issued probe mode, probe signaling type, task generation timestamp, and target probe node ID that executed the task.
[0078] By matching the task identifier in the response message with the records in the task library (for example, by executing an SQL query: SELECT * FROM task_library WHERE task_id=?), the platform can accurately reconstruct the original context information of this probe task, that is, to obtain which target device this probe is targeting (target device identifier), what probe method is used (probe mode), and the expected data granularity (probe signaling type).
[0079] S42: Using the task generation timestamp in the initial detection task record as a reference, normalize each local timestamp in the in-band telemetry data formation sequence, and align the local timestamp with the final timestamp of the response message arrival to reconstruct the timing state vector of the response message on the transmission path.
[0080] In this embodiment, after successfully associating the task, the platform begins time normalization and path reconstruction. The in-band telemetry data carried in the response message is a sequence of data tuples added sequentially by each network device along the path. Each tuple contains the local timestamp of the device when it collected the data. Since there may be slight deviations in the clocks of different devices in the network, directly using these raw timestamps to calculate the latency would be inaccurate.
[0081] Therefore, the platform first uses the task generation timestamp obtained from the task library ( This serves as the absolute time reference. Then, it parses the in-band telemetry data sequence. This sequence consists of data tuples added sequentially by multiple network devices along the path. For the first tuple in the sequence... The data tuple (corresponding to the first data tuple on the path) (Network device), read the local timestamp of the data collected by the device, and record it as The platform calculates... , obtained the The network device timestamp relative to the task start time base ( offset of ) This completes the normalization of the local timestamps of all devices.
[0082] This calculation process implicitly assumes the synchronization of the clocks of each device. Ideally, the network should use NTP or PTP for time synchronization.
[0083] Next, the platform records the final timestamp of the response message arriving at the analysis platform itself. By analyzing the entire sequence The increasing relationship and The platform can reconstruct the transmission sequence of packets in the network. For example, it can calculate the approximate absolute time for a packet to arrive at the i-th device on the path and infer the packet's dwell time between adjacent devices (approximately the dwell time between adjacent devices). The difference between the two and the total end-to-end latency ( All of this timing information together constitutes a timing state vector describing the timing state of a message at each hop along the path.
[0084] S43: Based on the time-series state vector, the detection mode, the port identifier and queue status in the in-band telemetry data, calculate and generate the fused network state information, which includes connectivity state features, hop-by-hop performance features and topology features.
[0085] In this embodiment, the platform finally performs multi-dimensional data fusion and feature extraction to generate a final, highly readable network state insight. This process integrates all the information from the first three steps: 1. Connectivity state feature calculation: Based on whether the response message was successfully received and its total end-to-end latency (obtained from the timing state vector), the connectivity state of the target device (e.g., online, offline, high latency) is directly determined. This is the most basic detection result.
[0086] 2. Hop-by-hop performance characteristic calculation: Using the time-series state vector, the transmission delay of a packet between each hop on the path is calculated (the timestamp of the next hop device minus the timestamp of the previous hop device). Combined with the queue status reported by each device in the in-band telemetry data (such as queue occupancy depth), the specific causes of the delay can be further analyzed (for example, if the delay of a hop suddenly increases and the queue depth reported by that device is very high, it indicates that there may be congestion at that point). Simultaneously, performance metrics such as overall latency jitter and packet loss rate (if sequence numbers are supported) can be calculated.
[0087] 3. Topology Feature Extraction: Parse the port identifiers and device identifiers in the in-band telemetry data sequence. Based on this information, the actual forwarding path traversed by this probe packet can be accurately plotted, i.e., a topology segment composed of device identifiers and connected ports.
[0088] For example, the path might be: → → → This enables real-time discovery and verification of network Layer 2 / Layer 3 topology.
[0089] Finally, the platform encapsulates the calculated connectivity state features (device level), hop-by-hop performance features (path level), and topology features (network level) into a structured, fused network state information object (e.g., a JSON or Protobuf data structure).
[0090] S5: Update the network behavior profile based on the network status information, and adjust the detection mode and detection signaling type in the dynamic detection strategy for the next cycle based on the updated network behavior profile.
[0091] Specifically, step S5 includes the following sub-steps: S51: Based on the connectivity state characteristics, update the state profile of the corresponding target device in the network behavior profile, and based on the hop-by-hop performance characteristics and the topology characteristics, update the performance baseline profile and structure profile of the corresponding network path in the network behavior profile.
[0092] In this embodiment, the platform uses the fused network state information generated in step S43 as input to update the three core components of the network behavior profile.
[0093] First, based on the status profile of the target device, the platform extracts connectivity status features from the network status information (such as the online / offline status and response latency in this probe).
[0094] Then, it queries the historical status records of the target device in the network behavior profile, using the latest N (e.g., the most recent 100) probe results stored in a time-series database or circular buffer. The platform uses algorithms such as a sliding window model or exponentially weighted moving average to recalculate the device's average online rate, average response latency, and jitter range, and updates its status profile accordingly. For example, if the device's response latency has been continuously increasing recently, the "baseline latency" value in its profile will be increased accordingly.
[0095] Secondly, based on the performance baseline profile of the network path, the platform extracts hop-by-hop performance features (such as latency, jitter, and packet loss rate for each hop). For this probe path (from the source probe to the target device), the platform maintains the statistical distribution of its historical performance indicators (such as mean, standard deviation, and percentiles). The hop-by-hop performance data obtained in this probe will be included in the historical data set of this path, and its performance baseline will be recalculated. For example, the "normal latency range" of the path will be updated to the 95% confidence interval of the historical data. This allows the system to perceive the long-term gradual trend of network performance.
[0096] Finally, based on the network path structure profile, the platform extracts topological features (i.e., the actual sequence of devices and ports through which packets pass). This actual path is compared with the expected paths or historically common paths stored in the profile. If a path change is detected (e.g., a switch no longer appears in the path, or a new intermediate device appears), the structure profile is updated to reflect the latest network topology connections, and may be marked as a "topology change event" for subsequent analysis.
[0097] S52: Based on the updated state profile, performance baseline profile, and structure profile, the detection mode adjustment parameters and detection signaling type adjustment parameters for the target device in the next cycle are calculated using a preset detection strategy adjustment model.
[0098] In this embodiment, the preset detection strategy adjustment model is a rule-based expert system. The model uses the updated state profile, performance baseline profile, and structural profile as input features.
[0099] Example of a rules engine: The model has a built-in series of policy adjustment rules.
[0100] For example, rule 1: IF (the status profile of a device shows that its recent offline count exceeds the threshold) THEN (increase the detection frequency in the next cycle, i.e., adjust the detection mode parameters).
[0101] Rule 2: IF (The performance baseline profile of a certain path shows a significant increase in its jitter variance) THEN (Upgrade the probe signaling type from BASIC_LIVENESS to ADVANCED_PERFORMANCE to capture more detailed performance data).
[0102] Rule 3: IF (Structure profile detects path change) THEN (Set the detection mode to high-frequency discovery mode for the next few cycles and use ADVANCED_PERFORMANCE signaling to quickly familiarize yourself with the new path).
[0103] Ultimately, the adjustment model outputs quantified adjustment parameters, such as adjusting the probe interval of target device A from 60 seconds to 30 seconds (probe mode adjustment parameter) and setting its signaling type to ADVANCED_PERFORMANCE (probe signaling type adjustment parameter). These parameters are designed to make the probe behavior in the next cycle more adaptable to the actual conditions of the current network.
[0104] S53: The detection mode adjustment parameters and the detection signaling type adjustment parameters are used as constraints to be input into the dynamic detection strategy generation process of the next cycle, so as to generate an optimized dynamic detection strategy that adapts to the current network state.
[0105] In this embodiment, the detection mode adjustment parameters and detection signaling type adjustment parameters calculated in step S52 are used as key constraints and input into the "S1: Dynamic Detection Strategy Generation" process in the next round (i.e. the next detection cycle).
[0106] Specifically, when the system generates new dynamic detection strategies for each target device, the decision-making logic of the strategy engine will be constrained by these adjustment parameters. For example, for target device A, when calculating its detection requirements, the strategy engine will prioritize the suggestions of the adjustment parameters: directly set its detection mode to the adjusted high-frequency mode (such as a 30-second interval) and lock its detection signaling type to ADVANCED_PERFORMANCE.
[0107] This means that the new strategy is no longer calculated from scratch, but rather adjusted based on the previous strategy, combined with the latest network profile and explicit optimization instructions. The resulting optimized dynamic detection strategy will be more accurate and efficient. Subsequently, the new strategy will be distributed to the target detection nodes through step S13, initiating a new cycle of detection, data collection, analysis, and optimization.
[0108] In this way, the entire system achieves a complete closed loop from "monitoring" to "analysis" to "optimization" and finally to "execution," enabling the detection behavior to continuously adapt to the dynamically changing network environment, thereby continuously improving the accuracy and efficiency of detection.
[0109] Example 2: This invention also provides a network device edge detection system based on a heartbeat detection mechanism, used to execute the aforementioned network device edge detection method based on a heartbeat detection mechanism. Referring to Figure 2, the detection system includes: a policy generation and allocation module 100, used to generate dynamic detection policies based on network behavior profiles and network context, and allocate them to target detection nodes in the edge detection device cluster. The dynamic detection policy includes target device identifier, detection signaling type, and detection mode.
[0110] The probe message construction module 200 is used by the target probe node to construct a probe message corresponding to the probe mode according to the dynamic probe strategy, and to encapsulate signaling data including the probe signaling type in the probe message and send it to the target device.
[0111] The telemetry data injection module 300 is used to carry out the forwarding of the probe message, identify the probe signaling type, and add the in-band telemetry data of the network device to the response message when forwarding the response message returned by the target device.
[0112] The status information fusion module 400 is used to receive a response message including the in-band telemetry data, associate the arrival information of the response message with the in-band telemetry data in terms of time and task, and generate fused network status information.
[0113] The profile update and adjustment module 500 is used to update the network behavior profile based on the network status information, and adjust the detection mode and detection signaling type in the dynamic detection strategy for the next cycle based on the updated network behavior profile.
[0114] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.
[0115] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0116] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
Claims
1. A network device edge detection method based on a heartbeat detection mechanism, characterized in that, The detection method includes the following steps: Based on network behavior profiles and network context, a dynamic detection strategy is generated and allocated to target detection nodes in an edge detection device cluster. The dynamic detection strategy includes a target device identifier, detection signaling type, and detection mode. The target detection node constructs a detection packet corresponding to the detection mode according to the dynamic detection strategy, encapsulates signaling data including the detection signaling type in the detection packet, and sends it to the target device. The network device carrying the forwarding of the detection packet identifies the detection signaling type and, when forwarding a response packet returned by the target device, adds the network device's in-band telemetry data to the response packet for reporting. The network device receives a response packet including the in-band telemetry data, associates the arrival information of the response packet with the in-band telemetry data in terms of time and task, and generates fused network status information. The network behavior profile is updated based on the network status information, and the detection mode and detection signaling type in the dynamic detection strategy for the next cycle are adjusted based on the updated network behavior profile.
2. The network device edge detection method based on a heartbeat detection mechanism according to claim 1, characterized in that, The step of generating dynamic detection strategies based on network behavior profiles and network context and allocating them to target detection nodes in the edge detection device cluster includes: acquiring the network behavior profile including historical detection data and network event logs, and the network context including network topology and device attribute information; calculating the detection requirements for different target devices based on the network behavior profile and the network context, and generating the dynamic detection strategy including target device identifier, detection signaling type, and detection mode; evaluating the cost of each detection node executing the dynamic detection strategy based on the load status and network location of each detection node in the edge detection device cluster, and allocating the dynamic detection strategy to the target detection node with the lowest evaluated cost.
3. The network device edge detection method based on a heartbeat detection mechanism according to claim 2, characterized in that, The target detection node constructs a detection packet corresponding to the detection mode according to the dynamic detection strategy, and encapsulates signaling data including the detection signaling type in the detection packet, and sends it to the target device. This includes: selecting a corresponding detection protocol from a preset protocol library based on the detection mode; generating an initial detection packet carrying a task identifier according to the selected detection protocol; writing the signaling instruction corresponding to the detection signaling type into the initial detection packet to generate an in-band telemetry enabled detection packet; and tagging the detection packet with a corresponding target VLAN tag according to the target device identifier, which is then sent by the target detection node.
4. The network device edge detection method based on a heartbeat detection mechanism according to claim 3, characterized in that, The network device carrying the probe message forwarding identifies the probe signaling type and, when forwarding the response message returned by the target device, adds the network device's in-band telemetry data to the response message for reporting. This includes: determining a collection template based on the probe signaling type, wherein the collection template defines data fields to be collected, the data fields including at least the port identifier, timestamp, and queue status involved by the network device when forwarding the response message; when forwarding the response message, collecting the real-time port identifier, timestamp, and queue status from the network device according to the definition of the collection template, and combining them to generate the in-band telemetry data; binding the task identifier in the response message with the in-band telemetry data and writing it into the selected in-band telemetry extension field in the response message, and reporting it through the network device.
5. The network device edge detection method based on a heartbeat detection mechanism according to claim 4, characterized in that, The step of receiving a response message including the in-band telemetry data, and associating the arrival information of the response message with the in-band telemetry data in terms of time and task to generate fused network state information includes: matching the corresponding initial probe task record in a preset policy task library based on the task identifier to obtain the target device identifier, the probe signaling type, and the probe mode; normalizing each local timestamp in the sequence of in-band telemetry data formation based on the task generation timestamp in the initial probe task record, and aligning the local timestamp with the final timestamp of the response message arrival to reconstruct the timing state vector of the response message on the transmission path; and calculating and generating the fused network state information including connectivity state features, hop-by-hop performance features, and topology features based on the timing state vector, the probe mode, the port identifier and queue status in the in-band telemetry data.
6. The network device edge detection method based on a heartbeat detection mechanism according to claim 5, characterized in that, The step of updating the network behavior profile based on the network state information and adjusting the detection mode and detection signaling type in the dynamic detection strategy for the next cycle based on the updated network behavior profile includes: updating the state profile of the corresponding target device in the network behavior profile based on the connectivity state characteristics, and updating the performance baseline profile and structure profile of the corresponding network path in the network behavior profile based on the hop-by-hop performance characteristics and the topology characteristics; calculating the detection mode adjustment parameters and detection signaling type adjustment parameters for the target device in the next cycle based on the updated state profile, the performance baseline profile, and the structure profile through a preset detection strategy adjustment model; and inputting the detection mode adjustment parameters and the detection signaling type adjustment parameters as constraints into the dynamic detection strategy generation process for the next cycle to generate an optimized dynamic detection strategy adapted to the current network state.
7. A network device edge detection system based on a heartbeat detection mechanism, used to execute a network device edge detection method based on a heartbeat detection mechanism as described in any one of claims 1 to 6, characterized in that, The detection system includes: a policy generation and allocation module, used to generate dynamic detection policies based on network behavior profiles and network context, and allocate them to target detection nodes in the edge detection device cluster, wherein the dynamic detection policy includes a target device identifier, a detection signaling type, and a detection mode; a detection packet construction module, used by the target detection node to construct a detection packet corresponding to the detection mode according to the dynamic detection policy, and encapsulate signaling data including the detection signaling type in the detection packet, and send it to the target device; and a telemetry data injection module, used to carry the forwarding of the detection packet and identify the network device. The detection signaling type is specified, and when forwarding the response message returned by the target device, the in-band telemetry data of the network device is added to the response message for reporting; the status information fusion module is used to receive the response message including the in-band telemetry data, associate the arrival information of the response message with the in-band telemetry data in terms of time and task, and generate fused network status information; the profile update and adjustment module is used to update the network behavior profile according to the network status information, and adjust the detection mode and detection signaling type in the dynamic detection strategy for the next cycle based on the updated network behavior profile.
Citation Information
Patent Citations
Network telemetering method and device for distinguishing scenes
CN118740647A
Routing state sensing method and device cooperatively monitored by distributed network probes
CN120567748A
Network monitoring method, device, equipment, storage medium and computer program product
CN120768811A
Network service fault detection method, device and system, electronic equipment and storage medium
CN121619222A