Network monitoring based edge computing node intelligent online monitoring method and system

CN122513306BActive Publication Date: 2026-09-18HANGZHOU WANGDING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610992220.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-18
Estimated Expiration
2046-07-06

AI Technical Summary

Technical Problem

然而,边缘节点通常资源受限、环境复杂,且承担着实时性要求较高的业务,其出站流量异常(如突发丢包、重传率升高、发送队列堆积等)会直接影响终端用户体验与服务可靠性

Benefits of technology

[0025] In several embodiments of this specification, the provided intelligent online monitoring method and system for edge computing nodes avoids the high resource consumption caused by continuous collection of packet-level data by deploying a lightweight monitoring agent within the edge computing node and only periodically recording aggregated metrics and internal status time-series snapshots. When an abnormal traffic pattern is detected, a preliminary diagnosis is first performed using existing snapshots. Introspection mode is triggered and packet-level data recording is temporarily enabled only when the fault type cannot be determined. Introspection mode automatically exits after a preset duration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122513306B_ABST
    Figure CN122513306B_ABST
Patent Text Reader

Abstract

The application relates to the field of information technology, in particular to an edge computing node intelligent online monitoring method and system based on network monitoring, which comprises the following steps: deploying a lightweight monitoring agent in an edge computing node, counting aggregate indexes of outbound traffic, and recording internal state time sequence snapshots at a preset period; online monitoring and analyzing the outbound traffic; when detecting an abnormal traffic mode, determining a time window of abnormality occurrence and an associated edge computing node; sending a first request to the edge computing node to obtain internal state time sequence snapshot data in the time window; analyzing the obtained time sequence snapshot data and combining outbound traffic characteristics to preliminarily diagnose a fault; if the preliminary diagnosis cannot determine that the fault belongs to a preset distinguishable type, sending a second request to the edge computing node to record packet-level data of each outbound packet in a time window specified by the second request; and receiving and obtaining a fault type according to the packet-level data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, specifically to an intelligent online monitoring method and system for edge computing nodes based on network monitoring. Background Technology

[0002] Edge computing deploys computing and storage resources at the network edge to reduce service latency and bandwidth consumption. However, edge nodes are typically resource-constrained, operate in complex environments, and handle services with high real-time requirements. Abnormal outbound traffic (such as sudden packet loss, increased retransmission rates, and congestion in the sending queue) directly impacts end-user experience and service reliability. Traditional network monitoring methods are mostly based on centralized data collection or passive log analysis. On the one hand, they struggle to continuously record fine-grained packet-level data on resource-limited edge nodes; on the other hand, the large volume of data often interferes with normal node operation. For example, conventional simple network management protocols or flow collection methods (such as NetFlow and sFlow) only provide aggregated traffic statistics, lacking the ability to correlate with process-level and container-level internal states, making it difficult to pinpoint the root cause of faults. Some solutions attempt to continuously enable deep packet inspection or kernel probes on edge nodes. While this can obtain detailed data, it significantly increases the CPU and memory overhead of edge nodes, potentially causing secondary performance issues, especially during peak traffic periods. Therefore, how to achieve online monitoring and fault diagnosis of outbound traffic anomalies without affecting the normal operation of edge nodes is a technical problem that urgently needs to be solved in the current field of edge computing operation and maintenance. Summary of the Invention

[0003] This specification describes a method and system for intelligent online monitoring of edge computing nodes based on network monitoring through several embodiments.

[0004] Firstly, embodiments of this specification provide an intelligent online monitoring method for edge computing nodes based on network monitoring, including the following steps:

[0005] A lightweight monitoring agent is deployed within the edge computing node. The monitoring agent collects aggregated metrics of outbound traffic by process or container dimension and records time-series snapshots of the internal state of the edge computing node at a preset period.

[0006] Online monitoring and analysis of outbound traffic from edge computing nodes; when abnormal traffic patterns are detected, the time window of the anomaly and the associated edge computing nodes are determined.

[0007] Send a first request to the edge computing node to obtain internal state time-series snapshot data within the time window;

[0008] The acquired time-series snapshot data is analyzed and combined with outbound traffic characteristics to perform preliminary fault diagnosis and determine whether the fault type belongs to a preset distinguishable type.

[0009] If the initial diagnosis cannot determine that the fault belongs to a preset distinguishable type, a second request is sent to the edge computing node to trigger the edge computing node to start introspection mode and record packet-level data of each outbound packet within the time window specified in the second request.

[0010] The system receives and obtains the fault type based on the packet-level data, and then automatically exits after the introspection mode has run for a preset duration.

[0011] Secondly, embodiments of this specification provide an intelligent online monitoring system for edge computing nodes based on network monitoring, including:

[0012] The acquisition module deploys a lightweight monitoring agent within the edge computing node. The monitoring agent collects aggregated metrics of outbound traffic by process or container dimension and records time-series snapshots of the internal state of the edge computing node at a preset period.

[0013] The monitoring module performs online monitoring and analysis of outbound traffic from edge computing nodes. When an abnormal traffic pattern is detected, it determines the time window of the anomaly and the associated edge computing nodes.

[0014] The first acquisition module sends a first request to the edge computing node to acquire internal state time-series snapshot data within the time window;

[0015] The first analysis module parses the acquired time-series snapshot data, combines it with outbound traffic characteristics to perform preliminary fault diagnosis, and determines whether the fault type belongs to a preset distinguishable type.

[0016] If the preliminary diagnosis fails to determine that the fault belongs to a preset distinguishable type, the second acquisition module sends a second request to the edge computing node, triggering the edge computing node to start introspection mode and record packet-level data of each outbound packet within the time window specified in the second request.

[0017] The second analysis module receives and obtains the fault type based on the packet-level data, and automatically exits after the introspection mode has run for a preset duration.

[0018] Thirdly, embodiments of this specification provide an electronic device, including a processor and a memory;

[0019] The processor is connected to the memory;

[0020] The memory is used to store executable program code;

[0021] The processor runs a program corresponding to the executable program code stored in the memory to perform the method described in any of the above aspects.

[0022] Fourthly, embodiments of this specification provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any of the above aspects.

[0023] Fifthly, embodiments of this specification provide a computer program product, including a computer program that, when executed by a processor, implements the methods described in any of the above aspects.

[0024] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:

[0025] In several embodiments of this specification, the provided intelligent online monitoring method and system for edge computing nodes avoids the high resource consumption caused by continuous collection of packet-level data by deploying a lightweight monitoring agent within the edge computing node and only periodically recording aggregated metrics and internal status time-series snapshots. When an abnormal traffic pattern is detected, a preliminary diagnosis is first performed using existing snapshots. Introspection mode is triggered and packet-level data recording is temporarily enabled only when the fault type cannot be determined. Introspection mode automatically exits after a preset duration.

[0026] The adoption of a phased, on-demand adaptive data collection strategy significantly reduces the CPU and memory overhead of edge nodes, avoiding interference with normal business operations, and is particularly suitable for resource-constrained edge environments. Joint analysis of outbound traffic characteristics and the internal status of edge nodes enables the differentiation of multiple preset distinguishable fault types.

[0027] Introspection mode uses dynamically loaded probes, operating only within a specified time window and automatically unloading them after completion, thus avoiding long-term occupation of kernel resources. The second request can carry filtering conditions, recording only the packets of interest, which helps reduce data volume.

[0028] Other features and advantages of various embodiments of this specification will be further revealed in the following detailed description and accompanying drawings. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram of intelligent online monitoring of edge computing nodes provided in this manual.

[0031] Figure 2This is a schematic diagram of the intelligent online monitoring method for edge computing nodes provided in this manual.

[0032] Figure 3 This is a flowchart illustrating the method for recording packet-level data for each outbound packet provided in this manual.

[0033] Figure 4 This is a flowchart illustrating the method for obtaining fault types based on package-level data, as provided in this manual.

[0034] Figure 5 This is a schematic diagram of the intelligent online monitoring system for edge computing nodes provided in this manual.

[0035] Figure 6 This is a schematic diagram of the electronic device provided in this manual. Detailed Implementation

[0036] The technical solutions of the embodiments of this specification will be explained and described below with reference to the accompanying drawings. However, the following embodiments are only preferred embodiments of this specification and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments in the implementation methods without creative effort are all within the protection scope of this specification.

[0037] The terms "first," "second," "third," etc., in the description, claims, and accompanying drawings are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0038] In the following description, terms such as “inner,” “outer,” “upper,” “lower,” “left,” and “right” are used only to facilitate the description of the embodiments and to simplify the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this specification.

[0039] All data involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0040] Before introducing the technical solutions described in this manual, the application scenarios and related technologies of the technical solutions will be introduced.

[0041] This embodiment relates to the fields of edge computing and network operation and maintenance technology, and is particularly suitable for detecting abnormal outbound traffic and locating the root cause of faults at edge nodes with limited resources and high real-time requirements.

[0042] Currently, edge computing nodes are widely deployed at the network edge, undertaking critical tasks such as data aggregation, real-time processing, and low-latency forwarding. However, edge nodes generally suffer from limited computing resources (CPU, memory) and storage resources, as well as complex and variable operating environments. Abnormal outbound traffic (such as sudden packet loss, spiked TCP retransmission rates, and send queue congestion) directly impacts end-user experience and the stability of upper-layer services. To ensure service quality, the operation and maintenance system needs to continuously monitor the outbound traffic of edge nodes and quickly pinpoint the cause of failures.

[0043] In existing technologies, monitoring of network traffic anomalies mainly falls into two categories. One is coarse-grained monitoring based on aggregated statistics, such as collecting interface traffic counts via SNMP or collecting flow-level statistics via NetFlow / sFlow. This type of method has low resource overhead, but it only provides aggregated traffic metrics and cannot correlate with process-level or container-level internal states (such as CPU load distribution, TCP send queue changes, scheduling delays, etc.). It also struggles to distinguish whether the fault is caused by external network interference, slow application-layer writes, kernel resource contention, or abnormal processes within the node itself, resulting in severely insufficient root cause localization capabilities. The other type of method is fine-grained deep packet inspection or kernel probe monitoring, which involves continuously recording packet-level data or collecting full kernel events on the node. While this method can obtain detailed information for each packet, its CPU and memory overhead is extremely high. During peak traffic periods or on resource-constrained edge nodes, it can easily cause secondary performance problems, or even render the node itself unusable.

[0044] Furthermore, most existing monitoring systems employ fixed data collection strategies, failing to adaptively adjust data granularity based on anomaly types. This leads to either insufficient data for diagnosis or excessive data collection resulting in wasted resources. Therefore, achieving low-overhead, accurate, and traceable online monitoring and fault diagnosis of outbound traffic anomalies while ensuring the normal operation of edge nodes has become a pressing technical challenge in this field.

[0045] This embodiment provides an intelligent online monitoring method and system for edge computing nodes 10 based on network monitoring. Please refer to the appendix. Figure 1 Through a phased, on-demand adaptive monitoring architecture, aggregated metrics and internal state time-series snapshots are collected with minimal overhead under normal circumstances. Upon detecting an anomaly, existing snapshots are used first for rapid preliminary diagnosis. Only when the fault type cannot be determined initially will lightweight package-level data recording (introspection mode) be temporarily triggered within a specified time window, and the process will automatically exit upon completion.

[0046] Furthermore, for complex faults that cannot be distinguished by initial diagnosis, a deep introspection mechanism is supported to further collect process-level fine-grained indicators to pinpoint the root cause. This avoids the resource consumption associated with long-term packet-level monitoring while ensuring that detailed data sufficient to locate faults is available when anomalies occur, achieving an optimal balance between monitoring accuracy and operational overhead, and significantly improving the automation level and self-healing capabilities of edge computing clusters.

[0047] This embodiment provides an intelligent online monitoring method for edge computing node 10 based on network monitoring. Please refer to the appendix. Figure 2 The steps include:

[0048] Step S1: Deploy a lightweight monitoring agent within the edge computing node 10. The monitoring agent collects aggregated metrics of outbound traffic by process or container dimension and records the internal state time-series snapshot 11 of the edge computing node 10 at a preset period.

[0049] The monitoring agent performs the following steps:

[0050] Aggregated metrics for outbound traffic statistics by process or container dimension, the aggregated metrics include one or more of the following: number of bytes sent, number of transmission control protocol retransmissions, and average round-trip time;

[0051] The internal state timing snapshot 11 of the edge computing node 10 is recorded at a preset period. The internal state timing snapshot 11 includes the CPU load, run queue length, memory usage, transmission control protocol send queue size, and network card packet loss count. The snapshot is stored cyclically in a local circular buffer and retained for a preset duration.

[0052] Deploy an independent, lightweight monitoring agent in each edge computing node 10 that needs to be monitored. This agent runs as a daemon, consuming minimal CPU and memory resources and not affecting the normal business operations carried on the edge node. The core work of the monitoring agent consists of two parallel tasks: first, to collect aggregated metrics of outbound traffic by process or container; and second, to record time-series snapshots of the node's internal state at preset intervals.

[0053] Aggregated metrics include one or more of the following: number of bytes sent, TCP retransmission count, and average round-trip time (RTT). The monitoring agent distinguishes outbound traffic from different processes or containers by reading network statistics interfaces provided by the operating system kernel (such as ` / proc / net / tcp`, `netstat`, or eBPF hooks in Linux). For example, in a scenario where a video streaming container and a log upload process are running simultaneously on the same edge node, the monitoring agent will separately count the number of bytes sent, TCP retransmission count, and average RTT of the video streaming container in the most recent sampling period, while independently counting the corresponding metrics for the log upload process. If the TCP retransmission rate of the video streaming container suddenly increases while the log upload process remains normal, it can be preliminarily determined that the problem does not originate from the node's global network, but may be related to a specific connection of the video streaming container or a peer service.

[0054] Internal state time-series snapshot 11 includes CPU load, run queue length, memory usage, TCP send queue size, and network interface card (NIC) packet loss count. The monitoring agent retrieves this internal state by calling the operating system interface at the same period as the aggregated metrics (or at a separately configured period). For example, it obtains CPU load (1-minute, 5-minute, and 15-minute average load) by reading ` / proc / loadavg`, the number of processes waiting to execute in the run queue by reading ` / proc / stat`, memory usage by ` / proc / meminfo`, the send queue size for each TCP connection by `ss -ti` or `netstat`, and the NIC packet loss count by ` / proc / net / dev`. The monitoring agent organizes this internal state data into time-series snapshots and stores them cyclically in a local circular buffer. The size of the circular buffer is configurable, for example, retaining the most recent 60 snapshots.

[0055] The sampling period can be set to 10 seconds or 30 seconds, depending on the node's business sensitivity and resource constraints. The monitoring agent aggregates metrics using timestamps as indexes and lightweightly records them in local memory or a circular buffer, without writing them to persistent storage to avoid disk I / O interference.

[0056] Step S2: Perform online monitoring and analysis of the outbound traffic of edge computing node 10. When an abnormal traffic pattern is detected, determine the time window of the abnormality and the associated edge computing node 10.

[0057] Methods for online monitoring and analysis of outbound traffic of edge computing node 10 include:

[0058] The detection is based on the bandwidth utilization, transmission control protocol retransmission rate, packet loss rate, and deviation of sudden changes in traffic from historical baselines in outbound traffic.

[0059] When the deviation exceeds a preset threshold, it is determined to be an abnormal traffic pattern, and the time window of the abnormality and the associated edge computing node 10 identifier are recorded.

[0060] An independent online monitoring and analysis module 20 (which can be deployed on the central control console 30 or a monitoring server that collaborates with the edge computing nodes 10) is responsible for real-time or near real-time analysis of the outbound traffic of each edge computing node 10. This module continuously evaluates the outbound traffic patterns of each edge computing node 10 by receiving aggregated metrics (number of bytes sent, number of TCP retransmissions, average round-trip latency, etc.) reported by the lightweight monitoring agent on the edge computing node 10, or by directly connecting to the traffic mirroring of the edge gateway.

[0061] A dynamic historical baseline model is maintained for each edge computing node 10. The baseline can be calculated using an exponentially weighted moving average (EWMA) or a simple periodic sliding window (such as the median of the same period in the past 24 hours). Detection metrics include at least bandwidth utilization, transmission control protocol retransmission rate, packet loss rate, and burst traffic changes. Preset thresholds can be flexibly configured according to business scenarios. For example, the TCP retransmission rate deviation threshold can be set to 3% (i.e., exceeding the baseline by 3 percentage points), the bandwidth utilization deviation threshold can be set to twice the baseline value, and the burst detection threshold can be set to three times the standard deviation.

[0062] The time window is determined by starting from the moment the deviation first exceeds the threshold and ending when the deviation falls back below the threshold, forming a continuous abnormal interval. The time window needs to record the start and end timestamps (accurate to the second or millisecond) and associate them with the unique identifier of the corresponding edge computing node (such as node ID or IP address).

[0063] Step S3: Send a first request to the edge computing node 10 to obtain the internal state time sequence snapshot 11 data within the time window.

[0064] After detecting an abnormal traffic pattern and determining the time window of the anomaly (denoted as T_start to T_end) and the associated edge computing node 10 identifier, the online monitoring and analysis module 20 proactively initiates a first request to the edge computing node 10. Internal state time-series snapshot 11 data corresponding to the abnormal time window is extracted from the circular buffer stored locally on the edge computing node 10 for subsequent preliminary root cause correlation analysis.

[0065] The first request includes at least the following key parameters: the identifier of the target edge computing node 10 (such as node ID, IP address or hostname), the start time T_start and end time T_end of the request time window, optional data format requirements (such as JSON or Protobuf) and compression flags.

[0066] A lightweight monitoring agent running on edge computing node 10 listens for the first request from online monitoring and analysis module 20 (typically via HTTP / gRPC or message queue communication). Upon receiving the first request, the monitoring agent first queries the local circular buffer based on the time window parameter. Since the circular buffer stores internal state time-series snapshots 11 recorded at preset intervals (e.g., 30 seconds), each snapshot has a precise timestamp, the monitoring agent can quickly locate and extract all snapshot records whose timestamps fall within the [T_start, T_end] interval.

[0067] If the circular buffer completely covers the requested time window, all relevant snapshots are returned. If some historical data has been overwritten due to the capacity limit of the circular buffer (e.g., the time when the exception occurred was far from the time when the request was initiated, exceeding the preset retention period), the monitoring agent returns the largest available subset of snapshots that still exist in the buffer and whose timestamps fall within the window, and includes a data integrity indicator (such as "partial data" or "complete data") in the response.

[0068] The online monitoring and analysis module 20 completed anomaly detection at 9:08:30 AM, determining the anomaly time window to be [9:05:30, 9:08:20], and associated it with edge computing node 10 "EdgeNode-12". The module then constructed a first request, including the node identifier "EdgeNode-12", the start time 9:05:30, and the end time 9:08:20, and sent it to the monitoring agent on "EdgeNode-12" via the gRPC interface.

[0069] Upon receiving the request, the monitoring agent queries its local circular buffer. This buffer is configured to store the 120 most recent snapshots, with each snapshot spaced 30 seconds apart, meaning it retains data from the last 60 minutes (9:05:30 is only 3 minutes past the current 9:08:30, so all related snapshots are untouched). The monitoring agent extracts 6 internal status time-series snapshots (11) with timestamps of 9:05:30, 9:06:00, 9:06:30, 9:07:00, 9:07:30, and 9:08:00 (the 9:08:30 snapshot has not yet been recorded). Each snapshot contains the following data example:

[0070] 9:07:00 Snapshot: CPU load (1 minute) is 2.5, run queue length is 4, memory usage is 1.8GB / 4GB, TCP send queue size is 0.5KB (connection count 12), and network card packet loss count is 85 (an increase of 12 from the previous snapshot).

[0071] 9:07:30 Snapshot: CPU load increased to 3.2, run queue length 6, memory usage 1.9GB, TCP send queue size 8KB (significantly increased), and network card packet loss count accumulated to 102 (an increase of 17).

[0072] 9:08:00 Snapshot: CPU load 3.0, run queue length 5, memory usage 1.9GB, TCP send queue size 15KB, network card packet loss count cumulative 124 (increased by 22).

[0073] These 6 snapshot data are encapsulated into a compressed response message and sent back to the online monitoring and analysis module 20.

[0074] Step S4: Analyze the acquired time-series snapshot data, combine it with outbound traffic characteristics to perform preliminary fault diagnosis, and determine whether the fault type belongs to the preset distinguishable type.

[0075] The method for analyzing the acquired time-series snapshot data and combining it with outbound traffic characteristics to perform preliminary fault diagnosis and determine the fault type includes:

[0076] When the transmission control protocol send queue size is normal but the outbound retransmission rate increases, the fault type is inferred to be external network interference.

[0077] If parsing the time-series snapshot data reveals that the TCP sending queue size (i.e., the amount of data to be sent for each TCP connection or the total amount of data to be sent) remains at a low level (e.g., less than 1.5 times the normal baseline), and the CPU load, memory usage, and running queue length within the node are all normal, but the outbound traffic characteristics show a significantly increased TCP retransmission rate (e.g., exceeding 3%), it indicates that the data packets have been successfully delivered from the node kernel to the network card and sent, but have not been acknowledged by the other end. The reason is likely due to packet loss, congestion, or interference from intermediate devices in the external network, rather than a problem with the node itself.

[0078] When the transmission control protocol send queue is overflowing and the system memory is sufficient, the fault type is inferred to be slow application layer writes.

[0079] Analysis of the time-series snapshot data revealed that the TCP send queue size was continuously increasing or remained at a high level for an extended period (e.g., more than three times the normal value), while the node's system memory usage did not reach the warning threshold (e.g., remaining memory greater than 20%), and the CPU load was not necessarily high. This indicates that the kernel is ready to send data, but the application layer's speed of writing data to the socket cannot keep up with the consumption of the send queue (or the send queue is quickly filled but sending is slow). This is usually because the application process itself generates data slowly, or there are blocking issues in the application layer logic (such as slow disk I / O, lock waits, etc.), rather than insufficient network or kernel resources.

[0080] When the length of the transmission queue momentarily exceeds the first threshold and the packet loss count of the network card increases at a millisecond time point, the fault type is inferred to be micro-burst packet loss.

[0081] The sampling period of a conventional time-series snapshot (e.g., 30 seconds) is insufficient to capture millisecond-level instantaneous jitter. Therefore, this embodiment allows the monitoring agent to maintain an additional high-precision ring counter or utilize eBPF (Extended Berkeley Packet Filter) to count the instantaneous queue peaks that occur within the sampling interval while recording internal state time-series snapshot 11. When parsing the snapshot data, if it is found that within any millisecond-level observation window, the TCP send queue length instantaneously exceeds the first threshold (e.g., exceeding twice the network card bandwidth-delay product or an absolute value greater than 100KB), and the network card packet loss count increases during the same time period (indicating that instantaneous queue overflow caused packet loss), then the fault type is inferred to be "micro-burst packet loss". Micro-burst packet loss is common in scenarios such as video keyframes and real-time log batch reporting, and its characteristic is that the average bandwidth is normal but the instantaneous peak is extremely high.

[0082] When the difference in the run queue length of multiple CPU cores exceeds the second threshold, the fault type is inferred to be a scheduling affinity configuration error.

[0083] In general-purpose operating systems like Linux, if the run queue lengths (i.e., the number of tasks running or waiting to run on each core) of multiple CPU cores differ significantly—for example, if the run queue length of one core is 8 while the average length of other cores is only 1, and the difference exceeds a second threshold (e.g., set to 5)—it indicates that the CPU affinity of processes or threads is bound to a few cores, or that interrupt load is unevenly distributed. This can cause sent tasks to concentrate on a few cores, creating a bottleneck and thus affecting outbound traffic performance.

[0084] If the preliminary diagnosis result belongs to any of the above-mentioned preset distinguishable types, output the fault type and end; otherwise, output that the fault cannot be determined to belong to the preset distinguishable type.

[0085] If none of the above rules are met—that is, the TCP send queue size and retransmission rate are normal, the queue is backlogged but memory is insufficient or other complex situations exist, no micro-burst packet loss is detected, or the CPU is running a load-balanced queue—then the cause of the anomaly does not belong to these common fault types that are easily distinguishable through aggregated snapshots. In this case, the analysis module outputs "Unable to determine if the fault belongs to the preset distinguishable type."

[0086] Step S5: If the preliminary diagnosis cannot determine that the fault belongs to a preset distinguishable type, a second request is sent to the edge computing node 10 to trigger the edge computing node 10 to start introspection mode and record the packet-level data of each outbound packet within the time window specified by the second request.

[0087] When the initial fault diagnosis outputs "Unable to determine if the fault belongs to the preset distinguishable type," it indicates that the aggregated indicators and internal state time-series snapshot 11 alone are insufficient to locate the root cause, and more granular packet-level data is needed. Therefore, in this embodiment, the online monitoring and analysis module 20 sends a second request to the edge computing node 10, requesting the node to temporarily activate introspection mode and record detailed data for each outbound packet within a specified time window. Introspection mode is a deep data acquisition mechanism that starts on demand and closes periodically, helping to minimize interference with the normal operation of the edge computing node 10.

[0088] Please see the appendix Figure 3 The method for sending a second request to the edge computing node 10 to trigger the edge computing node 10 to enter introspection mode and record packet-level data of each outbound packet within the time window specified in the second request includes:

[0089] Step S51: The second request includes the duration of introspection and the filtering conditions to be tracked.

[0090] The second request constructed by the online monitoring and analysis module 20 carries at least two key parameters: first, the introspection duration (e.g., 10 seconds, 30 seconds, or 60 seconds), used to control the time window length for packet-level data recording; and second, the filtering conditions to be tracked, used to limit which outbound packets are recorded, avoiding the massive data overhead of recording the entire traffic. The filtering conditions can be combined based on a 5-tuple (source IP, destination IP, source port, destination port, protocol), process identifier (PID), container ID, or specific error codes. For example, only outbound packets from a suspicious process with PID=1234 can be recorded, or only packets whose target Internet Protocol address (IP) belongs to a specific subnet can be recorded.

[0091] The second request may also include a return address (if an active push method is used) or a request for edge computing node 10 to wait for retrieval after recording is complete. The introspection duration is typically set to several seconds to one minute, sufficient to capture a complete cycle of abnormal traffic patterns while avoiding prolonged high-overhead operation. The return address can be the address of the online monitoring and analysis module 20 or the address of the central console 30.

[0092] Step S52: After receiving the second request, the edge computing node 10 dynamically loads the extended packet filter probe and attaches it to the kernel function dev_queue_xmit and the Transmission Control Protocol retransmission event.

[0093] The monitoring agent resident on edge computing node 10 is responsible for receiving the second request. Upon receiving the request, the monitoring agent first parses the introspection duration and filtering conditions, and then dynamically loads a set of pre-compiled extended Berkeley package filter probes or kernel modules. These probes are attached to key kernel functions, including:

[0094] `dev_queue_xmit`: This function is the entry point for the network device's queue transmission. Each outbound packet passes through here before entering the network interface card's transmission queue. Attaching a probe allows you to capture the packet's metadata as it is being transmitted.

[0095] Transmission Control Protocol (TCP) retransmission events: These are specifically recorded by attaching to tcp_retransmit_skb or a similar TCP retransmission handling function.

[0096] The dynamic loading method means that these probes do not exist under normal circumstances, resulting in zero overhead for edge computing nodes 10; they are only loaded temporarily when introspection is required, and the probes are designed to be lightweight, extracting only necessary fields and not performing complex packet content parsing to reduce CPU consumption.

[0097] Step S53: Within the specified time window, record the sending timestamp, target Internet Protocol address, process identifier, and error code returned by the kernel sending path for each outbound packet.

[0098] During the probe's active window (from the second request confirmation to the end of the introspection duration), whenever `dev_queue_xmit` is called or a TCP retransmission event occurs, the probe intercepts key information about the packet and records it in a circular buffer within the kernel. Each record contains the following fields: Sending timestamp: accurate to microseconds or nanoseconds, used for timing analysis. Destination Internet Protocol address (IP): 32-bit for IPv4 and 128-bit for IPv6, helping to distinguish the destination of traffic. Process identifier (PID): obtained through the `current` macro in the kernel or socket owner information, indicating which process sent the packet. Error code returned by the kernel sending path: for example, error values ​​returned after calling `dev_queue_xmit`, common ones include -ENOBUFS (insufficient NIC send buffer), -ENOMEM (memory allocation failure), -EAGAIN (temporarily unavailable), etc. This directly reflects the underlying problems encountered by the kernel when sending packets. For TCP retransmission events, the retransmission sequence number and retransmission reason are additionally recorded. The filtering conditions take effect here: only packets that meet the conditions will be recorded, such as PID matching, target IP matching, etc.

[0099] Step S54: Compress the recorded packet-level data and send it back.

[0100] After the introspection period ends (or the buffer size limit is reached early), the monitoring agent automatically unloads the probe and stops packet-level data collection. Subsequently, the monitoring agent reads all recorded packet-level data from the kernel ring buffer, assembles it into a structured format (such as a JSON array or Protocol Buffers), compresses it using a lightweight compression algorithm (such as gzip or LZ4), and sends the compressed data back to the online monitoring and analysis module 20 that sent the second request over the network. After the data return is complete, the edge computing node 10 cleans up the temporary buffer, releases resources, and returns to the normal probe-free mode. If the data return fails (e.g., due to network interruption), the data is temporarily stored locally for a certain period, waiting for a retry.

[0101] Step S6: Receive and obtain the fault type based on the packet-level data. The introspection mode will automatically exit after running for a preset duration.

[0102] Please see the appendix Figure 4 The method for receiving and obtaining the fault type based on the packet-level data includes:

[0103] Step S61: The fault type is determined to be insufficient network card transmit buffer or exhaustion of driver resources.

[0104] Each packet record contains an error code returned by the kernel function `dev_queue_xmit` or the TCP retransmission path. `-ENOBUFS` indicates that the network card's transmit buffer is full and cannot accept new packets; `-ENOMEM` indicates that the kernel failed to allocate the necessary skb (socket buffer) memory for transmission, usually related to driver resources or system memory fragmentation. If, within a time window, the frequency of these two error codes is significantly higher than normal (e.g., exceeding 1% of the total number of packets or an absolute number greater than 10), the fault type is determined to be "insufficient network card transmit buffer" or "driver resource exhaustion".

[0105] For example, of the 1200 packet records returned, 1100 records returned the error code -ENOBUFS, accounting for over 90%. Simultaneously, checking the arrival timestamps of the acknowledgment characters from the other end revealed that a large number of packets were not acknowledged, but this was not a network latency issue. Based on this, the analysis module determined the fault type to be "insufficient network interface card (NIC) transmit buffer" and recommended checking the NIC configuration, ring buffer size, or driver version.

[0106] Step S62: If the error code is normal but the interval between the sending timestamp and the arrival timestamp of the acknowledgment character from the other end is significantly greater than the historical average, then the fault type is determined to be external network path delay.

[0107] The packet-level data only contains the sending timestamp, while the ACK arrival timestamp needs to be obtained through additional mechanisms: one way is to simultaneously mount a TCP input path probe (such as tcp_v4_rcv) in the introspection mode of the edge computing node 10 to record the arrival time of each ACK and match it with the sequence number of the original packet; another way is to obtain it through peer feedback or passive measurement. When all error codes are 0 (indicating that there is no abnormality in the kernel sending path), the time difference (i.e., round-trip time) between the sending timestamp of each packet and the arrival timestamp of the corresponding ACK is calculated. If this time difference is significantly greater than the historical average round-trip time of the edge computing node 10 during normal periods (e.g., more than 3 times the standard deviation), it indicates that although the data packet was successfully sent locally, it encountered additional queuing, routing detours, or intermediate device processing delays on the network path, and the fault type is determined to be "external network path delay".

[0108] Step S63: If the packet sending rate corresponding to the same process identifier drops sharply within the time window and the central processing unit scheduling delay increases at the same time, the fault type is determined to be an internal process fault of the node.

[0109] The packet-level data records the process identifier (PID) for each packet, allowing us to calculate the number of packets sent per unit time (packet sending rate) by PID. If, within a time window, the packet sending rate of a certain PID suddenly drops from its normal value to below 20% of the normal rate, and scheduling latency metrics collected from internal state time-series snapshot 11 or introspection mode (such as / proc / sched_debug or the time from ready to running recorded by probes) show a significant increase in the CPU scheduling latency of that process (e.g., greater than 5 milliseconds, while normally less than 1 millisecond), it indicates that the process itself failed to be scheduled and executed by the CPU in a timely manner, resulting in the inability to generate or send data. In this case, the root cause of the failure lies in the abnormal process state within the node, rather than network or external factors.

[0110] For example, introspection data showed that the packet sending rate of process PID=5678 plummeted from 2000 packets / second to 50 packets / second within the abnormal window. Simultaneously, additional scheduling latency recorded by the probe showed that the process's average wait time jumped from 0.5 milliseconds to 25 milliseconds. The analysis module determined the fault type to be "internal process fault within the node" and triggered further deep introspection to further investigate the cause (such as infinite loops, lock contention, or memory thrashing).

[0111] Step S64: Associate and store the final fault type with the time window and the edge computing node 10 identifier, and generate a diagnostic report.

[0112] After completing the above diagnosis, the analysis module binds the determined fault type (such as "insufficient network card transmit buffer", "external network path delay", or "internal process failure of the node") with the abnormal time window recorded in step S2 and the associated edge computing node 10 identifier to form a complete diagnostic record. This record can be stored in a structured database or log file, and a readable diagnostic report is generated. The report includes at least: node identifier, abnormal time window, final fault type, key evidence summary (such as error code statistics, latency comparison, packet transmission rate curve, etc.), and suggested remedial measures (such as "increase the network card ring buffer" or "check for process dead loops"). The diagnostic report can be pushed to the central console 30 or consumed by the automated operation and maintenance system.

[0113] For example, the online monitoring and analysis module 20 ultimately outputs the following diagnostic report:

[0114] Node identifier: EdgeNode-12

[0115] Abnormal time window: 2025-01-15 09:05:30 - 09:08:20

[0116] Final Fault Type: Insufficient network card transmit buffer

[0117] Evidence Summary: 1200 outbound packets were recorded within 30 seconds, of which 1100 returned error code -ENOBUFS, accounting for 91.7%; the network card packet loss count increased by 450 within the window.

[0118] Repair suggestions: Check the network card driver configuration and increase tx_ring_size (e.g., ethtool -G eth0 tx2048); or upgrade the driver version.

[0119] On the other hand, when the fault type is an internal node process fault, the following steps are performed:

[0120] A third request is sent to the edge computing node 10 to trigger a deep introspection mode for the process corresponding to the process identifier;

[0121] The deep introspection mode includes collecting the process's CPU utilization, user mode and kernel mode time ratio, run queue waiting time, memory page fault rate, virtual memory region changes, page faults per second, system call latency distribution and return values ​​on the Transmission Control Protocol (TCP) transmission path, thread context switching count and the ratio of voluntary to involuntary switching.

[0122] The collected data is compared with the preset process health baseline. If the CPU utilization rate is consistently higher than the first threshold and the user mode ratio is too high, it is determined to be an application-layer dead loop or a computationally intensive anomaly.

[0123] If the CPU utilization is normal but the system call latency increases significantly, it is determined to be kernel resource contention or lock contention.

[0124] If the memory page fault rate or the number of page faults suddenly increases, it is determined to be an abnormal memory access mode or swap partition turbulence.

[0125] If the proportion of involuntary context switching increases abnormally, it is determined that the task is being preempted by a higher priority task or that the CPU resources are insufficient.

[0126] The deep diagnostic results are associated with the process identifier, the diagnostic report is updated, and repair suggestions are generated.

[0127] When the fault type is determined to be "internal process fault", it means that the root cause of the anomaly is that a specific process within the edge computing node 10 has exhibited abnormal behavior, rather than a network or external environment problem.

[0128] At this point, the sudden drop in packet sending rate and increase in scheduling latency in packet-level data alone are insufficient to pinpoint the specific cause (e.g., application-level infinite loop, kernel lock contention, memory thrashing, or resource preemption). Therefore, this embodiment further provides a deep introspection mechanism, which triggers finer-grained data collection for the process via a third request and compares it with a preset health baseline, thereby achieving accurate and in-depth fault diagnosis.

[0129] After arriving at the preliminary conclusion of "internal process failure within the node," the online monitoring and analysis module 20 immediately constructs a third request and sends it to the target edge computing node 10. The third request includes at least: the target node identifier, the target process identifier (PID), and an optional deep introspection duration (typically set to 10 to 30 seconds to avoid excessive overhead). Similar to the second request, the third request adopts an on-demand triggering principle, enabling deep introspection only when truly necessary.

[0130] Upon receiving a third request, the monitoring agent on edge computing node 10 dynamically loads a more granular set of data collection probes or calls process-level statistical interfaces provided by the operating system (such as perf_event in Linux, process context switch tracing in eBPF, and dynamic reading under / proc / [pid] / ). These probes are specifically designed to collect data from the specified PID and will not affect the normal operation of other processes.

[0131] The deep introspection mode collects the following metrics:

[0132] Central Processing Unit (CPU) Utilization: The percentage of CPU time used by this process, distinguishing between user mode and kernel mode time.

[0133] Run queue wait time: The time a process waits from the ready state to being scheduled to run, reflecting scheduling latency.

[0134] Memory page fault rate: includes minor page faults (already in memory) and major page faults (requiring swapping from disk). It is usually measured in "faults per second".

[0135] Virtual Memory Region (VMA) Changes: Monitor the number and frequency of changes in the virtual memory regions of the process. Frequent changes may indicate abnormal memory allocation / release.

[0136] Page faults per second: Similar to page fault rate, focus on tracking major page faults.

[0137] The latency distribution and return values ​​of system calls on the Transmission Control Protocol (TCP) send path: for example, the time taken by system calls such as write, send, and sendto (from entry to return) and the return value of each call (number of bytes successfully sent or error code). This data can be obtained by hooking the sys_sendmsg or tcp_sendmsg paths in the kernel.

[0138] Thread context switching frequency and the ratio of voluntary to involuntary switching: Voluntary context switching typically occurs when a process actively blocks (e.g., waiting for I / O); involuntary context switching occurs when the time slice expires or is preempted by a higher-priority task. An excessively high ratio often indicates CPU resource strain or improper priority settings.

[0139] Of the aforementioned metrics, some can be read at high speed from / proc / [pid] / stat, / proc / [pid] / status, and / proc / [pid] / schedstat; more granular latency distribution requires dynamic collection using eBPF probes. The monitoring agent aggregates the collected data and sends it back to the online monitoring and analysis module 20.

[0140] Before receiving in-depth introspection data, the online monitoring and analysis module 20 has already maintained a dynamic or static health baseline for common business processes. The baseline can be obtained through indicators collected during historical normal operation (such as the median and standard deviation of the same period in the past 7 days) or manually preset thresholds. The comparison method uses deviation detection and threshold discrimination.

[0141] If the CPU utilization rate consistently exceeds the first threshold and the user-mode ratio is too high, it is determined to be an application-layer infinite loop or a computationally intensive anomaly. The first threshold can be set to 80% (or adjusted according to business needs). If the CPU utilization rate of the target process consistently exceeds 80% during deep introspection, and the user-mode ratio exceeds 90% of the total CPU time (i.e., the kernel-mode ratio is very low), it indicates that the process spends most of its time executing in the application-layer code and does not get stuck in a large number of system calls. A typical scenario is that the application layer has an infinite loop (such as while(1) non-blocking) or undertakes a computational task beyond expectations (such as compression, encryption, etc.). After the determination, the output is "Application-layer infinite loop or computationally intensive anomaly".

[0142] If CPU utilization is normal but system call latency is significantly increased, it is determined to be kernel resource contention or lock contention. If the target process's CPU utilization is below a first threshold (e.g., 50%), but the latency of system calls (such as write) on the TCP send path is significantly greater than the historical baseline (e.g., average latency increases from 10 microseconds to 5 milliseconds), and the return value is sometimes -EAGAIN or partial write, it indicates that the process is waiting for certain resources (such as kernel locks, memory allocation, socket buffer space) in the kernel for too long. Common causes include lock contention in the kernel (such as sock_lock), resource contention in the file system or network subsystem, and it is determined to be "kernel resource contention or lock contention".

[0143] If the rate of page faults or the number of page faults suddenly increases, it is determined to be an abnormal memory access pattern or swap thrashing. If the collected major page fault rate exceeds the baseline (e.g., normally 0 times / second, now 50 times / second), or the number of page faults per second (including minor page faults) suddenly increases by tens of times, it indicates that the process is frequently accessing unmapped or swapped-to-disk memory pages. This usually occurs when insufficient memory leads to swap usage, or when the process has dangling pointers, or accesses freed areas after a memory leak (which may manifest as page faults before triggering a segmentation fault). This is determined to be "abnormal memory access pattern or swap thrashing".

[0144] If the proportion of involuntary context switches increases abnormally, it is determined that the process is being preempted by a higher-priority task or that CPU resources are insufficient. Involuntary context switch proportion = number of involuntary context switches / (voluntary context switches + involuntary context switches). A normal proportion is usually below 20%. If this proportion exceeds 50% during deep introspection and the absolute value is high (e.g., more than 1000 involuntary context switches per second), it indicates that the target process is frequently being forcibly deprived of CPU by other tasks. The reasons may include the presence of higher-priority real-time tasks, insufficient CPUs causing long queues, or the implementation of preemptive scheduling policies such as SCHED_FIFO. This is determined as "preemption by a higher-priority task or insufficient CPU resources."

[0145] The above-mentioned in-depth assessment results (e.g., "application-layer infinite loop or compute-intensive anomaly") are correlated with the node identifier, anomaly time window, and process identifier in the original diagnostic report, and an in-depth diagnostic field is added. Simultaneously, executable repair suggestions are generated based on the fault type.

[0146] For application-level infinite loops: it is recommended to restart the process and notify the development team to add a timeout exit or watchdog mechanism.

[0147] For kernel resource contention: it is recommended to optimize the application layer system call mode (such as batch sending), adjust kernel parameters, or upgrade the driver.

[0148] For memory access errors: it is recommended to increase memory, disable swap, or fix memory leak code.

[0149] For insufficient CPU resources: it is recommended to adjust process priority, increase CPU quota, or horizontally expand nodes.

[0150] The diagnostic report (including preliminary diagnosis, package-level data diagnosis, and deep introspection diagnosis) is stored in the database and pushed to the central console 30 or the automated repair engine. The probe is automatically unloaded after the preset time (e.g., 20 seconds) in deep introspection mode, and the edge computing node 10 resumes its lightweight monitoring state, which will not have a lasting impact on subsequent business.

[0151] On the other hand, this embodiment provides an intelligent online monitoring system for edge computing nodes 10 based on network monitoring. Please refer to the appendix. Figure 5 ,include:

[0152] The acquisition module 100 deploys a lightweight monitoring agent within the edge computing node 10. The monitoring agent collects aggregated metrics of outbound traffic by process or container dimension and records the internal state time-series snapshot 11 of the edge computing node 10 at a preset period.

[0153] The monitoring module 200 performs online monitoring and analysis of the outbound traffic of the edge computing node 10. When an abnormal traffic pattern is detected, it determines the time window of the abnormality and the associated edge computing node 10.

[0154] The first acquisition module 300 sends a first request to the edge computing node 10 to acquire the internal state time-series snapshot 11 data within the time window;

[0155] The first analysis module 400 parses the acquired time-series snapshot data, combines it with outbound traffic characteristics to perform preliminary fault diagnosis, and determines whether the fault type belongs to a preset distinguishable type.

[0156] If the preliminary diagnosis fails to determine that the fault belongs to a preset distinguishable type, the second acquisition module 500 sends a second request to the edge computing node 10, triggering the edge computing node 10 to start introspection mode and record the packet-level data of each outbound packet within the time window specified by the second request.

[0157] The second analysis module 600 receives and obtains the fault type based on the packet-level data, and automatically exits after the introspection mode has run for a preset duration.

[0158] Please see Figure 6 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this specification.

[0159] like Figure 6 As shown, the electronic device 1100 may include: at least one processor 1101, at least one network interface 1104, a user interface 1103, a memory 1105, and at least one communication bus 1102. The communication bus 1102 can be used to connect and communicate with the various components mentioned above. The user interface 1103 may include buttons, and optionally may include standard wired or wireless interfaces. The network interface 1104 may include, but is not limited to, a Bluetooth module, an NFC module, or a Wi-Fi module. The processor 1101 may include one or more processing cores. The processor 1101 connects to various parts within the electronic device 1100 using various interfaces and lines, and performs various functions of the routing device and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1105, and by calling data stored in the memory 1105. Optionally, the processor 1101 may be implemented using at least one hardware form of DSP, FPGA, or PLA. The processor 1101 may integrate one or more combinations of CPU, GPU, and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content that the display screen needs to show; and the modem is used for wireless communication.

[0160] It is understandable that the aforementioned modem may not be integrated into the processor 1101, but may be implemented using a separate chip.

[0161] The memory 1105 may include RAM or ROM. Optionally, the memory 1105 may include a non-transitory computer-readable medium. The memory 1105 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 1105 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 1105 may also be at least one storage device located remotely from the aforementioned processor 1101. As a computer storage medium, the memory 1105 may include an operating system, a network communication module, a user interface module, and application programs. The processor 1101 may be used to call the application programs stored in the memory 1105 and execute the methods in the above-described embodiments.

[0162] This specification also provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform multiple steps as described in the above embodiments. If the constituent modules of the above-described electronic device are implemented as software functional units and sold or used as independent products, they can be stored in the computer-readable storage medium.

[0163] This specification also provides a computer program product, including a computer program that, when executed by a processor, implements the multiple steps described in the above embodiments.

[0164] Where there is no conflict, the technical features in this embodiment and implementation scheme can be combined arbitrarily.

[0165] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes multiple computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this specification are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center integrating multiple available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).

[0166] When implemented through hardware or firmware, the aforementioned method flow is programmed into the hardware circuit to obtain the corresponding hardware circuit structure and achieve the corresponding function. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit, whose logic function is determined by the user programming the device. Designers can program a digital system onto a PLD themselves, eliminating the need for chip manufacturers to design and fabricate dedicated integrated circuit chips. Furthermore, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, similar to the software compiler used in program development. The original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There is not just one HDL, but many. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of the aforementioned hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logic method flow can be easily obtained.

[0167] The embodiments described above are merely preferred embodiments of this specification and are not intended to limit the scope of this specification. Any modifications and improvements made by those skilled in the art to the technical solutions of this specification without departing from the spirit of this specification should fall within the protection scope defined by the claims of this specification.

Claims

1. An intelligent online monitoring method for edge computing nodes based on network monitoring, characterized in that, Including the following steps: A lightweight monitoring agent is deployed within the edge computing node. The monitoring agent collects aggregated metrics of outbound traffic by process or container dimension and records time-series snapshots of the internal state of the edge computing node at a preset period. Online monitoring and analysis of outbound traffic from edge computing nodes; when abnormal traffic patterns are detected, the time window of the anomaly and the associated edge computing nodes are determined. Send a first request to the edge computing node to obtain internal state time-series snapshot data within the time window; The acquired time-series snapshot data is analyzed and combined with outbound traffic characteristics to perform preliminary fault diagnosis and determine whether the fault type belongs to a preset distinguishable type. If the initial diagnosis cannot determine that the fault belongs to a preset distinguishable type, a second request is sent to the edge computing node to trigger the edge computing node to start introspection mode and record packet-level data of each outbound packet within the time window specified in the second request. The system receives and obtains the fault type based on the packet-level data, and then automatically exits after the introspection mode has run for a preset duration.

2. The intelligent online monitoring method for edge computing nodes based on network monitoring according to claim 1, characterized in that, The monitoring agent performs the following steps: Aggregated metrics for outbound traffic statistics by process or container dimension, the aggregated metrics include one or more of the following: number of bytes sent, number of transmission control protocol retransmissions, and average round-trip time; The internal state time-series snapshots of the edge computing node are recorded at a preset period. The internal state time-series snapshots include the CPU load, run queue length, memory usage, transmission control protocol send queue size, and network card packet loss count. The snapshots are stored cyclically in a local circular buffer and retained for a preset duration.

3. The intelligent online monitoring method for edge computing nodes based on network monitoring according to claim 1, characterized in that, Methods for online monitoring and analysis of outbound traffic from edge computing nodes include: The detection is based on the bandwidth utilization, transmission control protocol retransmission rate, packet loss rate, and deviation of sudden changes in traffic from historical baselines in outbound traffic. When the deviation exceeds a preset threshold, it is determined to be an abnormal traffic pattern, and the time window of the abnormality and the associated edge computing node identifier are recorded.

4. The intelligent online monitoring method for edge computing nodes based on network monitoring according to claim 1, characterized in that, The method for analyzing the acquired time-series snapshot data and combining it with outbound traffic characteristics to perform preliminary fault diagnosis and determine the fault type includes: When the transmission control protocol send queue size is normal but the outbound retransmission rate is high, the fault type is inferred to be external network interference. When the Transmission Control Protocol (TCP) send queue is overflowing and the system memory is sufficient, the fault type is inferred to be slow application layer writes. When the length of the transmission queue momentarily exceeds the first threshold and the network card packet loss count increases at a millisecond time point, the fault type is inferred to be micro-burst packet loss. When the difference in the run queue length of multiple CPU cores exceeds the second threshold, the fault type is inferred to be a scheduling affinity configuration error. If the preliminary diagnosis result belongs to any of the above-mentioned preset distinguishable types, output the fault type and end; otherwise, output that the fault cannot be determined to belong to the preset distinguishable type.

5. The intelligent online monitoring method for edge computing nodes based on network monitoring according to claim 1, characterized in that, The method of sending a second request to the edge computing node to trigger the edge computing node to enter introspection mode and recording packet-level data of each outbound packet within the time window specified in the second request includes: The second request includes the duration of introspection and the filtering conditions to be tracked; After receiving the second request, the edge computing node dynamically loads the extended packet filter probe and attaches it to the kernel function dev_queue_xmit and the Transmission Control Protocol retransmission event; Within a specified time window, record the sending timestamp, target Internet Protocol address, process identifier, and error code returned by the kernel sending path for each outbound packet; The recorded packet-level data is compressed and then sent back.

6. The intelligent online monitoring method for edge computing nodes based on network monitoring according to claim 1, characterized in that, A method for receiving and obtaining the fault type based on the packet-level data includes: Parse the kernel send path error code in the packet-level data. If -ENOBUFS or -ENOMEM is detected, the fault type is determined to be insufficient network card send buffer or exhaustion of driver resources. If the error code is normal but the interval between the sending timestamp and the arrival timestamp of the acknowledgment character from the other end is significantly greater than the historical average, then the fault type is determined to be external network path delay. If the packet sending rate corresponding to the same process identifier drops sharply within the time window and the central processing unit scheduling delay increases at the same time, the fault type is determined to be an internal process fault of the node. The final fault type is associated with and stored in relation to the time window and edge computing node identifier, and a diagnostic report is generated.

7. The intelligent online monitoring method for edge computing nodes based on network monitoring according to claim 6, characterized in that, When the fault type is an internal node process fault, perform the following steps: Send a third request to the edge computing node to trigger a deep introspection mode for the process corresponding to the process identifier; The deep introspection mode includes collecting the process's CPU utilization, user mode and kernel mode time ratio, run queue waiting time, memory page fault rate, virtual memory region changes, page faults per second, system call latency distribution and return values ​​on the Transmission Control Protocol (TCP) transmission path, thread context switching count and the ratio of voluntary to involuntary switching. The collected data is compared with the preset process health baseline. If the CPU utilization rate is consistently higher than the first threshold and the user mode ratio is too high, it is determined to be an application-layer dead loop or a computationally intensive anomaly. If the CPU utilization is normal but the system call latency increases significantly, it is determined to be kernel resource contention or lock contention. If the memory page fault rate or the number of page faults suddenly increases, it is determined to be an abnormal memory access mode or swap partition turbulence. If the proportion of involuntary context switching increases abnormally, it is determined that the task is being preempted by a higher priority task or that the CPU resources are insufficient. The deep diagnostic results are associated with the process identifier, the diagnostic report is updated, and repair suggestions are generated.

8. An intelligent online monitoring system for edge computing nodes based on network monitoring, characterized in that: include: The acquisition module deploys a lightweight monitoring agent within the edge computing node. The monitoring agent collects aggregated metrics of outbound traffic by process or container dimension and records time-series snapshots of the internal state of the edge computing node at a preset period. The monitoring module performs online monitoring and analysis of outbound traffic from edge computing nodes. When an abnormal traffic pattern is detected, it determines the time window of the anomaly and the associated edge computing nodes. The first acquisition module sends a first request to the edge computing node to acquire internal state time-series snapshot data within the time window; The first analysis module parses the acquired time-series snapshot data, combines it with outbound traffic characteristics to perform preliminary fault diagnosis, and determines whether the fault type belongs to a preset distinguishable type. If the preliminary diagnosis fails to determine that the fault belongs to a preset distinguishable type, the second acquisition module sends a second request to the edge computing node, triggering the edge computing node to start introspection mode and record packet-level data of each outbound packet within the time window specified in the second request. The second analysis module receives and obtains the fault type based on the packet-level data, and automatically exits after the introspection mode has run for a preset duration.

9. An electronic device, characterized in that, Including the processor and memory; The processor is connected to the memory; The memory is used to store executable program code; The processor runs a program corresponding to the executable program code stored in the memory to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Communication equipment fault intelligent diagnosis method and device, equipment and medium

    CN120658571A

  • Chemical process anomaly type identification method based on multi-dimensional causal association index

    CN121561740A