Network monitoring data collection method and system

CN122741384APending Publication Date: 2026-09-11BEIJING YIHE TIANRUN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611018242.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

采集节点间的状态同步缺乏高效的原子操作支持,例如共享可达性映射的更新必须结合外部锁或事务,增加了实现复杂度并降低了吞吐量

Benefits of technology

[0057]系统启动时将设备信息、监控指标与轮询计划预加载至内存任务缓存,并预计算存储路径嵌入任务对象,避免了运行期间频繁数据库查询,显著降低I/O延迟与数据库压力。轮询调度器直接从内存读取任务驱动采集,结合订阅数据库增量变更通知仅对受影响条目实时更新,无需全量扫描,大幅提升调度响应速度与系统吞吐量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122741384A_ABST
    Figure CN122741384A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of network monitoring, and more particularly to a network monitoring data collection method and system, which loads configuration to the memory and subscribes to incremental update when the system starts, dynamically adjusts the frequency based on the reachability state before collection, uses guard object management when borrowing UDP session, stores data according to the routing prompt after collection, uses atomic operation to update the count and expose the telemetry interface when brushing out, improves the collection efficiency and reliability, and reduces the resource consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network monitoring technology, and in particular to a method and system for acquiring network monitoring data. Background Technology

[0002] Current data acquisition methods in the network monitoring field generally adopt a centralized scheduling architecture. The system typically relies on a relational database as the core hub for task storage and status synchronization, with all device information, monitoring metrics, and polling plans stored uniformly in the database tables. The scheduler periodically polls the database to obtain tasks to be executed and issues acquisition instructions to each acquisition node or thread. After the acquisition node completes data acquisition, it directly writes the raw counter values ​​into the raw data record table in the database, and updates the cumulative count and time consumption through database transactions. This type of solution can operate stably when the number of devices is small and the acquisition frequency is low, with low system complexity and ease of maintenance.

[0003] Task scheduling relies entirely on a database polling mechanism, causing the scheduling cycle to be limited by database query performance. As the scale of equipment increases or the polling frequency rises, frequent database read and write operations quickly become a system bottleneck, increasing response latency and consuming a large amount of database connection resources. At the same time, dynamic changes in equipment status (such as network interruption or recovery) cannot be detected in real time, and the scheduler must wait for the next polling cycle to adjust the task execution strategy, resulting in blind spots or redundant detection.

[0004] Data write and update operations typically rely on database row-level locking or optimistic locking mechanisms, which can easily lead to lock contention or retry overhead in concurrent data collection scenarios. State synchronization between collection nodes lacks efficient atomic operation support; for example, updates to shared reachability maps must be combined with external locks or transactions, increasing implementation complexity and reducing throughput. Furthermore, session management often employs simple timeout recycling or fixed connection pool strategies, failing to differentiate handling of network anomalies (such as timeouts and fatal errors), resulting in low session resource utilization and even memory leaks caused by residual abnormal sessions. Summary of the Invention

[0005] The embodiments of the present invention provide a method and system for collecting network monitoring data, which can solve the problems in the prior art.

[0006] A first aspect of the present invention provides a method for collecting network monitoring data, comprising:

[0007] When the system starts, it loads device information, monitoring metrics set and polling plan from the database into the memory task cache, and pre-calculates the data storage path for each task and embeds it into the corresponding task object as a routing prompt. During operation, the polling scheduler reads task-driven data collection from the memory task cache, subscribes to database incremental change notifications, and performs add, delete, modify and update operations on the affected entries.

[0008] Before executing a task, read the device status in the shared reachability map using atomic operations. If the status is normal, execute at the planned frequency. If the status is unreachable, determine whether the probe point has been reached based on the frequency reduction interval. If the probe is reached, execute the probe and update the status atomically based on the result. If the status is not reached, skip the task.

[0009] When borrowing a UDP session, a guard object is created and marked as failed by default. After collection is completed, it is marked as successful or retained as failed based on the result. When the guard is destroyed, it is categorized and evicted: if successful, it is returned to the pool and the health count is reset; if it times out, it is closed, rebuilt and added to the pool; if it is a fatal error, it is destroyed and removed. When evicting, the session generation counter is checked to be consistent with the borrowing record before execution.

[0010] During data collection, the raw counter values ​​are obtained from the device and stored in the raw data table according to the routing prompts in the task object. The derived indicators are dynamically calculated by the query layer view based on the raw data. When the data is refreshed, the cumulative count and time are updated lock-free through atomic comparison and exchange operations, generating refresh evidence records and exposing the telemetry interface.

[0011] The system loads device information, monitoring metrics sets, and polling plans from the database into the in-memory task cache, and pre-calculates the data storage path for each task, embedding it into the corresponding task object as a routing hint, including:

[0012] Read device records from the database containing the device's unique identifier, network address, transport layer protocol type, and authentication credentials, and construct a device object hash table using the device's unique identifier as the key.

[0013] Read monitoring indicator definition records from the database, which contain indicator names, identifiers of objects to be collected, and numerical types. Establish an association index structure from the device object to the indicator definition set using the device model identifier as the key.

[0014] Read the polling plan record containing device identifier, polling period and retry parameters from the database, and establish a correspondence between device identifier and device object;

[0015] The device identifiers in the device object diagram and the corresponding monitoring indicator definition set for the device model are expanded by Cartesian according to the indicator dimension to generate one-to-one corresponding collection task entries. Each entry encapsulates the device identifier, network address, protocol type, single object identifier to be collected, and polling parameters.

[0016] Based on the device identifier of the data collection task entry, the storage configuration metadata is queried to read the target table name and partition information. The target table name prefix, device identifier hash value and time partition sharding key are concatenated to form a complete storage path string. The complete storage path is bound to the corresponding data collection task entry as a routing hint field, so that the runtime refresh component can directly complete the storage location based on the routing hint without performing path parsing query.

[0017] Subscribe to database incremental change notifications and perform add, delete, modify, and update operations on affected entries, including:

[0018] Establish a change monitoring connection with the database based on a publish-subscribe mechanism, and register row-level change monitoring for the device configuration table, monitoring template table, and polling plan table with the database, covering insert, update, and delete operation types;

[0019] After receiving the change notification pushed by the database, the operation type, the changed table name and the affected primary key value are parsed from the event payload. The corresponding set of entries in the memory task cache is located by mapping the pre-maintained primary key to the cache key.

[0020] If the operation type is insert, the complete record is read back from the database based on the primary key value to construct a new collection task entry, which is then inserted into the memory task cache and the primary key mapping is updated.

[0021] If the operation type is update, the changed record is read back from the database based on the primary key value, and the corresponding entry in the memory task cache is replaced and updated at the field level. If the entry is being held by the collection thread, the new version is temporarily stored using the copy-on-write strategy and the replacement is completed after the thread is released.

[0022] If the operation type is deletion, the corresponding entry is located through the primary key mapping, the entry is removed from the memory task cache, and the primary key mapping is unregistered;

[0023] If a readback operation times out or returns incomplete data, the known complete state of that entry in the memory task cache is preserved and an alarm notification is generated. The normal scheduling and execution of other entries are not affected.

[0024] Before executing the task, read the device status in the shared reachability map using atomic operations to determine whether the probe point has been reached, including:

[0025] A shared reachability mapping table is maintained in memory, with the key being a unique device identifier and the value being a reachability state structure. The reachability state structure contains a reachability Boolean flag, the first unreachable timestamp, and the most recent probe timestamp.

[0026] Before starting to process a single device task, the polling scheduler reads the current value of the reachability Boolean flag in the reachability state structure corresponding to the device using an atomic load instruction;

[0027] If the flag is true, the collection polling interaction will be executed directly according to the pre-set frequency of the task entry. After receiving a valid response, the reachability Boolean flag will be kept true by atomic writing.

[0028] If the flag is false, the current timestamp is obtained, the time interval between the current time and the most recent detection timestamp is calculated, and the value is compared with the preset frequency reduction interval threshold parameter in the task entry.

[0029] If the time interval is less than the frequency reduction interval threshold parameter, the current round of scheduling is abandoned, the current worker thread slot and the allocated network receive buffer are released, and the next device task is processed.

[0030] If the time interval is not less than the frequency reduction interval threshold parameter, a lightweight probe request is sent to the device by borrowing a UDP session from the session pool. The result is determined based on whether a valid response is received within the preset timeout threshold, and the reachability status is updated accordingly.

[0031] Based on the detection results, the reachability flag is updated, and the detection frequency is controlled according to the frequency reduction interval threshold, including:

[0032] When the detection is successful, the reachability Boolean flag in the reachability state structure is set to true by atomic comparison and swap operation, the first unreachable timestamp field value is cleared and the most recent detection timestamp is updated to the current time. At the same time, the frequency reduction interval threshold parameter is reset to the preset initial frequency reduction interval threshold so that the device can perform acquisition interaction at the original planned frequency in subsequent scheduling.

[0033] When a detection fails, the reachability Boolean flag is kept false by comparing and swapping atomic operations, the most recent detection timestamp atom is updated to the current time, and the current frequency reduction interval threshold parameter is kept unchanged for use in the next scheduling cycle.

[0034] If the current down-frequency interval threshold is the initial value, after the first detection fails, the down-frequency interval threshold will be expanded to the second-level down-frequency interval threshold by the backoff multiplier factor. After each subsequent detection failure, the backoff multiplier factor will be multiplied to obtain the next-level down-frequency interval threshold. If the product exceeds the preset upper limit value, the down-frequency interval threshold will be clamped to the upper limit value.

[0035] If the device is marked as unreachable again due to continuous failures after a successful detection, it will re-execute the step-by-step backoff growth strategy starting from the initial frequency reduction interval threshold until the upper limit of the frequency reduction interval threshold is reached, so as to ensure that the detection frequency is adaptively adjusted in the network environment where the device is repeatedly connected and disconnected.

[0036] When borrowing a UDP session, a guard object is created and marked as failed by default. Guards are categorized and evicted upon destruction, including:

[0037] The idle session handle is obtained from the UDP session connection pool in a mutually exclusive manner. When constructing the guard object, the generation counter value of the current session is read in atomic operation and recorded in the borrow generation number field of the guard object. The internal status flag field is initially set to the failure value. The guard object takes over the life cycle management of the session handle.

[0038] The caller uses the borrowed session to send a data acquisition request packet to the target device and enters a blocking wait. After receiving the data packet within the timeout threshold, it performs an integrity check on the returned data packet. If the check passes, the internal status flag is set to a success value. If a timeout, check failure, or error status code occurs, the failure value is retained without modification.

[0039] When the guard object is destructed, the current generation value of the session is first read atomically and compared with the borrowed generation number. If the two are not equal, it means that the session has been rebuilt and borrowed to another caller during the borrowing period, and the current eviction logic is abandoned.

[0040] If the generation number is equal and the status is successful, the session is returned to the connection pool and the consecutive success count is incremented. If the generation number is equal and the error type is timeout, the socket is closed, a new session is added to the connection pool, and the consecutive success count is reset. If the generation number is equal and the error type is socket-level error, the socket is closed, all resources are released, and the session slot is removed from the connection pool.

[0041] When a new session is added to the connection pool, a new generation counter value is assigned to prevent eviction operations from accidentally affecting newly borrowed sessions, including:

[0042] The connection pool maintains a global monotonically increasing generation number allocation counter. Each time a new underlying session object is created, the new generation counter value is obtained by atomic increment operation and written to the generation counter field of the session object. For session objects with the same device identifier, strictly increasing generation counter values ​​are allocated sequentially according to the creation time to distinguish different generation session objects.

[0043] Before executing the eviction decision, the guard object destruction logic reads the current generation counter value of the underlying session object again with an atomic load instruction with full memory barrier semantics, and compares it with the borrow generation number field value recorded during guard construction. The full memory barrier ensures that all previous release write operations of other threads are visible to this read.

[0044] If the current generation value is equal to the borrowed generation number, it indicates that the session has not been destroyed and rebuilt from the time of borrowing to the time of destruction, and the guard's expulsion decision can be safely applied to the session object.

[0045] If the current generation value is not equal to the borrowed generation number, it indicates that the session has been destroyed and replaced by a new session during the borrowing period, and the new session has been borrowed for use. In this case, the eviction should be abandoned to prevent the invalid destruction or incorrect return of a valid new session that is currently in use, and to prevent delayed timeout evictions from causing accidental damage to new sessions that are in normal use.

[0046] A second aspect of the present invention provides a network monitoring data acquisition system, comprising:

[0047] The scheduling unit is used to load device information, monitoring indicator set and polling plan from the database to the memory task cache when the system starts up, and pre-calculate the data storage path for each task and embed it into the corresponding task object as a routing prompt. During operation, the polling scheduler reads the task-driven collection from the memory task cache, subscribes to the database incremental change notification, and performs add, delete, modify and update operations on the affected entries.

[0048] The reachable unit is used to read the device status in the shared reachability map with atomic operations before executing the task. If it is normal, it will be executed at the planned frequency. If it is unreachable, it will determine whether the probe point has been reached based on the frequency reduction interval. If it is reached, the probe will be executed and the status will be updated atomically according to the result. If it is not reached, it will be skipped.

[0049] Session unit is used to create guard objects when borrowing a UDP session and mark them as failed by default. After collection is completed, it is marked as successful or retained as failed based on the result. When the guard is destroyed, it is classified and expelled: if successful, it is returned to the pool and the health count is reset; if it times out, it is closed, rebuilt and added to the pool; if it is a fatal error, it is destroyed and removed.

[0050] The verification unit is used to verify that the session generation counter matches the borrowing record before execution during expulsion.

[0051] The data acquisition unit is used to obtain the raw counter value from the device during data acquisition, store it in the raw data table according to the routing prompts in the task object, and dynamically calculate the derived indicators based on the raw data by the query layer view. When the data is refreshed, the cumulative count and time are updated lock-free through atomic comparison and exchange operations, generating refresh evidence records and exposing the telemetry interface.

[0052] A third aspect of the present invention provides an electronic device, comprising:

[0053] processor;

[0054] Memory used to store processor-executable instructions;

[0055] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0056] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0057] When the system starts, device information, monitoring metrics, and polling plans are preloaded into the memory task cache, and the storage path is pre-calculated and embedded into the task object. This avoids frequent database queries during runtime, significantly reducing I / O latency and database pressure. The polling scheduler directly reads task-driven data from memory and, combined with subscribing to incremental database change notifications, only updates affected entries in real time, eliminating the need for a full scan and greatly improving scheduling response speed and system throughput.

[0058] Before task execution, the device status in the shared reachability map is read through atomic operations. Under normal conditions, data is collected at the planned frequency; when unreachable, the timing of probing is determined based on the frequency reduction interval. If the device is reached, probing is performed and the status is updated atomically; otherwise, the process is skipped. This mechanism avoids invalid polling of unreachable devices, saving network bandwidth and computing resources, and preventing false positives through frequency reduction probing, ensuring the real-time performance and accuracy of state transitions.

[0059] When borrowing a UDP session, a guard object is created and marked as failed by default. After collection is complete, the mark is updated based on the results. When a guard is destroyed, it is categorized and removed: successfully borrowed guards are returned to the pool and their health count is reset; guards that time out are closed, rebuilt, and added back; and guards with fatal errors are completely removed. An intergenerational counter and borrowing record verification are also introduced to ensure thread safety during session return. This mechanism effectively prevents session leakage and resource exhaustion, improves session reuse, automatically isolates abnormal connections, and ensures the stability of the collection chain.

[0060] The original counter values ​​are directly stored in the original data table based on the routing hints in the task object. Derived metrics are dynamically calculated by the query layer view based on the original data, eliminating the need to pre-store redundant fields, reducing storage costs and maintaining data flexibility. When data is flushed, the cumulative count and time are updated lock-free through atomic comparison and exchange operations, avoiding lock contention, generating flushing evidence records and exposing the telemetry interface for real-time monitoring of system operation status and anomaly tracing. Attached Figure Description

[0061] Figure 1 This is a flowchart illustrating the network monitoring data acquisition method.

[0062] Figure 2 Flowchart for distributed singleton protection;

[0063] Figure 3 Flowchart for automatic identification and application of device templates. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0065] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0066] refer to Figures 1 to 3 The network monitoring data acquisition method of this invention includes:

[0067] When the system starts, it loads device information, monitoring metrics set and polling plan from the database into the memory task cache, and pre-calculates the data storage path for each task and embeds it into the corresponding task object as a routing prompt. During operation, the polling scheduler reads task-driven data collection from the memory task cache, subscribes to database incremental change notifications, and performs add, delete, modify and update operations on the affected entries.

[0068] Before executing a task, read the device status in the shared reachability map using atomic operations. If the status is normal, execute at the planned frequency. If the status is unreachable, determine whether the probe point has been reached based on the frequency reduction interval. If the probe is reached, execute the probe and update the status atomically based on the result. If the status is not reached, skip the task.

[0069] When using a UDP session, a guard object is created and marked as failed by default. After data collection is complete, it is marked as successful or retained as failed based on the results. When a guard is destroyed, it is categorized and evicted.

[0070] Successfully returned to the pool and the health count reset; if timeout occurs, the pool will be closed and rebuilt; if a fatal error occurs, the pool will be destroyed and removed.

[0071] Upon eviction, the session generation counter is checked to ensure it matches the borrowing record before execution.

[0072] During data collection, the raw counter values ​​are obtained from the device and stored in the raw data table according to the routing prompts in the task object. The derived indicators are dynamically calculated by the query layer view based on the raw data. When the data is refreshed, the cumulative count and time are updated lock-free through atomic comparison and exchange operations, generating refresh evidence records and exposing the telemetry interface.

[0073] In one optional implementation, the method of this application is applied to a network monitoring data acquisition system, which includes: a memory task caching module, a pessimistic RAII session management module, a distributed singleton protection module, a device template matching module, an adaptive polling control module, a raw data acquisition pipeline, a data transmission evidence telemetry module, and a query layer derived view module.

[0074] The memory task caching module performs an initialization process upon system startup: it reads all device information, monitoring templates, metric OID sets, and polling plans (period, timeout, etc.) from the relational database to construct a complete memory task object graph. It should be noted that in this system, the "monitoring template" corresponds to a metric pack in the specific code implementation, and is managed through a pack registry and pack catalog.

[0075] During the loading phase, path information such as the target table name and partition key for data storage is pre-parsed and concatenated based on device and metric dimensions, and directly embedded into the task object as a prompt field. This way, during runtime data collection and writing to the persistence layer, there's no need for conditional branches or table lookups to determine the storage location; the pre-calculated path is used directly, maximizing code execution efficiency. At runtime, the polling scheduler uses configurable threads / coroutines concurrently to directly traverse tasks from the aforementioned memory cache, driving data collection without accessing the database throughout the entire process.

[0076] Cache updates rely on an incremental cache synchronization mechanism:

[0077] Leveraging the database's LISTEN / NOTIFY feature, the database proactively pushes change events when administrators add, delete, or modify devices, templates, or polling schedules. The caching module subscribes to these events and only adds, deletes, or modifies a small number of affected cache entries, achieving surgical updates without requiring a full reload.

[0078] It also features fault-tolerant refresh:

[0079] If the data of a single device is corrupted or abnormal during an update, the system retains the device's last known good cache state and issues an alarm. The caches of other devices are completely unaffected, thus avoiding a global failure.

[0080] Pessimistic RAII Session Fortress and Three-Way Expulsion Mechanism

[0081] The system constructs a UDP session connection pool. Each time a session is borrowed from the pool, instead of returning a raw session handle directly, it is encapsulated in a RAII guard object.

[0082] By default, the guard is constructed with the session's internal state marked as "may fail".

[0083] After the data collection operation is completed, the caller must explicitly call `mark_success()` based on the result to mark the session as successful. If an exception is thrown during the operation or the function returns prematurely, causing the guard to be destroyed without being marked as successful, the default failure status is retained.

[0084] The destruction of the guards triggers a three-way expulsion category:

[0085] Success path: If a success has been marked, the session is considered healthy, returned to the connection pool, and its health counter (such as the number of consecutive successes) is reset.

[0086] Timeout Path: If a failure occurs due to an I / O timeout, the session is marked as "preventive rebuild," and a one-hit 'em" strategy is implemented—that is, the session is not returned to the pool, but is directly closed and a new session is created and added to the pool. This effectively prevents potentially delayed UDP response packets from being misread in the future and polluting subsequent requests (socket poisoning).

[0087] Fatal error path: If a socket error or protocol stack error occurs, the session is immediately destroyed and completely removed from the pool.

[0088] To prevent race conditions, each underlying session is assigned a monotonically increasing generation counter.

[0089] When a session is rebuilt, the new session acquires a new generation. During an eviction operation, it is necessary to verify whether the current generation of the session is consistent with the record at the time of borrowing. If they do not match (indicating that the session handle has been rebuilt and re-borrowed), the eviction is abandoned to ensure that a new session is not mistakenly deleted.

[0090] This combined approach enables deterministic lifecycle management of UDP sessions, eliminating the risks of socket poisoning and resource leaks.

[0091] In a preferred embodiment of the present invention, the system may also provide a service singleton guarantee mechanism without external dependencies. After the service process starts, it immediately uses the service name (e.g., "network-poller") as input and calculates an integer key using a deterministic hash algorithm. Then, it sends a request to the shared relational database to attempt to acquire a session-level advisory lock, such as PostgreSQL's pg_try_advisory_lock. This lock is bound to the lifecycle of the database connection. If successfully acquired, the process continues to execute all data collection tasks and keeps the database connection active. If the lock is already held by another instance, acquisition fails, the current process logs its exit, and then immediately exits. When the process holding the lock crashes or is terminated, its database connection is broken, the database automatically releases the session-level advisory lock, and other waiting or later-starting instances can immediately acquire the lock and take over the work. The entire process does not require any external coordination services and has crash safety features.

[0092] Metric Pack System and Automatic Identification Applications:

[0093] In a specific embodiment of the present invention, this device template system is also referred to as a metric pack system, which manages template metadata and metric definitions through a pack registry and a pack catalog. For consistency, the term "device template system" will continue to be used below.

[0094] The system pre-maintains a device template knowledge base in the database tables. Each template record contains matching conditions (system object identifier, SysOID, such as `1.3.6.1.4.1.9...`) and the corresponding set of collected indicators (OID list, calculation factors, etc.). When the system starts, the entire template knowledge base is loaded into the memory cache.

[0095] When the polling scheduler encounters a device for the first time (based on its unique identifier) ​​and the device is not yet associated with any specific data collection metrics in the cache, an automatic identification process is triggered: the data collection engine sends an SNMP GET request to the device to obtain its SysOID; after obtaining the SysOID, it performs best prefix matching in the in-memory template library to find the corresponding template; once a match is found, all metrics defined in that template are instantiated and associated with the device in the in-memory task cache. Subsequent polling of that device will then be performed according to the matched metric set. The entire process is completely transparent to the administrator and is plug-and-play. Adding support for new device models only requires inserting a template record into the database, which takes effect through incremental synchronization, without requiring any code modifications or service restarts.

[0096] Adaptive polling gates and shared reachability states:

[0097] The system maintains a shared reachability mapping table (hash table, with device IDs as keys and state structures and atomic flags as values) in memory. This mapping table is visible to all polling worker threads.

[0098] In the polling scheduling process, before executing a specific device task, the reachability status of the device is read using lock-free atomic operations:

[0099] If the status is "normal", then polling will be performed immediately at the planned frequency.

[0100] If the status is "unreachable", read the current time and compare it with the last probe time on the device to determine if the frequency reduction interval is met (e.g., 30 seconds at normal frequency, 5 minutes after frequency reduction). If not met, skip this round and release the relevant resources; if met, perform a probe poll.

[0101] After the probe polling is completed, the state is updated using atomic operations based on the result: if successful, it is changed to "normal" and the original high-frequency polling is restored in the next scheduling; if it fails, it remains "unreachable" and the last probe timestamp is updated, and the frequency reduction continues.

[0102] In this way, the system converts the vast majority of invalid waits into no-operations, significantly reducing resource consumption caused by unreachable devices, while maintaining a reasonable rate of awareness of device recovery.

[0103] After receiving the counters returned by the device (such as the number of bytes received / sent via the interface, the number of error packets, etc.), the acquisition engine does not perform any arithmetic transformations or business calculations. It simply stores the raw counter values ​​along with the timestamp, device identifier, and interface index into the raw data table of the database. All formula definitions and threshold calculations are completed by the database view or materialized view of the upper query layer. For example, bandwidth utilization is calculated in real time by dividing the difference in bytes between the current and previous sampling points by the time interval and the interface rate through the view. When it is necessary to correct the calculation formula, only the view definition needs to be modified, and the historical data can be retrieved again to obtain the new indicator value without the need for backfilling or re-acquisition.

[0104] At the end of the data acquisition pipeline is an asynchronous buffer flushing component. Each time this component writes a batch of data (typically hundreds to thousands of samples) to the persistent layer, it constructs a `FlushEvidence` record, which includes: the total number of samples flushed, the number of successful writes, the number of failed writes, the time taken, and error information. The cumulative values ​​of these statistics (total flushes, total successful samples, and total failed samples) are updated lock-free by multiple flushing tasks using CPU atomic compare-and-swap (CAS) instructions, completely avoiding mutex lock overhead. The internal telemetry API exposes these atomic counters and the most recent flush evidence record, allowing operators or built-in health checks to understand the throughput and error status of the data pipeline in real time, without any blockage of the acquisition hot path.

[0105] The modules described above work together seamlessly: memory caching enables extremely fast polling scheduling with no database dependency; pre-computed route hints eliminate runtime path judgment, enabling direct data persistence to disk; adaptive polling gates reduce unnecessary connection attempts, protecting the RAII session pool from being exhausted by unreachable devices; RAII guards guarantee the deterministic state of sessions, working with the adaptive polling gates to improve system robustness under anomalies; device template self-identification further reduces the burden of cache maintenance; singleton protection ensures that the aforementioned memory cache and state uniqueness do not conflict; and raw data policies and flashing evidence telemetry provide the system with complete auditing and self-monitoring capabilities. This comprehensive approach addresses the inherent shortcomings of traditional architectures.

[0106] For memory cache synchronization, if you don't want to rely on the database's LISTEN / NOTIFY mechanism, you can incrementally refresh the cache by periodically polling the database change timestamp column, but the real-time performance is slightly worse. For RAII session management, the three-way classification can be simplified to two-way (successful return / failure destruction), but this will lose the one-hit rebuild optimization for timeouts and increase the frequency of resource rebuilds. For singleton protection, in databases without advisory lock support, you can use a unique constraint table row to simulate a lock, but this requires additional expiration cleanup logic.

[0107] The system loads device information, monitoring metrics sets, and polling plans from the database into the in-memory task cache, and pre-calculates the data storage path for each task, embedding it into the corresponding task object as a routing hint, including:

[0108] Read device records from the database containing the device's unique identifier, network address, transport layer protocol type, and authentication credentials, and construct a device object hash table using the device's unique identifier as the key.

[0109] Read monitoring indicator definition records from the database, which contain indicator names, identifiers of objects to be collected, and numerical types. Establish an association index structure from the device object to the indicator definition set using the device model identifier as the key.

[0110] Read the polling plan record containing device identifier, polling period and retry parameters from the database, and establish a correspondence between device identifier and device object;

[0111] The device identifiers in the device object diagram and the corresponding monitoring indicator definition set for the device model are expanded by Cartesian according to the indicator dimension to generate one-to-one corresponding collection task entries. Each entry encapsulates the device identifier, network address, protocol type, single object identifier to be collected, and polling parameters.

[0112] Based on the device identifier of the data collection task entry, the storage configuration metadata is queried to read the target table name and partition information. The target table name prefix, device identifier hash value and time partition sharding key are concatenated to form a complete storage path string. The complete storage path is bound to the corresponding data collection task entry as a routing hint field, so that the runtime refresh component can directly complete the storage location based on the routing hint without performing path parsing query.

[0113] When reading device records from the database, four core fields are extracted for each record: unique device identifier, network address, transport layer protocol type, and authentication credentials. The unique device identifier is used to uniquely identify a managed device throughout the entire data collection lifecycle. The network address records the routable address of the device in the network topology (either IPv4 or IPv6). The transport layer protocol type identifies the communication protocol to be selected when establishing subsequent sessions (e.g., SNMP over UDP, gRPC over TCP). The authentication credentials encapsulate the authentication materials required for handshaking with the device, such as the community string, username / password, or certificate fingerprint. After reading, a hash table is built in memory using the unique device identifier as the key and the device object encapsulating the above four fields as the value. This allows any subsequent operations requiring rapid device attribute location by device identifier to be completed in constant time, avoiding repeated database lookups during the task deployment phase.

[0114] When reading monitoring metric definition records from the database, each record contains three fields: metric name, identifier of the object to be collected, and value type. The metric name is the semantic label of the metric within the system, such as the number of bytes of inbound traffic to the interface or the percentage of CPU utilization. The identifier of the object to be collected is the resource path that can be directly addressed on the device side, such as an SNMP OID string or a gRPC path. The value type describes the data semantics of the metric, distinguishing between counter types (monotonically increasing, requiring differential calculation rate) and meter types (instantaneous values, directly usable). After reading, using the device model identifier as the key, all metric definition records under the same model are aggregated into a set, establishing an association index structure from device objects to the metric definition set. The significance of this index structure is that multiple devices of the same model share the same metric definition set, eliminating the need to store a separate copy for each device, saving memory while ensuring model-level metric configuration consistency.

[0115] When reading polling plan records from the database, each record contains a device identifier, polling period, and retry parameters. The polling period determines the time interval at which the scheduler adds the corresponding device's acquisition task to the execution queue, while the retry parameters include the maximum number of retries and the base latency value required for the backoff strategy. After reading, the polling plan records are mapped to existing device objects in the hash table based on their device identifiers. This allows each device object to directly carry its scheduling parameters during subsequent task deployment phases without needing to perform a cross-structure lookup at runtime.

[0116] The task expansion phase involves expanding the device dimension and the metric dimension using a Cartesian product. Each device in the device object graph is traversed, and its corresponding metric definition set is searched in the associated index structure based on its device model identifier. For each metric definition record in this set, an expansion operation is performed, generating an independent data collection task entry. Each data collection task entry encapsulates the following fields: device identifier, network address, protocol type, single target object identifier, and polling parameters. The emphasis here on the "single target object identifier" is a key design decision—decomposing each metric into an independent task entry, rather than merging all metrics of the same device into a batch task, allows the scheduler to perform fine-grained frequency control, timeout management, and failure isolation at the metric granularity. The failure to collect a single metric will not contaminate the collection results of other metrics for the same device. After expansion, the number of task entries in the in-memory task cache equals the sum of the number of metrics for all devices, forming a flattened task list for the polling scheduler to directly traverse.

[0117] The routing hints are a core mechanism for reducing runtime path resolution overhead. For each generated data collection task entry, the storage configuration metadata is queried based on its device identifier to retrieve the target table name prefix and partition information. The target table name prefix is ​​the logical namespace prefix of the original data table in the storage layer, such as "raw_metrics"; the device identifier hash value is obtained by performing a fixed-length hash operation on the unique device identifier, used to route data from different devices to different physical shards in horizontal sharding scenarios, avoiding hotspot concentration; the time partition shard key is generated as a partition identifier string according to the expected write time window of the data collection task at a predetermined granularity (such as by day or by hour). The target table name prefix, device identifier hash value, and time partition shard key are concatenated with a predefined separator to form a complete storage path string, such as a hierarchical path format like "raw_metrics / shard_3 / 2024_06_15".

[0118] The complete storage path string is directly bound to the corresponding data collection task object as a routing hint field, making this field an inherent attribute of the task object and flowing with the task through scheduling, execution, and flushing stages. When the runtime flushing component writes the collected raw counter values ​​to the storage layer, it can directly read the routing hint field on the task object to locate the target table and partition, completely bypassing the path resolution query operation that originally required accessing configuration metadata. This design is highly effective in high-frequency data collection scenarios: if thousands of tasks trigger flushing per second, and each flush requires a path resolution query, the metadata storage will experience read pressure equal to the data collection frequency; after pre-computing and embedding the routing hint, metadata queries are only performed during system startup or configuration changes, reducing the read pressure on the metadata storage to near zero during runtime.

[0119] The handling of time partition keys requires special explanation. Since route hints are pre-calculated at startup, and the data collection task will continue to execute for a considerable period, time partition keys may cross partition boundaries. To address this, the time partition portion of the route hint can be designed with a lazy update mechanism: when the refresh component uses the route hint, it first checks if the time partition key is still within the current time window. If it has expired, it reassembles the partition key of the current time window locally and overwrites the route hint field on the task object. This operation only involves string concatenation and does not require access to metadata storage, thus completely avoiding path resolution queries.

[0120] After subscribing to database incremental change notifications, when device records, metric definitions, or polling plans change, add, delete, modify, and update the affected task entries. The update process also includes recalculating the routing hints to ensure that the routing hints in the memory task cache are consistent with the latest storage configuration metadata, and that the flushing component will not write data to expired or incorrect storage paths due to configuration changes.

[0121] Subscribe to database incremental change notifications and perform add, delete, modify, and update operations on affected entries, including:

[0122] Establish a change monitoring connection with the database based on a publish-subscribe mechanism, and register row-level change monitoring for the device configuration table, monitoring template table, and polling plan table with the database, covering insert, update, and delete operation types;

[0123] After receiving the change notification pushed by the database, the operation type, the changed table name and the affected primary key value are parsed from the event payload. The corresponding set of entries in the memory task cache is located by mapping the pre-maintained primary key to the cache key.

[0124] If the operation type is insert, the complete record is read back from the database based on the primary key value to construct a new collection task entry, which is then inserted into the memory task cache and the primary key mapping is updated.

[0125] If the operation type is update, the changed record is read back from the database based on the primary key value, and the corresponding entry in the memory task cache is replaced and updated at the field level. If the entry is being held by the collection thread, the new version is temporarily stored using the copy-on-write strategy and the replacement is completed after the thread is released.

[0126] If the operation type is deletion, the corresponding entry is located through the primary key mapping, the entry is removed from the memory task cache, and the primary key mapping is unregistered;

[0127] If a readback operation times out or returns incomplete data, the known complete state of that entry in the memory task cache is preserved and an alarm notification is generated. The normal scheduling and execution of other entries are not affected.

[0128] When establishing a change monitoring connection with the database, a publish-subscribe mechanism is used instead of polling queries, fundamentally eliminating the latency and resource overhead caused by periodic full table scans. Specifically, after the system starts and the in-memory task cache is initialized, a dedicated long-lived connection is immediately initiated to the database. This connection is maintained independently of the regular read-write connection pool to ensure that the real-time nature of change notifications is not affected by business query load. Row-level change monitoring is registered with the database through this connection, covering three core configuration tables: the device configuration table, the monitoring template table, and the polling plan table. During registration, the monitored operation types are explicitly declared as insert, update, and delete. After any of these operations occur in any row of the database, the change event will be proactively pushed to this monitoring connection without the collection side initiating a query. This push model reduces the latency between configuration changes being written to the database and taking effect in the in-memory task cache to milliseconds, significantly better than polling-based synchronous solutions.

[0129] Upon receiving a change notification from the database, the event payload is first parsed in a structured manner. The event payload typically carries an operation type field, a changed table name field, and the primary key value of the affected rows in a serialized format. Parsing the operation type is used for subsequent branching, parsing the changed table name determines whether the change affects the device, template, or plan dimension, and parsing the primary key value precisely locates the set of entries in the in-memory task cache that need to be changed. To avoid performing a linear scan of the entire in-memory task cache every time a notification is received, the system pre-maintains a mapping table from primary keys to cache keys. This mapping table is written synchronously when a task entry is inserted into the in-memory cache and deregistered synchronously when an entry is deleted. Using this mapping table, one or more corresponding cache entries can be located from the primary key value in constant time, significantly reducing the time complexity of change processing, especially when the in-memory task cache is large.

[0130] When the operation type is insert, the change notification only carries the primary key value of the new row and does not include the complete business fields. Therefore, a readback request needs to be initiated to the database based on this primary key value to obtain the complete device information, monitoring indicator set, or polling plan record, and a new collection task entry object is constructed based on this. During the construction process, the data storage path is pre-calculated synchronously, and the routing hints are embedded in the task object so that the storage path does not need to be deduced again during subsequent collection execution. After the construction is completed, the new entry is atomically written to the in-memory task cache, and the corresponding relationship is registered in the primary key to cache key mapping table to ensure that subsequent update or deletion notifications for this primary key can be correctly located. The entire insert process is transparent to other running collection tasks and will not cause global lock contention.

[0131] When the operation type is update, the processing logic needs to consider concurrency safety issues. The complete record after the change is read back based on the primary key value, and then a field-level replacement is attempted on the corresponding entry in the memory task cache. The advantage of field-level replacement instead of overall replacement is that only the actually changed fields are updated, reducing unnecessary memory write operations and lowering the probability of read-write conflicts with the collection thread. However, if the entry is currently held by a collection thread—for example, in the execution path of SNMP message transmission / reception or result parsing—it cannot be modified directly in place, otherwise it may cause the collection thread to read inconsistent data in an intermediate state. In this case, a copy-on-write strategy is adopted: the new version after the change is temporarily stored in a slot to be replaced. After the collection thread completes its current task and releases its holding reference to the entry, the release operation triggers a slot check. If the version to be replaced exists in the slot, the final replacement is completed and the slot is cleared. The copy-on-write strategy ensures that the collection thread sees a consistent task view throughout the entire execution cycle, and also prevents the processing flow of change notifications from being blocked due to waiting for thread release.

[0132] When the operation type is deletion, after locating the corresponding cached entry through the primary key mapping table, it is removed from the memory task cache, and the corresponding relationship in the primary key mapping is simultaneously unregistered. The removal operation also needs to consider concurrency scenarios: if the entry to be deleted is currently held by the collection thread, it is marked as pending deletion. When the collection thread completes the current round of collection and returns the ownership of the entry, it detects the pending deletion mark and automatically removes the entry without including it in the next round of scheduling. This delayed deletion mechanism avoids forcibly interrupting the ongoing collection operation, ensuring the integrity of the single collection result, and also prevents the configuration of deleted devices from remaining in memory for a long time.

[0133] Readback operations carry a risk of timeout during network fluctuations or high database loads, and may also result in incomplete data returns due to transaction isolation levels. To address this, a conservative strategy is adopted: discard the readback result, retain the known complete state of the entry in the in-memory task cache, and do not modify any fields. Simultaneously, an alarm notification is generated, recording the table name, primary key value, reason for failure, and time of occurrence, for maintenance personnel to investigate. The alarm notification is exposed via a telemetry interface and can be subscribed to and consumed by external monitoring platforms. Retaining the known complete state means that the entry will continue to participate in scheduling according to the old configuration. Although there may be temporary differences from the latest configuration in the database, this is less harmful than building an incorrect task entry with incomplete data. After an alarm is generated, the system can be configured with a retry mechanism to attempt to perform a readback on the primary key again after a certain delay. If the retry succeeds, normal insert, delete, update, and delete operations are performed. If multiple retries fail, the alarm status is maintained pending manual intervention. The entire fault handling process is strictly isolated within the affected entry; the scheduling and execution of other entries are completely unaffected, ensuring the overall availability of the monitoring system in the event of local configuration synchronization anomalies.

[0134] The publish / subscribe connection itself also needs to have reconnection capabilities. When the listening connection is broken due to network interruption or database restart, it needs to automatically initiate a reconnection and re-register row-level change listeners. After a successful reconnection, to compensate for any missed change events during the disconnection, a full reconciliation needs to be performed: the current state of the three configuration tables in the database is compared with the current state of the in-memory task cache, and compensatory insert, delete, and update operations are performed on the discrepancies to ensure that the in-memory task cache and the database are eventually consistent. The reconciliation process also follows the copy-on-write and delayed deletion rules, so it does not affect ongoing data collection tasks.

[0135] Before executing the task, read the device status in the shared reachability map using atomic operations to determine whether the probe point has been reached, including:

[0136] A shared reachability mapping table is maintained in memory, with the key being a unique device identifier and the value being a reachability state structure. The reachability state structure contains a reachability Boolean flag, the first unreachable timestamp, and the most recent probe timestamp.

[0137] Before starting to process a single device task, the polling scheduler reads the current value of the reachability Boolean flag in the reachability state structure corresponding to the device using an atomic load instruction;

[0138] If the flag is true, the collection polling interaction will be executed directly according to the pre-set frequency of the task entry. After receiving a valid response, the reachability Boolean flag will be kept true by atomic writing.

[0139] If the flag is false, the current timestamp is obtained, the time interval between the current time and the most recent detection timestamp is calculated, and the value is compared with the preset frequency reduction interval threshold parameter in the task entry.

[0140] If the time interval is less than the frequency reduction interval threshold parameter, the current round of scheduling is abandoned, the current worker thread slot and the allocated network receive buffer are released, and the next device task is processed.

[0141] If the time interval is not less than the frequency reduction interval threshold parameter, a lightweight probe request is sent to the device by borrowing a UDP session from the session pool. The result is determined based on whether a valid response is received within the preset timeout threshold, and the reachability status is updated accordingly.

[0142] A shared reachability mapping table is constructed in memory, using the device's unique identifier as the key and reachability state structures as values. This mapping table is initialized along with the task cache during process startup, and all worker threads access it through the same memory address space, eliminating the need for cross-process or cross-node communication. The reachability state structure contains three fields: a reachability boolean flag to record whether the device is currently reachable; a first unreachable timestamp to record the moment the device was first determined to be unreachable, used as a reference for upper-layer alarms or timeout circuit breaker logic; and a most recent probe timestamp to record the moment the last active probe was completed, used to control the frequency of probe reduction. These three fields together describe the complete historical state of a device in the reachability dimension, allowing scheduling decisions to be made solely from memory reads without initiating any network interactions.

[0143] Before processing a single device task, the polling scheduler uses an atomic load instruction to read the current value of the reachability boolean flag in the reachability state structure corresponding to that device. The atomic load instruction guarantees the indivisibility of the read operation; even if other worker threads are concurrently performing write operations on the same field, the read result will not tearing, thus avoiding reliance on mutex locks. The core value of this design is that reachability checks are on a hot path; each device task triggers a read in each scheduling cycle, and any additional lock contention will directly affect overall throughput. By using lock-free atomic instructions to compress the read overhead to the level of a single CPU instruction, the scheduler's throughput in high-density device scenarios is not hampered by reachability checks.

[0144] When the reachability boolean flag is true, it indicates that the device is in a normally reachable state, and the collection polling interaction is executed directly according to the pre-set frequency in the task entry. After the collection interaction is completed, if a valid response is received from the target device, an atomic write instruction is used to keep the reachability boolean flag true. The expression "keep" rather than "update" here means that even if the flag is already true, an atomic write is still performed. The purpose is to ensure that the memory visibility semantics of the write operation are satisfied in a multi-core environment, preventing compiler or processor out-of-order optimizations from causing other threads to observe stale values. The processing logic under the normal path is simple and does not involve any additional timestamp reads and writes, minimizing the computational overhead of the hot path.

[0145] When the reachability Boolean flag read is false, proceed to the down-frequency detection branch. Obtain the current timestamp, denoted as... Simultaneously, the most recent probe timestamp is read from the reachability state structure and denoted as... The time interval is obtained by calculating the difference between the two. .Will With respect to the preset frequency reduction interval threshold parameter in the task entry Numerical comparisons are performed to determine whether a probe needs to be initiated in this round of scheduling. Frequency reduction interval threshold parameter. During the task entry loading phase, task objects are read from the database and embedded. Different devices or different business scenarios can be configured with different... This allows for fine-grained control of the detection frequency. For example, a smaller value can be set for core backbone equipment. This allows it to continue probing at a high frequency after becoming unreachable; for edge access devices, a larger [configuration value] can be set. This avoids a network storm caused by a large number of unreachable devices launching simultaneous probes.

[0146] like If the target device has not yet been reached, the current round of scheduling is abandoned. The abandonment operation involves two resource release actions: releasing the current worker thread slot so that it can be immediately occupied by other pending tasks; and releasing the allocated network receive buffer, returning the buffer to the buffer pool for reuse by subsequent tasks. These two release operations ensure that the decision to skip the probe does not cause any resource leaks, and the scheduler then continues processing the next device task in the queue. The overall scheduling loop does not become blocked or idle due to the existence of unreachable devices.

[0147] like If the probe point is reached, a lightweight probe request needs to be initiated. The probe request is carried out by borrowing a UDP session from the session pool. The borrowing operation itself follows a guard object mechanism: a guard object is created during borrowing and marked as failed by default, ensuring that even if the subsequent process is abnormally interrupted, the session return or destruction logic can be correctly handled when the guard is destroyed. Lightweight probe requests differ from full data collection interactions in that their message size is small, round-trip latency is short, and they are only used to verify the network reachability of the target device. They do not carry a complete SNMP OID list or other large query parameters, thus minimizing the impact of the probe on network bandwidth and the target device's CPU.

[0148] After the probe is sent, within the preset timeout threshold Waiting for a response. If... If a valid response is received, the device is determined to be reachable again: the reachability Boolean flag is updated to true using an atomic write instruction, and the most recent probe timestamp is also updated. For the current moment, the first unreachable timestamp field retains its original value for subsequent statistical analysis. If in If no valid response is received, the device is determined to remain in an unreachable state: An atomic write instruction is used to maintain the reachability Boolean flag at a false value, and the most recent probe timestamp is updated. Updated to the time when this probe was initiated, so that the next scheduling cycle can be recalculated based on the correct reference time. To avoid the detection interval from shrinking uncontrollably to a value much smaller than the specified value due to outdated timestamps. The degree of.

[0149] Throughout the reachability determination and detection process, all reads and writes to the reachability Boolean flag are completed using atomic instructions. While write operations to the timestamp field do not require strict atomicity (a single worker thread holds write privileges), the scheduler's task sharding mechanism ensures that only one worker thread is processing the same device at any given time, thus logically eliminating concurrent write conflicts to the timestamp field. This design, combining atomicity guarantees with logical mutual exclusion, maintains high performance while avoiding excessive use of heavyweight synchronization primitives, enabling the entire acquisition and scheduling framework to maintain stable low latency even when facing concurrent polling from thousands of devices.

[0150] Based on the detection results, the reachability flag is updated, and the detection frequency is controlled according to the frequency reduction interval threshold, including:

[0151] When the detection is successful, the reachability Boolean flag in the reachability state structure is set to true by atomic comparison and swap operation, the first unreachable timestamp field value is cleared and the most recent detection timestamp is updated to the current time. At the same time, the frequency reduction interval threshold parameter is reset to the preset initial frequency reduction interval threshold so that the device can perform acquisition interaction at the original planned frequency in subsequent scheduling.

[0152] When a detection fails, the reachability Boolean flag is kept false by comparing and swapping atomic operations, the most recent detection timestamp atom is updated to the current time, and the current frequency reduction interval threshold parameter is kept unchanged for use in the next scheduling cycle.

[0153] If the current down-frequency interval threshold is the initial value, after the first detection fails, the down-frequency interval threshold will be expanded to the second-level down-frequency interval threshold by the backoff multiplier factor. After each subsequent detection failure, the backoff multiplier factor will be multiplied to obtain the next-level down-frequency interval threshold. If the product exceeds the preset upper limit value, the down-frequency interval threshold will be clamped to the upper limit value.

[0154] If the device is marked as unreachable again due to continuous failures after a successful detection, it will re-execute the step-by-step backoff growth strategy starting from the initial frequency reduction interval threshold until the upper limit of the frequency reduction interval threshold is reached, so as to ensure that the detection frequency is adaptively adjusted in the network environment where the device is repeatedly connected and disconnected.

[0155] In the atomic update mechanism of reachability state, the probe result directly determines how the fields of the reachability state structure are updated. The reachability state structure contains three core fields: a reachability boolean flag, a timestamp of the first unreachable event, and a timestamp of the most recent probe. It also maintains a currently valid frequency reduction interval threshold parameter. All modifications to these fields are performed through atomic compare-and-swap operations to avoid data races or state fragmentation issues during multi-threaded concurrent scheduling.

[0156] Upon successful detection, an atomic compare-and-swap operation sets the reachability Boolean flag from its current value (whether true or false) to true. If the device was previously unreachable, the first unreachable timestamp field is simultaneously cleared and reset to null or zero after successful flag setting, indicating that the device has regained reachability and the duration of unreachability has been reset to zero. The most recent probe timestamp field is updated to the current time. This records the time reference for this successful detection, which will be used to calculate the time interval in subsequent scheduling cycles. Simultaneously, the frequency reduction interval threshold parameter... Reset to the preset initial down-frequency interval threshold This design ensures that the device resumes data collection at the originally planned frequency during subsequent scheduling, completely resetting the backoff strategy. This design guarantees that once normal communication is restored, the scheduler can immediately collect data from the device at the highest frequency, avoiding data collection delays caused by residual historical backoff states.

[0157] When a probe fails, the reachability Boolean flag is kept false by an atomic compare-and-swap operation and is not flipped. The most recent probe timestamp field is atomically updated to the current time. This is so that the next scheduling cycle can calculate the time interval based on the latest detection time when determining whether the detection point has been reached. ,in This is the most recent probe timestamp after this update. Current down-frequency interval threshold parameter. The time interval remains unchanged and is available for use in the next scheduling cycle. This means that after a probe fails, the scheduler will determine whether the time interval is satisfied in the next polling cycle. In this case, the threshold used is consistent with that used before the current detection, maintaining continuity.

[0158] The backoff growth strategy for the frequency reduction interval threshold employs an exponential backoff mechanism, using a backoff multiplier factor. The detection interval is gradually increased based on the current frequency reduction interval threshold. Equal to the initial down-frequency interval threshold Then, after the first detection fails, Updated to the second-level down-frequency interval threshold After each subsequent detection failure, the current down-frequency interval threshold is multiplied by the backoff multiplier factor to obtain the next level of down-frequency interval threshold. ,in This indicates the current backoff level index. If the calculated product exceeds a preset upper limit... Then the down-frequency interval threshold will be clamped to That is, the actual effective down-frequency interval threshold is The clamping operation ensures that the detection interval does not expand indefinitely, and can continue to detect at a reasonable minimum frequency even in scenarios where the network is unreachable for a long time, thus preventing the device from being undetected for a long time after it is restored.

[0159] retreat multiplier factor A typical value for this is 2, meaning the detection interval doubles after each detection failure, consistent with the standard exponential backoff algorithm. Initial down-frequency interval threshold. It can be configured according to the actual network environment, for example, set to a certain multiple of the normal collection cycle. Upper limit value. This setting is based on the business's tolerance for device recovery latency, typically not exceeding a few minutes, to balance the conflict between detection overhead and recovery response speed. Backoff Level Index Explicit storage is not required because the current down-frequency interval threshold is already set. The system itself already implicitly contains information about the level at which the retreat occurs; each failure directly affects... Simply perform the multiplication and clamp the bit.

[0160] In network environments with repeated device on / off cycles, this adaptive backoff strategy can effectively handle jitter scenarios. When a device is successfully detected, the frequency reduction interval threshold is set. Reset to If the device is subsequently marked as unreachable again due to continuous data acquisition failures, the backoff process will begin from [the point where the device is marked as unreachable]. Restart, press The backoff history increases sequentially until it reaches an upper limit. This means that each time a device recovers and is successfully detected, the backoff history is completely erased, and if it loses contact again, it will go through the entire backoff growth process again. This "success resets to zero, failure accumulates" strategy can quickly respond to recovery events when devices frequently disconnect, while gradually reducing the detection frequency to save network and computing resources when the device remains unreachable.

[0161] At the implementation level, the frequency reduction interval threshold parameter in the reachability state structure Use atomic integer or atomic floating-point data types to ensure thread safety for read and write operations. Backoff multiplier factor. Initial down-frequency interval threshold and upper limit Configuration parameters are loaded from the configuration file or database at system startup and remain unchanged during runtime, allowing for safe reading without locking. Atomic update operations of the probe results and updates of the frequency reduction interval threshold are executed sequentially within the same logical path, avoiding inconsistencies between the two. When multiple acquisition threads concurrently probe the same device, atomic comparison and swap operations ensure that only the first thread to complete the probe successfully updates the reachability boolean flag; subsequent threads' update operations will retry or abort due to comparison failures, thus guaranteeing the consistency of the state structure.

[0162] At the telemetry interface level, the frequency reduction interval threshold The current value and the change in backoff level can be exposed as observable indicators, allowing maintenance personnel to monitor the detection frequency distribution of each device in real time and assist in judging network quality and device stability.

[0163] When borrowing a UDP session, a guard object is created and marked as failed by default. Guards are categorized and evicted upon destruction, including:

[0164] The idle session handle is obtained from the UDP session connection pool in a mutually exclusive manner. When constructing the guard object, the generation counter value of the current session is read in atomic operation and recorded in the borrow generation number field of the guard object. The internal status flag field is initially set to the failure value. The guard object takes over the life cycle management of the session handle.

[0165] The caller uses the borrowed session to send a data acquisition request packet to the target device and enters a blocking wait. After receiving the data packet within the timeout threshold, it performs an integrity check on the returned data packet. If the check passes, the internal status flag is set to a success value. If a timeout, check failure, or error status code occurs, the failure value is retained without modification.

[0166] When the guard object is destructed, the current generation value of the session is first read atomically and compared with the borrowed generation number. If the two are not equal, it means that the session has been rebuilt and borrowed to another caller during the borrowing period, and the current eviction logic is abandoned.

[0167] If the generation number is equal and the status is successful, the session is returned to the connection pool and the consecutive success count is incremented. If the generation number is equal and the error type is timeout, the socket is closed, a new session is added to the connection pool, and the consecutive success count is reset. If the generation number is equal and the error type is socket-level error, the socket is closed, all resources are released, and the session slot is removed from the connection pool.

[0168] In network monitoring and data collection scenarios, the borrowing and returning of UDP sessions involves two core issues: concurrency security and resource lifecycle management. To address this, a Session Guard object is introduced as a lifecycle proxy for session handles. The session state marking and eviction logic are encapsulated in the construction and destruction process of the Session Guard object, thereby ensuring that session resources are properly handled regardless of how the data collection process ends.

[0169] When borrowing a session from the UDP session connection pool, the idle queue of the connection pool is accessed in a mutually exclusive manner to retrieve an idle session handle. This mutual exclusion ensures that only one caller can complete the borrowing operation at a time, preventing multiple data collection tasks from concurrently holding the same session handle. After obtaining the handle, the current value of the generation counter field maintained on the session object is immediately read using an atomic load operation, and this value is written to the borrowing generation number field of the guard object. The generation counter is a monotonically increasing integer value, atomically incremented by the reconstruction logic each time a session is closed and rebuilt; its purpose is to assign a unique identifier to each round of session lifecycle. After the guard object is constructed, its internal status flag field is initialized to a failure value. At this point, the guard object officially takes over the lifecycle management of the session handle; the caller no longer directly holds a handle reference, and all subsequent operations are completed through the access interface provided by the guard object.

[0170] The caller sends a data acquisition request packet to the target device through the session handle held by the guard object, and then enters a blocking wait state, waiting for a response from the peer or a timeout event. During the wait, the underlying socket's receive buffer continuously listens for data packets from the target device's IP and port. If a data packet is received within the timeout threshold, an integrity check is performed on the returned data packet. The check includes whether the data packet length meets expectations, whether the protocol version and request identifier in the packet header are consistent with the current request, and whether the checksum field passes verification. After all checks pass, the guard object's internal status flag field is updated from a failure value to a success value. If no response is received after the timeout threshold expires during the wait, or a data packet is received but the integrity check fails, or the received data packet carries an error status code, no modification is made to the status flag field, and its initial failure value is retained. This design makes the semantics of the status flag clear: a success value is only written after a valid response is clearly received, and any abnormal path is covered by a failure value, eliminating the need to perform a separate status write-back operation in each abnormal branch.

[0171] When a guard object is destroyed, the current value of the generation counter on the session handle is first read again using an atomic load operation and compared with the borrowing generation number recorded during the guard object's construction. If the two are not equal, it means that during this borrowing period, the session has been closed and rebuilt by another path (such as concurrent timeout reconstruction logic). The generation counter has been incremented during reconstruction, and the handle reference held by the current guard object no longer points to a valid current session. Continuing to execute the eviction logic will result in an erroneous operation on resources that have been released or borrowed elsewhere. Therefore, if the generation numbers are not equal, the eviction logic is abandoned directly, the guard object completes its destruction and exits, and no write operation is performed on the connection pool. This generation verification mechanism solves the race condition problem of "session being replaced during borrowing" in concurrent scenarios with extremely low overhead, without holding a global lock throughout the entire acquisition process.

[0172] When the generation number matches, different eviction branches are entered based on the value of the internal status flag field and the error type field of the guarded object. When the status flag is a success value, the session handle is returned to the idle queue of the connection pool, and the consecutive success count field maintained on the session object is incremented. The consecutive success count is used by the upper-layer health assessment logic to determine whether the session is in a stable and available state. When the number of consecutive successes accumulates to a preset threshold, strategies such as connection pool expansion or priority promotion can be triggered.

[0173] When the generation counters are equal and the status flag is a failure value with a timeout record in the error type field, the timeout reconstruction process is executed: The underlying UDP socket held by the current session is closed, releasing the file descriptor resources it occupies. Then, a new UDP socket is created with the same destination address and port parameters, constructing a new session object. The generation counter of the new session is atomically initialized to the old session's generation value plus one. The new session object is then added to the connection pool's idle queue, and the consecutive success count of the new session is reset to zero. Timeout scenarios typically indicate that the network path is temporarily unavailable or the peer's response is delayed. Reconstructing the socket can clear any residual error states from the old socket, providing a clean transmission endpoint for subsequent data collection.

[0174] When the intercalary intercalary code matches and the error type field records a socket-level error, it indicates that an unrecoverable error has occurred on the underlying socket. This could be due to an ICMP port unreachable message causing the socket to enter an error state, or a fatal error code being returned by the operating system. In this case, the socket is closed, all resources occupied by the session object are released, and the session entry is removed from the connection pool's slot index. No new sessions are added to the connection pool. The removal of connection pool slots is also performed in a mutually exclusive manner to ensure that the slot count matches the actual number of available sessions. The upper-layer connection pool management logic can asynchronously trigger a session replenishment process when it detects that the number of slots is below the minimum threshold, adding new sessions to the connection pool as needed.

[0175] The core value of the entire guard object mechanism lies in combining the "snapshot when borrowing, snapshot verified when returning" model with the "default failure, explicit success" status marking model. This allows each execution path of the collection logic—whether it completes normally, times out, fails verification, or has a socket error—to automatically trigger the correct resource disposal logic when the guard object is destroyed. This eliminates the need for the caller to manually handle session return or destruction at each return point, fundamentally eliminating the risk of resource leakage and state inconsistency.

[0176] When a new session is added to the connection pool, a new generation counter value is assigned to prevent eviction operations from accidentally affecting newly borrowed sessions, including:

[0177] The connection pool maintains a global monotonically increasing generation number allocation counter. Each time a new underlying session object is created, the new generation counter value is obtained by atomic increment operation and written to the generation counter field of the session object. For session objects with the same device identifier, strictly increasing generation counter values ​​are allocated sequentially according to the creation time to distinguish different generation session objects.

[0178] Before executing the eviction decision, the guard object destruction logic reads the current generation counter value of the underlying session object again with an atomic load instruction with full memory barrier semantics, and compares it with the borrow generation number field value recorded during guard construction. The full memory barrier ensures that all previous release write operations of other threads are visible to this read.

[0179] If the current generation value is equal to the borrowed generation number, it indicates that the session has not been destroyed and rebuilt from the time of borrowing to the time of destruction, and the guard's expulsion decision can be safely applied to the session object.

[0180] If the current generation value is not equal to the borrowed generation number, it indicates that the session has been destroyed and replaced by a new session during the borrowing period, and the new session has been borrowed for use. In this case, the eviction should be abandoned to prevent the invalid destruction or incorrect return of a valid new session that is currently in use, and to prevent delayed timeout evictions from causing accidental damage to new sessions that are in normal use.

[0181] One of the core challenges facing connection pools in managing the lifecycle of underlying session objects is ensuring the accuracy of eviction decisions amidst the intertwined execution of multi-threaded concurrent borrowing, returning, and destruction operations. When a session is destroyed and rebuilt due to a timeout or fatal error, the newly created session object may be immediately borrowed by another thread. If the guard object that originally held the session reference delays triggering its destruction logic, its eviction decision will apply to a completely different session object than when it was borrowed, leading to false positives. The generational counter mechanism is designed to solve this problem. By assigning a globally unique and monotonically increasing generational identifier to each new session, it enables guard objects to reliably determine during destruction whether the session they hold is still the same instance as when it was borrowed.

[0182] Internally, the connection pool maintains a global generation number allocation counter. This counter is initialized to a specific starting value (e.g., 1) and monotonically increases throughout the connection pool's lifecycle, without wrapping around or resetting. Whenever a new underlying session object needs to be created, a new value is retrieved from this global counter using an atomic increment operation (fetch-and-add semantics), and written to the generation counter field within the new session object. Because the atomic increment operation guarantees that each retrieved value is strictly greater than the previous one, for multiple session objects created sequentially under the same device identifier, their generation counter values ​​strictly increment in the order of creation, ensuring that the generation values ​​of any two different session instances will not be the same. This characteristic is the foundation for subsequent eviction decisions to distinguish between the "original session" and the "new session."

[0183] When a guard object is constructed, in addition to recording the reference to the borrowed session object, it also needs to synchronously record the current generation counter value of the session object and store it in the guard's own borrowed generation number field. This recording action occurs when the session is successfully borrowed and the guard object completes initialization, at which point the generation value of the session is deterministic and stable. The borrowed generation number field remains unchanged throughout the guard's entire lifecycle, serving as a "generation snapshot at the time of borrowing" for use by the destruction logic. The guard object also marks the session state as failed by default, updating the state to success only after the collection task is successfully completed. This default failure semantics ensures that even if the guard is destroyed prematurely under an abnormal path, an incomplete collection session will not be incorrectly returned to the available pool.

[0184] Before executing an eviction decision, the destruction logic of the guarded object needs to reread the current generation counter value of the underlying session object using an atomic load instruction with full memory barrier semantics. Full memory barrier semantics are chosen over the weaker, more lenient load semantics because it's necessary to ensure that all write operations performed on the session object by other threads (including destroying the old session, writing the new session generation value, and other release writes) are visible to the current thread before this read operation is performed. If a lenient atomic load were used, processor or compiler instruction reordering might lead to reading an outdated generation value, resulting in an incorrect eviction decision. Full memory barrier semantics are a necessary guarantee of correctness here, not an optional performance optimization.

[0185] Comparing the current generation value obtained from the atomic load with the borrowed generation number recorded during the guard's construction can result in two scenarios. If they are equal, it means the generation value of the underlying session object has never changed from the time of borrowing to the time of destruction; that is, the session was not destroyed or recreated during the entire borrowing period. In this case, the session reference held by the guard still points to the same instance at the time of borrowing, and eviction can be safely applied: if the collection is successful during borrowing, the session is returned to the availability pool and the health count is reset; if a timeout occurs, the session is closed and a new session is recreated and added to the pool; if a fatal error occurs, the session is destroyed and removed from the pool without being recreated.

[0186] When the two are not equal, it indicates that the underlying session object has undergone at least one destruction and reconstruction during the borrowing period. Specifically, if the current generation value is strictly greater than the borrowed generation number, it means that the connection pool has allocated and written a higher generation value for the device identifier, i.e., a new session instance already exists. Furthermore, since this new session has been borrowed and is being used by another thread, if the guard still performs the eviction operation according to the original judgment logic (e.g., misjudging the new session as timed out and closing it, or misjudging it as successful and returning it early), it will directly disrupt the ongoing data acquisition task, leading to data loss or inconsistent session states. Therefore, when the generation values ​​do not match, the guard's destruction logic abandons this eviction operation, does not perform any modifications to the underlying session object, and directly completes the destruction. This abandonment operation is safe because the original session has already completed resource reclamation in the previous destruction process, and the lifecycle of the new session is managed by the guard object that borrowed it, so there is no risk of resource leakage.

[0187] The generation counter mechanism can also effectively prevent a special type of race condition scenario: a data acquisition task experiences extremely long response delays due to network jitter, and the destruction of the guarded object is postponed until the session has timed out and been destroyed, and a new session is created and borrowed. In implementations without generation verification, the delayed eviction decision will directly affect the new session, causing the interruption of normal data acquisition tasks. With the introduction of the generation counter, the guards of delayed destruction can detect changes in the generation value during the verification phase, thus safely skipping eviction and returning the complete lifecycle of the new session to its rightful holder.

[0188] In practical deployments, the global generation number allocation counter typically uses a 64-bit unsigned integer data type. Under normal network monitoring and data acquisition scenarios, even with a very high frequency of continuously newly established sessions, this counter will not overflow within the expected device operating cycle, thus ensuring that the assumption of global uniqueness of generation values ​​holds true in engineering practice. Both the generation counter field and the borrowed generation number field adopt an aligned memory layout compatible with atomic operations, ensuring that read and write operations meet atomicity requirements at the hardware level without introducing additional lock overhead. The entire generation verification process adds only one atomic load and one integer comparison operation to the hot path of guarded destruction, with negligible impact on acquisition performance, while fundamentally eliminating the security risk of accidental expulsion in multi-threaded concurrent scenarios.

[0189] A second aspect of the present invention provides a network monitoring data acquisition system, comprising:

[0190] The scheduling unit is used to load device information, monitoring indicator set and polling plan from the database to the memory task cache when the system starts up, and pre-calculate the data storage path for each task and embed it into the corresponding task object as a routing prompt. During operation, the polling scheduler reads the task-driven collection from the memory task cache, subscribes to the database incremental change notification, and performs add, delete, modify and update operations on the affected entries.

[0191] The reachable unit is used to read the device status in the shared reachability map with atomic operations before executing the task. If it is normal, it will be executed at the planned frequency. If it is unreachable, it will determine whether the probe point has been reached based on the frequency reduction interval. If it is reached, the probe will be executed and the status will be updated atomically according to the result. If it is not reached, it will be skipped.

[0192] Session unit is used to create guard objects when borrowing a UDP session and mark them as failed by default. After collection is completed, it is marked as successful or retained as failed based on the result. When the guard is destroyed, it is classified and expelled: if successful, it is returned to the pool and the health count is reset; if it times out, it is closed, rebuilt and added to the pool; if it is a fatal error, it is destroyed and removed.

[0193] The verification unit is used to verify that the session generation counter matches the borrowing record before execution during expulsion.

[0194] The data acquisition unit is used to obtain the raw counter value from the device during data acquisition, store it in the raw data table according to the routing prompts in the task object, and dynamically calculate the derived indicators based on the raw data by the query layer view. When the data is refreshed, the cumulative count and time are updated lock-free through atomic comparison and exchange operations, generating refresh evidence records and exposing the telemetry interface.

[0195] A third aspect of the present invention provides an electronic device, comprising:

[0196] processor;

[0197] Memory used to store processor-executable instructions;

[0198] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0199] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0200] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for collecting network monitoring data, characterized in that, include: When the system starts, it loads device information, monitoring metrics set and polling plan from the database into the memory task cache, and pre-calculates the data storage path for each task and embeds it into the corresponding task object as a routing prompt. During operation, the polling scheduler reads task-driven data collection from the memory task cache, subscribes to database incremental change notifications, and performs add, delete, modify and update operations on the affected entries. Before executing a task, read the device status in the shared reachability map using atomic operations. If the status is normal, execute at the planned frequency. If the status is unreachable, determine whether the probe point has been reached based on the frequency reduction interval. If the probe is reached, execute the probe and update the status atomically based on the result. If the status is not reached, skip the task. When borrowing a UDP session, a guard object is created and marked as failed by default. After collection is completed, it is marked as successful or retained as failed based on the result. When the guard is destroyed, it is categorized and evicted: if successful, it is returned to the pool and the health count is reset; if it times out, it is closed, rebuilt and added to the pool; if it is a fatal error, it is destroyed and removed. When evicting, the session generation counter is checked to be consistent with the borrowing record before execution. During data collection, the raw counter values ​​are obtained from the device and stored in the raw data table according to the routing prompts in the task object. The derived indicators are dynamically calculated by the query layer view based on the raw data. When the data is refreshed, the cumulative count and time are updated lock-free through atomic comparison and exchange operations, generating refresh evidence records and exposing the telemetry interface.

2. The method according to claim 1, characterized in that, The system loads device information, monitoring metrics sets, and polling plans from the database into the in-memory task cache, and pre-calculates the data storage path for each task, embedding it into the corresponding task object as a routing hint, including: Read device records from the database containing the device's unique identifier, network address, transport layer protocol type, and authentication credentials, and construct a device object hash table using the device's unique identifier as the key. Read monitoring indicator definition records from the database, which contain indicator names, identifiers of objects to be collected, and numerical types. Establish an association index structure from the device object to the indicator definition set using the device model identifier as the key. Read the polling plan record containing device identifier, polling period and retry parameters from the database, and establish a correspondence between device identifier and device object; The device identifiers in the device object diagram and the corresponding monitoring indicator definition set for the device model are expanded by Cartesian according to the indicator dimension to generate one-to-one corresponding collection task entries. Each entry encapsulates the device identifier, network address, protocol type, single object identifier to be collected, and polling parameters. Based on the device identifier of the data collection task entry, the storage configuration metadata is queried to read the target table name and partition information. The target table name prefix, device identifier hash value and time partition sharding key are concatenated to form a complete storage path string. The complete storage path is bound to the corresponding data collection task entry as a routing hint field, so that the runtime refresh component can directly complete the storage location based on the routing hint without performing path parsing query.

3. The method according to claim 2, characterized in that, Subscribe to database incremental change notifications and perform add, delete, modify, and update operations on affected entries, including: Establish a change monitoring connection with the database based on a publish-subscribe mechanism, and register row-level change monitoring for the device configuration table, monitoring template table, and polling plan table with the database, covering insert, update, and delete operation types; After receiving the change notification pushed by the database, the operation type, the changed table name and the affected primary key value are parsed from the event payload. The corresponding set of entries in the memory task cache is located by mapping the pre-maintained primary key to the cache key. If the operation type is insert, the complete record is read back from the database based on the primary key value to construct a new collection task entry, which is then inserted into the memory task cache and the primary key mapping is updated. If the operation type is update, the changed record is read back from the database based on the primary key value, and the corresponding entry in the memory task cache is replaced and updated at the field level. If the entry is being held by the collection thread, the new version is temporarily stored using the copy-on-write strategy and the replacement is completed after the thread is released. If the operation type is deletion, the corresponding entry is located through the primary key mapping, the entry is removed from the memory task cache, and the primary key mapping is unregistered; If a readback operation times out or returns incomplete data, the known complete state of that entry in the memory task cache is preserved and an alarm notification is generated. The normal scheduling and execution of other entries are not affected.

4. The method according to claim 1, characterized in that, Before executing the task, read the device status in the shared reachability map using atomic operations to determine whether the probe point has been reached, including: A shared reachability mapping table is maintained in memory, with the key being a unique device identifier and the value being a reachability state structure. The reachability state structure contains a reachability Boolean flag, the first unreachable timestamp, and the most recent probe timestamp. Before starting to process a single device task, the polling scheduler reads the current value of the reachability Boolean flag in the reachability state structure corresponding to the device using an atomic load instruction; If the flag is true, the collection polling interaction will be executed directly according to the pre-set frequency of the task entry. After receiving a valid response, the reachability Boolean flag will be kept true by atomic writing. If the flag is false, the current timestamp is obtained, the time interval between the current time and the most recent detection timestamp is calculated, and the value is compared with the preset frequency reduction interval threshold parameter in the task entry. If the time interval is less than the frequency reduction interval threshold parameter, the current round of scheduling is abandoned, the current worker thread slot and the allocated network receive buffer are released, and the next device task is processed. If the time interval is not less than the frequency reduction interval threshold parameter, a lightweight probe request is sent to the device by borrowing a UDP session from the session pool. The result is determined based on whether a valid response is received within the preset timeout threshold, and the reachability status is updated accordingly.

5. The method according to claim 4, characterized in that, Based on the detection results, the reachability flag is updated, and the detection frequency is controlled according to the frequency reduction interval threshold, including: When the detection is successful, the reachability Boolean flag in the reachability state structure is set to true by atomic comparison and swap operation, the first unreachable timestamp field value is cleared and the most recent detection timestamp is updated to the current time. At the same time, the frequency reduction interval threshold parameter is reset to the preset initial frequency reduction interval threshold so that the device can perform acquisition interaction at the original planned frequency in subsequent scheduling. When a detection fails, the reachability Boolean flag is kept false by comparing and swapping atomic operations, the most recent detection timestamp atom is updated to the current time, and the current frequency reduction interval threshold parameter is kept unchanged for use in the next scheduling cycle. If the current down-frequency interval threshold is the initial value, after the first detection fails, the down-frequency interval threshold will be expanded to the second-level down-frequency interval threshold by the backoff multiplier factor. After each subsequent detection failure, the backoff multiplier factor will be multiplied to obtain the next-level down-frequency interval threshold. If the product exceeds the preset upper limit value, the down-frequency interval threshold will be clamped to the upper limit value. If the device is marked as unreachable again due to continuous failures after a successful detection, it will re-execute the step-by-step backoff growth strategy starting from the initial frequency reduction interval threshold until the upper limit of the frequency reduction interval threshold is reached, so as to ensure that the detection frequency is adaptively adjusted in the network environment where the device is repeatedly connected and disconnected.

6. The method according to claim 1, characterized in that, When borrowing a UDP session, a guard object is created and marked as failed by default. Guards are categorized and evicted upon destruction, including: The idle session handle is obtained from the UDP session connection pool in a mutually exclusive manner. When constructing the guard object, the generation counter value of the current session is read in atomic operation and recorded in the borrow generation number field of the guard object. The internal status flag field is initially set to the failure value. The guard object takes over the life cycle management of the session handle. The caller uses the borrowed session to send a data acquisition request packet to the target device and enters a blocking wait. After receiving the data packet within the timeout threshold, it performs an integrity check on the returned data packet. If the check passes, the internal status flag is set to a success value. If a timeout, check failure, or error status code occurs, the failure value is retained without modification. When the guard object is destructed, the current generation value of the session is first read atomically and compared with the borrowed generation number. If the two are not equal, it means that the session has been rebuilt and borrowed to another caller during the borrowing period, and the current eviction logic is abandoned. If the generation number is equal and the status is successful, the session is returned to the connection pool and the consecutive success count is incremented. If the generation number is equal and the error type is timeout, the socket is closed, a new session is added to the connection pool, and the consecutive success count is reset. If the generation number is equal and the error type is socket-level error, the socket is closed, all resources are released, and the session slot is removed from the connection pool.

7. The method according to claim 6, characterized in that, When a new session is added to the connection pool, a new generation counter value is assigned to prevent eviction operations from accidentally affecting newly borrowed sessions, including: The connection pool maintains a global monotonically increasing generation number allocation counter. Each time a new underlying session object is created, the new generation counter value is obtained by atomic increment operation and written to the generation counter field of the session object. For session objects with the same device identifier, strictly increasing generation counter values ​​are allocated sequentially according to the creation time to distinguish different generation session objects. Before executing the eviction decision, the guard object destruction logic reads the current generation counter value of the underlying session object again with an atomic load instruction with full memory barrier semantics, and compares it with the borrow generation number field value recorded during guard construction. The full memory barrier ensures that all previous release write operations of other threads are visible to this read. If the current generation value is equal to the borrowed generation number, it indicates that the session has not been destroyed and rebuilt from the time of borrowing to the time of destruction, and the guard's expulsion decision can be safely applied to the session object. If the current generation value is not equal to the borrowed generation number, it indicates that the session has been destroyed and replaced by a new session during the borrowing period, and the new session has been borrowed for use. In this case, the eviction should be abandoned to prevent the invalid destruction or incorrect return of a valid new session that is currently in use, and to prevent delayed timeout evictions from causing accidental damage to new sessions that are in normal use.

8. A network monitoring data acquisition system, used to implement the method as described in any one of claims 1-7, characterized in that, include: The scheduling unit is used to load device information, monitoring indicator set and polling plan from the database to the memory task cache when the system starts up, and pre-calculate the data storage path for each task and embed it into the corresponding task object as a routing prompt. During operation, the polling scheduler reads the task-driven collection from the memory task cache, subscribes to the database incremental change notification, and performs add, delete, modify and update operations on the affected entries. The reachable unit is used to read the device status in the shared reachability map with atomic operations before executing the task. If it is normal, it will be executed at the planned frequency. If it is unreachable, it will determine whether the probe point has been reached based on the frequency reduction interval. If it is reached, the probe will be executed and the status will be updated atomically according to the result. If it is not reached, it will be skipped. Session unit is used to create guard objects when borrowing a UDP session and mark them as failed by default. After collection is completed, it is marked as successful or retained as failed based on the result. When the guard is destroyed, it is classified and expelled: if successful, it is returned to the pool and the health count is reset; if it times out, it is closed, rebuilt and added to the pool; if it is a fatal error, it is destroyed and removed. The verification unit is used to verify that the session generation counter matches the borrowing record before execution during eviction. The data acquisition unit is used to obtain the raw counter value from the device during data acquisition, store it in the raw data table according to the routing prompts in the task object, and dynamically calculate the derived indicators based on the raw data by the query layer view. When the data is refreshed, the cumulative count and time are updated lock-free through atomic comparison and exchange operations, generating refresh evidence records and exposing the telemetry interface.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.