Observable monitoring method for edge computing network

By identifying and coordinating eBPF probe conflicts in the edge computing network, generating probe conflict mapping tables and sampling scheduling sequences, the performance and stability problems caused by probe conflicts in the prior art are solved, and low overhead and high-precision network observability and fault diagnosis are achieved to ensure the stability and service quality of edge services.

CN120371650AActive Publication Date: 2025-07-25CHINA TOWER CO LTD

Patent Information

Application Number
CN202510842517.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-25
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

In the scenarios of tight resources and complex environment, existing edge computing network monitoring technology cannot effectively solve the performance and stability problems caused by probe conflicts, and it is difficult to balance the comprehensiveness of monitoring coverage and system overhead, and it is impossible to promptly detect abnormal events such as network congestion, equipment failure and security attacks.

Method used

By collecting the running status data of multiple eBPF probes on edge devices, identifying monitoring conflicts, generating probe conflict mapping tables, and combining probe conflict mapping tables to coordinate sampling permissions between probes, generating probe sampling scheduling sequences and data fingerprint libraries, predicting key event distributions based on historical event patterns, adjusting time allocations, and realizing distributed sampling and fault diagnosis.

Benefits of technology

Significantly reduce the system overhead of multi-probe monitoring, achieve low overhead and high-precision network observability, can promptly detect and diagnose network abnormalities, and ensure the service quality of edge services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371650A_ABST
    Figure CN120371650A_ABST
Patent Text Reader

Abstract

The invention discloses an observable monitoring method for an edge computing network, which comprises the following steps of: acquiring the running states of a plurality of eBPF probes on edge equipment, analyzing the access condition of the eBPF probes to the same kernel mounting point, identifying and monitoring conflicts and generating a probe conflict mapping table; based on the conflict relation, an improved distributed mutual exclusion algorithm is combined with a data fingerprint mechanism, sampling authority coordination between probes is carried out, repeated collection is avoided, and a probe sampling scheduling sequence is generated; predicting a future key event according to a historical event mode, and dynamically adjusting a time sequence interleaving strategy of each probe to form an optimized time slice distribution table; and each probe executes cooperative sparse sampling according to the time slice distribution table and the data fingerprint database, and reconstructs complete monitoring information from sparse data based on a compressed sensing theory to realize fault diagnosis. According to the invention, the system overhead of multi-probe monitoring can be obviously reduced, and low-overhead and high-precision network observability is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to network technologies, and in particular to an observable monitoring method for an edge computing network based on collaborative conflict resolution. Background Art

[0002] With the rapid development of 5G, Internet of Things (IoT), and artificial intelligence technologies, the computing mode is evolving from traditional centralized cloud computing to edge computing. Edge computing pushes computing and data storage to the network edge, providing services in a way closer to the data source, thus having key advantages such as low latency, high bandwidth, and protection of data privacy. However, edge nodes are usually resource-constrained, have heterogeneous environments, and dynamic and variable topological structures, which pose great challenges to ensuring the stability, performance, and security of the services running on them. Establishing an efficient, accurate, and low-overhead observable monitoring system for edge computing networks is crucial for timely detecting and diagnosing abnormal events such as network congestion, device failures, and security attacks, and ensuring the quality of service (QoS) of edge services. It is the core supporting technology for promoting the large-scale implementation and reliable application of edge computing technologies, and has great research significance and application value.

[0003] Currently, various paths have been developed for monitoring technologies of computing systems. Traditional network monitoring mostly relies on the Simple Network Management Protocol (SNMP) to obtain basic metrics of devices through polling, or uses technologies such as NetFlow and sFlow for traffic statistics. These methods can provide a macroscopic network view. At the system level, the industry usually constructs observability by collecting and analyzing system logs (Logging), aggregating metrics (Metrics), and distributed tracing (Tracing). In recent years, with the maturity of eBPF (Extended Berkeley Packet Filter) technology, its ability to be programmable in the kernel state and safely and efficiently capture system events without modifying the kernel source code makes it an ideal tool for achieving fine-grained system observability. Existing eBPF monitoring solutions are usually designed for specific targets. For example, eBPF probes are deployed at kprobe or tracepoint mount points to monitor specific kernel function calls, network packet sending and receiving, disk I / O, etc., and the collected data is reported to the user space for processing and analysis. These technologies provide unprecedented capabilities for deeply understanding the internal operating state of the system.

[0004] However, when existing monitoring technologies are applied to edge computing scenarios with resource constraints and complex environments, there are still some problems. Regarding the probe conflict problem, allowing it to deteriorate or adopting a simple static priority strategy cannot perform fair and efficient dynamic arbitration at runtime, thus unable to fundamentally solve the performance and stability problems brought by conflicts. Summary of the Invention

[0005] Objective of the invention: to provide an observable monitoring method for edge computing networks, aiming to solve at least one problem existing in the prior art.

[0006] Technical solution: an observable monitoring method for edge computing networks, comprising: Collect the running state data of multiple eBPF probes on edge devices, identify the listening conflicts for the same kernel event, and generate a probe conflict mapping table; Read real-time event requests, coordinate the sampling permissions among probes in combination with the probe conflict mapping table, and generate a probe sampling scheduling sequence and a data fingerprint library; Predict the distribution of key events according to historical event patterns and the probe sampling scheduling sequence, adjust the time allocation, and form a time slice allocation table; Execute distributed sampling on the kernel event stream according to the time slice allocation table, query the data fingerprint library and adjust the sampling strategy accordingly, and generate a sampling data set; Generate complete monitoring data based on the sampling data set and perform pattern matching and fault location accordingly.

[0007] According to one aspect of the present application, collecting the running state data of multiple eBPF probes on edge devices, identifying the listening conflicts for the same kernel event, and generating a probe conflict mapping table includes: Construct a probe registration table; For each active probe in the probe registration table, read the real-time performance metrics of each probe from the shared memory and analyze the listening situations of multiple probes for the same kernel event; When it is detected that multiple probes are mounted on the same mount point, calculate the degree of resource competition, and generate a probe conflict mapping table including conflicting probe pairs, conflict event types, and competition intensity.

[0008] Preferably, generating a probe conflict mapping table includes: Construct a probe conflict matrix and record the historical conflict frequency and resource competition intensity of each pair of probes; Calculate the conflict similarity between probes based on the probe conflict matrix, and cluster the probes with similar conflict patterns and low conflict frequencies into probe groups; Assign a group identifier to each probe group, and generate a grouped probe conflict mapping table including probe identifiers, the group identifiers to which they belong, and a list of group members.

[0009] Preferably, when updating the probe conflict mapping table, predict the future degree of resource competition, specifically: Extract the probe monitoring data for consecutive sampling periods from the real-time performance metrics read from the shared memory, and construct a multi-dimensional load feature vector sequence.

[0010] Calculate the first-order difference and the second-order difference based on the multi-dimensional load feature vector sequence to obtain the load change acceleration.

[0011] Integrate the current resource competition degree, first-order difference, and load change acceleration of the fusion probe to generate a predicted competition intensity.

[0012] Update the probe conflict mapping table using the predicted competition intensity.

[0013] Preferably, the process of generating the predicted competition intensity further includes introducing a pattern recognition mechanism, specifically: Analyze the autocorrelation and variance of the multi-dimensional load feature vector sequence to identify periodic or bursty load patterns.

[0014] According to the identified load patterns, dynamically select the corresponding prediction model or adjust the weights.

[0015] Generate the predicted competition intensity using the adapted prediction model or weights.

[0016] Preferably, the process of generating the predicted competition intensity further includes: Continuously monitor the deviation between the predicted competition intensity and the subsequent actually observed load values to form a predicted residual sequence.

[0017] Analyze the temporal distribution characteristics of the predicted residual sequence, distinguish isolated pulsed deviations from continuous step deviations, and accordingly determine the load dynamic pattern.

[0018] According to the determined load dynamic pattern, dynamically adjust the process noise covariance matrix inside the predictor.

[0019] When it is determined to be a step deviation, increase the value of the process noise covariance matrix.

[0020] According to one aspect of the present application, perform sampling permission coordination between probes in combination with the probe conflict mapping table, including probe competition relationship judgment, specifically: Listen to the sampling requests in the real-time event request queue, and judge whether there is a resource competition relationship between the probes that initiate the requests according to the probe conflict mapping table; For probes with a competition relationship, first calculate the hash value of the event to be collected, check whether the same data has been collected by other probes, and if so, directly reuse it, otherwise collect it again.

[0021] Preferably, performing sampling permission coordination between probes in combination with the probe conflict mapping table includes: According to the group identifier in the grouped probe conflict mapping table, perform local ticket number competition within the group to which the probe belongs, and generate the optimal probe within the group as the group representative; The representatives of each group calculate the comprehensive priority based on the monitored event type, waiting duration, and historical contribution degree; Inter-group arbitration is performed by comparing the comprehensive priorities of each group's representative to determine the final sampling executor and grant the sampling permission; After sampling is completed, the holder of the permission queries the waiting queue and directly transfers the permission to the next waiting person with the highest priority, updating the probe sampling scheduling sequence.

[0022] Among them, calculating the comprehensive priority includes: Extract the event type from the real-time event request, and query the preset event importance scoring table to obtain the basic priority score; Read the waiting duration of the group representative, and when it exceeds the dynamic threshold, improve its timeliness priority according to the preset function; Statistical ratio of the historical effective data output to resource consumption of the group representative, and calculate the resource efficiency factor; Weightedly fuse the basic priority score, timeliness priority, and resource efficiency factor to generate a comprehensive priority for inter-group arbitration.

[0023] Directly transferring the permission to the next waiting person with the highest priority includes: The probe that has completed sampling traverses the waiting queue and selects the next executor according to the comprehensive priority; Directly write the sampling permission token into the permission status bit of the selected probe, bypassing the re-competition process; Send a wake-up signal to the selected probe through the event notification descriptor to activate it from the sleep state; After the awakened probe reads the permission status bit and confirms that it has obtained the permission, it immediately performs sampling without waiting.

[0024] According to one aspect of the present application, the process of re-sampling, generating the probe sampling scheduling sequence and the data fingerprint library is specifically: Probes within the same conflict group obtain sampling permissions through a distributed mutual exclusion algorithm, The probe that has obtained the permission performs data collection, and generates a probe sampling scheduling sequence according to the sampling order; After the collection is completed, the hash value and time stamp of the event data are stored in the shared storage to form a data fingerprint library.

[0025] According to one aspect of the present application, obtaining sampling permissions through a distributed mutual exclusion algorithm includes: When the probe receives a sampling request, set its own initial selection array position; Traverse the ticket number arrays of all other probes within the conflict group, obtain the current ticket numbers of each probe, find the maximum value and add 1 as its own ticket number, and update it to its own selection array position; Perform a two - layer check on each conflict probe: Ensure that other probes complete the ticket number selection through spin - waiting, compare the ticket numbers. If the ticket number of other probes is smaller or the ticket numbers are the same but the probe identifier is smaller, continue to wait until the highest - priority sampling permission is obtained.

[0026] According to one aspect of the present application, store the hash value and timestamp of the event data in shared storage, set a validity period and clean it regularly, and form a data fingerprint library including: Perform two - level hash calculation on the event data. Use the first hash algorithm to calculate a 64 - bit hash value for the event type, timestamp, and key parameters. After combining this hash value with the probe identifier, calculate a 32 - bit fingerprint through the second hash algorithm to obtain the event data fingerprint; Store the event data fingerprint in a hash - table structure, form a fingerprint storage index, and perform addition, deletion, and maintenance according to activity to obtain the data fingerprint library.

[0027] Preferably, the process of forming the data fingerprint library further includes establishing a hierarchical adaptive verification mechanism, specifically: Construct a hierarchical verification cost model that includes various verification costs and depths.

[0028] Collect the error rate data of historical verifications from the data fingerprint library.

[0029] Generate a dynamic verification policy configuration based on the error rate data and the hierarchical verification cost model. This configuration determines the execution ratio of each verification level in subsequent verification operations.

[0030] Execute the dynamic verification policy configuration within the preset total system resource budget.

[0031] Preferably, the hierarchical adaptive verification mechanism further includes an error source location method, specifically: Statistically analyze the probe source distribution of error events in the data fingerprint library, and identify the probes with abnormal error rates as high - risk probes.

[0032] Allocate differentiated key verification resources to high - risk probes to adjust the dynamic verification policy configuration.

[0033] Construct an error propagation model. When a probe makes an error within a specific time window, temporarily enhance the verification intensity of the events of its neighboring probes.

[0034] According to one aspect of the present application, predict the critical event distribution based on historical event patterns and probe sampling scheduling sequences, adjust the time allocation, and form a time - slice allocation table including: Analyze the historical sampling patterns of each probe based on the probe sampling scheduling sequence, extract the event interval distribution from the historical event pattern data, and construct the event timing characteristics; Train a prediction model using event time series features, update the state transition probability in real time, and predict high-frequency events within a future time window when a burst mode is detected to obtain a critical event prediction set; Divide the monitoring window into multiple time slices. Under normal circumstances, allocate staggered time slices to the probes. Identify the time periods that require intensive sampling based on the critical event prediction set, and expand the time slices of the relevant probes during these time periods to generate a time slice allocation table.

[0035] Preferably, the process of adjusting time allocation to form a time slice allocation table is deepened by a multi-resource collaborative regulation mechanism, specifically: Analyze historical resource usage data to establish a resource correlation matrix among multiple system resources. Combine historical event patterns and probe sampling scheduling sequences to generate a multi-resource demand prediction vector covering multiple resource types.

[0036] Fuse the multi-resource demand prediction vector and the resource correlation matrix, calculate and generate a collaborative regulation execution plan. The collaborative regulation execution plan replaces the time slice allocation table and is used to guide subsequent collaborative sparse sampling.

[0037] Preferably, the multi-resource collaborative regulation mechanism also includes a regulation fallback function, specifically: After executing the collaborative regulation execution plan, continuously monitor key performance indicators.

[0038] Establish and maintain a regulation decision snapshot cache for storing the recent regulation decisions.

[0039] When the key performance indicators deteriorate and exceed a predetermined threshold, use the regulation decision snapshot cache to restore the system to the previous stable state.

[0040] Preferably, the multi-resource collaborative regulation mechanism also has the ability of policy self-evolution, specifically: Collect and evaluate the performance feedback data of historical regulation operations to form a regulation risk assessment model.

[0041] Based on the regulation risk assessment model, continuously optimize the core parameters used to generate the collaborative regulation execution plan.

[0042] Introduce policy version management to version the optimized core parameters and support parallel testing and comparison of new and old policies.

[0043] According to one aspect of the present application, obtaining a critical event prediction set includes: Read the event interval data in the event time series features, map it to a predefined quantization interval through logarithmic transformation to generate a quantized event interval sequence; Based on the quantized event interval sequence, update the state transition matrix of the Markov chain and calculate the new state transition matrix; Monitor the number of occurrences of consecutive identical intervals in the monitoring quantization event interval sequence. When the same interval appears continuously for a preset number of times, trigger burst detection, calculate the burst intensity, and mark it as an emergency event; Combine the new state transition matrix and the emergency event mark to calculate the probability distribution of the occurrence of each event type within the future time window, improve the prediction probability of emergency events, and generate a critical event prediction set.

[0044] According to one aspect of the present application, based on the critical event prediction set, calculate the time slice allocation table for each probe, including: Divide the monitoring window into a fixed number of basic time slices according to the number of active probes, and generate an initial bit mask with evenly staggered time slices for each probe; When the critical event prediction set indicates that there is a high-priority event in a certain time window, identify the probes responsible for monitoring this event, expand the mask bits of these probes within the prediction window, and detect whether there is a conflict. If there is a conflict, perform priority arbitration according to the event prediction confidence level to make a reasonable allocation of each time slice, and output the time slice allocation table.

[0045] According to one aspect of the present application, generate a sampling data set, including: Each probe activates the listening of the kernel event stream within the time window allocated in the time slice allocation table, and captures the corresponding kernel event data within the specified time slice; When the predicted critical event actually occurs, the relevant probes activate the compensation sampling mode, ignore the time slice limit for intensive sampling, and capture the complete critical event; Summarize the kernel event data and compensation sampling data collected by each probe within the allocated time slice, and count the ratio of the actual sampling points to the theoretical maximum sampling points to form a sampling data set.

[0046] According to one aspect of the present application, the process of coordinating sampling permissions through a distributed mutual exclusion algorithm is also regulated by an emergency status flag bit, and includes: When any eBPF probe matches the predefined high-risk event characteristics, set the emergency status flag bit by this probe; All the remaining probes waiting for sampling permissions immediately abandon their current permission requests after detecting that the emergency status flag bit is set; The probe that sets the emergency status flag bit then bypasses all scheduling coordination and obtains immediate sampling permissions; After this probe completes sampling, it clears the emergency status flag bit to restore the normal scheduling coordination process.

[0047] Beneficial effects: It can significantly reduce the system overhead of multi-probe monitoring and achieve low-overhead and high-precision network observability. Description of the Drawings

[0048] Figure 1 is the flow chart of the present invention.

[0049] Figure 2 is the flow chart of generating the probe conflict mapping table of the present invention.

[0050] Figure 3 is the flow chart of judging the probe competition relationship of the present invention.

[0051] Figure 4 is the flow chart of generating the probe sampling scheduling sequence and the data fingerprint library of the present invention. Detailed Embodiments

[0052] As Figure 1 shown, first of all, existing solutions generally lack an effective mechanism for managing runtime conflicts when multiple eBPF probes coexist. In a real edge environment, multiple monitoring or security applications from different services and different manufacturers may be deployed simultaneously, and they may create multiple eBPF probes and attach them to the same critical kernel events (such as tcp_sendmsg, block_rq_issue, etc.). This uncoordinated concurrent listening will cause fierce CPU resource preemption and redundant data collection, resulting in a sharp increase in system overhead, serious performance jitter, and even system instability in extreme cases. Existing technologies either ignore such probe conflict problems and let them deteriorate, or adopt simple static priority strategies, which cannot perform fair and efficient dynamic arbitration at runtime, thus unable to fundamentally solve the performance and stability problems brought by conflicts. Secondly, it is difficult for existing technologies to achieve an effective balance between the comprehensiveness of monitoring coverage and system overhead. In order to capture common and sudden performance anomalies or security attacks in the edge network, the monitoring system needs to have a high-time-resolution sampling ability. However, if high-frequency sampling is continuously performed, it will generate unbearable performance overhead on edge nodes where resources are already stretched. On the contrary, adopting a low-frequency static sampling strategy will almost certainly miss a large number of fleeting key events, resulting in blindness in fault diagnosis. Existing methods lack an intelligent mechanism that can predict system load and event patterns and adaptively adjust the sampling strategy accordingly.

[0053] According to one aspect of the present application, an observable monitoring method for an edge computing network based on collaborative conflict resolution S1. Real-time collect the operation status data and kernel event stream information of multiple eBPF probe instances on edge devices. By analyzing the access frequency and resource occupancy of each probe to the same kernel mount point, construct a probe conflict mapping table, and at the same time calculate resource competition indicators such as CPU occupancy rate and memory usage, providing a decision basis for subsequent conflict resolution.

[0054] S11. Traverse all loaded eBPF probe instances in the edge device kernel. By reading the file descriptors returned by / sys / kernel / debug / tracing / kprobe_events and the bpf system call, extract the mounting point information, probe type, and priority configuration of each probe, and construct a probe registry.

[0055] S12. For each active probe in the probe registry, read its real-time performance metrics through the shared memory of the BPF_MAP_TYPE_PERCPU_ARRAY type, including the number of triggers per second, the execution time per single execution, the CPU cycle consumption, etc., to form a probe performance metric set.

[0056] S13. Analyze the listening situations of multiple probes in the probe performance metric set for the same kernel events (such as block_rq_issue, tcp_sendmsg, etc.). When it is detected that two or more probes are mounted on the same tracepoint or kprobe, calculate their resource competition degree, and generate a probe conflict mapping table containing the conflicting probe pairs, the conflicting event types, and the competition intensity.

[0057] For example, if multiple probes listen to the tcp_sendmsg event simultaneously, it is considered that there is a listening conflict. At this time, it is necessary to calculate the resource competition degree between them. As an example, the competition intensity C i between two probes P j at the same mounting point can be calculated by the following formula: ij C C ij = w t * { (N i + N j ) / N total}+ w c * { T i * N i + T j * N j ) / ∑ n k=1 (T k * N k )}; C ij is the competition intensity between probe P i and P j . N i , N j are the number of triggers of probe P i , P j per unit time respectively. N total is the total number of triggers of this mounting point per unit time. T i , T jThey are the average time taken for a single execution of probe P i and P j respectively. n is the total number of active probes in the system. w t and w c are configurable weight factors, representing the degree of attention to the two dimensions of trigger frequency and execution time respectively. For example, they can both be set to 0.5.

[0058] Summarize all identified conflict relationships and their quantified competition intensities to generate a probe conflict mapping table. The data structure of this table can be a set of {(ProbeID_i, ProbeID_j), EventType, CompetitionStrength}, which clearly records which probe pairs, in which types of events, and the degree of competition.

[0059] S14. Collect system-level resource usage data, including CPU usage rate in / proc / stat, memory occupancy in / proc / meminfo, and hardware performance counter data obtained through perf_event_open. Combine the probe performance metric set to calculate the resource occupancy ratio of each probe and output the resource competition metric.

[0060] S2. Based on the conflict relationships identified in the probe conflict mapping table and the real-time event request queue, use an improved distributed Bakery algorithm to coordinate the sampling permissions among probes. Introduce a data fingerprint mechanism to avoid multiple probes from repeatedly collecting the same event, generate a probe sampling scheduling sequence, and maintain a data fingerprint library to ensure the efficient collaborative work of each probe.

[0061] Through a distributed mutual exclusion and collaboration mechanism, ensure that at any moment, for a conflict event, only one probe can obtain the final sampling permission, and at the same time, avoid redundant collection through a data reuse mechanism.

[0062] S21. Read the probe conflict mapping table to identify the probe groups with current resource competition, and initialize the distributed coordination data structure for each group of conflicting probes, including an atomic ticket number array, a selection flag bit array, and a data hash storage area, and establish a conflict coordination context.

[0063] S22. Listen to the sampling requests in the real-time event request queue. When a probe needs to collect data, first calculate the hash value of the event to be collected, traverse the data fingerprint library to check if the same data has been collected by other probes. If so, directly reuse it to avoid repeated collection and update the data reuse statistics.

[0064] S23. For events that need to be newly collected, execute the improved Bakery algorithm to obtain sampling permissions: The probe first sets the selection flag bit, reads the current ticket numbers of all conflicting probes, selects the maximum value plus 1 as its own ticket number, and then clears the selection flag bit; determines the sampling order by comparing the ticket numbers and the probe IDs, waits to obtain the sampling permissions, and then executes data collection to generate the probe sampling scheduling sequence.

[0065] S24. The probes that successfully collect data store the hash value and timestamp of the event data in the shared data fingerprint library, set the fingerprint validity period to 100 milliseconds, and automatically clear it after timeout to ensure that the fingerprint library will not grow infinitely, while providing data reuse opportunities for other probes.

[0066] This process transforms the disordered resource competition when multiple eBPF probes concurrently access the same kernel mount point into an orderly, fair, and efficient collaborative scheduling process.

[0067] S3. Combine the probe sampling scheduling sequence and historical event pattern data, use the lightweight Markov chain model to predict the distribution of key events in the future time window, dynamically adjust the timing interleaving strategy of the probes accordingly, and output the optimized time slice allocation table and the key event prediction set to achieve the balance between sampling coverage and system overhead.

[0068] S31. Analyze the historical sampling patterns of each probe based on the probe sampling scheduling sequence, extract the event interval distribution within the most recent 1000 sampling periods from the historical event pattern data, construct an event interval histogram with 16 logarithmic intervals as buckets, and form the event timing characteristics.

[0069] S32. Use the event timing characteristics to train the lightweight Markov chain model, adopt the exponential moving average algorithm to update the state transition probability in real time, trigger the burst mode detection when three consecutive event intervals fall into the same bucket, predict the high-frequency events that may occur within the next 10 milliseconds, and output the key event prediction set.

[0070] S33. Dynamically calculate the time slice allocation strategy for each probe according to the resource competition index and the key event prediction set: The basic allocation uses a 32-bit mask to represent the time slices within a 10-millisecond window (each bit represents approximately 0.3 milliseconds). Under normal circumstances, the probe obtains the interleaved time slices. When a key event is predicted, the time slices of the relevant probes are temporarily extended to generate the optimized time slice allocation table.

[0071] S34. Set the trigger threshold and duration of the compensation mechanism. When the event confidence level in the key event prediction set exceeds 0.8, allow the relevant probes to break through the timing interleaving limit for 100 microseconds of full sampling, and then automatically resume to the normal time slice allocation mode to ensure that the system will not be in a high-load state for a long time.

[0072] Based on this prediction result, the system can dynamically adjust the time slice allocation table of each probe, and intelligently tilt the precious sampling resources (i.e., CPU time slices) from the period with stable network behavior to the high-incidence period when key events (such as network congestion and application anomalies) are about to occur, and trigger a compensation sampling mechanism to ensure data integrity. Without sacrificing the ability to capture key fault features, the system overhead of daily continuous monitoring is minimized to the greatest extent possible.

[0073] S4. Perform distributed sampling on the kernel event stream according to the optimized time slice allocation table. Each probe collaboratively collects data according to the allocated time slice and the predicted key events, and refers to the data fingerprint library to avoid redundant sampling, finally forming a sparse sampling data set and its corresponding sampling coverage rate matrix.

[0074] S41. Each probe activates the listening of the kernel event stream within the specified time slice according to the time window allocated in the optimized time slice allocation table, and captures the corresponding kernel event data through the BPF_PROG_TYPE_TRACEPOINT program of eBPF, including information such as timestamps, event types, and key parameters.

[0075] S42. During the sampling process, the probe first queries the data fingerprint library to verify whether the event to be collected has been processed by other probes. If a matching fingerprint is found, the actual collection step is skipped, and the existing data is directly referenced and the reference count is updated, reducing the overhead of repeated kernel-space to user-space data copying.

[0076] S43. When the predicted events in the key event prediction set actually occur, the relevant probes immediately activate the compensation sampling mode, temporarily ignoring the time slice limit for intensive sampling to ensure the complete capture of key events. The supplementary data collected is marked with a special flag to distinguish it from the regular sampling data.

[0077] S44. Aggregate the event data collected by each probe within the allocated time slice, calculate the ratio of the actual number of sampling points to the theoretical maximum number of sampling points in each time window, construct a sampling coverage rate matrix reflecting data integrity, and integrate all the collected raw data into a sparse sampling data set.

[0078] S5. Input the sparse sampling data set and the sampling coverage rate matrix into the data reconstruction module based on the compressed sensing theory. Combining with the pre-computed sparse basis, use the improved orthogonal matching pursuit algorithm to recover the complete monitoring information from partial observations, generating the reconstructed complete monitoring data and its reconstruction confidence score.

[0079] S51. Load the pre-computed sparse basis matrix (64×256 dimensions) and measurement matrix (32×256 dimensions) optimized for kernel event features. These matrices are obtained through offline training based on a large amount of historical monitoring data and can effectively represent the sparse characteristics of kernel events. For example, the K-SVD algorithm is used to train 100,000 pieces of normal and abnormal network event data collected to generate an over-complete dictionary as the sparse basis matrix.

[0080] S52. Group the sparse sampling data set according to event types and time windows, apply the measurement matrix to the data in each group for dimensionality reduction projection, and at the same time mark the missing positions with reference to the sampling coverage matrix to generate compressed measurement vectors.

[0081] S53. Execute the improved orthogonal matching pursuit (OMP) algorithm for signal reconstruction: First, calculate the inner product of the compressed measurement vector and each atom of the sparse basis to find the most matching basis vector, iteratively update the sparse representation coefficients and calculate the residual. When the residual norm is lower than the preset threshold or reaches the maximum number of iterations, terminate to obtain the sparse representation coefficients.

[0082] In this embodiment, after each iteration, the kurtosis value of the current residual signal is calculated. Kurtosis is a statistic that measures the impulsiveness of a signal. When the kurtosis value is much greater than 0, it indicates that the residual still contains significant, non-Gaussian sparse signal components and the iteration should continue; when the kurtosis value approaches 0, it indicates that the residual is close to Gaussian white noise and the main signal components have been extracted, and the iteration can be terminated. This improvement enables the OMP algorithm to intelligently stop the iteration according to the actual sparse characteristics of the signal, thus achieving a better balance between reconstruction accuracy and computational efficiency and solving the drawbacks brought by a fixed number of iterations.

[0083] S54. Use the sparse representation coefficients and the pre-computed sparse basis for signal reconstruction to restore the complete monitoring data time series. At the same time, calculate the credibility of each reconstructed data point based on the residual size and the number of iterations during the reconstruction process, and output the reconstructed complete monitoring data and the corresponding reconstruction confidence score.

[0084] In the practical application of edge computing, sampling can be performed with fewer local CPU and memory resources, and the monitoring data can be transmitted back to the central platform with a smaller network bandwidth (due to less data transmitted), while still obtaining a monitoring view close to full-scale acquisition.

[0085] S6. Comprehensively analyze the reconstructed complete monitoring data and the reconstruction confidence score, perform pattern matching with the preset fault feature library, identify network abnormal behaviors and locate the fault source, and finally output a fault diagnosis report including the fault type, impact range, and repair suggestions, as well as in-depth root cause analysis results.

[0086] S61. Input the reconstructed complete monitoring data into the anomaly detection engine, calculate the statistical features of key metrics such as network latency, packet loss rate, and bandwidth utilization in a sliding window manner, and trigger an anomaly alarm when the metric deviates from the baseline by more than 3 standard deviations, generating an anomaly event sequence.

[0087] S62. Weight the credibility of the anomaly event sequence in combination with the reconstruction confidence score, prioritize the processing of high-confidence anomaly events, identify the fault type (such as network congestion, device failure, configuration error, etc.) by pattern matching with a predefined fault feature library, and output a preliminary fault classification result.

[0088] S63. Based on the fault classification result and the complete monitoring data, construct a fault propagation graph, analyze the correlation relationship of anomaly events in the time and space dimensions, determine the root cause and impact scope of the fault by tracing the fault source and impact path, and generate a detailed root cause analysis result.

[0089] S64. Integrate the fault classification result and the root cause analysis result, combine the topology structure and business priority of the edge computing environment, generate a comprehensive fault diagnosis report including fault description, impact assessment, repair suggestions, and preventive measures, and provide actionable decision support for operation and maintenance personnel.

[0090] According to one aspect of the present application, S23. The improved Bakery algorithm is specifically as follows: When the probe receives a sampling request, first read the shared data structure of the current probe group from the conflict coordination context, set the position of its own choosing array to 1 through the atomic operation atomic_set, indicating entering the ticket number selection stage, and at the same time execute a memory barrier instruction to ensure that all CPU cores can see this state change, generating a selection status flag.

[0091] After the selection status flag takes effect, the probe traverses the ticket array of all other probes in the conflict group, reads the current ticket number of each probe through atomic_read, finds the maximum value max_ticket, writes (max_ticket + 1) to its own ticket array position through atomic_set, and then clears the choosing flag to complete the ticket number acquisition, outputting the probe ticket number sequence.

[0092] The probe holding the probe ticket number sequence enters the waiting stage and performs two-layer checks on each conflicting probe: first, spin wait (cpu_relax) to ensure that other probes complete the ticket number selection (choosing[i] == 0), and then compare the ticket numbers. If the ticket number of other probes is smaller or the ticket numbers are the same but the probe ID is smaller, continue to wait until the highest priority is obtained, forming a sampling permission token.

[0093] The probe that obtains the sampling permission token performs the actual data collection operation. After the collection is completed, it sets its value in the ticket array to 0 through atomic_set to release the sampling permission. At the same time, it stores the hash value of the collected data in the shared buffer. Other waiting probes re-evaluate their priorities after detecting that the ticket number is cleared to zero, and finally output the updated probe sampling scheduling sequence.

[0094] In another embodiment of the present application, the calculation of the priority is specifically Priority(i) = α·EventPriority(i) + β·ResourceEfficiency(i) + γ·HistoricalYield(i); Where: EventPriority(i) is the real-time importance score of the monitoring event of probe i; ResourceEfficiency(i) is the resource efficiency of probe i (collection value / resource consumption); HistoricalYield(i): the historical effective data output rate of probe i; α is the event importance weight, β is the resource efficiency weight, and γ is the historical data output weight; The ticket number generation policy is Ticket(i) = BaseTicket + Priority(i) × TicketRange, where BaseTicket is the base ticket number and TicketRange is the coefficient used to amplify the influence of the priority.

[0095] According to one aspect of the present application, S24, the management of the data fingerprint library is specifically as follows: Perform two-level hash calculation on the event data to be stored. First, use the xxHash algorithm to calculate the 64-bit hash value of the event type, timestamp (precision reduced to microseconds), and the first 32 bytes of key parameters. Then, combine this hash value with the probe ID and calculate the final 32-bit fingerprint through CRC32 to ensure that the same data of different probes generates the same fingerprint, and output the event data fingerprint.

[0096] Store the event data fingerprint in a hash table structure based on the open addressing method. The table size is 4096 slots. Each slot contains a fingerprint value, a timestamp, a reference count, and a pointer to the next node (next). When a hash conflict occurs, linear probing is used to find the next empty slot. At the same time, maintain a doubly linked list sorted by timestamp for timeout cleaning to form a fingerprint storage index.

[0097] Start an independent cleaning thread to scan the fingerprint storage index every 10 milliseconds. Start checking from the head of the timestamp linked list, mark fingerprint entries with a survival time exceeding 100 milliseconds as expired. If the reference count is 0, immediately delete and release the slot. If the reference count is greater than 0, force deletion after a 50-millisecond delay to ensure that the size of the fingerprint database is controllable, and output the set of active fingerprints.

[0098] Protect the concurrent access to the set of active fingerprints through a read-write spin lock. The read operation (querying fingerprints) acquires a read lock to allow multiple probes to query simultaneously. The write operation (inserting / deleting fingerprints) acquires a write lock to ensure exclusive access. Use the try_lock mechanism to avoid long-term blocking. If the lock acquisition fails, add the operation request to the pending queue, and finally maintain a consistent data fingerprint database.

[0099] According to another aspect of the present application, the processing process is described through an example: When a probe (such as P i ) needs to collect a kernel event, the following process is executed: First, perform a competition relationship judgment and data reuse check. Probe P i According to the probe conflict mapping table generated in the first embodiment, judge whether there is a competition relationship for the event currently to be collected. If there is, probe P i First calculate the data fingerprint of the event to be collected (the calculation method will be detailed later), and query a globally shared data fingerprint database. If the fingerprint already exists, it means that other conflicting probes have collected the same or highly similar event data in the near future (for example, within 100 milliseconds). At this time, P i Will directly reuse this data and abandon this collection request, thus avoiding repeated kernel operations and data transmissions. If the data fingerprint does not exist, it indicates that a new data collection is required. At this time, start a distributed mutual exclusion algorithm to obtain a unique sampling permission.

[0100] In this embodiment, preferably, an improved distributed Bakery algorithm is adopted. The process is as follows: Enter the selection stage: Probe P i In a globally shared boolean choosing array of size n (the number of probes in the conflict group), set choosing[i] to 1 through an atomic operation. This indicates that P i Is selecting a ticket number.

[0101] Obtain the ticket number: P i Traverse the globally shared integer ticket array of size n, find the maximum value of all current ticket numbers, and add 1 as its own ticket number, and write it into ticket[i] through an atomic operation. Then, set choosing[i] to 0, indicating that the ticket number selection is complete.

[0102] Waiting for authorization: P i Enter the waiting loop. It traverses all other probes P within the conflict group j (where j = i) and performs two-level checks: First-level check: Wait for choosing[j] to become 0 to ensure that P j has completed its ticket selection.

[0103] Second-level check: Wait for the condition (ticket[j] == 0) OR ((ticket[j], j) > (ticket[i], i)) to hold. Here, (ticket[j], j) > (ticket[i], i) means a lexicographical comparison based on the ticket number and the probe ID size. That is, if the ticket number of probe P j is smaller than that of P i , or the ticket numbers are the same but the ID of P j is smaller than that of P i , then P i must wait. ticket[j] == 0 means that probe P j is not currently requesting sampling.

[0104] Perform acquisition and release: When P i passes the checks for all other probes, it obtains the exclusive sampling permission and can perform the actual data acquisition operation. After the acquisition is completed, P i sets ticket[i] to 0 through an atomic operation to release the permission.

[0105] Through the above distributed mutual exclusion algorithm, a fair and conflict-free probe sampling scheduling sequence will ultimately be formed.

[0106] After obtaining the sampling permission, the probe needs to store the fingerprint of the data into the shared data fingerprint library after collecting the data. The fingerprint generation and storage mechanism is as follows: Fingerprint calculation: Two-level hashing is used to reduce the conflict probability.

[0107] It should be noted that in the previous embodiment, this hash value was combined with the probe ID, and the final 32-bit fingerprint was calculated through CRC32 to ensure that the same data from different probes generates different fingerprints.

[0108] If the fingerprint contains the probe ID, different probes collecting the same data will generate different fingerprints, and the purpose of data reuse cannot be achieved. Therefore, it is processed as follows: Level 1: Use a high-speed hashing algorithm (such as xxHash) to calculate a 64-bit hash value H64 for the core event data: H64 = xxHash(EventTypel || Timestamp\mus || KeyParams1..32); where Timestamp is the timestamp reduced to microsecond precision, KeyParams1..32 are the first 32-byte key parameters of the event, EventTypel is the event type field, and mus represents the timestamp that accurately represents the event occurrence time to the microsecond level.

[0109] Level 2: Calculate a 32-bit fingerprint F32 for this 64-bit hash value, for example, using the CRC32 algorithm: F32 = CRC32(H64), and this F32 is the final fingerprint of the event data.

[0110] Storage and management: Store the calculated event data fingerprint F32, the current timestamp, and a reference count in a shared hash table implemented based on open addressing. At the same time, set a validity period for each fingerprint item, for example, 100 milliseconds. An independent cleaning thread will regularly scan the hash table and remove expired fingerprint items with a reference count of zero to ensure that the fingerprint library does not grow infinitely and maintain its timeliness.

[0111] In another embodiment of this application, when a probe needs to collect data, the data stream flows through the following path: The probe first checks whether it has a pre-allocated time slot. If it does and the time matches, it directly performs sampling. Otherwise, the probe submits a sampling request to the group it belongs to and participates in the local ticket number competition within the group. The probe that obtains the smallest ticket number within the group becomes the temporary representative of the group.

[0112] The representatives of each group carry the sampling requirements of their respective groups to participate in the second-level competition. The system selects the final sampling executor based on the comprehensive priority. After the probe that obtains the permission completes the data collection, it writes the result into the shared buffer, and at the same time checks the waiting queue and wakes up the next waiter.

[0113] Specifically, the system first dynamically divides the probes into several groups according to the historical conflict patterns and resource usage characteristics of the probes. Each group contains 3 to 7 probes. The principle of grouping is to gather probes with low conflict frequencies in the same group, thereby reducing the intensity of competition within the group.

[0114] When a probe needs to sample, it first conducts local competition within the group it belongs to. Each group maintains an independent ticket number sequence, and the probe only needs to check the status of other members within the group. This reduces the operation that originally needed to check n probes to only about 5, significantly reducing the coordination overhead.

[0115] Each group selects a temporary representative through local competition, and this representative participates in the inter-group competition at the second level with the sampling requests within the group. The inter-group competition adopts a priority-based scheduling strategy, which is no longer a simple comparison of ticket numbers, but comprehensively considers the event type, waiting time, and historical contribution degree to calculate the dynamic priority.

[0116] Balances fairness and efficiency: relatively fair rotation is maintained within the group, while intelligent scheduling is performed between groups according to the importance of tasks.

[0117] In another embodiment of the present application, step S24 can also be: The probe that successfully collects data needs to convert the event information into a data fingerprint that is independent of the probe identity and can be used for efficient deduplication, and store it in the globally shared data fingerprint library. This process aims to provide a reliable basis for data reuse for other conflicting probes, and its preferred implementation methods include: The first level (content hash): Use a high-speed hash algorithm (such as xxHash) to calculate a 64-bit intermediate hash value for the core content of the event. The core content preferably includes the event type, a timestamp reduced to microsecond precision, and key parameters that can uniquely identify the event (such as the source / destination IP and port of a network packet, the block address of disk I / O, etc., and the first 32 bytes can be intercepted). The timestamp precision is reduced to microseconds to generate consistent hash values even when different probes capture almost simultaneously on different CPU cores.

[0118] The second level (final fingerprint): In order to obtain a shorter fingerprint suitable for use as a key in the hash table, further apply a second hash algorithm (such as CRC32) to the above 64-bit intermediate hash value to calculate a 32-bit final fingerprint.

[0119] Store the calculated 32-bit fingerprint, together with its corresponding exact timestamp, reference count (initially 1), and an optional pointer to the data storage location, into a globally shared hash table implemented based on open addressing. The size of this hash table can be preset (such as 4096 slots), and linear probing is used to solve hash conflicts when they occur.

[0120] Start an independent low-priority cleaning thread, which scans the fingerprint library at a fixed period (such as every 10 milliseconds). It marks the fingerprint items whose survival time exceeds the preset expiration period (TTL, such as 100 milliseconds) as expired according to the timestamps stored in each fingerprint item. For expired items with a reference count of zero, they are immediately removed from the hash table; for items with a reference count greater than zero, a longer forced deletion time can be set to prevent ongoing read operations from failing. This mechanism ensures the timeliness of the fingerprint library and does not grow indefinitely.

[0121] Preferably, xxHash can be replaced by hash algorithms such as MurmurHash3. In terms of storage structure, in addition to open addressing, hash tables can also be implemented using chaining. In scenarios with extremely high conflicts, optionally, a more complex lock-free hash table design can be adopted to eliminate the overhead of lock contention. In terms of the cleaning mechanism, in addition to an independent cleaning thread, amortized cleaning can also be used, that is, during each write operation, several potentially expired slots are randomly selected and cleaned.

[0122] According to one aspect of the present application, the real-time update process of the Markov chain model is specifically as follows: Read the original event interval data (in nanoseconds) in the event time series features, and map it to 16 predefined buckets through logarithmic transformation log2(interval), where bucket[0] represents [0, 1] microseconds, and bucket

[15] represents [32768, 65536] microseconds. Values outside the range are truncated to generate a quantized event interval sequence.

[0123] Update the state transition matrix of the Markov chain based on the quantized event interval sequence, using the Exponential Moving Average (EMA) algorithm: new transition probability = α × observed frequency + (1 - α) × historical probability, where α = 0.125 ensures the smooth decay of historical information, and normalization is immediately performed after each state transition update to ensure that the sum of probabilities is 1, and the real-time state transition matrix is output.

[0124] Monitor the number of consecutive occurrences of the same bucket in the quantized event interval sequence. When the same bucket appears continuously 3 times or more, burst detection is triggered. Calculate the burst intensity score = consecutive times × the historical rarity of this bucket. When the score exceeds the threshold of 10, it is marked as an emergency event, record the start time and expected duration of the burst, and generate an emergency event mark.

[0125] Combine the real-time state transition matrix and the emergency event mark, and calculate the probability distribution of each bucket within the next 10 milliseconds through the forward algorithm. For the event types marked as bursts, their predicted probabilities are additionally increased by 50%. Finally, output a key event prediction set including event types, prediction time windows, and confidence levels (in the range of 0 - 1).

[0126] According to one aspect of the present application, S33, the dynamic time slice allocation process, is specifically as follows: Divide the 10 - millisecond monitoring window into 32 basic time slices according to the number N of active probes in the system. Each time slice is approximately 0.3125 milliseconds. Generate an initial 32 - bit mask for each probe through the round - robin algorithm. For example, when there are 3 probes, assign staggered patterns of 0x92492492 (001001...), 0x49249249 (010010...), 0x24924924 (100100...) respectively to ensure uniform distribution, and output the basic time - slice mask.

[0127] When the critical - event prediction set indicates that there is a high - priority event in a certain time window, identify the set of probes responsible for monitoring this event type. Expand the mask bits of these probes to all 1s through the OR operation within the predicted time window, and at the same time compress the time slices of other probes proportionally (up to 50% of the original), generating an extended time - slice mask.

[0128] Detect conflict situations in the extended time - slice mask. When the extended requests of multiple probes cause some time slices to be over - allocated (occupied by more than 2 probes simultaneously), start priority arbitration: retain the extended permissions of the probes with the highest event - prediction confidence, and the remaining probes return to the basic allocation. Ensure that each time slice is shared by at most 2 probes through iterative adjustment, and finally output the coordinated and optimized time - slice allocation table.

[0129] Learn the temporal patterns of event occurrences from historical data and use this pattern to predict the future, so as to prospectively tilt the limited sampling resources to the time windows and probes more likely to have critical events.

[0130] Specifically, this step can be decomposed into: First, construct event - temporal features. Based on the probe - sampling scheduling sequence generated in Embodiment 2, analyze the timestamps of the successfully sampled events recorded therein, and calculate the time intervals (intervals) between consecutive events. For the convenience of model processing, perform logarithmic transformation and quantization on the original interval data at the nanosecond level, and map it to 16 predefined discrete buckets, generating a quantized event - interval sequence. For example, bucket[k] can represent events with time intervals in the range of [2k, 2k + 1) microseconds.

[0131] Second, use the prediction model to generate a critical - event prediction set. In this embodiment, preferably adopt a lightweight Markov - chain model.

[0132] Using the above - mentioned quantized event - interval sequence as input, construct a 16×16 state - transition matrix M. The element M in the matrix ijRepresents the probability that the event interval transfers from bucket[i] to bucket[j]. This matrix is updated in real-time using the Exponential Moving Average (EMA) algorithm to adapt to changes in the event pattern: M ij,new =α • O ij +(1 - α) • M ij,old where: M ij,new and M ij,old are the updated and pre-updated transition probabilities respectively. O ij is an indicator (1 if the currently observed transition is i→j, 0 otherwise) of whether the current observed transition is from i to j. α is the smoothing factor, which can be set to 0.125 to balance the importance of historical data and current observations.

[0133] Monitor the quantified event interval sequence. When the same bucket (e.g., bucket[0] representing a very short time interval) appears continuously for a preset number of times (e.g., 3 times), trigger the burst mode detection. A burst intensity score can be calculated, and when the score exceeds the threshold, mark the event type as a sudden event.

[0134] Combining the updated state transition matrix M and the sudden event marking, predict the probability distribution of the occurrence of each bucket (i.e., various event intervals) within a short future time window (e.g., 10 milliseconds) through the forward algorithm. For the types marked as sudden events, their predicted probabilities will be additionally increased (e.g., increased by 50%). Finally, output a key event prediction set, which includes the event type, prediction time window, and confidence score.

[0135] Then, dynamically calculate the time slice allocation table according to the prediction results. Divide a monitoring window (e.g., 10 milliseconds) into a fixed number of basic time slices (e.g., 32), and each time slice is about 0.3 milliseconds. The time slice allocation of each probe is represented by a 32-bit mask (bitmask).

[0136] Under normal circumstances, generate a uniformly interleaved initial bitmask for N active probes to ensure equal sampling opportunities. For example, for 3 probes, its mask can be the result of circular shifting, ensuring that at most one probe is active in each time slice.

[0137] When the key event prediction set indicates that a high-priority event is very likely to occur within a certain time window, identify the probe responsible for this event. Expand the mask bit of this probe to all 1s through the OR operation within the corresponding time window, so that it obtains temporary dense sampling permission. At the same time, the time slices of other irrelevant probes can be compressed proportionally to balance the total system load. If the expansion requests of multiple probes cause conflicts, then perform priority arbitration according to the confidence of the event prediction, and retain the expansion permission of the probe with the highest confidence.

[0138] Through the above steps, an optimized time slice allocation table that dynamically changes over time is finally generated to guide the collaborative sampling in the next stage.

[0139] According to one aspect of the present application, S53, the improved OMP algorithm is specifically as follows: Calculate the inner product of the compressed measurement vector with each atom in the pre-computed sparse basis. Use the SIMD instruction set (such as AVX2) to parallelize the dot product operations of multiple atoms. Divide the 64×256 sparse basis into 8 sub-blocks for parallel calculation. Find the atom index with the largest absolute value of the inner product, and record it as the best matching atom and its correlation coefficient.

[0140] Update the residual vector based on the best matching atom and the correlation coefficient, and adopt a delayed update strategy: first accumulate the contributions of multiple atoms into a temporary buffer, update the actual residual every 4 atoms to reduce the memory access overhead. At the same time, maintain the orthogonal basis of the selected atom set, and ensure that the newly added atom is orthogonal to the selected atoms through the Gram-Schmidt orthogonalization process to generate an updated residual vector.

[0141] Dynamically adjust the sparsity parameter according to the energy distribution of the updated residual vector. When the residual energy is concentrated in a few components, reduce the sparsity to accelerate convergence. When the residual energy is dispersed, increase the sparsity to improve the reconstruction accuracy. Calculate the kurtosis of the residual as the adjustment basis, and output the adaptive sparsity parameter.

[0142] Combine the L2 norm of the updated residual vector, the current iteration count, and the adaptive sparsity parameter to perform multi-condition convergence determination: if the residual norm is lower than 0.1% of the input signal energy, or the iteration count exceeds twice the sparsity, or the improvement in three consecutive iterations is less than 1%, then terminate the iteration and output the final sparse representation coefficients.

[0143] According to another aspect of the present application, in this embodiment, this step is the final execution link of all the foregoing scheduling and prediction strategies. Each eBPF probe acts as an independent execution unit and collaborates according to a globally consistent strategy.

[0144] Specifically, each probe performs the following operations: First, perform distributed sampling based on time slices. Each probe loads the optimized time slice allocation table generated in Embodiment III. During the time slice allocated to it, activate the listening and data capture of the kernel event stream. During the unallocated time slices, the probe is in a dormant or low-power state, thereby reducing the overall system overhead.

[0145] Secondly, perform compensatory sampling. When a critical event marked by the critical event prediction set actually occurs, the probe responsible for monitoring this event will immediately activate the compensatory sampling mode. In this mode, the probe will temporarily ignore the current time slice limit and perform intensive sampling for a short period (e.g., 100 microseconds) to ensure that critical and high-value event data can be completely captured. The collected compensatory data will be marked with a special tag for distinction.

[0146] Finally, summarize the data and construct a sampling coverage rate matrix. Summarize the regular event data collected by all probes within their respective time slices, as well as the critical event data collected in the compensatory mode. At the same time, calculate the ratio of the actual number of event points collected to the maximum number of event points that can be collected theoretically within each monitoring window, and construct a sampling coverage rate matrix. This matrix and the summarized sparse sampling data set will be used together as the input for subsequent steps (such as data reconstruction in S500).

[0147] Through the method provided by the embodiments of the present invention, by performing distributed coordination on the sampling conflicts among eBPF probes, dynamically adjusting the sampling strategy in combination with event prediction, and performing collaborative sparse sampling according to the generated strategy, the technical effect of significantly reducing the system performance overhead brought by multi-probe concurrent monitoring while ensuring the integrity of critical events is achieved, thereby solving the contradiction problem in the prior art that it is difficult for eBPF monitoring schemes to balance monitoring comprehensiveness and low system overhead.

[0148] This embodiment aims to elaborate in detail how to specifically calculate the resource competition degree among multiple eBPF probes when they are detected to be mounted on the same kernel mount point in step S13, so as to provide a quantifiable and accurate decision basis for generating the probe conflict mapping table.

[0149] In this embodiment, probe P i and P j The resource competition degree C ij on the same mount point is defined as a comprehensive index that combines the considerations of the CPU time consumption ratio and the event trigger frequency conflict. Its specific calculation formula is as follows: C ij =w cpu •{(N i *T i +N j *T j ) / ∑ k∈S (N k *T k )}+w freq {(N i +N j ) / N event}; Among them, the definitions of the parameters in the formula are as follows: C ij represents the resource competition degree between probes P i and P j . Its value range is [0, 1]. The larger the value, the more intense the competition. S represents the set of all active probes at the current conflict mount point. N i , N j , N k respectively represent the number of times probes P i , P j , P k are triggered within a unit statistical period (e.g., 1 second). T i , T j , T k respectively represent the average CPU time consumption (e.g., in nanoseconds) for a single execution of probes P i , P j , P k . N event represents the total number of times the events corresponding to this kernel mount point occur within the same unit statistical period. w cpu and w freq : respectively represent the weight factors of the CPU competition component and the frequency competition component. Both are positive numbers and their sum is 1 (e.g., w cpu = 0.6, w freq = 0.4). This allows for focused adjustment of different types of competition according to the monitoring target.

[0150] The first part of this formula measures the proportion of the total CPU time consumption of two probes in the total consumption of all conflicting probes, reflecting their direct competition intensity for CPU resources. The second part measures the proportion of the sum of the trigger times of two probes in the total number of events, reflecting the overlap degree or conflict frequency of their monitoring. The calculated C ij value will be written into the probe conflict mapping table together with the conflict probe pair {P i , P j} and the conflict event type.

[0151] The standard OMP algorithm usually uses a fixed sparsity K (or a fixed number of iterations) as the termination condition. In this embodiment, this termination condition is dynamically changed. After updating the residual vector rk in each round or every few rounds of iteration, a process of residual evaluation and parameter adjustment is added.

[0152] After the k-th iteration, calculate the kurtosis value Kurt(r k ) of the current residual vector rk. Kurtosis is a statistic that measures the sharpness of a probability distribution, and its calculation formula is: Kurt(r k ) = {[∑ n i=1 (rk,i -r * k ) 4 ) / n] / ∑ n i=1 ((r k,i -r * k ) 2 ) 2 / n}-3; where n is the dimension of the residual vector, and r k,i is the i-th component of the vector, and r * k is the mean of the vector. Using the excess kurtosis, i.e., the form after subtracting 3, the kurtosis of the normal distribution is 0.

[0153] Adjust the iterative strategy of the OMP algorithm based on the kurtosis value. Gaussian trend judgment (Kurt(r k )≈0): When the kurtosis value is close to 0 (e.g., -0.5 < Kurt(r k ) < 0.5), it indicates that the residual signal is becoming more and more like Gaussian white noise, and the useful signal components have been basically extracted. At this time, the iteration should be prepared to terminate.

[0154] Pulse characteristic judgment (Kurt(r k )≫0): When the kurtosis value is much greater than 0 (e.g., Kurt(r k ) > 2), it indicates that there are still significant, pulse-like un-recovered signal components in the residual signal. At this time, the iteration should continue or even be accelerated.

[0155] Dynamically adjust the sparsity / iteration upper limit: Set an initial maximum number of iterations K max (i.e., the initial sparsity). During the iteration, dynamically adjust an effective iteration upper limit Keff according to the kurtosis value: After the k-th iteration, if Kurt(r k ) first enters the interval close to 0, then update the effective iteration upper limit Keff to k + Δk, where Δk is a small additional number of iterations (e.g., 2 or 3), allowing the algorithm to terminate after a final fine-tuning.

[0156] If Kurt(r k ) remains large, then Keff remains K max , and the algorithm continues to iterate normally.

[0157] The termination condition of OMP becomes: when the number of iterations k reaches Keff, or the residual norm is lower than the preset threshold, the algorithm terminates.

[0158] Through this method, the OMP algorithm can intelligently judge the convergence state of the reconstruction process, stop in time when the residual approaches the noise, avoid overfitting and unnecessary calculations; ensure the continuity of iteration when there is still a strong signal in the residual, so as to achieve a better balance between efficiency and accuracy.

[0159] According to one aspect of the present application, the specific construction method of the lightweight Markov chain model is as follows: In order to improve the prediction accuracy and domain relevance of the model, the model state (State) in this embodiment is defined as a binary tuple: State = (EventType, IntervalBucket).

[0160] State space definition: EventType(E): represents the specific kernel event type, such as tcp_sendmsg, block_rq_issue, etc. Assume that the system focuses on M key event types, then E ∈ {E1, E2,..., EM}.

[0161] IntervalBucket(I): represents the time interval between the current event and its previous event of the same type, and is mapped to one of 16 discrete buckets after logarithmic transformation, that is, I ∈ {I0, I1,..., I15}.

[0162] Therefore, the size of the state space S of the entire model is M × 16. Each state si,j = (Ei, Ij) represents that the event Ei occurs, and the time interval between it and the previous Ei event is within the range of bucket Ij.

[0163] Construction and initialization of the state transition matrix: The state transition matrix P is a square matrix with a dimension of (M × 16) × (M × 16).

[0164] The matrix element Puv represents the probability that the system transfers from state u to state v. Specifically, P(Ei, Ij),(Ek, Il) represents the probability that the next observed event is Ek and its interval belongs to Il after observing the event Ei and its interval belongs to Ij.

[0165] Initialization: When the system is cold-started, the transition matrix P can be initialized as a uniform distribution matrix, that is, all transition probabilities are equal. Preferably, a more general initial transition matrix can be pre-trained using a large amount of offline collected historical monitoring data.

[0166] Model usage and prediction: Online update: Real-time monitoring of kernel event streams. Whenever a new event (Ek, Il) occurs after an event (Ei, Ij), a state transition observation is made. At this time, the corresponding elements P(Ei, Ij), (Ek, Il) and their rows of the matrix P are updated using the exponential moving average (EMA) algorithm mentioned in the solution.

[0167] Prediction: When the system is currently in state scurr=(Ec, Ic), the row vector corresponding to scurr in the state transition matrix P, i.e. P[scurr,:], gives the complete probability distribution of the next state. By selecting several states with the highest probability values in the row vector, a key event prediction set containing event type, expected time window and confidence (i.e. probability value) can be generated to provide decision support for the dynamic time slice allocation in step S33.

[0168] According to another aspect of the present application, this embodiment also provides for the introduction of a complementary mechanism with an emergency preemption channel, which enables the system to respond instantly to the highest priority events while ensuring the efficiency of conventional scheduling.

[0169] In the system initialization phase, in addition to defining the corresponding intra-group ticket number array and intra-group selection flag array for each probe conflict group, and defining the inter-group ticket number array and inter-group selection flag array for each group representative, this embodiment additionally defines a key global shared data structure: Emergency Status Flag (EmergencyStatusFlag): Define a globally shared integer variable that can be read and written through atomic operations, such as global_emergency_flag. This flag is initialized to 0, indicating that the system is in a normal scheduling state. When its value is 1, it means that the system has entered an emergency state and the preemption channel needs to be activated.

[0170] In normal state (i.e. when global_emergency_flag is 0), the system strictly follows the hierarchical Bakery algorithm for scheduling. The probe first competes for group representative authority in the group to which it belongs through the first-level Bakery algorithm, and the winner then participates in the second-level inter-group Bakery algorithm to compete for the final kernel event processing right. This process is designed to reduce scheduling delays under normal operations.

[0171] When any probe in the system (regardless of the group it belongs to or the current ticket size) identifies a predefined high-risk event feature through its own monitoring logic (for example, by matching with the built-in attack signature library or critical kernel error features), it will immediately perform the preemptive activation operation.

[0172] This probe sets the value of the global global_emergency_flag from 0 to 1 through a single atomic write operation. This operation ensures that all CPU cores can immediately see the change in state.

[0173] For all non-emergency probes that are participating in the polling wait of the Bakery algorithm at all levels, their waiting logic is modified to give the highest priority to checking the emergency status flag bit.

[0174] Modified waiting logic: In the spin-wait logic of a probe, before checking the ticket numbers and selection flag bits of other probes, it first atomically reads the value of global_emergency_flag.

[0175] If the value read is 1, the probe recognizes that the system has entered an emergency state. At this time, it must immediately perform a cooperative avoidance operation: through a single atomic write operation, reset its ticket number value in the corresponding ticket number array to 0 and immediately jump out of the current waiting loop, essentially abandoning the current queuing request.

[0176] If the value read is 0, it continues to execute the original Bakery algorithm waiting logic.

[0177] This mechanism ensures that once an emergency signal is issued, all the probes that are queuing will quickly clear the runway and make way for the execution of the emergency probe.

[0178] The probe that sets the emergency flag bit, after successfully activating the preemptive channel, can bypass all the queuing and waiting logics of the hierarchical Bakery algorithm and immediately obtain exclusive processing rights for kernel events to execute its high-risk event handler (for example, recording a detailed scene, blocking malicious connections, etc.).

[0179] After this emergency probe completes its critical task, it is responsible for restoring the system to its normal state. It will set the value of global_emergency_flag from 1 back to 0 through a single atomic write operation.

[0180] After the flag bit is cleared, those probes that gave up queuing due to avoidance will find that the system has returned to normal in their subsequent retry logic, and thus can start a new round of hierarchical scheduling requests. Thus, the entire system seamlessly switches from the emergency state back to the efficient normal scheduling state.

[0181] Through the above steps, a fusion solution is finally formed in this embodiment: in the normal state, the system realizes efficient and low-latency fair scheduling through the hierarchical Bakery algorithm; while in the emergency state, the instantaneous response to the highest-priority events is ensured through the emergency preemption channel mechanism, solving the entire problem chain and achieving the dynamic balance between the efficiency and response ability of the system.

[0182] According to one aspect of the present application, a method for prospectively predicting the degree of resource competition among eBPF probes is described in detail. The method includes the following steps: Step S101, extract the probe monitoring data of consecutive sampling periods from the real-time performance metrics read from the shared memory to construct a multi-dimensional load feature vector sequence.

[0183] In this embodiment, the multi-dimensional load feature vector is a quantitative description of the resource consumption state of a single eBPF probe at a certain moment, and is used for subsequent dynamic analysis. The dimension of the vector can be configured according to the monitoring fineness requirements.

[0184] Specifically, a preferred three-dimensional load feature vector L(t) includes: 1) normalized trigger frequency; 2) normalized CPU execution time; 3) normalized memory access times. Perform maximum-minimum normalization processing on the original monitoring data of consecutive N sampling periods (for example, N = 16) to make its value fall into the interval [0, 1], forming the multi-dimensional load feature vector sequence {L(t - N + 1),..., L(t)}.

[0185] Step S102, in order to solve the problem of poor adaptability of traditional prediction models in the edge computing dynamic load environment, a dual-mode residual-driven adaptive Kalman filter is used to generate the predicted competition intensity. The prediction residual refers to the difference between the load value predicted by the filter and the actually observed load value at the next moment.

[0186] In the Kalman filter, the process noise covariance matrix Q represents the quantification of the uncertainty of the system model. The larger the Q value, the less trust in the existing model, and thus more dependence on new observation data for correction.

[0187] Initialize a Kalman filter with the state vector x(t) of [load value, speed, acceleration]. At each sampling period, the filter executes a prediction-update loop, compares the predicted value generated in the prediction step with the actually observed load value at the end of this period to obtain the prediction residual e(t), and forms the prediction residual sequence with consecutive residuals.

[0188] Perform real-time analysis on the prediction residual sequence to determine the load dynamic mode.

[0189] When the signs of M consecutive (preferably M = 3) residual values are the same and the absolute value of their mean exceeds K times (preferably K = 2.5) the historical residual standard deviation, the current load dynamic mode is determined to be the step mode.

[0190] If the step mode condition is not met, but the absolute value of a single residual exceeds L times (preferably L = 4) the historical residual standard deviation, it is determined to be the pulse mode.

[0191] According to the determined load dynamic mode, the process noise covariance matrix Q inside the predictor is dynamically adjusted. When it is determined to be the step mode, a dynamic injection operation is performed on the process noise covariance matrix Q. Specifically, the diagonal element values are instantaneously increased by one order of magnitude (for example, multiplied by 10). This operation will significantly increase the Kalman gain of the filter and accelerate the convergence of the predictor to the structural changes of the system. When it is determined to be the pulse mode or the normal mode, the value of the Q matrix is kept stable.

[0192] The filter after adaptive adjustment outputs the predicted load value for the next cycle, and this value is the predicted competition intensity.

[0193] The determination of the pulsed deviation is as follows: |residual(t)| > k1×σ and duration < T1; residual(t) is the predicted residual at time t; k1 is the pulse detection sensitivity coefficient; duration is the deviation duration; T1 is the upper limit of the pulse duration; σ is the standard deviation of the predicted residual.

[0194] The determination of the step deviation is mean(residual[t - w:t]) > k2×σ and duration > T2; residual[t - w:t] is the residual sequence within the time window [t - w, t]; w is the sliding window width; k2 is the step detection sensitivity coefficient; T2 is the lower limit of the step duration.

[0195] In the process noise covariance matrix, Q_new = Q_old × (1 + λ×deviation_score); where deviation_score is calculated according to the deviation type and intensity. λ is a coefficient, Q_new is the updated covariance matrix, and Q_old is the old covariance matrix.

[0196] Step S103, update the probe conflict mapping table using the predicted competition intensity to identify high - risk conflicts in advance.

[0197] In the data structure of the probe conflict mapping table, in addition to recording the current conflict relationship, it also includes a predicted risk level field. The predicted competition intensity generated in step S102 will be used to update this field. For example, when the predicted competition intensity exceeds 150% of the current competition intensity, the risk level of this conflict pair is marked as high, so as to provide forward-looking input for subsequent scheduling decisions.

[0198] Through the above steps, the present invention can be upgraded from simple current state monitoring to accurate prediction of future conflict risks, providing a solid guarantee for the stability and resource utilization rate of the entire system.

[0199] In another embodiment of the present application, steps S102 and S103 may also be: Calculate the first-order difference and the second-order difference based on the multi-dimensional load feature vector sequence to obtain the load change acceleration, and generate the predicted competition intensity.

[0200] For the multi-dimensional load feature vector sequence L(t), calculate its first-order difference (velocity) and second-order difference (acceleration).

[0201] The first-order difference Δfirst(t)=L(t)-L(t - 1); The second-order difference Δsecond(t)=Δfirst(t)-Δfirst(t - 1); where the second-order difference Δsecond(t) is the load change acceleration.

[0202] Adopt a weighted fusion prediction model to fuse the current resource competition degree of the probe, the first-order difference, and the load change acceleration to generate the predicted competition intensity Cpredicted(t + 1) for the next period.

[0203] Preferably, the fusion formula can be: Cpredicted(t + 1)=wc•Ccurrent(t)+wv•Δfirst(t)+wa•Δsecond(t). Ccurrent(t) is the current resource competition degree of the probe at time t. wc, wv, and wa are the fusion weights of the current state, velocity, and acceleration respectively. In a basic implementation, it can be set as wc = 1.0, wv = 1.0, wa = 0.5. The weights represent the extent to which the prediction of the future state depends on the current position, current velocity, and current acceleration.

[0204] This embodiment assumes that the short-term behavior of the system can be approximated by its current state and change trend (velocity and acceleration). The calculation overhead is extremely small and the response is rapid, which is very suitable for real-time short-term trend prediction on resource-constrained edge nodes.

[0205] Preferably, the fusion weights wc, wv, and wa are not fixed values but can be dynamically adjusted according to a pattern recognition mechanism.

[0206] The system can analyze the autocorrelation and variance of the multi-dimensional load feature vector sequence in parallel. When the autocorrelation coefficient is greater than 0.8, it is recognized as a periodic pattern; when the variance exceeds a preset threshold, it is recognized as a burst pattern.

[0207] Adjust the weights according to the recognized load pattern. For example: In the periodic pattern, the load change is highly predictable, and the acceleration weight wa can be appropriately reduced (e.g., reduced to 0.25) to avoid overfitting.

[0208] In the burst pattern, the system state changes drastically, and a quick response is required. The speed weight wv and the acceleration weight wa can be appropriately increased (e.g., increased to 1.2 and 0.75). In this way, the prediction model can adapt to different load dynamics and further improve the prediction accuracy.

[0209] The predicted competition intensity Cpredicted will be used to update a dedicated field in the probe conflict mapping table, such as the prediction risk level. When Cpredicted exceeds a warning threshold (e.g., exceeding 150% of the current competition intensity), the system marks the conflict pair as high risk, thus triggering subsequent more proactive resource scheduling or isolation strategies.

[0210] According to one aspect of the present application, a method for introducing active and intelligent data quality assurance capabilities into a data fingerprint library includes: Step S201, construct a hierarchical verification cost model including various verification costs and depths.

[0211] In this embodiment, the data verification operation is divided into three levels, namely quick verification, standard verification, and in-depth verification, which are respectively: Execute the xxHash64 algorithm on the event core data to generate a hash value for comparison. This method is the fastest and is used for preliminary deduplication.

[0212] Apply the CRC32 algorithm to the key fields of the event to provide higher reliability than quick verification.

[0213] Execute lightweight semantic checks. For example, check whether the IP address is within a legal subnet or whether the port number is a port of a known service. This method has the highest cost but can detect logical errors.

[0214] By offline calibrating the CPU cycle consumption of each level of operation, establish a relative verification cost model. For example, the cost ratio is 1:4:16.

[0215] Step S202: Generate a dynamic verification policy configuration based on the historically verified error rate data and the hierarchical verification cost model. Continuously count the comprehensive error rate within a sliding time window (e.g., the most recent 1000 verifications).

[0216] When the comprehensive error rate exceeds 2 times the preset target value, trigger a policy upgrade. Specifically, increase the execution ratio of in-depth verification to 2 times the original, and at the same time, correspondingly reduce the ratio of quick checks to ensure that the total verification cost is within the budget.

[0217] When the comprehensive error rate is lower than half of the preset target value, trigger a policy downgrade. Specifically, increase the execution ratio of quick checks to 2 times the original, and reduce the ratio of in-depth verification to save system resources. The output of this process is a dynamically adjusted verification level execution ratio configuration [R_fast, R_standard, R_deep], where R_fast is the quick verification ratio, R_standard is the standard verification ratio, and R_deep is the in-depth verification ratio.

[0218] Step S203: Statistically analyze the probe source distribution of error events, identify high-risk probes, and construct an error propagation model.

[0219] The system maintains a statistical table of the error contributions of each probe. When the error rate of a certain probe exceeds 3 times the average error rate of all probes, this probe is marked as a high-risk probe.

[0220] When executing the policy in step S202, for events of probes marked as high-risk probes, the probability of in-depth verification will be multiplied by an amplification factor (e.g., 4 times).

[0221] When an event of a probe is confirmed as an error, the system will temporarily increase the verification intensity of events collected by other probes within 100 milliseconds before and after in the time dimension. For example, upgrade its standard verification to in-depth verification.

[0222] Through the method described in this embodiment, the data fingerprint library is no longer a passive deduplication tool, but an intelligent module with active quality monitoring and adaptive capabilities.

[0223] According to one aspect of the present application, an intelligent resource scheduling method is described in detail for responding to the prediction results of Embodiment 1. It includes: Step S301: Analyze historical resource usage data and establish a resource correlation matrix among various system resources.

[0224] The system collects historical time series data of various resources such as CPU utilization, memory bandwidth, and network IO throughput. Using the Pearson correlation coefficient algorithm, it calculates the correlation between each pair of resources within a sliding window (e.g., the most recent 1000 sampling points) to form the resource correlation matrix. For example, if the correlation coefficient between the CPU and memory bandwidth is greater than 0.7, they are marked as a strongly correlated resource pair.

[0225] Step S302, fuse the predicted competition intensity and the resource correlation matrix, and calculate and generate a collaborative adjustment execution plan.

[0226] The execution plan is the final scheduling instruction. Its input is the predicted competition intensity. The system converts it into the required amounts of specific resources in the future. Combining the resource correlation matrix, the system performs collaborative adjustment planning. For example, if it is predicted that the CPU demand will surge and the CPU and memory bandwidth are strongly correlated, the execution plan will not only include increasing the CPU scheduling priority or binding to large cores, but also reserving or increasing the memory bus frequency. The plan specifies the adjustment order, adjustment amount, and delay time.

[0227] Step S303, after executing the collaborative adjustment execution plan, continuously monitor the key performance indicators and implement the adjustment rollback function.

[0228] Before performing any adjustment operation, the system saves the current key configuration parameters in an adjustment decision snapshot cache with a size of 5. After execution, within a 100 - millisecond monitoring window, it continuously collects metrics such as application throughput and request latency.

[0229] When it is detected that the throughput drops by more than 10% or the latency increases by more than 20 milliseconds, the rollback mechanism is automatically triggered. The system restores the most recent stable configuration from the snapshot cache.

[0230] Step S304, collect and evaluate the performance feedback data of historical adjustment operations to achieve the ability of policy self - evolution.

[0231] The system records the scenario characteristics and final effects of each adjustment operation (including success, failure, and rollback) for training an adjustment risk assessment model.

[0232] Preferably, a lightweight reinforcement learning model can be used. Among them, the system state is the quantified load pattern and prediction trend, the action is the specific adjustment strategy, and the reward is the performance improvement brought by the adjustment. Through continuous learning, the model can continuously optimize the ability to select the best adjustment strategy in specific scenarios and achieve the self - evolution of core parameters.

[0233] According to one aspect of the present application, the process of generating the predicted competition intensity is replaced by a prediction method based on multi - scale analysis, specifically: Apply discrete wavelet transform to the multi-dimensional load feature vector sequence, decompose it into at least one low-frequency approximation component and one high-frequency detail component, and form sub-band signals of multiple different scales.

[0234] Configure an independent predictor for each of the sub-band signals, perform parallel multi-scale prediction, and generate a set of multi-scale prediction results. Fuse and reconstruct the set of multi-scale prediction results to generate the final predicted competition intensity.

[0235] Among them, configuring an independent predictor for each of the sub-band signals specifically means: configuring a dual-mode residual-driven adaptive filter as described in Embodiment 1 for each of the sub-band signals as its independent predictor.

[0236] Among them, fusing and reconstructing the multi-scale prediction results specifically means: calculating the real-time energy of each sub-band signal, and generating a set of normalized energy weights accordingly. Before reconstruction, use the energy weights to weight the set of multi-scale prediction results. Apply inverse discrete wavelet transform to the weighted multi-scale prediction results to complete the reconstruction.

[0237] In some scenarios with extremely high requirements for prediction accuracy, the prediction module in the above embodiment can be replaced by the higher-order multi-scale prediction module described in this embodiment. To solve the problem that a single model cannot take into account multi-scale dynamics. Specifically, the method includes: Step S401, apply discrete wavelet transform (DWT) to the multi-dimensional load feature vector sequence, and decompose it into sub-band signals of multiple different scales.

[0238] In this embodiment, preferably use the Daubechies4 (db4) wavelet basis with higher computational efficiency to perform 3-layer wavelet decomposition on the generated load feature vector sequence. The output of this operation is a set (here 4 sets) of mutually orthogonal sub-band signals {A3, D3, D2, D1}, which respectively represent different scale dynamics from long-term trends to short-term mutations.

[0239] Step S402, configure an independent dual-mode adaptive filter as described in Embodiment 1 for each sub-band signal, and perform parallel multi-scale prediction.

[0240] This step is the core structure of this embodiment. The system creates a filter bank, and instantiates an adaptive Kalman filter for each sub-band signal in {A3, D3, D2, D1}. Each filter independently predicts the sub-band signal input to it and can be specifically configured. For example, the decision threshold of the step mode of the filter processing the low-frequency component A3 can be set to be more sensitive.

[0241] Step S403: Calculate the real-time energy spectrum of each sub-band signal, generate weighting coefficients based on the spectrum, and perform weighted fusion and reconstruction on the prediction results of each sub-scale.

[0242] For each subband signal, calculate its energy E in the most recent time window i =∑|c j ∣ 2 , where c j is the coefficient of the subband signal.

[0243] Normalize the energy of each subband to obtain the energy weight Wi=E of each scale i / ∑E k .

[0244] The sub-scale prediction results output by each parallel filter are multiplied by the corresponding energy weight Wi. Subsequently, the weighted sub-scale prediction results are reconstructed using the inverse discrete wavelet transform (IDWT) to obtain the final prediction competition intensity that integrates the dynamics of all scales.

[0245] Through this embodiment, the system can dynamically focus the prediction attention on the load scale with the most severe current fluctuations, and the output prediction results can be directly used as high-quality input for the collaborative adjustment mechanism in the third embodiment. In summary, in order to solve the runtime conflicts and performance problems caused by uncoordinated concurrent monitoring of multiple eBPF probes in the prior art, the present application first collects and analyzes the running status of all probes in the system, and accurately constructs a probe conflict mapping table by quantitatively calculating the resource competition intensity for the same kernel event. On this basis, the present application designs a two-layer collaborative conflict resolution mechanism. The first layer is based on the pre-filtering of the data fingerprint library. Before requesting sampling, the probe first checks whether the event to be collected has been processed by other conflicting probes in the recent period. If it has been processed, the data is directly reused, which eliminates a large number of unnecessary repeated collections from the root, greatly reducing the probability of conflict. The second layer is for events that must be newly collected, and a distributed Bakery algorithm is introduced to conduct fair arbitration of sampling rights. At any time, only one probe in the conflict group can obtain the collection right, which transforms the original chaotic resource preemption mode into an orderly and deterministic queuing mode, thereby completely avoiding the system performance jitter and data storm caused by probe fighting, and ensuring the stability and fairness of the monitoring system.

[0246] To solve the dilemma in the prior art of having difficulty in choosing between comprehensive monitoring coverage and system overhead, a lightweight Markov chain model is used to learn the time interval distribution of the historical event stream, so as to be able to prospectively predict the event bursts or key anomalies that may occur within the future time window. Based on this prediction result, the time slice resources allocated to each probe are dynamically adjusted. When the system is running stably, the probes work in a low-frequency interleaved mode to save overhead; while when it is predicted that a key event is about to occur, the sampling time window of the relevant probes is temporarily extended for intensive sampling to ensure that key information is not lost. This intelligent scheduling breaks the limitation of the static sampling rate. Furthermore, to make up for the information incompleteness caused by sparse sampling, a signal reconstruction technique based on compressed sensing is introduced, which can highly probably recover the complete monitoring data time series from limited and sparse sampling data points. This application can run with extremely low system overhead for most of the time, and when necessary, it can increase the monitoring granularity as needed, achieving the technical effect of obtaining high-fidelity panoramic data with low-cost sparse sampling and perfectly solving the contradiction between coverage rate and system overhead.

[0247] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.

Claims

1. An observable monitoring method for an edge computing network, characterized in that, include: Collect the running status data of multiple eBPF probes on the edge device, identify the monitoring conflicts for the same kernel event, and generate a probe conflict mapping table; Read real-time event requests, coordinate sampling permissions between probes in combination with the probe conflict mapping table, and generate probe sampling scheduling sequences and data fingerprint libraries; Predict key event distribution based on historical event patterns and probe sampling scheduling sequences, adjust time allocation, and form a time slice allocation table; Perform distributed sampling on the kernel event stream according to the time slice allocation table, query the data fingerprint library and adjust the sampling strategy accordingly to generate a sampling data set; Generate complete monitoring data based on the sampled data set and perform pattern matching and fault location based on it.

2. The method according to claim 1, characterized in that Collect the running status data of multiple eBPF probes on the edge device, identify the monitoring conflicts for the same kernel event, and generate a probe conflict mapping table including: Build a probe registry; For each active probe in the probe registry, read the real-time performance indicators of each probe from the shared memory and analyze the monitoring of multiple probes on the same kernel event; When multiple probes are detected to be mounted on the same mount point, the resource contention degree is calculated, and a probe conflict mapping table including conflicting probe pairs, conflicting event types, and contention intensity is generated.

3. The method according to claim 1, characterized in that Combined with the probe conflict mapping table, the sampling authority between probes is coordinated, including the judgment of the competition relationship between probes, specifically: Monitor the sampling requests in the real-time event request queue, and determine whether the probe initiating the request has resource competition according to the probe conflict mapping table; For competing probes, the hash value of the event to be collected is first calculated to check whether other probes have collected the same data. If so, the data is reused directly; otherwise, the data is collected again.

4. The method according to claim 3, wherein The specific process of re-collecting and generating the probe sampling scheduling sequence and data fingerprint library is as follows: Probes in the same conflict group obtain sampling permissions through a distributed mutual exclusion algorithm. The probes that have obtained the permission perform data collection and generate a probe sampling scheduling sequence according to the sampling order; After the collection is completed, the hash value and timestamp of the event data are stored in the shared storage to form a data fingerprint library.

5. The method according to claim 4, wherein The sampling permission is obtained through a distributed mutual exclusion algorithm, including: When the probe receives a sampling request, it sets its own initial selection array position; Traverse the ticket number arrays of all other probes in the conflict group, obtain the current ticket number of each probe, find the maximum value and add 1 as your own ticket number, and update it to your own selection array position; Two levels of checks are performed on each conflicting probe: spin-wait to ensure that other probes have completed ticket number selection, compare ticket numbers, and continue waiting if other probes have smaller ticket numbers or the ticket numbers are the same but the probe ID is smaller, until the highest priority sampling permission is obtained.

6. The method according to claim 4, characterized in that, Store the hash value and timestamp of the event data in shared storage, set the validity period and clean it up regularly to form a data fingerprint library including: Perform two-level hash calculation on the event data, use the first hash algorithm to calculate a 64-bit hash value for the event type, timestamp and key parameters, combine this hash value with the probe identifier and use the second hash algorithm to calculate a 32-bit fingerprint to obtain the event data fingerprint; Store the event data fingerprints in a hash table structure to form a fingerprint storage index, and perform addition, deletion, and maintenance according to the activity to obtain a data fingerprint library.

7. The method according to claim 1, wherein Predict the critical event distribution based on historical event patterns and probe sampling scheduling sequences, adjust the time allocation, and form a time slice allocation table, including: Analyze the historical sampling patterns of each probe based on the probe sampling scheduling sequence, extract the event interval distribution from the historical event pattern data, and construct event timing characteristics; Use the event timing characteristics to train a prediction model, update the state transition probability in real time, and predict high-frequency events within a future time window when a burst pattern is detected to obtain a critical event prediction set; Divide the monitoring window into multiple time slices. Under normal circumstances, allocate staggered time slices to the probes. Identify the time periods that require intensive sampling according to the critical event prediction set, and extend the time slices of relevant probes during this time period to generate a time slice allocation table.

8. The method according to claim 7, characterized in that Obtain a critical event prediction set, including: Read the event interval data in the event timing characteristics, map it to a predefined quantization interval through logarithmic transformation, and generate a quantized event interval sequence; Based on the quantized event interval sequence, update the state transition matrix of the Markov chain and calculate a new state transition matrix; Monitor the number of consecutive occurrences of the same interval in the quantized event interval sequence. When the same interval appears continuously for a preset number of times, trigger burst detection, calculate the burst intensity, and mark it as an emergency event; Combine the new state transition matrix and the emergency event mark, calculate the probability distribution of each event type within a future time window, improve the prediction probability of emergency events, and generate a critical event prediction set.

9. The method according to claim 7, wherein According to the critical event prediction set, calculate the time slice allocation table for each probe, including: Divide the monitoring window into a fixed number of basic time slices according to the number of active probes, and generate an initial bit mask with evenly staggered time slices for each probe; When the critical event prediction set indicates that there are high-priority events in a certain time window, identify the probes responsible for monitoring this event, extend the mask bits of these probes within the prediction window, and detect whether there are conflicts; If there is a conflict, perform priority arbitration according to the event prediction confidence to ensure reasonable allocation of each time slice and output the time slice allocation table.

10. The method according to claim 2, characterized in that, The process of updating the probe conflict mapping table is specifically as follows: Continuously monitor the deviation between the predicted competition intensity and the subsequent actually observed load value to form a prediction residual sequence; Analyze the timing distribution characteristics of the prediction residual sequence, distinguish isolated pulse-like deviations from continuous step-like deviations, and determine the load dynamic pattern accordingly; According to the determined load dynamic pattern, dynamically adjust the process noise covariance matrix inside the predictor; When it is determined to be a step-like deviation, increase the value of the process noise covariance matrix.

11. The method according to claim 2, characterized in that, The process of updating the probe conflict mapping table is specifically as follows: Extract the probe monitoring data of consecutive sampling periods from the real-time performance metrics read from the shared memory to construct a multi-dimensional load feature vector sequence; Calculate the first-order difference and second-order difference of the multi-dimensional load feature vector sequence to obtain the load change acceleration; Fuse the current resource competition degree, first-order difference, and load change acceleration of the probe to generate a predicted competition intensity; Use the predicted competition intensity to update the probe conflict mapping table.

Citation Information

Patent Citations

  • Distributed system performance tracking method based on eBPF and SkyWalking technologies

    CN117950953A

  • Streaming data parallel query optimization method and system

    CN118394787A

  • Business request exception delimitation analysis method and system, electronic equipment and storage medium

    CN119088612A

Cited By

  • Efficient acquisition and processing system for safety production monitoring data of electric power steel structure

    CN121056487A

  • Internet of Things equipment data storage method and device based on Markov chain, and medium

    CN121356747A

  • QoS regulation and control method and system based on IO demand prediction

    CN122332135A

  • QoS regulation method and system based on IO demand prediction

    CN122332135B