Distributed collaborative learning and elastic response method, system and device based on eBPF and medium

By constructing a dual-loop collaborative architecture of user-mode learning and kernel-mode detection response in a cloud-native environment, the problems of rule rigidity and detection lag in cloud-native environments are solved, achieving real-time closed-loop security protection and improving real-time performance and adaptability.

CN121706083APending Publication Date: 2026-03-20GUANGZHOU ELECTRIC POWER COMM NETWORK LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511521392.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing technologies struggle to address dynamic threats in cloud-native environments, exhibiting issues such as rigid rules, delayed detection, and fragmented responses. They fail to achieve real-time closed-loop processing, resulting in an inability to effectively address dynamically changing threats in cloud-native environments.

Method used

By constructing a dual-loop collaborative architecture of user-mode learning and kernel-mode detection and response, a normal behavior baseline model is built in user mode and then synchronized to kernel mode after being lightweighted. In kernel mode, microsecond-level behavior deviation detection and risk assessment are performed, and elastic response decision-making and execution are achieved in combination with asset criticality.

Benefits of technology

It achieves real-time closed-loop security protection in cloud-native environments, improving real-time performance, adaptability, and protection accuracy, and has the ability to handle unknown threats in real time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121706083A_ABST
    Figure CN121706083A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of distributed collaborative learning and elastic response, in particular to an eBPF-based distributed collaborative learning and elastic response method, system, equipment and medium, which comprises the following steps of: performing global analysis in a user mode based on system behavior data acquired by a kernel mode eBPF program to construct a normal behavior baseline model; the baseline model is synchronized to a kernel mode after being lightened; in the kernel mode, deviation detection is carried out on system behaviors generated in real time based on the synchronous baseline model, and a risk score is output; determining a response level in combination with the risk score and a preset asset criticality; and in a kernel mode, executing a response action corresponding to the response level. The method has the beneficial effects that the inherent contradiction between the response speed and the detection precision in the prior art is successfully solved by constructing the fast and slow double-loop collaborative security system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of distributed collaborative learning and elastic response, and particularly relates to a distributed collaborative learning and elastic response method, system, device and medium based on eBPF. BACKGROUND

[0002] With the popularization of cloud computing technology, cloud native architecture has become the mainstream paradigm of modern application development, and its core is to build large-scale distributed systems through containerized microservices. In this context, eBPF technology has been widely used because it can provide non-invasive and high-performance system observability in the kernel state. Currently, runtime security solutions based on eBPF mainly fall into two categories: one is represented by Falco, which relies on pre-defined static rules for detection; the other is an anomaly detection scheme with offline machine learning as the core. These schemes provide a preliminary technical path for the security monitoring of cloud native environments.

[0003] Although existing technologies have made some progress in runtime security, they still have significant limitations in dealing with the dynamics and complexity of cloud native environments: on the one hand, static rule matching schemes have the risk of rule rigidity and bypassing, and cannot adapt to unknown threats and zero-day attacks; on the other hand, offline learning schemes have difficulty meeting the real-time response requirements of seconds or even sub-seconds due to the long data processing link. In addition, existing schemes generally face the problem of detection and response fragmentation, and cannot achieve immediate closed-loop in the kernel state, resulting in a potential attack window period. These problems make it difficult for existing technologies to effectively deal with dynamically changing threats in cloud native environments. SUMMARY

[0004] To solve the above technical problems, the present application provides the following technical solutions: In a first aspect, the present application provides a distributed collaborative learning and elastic response method based on eBPF, comprising performing global analysis in the user state based on system behavior data collected by the kernel state eBPF program to construct a normal behavior baseline model, and synchronizing the baseline model to the kernel state after lightening it. In the kernel state, performing deviation detection on the system behavior generated in real time based on the synchronized baseline model, and outputting a risk score. Determining a response level in combination with the risk score and a preset asset criticality. In the kernel state, performing a response action corresponding to the response level.

[0005] As a preferred scheme of the distributed collaborative learning and elastic response method based on eBPF of the present application, wherein: performing global analysis in the user state based on system behavior data collected by the kernel state eBPF program to construct a normal behavior baseline model, and synchronizing the baseline model to the kernel state after lightening it, comprises: In the kernel state, system behavior events are captured by eBPF programs, and features are extracted to generate behavior feature vectors; In the user state central analyzer, behavior feature vectors from multiple nodes are aggregated, and an online learning algorithm is used to build a normal behavior baseline model; The normal behavior baseline model is distilled into lightweight detection rules and incrementally synchronized to the kernel state eBPF storage mapping of each node.

[0006] As a preferred scheme of the eBPF-based distributed collaborative learning and elastic response method of the present application, wherein: deviation detection is performed on real-time generated system behavior, and a risk score is output, including, Calculate the instantaneous risk score of the current system behavior, wherein the instantaneous risk score at least integrates the global rarity of the behavior and the unexpectedness of the behavior in the learned behavior sequence context; The instantaneous risk score of the current system behavior is recursively combined with the historical risk score of the process, to generate a final risk score and output.

[0007] As a preferred scheme of the eBPF-based distributed collaborative learning and elastic response method of the present application, wherein: before performing a response action corresponding to a response level, including, In the kernel state, the number of times each process triggers a response action in a unit of time is recorded through an eBPF storage mapping structure; When the number of times exceeds a preset threshold, the execution of non-terminating response actions on the process is limited.

[0008] As a preferred scheme of the eBPF-based distributed collaborative learning and elastic response method of the present application, wherein: the online learning algorithm is a hidden Markov model.

[0009] As a preferred scheme of the eBPF-based distributed collaborative learning and elastic response method of the present application, wherein: the baseline model is synchronized to the kernel state after being lightened, including, The high-probability state transition path in the hidden Markov model is converted into a key-value pair rule; Wherein, the key of the key-value pair rule is the identifier of the predecessor behavior event, and the value is the compressed representation of the set of allowed successor behavior event identifiers.

[0010] As a preferred scheme of the eBPF-based distributed collaborative learning and elastic response method of the present application, wherein: the compressed representation is a Bloom filter.

[0011] In a second aspect, the present application provides an eBPF-based distributed collaborative learning and elastic response system, comprising: a synchronization module, configured to perform global analysis in a user mode based on system behavior data collected by a kernel-mode eBPF program, to construct a normal behavior baseline model, and to synchronize the baseline model to the kernel mode after lightening the baseline model; a detection output module, configured to perform deviation detection on real-time generated system behavior based on the synchronized baseline model in the kernel mode, and to output a risk score; a determination module, configured to determine a response level in combination with the risk score and a preset asset criticality; an execution module, configured to execute a response action corresponding to the response level in the kernel mode.

[0012] In a third aspect, the present application provides a computer device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0013] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the above method when executed by a processor.

[0014] Compared with the prior art, the present application has the following beneficial effects: by constructing a fast-slow dual-loop collaborative security system, the inherent contradiction between response speed and detection accuracy in the prior art is successfully solved. Among them, the kernel-mode real-time response closed loop (fast loop) ensures the microsecond-level immediate disposal capability for unknown threats, and the global intelligent learning closed loop (slow loop) realizes the continuous self-evolution of the security baseline. The two work collaboratively, so that the system not only has the extreme real-time and precise self-adaptive ability that the traditional scheme cannot have in the cloud native dynamic environment, but also fundamentally improves the protection level of runtime security. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0016] Figure 1 It is a system framework schematic diagram.

[0017] Figure 2 It is a flowchart.

[0018] Figure 3 It is a fast-slow dual-loop security system schematic diagram. DETAILED DESCRIPTION

[0019] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the present application.

[0020] Embodiment 1, refer to Figure 1 For the first embodiment of the present application, the embodiment provides an eBPF-based distributed collaborative learning and elastic response method, comprising: S100: Based on the system behavior data collected by the kernel eBPF program, global analysis is performed in the user state to construct a normal behavior baseline model, and the baseline model is lightened and synchronized to the kernel state; S200: In the kernel state, based on the synchronized baseline model, the deviation of the real-time generated system behavior is detected, and the risk score is output; S300: Combined with the risk score and the preset asset criticality, the response level is determined; S400: In the kernel state, the response action corresponding to the response level is executed.

[0021] It should be noted that in the cloud native architecture, the life cycle of the microservice instance is short and dynamic, and the east-west communication generally adopts encrypted transmission, which makes the traditional runtime security scheme based on static rules or offline analysis face serious challenges. Specifically, the static rule scheme cannot adapt to the dynamic environment and is difficult to detect unknown threats, while the offline learning scheme will cause a delay of minutes or even hours between detection and response due to the link delay of data collection, transmission and analysis, which cannot meet the real-time security protection requirements. Further, the detection module and the response module of the existing scheme are usually in a fragmented state, and it is difficult to realize the instant closed loop of "detection-decision-response" at the kernel level, resulting in a obvious response window period of security protection.

[0022] Therefore, in view of the above-mentioned problems of rigid rules, detection lag and response fragmentation, the present application constructs a user state learning-kernel state detection response double-loop collaborative architecture through the steps of S100-S400: that is, in the user state, an evolvable behavior baseline is constructed through global data analysis, and it is lightened and synchronized to the kernel; in the kernel state, based on the baseline, microsecond-level behavior deviation detection and risk assessment are realized, and combined with the asset criticality, graded elastic response decision and execution are realized, so as to complete the security closed loop from threat perception to instant disposal at the kernel level, effectively improving the real-time, adaptability and protection accuracy of runtime security in the cloud native environment.

[0023] Embodiment 2, refer toFigures 1-3 This is one embodiment of the present invention. Based on the above embodiment, a distributed collaborative learning and elastic response method based on eBPF is provided.

[0024] In this embodiment of the application, step S100 involves performing a global analysis in user space based on system behavior data collected by the kernel-mode eBPF program to construct a normal behavior baseline model, and then synchronizing the lightweight baseline model to the kernel mode, including the following steps A1-A3: A1: In kernel mode, system behavior events are captured by the eBPF program, and features are extracted to generate behavior feature vectors.

[0025] Understandably, extracting features to generate behavioral feature vectors aims to transform raw, unstructured kernel events (e.g., a long list of system call parameters) into structured, computable data. The specific extraction process involves the eBPF program attaching to specific kernel function entry or exit points to read raw parameters from the function context (e.g., the struct pt_regs register set). Further, to safely read string-type parameters (e.g., filenames) from user space, the helper function bpf_probe_read_user_str() is called to perform on-the-fly hashing of the read content. For network events, the program extracts the connection's four-tuple information (source / destination IP, source / destination port) from kernel network data structures (e.g., struct sock*). Furthermore, to reduce the amount of data reported to user space, this method implements efficient data aggregation in kernel space: using eBPF Maps (e.g., BPF_MAP_TYPE_HASH or BPF_MAP_TYPE_LRU_HASH), the eBPF program can perform real-time counting or frequency statistics of similar events occurring within a specific time window (e.g., within 1 second) in the kernel. For example, it is possible to count how often a process executes a specific system call.

[0026] Preferably, the pre-aggregation performed in the kernel can significantly reduce the amount of data that needs to be passed to user space via BPF_RINGBUF, thereby optimizing overall performance.

[0027] It should be noted that the structure of the behavioral feature vector is explicitly defined, and its components include, but are not limited to, the following fields: event_type: An enumeration of event types used to distinguish different behaviors, such as system calls (e.g., SYSCALL_EXECVE) or network connections (e.g., NET_CONNECT_IPV4).

[0028] syscall_id: The system call number for the system call event.

[0029] process_info: Contains key process identifiers such as TGID (Thread Group ID), PID (Process ID), and process name hash value.

[0030] `arg_features`: Provides characteristic representations of key parameters in system calls or functions. To avoid processing and storing large amounts of raw strings in the kernel, this approach uses methods such as hashing or pattern matching. For example, it performs hash calculations on filenames or IP address parameters, or uses pattern matching to determine whether an IP address belongs to a private network segment.

[0031] `process_lineage_hash`: Process lineage hash. This field contains the hash value of the combination of the current process's parent process ID (PPID) and session ID. This is crucial for accurately identifying anomalous process lineages. For example, a normal web service process (such as Nginx) should not spawn a shell process (such as ` / bin / bash`). Such threats can be easily detected through lineage analysis.

[0032] `file_attributes`: File attribute characteristics. For file access events, this field contains not only a hash of the file path but also encoded characteristic values ​​of file metadata (such as file permission mode, owner UID / GID, last modified time, etc.). This effectively detects unauthorized privilege escalation (such as a regular user process attempting to modify a configuration file with root privileges) or tampering with critical files.

[0033] `net_packet_metadata`: Network packet metadata. For network connection events, in addition to the connection four-tuple information, this field also contains meta-features extracted from the TCP / IP header. For example, analyzing the randomness of the Initial TCP Sequence Number (ISN) can identify potential TCP session hijacking or network scanning behavior; fingerprinting of options such as TCP window size and MSS (Maximum Segment Size) helps identify specific scanning tools or malware families.

[0034] return_value: The return value of the system call, used to determine whether the operation was successful.

[0035] In an optional implementation, the feature extraction in step A1 to generate the behavioral feature vector can also be based on container runtime interface feature extraction. That is, the container runtime interface captures container lifecycle events (such as start, stop, restart) and resource usage (such as CPU, memory, network bandwidth), and uses the event timestamp, container ID, Pod name, resource indicator sampling results and their changing trends as features to generate a behavioral feature vector for subsequent behavior analysis and anomaly detection.

[0036] In another optional implementation, the feature extraction in step A1 to generate the behavioral feature vector can also be achieved by using a user-space agent-assisted feature extraction method. That is, a lightweight user-space agent is deployed within the node, and ptrace or libseccomp is used to intercept process system calls and analyze their type, parameters, frequency, etc. At the same time, inotify or fanotify is used to monitor file system operations, extract information such as file path, type, and operation mode, and construct a behavioral feature vector to achieve feature extraction and monitoring of system behavior.

[0037] A2: In the central analyzer of the user space, behavioral feature vectors from multiple nodes are aggregated, and an online learning algorithm is used to build a baseline model of normal behavior.

[0038] Understandably, the purpose of building a normal behavior baseline model is not to find malicious features, but to define normal behavior patterns, that is, any behavior that deviates from the normal pattern is considered abnormal.

[0039] It should be noted that behavioral feature vectors are efficiently transmitted to the node agents in user space via BPF_RINGBUF. The agent is responsible for the initial aggregation and compression of these vectors, and, combined with the local node's Container Runtime Interface (CRI) information, injects key local context into this behavioral data, such as container IDs and Pod names. This is then asynchronously reported to the cluster-level central analyzer. Furthermore, the node agent reports the processed data stream to the system's cluster-level central analyzer. The analyzer, based on metadata such as Pod tags, aggregates data streams from different nodes belonging to the same logical service (e.g., payment-service). Then, its built-in behavioral model digital twin unit, based on this, uses the aggregated full-cluster data to run more complex online learning algorithms (such as Hidden Markov Models (HMM), incremental clustering, etc.), thereby performing a global, long-term model of the normal behavior of a distributed service.

[0040] In an optional implementation, the construction of the normal behavior baseline model using an online learning algorithm in step A2 can also employ a real-time feature aggregation method based on online clustering. Specifically, it can use online clustering algorithms such as Mini-Batch K-Means or StreamKM++ to perform real-time clustering of the behavioral feature vectors collected by the kernel-mode eBPF program. Each time a new feature vector is collected, the clustering model is dynamically updated, adjusting the number and position of the clusters. The normal behavior baseline model is constructed using the cluster centers and radii, and the key parameters of the model are synchronized to the kernel-mode eBPF storage mapping, achieving real-time aggregation and modeling of behavioral features.

[0041] In another optional implementation, the online learning algorithm used to construct the normal behavior baseline model in step A2 can also employ a distributed baseline update method based on stochastic gradient descent. Specifically, the stochastic gradient descent (SGD) algorithm is used to incrementally learn the behavior feature vectors in the user-state central analyzer, constructing and updating the parameters of the linear or neural network model. The updated model parameters are encoded into lightweight detection rules and distributed to the kernel-state eBPF storage mapping of each node via an incremental synchronization mechanism. The kernel uses the updated model parameters to evaluate the behavior feature vectors in real time, calculates their deviation from the normal behavior baseline, and thus detects abnormal behavior.

[0042] A3: Distill the normal behavior baseline model into lightweight detection rules and incrementally synchronize them to the kernel-mode eBPF storage mapping on each node.

[0043] It should be noted that when the online learning algorithm uses a Hidden Markov Model (HMM), synchronizing the baseline model to the kernel state after lightweighting means converting the high-probability state transition paths in the Hidden Markov Model into key-value pair rules. In these rules, the key is the identifier of the preceding behavioral event, and the value is a compressed representation of the set of allowed subsequent behavioral event identifiers. The further compressed representation is a Bloom filter.

[0044] Specifically, the macroscopic operational phases of a service (such as initialization, port listening, and request processing) are defined as the hidden states of the Hidden Model (HMM). On one hand, to achieve automated state identification and model self-evolution, incremental density clustering algorithms (such as streaming DBSCAN) are used to cluster continuously incoming behavioral feature vectors online. Each cluster represents a cohesive behavioral pattern and is defined as a hidden state of the HMM. On the other hand, to ensure that the baseline model can evolve synchronously with service version iterations or functional changes, a dynamic state adjustment mechanism is introduced: when the sample variance within a certain state (cluster) is too large, state splitting is triggered, refining it into multiple more precise new states; conversely, when the centers of different states are too close, state merging is triggered to simplify the model and prevent overfitting. Ideally, through this mechanism, the state space of the HMM can not only achieve dynamic optimization but also accurately characterize the actual operational phases of the service. That is, when the observed behavioral sequence switches between different states in real time, the system can infer the occurrence of state transitions.

[0045] Furthermore, the observations of the HMM are directly defined as "behavioral feature vectors" captured from the kernel state. The central analyzer continuously aggregates behavioral sequences from each node and trains the model using online learning algorithms (such as incremental expectation-maximization). For a newly arriving behavioral sequence, the algorithm calculates its most probable hidden state path and makes small, incremental adjustments to the core parameters of the HMM—the state transition probability matrix (A) and the observation emission probability matrix (B)—based on this path. Ideally, this process does not require full retraining of historical data, thus achieving smooth evolution and continuous self-learning of model parameters, enabling the behavioral baseline to adaptively serve the normal evolution of behavior.

[0046] Furthermore, to apply the complex Hidden Markov Model (HMM) to real-time kernel-level detection, it must be transformed into a form that can be efficiently executed by the eBPF program. Specifically, this involves converting the behavioral sequence knowledge inherent in the HMM into deterministic key-value pair rules. The process is as follows: traverse the high-probability state transition paths in the HMM; for each state S, extract its high-probability observation set (i.e., common behaviors); for each behavioral feature vector in this observation set (calculate its hash value as the key), encode the set of hash values ​​of the subsequent behavioral feature vectors corresponding to all its high-probability successor states into a Bloom filter, which serves as the value. Finally, the generated...<Key, Value> The set of key-value pairs is a lightweight detection rule.

[0047] Furthermore, to minimize synchronization overhead, a differentiated synchronization mechanism is employed. The central analyzer compares the new rule set with the currently effective old version, generating a patch file containing only add, delete, and modify commands. Upon receiving this patch, the node agent atomically and precisely updates the eBPF Maps in the kernel via BPF system calls.

[0048] Through the above process, a closed-loop empowerment is achieved, from user-space intelligence to kernel-space awareness.

[0049] In an optional implementation, step A3, which distills the normal behavior baseline model into lightweight detection rules, can also employ a decision tree-based rule distillation approach. This involves analyzing the normal behavior baseline model using a decision tree algorithm to extract key decision paths and node conditions from the decision tree as detection rules. These rules are then converted into simple conditional statements, such as "if-else" structures, and encoded into lightweight detection rules, which are synchronized to the kernel-mode eBPF storage mapping of each node. In kernel mode, matching these conditional statements quickly evaluates whether the behavior feature vector conforms to the normal behavior baseline.

[0050] In another optional implementation, the distillation of the normal behavior baseline model into lightweight detection rules in step A3 can also employ a Naive Bayes-based rule distillation method. This involves using a Naive Bayes classifier to distill the normal behavior baseline model, calculating the conditional probability distribution of each feature vector dimension under normal behavior. These conditional probability distributions are then simplified into discrete probability intervals or threshold ranges, forming lightweight detection rules. These rules are synchronized to the kernel-state eBPF storage mapping of each node. In the kernel state, the probability values ​​of each dimension of the feature vector are calculated and compared with preset threshold ranges to determine whether the behavior conforms to the normal behavior baseline.

[0051] In this embodiment of the application, step S200, in kernel mode, performs deviation detection on the real-time generated system behavior based on the synchronized baseline model and outputs a risk score, including the following steps B1-B2: B1: Calculate the instantaneous risk score of the current system behavior, where the instantaneous risk score combines at least the global rarity of the behavior and the unexpectedness of the behavior in the context of the learned behavior sequence.

[0052] Understandably, global rarity is used to quantify the absolute frequency of an event across the entire cluster. An event that occurs with extremely low frequency in the global data will receive a high rarity score. Unexpectedness (judged based on the model trained by S100, such as HMM) is used to quantify the irrationality of an event in the context of a specific behavioral sequence. An event that deviates significantly from the learned normal behavioral sequence will receive a high unexpectedness score.

[0053] Specifically, rarity The calculation formula is as follows: In the formula: Global rarity; For the event Frequency of occurrence in global data; It is a small constant set to prevent division by zero errors, where The value is obtained by the eBPF program by consulting the eBPF Map generated in step S100, which stores the global event frequencies.

[0054] Regarding the degree of surprise The calculation method is as follows: the eBPF program uses... The feature hash is used as the key to look up the corresponding set of successor events in the eBPF Map (such as a Bloom filter). If the current event... If the characteristic hash does not exist in the set, it is considered an unexpected event, and a value is assigned. =1; otherwise, =0.

[0055] Furthermore, the formula for calculating the instantaneous risk score is as follows: In the formula: The synergistic risk factor is a core innovation of the model. This interaction term quantifies the synergistic amplification effect that occurs when the attributes of rarity and unexpectedness occur simultaneously. An event that is both rare and unexpected carries a risk far greater than the sum of the individual risks of the two attributes, and this term accurately captures such high-risk signals.

[0056] This is a context fluctuation factor, which reflects the process to which the event belongs. The stability of behavior over a recent period is determined based on the recent anomaly history maintained by the kernel for that process. The specific calculation and update mechanism is as follows: the kernel maintains an exponential moving average (EMA) of the recent anomaly risk score for each process (keyed by TGID) in a dedicated eBPF Map. Whenever a new instantaneous risk score is obtained... When this occurs, the EMA value corresponding to the process will be updated accordingly. It is then defined as a function of the EMA value, for example ,in It is a configurable amplification factor. Thus, a historically behaving process (low EMA value) will have a significantly higher EMA value. A value close to 1 can effectively suppress potential noise interference; conversely, a process that has recently exhibited frequent anomalies (high EMA value) has... The value will be significantly greater than 1, thus amplifying the score of any newly occurring anomalies. To prevent processes from being permanently labeled "high-risk" due to historical anomalies, leading to a continuous amplification of false positives for subsequent normal behavior, this method introduces a risk decay mechanism for the calculation of the context fluctuation factor. This mechanism is implemented in the kernel, specifically: for a process that has not generated any instantaneous risk score (InstScore) exceeding a preset low threshold within a specific time window (e.g., the past hour), its corresponding EMA value will no longer remain unchanged, but will be multiplied by a decay coefficient δ (e.g., 0.999) slightly less than 1. This means that for a process that has previously exhibited abnormal behavior, if its subsequent behavior continues to remain normal, its historical risk accumulation value will be gradually forgotten or cooled over time, thus reducing its risk. The value smoothly regresses to the baseline level of 1, thereby improving the robustness of the model and enabling it to evaluate each behavioral event more fairly and dynamically.

[0057] , , Risk weight coefficients: These are configurable parameters issued by the central analyzer based on the global situation, used to adjust the model's sensitivity. To achieve intelligent and adaptive optimization of the weight coefficients, the central analyzer integrates a reinforcement learning from human feedback (RLHF) module. Its workflow is as follows: Feedback collection: After the system generates an alarm and reports it to the user space, security operations personnel can analyze the alarm and mark it as "True Positive", "False Positive" or "needs further observation".

[0058] Rewards and Penalties: These manually labeled events will serve as feedback signals to the reinforcement learning module. A sequence of events labeled as a "real attack" will generate a positive reward signal, while a "benign false positive" will generate a negative penalty signal.

[0059] Policy optimization: Based on these reward and punishment signals, the reinforcement learning module aims to maximize the detection rate of "real attacks" and minimize the "benign false positives" rate by automatically and incrementally fine-tuning the risk weight coefficients and the memory weight α (hereinafter referred to as the memory weight) through algorithms such as policy gradient.

[0060] Through this closed-loop feedback mechanism, risk assessment can continuously learn from real-world offensive and defensive confrontations, and its sensitivity and accuracy will continue to improve over time.

[0061] In an optional implementation, step B1, which calculates the instantaneous risk score of the current system behavior, can also employ a decision tree-based rule-matching risk assessment method. This involves using a decision tree model to evaluate the system behavior feature vector. Based on the input feature vector, matching is performed along the branches of the decision tree until a leaf node is reached. The risk score stored in the leaf node is the instantaneous risk score of the current system behavior.

[0062] In another optional implementation, step B1, which calculates the instantaneous risk score of the current system behavior, can also employ a probabilistic risk assessment method based on Naive Bayes. This involves using a Naive Bayes classifier to calculate the probability of the system behavior feature vector. Based on the trained model, the probability that the current behavior feature vector belongs to normal or abnormal behavior is calculated. By setting a threshold, when the probability of abnormal behavior exceeds this threshold, the abnormal probability is used as the instantaneous risk score of the current system behavior.

[0063] B2: The instantaneous risk score of the current system behavior is weighted and recursively combined with the historical risk score of the process to generate the final risk score and output it.

[0064] Understandably, in order to address the problem that instantaneous assessments are insensitive to "slow attacks" (a slow attack chain consisting of a series of low-risk events), a time-series risk accumulation mechanism has also been introduced.

[0065] It should be noted that the final risk score is calculated using the following formula: In the formula: Historical risk score: Represents the final score of the previous event for this process. This value is cached in a dedicated eBPF Map, realizing process-level risk memory.

[0066] α is the memory weight: a coefficient between 0 and 1 used to control the decay rate of historical risks, and determines the model's perception length and sensitivity to attack chains.

[0067] In an alternative implementation, the method for obtaining the pipeline hazard point in step S200 can also be by analyzing the interaction between the temperature field and the stress field, such as the thermal stress concentration effect, and combining the material's thermal expansion coefficient and constraint conditions to calculate the thermal stress distribution, that is: simulating the superposition effect of the temperature field gradient and mechanical stress to identify the region of stress abrupt change under the thermo-mechanical coupling effect.

[0068] In another alternative implementation, the method for obtaining the pipeline hazard point in step S200 can also be by simulating the dynamic loads in the actual operation of the pipeline, such as pressure fluctuations and transient temperature changes, and combining fatigue damage accumulation theory, such as Miner's rule, to identify the area with the fastest damage accumulation rate and determine the hazard point.

[0069] In this embodiment of the application, step S300 combines risk scoring and preset asset criticality to determine the response level, including the following steps C1-C2: C1: Based on a dynamic risk decision matrix, risk scores and asset criticality CoA are used as inputs and mapped to a predefined response level ResponseLevel; Specifically, the functional relationship can be formalized as: C2: Asset criticality is obtained from the metadata of the cloud-native orchestration platform and synchronized to the kernel space through the node agent.

[0070] It should be noted that, in order to ensure that decisions are accurately aligned with the business context, this method defines and designs the CoA as follows: First, at the definition level, asset criticality is seamlessly integrated with the metadata of cloud-native orchestration platforms (such as Kubernetes). This means users can intuitively define the security level of workloads by adding standard tags or annotations based on business logic. For example, a Pod for core payment operations can be marked with `security.asset.criticality=high`, while a regular log processing Pod can be marked with `security.asset.criticality=low`. Even better, this design allows security policy configuration to be naturally integrated into existing operations and development processes, achieving declarative management of security and business attributes.

[0071] Secondly, at the synchronization level, this method implements a real-time, precise kernel-mode synchronization pipeline through node agents. Furthermore, to ensure stable association in highly dynamic container environments, this method uses a native and stable kernel identifier—the cgroup ID—as the unique identifier for the asset. Its synchronization process is as follows: Continuous monitoring: Agents deployed on each node use the Watch mechanism of the Kubernetes API to monitor changes in Pod metadata on their respective nodes in real time.

[0072] Parsing and Mapping: Once a change is detected in a tag or annotation that is critical to an asset, the Agent immediately parses its value and converts it into a numerical level for internal use.

[0073] Identity transformation: The agent queries the local container runtime through Pod metadata to obtain the corresponding container ID, and then resolves its unique cgroup ID in the host operating system based on the container ID. This step is crucial; it successfully transforms user-defined, business-oriented logical concepts into physical identifiers that the kernel can recognize and process.

[0074] Kernel-mode update: Finally, the Agent writes a key-value pair containing {cgroup ID, CoA_Level} into an eBPF Map dedicated to this purpose.

[0075] The above process completes a closed loop from user definition to kernel awareness. When the eBPF detection program in the kernel captures the behavior of any process, it can perform an efficient key-value lookup in the eBPF Map using the cgroup ID to which the process belongs, thereby obtaining the criticality level of the asset in real time and accurately. This mechanism ensures that the technical risk score can be instantly combined with the business impact, providing a solid foundation for generating a response level that matches the business impact.

[0076] Furthermore, the risk decision matrix is ​​not statically configured, but rather a policy object that can be dynamically managed by the central analyzer. A default decision matrix configuration is shown in Table 1. To address the evolving threat landscape, the central analyzer can update the matrix mapping online based on global threat intelligence and atomically distribute it to all nodes through a two-phase commit mechanism: First, the node agent writes the new matrix completely into a backup eBPF Map; then, by atomically switching a global pointer, the kernel detection program instantly points to the new matrix. This two-phase commit mechanism ensures that there is no decision confusion or response delay during policy updates.

[0077] Table 1 Risk Decision Matrix

[0078] Furthermore, to balance the real-time nature of responses with the prudence of high-risk operations, the response scheduler employs a hierarchical decision-making model, intelligently distributing decision-making power and execution flow to different levels of the system. Specifically: After step S200 generates the Score, the node agent of the incident node will consult the latest decision matrix cached locally to obtain a preliminary ResponseLevel. Based on this ResponseLevel, the system will automatically select one of the following two different decision paths: Path 1: Local Autonomous Response (for low to medium risk): For low to medium impact operations such as "logging," "alerting," "observing," and "suppressing," the node agent is authorized to directly confirm and trigger immediate kernel-mode execution locally. This path ensures minimal latency control for the vast majority of common threats by completing the decision-making loop at the edge node.

[0079] Path Two: Centralized Confirmation and Response (For High-Risk Operations): For highly destructive operations such as "isolation" or "termination," the system automatically adds a "security lock." The node agent will pause local processing, report the full event to the central analyzer for final assessment, and can be subject to manual confirmation from security operations personnel. The node agent will only execute the action after the central analyzer or administrator issues the final confirmation instruction. This path provides a final safeguard against erroneous operations for critical business operations, ensuring the prudence and accuracy of high-risk decisions.

[0080] Furthermore, each response level determined by the risk decision matrix corresponds to a predefined, multi-stage, resilient response strategy, rather than an isolated action. The specific details of each level's strategy are as follows: Record: Enriches event tracking in the kernel, recording the complete context for post-event forensics and model optimization.

[0081] Alert: Sends structured alert events to the node agent and can be configured to provide real-time notifications for critical events.

[0082] Observe: Dynamically increase the granularity of monitoring all system behaviors related to the process or Pod, and enter the key monitoring mode.

[0083] Mitigate: A gentle, non-destructive intervention to address suspicious behavior. Examples include dynamically injecting error codes (such as EPERM) into specific system calls via eBPF, injecting TCP RST packets into suspicious network connections, or redirecting their DNS requests to a security sandbox.

[0084] Isolate: This response is not a single action, but a multi-stage strategy. For example, it first enters a 30-second "observation" mode; if the abnormal behavior persists, the network socket is dynamically redirected to an isolated network namespace using eBPF's sockmap / sockhash helper functions; only if malicious behavior is still detected does it finally escalate to hard isolation.

[0085] Terminate: Send a SIGKILL signal to the process as a last resort by using the bpf_send_signal(SIGKILL) helper function.

[0086] Preferably, this step combines a dynamic decision matrix, a hierarchical decision process, and a strategic response, extending the traditional risk assessment based on a single threshold into a process that can perform automated response orchestration based on quantified risk and business context.

[0087] In an alternative implementation, the method for obtaining the pipeline hazard point in step S300 can also be by analyzing the interaction between the temperature field and the stress field, such as the thermal stress concentration effect, and combining the material's thermal expansion coefficient and constraint conditions to calculate the thermal stress distribution, that is: simulating the superposition effect of the temperature field gradient and mechanical stress to identify the region of stress abrupt change under the thermo-mechanical coupling effect.

[0088] In another alternative implementation, the method for obtaining the pipeline hazard point in step S300 can also be by simulating the dynamic loads in the actual operation of the pipeline, such as pressure fluctuations and transient temperature changes, and combining fatigue damage accumulation theory, such as Miner's rule, to identify the area with the fastest damage accumulation rate and determine the hazard point.

[0089] In this embodiment of the application, step S400, in kernel mode, executes a response action corresponding to the response level, including the following steps D1-D2: Understandably, to ensure the stability and controllability of response actions and prevent a "response storm" triggered by a large number of abnormal events in a short period of time under certain attack scenarios (such as brute-force attacks and fork bombs), which could impact system performance, this method designs a response rate limiting and circuit breaker mechanism in kernel space. The core of this mechanism is an eBPFMap (BPF_MAP_TYPE_PERCPU_HASH) named response_limiter_map. This map uses the process's TGID as the key, and the value is a data structure containing {last_response_timestamp, event_count}.

[0090] D1: In kernel mode, the number of times each process triggers a response action per unit time is recorded through the eBPF storage mapping structure; D2: When the number of times exceeds the preset threshold, non-terminating response actions are restricted from being performed on the process.

[0091] For example, if the difference between the current time and the last_response_timestamp is less than a preset time window (e.g., 1 second), and the event_count has exceeded the threshold (e.g., 100 times), then this response action will be skipped, and only the counter will be incremented.

[0092] It should be noted that if a process triggers a higher number of responses within a very short period of time, the system will suspend all non-termination responses to that process for a short period of time (e.g., 5 seconds) to prevent resource abuse, while ensuring that the highest priority termination instructions can still be executed, and sending a "response storm" warning to user space to prevent potential kernel resource exhaustion risks.

[0093] Specifically, after the response level is determined, the system enters the kernel-mode immediate response execution phase, which is the final link in the security closed loop. To ensure that instructions are executed reliably, this method employs a sophisticated approach: the decision result (S300) is stored as an instruction in a dedicated eBPF Map, and response probes pre-positioned in critical kernel paths continuously check this Map. Once an instruction is detected, execution is immediately activated. To guarantee reliability, the return values ​​of all critical eBPF helper functions are rigorously checked. If execution fails, an event containing error details is immediately reported to the user-mode node Agent, thus forming an auditable and highly robust operational closed loop from instruction issuance to execution confirmation.

[0094] The specific response execution path is as follows: (1) Record Scenario: This response level is applied to the auditing and data collection of low-risk statistical bias behaviors. Its function is to capture behavioral samples that have optimization value for the baseline model, serving as data input for incremental learning and continuous evolution of the baseline model.

[0095] Technical Implementation: The response probe in the kernel encapsulates the event context into a structured data body, containing an event type identifier (EventType: RECORD) and a low-priority flag for subsequent processing. This data structure is then reported to the node agent in user space via the perf buffer through the bpf_perf_event_output() helper function. Upon receiving this data, the node agent routes it to the model optimization data stream based on its event type identifier. Data is asynchronously and batch-collected in this channel and ultimately transmitted to the central analyzer for consumption by the learning algorithm.

[0096] (2) Alert Scenario: When a low-risk anomaly is detected, it is necessary to notify the user to pay attention, log the anomaly, or link with external systems (such as SIEM).

[0097] Technical Implementation: The probe fills a structured event data body with information such as the current behavioral feature vector, anomaly score, and asset criticality CAR, and then calls the bpf_perf_event_output() helper function to report this event to the node agent in user space through the perf buffer.

[0098] System call: A web server process executed an uncommon setuid system call.

[0099] Network: An application service Pod connects to a new, previously unseen internal service IP.

[0100] File access: An operations and maintenance tool attempts to read a key file located in a non-standard path.

[0101] (3) Observe Scenario: To detect suspicious behavior of medium risk or requiring further analysis, and to conduct temporary, more granular enhanced monitoring.

[0102] Technical Implementation: The kernel probe first performs an "alert" operation. Subsequently, after receiving the "observe" command, the user-space node agent does not load a completely new eBPF probe. Instead, it activates enhanced functionality by updating an eBPF Map named "Enhanced Monitoring Process List." Specifically, the agent adds the target process's TGID to this map. When the pre-built eBPF probe in the kernel executes, it checks whether the current process's TGID exists in this map. If it does, the probe activates additional monitoring logic, such as calling bpf_get_stack() or bpf_get_stackid() to obtain the complete kernel-mode stack trace and reporting richer context data, achieving dynamic, on-demand enhanced monitoring of specific processes.

[0103] System call: A core database process executed a suspicious ptrace system call. The enhanced probe will immediately enable kernel-mode stack tracing (bpf_get_stack()) and report the complete call chain for context analysis.

[0104] Network: An application service Pod initiates a connection, but its intent is unknown. The enhanced probe loads a bpf_sock_ops program attached to the socket layer and begins recording metadata (such as size and interval) for all sent and received packets of this connection.

[0105] File Access: A process attempts to read multiple sensitive files. The enhanced probe loads a probe attached to the VFS layer and begins recording all subsequent file access paths of the process.

[0106] (4) Mitigate Scenario: To provide precise intervention for specific malicious behaviors that have been identified as medium to high risk but do not yet require termination of the entire process.

[0107] Technical Implementation: System calls: For specific system calls that are not allowed to be executed (such as bpf, kexec_load), the eBPF probe attached to the system call entry point (sys_enter) will call bpf_override_return(), which will directly return an error code (such as EPERM) to prevent the execution of the system call.

[0108] Network: If the abnormal behavior is a suspicious network connection, the eBPF program attached to the network protocol stack will perform connection reset, silent packet loss, or kernel-mode rate limiting.

[0109] File access: If the abnormal behavior is a suspicious file access, the eBPF probe attached to the VFS layer function (such as vfs_write) will call bpf_override_return() to return an EPERM (operation not allowed) error code to the process, thus preventing the file operation.

[0110] (5) Isolate Scenario: When high-risk behavior is detected, it is necessary to immediately limit the scope of the process's influence by placing it in a restricted sandbox to prevent it from interacting with other parts of the system, while keeping the process alive for subsequent analysis.

[0111] Isolation Environment Management: To achieve rapid isolation, the node agent pre-creates or configures a dedicated, network-isolated "isolated" cgroup or network namespace during its startup phase. The agent writes the identifier of this isolation environment (such as the cgroup's inode number or the network namespace's ID) into a global eBPF Map for the kernel program to consult at any time.

[0112] Technical Implementation: Network isolation: When isolation is required, the eBPF probe attached to the network-related kernel function reads the preset isolation cgroup identifier from the eBPF Map mentioned above, and then calls the bpf_sk_assign() helper function to forcibly assign all network sockets created by the target process to the isolation cgroup, thereby instantly cutting off its network connection.

[0113] File system isolation: The process can be prevented from accessing the file system by returning an error code to all subsequent file access system calls such as open and openat using bpf_override_return().

[0114] Process communication isolation: You can use bpf_override_return() to return errors for inter-process communication-related system calls such as kill and ptrace of the process, preventing it from affecting other processes.

[0115] (6) Terminate Scenario: The highest level of response, used to handle serious anomalous behavior that has been identified as high-risk or occurs on core critical assets (high CAR).

[0116] Technical Implementation: To balance the needs of immediate handling and post-event evidence collection, the Terminate response employs a safe mode of termination after a snapshot. When the eBPF probe detects a "terminate" command for the current process (TGID), it does not immediately send a SIGKILL signal, but instead performs the following two steps: Transient State Snapshot: The probe first places the process into a brief, quasi-isolated state, for example, temporarily blocking all new system calls via `bpf_override_return()`. Within this extremely short window (typically a few hundred milliseconds, and adjustable through configuration to balance the completeness of forensic information with the immediacy of the response), a dedicated eBPF probe is activated to capture a snapshot of the process. This snapshot data may include: kernel and user-space call stacks obtained via `bpf_get_stack()`, a list of currently open file descriptors, active network connection information, and fragments of key memory regions. This information is then rapidly reported to the node agent.

[0117] Final termination signal: After the snapshot acquisition command is issued, the probe calls the bpf_send_signal(SIGKILL) helper function to send the final termination signal to the process.

[0118] To ensure successful operation, the eBPF program checks the return value of bpf_send_signal(). If the call fails, a failure event containing an error code will be reported to the node agent for logging and alerting, forming a complete operation confirmation and auditing loop.

[0119] System call: The execve shell execution behavior was detected in the core authentication service.

[0120] Network: Behavior of connecting to a known malicious C2 server was detected in the database Pod.

[0121] File access: Batch, high-entropy file write operations in ransomware pattern were detected.

[0122] In an alternative implementation, the method for obtaining the pipeline hazard point in step S400 can also be by analyzing the interaction between the temperature field and the stress field, such as the thermal stress concentration effect, and combining the material's thermal expansion coefficient and constraint conditions to calculate the thermal stress distribution, that is: simulating the superposition effect of the temperature field gradient and mechanical stress to identify the region of stress abrupt change under the thermo-mechanical coupling effect.

[0123] In another alternative implementation, the method for obtaining the pipeline hazard point in step S400 can also be by simulating the dynamic loads in the actual operation of the pipeline, such as pressure fluctuations and transient temperature changes, and combining fatigue damage accumulation theory, such as Miner's rule, to identify the area with the fastest damage accumulation rate and determine the hazard point.

[0124] In summary, this method successfully resolves the inherent contradiction between response speed and detection accuracy in existing technologies by constructing a fast-slow dual-loop collaborative security system. Specifically, the kernel-level real-time response closed loop (fast loop) ensures microsecond-level instantaneous handling of unknown threats, while the global intelligent learning closed loop (slow loop) enables continuous self-evolution of the security baseline. Working together, these two loops enable the system in cloud-native dynamic environments to not only possess the extreme real-time performance and precise adaptive capabilities that are difficult to achieve simultaneously with traditional solutions, but also fundamentally improve the level of runtime security protection.

[0125] Example 3 illustrates a schematic scheme of a distributed collaborative learning and resilient response method based on eBPF. It should be noted that the technical solution of this eBPF-based distributed collaborative learning and resilient response system belongs to the same concept as the technical solution of the eBPF-based distributed collaborative learning and resilient response method described above. Details not described in detail in the technical solution of the eBPF-based distributed collaborative learning and resilient response system in this embodiment can be found in the description of the technical solution of the eBPF-based distributed collaborative learning and resilient response method described above.

[0126] This embodiment also provides a distributed collaborative learning and resilient response system based on eBPF, including: A synchronization module is built, which is based on system behavior data collected by the kernel-mode eBPF program, performs global analysis in user mode to build a normal behavior baseline model, and then synchronizes the lightweight baseline model to kernel mode. The detection output module, similar to that in kernel mode, performs deviation detection on the real-time generated system behavior based on a synchronous baseline model and outputs a risk score. The determination module, similar to combining risk scoring and pre-defined asset criticality, determines the response level; The execution module, similar to the kernel mode, executes the response actions corresponding to the response level.

[0127] This embodiment also provides an electronic device suitable for distributed collaborative learning and resilient response based on eBPF, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the distributed collaborative learning and resilient response method based on eBPF as proposed in the above embodiment.

[0128] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements the distributed collaborative learning and resilient response method based on eBPF as proposed in the above embodiments.

[0129] The storage medium proposed in this embodiment and the implementation of the distributed collaborative learning and elastic response method based on eBPF proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0130] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.

[0131] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A distributed collaborative learning and resilient response method based on eBPF, characterized in that: include, Based on system behavior data collected by kernel-mode eBPF programs, a global analysis is performed in user mode to construct a normal behavior baseline model, and the baseline model is then lightweighted and synchronized to kernel mode. In kernel mode, based on the synchronized baseline model, deviations in the real-time generated system behavior are detected, and a risk score is output. Based on the risk score and the pre-defined asset criticality, the response level is determined; In kernel mode, the response action corresponding to the response level is executed.

2. The distributed collaborative learning and resilient response method based on eBPF as described in claim 1, characterized in that: The process involves performing global analysis in user space based on system behavior data collected by kernel-mode eBPF programs to construct a normal behavior baseline model, and then synchronizing this lightweight baseline model to kernel space. In kernel mode, system behavior events are captured by the eBPF program, and features are extracted to generate behavior feature vectors; In the central analyzer in user mode, the behavioral feature vectors from multiple nodes are aggregated, and the normal behavior baseline model is constructed using an online learning algorithm; The normal behavior baseline model is distilled into lightweight detection rules and incrementally synchronized to the kernel-mode eBPF storage mapping of each node.

3. The distributed collaborative learning and resilient response method based on eBPF as described in claim 2, characterized in that: The system performs deviation detection on real-time generated system behavior and outputs a risk score. include, Calculate the instantaneous risk score of the current system behavior, wherein the instantaneous risk score combines at least the global rarity of the behavior and the unexpectedness of the behavior in the context of the learned behavior sequence; The instantaneous risk score of the current system behavior is weighted and recursively combined with the historical risk score of the process to generate the final risk score and output it.

4. The distributed collaborative learning and resilient response method based on eBPF as described in claim 3, characterized in that: Before executing the response action corresponding to the response level, including, In kernel mode, the number of times each process triggers a response action per unit time is recorded through the eBPF storage mapping structure; When the number of times exceeds a preset threshold, non-terminating response actions are restricted from being performed on the process.

5. A distributed collaborative learning and resilient response method based on eBPF as described in any one of claims 1-4, characterized in that: The online learning algorithm is a Hidden Markov Model.

6. The distributed collaborative learning and resilient response method based on eBPF as described in claim 5, characterized in that: The process of synchronizing the lightweight baseline model to the kernel state includes, The high-probability state transition paths in the Hidden Markov Model are transformed into key-value pair rules; In this key-value pair rule, the key is the identifier of the preceding behavior event, and the value is a compressed representation of the set of allowed subsequent behavior event identifiers.

7. The distributed collaborative learning and resilient response method based on eBPF as described in claim 6, characterized in that: The compression is referred to as a Bloom filter.

8. A distributed collaborative learning and resilient response system based on eBPF, employing the method described in any one of claims 1-7, characterized in that, include: A synchronization module is constructed, which is based on system behavior data collected by the kernel-mode eBPF program, performs global analysis in user mode to construct a normal behavior baseline model, and then synchronizes the baseline model to kernel mode after it is lightweighted. The detection output module, similar to that in kernel mode, performs deviation detection on the real-time generated system behavior based on the synchronized baseline model and outputs a risk score; The determination module, by combining the risk score and the preset asset criticality, determines the response level; The execution module, similar to the kernel mode, executes the response action corresponding to the response level.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Escape type malicious software detection method and system

    CN118296602A

  • Linux kernel scheduler parameter optimization system and method

    CN119271382A

  • Method and system for realizing CUDA call tracking based on eBPF

    CN120723587A