System state acquisition method, device and system and nonvolatile storage medium
By using eBPF and LLM technologies in the trusted computing environment, combined with observation data analysis at the microservice layer, container network layer, and infrastructure layer, the problem of efficient observation of cloud-native applications is solved, and accurate and efficient determination of system status is achieved.
Patent Information
- Application Number
- CN202510806382.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-26
AI Technical Summary
In the trusted computing environment, existing technologies cannot efficiently observe cloud-native applications, resulting in the inability to determine the system status in a timely and accurate manner.
The extended Berkeley filter proxy module (eBPF) and large language model (LLM) running in the kernel space are used to obtain observation data at the microservice layer, container network layer, and infrastructure layer respectively, and the data is filtered and analyzed through the neural network model to summarize the system operation status.
It enables comprehensive observation of cloud-native applications, reduces resource overhead, and improves the accuracy and efficiency of system status determination.
Smart Images

Figure CN120704987A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cloud computing, and more specifically, to a method, device, system, and non-volatile storage medium for acquiring system status. Background Art
[0002] Related technologies for observing cloud-native applications deployed in trusted computing environments cannot fully monitor these applications, and the resource overhead of various types of observations is high. Therefore, related technologies cannot efficiently monitor applications deployed in trusted computing environments, resulting in an inability to accurately determine system status in a timely manner.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, system and non-volatile storage medium for obtaining system status, so as to at least solve the technical problem of being unable to accurately determine the system status due to the inability to efficiently observe applications deployed to the trusted innovation environment in related technologies.
[0005] According to one aspect of an embodiment of the present application, a method for obtaining system status is provided, including: obtaining first observation data of a microservice application in the microservice layer of the system at the microservice layer of the system; obtaining second observation data of a node application in the container network layer of the system through a first extended Berkeley filter proxy module, wherein the first extended Berkeley filter proxy module is set in the kernel space of each node in the cluster node; obtaining third observation data of the node in the infrastructure layer of the system through a second extended Berkeley filter proxy module, wherein the second extended Berkeley filter proxy module is set in the kernel space; and summarizing and analyzing the first observation data, the second observation data, and the third observation data to obtain the operating status of the system.
[0006] Optionally, obtaining first observation data of a microservice application in the microservice layer of the system includes: determining first tracking information of a network data packet in the system, wherein the first tracking information is used to indicate transmission path information of the network data packet in the microservice layer; using a neural network model to filter the initial first observation data, remove noise data in the first observation data, and obtain the first observation data; using a neural network model to process the first observation data and the tracking information to obtain a microservice call pattern recognition result, wherein the microservice call pattern recognition result is used to indicate whether there is an abnormal microservice call pattern in the microservice layer of the system.
[0007] Optionally, determining the first tracing information of the east-west network data packet of the microservice application includes: adding a distributed tracing context to the header of the network data packet of the microservice application; determining the initial first tracing information of the network data packet based on the distributed tracing context; and using a neural network model to filter the initial first tracing information, remove noise information in the initial first tracing information, and obtain the first tracing information.
[0008] Optionally, the method further includes: identifying the tracing context information of the network data packet in the system through the first extended Berkeley filter proxy module to obtain second tracing information of the network data packet, wherein the second tracing information includes transmission path information of the network data packet in the container network layer.
[0009] Optionally, after obtaining the second observation data of the node application through the first extended Berkeley filter proxy module, the method also includes: using a neural network model to process the second observation data and the second tracking information to determine the fault type of the system and the fault cause corresponding to the fault type, wherein the fault type includes at least one of the following: network congestion, packet loss, and network delay higher than a preset delay threshold.
[0010] Optionally, obtaining the third observation data of each node through the second extended Berkeley filter proxy module includes: obtaining the third observation data of each node through the second extended Berkeley filter proxy module and the data exporter; using a neural network model to process the third observation data to obtain the correlation relationship between internal data of the third observation data.
[0011] Optionally, the first observation data includes at least one of the following: indicator data or log data of the microservice application, the second observation data includes communication performance indicator data of the node application, and the third observation data includes perception container telemetry data of the node.
[0012] According to another aspect of an embodiment of the present application, a device for obtaining system status is also provided, including: a first processing module for obtaining first observation data of a microservice application in the microservice layer of the system at the microservice layer of the system; a second processing module for obtaining second observation data of a node application in the container network layer of the system through a first extended Berkeley filter proxy module, wherein the first extended Berkeley filter proxy module is set in the kernel space of each node in the cluster node; a third processing module for obtaining third observation data of the node in the infrastructure layer of the system through a second extended Berkeley filter proxy module, wherein the second extended Berkeley filter proxy module is set in the kernel space; a fourth processing module for summarizing and analyzing the first observation data, the second observation data and the third observation data to obtain the operating status of the system.
[0013] According to another aspect of an embodiment of the present application, a non-volatile storage medium is provided, in which a program is stored. When the program is running, a method for obtaining a system status of a device where the non-volatile storage medium is located is controlled.
[0014] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the method for acquiring the system status is executed when the program is running.
[0015] According to another aspect of an embodiment of the present application, a computer program product is provided, including a computer program, which implements a method for acquiring a system status when executed by a processor.
[0016] In an embodiment of the present application, the first observation data of the microservice application in the microservice layer of the system is obtained; the second observation data of the node application in the container network layer of the system is obtained through the first extended Berkeley filter proxy module, wherein the first extended Berkeley filter proxy module is set in the kernel space of each node in the cluster node; the third observation data of the node in the infrastructure layer of the system is obtained through the second extended Berkeley filter proxy module, wherein the second extended Berkeley filter proxy module is set in the kernel space; the first observation data, the second observation data and the third observation data are summarized and analyzed to obtain the operating status of the system, by respectively obtaining observation data at different levels and collecting observation data in the kernel space through the extended Berkeley filter proxy module, the purpose of comprehensively collecting observation data and reducing resource overhead in the process of collecting observation data is achieved, thereby achieving the technical effect of efficiently observing the application, and thus solving the technical problem of being unable to accurately determine the system status due to the inability to efficiently observe the application deployed to the trust innovation environment in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 1 is a schematic diagram of the structure of a computer terminal (mobile terminal) provided according to an embodiment of the present application;
[0019] Figure 2 This is a flow chart of a method for obtaining system status according to an embodiment of the present application;
[0020] Figure 3This is a schematic diagram of an observable agent architecture based on an extended Berkeley filter provided according to an embodiment of the present application;
[0021] Figure 4 This is a schematic diagram of a cloud-native microservices full-link observability architecture provided according to an embodiment of the present application;
[0022] Figure 5 This is a performance indicator comparison diagram provided according to an embodiment of the present application;
[0023] Figure 6 It is a structural diagram of a system status acquisition device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:
[0027] eBPF (Extended Berkeley Packet Filter): A powerful kernel technology originally designed for network packet filtering but now extended to many other uses. It allows users to run small, efficient programs in the kernel without modifying the kernel source code or loading kernel modules.
[0028] eBPF has attracted significant attention for load balancing, firewalling, and network security. An eBPF program is a series of 64-bit instructions verified by a compiler for safety and high performance. It is just-in-time compiled using tools such as Clang LLVM on the host for instrumentation on the target. eBPF programs export kernel-level metrics by writing them to a memory map located in kernel space but accessible from user space. This approach reduces the need to copy data from kernel space to user space, enabling higher throughput with lower overhead. This lower overhead enables running eBPF scripts dynamically in production environments to debug real systems without impacting users.
[0029] OTel (OpenTelemetry): Provides a standardized Software Development Kit (SDK), Application Programming Interface (API), and tools for ingesting, transforming, and transmitting data to observability backends.
[0030] Sidecar: A container architecture pattern that runs alongside the main application container, providing logging, monitoring, and proxy services to support the operation, maintenance, and scalability of the main application. It is commonly used in microservices and containerized environments.
[0031] LLM (Large Language Model): Based on deep learning technology, it can process and generate natural language text, has strong language understanding and generation capabilities, and can be used for intelligent interaction and data processing tasks in various fields.
[0032] Observability in cloud-native applications refers to metrics, logs, and traces. Metrics are time-varying numerical data representations that can be used to infer system behavior through mathematical modeling or prediction. Metrics are optimized for storage, compression, and retrieval, enabling longer retention and faster querying. Logs are immutable, time-stamped records of discrete events recorded over time. Finally, traces represent interactions between distributed applications, providing a high-level view of the request-response lifecycle. In a distributed system, a single request may traverse multiple services distributed across different servers or geographical locations. Implementing observability in monolithic applications is straightforward because their codebase and deployment reside on a single infrastructure endpoint. However, achieving the same task in a secure deployment is more complex. In a monolithic application, developers can instrument the application or host to collect observability data. However, in a secure deployment with numerous containers, distributed hosts, and a diverse microservices ecosystem, manually adding additional observability code to different microservices or operating system (OS) kernels is not feasible. Therefore, a key issue in ensuring observability in cloud-native applications is the overhead associated with maintaining and updating observability code across the diverse microservices ecosystem.
[0033] Kubernetes is a highly modular and extensible open source container orchestration platform that automates the deployment, scaling, and management of containerized applications. OTel provides per-language instrumentation libraries that can be used to export metrics, logs, and trace information from applications, supporting multiple languages and platforms, including Java, Go, Node.js, Python, and .NET. However, in an agile, multi-language environment, using OTel to instrument each application is cumbersome. In addition, it only understands the L7 metrics of the application and does not collect information at other layers, the network stack, or the kernel level, which is critical for cloud-native observability. Moreover, OTel aggregates observable data from multiple pods into services, which limits its ability to perform pod-level analysis.
[0034] Observability sidecars (helper containers that run alongside application containers in Kubernetes pods) reduce the overhead of maintaining multiple applications by redirecting all observability functionality from application containers to separate (sidecar) containers. Typically, sidecars deployed alongside each microservice handle network functions and perform L3, L4, and L7 network observability tasks. However, sidecar observability is limited to collecting network metrics, namely latency, traffic / error rates, and tracing information. Therefore, because sidecars increase container overhead on cloud-native clusters, they have the disadvantage of consuming valuable cluster resources, including compute, storage, and network resources. Furthermore, while sidecars are particularly useful for detecting problems in microservices, they cannot provide precise root cause analysis. For example, tracing information can be used to identify misbehaving services and analyze latency, but debugging the internals of application containers through sidecars is not possible. Therefore, to achieve comprehensive observability of containers, observability data must be collected directly from the underlying host operating system. While many agent-based approaches exist (e.g., OpenTelemetry, cAdvisor, AWS X-Ray, Cloudwatch ContainerInsights), deploying these solutions (typically in user space) incurs additional resource overhead, impacting the performance of deployed workloads.
[0035] It is important to note that, unlike monolithic architectures where components communicate through in-process calls, microservice architectures are subject to more service interaction failures because the interaction between microservices may occur over unreliable networks. In addition, containers in related technologies are designed to be short-lived and are started and destroyed quickly, which makes tracking their related data challenging. In addition, containers deployed on distributed hosts make it complicated to collect a comprehensive view of service behavior and performance. Finally, managing containers with Kubernetes requires monitoring additional components such as kube-apiserver, kube-controller-manager, agents, kubelet, and container network interface (CNI).
[0036] In addition, LLM has made significant progress in the field of natural language processing, demonstrating powerful language understanding, generation, and reasoning capabilities. However, in the observability scenario of trusted container clusters, the advantages of LLM have not yet been fully utilized to solve the problems existing in related technologies.
[0037] In order to solve the above problems and achieve comprehensive observability of containers, we must collect observable data directly from the underlying host operating system. Although there are many agent-based approaches, deploying these solutions (usually in user space) will incur additional resource overhead, affecting the performance of deployed workloads. In the embodiment of this application, a comprehensive intelligent observability solution based on eBPF covering indicators, logs and distributed tracing is developed for the trusted container environment, which is described in detail below.
[0038] It should be noted that the observability solutions implemented in the related art use user space programs to capture and analyze data, which results in significant performance overhead. The method proposed in the embodiment of the present application transfers these tasks to the eBPF program in the kernel space, thereby significantly reducing the overhead, while being able to attach in-depth contextual information to these data from the kernel. Moreover, the eBPF framework proposed in the embodiment of the present application can be dynamically injected into any deployed application without the need for recompilation or redeployment. The solution we proposed improves Cloudflare's existing eBPF exporter and incorporates new features into several .bpf.c scripts, including accept-latency, cachestat, llcstat, malloc, oomkill, runqlat, shrinklat, and tcpbacklog.
[0039] That is to say, the embodiment of the present application proposes a complete intelligent observability solution for the container environment, which utilizes the extended Berkeley Packet Filter (eBPF) in the Linux kernel. The proposed solution relies on a small event trigger based on eBPF to extend the operating system functionality without detecting the corresponding application code, while running at line speed in the kernel space (that is, running at the maximum speed transmission rate supported by the hardware device) without reducing the performance of the detected events. By utilizing an observable agent based on kernel space eBPF, the solution can directly collect in-depth contextual observable data without incurring additional overhead. In addition, the proposed solution allows for in-depth understanding of the context of the observed application without the burden of detection. With the help of LLM, intelligent data screening, analysis and visualization are achieved, reducing manual intervention and improving the intelligence level of the observable system. The framework proposed in the embodiment of the present application can be dynamically injected into deployed applications without recompilation or deployment, enhancing the flexibility and scalability of the system.
[0040] According to an embodiment of the present application, a method embodiment of a method for obtaining a system status is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0041] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for obtaining system status is shown. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more (illustrated as 102a, 102b, ..., 102n) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0042] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0043] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for obtaining the system status in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned method for obtaining the system status. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0044] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0045] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0046] In the above operating environment, the embodiment of the present application provides a method for obtaining the system status, such as Figure 2 As shown, the method includes the following steps:
[0047] Step S202: obtaining first observation data of a microservice application in the microservice layer of the system at the microservice layer of the system;
[0048] In the technical solution provided in step S202, the step of obtaining first observation data of a microservice application in the microservice layer of the system includes: determining first tracking information of a network data packet in the system, wherein the first tracking information is used to indicate transmission path information of the network data packet in the microservice layer; using a neural network model to filter the initial first observation data, remove noise data in the first observation data, and obtain the first observation data; using a neural network model to process the first observation data and the tracking information to obtain a microservice call pattern recognition result, wherein the microservice call pattern recognition result is used to indicate whether there is an abnormal microservice call pattern in the microservice layer of the system.
[0049] As an optional implementation, the following methods can be used to identify whether there are abnormal microservice call patterns in the system's microservice layer:
[0050] Determine the first trace information of the network packets in the system. This information details the transmission path of the packet at the microservice layer, including the source service, destination service, and service nodes passed through. To obtain this information, the first eBPF proxy module is deployed on each node of the Kubernetes cluster, directly capturing the header of the network packet in kernel space to read and parse key trace identifiers, such as the distributed tracing ID. This is achieved through standard fields such as X-Request-ID. These fields are updated by each service node as the packet traverses the microservice mesh, thereby recording the complete request-response cycle and the interaction between microservices.
[0051] LLM can then be used to screen and process the initially collected first observation data. This process aims to remove irrelevant or redundant data, that is, noise data in the first observation data, to ensure the accuracy and efficiency of subsequent analysis. The first observation data includes indicators such as the call frequency, response time, error rate of microservices, as well as network-level data such as packet size, sending timestamp, etc. LLM is trained to identify features that are closely related to microservice call patterns. For example, the average response time and standard deviation of a specific service call under normal circumstances. Based on this, the model can filter out data points that do not provide additional diagnostic value, such as service call records within the normal response time range, thereby obtaining refined first observation data.
[0052] Then, when processing the refined first observation data and trace information, LLM can be used to further analyze the first observation data and trace information to identify anomalies in the microservice call pattern. This process involves correlating the trace information with the first observation data, combining network path information and microservice interaction data. In this process, the LLM model constructs a dynamic graph of microservice calls, where nodes represent services and edges represent calls between services. The weights of edges are determined by metrics such as call frequency and latency. In this way, by comparing historical data, the model can identify call patterns that deviate from normal behavior. For example, a sudden surge in the number of calls to a service, a significant increase in average latency, or an abnormally high error rate may all be signs of service failure, network congestion, or a malicious attack.
[0053] In an exemplary embodiment, suppose the system monitors that the latency of calls from microservice A to microservice B suddenly increases from an average of 50 milliseconds to over 200 milliseconds, and the call frequency is 30% higher than usual. Simultaneously, the LLM notices that the response time of calls from microservice B to microservice C is normal, but calls to microservice D are failing significantly. Combining these observations with trace information, the model can infer that the root cause of the abnormal call pattern may be a problem in the specific path from microservice B to D, rather than a performance bottleneck in microservice B itself.
[0054] During this process, tracing information, acting as path identifiers, provides precise context for microservice calls, helping to locate problematic services and paths. First-observation data, acting as quantitative indicators, reveals anomalous characteristics of call patterns. The combination of these two makes it possible to accurately identify abnormal microservice call patterns within the system, providing key clues for rapid problem diagnosis and remediation. Through continuous monitoring and dynamic model updates, this mechanism can adapt to changes in the system environment, such as the introduction of new services or network configuration adjustments, to maintain its effectiveness in identifying abnormal call patterns.
[0055] In some embodiments of the present application, the above-mentioned network data packets include east-west data packets and north-south data packets, and distributed tracing context can be added and traced only in the header of the east-west data packets or only in the header of the north-south data packets.
[0056] As an optional implementation, the step of determining the first tracing information of the east-west network data packet of the microservice application includes: adding a distributed tracing context to the header of the network data packet of the microservice application; determining the initial first tracing information of the network data packet based on the distributed tracing context; and using a neural network model to filter the initial first tracing information, remove noise information in the initial first tracing information, and obtain the first tracing information.
[0057] In some embodiments of this application, the OTeI library can be used at the microservice layer to detect microservices, appending additional distributed tracing context to east-west network packet headers for propagation. Simultaneously, a neural network model (LLM) is introduced to perform preliminary screening and analysis of the metrics, logs, and trace information generated by OTeI. For example, LLM can identify abnormal microservice call patterns based on predefined business rules and historical data patterns, issuing early warnings.
[0058] Step S204: obtaining second observation data of the node application at the container network layer of the system through a first extended Berkeley filter proxy module, wherein the first extended Berkeley filter proxy module is set in the kernel space of each node in the cluster node;
[0059] In the technical solution provided in step S204, the method further includes: identifying the tracing context information of the network data packet in the system through the first extended Berkeley filter proxy module to obtain second tracing information of the network data packet, wherein the second tracing information includes transmission path information of the network data packet in the container network layer.
[0060] As an optional implementation, after obtaining the second observation data of the node application through the first extended Berkeley filter proxy module, a neural network model can also be used to process the second observation data and the second tracking information to determine the system fault type and the fault cause corresponding to the fault type, wherein the fault type includes at least one of the following: network congestion, packet loss, and network delay higher than a preset delay threshold.
[0061] In some embodiments of the present application, through the first extended Berkeley filter agent module, the second observation data of the node application can be collected directly in the kernel space, including but not limited to key performance indicators such as network interface throughput, data packet sending and receiving status, network latency, etc. These data are not processed by the user space agent, but are captured directly by the eBPF program, ensuring the efficiency and accuracy of data collection. Subsequently, the second tracing information is extracted, which includes the complete path of the network packet and the tracing ID of the request-response cycle, providing detailed context about the network activity.
[0062] LLM then processes the second observation data and the second trace information to identify potential system failure types. In this process, the second observation data provides quantitative indicators of system operation, while the second trace information provides specific details about the transmission of data packets in the network, including the network nodes passed, the delays experienced, and whether any packet loss occurred.
[0063] For example, suppose a network node in a cluster experiences a sustained high load, with a significant number of inbound and outbound packets on its network interface, potentially leading to network congestion. The eBPF agent module monitors the node's network interface and collects secondary observations, such as packet queue wait time, number of failed deliveries, and average network latency. Furthermore, by parsing secondary trace information residing in the network packet headers, it can identify the packet's origin and destination, as well as the intermediate nodes it traverses.
[0064] By learning from historical data, the LLM establishes a baseline for network congestion, packet loss, and latency under normal network conditions. When the model then receives the latest secondary observations, it can compare them with the baseline data to identify potential anomalies. For example, if the queue wait time for packets increases significantly, the packet loss rate exceeds a normal threshold, or the network latency exceeds a preset delay threshold, the model will interpret these anomalies as signs of network congestion.
[0065] In another exemplary embodiment, assume that the system detects frequent packet loss from microservice E to microservice F, which could lead to inter-service communication interruptions or data inconsistencies. The eBPF proxy module uses the second tracing information to identify all packets from E to F and collects second observation data, including packet size, send timestamp, and lack of receipt acknowledgment. The LLM processes this information to identify the pattern and frequency of packet loss. If the packet loss rate reaches or exceeds a preset threshold, the model diagnoses a network packet loss fault in the system.
[0066] Finally, if network latency exceeds a preset latency threshold, the eBPF agent module collects secondary observation data, including the timestamps of sent and received packets and the calculated latency. The LLM analyzes this latency data, compares it with historical data, and identifies services or paths with abnormal latency. If the model detects that the average latency from microservice G to microservice H is significantly higher than normal, this may indicate that the network connection between G and H is disrupted or has a bottleneck. The model will diagnose this as a network latency fault.
[0067] In this way, the eBPF framework and LLM work together to intelligently detect and identify network congestion, packet loss, and latency anomalies at the container network layer. This provides timely warnings to the operations team, facilitating rapid problem location and resolution, and improving system reliability and efficiency. With continuous model learning and optimization, the accuracy and responsiveness of this identification mechanism will continue to improve, further adapting to the complex and ever-changing container environment.
[0068] In some embodiments of the present application, at the container network layer, the RED indicators of the application, namely rate, error and duration, are monitored by the eBPF agent (that is, the first extended Berkeley filter agent module) deployed on the cluster nodes. The eBPF agent captures and exports throughput, latency and error codes. Using the Deepflow eBPF library, the agent provides a mechanism for automatically collecting context propagation and OpenTelemetry distributed tracing. On this basis, LLM analyzes the collected network data in real time, intelligently determines the root causes of network congestion, packet loss and other problems, and provides corresponding solution suggestions. For example, when an increase in network latency is detected, LLM can analyze whether it is due to excessive traffic of a certain microservice or a network configuration problem, and give suggestions for adjusting traffic distribution or optimizing network configuration.
[0069] Step S206: Acquire third observation data of the node at the infrastructure layer of the system through the second extended Berkeley filter proxy module, wherein the second extended Berkeley filter proxy module is set in the kernel space;
[0070] In some embodiments of the present application, the step of obtaining the third observation data of each node through the second extended Berkeley filter proxy module includes: obtaining the third observation data of each node through the second extended Berkeley filter proxy module and the data exporter; using a neural network model to process the third observation data to obtain the correlation relationship between the internal data of the third observation data.
[0071] As an optional implementation, at the infrastructure layer, an eBPF proxy (also known as the second Extended Berkeley Filter Proxy Module) can be deployed in conjunction with Cloudflare's eBPFexporter to obtain detailed, container-aware telemetry data about cluster nodes, which can help debug and troubleshoot anomalies. LLM deeply analyzes this underlying telemetry data, discovering potential correlations between the data and providing system administrators with a more comprehensive system health report.
[0072] Optionally, the second extended Berkeley filter proxy module can aggregate node data from various domains in the cluster nodes, where each domain can be viewed as a node set consisting of a subset of the nodes in the cluster nodes. Nodes in different domains cannot interact directly with each other. During the third observation data collection process, the extended Berkeley program installed in the kernel of each node writes the passively collected data into eBPF maps, which are then transferred to the second extended Berkeley filter proxy module by the eBPF exporter.
[0073] Step S208: Summarize and analyze the first observation data, the second observation data, and the third observation data to obtain the operating status of the system.
[0074] In the technical solution provided in step S208, the first observation data includes at least one of the following: metric data or log data of a microservice application; the second observation data includes communication performance metric data of a node application; and the third observation data includes the node's perceived container telemetry data. The perceived container telemetry data includes data that reflects the running status of applications within the container, container resource usage, and interaction information within the node's containers. The perceived container telemetry data can be used to monitor and analyze the health, performance, and security of the containerized environment.
[0075] In the solution for intelligently observing the trusted container environment, which combines the dynamic injection of the eBPF framework with LLM technology, provided in the embodiments of this application, a multi-layered data collection and analysis method is adopted to achieve a comprehensive understanding of the entire system. The first observation data, the second observation data, and the third observation data are respectively derived from the microservice layer, the container network layer, and the infrastructure layer. The data at each layer provides a specific perspective, and together they form a picture of the overall operating status of the system.
[0076] First, observation data, namely, metrics or log data from microservice applications, includes inter-service call frequency, response time, error rates, and application logs. This data can reveal the health and interaction patterns of microservices. Through the OTeI library, this information is attached to network packets for dissemination and centralized processing. During the analysis phase, the LLM model can interpret this data to identify abnormal patterns in microservice calls or service performance degradation, such as increased response latency, frequent errors, or unexpected service interaction behavior.
[0077] Secondary observation data involves communication performance metrics for node applications, primarily including network congestion, packet loss, and latency. This data is captured in real time by the eBPF agent in kernel space and directly reflects the efficiency of network interactions between containers and with the outside world. The LLM model leverages this secondary observation data to provide insight into potential network-level issues, such as congestion on network paths, unstable packet transmission, or abnormal latency. This provides timely warnings of risks of network performance degradation and aids operational decision-making.
[0078] Third-order observation data, or tertiary observation data, is the node's perceived container telemetry data. This data covers key information about the containerized environment, including the running status of applications within the container, container resource usage (such as CPU, memory, and disk I / O), and interactions between containers. This data not only reflects the resource consumption and performance of individual containers, but also provides insights into interactions between containers and between containers and the host. The LLM model leverages its powerful semantic understanding and pattern recognition capabilities to conduct in-depth analysis of tertiary observation data, assessing the overall health of the container environment, predicting potential resource bottlenecks or security vulnerabilities, and ensuring the efficient utilization and secure operation of system resources.
[0079] By combining primary, secondary, and tertiary observations, the LLM model can perform advanced data fusion and analysis to create a comprehensive view of the system's operational status. The model combines call patterns captured at the microservice layer with communication performance at the network layer and resource usage at the infrastructure layer, corroborating these observations and identifying system-level issues. For example, if the primary observation reveals increased latency in microservice calls, while the secondary observation reveals increased packet loss along a specific network path, and the tertiary observation reveals a surge in CPU utilization for a particular container, the LLM model can link these isolated observations and infer that excessive CPU resource consumption may be causing network latency and packet transmission issues.
[0080] Through this comprehensive analysis, the LLM model not only detects and diagnoses system failures or performance bottlenecks, but also predicts future challenges, providing data-driven decision-making for optimizing resource allocation strategies, strengthening security measures, and improving user experience. As data accumulates and the model optimizes itself, this analytical framework will become more intelligent and efficient, enabling intelligent operation, maintenance, and management of trusted container environments.
[0081] In some embodiments of the present application, Prometheus may be used to provide data storage for observable data sources, and a Grafana dashboard may be used for visualization.
[0082] In some embodiments of the present application, Figure 3 As shown in the figure, an observable agent architecture based on extended Berkeley filters is also provided. The agent architecture is also integrated with the Kubernetes infrastructure. Figure 3As can be seen, the cloud-native microservice full-link observability enhancement solution based on eBPF and LLM technology is composed of multiple eBPF programs running in the kernel space. These eBPF programs are triggered by system calls generated by other applications. The Prometheus collector running on the Worker eBPF node will regularly extract the collected data from the eBPF agent on the cluster. In the master node MasterNode, Control Plane and Network will be used to export relevant data. The master node and the worker node can interact with each other by using API communication. In addition Figure 3 and Figure 4 The eBPF Agents in this section refer to eBPF collection clients, which collect node data. If cross-domain transmission is not required, the collected data is sent directly to Prometheus in the same domain. If cross-domain transmission is required, data collected by the collection client in one domain is first sent to a proxy (eBPF proxy) and then to Prometheus.
[0083] It can be seen that the solution provided by the embodiment of the present application realizes the complete microservice observability of Kubernetes applications by combining the three observability layers of microservice layer, container network layer and infrastructure layer. It can collect indicators, trace information and detect performance anomalies from distributed containers, taking into account both scalability and efficiency. The eBPF-based Agent is deployed on each cluster node to dynamically expand the operating system functions without affecting the application life cycle. The observable data generated at each layer is exported to the Prometheus instance and visualized through the customized Grafana dashboard. The cloud native microservice full-link observability architecture within the working node is as follows: Figure 4 As shown. Furthermore, in the method provided by this application, microservices run in a container runtime environment and communicate through the Envoy sidecar proxy. The eBPFAgent (eBPF collection client) runs as a DaemonSet in the kernel space of all cluster nodes. The sidecar, collector, and application provide distributed tracing IDs and metrics to the Prometheus instance, with a crawling interval of 1 second.
[0084] By adopting the microservice layer of the system, the first observation data of the microservice application in the microservice layer of the system is obtained; the second observation data of the node application in the container network layer of the system is obtained through the first extended Berkeley filter proxy module, wherein the first extended Berkeley filter proxy module is set in the kernel space of each node in the cluster node; the third observation data of the node in the infrastructure layer of the system is obtained through the second extended Berkeley filter proxy module, wherein the second extended Berkeley filter proxy module is set in the kernel space; the first observation data, the second observation data and the third observation data are summarized and analyzed to obtain the operating status of the system, by respectively obtaining observation data at different levels and collecting observation data in the kernel space through the extended Berkeley filter proxy module, the purpose of comprehensively collecting observation data and reducing resource overhead in the process of collecting observation data is achieved, thereby realizing the technical effect of efficiently observing the application, and thus solving the technical problem of being unable to accurately determine the system status due to the inability to efficiently observe the application deployed to the trust innovation environment in the related technology.
[0085] In addition, the solution provided in the embodiment of the present application for intelligently observing the trusted container environment by combining the dynamic injection of the eBPF framework and LLM technology simultaneously achieves three effects: reducing the OpenTelemetry Collector overhead, eliminating the Sidecar in distributed tracing, and improving the accuracy of host-level indicators.
[0086] To reduce the overhead of the OpenTelemetry Collector, manually instrumented applications are the first step to achieve observability. In the embodiment of the present application, the application is instrumented by integrating the OTeI library to achieve the purpose of observability. The embodiment of the present application proposes a two-step approach to achieve complete application-level observability. First, the OTeI language SDK is used to generate indicators, logs, and tracing information, which requires application developers to instrument the application, but it is a reliable way to expose custom (organization-specific) observable data. Next, the eBPF Agent collects the indicators and tracing information generated by the OpenTelemetry instrumented application and exports it to Prometheus. During this process, LLM can compress and optimize the collected data in real time to reduce data transmission and storage usage. For example, LLM can identify duplicate or redundant data, merge or delete it, and improve data transmission and storage efficiency.
[0087] To eliminate sidecars in distributed tracing, the Kubernetes microservice architecture uses an ingress or reverse proxy (such as Envoy) as the entry point for all incoming network packets to form a network service mesh. OpenTelemetry uses this proxy to generate distributed tracing of applications. When Pods forward requests to other Pods in our network, the Pods' Envoy sidecar extracts and propagates the distributed tracing ID using the X-Request-ID header. By aggregating information from multiple sidecars, the round-trip process of the request-response cycle can be visualized in distributed tracing. Envoy's sidecar approach also tracks metrics such as throughput, latency, and errors. Sidecars like Envoy help provide Layer 7 (application layer) observability by being able to inspect network packets and decrypt them through Transport Layer Security (TLS). However, a complete network stack observability solution does not require the deployment of user space sidecars. The solution proposed in the embodiment of this application uses eBPF to natively parse all packets flowing in the network and extract the X-Request-ID header to generate a complete distributed trace. At the same time, Deepflow is used to demonstrate a kernel space distributed tracing method based on eBPF. During this process, LLM can intelligently analyze distributed tracing data, quickly locate fault points, and provide guidance for troubleshooting. For example, when request loss or high latency occurs, LLM can analyze the tracing data to identify the node or microservice causing the problem and provide corresponding repair suggestions.
[0088] To improve the accuracy of host-level metrics, most host-level observability agents are privileged applications that access the / proc virtual file system. The / proc folder contains runtime system information (such as hardware configuration). Prometheus NodeExporter is a wrapper that reads data from the / proc folder and serves it through an HTTP endpoint. Unlike NodeExporter, cAdvisor provides a breakdown of resource utilization for each container while exporting metrics, i.e., a container-aware exporter. Although NodeExporter and cAdvisor provide time series metrics by sampling data in the / proc folder, they may cause information loss depending on the choice of sampling interval. The eBPF-based solution proposed in the embodiment of the present application adopts an innovative sampling method. The eBPF Agent collects metrics directly by running the eBPF program in the Linux kernel. These metrics include comprehensive system performance, resource utilization, and network traffic information. The collected metrics are written to the eBPF Map, capturing all events without losing information. Using the eBPF exporter significantly reduces overhead, which is only 1% of the NodeExporter overhead. It also ensures comprehensive metric collection, enabling access to kernel-level metrics and making over 2,000 kernel-level metrics and tracepoints available, metrics that are inaccessible to other tools. LLM can perform in-depth analysis of these host-level metrics, providing more accurate system performance assessment and forecasting. For example, LLM can predict future resource usage trends based on historical metric data, helping administrators plan and adjust resources in advance.
[0089] A comparison diagram of the observation capabilities of the method provided in the embodiment of the present application and the OpenTelemetry, EnvoySidecar, cAdvisor, NodeExporter and other solutions in related technologies is shown in FIG. Figure 5 shown. Figure 5 Wherein, “Ouer solution” represents the method provided in the embodiment of the present application, “√” indicates that the performance of the method in this evaluation indicator meets the preset requirements, and “×” indicates that it does not meet the requirements.
[0090] The embodiment of the present application provides a device for obtaining system status. Figure 6 It is a structural diagram of the device. Figure 6It can be seen that the device includes: a first processing module 60, which is used to obtain the first observation data of the microservice application in the microservice layer of the system at the microservice layer of the system; a second processing module 62, which is used to obtain the second observation data of the node application in the container network layer of the system through the first extended Berkeley filter proxy module, wherein the first extended Berkeley filter proxy module is set in the kernel space of each node in the cluster node; a third processing module 64, which is used to obtain the third observation data of the node in the infrastructure layer of the system through the second extended Berkeley filter proxy module, wherein the second extended Berkeley filter proxy module is set in the kernel space; a fourth processing module 66, which is used to summarize and analyze the first observation data, the second observation data and the third observation data to obtain the operating status of the system.
[0091] In some embodiments of the present application, the step of the first processing module 60 obtaining the first observation data of the microservice application in the microservice layer of the system includes: determining the first tracking information of the network data packet in the system, wherein the first tracking information is used to indicate the transmission path information of the network data packet in the microservice layer; using a neural network model to filter the initial first observation data, remove noise data in the first observation data, and obtain the first observation data; using a neural network model to process the first observation data and the tracking information to obtain a microservice call pattern recognition result, wherein the microservice call pattern recognition result is used to indicate whether there is an abnormal microservice call pattern in the microservice layer of the system.
[0092] In some embodiments of the present application, the step of the first processing module 60 determining the first tracing information of the east-west network data packet of the microservice application includes: adding a distributed tracing context to the header of the network data packet of the microservice application; determining the initial first tracing information of the network data packet based on the distributed tracing context; using a neural network model to filter the initial first tracing information, remove noise information in the initial first tracing information, and obtain the first tracing information.
[0093] In some embodiments of the present application, the second processing module 62 is further used to: identify the tracing context information of the network data packet in the system through the first extended Berkeley filter proxy module, and obtain second tracing information of the network data packet, wherein the second tracing information includes transmission path information of the network data packet in the container network layer.
[0094] In some embodiments of the present application, after obtaining the second observation data of the node application through the first extended Berkeley filter proxy module, the second processing module 62 is further used to: use a neural network model to process the second observation data and the second tracking information to determine the fault type of the system and the fault cause corresponding to the fault type, wherein the fault type includes at least one of the following: network congestion, packet loss, and network delay higher than a preset delay threshold.
[0095] In some embodiments of the present application, the step of the third processing module 64 obtaining the third observation data of each node through the second extended Berkeley filter proxy module includes: obtaining the third observation data of each node through the second extended Berkeley filter proxy module and the data exporter; using a neural network model to process the third observation data to obtain the correlation relationship between the internal data of the third observation data.
[0096] In some embodiments of the present application, the first observation data includes at least one of the following: indicator data or log data of the microservice application, the second observation data includes communication performance indicator data of the node application, and the third observation data includes perception container telemetry data of the node.
[0097] It should be noted that the various modules in the above-mentioned system status acquisition device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.
[0098] According to an embodiment of the present application, a non-volatile storage medium is also provided, in which a program is stored, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the following system status acquisition method: in the microservice layer of the system, first observation data of the microservice application in the microservice layer of the system is obtained; second observation data of the node application in the container network layer of the system is obtained through the first extended Berkeley filter proxy module, wherein the first extended Berkeley filter proxy module is set in the kernel space of each node in the cluster node; third observation data of the node in the infrastructure layer of the system is obtained through the second extended Berkeley filter proxy module, wherein the second extended Berkeley filter proxy module is set in the kernel space; the first observation data, the second observation data and the third observation data are summarized and analyzed to obtain the operating status of the system.
[0099] According to an embodiment of the present application, an electronic device is also provided, including a memory and a processor, the processor being used to run a program stored in the memory, wherein the following method for obtaining the system status is executed when the program is running: in the microservice layer of the system, first observation data of the microservice application in the microservice layer of the system is obtained; second observation data of the node application in the container network layer of the system is obtained through a first extended Berkeley filter proxy module, wherein the first extended Berkeley filter proxy module is set in the kernel space of each node in the cluster node; third observation data of the node in the infrastructure layer of the system is obtained through a second extended Berkeley filter proxy module, wherein the second extended Berkeley filter proxy module is set in the kernel space; the first observation data, the second observation data and the third observation data are summarized and analyzed to obtain the operating status of the system.
[0100] According to an embodiment of the present application, a computer program product is also provided, including a computer program, which implements the following method for obtaining system status when executed by a processor: in the microservice layer of the system, first observation data of the microservice application in the microservice layer of the system is obtained; second observation data of the node application in the container network layer of the system is obtained through a first extended Berkeley filter proxy module, wherein the first extended Berkeley filter proxy module is set in the kernel space of each node in the cluster node; third observation data of the node in the infrastructure layer of the system is obtained through a second extended Berkeley filter proxy module, wherein the second extended Berkeley filter proxy module is set in the kernel space; the first observation data, the second observation data and the third observation data are summarized and analyzed to obtain the operating status of the system.
[0101] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0102] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0103] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0104] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0105] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0106] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for obtaining system status, characterized in that: include: At the microservice layer of the system, first observation data of a microservice application in the microservice layer of the system is obtained; Acquire second observation data of the node application at the container network layer of the system through a first extended Berkeley filter proxy module, wherein the first extended Berkeley filter proxy module is set in the kernel space of each node in the cluster node; Acquiring third observation data of the node at the infrastructure layer of the system through a second extended Berkeley filter proxy module, wherein the second extended Berkeley filter proxy module is set in the kernel space; The first observation data, the second observation data, and the third observation data are summarized and analyzed to obtain the operating status of the system.
2. The method for obtaining system status according to claim 1, characterized in that: Obtaining the first observation data of the microservice application in the microservice layer of the system includes: Determine first tracking information of a network data packet in the system, wherein the first tracking information is used to indicate transmission path information of the network data packet at the microservice layer; Using a neural network model to filter the initial first observation data, removing noise data in the first observation data, and obtaining the first observation data; A neural network model is used to process the first observation data and the tracking information to obtain a microservice call pattern recognition result, wherein the microservice call pattern recognition result is used to indicate whether there is an abnormal microservice call pattern in the microservice layer of the system.
3. The method for obtaining system status according to claim 2, characterized in that: Determining first tracing information of east-west network packets of the microservice application includes: Adding distributed tracing context to the network packet header of the microservice application; determining initial first tracing information of the network data packet according to the distributed tracing context; The neural network model is used to filter the initial first tracking information, remove noise information in the initial first tracking information, and obtain the first tracking information.
4. The method for obtaining system status according to claim 1, wherein: The method further comprises: The first extended Berkeley filter proxy module identifies tracing context information of the network data packet in the system to obtain second tracing information of the network data packet, wherein the second tracing information includes transmission path information of the network data packet in the container network layer.
5. The method for obtaining system status according to claim 4, characterized in that: After obtaining the second observation data of the node application through the first extended Berkeley filter proxy module, the method further includes: A neural network model is used to process the second observation data and the second tracking information to determine a fault type of the system and a fault cause corresponding to the fault type, wherein the fault type includes at least one of the following: network congestion, packet loss, and network delay higher than a preset delay threshold.
6. The method for obtaining system status according to claim 1, characterized in that: Acquiring the third observation data of each of the nodes through the second extended Berkeley filter proxy module includes: Acquire the third observation data of each of the nodes through the second extended Berkeley filter proxy module and the data exporter; The third observation data is processed using a neural network model to obtain the correlation relationship between internal data of the third observation data.
7. The method for obtaining system status according to claim 1, characterized in that: The first observation data includes at least one of the following: indicator data or log data of the microservice application, the second observation data includes communication performance indicator data of the node application, and the third observation data includes perception container telemetry data of the node.
8. A device for acquiring system status, characterized in that: include: A first processing module is configured to obtain, at the microservice layer of the system, first observation data of a microservice application in the microservice layer of the system; a second processing module, configured to obtain second observation data of the node application at the container network layer of the system through a first extended Berkeley filter proxy module, wherein the first extended Berkeley filter proxy module is set in the kernel space of each node in the cluster node; a third processing module, configured to obtain third observation data of the node at the infrastructure layer of the system through a second extended Berkeley filter proxy module, wherein the second extended Berkeley filter proxy module is set in the kernel space; The fourth processing module is used to summarize and analyze the first observation data, the second observation data and the third observation data to obtain the operating status of the system.
9. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the method for obtaining the system status according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the method for obtaining the system status according to any one of claims 1 to 7 is executed when the program is run.
11. A computer program product, characterized in that The invention comprises a computer program, which implements the method for obtaining the system status according to any one of claims 1 to 7 when being executed by a processor.