A distributed full-link observation system for an AI agent

CN122672979APending Publication Date: 2026-09-01SHANGHAI JIEYUE JIYUAN INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611132989.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0005]针对现有技术存在的不足,本申请提供一种面向AI智能体的分布式全链路观测系统,至少用以解决现有技术中多源行为数据分散且缺乏跨层关联,导致无法实现AI智能体从语义交互到底层执行的全链路统一观测的问题

Benefits of technology

[0009]与现有技术相比,本申请实施例提供的方案中,通过构建由边缘采集层与中心分析层组成的分布式全链路观测系统,实现了对AI智能体运行行为的跨层级统一采集与关联分析。在计算节点操作系统内核中加载基于扩展伯克利数据包过滤器的内核态探针,使得无需对业务代码进行修改即可对AI智能体相关进程的系统调用、网络通信及执行行为进行实时采集,并结合用户态的协议语义解析处理及上下文信息关联处理,将原始观测数据转化为携带上下文元数据的结构化观测数据,从而提升了数据采集的完整性以及对上层语义信息的表达能力。通过中心分析层对多节点的结构化观测数据进行统一汇聚,并结合外部系统审计数据进行跨源关联分析,使审计层的主体信息能够与系统底层执行行为建立对应关系,进而实现了从语义交互到内核执行的全链路贯通。基于该技术方案,能够有效消除现有技术中应用层监控与系统层监控之间的割裂状态,提高了AI智能体执行行为的可观测粒度与关联准确性,为行为分析、安全检测及故障定位提供一致的数据基础,从而提升了分布式环境下智能体系统的可观测性与运行可控性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122672979A_ABST
    Figure CN122672979A_ABST
Patent Text Reader

Abstract

The application provides a distributed full-link observation system for an AI agent, comprising: an edge collection layer configured to be distributedly deployed on multiple computing nodes, the edge collection layer collecting behavior data of an AI agent related process through a kernel mode probe loaded on a corresponding computing node, performing preprocessing on the collected data in a user mode, and outputting structured observation data carrying context metadata; and a central analysis layer in communication connection with the edge collection layer, configured to converge the structured observation data from the multiple computing nodes, and perform cross-source correlation analysis on the structured observation data and system audit data obtained externally, so as to establish a corresponding relationship between subject information identified in the system audit data and running behavior in the structured observation data. The application improves the observable granularity and correlation accuracy of the AI agent execution behavior, and improves the observability and running controllability of the agent system in a distributed environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of cloud-native observability and artificial intelligence security technology, and in particular to a distributed end-to-end observation system for AI agents. Background Technology

[0002] With the widespread application of AI agents driven by large language models in cloud-native environments, these agents are gradually acquiring autonomous decision-making and tool invocation capabilities. Their operation involves multiple levels, including model inference, external tool interaction, and underlying system execution. To ensure system stability and security, it is necessary to uniformly observe and analyze the behavior of the agent at different levels, thereby achieving traceability and monitoring of its execution process.

[0003] Existing observable methods for AI agents mainly include application-layer instrumentation-based monitoring schemes and traditional infrastructure-layer monitoring schemes. Application-layer instrumentation schemes typically rely on embedding SDKs in business code to collect model inference and interaction data, which suffers from strong invasiveness to business logic, complex deployment, and difficulty in covering all operational behaviors. While infrastructure monitoring schemes can collect system calls and resource metrics, they lack the ability to understand the semantic interaction process of the agent, making it difficult to establish a correlation between upper-layer intent and lower-layer execution. Furthermore, the independence of different data sources and the lack of a unified correlation mechanism make it difficult to reconstruct and analyze the complete execution chain of the agent in a distributed environment, thus creating a disconnect between semantics and system behavior.

[0004] Therefore, how to achieve unified collection, cross-source association, and centralized analysis of the entire chain of AI agents in a distributed environment, from semantic interaction to underlying execution, without intruding on business code, has become an urgent technical problem to be solved. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this application provides a distributed end-to-end observation system for AI agents, which at least solves the problem that the dispersed multi-source behavioral data and lack of cross-layer correlation in existing technologies make it impossible to achieve unified end-to-end observation of AI agents from semantic interaction to underlying execution.

[0006] To achieve the above objectives and other advantages, some embodiments of this application provide a distributed end-to-end observation system for AI agents, including:

[0007] The edge acquisition layer is configured to be distributed and deployed on multiple computing nodes. The edge acquisition layer collects behavioral data of AI agent-related processes running on the computing nodes by loading a kernel-mode probe based on the extended Berkeley packet filter into the operating system kernel of the corresponding computing node. In user mode, the collected data is preprocessed, including protocol semantic parsing and context information association, and structured observation data carrying context metadata is output.

[0008] The central analysis layer, which is communicatively connected to the edge acquisition layer, is used to aggregate structured observation data from multiple computing nodes and perform cross-source correlation analysis on the structured observation data and externally acquired system audit data to establish the correspondence between the subject information identified in the system audit data and the operational behavior in the structured observation data, thereby constructing a full-link observation model that reflects the AI ​​agent from semantic interaction to kernel execution.

[0009] Compared with existing technologies, the solution provided in this application, by constructing a distributed end-to-end observation system composed of an edge acquisition layer and a central analysis layer, achieves unified cross-level acquisition and correlation analysis of the operational behavior of AI agents. A kernel-mode probe based on an extended Berkeley packet filter is loaded into the operating system kernel of the computing nodes, enabling real-time acquisition of system calls, network communications, and execution behaviors of AI agent-related processes without modifying the business code. Combined with user-mode protocol semantic parsing and context information correlation processing, the raw observation data is transformed into structured observation data carrying context metadata, thereby improving the completeness of data acquisition and the ability to express upper-layer semantic information. The central analysis layer uniformly aggregates the structured observation data from multiple nodes and performs cross-source correlation analysis with external system audit data, enabling the main information of the audit layer to establish a correspondence with the underlying system execution behavior, thus achieving end-to-end connectivity from semantic interaction to kernel execution. Based on this technical solution, the disconnect between application-layer monitoring and system-layer monitoring in existing technologies can be effectively eliminated, improving the observable granularity and correlation accuracy of AI agent execution behavior, providing a consistent data foundation for behavior analysis, security detection and fault location, thereby enhancing the observability and operational controllability of agent systems in a distributed environment. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other implementation methods can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of the overall architecture of a distributed end-to-end observation system for AI agents provided in an embodiment of this application;

[0012] Figure 2 This illustration shows a schematic diagram of the multidimensional deviation analysis and risk assessment process based on full-link behavioral data in an embodiment of this application.

[0013] Figure 3 This paper illustrates a logical diagram of an asset interaction structure centered on an AI agent, as shown in an embodiment of this application.

[0014] Figure 4 The diagram illustrates the logic of establishing a full-link tracing structure based on a multi-layer observation model in an embodiment of this application. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0016] The following terms are used in this document.

[0017] AI agents are autonomous software entities based on artificial intelligence technology. They can perceive their environment by receiving input information and perform semantic understanding, decision-making, planning, and action execution based on that information to complete preset task objectives. AI agents are typically based on large language models or other intelligent models and possess a certain degree of autonomy and task execution capabilities. They can call external tools, access remote services, or interact with other intelligent agents, thereby forming a continuous task-oriented execution process.

[0018] The Extended Berkeley Packet Filter (eBPF) is a programmable technology mechanism that runs within the operating system kernel. This mechanism allows users to extend the functionality of the operating system kernel without modifying the kernel source code or loading kernel modules by loading a validated program into the kernel. eBPF programs typically execute in an event-driven manner and can be attached to critical locations such as system call paths, network protocol stacks, or kernel trace points, capturing and processing system behavior when relevant events are triggered. Based on this mechanism, efficient observation and data acquisition of the system's operation can be achieved in kernel space, making it widely applicable in scenarios such as system observability, performance analysis, and network monitoring.

[0019] A data acquisition probe is a program unit loaded onto the operating system kernel or user-space execution path based on the eBPF mechanism. It is used to capture target behaviors and generate corresponding behavioral event data. Data acquisition probes can be deployed in different locations depending on the target being acquired. For example, they can be mounted on the system call path to capture process execution and system call events, on the network data processing path to capture network communication events, or on the user-space function execution path to obtain application-layer interaction data. Through the configuration and scheduling of data acquisition probes, selective acquisition of specific processes, network connections, or protocol behaviors can be achieved, thereby obtaining multi-dimensional data related to agent behavior while ensuring system stability.

[0020] Cloud-native environments typically encapsulate applications using container technology, allowing applications and their dependencies to run as images on compute nodes. A container orchestration platform (such as a cluster scheduling and management system) provides unified scheduling, deployment, and lifecycle management for containers across multiple compute nodes. Combined with mechanisms such as service discovery, elastic scaling, and load balancing, this enables dynamic scaling and high availability of applications. In this environment, applications are often deployed as microservices, with different services communicating over a network, exhibiting distributed, dynamic, and highly automated operational characteristics.

[0021] Reference Figure 1 As shown, some embodiments of this application relate to a distributed end-to-end observation system for AI agents. Deployed in a cloud-native environment, the system includes a distributed edge acquisition layer and a central analysis layer connected to it, used to achieve end-to-end observation and analysis of the operational behavior of AI agents. The system includes:

[0022] The edge acquisition layer is configured to be distributed across multiple computing nodes. By loading a kernel-mode probe based on an extended Berkeley packet filter into the operating system kernel of the corresponding computing node, the edge acquisition layer collects behavioral data of AI agent-related processes running on the computing node. In user mode, it performs preprocessing on the collected data, including protocol semantic parsing and context information association, and outputs structured observation data carrying context metadata.

[0023] The central analysis layer, which communicates with the edge acquisition layer, is used to aggregate structured observation data from multiple computing nodes and perform cross-source correlation analysis between the structured observation data and externally acquired system audit data to establish the correspondence between the subject information identified in the system audit data and the operational behavior in the structured observation data, thereby constructing a full-link observation model that reflects the AI ​​agent from semantic interaction to kernel execution.

[0024] In this embodiment, the edge acquisition layer is deployed in a distributed manner across each compute node. Specifically, a containerized deployment approach can be adopted, deploying the edge acquisition component as a daemon set within a container orchestration platform, so that each compute node automatically runs a corresponding acquisition instance. Each acquisition instance is automatically loaded and run when the node starts up or joins the cluster, thereby achieving full coverage acquisition capability across all compute nodes in the cluster.

[0025] In each computing node, the edge acquisition layer collects process behavior data related to the AI ​​agent by loading kernel-mode probes based on the Extended Berkeley Packet Filter (eBPF) into the operating system kernel. For example, probes are attached to system call entry points (such as execve, open, connect, etc.), network communication paths (such as socket sending and receiving), and process scheduling-related locations to capture process execution behavior, network interaction behavior, and resource access behavior. These probes can be dynamically injected into the kernel and run based on the kernel's security verification mechanism, thereby obtaining the corresponding raw behavioral data without affecting the normal execution of business processes. Through this method, process execution information, network communication data, and runtime context information related to the AI ​​agent can be collected in real time.

[0026] After collecting the raw behavioral data, the edge acquisition layer preprocesses the data in user space. Preprocessing includes protocol semantic parsing and context information association. Specifically, the collected network communication data is reconstructed into byte streams and restored to application-layer protocol objects, from which semantic information such as request content and response data is extracted. Simultaneously, based on process identification information and combined with interfaces provided by the container orchestration platform, the corresponding container identifier, namespace, and service identity information are obtained, thereby establishing the association between process behavior and the runtime environment. Through the above processing, the raw collected data is transformed into structured observation data carrying context metadata, which identifies the executing entity, runtime environment, and source of behavior.

[0027] In terms of data output, the edge acquisition layer can use message queues or streaming mechanisms to send structured observation data to the central analysis layer. For example, structured observation data can be continuously sent in a streaming manner through a publish-subscribe messaging system or a transport protocol based on HTTP / GPPC, thereby achieving real-time data transmission and decoupled processing.

[0028] In this embodiment, the central analysis layer is deployed on a centralized processing node and communicates with edge acquisition layers on multiple computing nodes via a network. The central analysis layer receives structured observation data from each computing node and aggregates it uniformly. Specifically, data from different nodes can be accessed through a data receiving service and merged according to timestamps or source nodes to form a unified data stream. Simultaneously, the central analysis layer can also obtain system audit data from external systems, such as audit logs generated by a container orchestration platform, which record request subject information, operational behavior, and target resource identifiers.

[0029] The central analysis layer performs cross-source correlation analysis based on structured observation data and system audit data. By aligning the two types of data in the time dimension, matching them in the resource identification dimension, and analyzing the consistency in behavioral semantics, a correspondence is established between the subject information identified in the system audit data and the actual execution behavior recorded in the structured observation data. For example, within a preset time window, a scheduling request is correlated with the process execution behavior on the corresponding computing node to determine the initiating entity of the execution behavior. Based on the above correlation results, a full-link observation model reflecting the execution process of the AI ​​agent is constructed. The full-link observation model uniformly organizes data from different levels, enabling various behaviors generated by the AI ​​agent during task execution to be linked together according to logical relationships, thus forming a complete execution path expression. Through this model, the full-link observability of the AI ​​agent's operational behavior can be achieved, providing a data foundation for subsequent behavior analysis, security detection, and anomaly localization.

[0030] The solution provided in this application constructs a distributed end-to-end observation system consisting of an edge acquisition layer and a central analysis layer, achieving unified cross-level acquisition and correlation analysis of the operational behavior of AI agents. A kernel-mode probe based on an extended Berkeley packet filter is loaded into the operating system kernel of the computing nodes. This allows for real-time acquisition of system calls, network communications, and execution behaviors of AI agent-related processes without modifying the business code. Combined with user-mode protocol semantic parsing and context information correlation processing, the raw observation data is transformed into structured observation data carrying context metadata, thereby improving the completeness of data acquisition and the ability to express upper-layer semantic information. The central analysis layer uniformly aggregates the structured observation data from multiple nodes and performs cross-source correlation analysis with external system audit data. This enables the main information of the audit layer to establish a correspondence with the underlying system execution behavior, thus achieving end-to-end connectivity from semantic interaction to kernel execution. Based on this technical solution, the disconnect between application-layer monitoring and system-layer monitoring in existing technologies can be effectively eliminated, improving the observable granularity and correlation accuracy of AI agent execution behavior, providing a consistent data foundation for behavior analysis, security detection and fault location, thereby enhancing the observability and operational controllability of agent systems in a distributed environment.

[0031] In a preferred embodiment, the edge acquisition layer includes a kernel-mode acquisition module and a user-mode processing module, wherein:

[0032] The kernel-mode acquisition module is used to acquire system call events, network communication data, and data generated during the execution of user-mode encryption library functions of AI agent-related processes by attaching kernel-mode probes.

[0033] In this embodiment, the kernel-mode acquisition module can include three types of probes: process execution probes, user-mode probes, and socket filters. These probes are deployed on different execution paths to collect behavioral data at different levels. Specifically, the process execution probes are attached to locations related to the process lifecycle within the operating system kernel. For example, they are attached to the system call entry point `sys_enter_execve` related to program execution to capture new process creation behavior; to the system call entry points `sys_enter_clone` or `sys_enter_fork` related to process forking to capture process forking events; to the system call entry point `sys_enter_connect` related to network connection establishment to capture network connection establishment behavior; and to the system call entry point `sys_enter_bind` related to port binding to capture port binding behavior. When the process containing the AI ​​agent is created or a new program is executed, the process execution probes are triggered to obtain information such as the corresponding process identifier, parent process identifier, execution command, and execution path. When the process executes a system call, the process execution probes can also obtain the system call type and corresponding parameter information.

[0034] User-space probes can be attached to the execution path of user-space functions, especially the execution locations of encryption library functions related to network communication, based on the eBPF user-space probe mechanism. In one specific implementation, user-space probes can be attached to the entry points of key functions in common encryption libraries, such as the `SSL_write` and `SSL_read` functions in the OpenSSL library, corresponding to the encryption processing stage before data transmission and the decryption processing stage after data reception, respectively. They can also be attached to data write and read functions in the Go language's `crypto / tls` library, or corresponding data processing functions in other encryption libraries, thus achieving unified coverage of different programming languages ​​and encryption implementations. When the AI ​​agent interacts with external services through an encrypted communication protocol, the user-space probe is triggered during function execution. During the data transmission stage, before the encryption processing function is executed, the user-space probe reads the data to be sent from the memory register or buffer corresponding to the function call context; during the data reception stage, after the decryption processing function is executed, the user-space probe reads the decrypted data from the function call context. In this way, application-layer interaction content can be obtained without changing the communication protocol or introducing an intermediate proxy.

[0035] Socket filters can be deployed in the operating system's socket layer or kernel protocol stack based on the eBPF mechanism to filter and collect data as network packets pass through this layer. Preset filtering rules can be set at the socket layer to match data traffic corresponding to target port numbers or communication protocol types, thereby achieving selective collection of network communication data related to the AI ​​agent. When the process containing the AI ​​agent establishes a network connection, sends data, or receives data, the corresponding data packets are transmitted in the kernel protocol stack, and the socket filter is triggered on the data packet processing path. For data packets that meet the filtering rules, they can be copied and written to the kernel's data storage structure to achieve temporary storage of network data. Through the collection of the above information, raw behavioral event data describing the entire network communication process can be formed. Therefore, the kernel-mode acquisition module can generate raw behavioral data across kernel mode, user mode, and the network layer.

[0036] The user-space processing module is used to perform context information association processing, including: obtaining the corresponding container orchestration platform metadata based on the process identifier in the system call event, and establishing the association between the process identifier and the metadata.

[0037] In this embodiment, after acquiring the raw behavioral data, the user-space processing module performs context information association processing. Specifically, process identification information and network connection information can be extracted from the raw behavioral data, and the corresponding process runtime environment can be determined based on the process identification information. In a containerized deployment scenario, a runtime instance is a resource entity in a container orchestration platform used to host application execution. The corresponding runtime instance identifier can be obtained based on the process identification information, and further, by interacting with the container orchestration platform, metadata of the container orchestration platform corresponding to the runtime instance identifier can be obtained, thereby determining the runtime environment and business context to which the process belongs.

[0038] In specific implementations, the mapping relationship between running instances and business resources can be obtained by calling the interface services provided by the container orchestration platform or by listening to the node proxy service. Alternatively, a local cache can be maintained in user space to store the mapping relationship between running instance identifiers and the metadata of the container orchestration platform, and this mapping relationship can be dynamically updated. In one specific implementation, the container orchestration platform is a Kubernetes platform, and its metadata is Kubernetes metadata. Kubernetes metadata can include Pod names, namespaces, service accounts, tag information, annotation information, and information about the controller resources to which it belongs. In other implementations, the container orchestration platform can also be DockerSwarm, Apache Mesos, HashiCorp Nomad, or a container orchestration platform provided by a cloud vendor. Correspondingly, the metadata of the container orchestration platform can include structured information used to identify container instances, task instances, or service resources, such as service names, task identifiers, node identifiers, and resource scheduling information.

[0039] The user-space processing module can query the corresponding container orchestration platform's metadata in the local cache based on the process identification information carried in the original behavior data, and associate the query results with the original behavior data to supplement the business context information corresponding to the event. For example, when an AI agent runs in a container environment and initiates network communication behavior, the original behavior data may contain the corresponding process identification information. By querying the mapping relationship synchronized by the container orchestration platform, the Pod to which the process belongs and its business affiliation information can be determined, and the Pod name, namespace, and controller resource information can be added to the original behavior data, thereby establishing the association between the process identification and the container orchestration platform's metadata.

[0040] Through the above processing, the metadata of the original behavioral data is enriched, transforming it from containing only process-level execution information into structured observation data carrying container orchestration platform context information.

[0041] In a preferred embodiment, the user-space processing module is further configured to perform protocol semantic parsing processing, including:

[0042] The protocol reorganization is performed on network communication data and data generated during the execution of encryption library functions to generate corresponding application layer protocol objects, and semantic payload information is extracted from the application layer protocol objects; and the model resource consumption of AI agents is measured based on semantic payload information to obtain corresponding resource consumption indicators, and the correlation, semantic payload information and resource consumption indicators are encapsulated as attribute fields into structured observation data.

[0043] In this embodiment, the user-space processing module categorizes the collected data based on the session identifier of the network connection. The session identifier may include information such as source address, destination address, source port, destination port, and transport protocol type. For data within the same session, data segments are sorted and concatenated according to the arrival order or sequence number of data packets to recover a continuous byte stream.

[0044] After reassembling the byte stream, the user-space processing module parses the byte stream based on preset protocol parsing rules to identify the corresponding application layer protocol type and generate the corresponding application layer protocol object. In one implementation, when identified as an HTTP protocol, the byte stream can be segmented according to the request line, request header, and message body structure to construct an HTTP request object or response object; when identified as a database access protocol, an SQL request object can be generated based on the query statement structure; when identified as a communication protocol between an AI agent and external tools, a corresponding interaction protocol object can be generated according to a predefined message format.

[0045] In scenarios involving encrypted communication, the user-space processing module can also utilize data collected during the execution of user-space encryption library functions to directly obtain the decrypted application-layer data content, thereby avoiding protocol parsing of the ciphertext data. Based on this, the decrypted data is processed in a unified manner with network communication data to complete protocol reassembly and generate the corresponding application-layer protocol object.

[0046] The application layer protocol object is parsed to extract information that characterizes the semantics of AI agent interaction, such as request parameters, response results, call identifiers, and interaction content, as semantic payload information. In one specific implementation, when the application layer protocol object corresponds to a model call process, the call identifier information, interaction content, and resource usage information related to the model call can be obtained by parsing the fields in the request and response messages. Specifically, when parsing the response message, the usage field can be extracted, and further parsed based on the usage field to obtain the number of input tokens corresponding to prompt_tokens and the number of output tokens corresponding to completion_tokens. The total_tokens can also be extracted as supplementary information on the total token consumption. When parsing the request message, model identifier information can be extracted from the request path, request header, or request body to identify the target model currently being called.

[0047] The user-state processing module measures the model resource consumption of the AI ​​agent based on semantic payload information. Specifically, it determines the start and end times of a single model call based on the time information in the structured behavioral event data, and uses the time interval defined by these start and end times as a resource measurement window. It extracts computing resource usage information falling within this time interval from the structured behavioral event data, aggregates the corresponding computing resource usage information, and obtains the computing resource consumption value. Subsequently, based on the model identification information, the number of input tokens and output tokens are determined as digital asset consumption indicators, and the computing resource consumption value is determined as a physical resource consumption indicator. The digital asset consumption indicator and the physical resource consumption indicator are then correlated to generate the corresponding resource consumption indicator.

[0048] After completing the above processing, the user-state processing module encapsulates the relationships formed during context information association processing, the semantic payload information obtained from protocol semantic parsing, and the resource consumption indicators obtained from resource metering as attribute fields to generate corresponding structured observation data. Through this processing, the structured observation data not only characterizes the interactive semantics and operational context of the AI ​​agent but also the resource usage during model invocation, thus providing fundamental data support for subsequent cross-source correlation analysis and end-to-end observation.

[0049] In a preferred embodiment, the central analysis layer includes a correlation processing module, which is used for:

[0050] Based on a preset time window, the system audit data and structured observation data are matched in terms of time dimension. Based on the resource identifiers in the system audit data and the metadata carried in the structured observation data, the system audit data and structured observation data are matched in terms of spatial dimension. Based on the request behavior information in the system audit data and the execution behavior information in the structured observation data, the system audit data and structured observation data are matched in terms of behavior consistency. When the matching results of each dimension meet the preset conditions, the system audit data and structured observation data are associated with the subject information.

[0051] In this embodiment, the correlation processing module performs time-dimensional matching processing on system audit data and structured observation data within a preset time window. The system audit data includes a request timestamp, and the structured observation data includes an execution timestamp. The correlation processing module calculates the time difference based on the request and execution timestamps and determines whether the execution timestamp is later than the request timestamp and whether the time difference is less than the preset time window. When the above conditions are met, the corresponding structured observation data is determined to meet the time constraint conditions and is thus used as candidate correlation data. In some implementations, the time of data from different systems can also be aligned to a unified benchmark to eliminate the impact of inter-system clock deviations on the matching results.

[0052] Building upon the temporal dimension matching, the association processing module further performs spatial dimension matching based on the resource identifiers in the system audit data and the metadata carried in the structured observation data. The resource identifiers in the system audit data characterize the target resource pointed to by the operation request, while the metadata in the structured observation data characterizes the actual execution environment in which the runtime behavior occurs. Since data from different sources differ in their resource identifier representations, the association processing module can convert the runtime resource identifiers in the structured observation data into logical matching identifiers based on a preset resource identifier mapping relationship, and then compare the logical matching identifiers with the resource identifiers in the system audit data. When both are consistent in preset fields, it is determined that the runtime behavior occurred at the target resource, thus confirming that the system audit data and the structured observation data meet the spatial dimension consistency condition. In some implementations, the resource identifier mapping relationship can be dynamically obtained through container orchestration platform interfaces, node-side caches, or container runtime environment information, and continuously updated.

[0053] After completing time-dimension and spatial-dimension matching, the association processing module further performs behavior consistency matching based on request behavior information in system audit data and execution behavior information in structured observation data. Specifically, behavior feature parsing can be performed on request behavior information and execution behavior information to obtain behavior instruction information and execution parameter information respectively. The execution parameter information is then normalized to eliminate the influence of path differences, parameter escaping, shell wrapping, or default execution environment, thereby obtaining a standardized execution representation. Based on this, consistency matching is determined based on behavior instruction information and standardized execution representation. The corresponding data is determined to meet the behavior consistency condition when any of the following conditions are met: the behavior instruction information and the standardized execution representation are semantically consistent under a unified expression form; or the behavior instruction information is contained in the standardized execution representation; or the standardized execution representation conforms to the default execution environment characteristics corresponding to the behavior instruction information.

[0054] In the multidimensional matching process described above, the association processing module can filter candidate data layer by layer in the order of time dimension matching, spatial dimension matching, and behavioral consistency matching, thereby gradually narrowing down the candidate association range, improving association accuracy and reducing the probability of mismatch.

[0055] When the time dimension matching results, spatial dimension matching results, and behavioral consistency matching results all meet preset conditions, the association processing module establishes an association between the subject information in the system audit data and the operational behavior in the structured observation data. Specifically, the subject information extracted from the system audit data can be bound to the operational behavior recorded in the structured observation data to obtain the correspondence between the subject and the execution behavior. In some implementations, association identification data can also be generated based on the association relationship to identify the binding relationship between the subject and the operational behavior, and to support subsequent association queries and traceability analysis.

[0056] Through the above processing, comprehensive matching of data from different observation levels in the time, space and behavioral semantic dimensions is achieved. Without modifying the business system, cross-source association between the subject information in the system audit data and the operational behavior in the structured observation data is realized.

[0057] In a preferred embodiment, the central analysis layer further includes a behavioral baseline analysis module, which is used for:

[0058] Multidimensional feature modeling is performed on the historical normal behavior data of the AI ​​agent, and a corresponding behavioral baseline model is constructed based on the correlation between different feature dimensions. The multidimensional features include at least semantic interaction features, tool call features, resource usage features, and system execution features. The behavioral baseline model is used to analyze the degree of deviation between the real-time observation data generated by the AI ​​agent during operation and the behavioral baseline model.

[0059] The behavioral baseline model includes a semantic interaction baseline built based on historical semantic interaction data, a behavioral sequence baseline built based on historical tool call data, a resource consumption baseline built based on historical resource usage data, and a system execution behavior baseline built based on historical system call data.

[0060] Reference Figure 2 As shown, this embodiment illustrates the process of AI agent behavior baseline analysis and risk assessment. In this embodiment, behavioral data of the AI ​​agent during its historical operation can be collected and filtered. Task execution data that is in a stable operating state and has not triggered abnormal alarms is selected as historical normal behavior data to ensure the accuracy and representativeness of the constructed behavior baseline model. Historical normal behavior data may include semantic interaction data, tool call data, resource usage data, and system execution data. These data types are correlated based on a unified time identifier to form a historical behavior data set covering the complete task execution chain.

[0061] Before modeling historical normal behavior data, the data can be cleaned, aligned, and standardized. This includes removing abnormal noise data, standardizing data formats, and mapping and transforming fields from different data sources to construct a unified multidimensional behavior dataset. Based on this, the multidimensional behavior dataset is divided according to preset analysis dimensions to extract behavioral features from different dimensions. In this embodiment, the multidimensional features include at least semantic interaction features, tool invocation features, resource usage features, and system execution features. Semantic interaction features characterize the AI ​​agent's behavior at the interaction intent level; tool invocation features characterize the AI ​​agent's behavioral path characteristics during task execution; resource usage features characterize the AI ​​agent's resource usage during operation; and system execution features characterize the AI ​​agent's execution behavior characteristics within the underlying system environment.

[0062] In one specific embodiment, the behavioral baseline model may include multiple sub-baselines corresponding to each feature dimension. Specifically, it may include a semantic interaction baseline constructed based on historical semantic interaction data, a behavioral sequence baseline constructed based on historical tool call data, a resource consumption baseline constructed based on historical resource usage data, and a system execution behavior baseline constructed based on historical system call data. For the semantic interaction baseline, multi-dimensional semantic features of the AI ​​agent during historical normal periods can be extracted. These multi-dimensional semantic features include at least system prompts, user input, thought chain reasoning steps, and model response content. Furthermore, a pre-defined embedding model is used to transform the multi-dimensional semantic features into high-dimensional semantic vectors, and the centroid vector and distribution radius are calculated based on the high-dimensional semantic vectors to characterize the distribution center and discrete range of normal semantic behavior.

[0063] For a behavioral sequence baseline, a behavioral state space can be defined to characterize the AI ​​agent's behavior during task execution, and historical tool call data can be mapped to corresponding behavioral state sequences. Based on this, the frequency of transitions between different behavioral states is statistically analyzed, and state transition probabilities are calculated to construct a state transition matrix representing the transition relationships between behavioral states, thereby performing sequence modeling of the AI ​​agent's normal behavioral paths. Through the above processing, a behavioral sequence baseline reflecting the probability distribution characteristics of normal behavioral paths and the characteristics of state transition patterns can be formed.

[0064] For the resource consumption baseline, resource usage data such as token consumption, request call frequency, CPU utilization, memory usage, and task execution time per unit time can be extracted from historical normal operation phases. This data is then statistically summarized based on time windows to form a resource usage characteristic sequence. Based on this, the statistical distribution characteristics of the resource usage features are extracted to characterize the resource usage range and fluctuation characteristics of the AI ​​agent under normal operating conditions.

[0065] For the system execution behavior baseline, system call behavior data triggered by the AI ​​agent during the historical normal operation phase can be obtained, and the system call behavior data can be classified and structured to extract features such as call type, call frequency, call order and call context, thereby constructing a system execution behavior baseline that reflects the underlying execution mode and risk characteristics.

[0066] After constructing the behavioral baseline model, the behavioral baseline analysis module is also used to analyze the deviation between the real-time observation data generated by the AI ​​agent during operation and the behavioral baseline model. Specifically, it can acquire real-time observation data generated by the AI ​​agent during task execution, covering multiple observation dimensions such as the semantic interaction layer, tool invocation layer, resource usage layer, and system execution layer, and then perform correlation and comparative analysis with the behavioral baseline model.

[0067] In terms of semantic interaction features, the behavior baseline analysis module can extract semantic features based on real-time observation data and transform these features into real-time semantic vectors using a pre-defined embedding model. Semantic features include at least system prompts, user input, inference process information, and model response content, thus forming a vectorized representation of the current semantic behavior. Further, the similarity between the real-time semantic vector and the centroid vector in the semantic interaction baseline is calculated to characterize the degree of closeness of the current semantic behavior to the historical normal semantic distribution. In one implementation, cosine similarity can be used for measurement, and its calculation method is as follows:

[0068]

[0069] in, This represents the similarity value between the real-time semantic vector and the centroid vector; This represents the real-time semantic vector corresponding to the current semantic interaction; Represents the centroid vector in the semantic interaction baseline; The symbol represents the magnitude of the vector; "·" indicates the vector dot product operation. The similarity value reflects the degree of proximity between the current semantic behavior and the historical normal semantic center.

[0070] Based on semantic similarity values, deviation analysis can be performed on the semantic behavior of AI agents. When semantic similarity decreases significantly during continuous interactions, or undergoes a sudden change compared to the historical normal range, it indicates that the current semantic vector deviates more from the centroid vector, suggesting an abnormal change in the AI ​​agent's focus or expression pattern. In this case, it can be determined that the AI ​​agent exhibits semantic drift behavior, and its deviation index in the semantic interaction feature dimension can be determined accordingly.

[0071] In terms of tool call features, the behavior baseline analysis module maps tool call behaviors in real-time observation data to real-time behavior state sequences. Based on the state transition matrix in the behavior sequence baseline, it performs probabilistic modeling of the real-time behavior state sequences to calculate their overall occurrence probability. Furthermore, it compares the overall occurrence probability with the probability distribution characteristics of historical normal behavior paths. When the overall occurrence probability is lower than a preset probability threshold or does not belong to the normal distribution range, it determines that the behavior path has deviated. Simultaneously, it can also compare the state transition paths in the real-time behavior state sequences with historical high-frequency state transition patterns. When low-probability paths, no-occurrence paths, or abnormal loop structures exist, it determines the corresponding degree of deviation and quantifies the deviation index in conjunction with the path anomaly degree.

[0072] Based on this, different types of deviation information can be fused according to the degree of deviation of the overall occurrence probability relative to the probability distribution characteristics, and the matching difference of the real-time behavior state sequence on the state transition path relative to the state transition pattern characteristics, so as to determine the deviation index of the AI ​​agent in the tool call feature dimension.

[0073] In terms of resource usage characteristics, the behavioral baseline analysis module extracts resource usage features based on real-time observation data. These features include at least token consumption, request call frequency, computational resource usage, and task execution time. Furthermore, the resource usage features can be normalized and compared with the normal resource usage range represented by the resource consumption baseline. When resource indicators exceed the normal range or their trends significantly deviate from historical patterns, a corresponding deviation index is determined based on the magnitude and duration of the deviation. Additionally, multiple resource indicators can be weighted and fused to obtain a deviation index for the resource usage characteristic dimension.

[0074] In the system execution characteristic dimension, the behavior baseline analysis module extracts system call behavior characteristics based on system call monitoring data. These characteristics include call type, call frequency, call order, and call context information. Furthermore, the system call behavior characteristics are matched with the system execution behavior baseline to assess the degree of difference in call patterns and execution paths of the current system call behavior. In one implementation, risk scores can be assigned to different types of system calls based on a preset risk strategy, and these scores are accumulated based on their frequency of occurrence to determine the deviation index of the system execution characteristic dimension.

[0075] After obtaining the deviation indicators for each dimension, the behavioral baseline analysis module can also perform fusion analysis on the deviation information from different dimensions. Specifically, it can normalize the deviation indicators for each dimension and then perform a weighted combination of the normalized deviation indicators based on preset weight coefficients to obtain a risk assessment score that reflects the current operating status of the AI ​​agent.

[0076] In some implementations, the scores of each dimension can be fused and calculated in the following manner:

[0077]

[0078] in, This indicates that the AI ​​agent is at any moment The risk assessment score is used to characterize the degree of risk in its current security status; The semantic deviation score represents the semantic interaction dimension; The resource anomaly score represents the resource consumption dimension; This represents the risk score for the system execution dimension; The sequence anomaly score represents the behavioral sequence dimension; These represent the weight coefficients for the corresponding dimensions, used to adjust the contribution of different dimensions in the comprehensive risk assessment.

[0079] Based on the aforementioned fusion calculation, multi-dimensional deviation information can be transformed into a unified risk assessment score, thereby reflecting the overall safety level of the AI ​​agent's current operating status. Simultaneously, a weighting adjustment mechanism ensures that the influence of each dimension on the final risk assessment result is controllable, thus adapting to the security needs of different application scenarios.

[0080] Through the above-mentioned multi-dimensional correlation comparison and analysis mechanism, the deviation of the AI ​​agent's operating behavior from the historical normal behavior pattern can be reflected throughout the entire link. It can not only identify semantic layer anomalies and behavior path anomalies, but also detect resource consumption anomalies and underlying execution risks, thereby improving the accuracy of anomaly detection and the comprehensiveness of risk assessment.

[0081] In a preferred embodiment, the central analysis layer further includes an asset management module, which is used for:

[0082] Based on a pre-defined asset model, structured observation data and system audit data are analyzed to identify multiple types of asset objects related to AI agents, establish relationships between different asset objects, and construct a topological structure between asset objects. Among them, asset objects include at least agent asset objects, tool asset objects, and model service asset objects.

[0083] Specifically, the pre-defined asset model is used to define the classification rules for various asset objects in the AI ​​agent ecosystem and their corresponding identification features. Specifically, the asset model may include asset type definition rules, feature matching rules, and identifier extraction rules. The asset type definition rules specify the classification methods for different asset objects, the feature matching rules describe the behavioral or identifier features of various asset objects in structured observation data and system audit data, and the identifier extraction rules extract key information from the data that can uniquely identify asset objects. In this embodiment, three core asset objects are defined in the AI ​​agent ecosystem: agent asset objects, tool asset objects, and model service asset objects. During the parsing process, based on the pre-defined asset model, multi-dimensional feature information in structured observation data and system audit data can be matched and analyzed to achieve automatic identification of asset objects. Network communication behavior features, system call behavior features, tool call behavior features, and resource access features can be extracted from structured observation data, and combined with subject information, resource identifier information, and operation context information in system audit data to classify and identify different asset objects.

[0084] Among them, intelligent agent asset objects are used to represent workloads with autonomous reasoning capabilities. These objects can proactively initiate requests to large language models and invoke external tools to complete task execution based on the model's response. The identification of intelligent agent asset objects can be achieved in several ways. One approach is based on network communication behavior characteristics in structured observation data. When a target workload is detected frequently initiating outbound HTTPS connections to the model server endpoint, and its communication traffic exhibits a long text-based request-response interaction pattern, the workload can be identified as an intelligent agent asset object. Network communication behavior characteristics can be extracted by parsing network protocol data to obtain request paths, communication frequencies, and data payload characteristics. Another approach combines subject information and operational context information from system audit data for identification. The corresponding intelligent agent instance is determined by matching preset tags or annotations in the container orchestration platform. These tags or annotations can be obtained through system audit logs or platform metadata interfaces.

[0085] Furthermore, dependencies can be constructed based on resource access characteristics and tool call behavior characteristics in structured observation data. When a target node is in the central position in the service topology and connects to model services through network communication behavior characteristics and multiple tool services through tool call behavior characteristics, it can be determined to be an intelligent agent asset object.

[0086] Tool asset objects represent execution units invoked by intelligent agents, carrying the actual operational behaviors generated by the AI ​​agent during task execution. In specific implementations, tool asset objects can include various subtypes, such as infrastructure tools, data tools, external service interfaces, and model context protocol servers. Infrastructure tools can include container orchestration interfaces or cloud platform command tools; data tools can include databases or vector storage systems; external service interfaces can include third-party service APIs; and model context protocol servers can include tool services that conform to the model context protocol.

[0087] During the identification process, tool call behavior characteristics and system call behavior characteristics can be used for identification based on structured observation data. Tool call behavior characteristics include information such as call instructions, call target identifiers, call types, call order, and call results. Based on these characteristics, different types of tool asset objects can be identified. For example, when the call target identifier corresponds to a container orchestration interface or cloud platform command tool, it can be identified as an infrastructure tool; when the call target identifier corresponds to a database service or vector storage service, it can be identified as a data tool; when the call target identifier corresponds to a third-party service interface, it can be identified as an external service interface tool; and when the call behavior conforms to the model context protocol interaction mode, it can be identified as a model context protocol server tool.

[0088] Furthermore, tool asset objects can be verified and their types refined based on system call behavior characteristics. System call behavior characteristics include process execution, file read / write, network socket operations, and resource access behaviors. For example, when system call behaviors related to database access are detected, identified data-related tools can be confirmed; when system call behaviors related to command execution or process creation are detected, infrastructure-related tools can be confirmed; when continuous data read / write or interactive behaviors are detected and match the tool call context, the type of the corresponding tool asset object can be further verified.

[0089] Model service asset objects are used to represent service endpoints that provide inference capabilities for AI agents. In specific implementations, model service asset objects can include public cloud model services and privately deployed model services. Identification of model service asset objects can be based on network communication behavior characteristics and resource access characteristics in structured observation data. For example, public cloud model services can be identified by matching DNS domain names to a pre-defined whitelist, model service endpoints can be identified by resolving communication targets using TLS server name indications, or privately deployed model services can be identified based on internal service discovery mechanisms combined with resource identification information in system audit data.

[0090] Furthermore, during the asset identification process, key fields for identifying asset objects can be extracted from structured observation data and system audit data based on identifier extraction rules. These fields include process identifiers, container identifiers, service addresses, domain name information, and interface paths. Based on these key fields, asset objects can be uniquely identified and deduplicated, thereby completing the standardized representation of asset objects.

[0091] After identifying asset objects, the asset management module establishes relationships between different asset objects and constructs a topological structure among them. In one specific embodiment, refer to... Figure 3 As shown, an AI agent is used as the central node to model the interaction relationships between different asset objects. The AI ​​agent establishes interaction relationships with model service asset objects, tool asset objects, and other agent asset objects, thereby forming a multi-type asset interaction structure around the AI ​​agent.

[0092] During the interaction, AI agents communicate with corresponding objects through different types of interaction protocols, forming a multi-protocol interaction system. Specifically, AI agents interact with model services through inference requests, using the interaction protocol between agents and large language models; model services include public and private model services. AI agents interact with tool assets through tool calls, using the interaction protocol between agents and tools. AI agents interact with other agents through agent collaboration, used for task distribution and execution result feedback, using the interaction protocol between agents. In scenarios supporting model context protocols, AI agents can also interact with tool services through model context protocols, thereby achieving a unified tool call and context passing mechanism. Model context protocols can support HTTP-based or standard input / output-based transmission methods.

[0093] In the aforementioned multi-protocol interaction process, the interaction data generated by different protocols can be uniformly parsed and processed, transforming it into structured interaction metadata. Interaction metadata can include source entity information, target entity information, interaction type information, and time information. By parsing the interaction metadata, the initiating entity, target, and direction of the interaction can be determined, thereby establishing the calling relationship between different asset objects.

[0094] Data belonging to the same execution chain can be categorized based on association identifiers in the interaction metadata, and interaction events can be organized chronologically within that execution chain to determine the sequence and dependencies between different interaction behaviors. Based on this, the association between model service asset objects and tool asset objects can be established according to the triggering relationship between model response results and subsequent tool invocation behaviors, reflecting the process of tool execution driven by model inference results.

[0095] After establishing the aforementioned relationships, the relationships between different asset objects can be expressed in a structured way. Specifically, AI agents, model services, tool components, and other agents can be abstracted as nodes in the topology, and inference request relationships, tool call relationships, and agent collaboration relationships can be abstracted as connecting edges between nodes. At the same time, interaction relationships based on model context protocols can be included as a type of connection relationship, thereby constructing an asset topology centered on AI agents.

[0096] The above methods can transform the interaction data scattered in the multi-protocol interaction process into a unified asset object and its relationship, and further construct a structured asset topology model, thereby realizing a holistic description of the interaction relationship between multiple types of assets during the operation of the AI ​​agent.

[0097] In a preferred embodiment, the asset management module is further configured to:

[0098] The system continuously monitors the operational status of asset objects based on structured observation data and system audit data, and dynamically updates the relationships and topology between asset objects.

[0099] Specifically, the asset management module can stream structured observation data from the edge acquisition layer and, combined with subject information, resource identification information, and operation context information from system audit data, perceive the operational status of each asset object in real time. Operational status can include: the activity status of the AI ​​agent, the availability status of model services, the invocation status of tool components, and the level of interaction activity between asset objects.

[0100] During continuous monitoring, the current behavior of asset objects can be determined based on network communication behavior characteristics, tool call behavior characteristics, system call behavior characteristics, and resource access characteristics in structured observation data. This, combined with scheduling events and operation records in system audit data, allows for the updating of the asset object's lifecycle status. For example, when a new process identifier or container identifier is detected, it can be determined that a new asset object has been added; when a corresponding asset object does not generate any interactive behavior within a preset time window, it can be determined that the asset object is inactive or has exited operation.

[0101] The asset management module can dynamically maintain the relationships between asset objects based on continuously received interaction metadata. Specifically, it continuously updates the calling relationships between different asset objects based on the source entity information, target entity information, and interaction type information in the interaction metadata; when a new interaction relationship is detected, the corresponding relationship is added to the relationship set; when an existing interaction relationship is detected and does not reappear within a preset time window, the relationship can be decayed or marked as invalid.

[0102] Based on this, the topology between asset objects can be dynamically updated. Newly added asset objects are added to the topology as new nodes, and corresponding connecting edges are established based on the updated relationships. At the same time, for asset objects that are no longer valid or active, and their relationships, they can be deleted from the topology or marked as inactive nodes, thereby maintaining the consistency between the topology and the actual operating state.

[0103] In one implementation, the topology can be updated periodically based on a time window mechanism. By statistically analyzing the interaction data within different time windows, a dynamic topology reflecting the current operating status can be constructed. In another implementation, an event-driven mechanism can be used to trigger a local update of the topology when a change in the state of an asset object or a change in its relationship is detected, thereby improving update efficiency.

[0104] Through the above methods, continuous monitoring of the operational status of asset objects and dynamic maintenance of asset relationships and topology are achieved. This enables the constructed asset topology to reflect the changes in the interaction relationships between multiple types of assets during the operation of the AI ​​agent in real time, thereby providing a real-time and accurate data foundation for subsequent behavioral security analysis.

[0105] In a preferred embodiment, the central analysis layer further includes a security analysis module, which is used for:

[0106] An asset context is constructed based on the asset objects and their topological relationships determined by the asset management module. The structured observation data is then analyzed based on the asset context to identify abnormal behaviors in the operation of the AI ​​agent and to generate alarm information when abnormal behaviors are identified.

[0107] Specifically, the asset context can be generated based on the set of asset objects and their topology output by the asset management module. This context reflects the overall resource relationships and interaction dependencies of the AI ​​agent in the current operating environment. The asset context can include the type information, identification information, ownership relationships, invocation relationships, and positional relationships of asset objects within the topology. For example, it can record the model service asset objects corresponding to the AI ​​agent, the tool asset objects it depends on, and the collaborative relationships with other agents, thus forming a contextual environment description surrounding the AI ​​agent. Based on this, the security analysis module can map structured observation data to the asset context. Specifically, based on the source subject information, target subject information, and interaction type information extracted from the structured observation data, the corresponding interactive behaviors can be associated with the corresponding asset nodes and connections in the topology, thereby transforming the observation data into behavioral data with contextual semantics.

[0108] In one implementation, real-time observed interaction behaviors can be compared and analyzed based on the normal interaction patterns defined in the asset context. When an interaction deviates from the existing call relationship, exhibits an abnormal call path, or accesses an unauthorized asset object, it can be identified as abnormal behavior. For example, when an AI agent directly accesses a new tool asset object without establishing a relationship, or bypasses the model service to directly trigger a high-risk tool call, it can be identified as abnormal call behavior.

[0109] In another implementation, the topology within the asset context can be used to perform path analysis on interactive behaviors. When behavioral links that do not conform to the expected execution path are detected, anomalies can be identified. For example, when an AI agent exhibits abnormal call sequences, cyclic calls, or call paths that cross multiple unrelated asset nodes during execution, it can be identified as abnormal behavior. Furthermore, resource access characteristics and system call behavior characteristics from structured observation data can be combined to detect behaviors involving access to sensitive resources or high-risk operations.

[0110] After identifying abnormal behavior, the security analysis module can generate corresponding alarm information. Abnormal behavior can include abnormal call relationships, abnormal call paths, unauthorized access behavior, or high-risk system operation behavior. When an alarm is triggered, key information related to the abnormal behavior can first be extracted based on structured observation data and asset context. Key information can include the type of abnormal behavior, the identifier of the involved asset object, the time of the anomaly, the interaction type, and the corresponding call path information. The call path information can be extracted based on the asset topology to describe the propagation path of the abnormal behavior between asset nodes, such as the complete call chain from the AI ​​agent to the model service and then to the tool component. Alarm information can be structured to form alarm event data in a unified format. Alarm event data can further include contextual information of the abnormal behavior, such as the characteristics of the request parameters that triggered the abnormal behavior, the characteristics of the execution result, or the characteristics of the system call behavior, thereby improving the interpretability of the alarm information.

[0111] Furthermore, alarm information can be categorized based on asset attribute information within the asset context. Specifically, alarm information can be risk-assessed based on the importance or sensitivity level of the involved asset object within the topology. For example, when abnormal behavior involves core AI agents, critical model services, or highly sensitive tool components, the corresponding alarm can be marked as high-risk; when abnormal behavior only involves low-sensitivity asset objects or edge call behavior, the alarm can be marked as medium-risk or low-risk.

[0112] In one implementation, the alarm level can be comprehensively determined by combining the severity of the abnormal behavior and its propagation range within the topology. For example, when abnormal behavior crosses multiple critical nodes or forms a long propagation path within the topology, the corresponding alarm level can be increased; when abnormal behavior is limited to only local nodes, its alarm level can be decreased. After alarm information generation and classification are completed, the alarm information can be output to a preset alarm processing system. Output methods may include writing to the log system, pushing to the monitoring platform, or sending to the operation and maintenance alarm channel, thereby supporting subsequent risk handling and response operations.

[0113] In a preferred embodiment, the central analysis layer further includes a query analysis module, which is used for:

[0114] The system audit data is persistently stored in the audit metadata table, and the structured observation data is persistently stored in the behavior log table;

[0115] Based on preset tracking identifiers, an association mapping relationship is established between the audit metadata table and the behavior record table, and the scheduling intent of the AI ​​agent is bound to multi-level execution behavior through tracking identifiers;

[0116] In response to a query request carrying a query identifier, locate the target scheduling record in the audit metadata table and extract the corresponding tracking identifier from the target scheduling record;

[0117] Recursively search the behavior record table based on the tracking identifier to obtain the corresponding end-to-end execution behavior data;

[0118] Based on the full-link execution behavior data, generate and output full-link structured information reflecting the multi-level execution behavior of the AI ​​agent.

[0119] In this embodiment, the execution process of the AI ​​agent is first modeled as a link to construct a full-link tracing structure to support query analysis. Specifically, observation data from the semantic interaction layer, agent behavior layer, platform scheduling layer, process execution layer, and process chain layer can be obtained based on a multi-layer observation model, and corresponding event information can be extracted from the observation data to construct link nodes. Link nodes are used to represent the behavior of the AI ​​agent at different stages of the execution process, including semantic interaction behavior nodes, agent behavior nodes, scheduling behavior nodes, and process execution nodes, and are represented by a unified data structure, which includes node identifiers, time information, and event attribute information.

[0120] Reference Figure 4 As shown, the link nodes of the different observation layers can be organized according to the execution process to form a full-link tracing structure. In the semantic interaction layer, semantic interaction layer link nodes can be constructed based on model request and response data, recording information such as tracing identifiers, model names, and token consumption. In the agent behavior layer, corresponding link nodes can be constructed based on the agent's decision-making behavior and tool invocation behavior. In the platform scheduling layer, scheduling nodes can be constructed based on audit log data to record user identity, scheduling identifiers, and target resource information. In the process execution layer, execution nodes can be constructed based on command execution behavior, recording executed commands, execution status, and resource consumption information. In the process chain layer, child nodes can be constructed based on the parent-child relationship between processes, forming a recursively expanded execution link structure.

[0121] During the link node construction process, preset tracking identifiers can be set for the link nodes. These tracking identifiers are used to identify the same execution link and are passed and inherited between different observation layers. Specifically, corresponding tracking identifiers can be generated during the semantic interaction phase and passed on during subsequent agent behavior, platform scheduling, and process execution, enabling all link nodes in the same execution link to share a unified identifier, thereby achieving cross-level link connectivity. Through this method, the association between tracking identifiers and the full-link tracking structure can be realized, allowing link nodes in different observation layers to form a complete link under a unified identifier.

[0122] After completing the construction of the full-link tracing structure and the association of tracing identifiers, the query and analysis module can perform persistent storage processing on the link node data. Specifically, the link node data corresponding to the platform scheduling layer can be persistently stored in the audit metadata table, and the link node data corresponding to the process execution layer and the process chain layer can be persistently stored in the behavior record table. Among them, the audit metadata table is used to record the context information of scheduling behavior, and the behavior record table is used to record the execution behavior and its process expansion relationship.

[0123] Furthermore, a mapping relationship can be established between the audit metadata table and the behavior record table based on the tracking identifier. The tracking identifier is a unified identifier that runs through different observation layers and is used to identify the same execution link. During persistent storage, the tracking identifier can be written as a correlation field into the audit metadata table and the behavior record table, thereby achieving a unified association between the two types of data. At the same time, the tracking identifier binds the scheduling intent represented by the AI ​​agent at the platform scheduling layer with the multi-level execution behaviors represented by the process execution layer and the process chain layer, so that a consistent link relationship is formed between the scheduling behavior and the actual execution path.

[0124] Upon receiving a query request carrying a query identifier, the query analysis module can first search the audit metadata table to locate the target scheduling record. The query identifier can be a scheduling identifier or a session identifier, used to identify the target execution behavior to be queried. When the query identifier corresponds to a single scheduling operation, a corresponding target scheduling record can be located, and the corresponding tracking identifier and scheduling context information can be extracted from that target scheduling record; when the query identifier corresponds to a session dimension, a set of multiple target scheduling records can be obtained.

[0125] After obtaining the target scheduling record, a recursive search can be performed in the behavior record table based on the tracing identifier extracted from the target scheduling record to obtain multi-level execution behavior data corresponding to the tracing identifier. According to the mapping relationship between process identifiers and parent process identifiers recorded in the behavior record table, the execution behavior can be traced back level by level, recursively obtaining the execution records of child processes derived from the initial execution behavior, thereby reconstructing the complete process execution chain structure. During the search process, command execution information, execution result information, resource consumption information, and output content corresponding to each level of execution behavior can be obtained simultaneously, thus forming a complete set of execution chain node data.

[0126] After acquiring end-to-end execution behavior data, the target scheduling records and corresponding execution link node data can be integrated and processed to generate end-to-end structured information reflecting the multi-level execution behavior of the AI ​​agent. This end-to-end structured information can include scheduling context information and corresponding execution link structure information, where the execution link structure information describes the hierarchical relationships and dependencies between execution behaviors at each level in a tree structure. Through this method, data originally scattered across different data tables and observation layers can be uniformly reconstructed, thereby outputting structured link information oriented towards single-execution or session-based execution, supporting application scenarios such as fault location, behavior auditing, and execution process analysis.

[0127] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0128] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or modules recited in a system claim may also be implemented by a single unit or device via software or hardware. Terms such as "first," "second," etc., are used only for distinguishing descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.

[0129] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A distributed end-to-end observation system for AI intelligent agents, characterized in that, include: The edge acquisition layer is configured to be distributed and deployed on multiple computing nodes. The edge acquisition layer collects behavioral data of AI agent-related processes running on the computing nodes by loading a kernel-mode probe based on the extended Berkeley packet filter into the operating system kernel of the corresponding computing node. In user mode, the collected data is preprocessed, including protocol semantic parsing and context information association, and structured observation data carrying context metadata is output. The central analysis layer, which is communicatively connected to the edge acquisition layer, is used to aggregate structured observation data from multiple computing nodes and perform cross-source correlation analysis on the structured observation data and externally acquired system audit data to establish the correspondence between the subject information identified in the system audit data and the operational behavior in the structured observation data, thereby constructing a full-link observation model that reflects the AI ​​agent from semantic interaction to kernel execution.

2. The distributed end-to-end observation system for AI agents according to claim 1, characterized in that, The edge acquisition layer includes a kernel-mode acquisition module and a user-mode processing module, wherein: The kernel-mode acquisition module is used to acquire system call events, network communication data, and data generated during the execution of user-mode encryption library functions of the AI ​​agent-related processes by attaching the kernel-mode probe. The user-mode processing module is used to perform the context information association processing, including: obtaining the metadata of the corresponding container orchestration platform based on the process identifier in the system call event, and establishing the association between the process identifier and the metadata.

3. The distributed end-to-end observation system for AI agents according to claim 2, characterized in that, The user-space processing module is also used to perform the protocol semantic parsing processing, including: The network communication data and the data generated during the execution of the encryption library functions are reassembled to generate corresponding application layer protocol objects, and semantic payload information is extracted from the application layer protocol objects; and Based on the semantic payload information, the model resource consumption of the AI ​​agent is measured to obtain the corresponding resource consumption index, and the correlation, the semantic payload information and the resource consumption index are encapsulated as attribute fields into the structured observation data.

4. The distributed end-to-end observation system for AI agents according to claim 1, characterized in that, The central analysis layer includes a correlation processing module, which is used for: The system audit data and the structured observation data are matched in terms of time dimension based on a preset time window, and in terms of spatial dimension based on the resource identifier in the system audit data and the metadata carried in the structured observation data. In terms of behavioral consistency based on the request behavior information in the system audit data and the execution behavior information in the structured observation data, when the matching results of each dimension meet the preset conditions, the association between the subject information in the system audit data and the running behavior in the structured observation data is established.

5. The distributed end-to-end observation system for AI intelligent agents according to claim 1, characterized in that, The central analysis layer also includes a behavioral baseline analysis module, which is used for: Multidimensional feature modeling is performed on the historical normal behavior data of the AI ​​agent, and a corresponding behavioral baseline model is constructed based on the correlation between different feature dimensions. The multidimensional features include at least semantic interaction features, tool call features, resource usage features, and system execution features. The behavioral baseline model is used to analyze the degree of deviation between the real-time observation data generated by the AI ​​agent during operation and the behavioral baseline model.

6. The distributed end-to-end observation system for AI agents according to claim 5, characterized in that, The behavioral baseline model includes a semantic interaction baseline built based on historical semantic interaction data, a behavioral sequence baseline built based on historical tool call data, a resource consumption baseline built based on historical resource usage data, and a system execution behavior baseline built based on historical system call data.

7. The distributed end-to-end observation system for AI agents according to claim 1, characterized in that, The central analysis layer also includes an asset management module, which is used for: Based on a preset asset model, the structured observation data and the system audit data are parsed to identify multiple types of asset objects related to the AI ​​agent, establish the association between different asset objects, and construct the topological structure between the asset objects. The asset objects include at least agent asset objects, tool asset objects, and model service asset objects.

8. The distributed end-to-end observation system for AI intelligent agents according to claim 7, characterized in that, The asset management module is also used for: The operational status of the asset objects is continuously monitored based on the structured observation data and the system audit data, and the relationships between the asset objects and the topology are dynamically updated.

9. The distributed end-to-end observation system for AI agents according to claim 7, characterized in that, The central analysis layer also includes a security analysis module, which is used for: An asset context is constructed based on the asset objects and their topological relationships determined by the asset management module. The structured observation data is analyzed based on the asset context to identify abnormal behaviors in the operation of the AI ​​agent, and alarm information is generated when the abnormal behavior is identified.

10. The distributed end-to-end observation system for AI agents according to claim 1, characterized in that, The central analysis layer also includes a query analysis module, which is used for: The system audit data is persistently stored in the audit metadata table, and the structured observation data is persistently stored in the behavior record table; Based on the preset tracking identifier, an association mapping relationship is established between the audit metadata table and the behavior record table, and the scheduling intention of the AI ​​agent is bound to the multi-level execution behavior through the tracking identifier; In response to a query request carrying a query identifier, the target scheduling record is located in the audit metadata table, and the corresponding tracking identifier is extracted from the target scheduling record; Based on the tracking identifier, a recursive search is performed in the behavior record table to obtain the corresponding end-to-end execution behavior data; Based on the full-link execution behavior data, full-link structured information reflecting the multi-level execution behavior of the AI ​​agent is generated and output.