Monitoring system for AI Agent risk behavior

By collecting and analyzing risk behavior signals of AI agents in real time and in parallel, and combining multi-agent systems and adaptive protection mechanisms, the problem of insufficient recognition of dynamic behavior of AI agents in existing technologies is solved, and efficient and comprehensive risk monitoring and protection are achieved.

CN121902138APending Publication Date: 2026-04-21浙江实在智能科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
浙江实在智能科技有限公司
Filing Date
2026-03-23
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies cannot effectively identify the dynamic behavior and operational intentions of AI agents, making it difficult to promptly block abnormal operations. They also have limited coverage and insufficient full-process risk protection capabilities against unknown threats and multi-model collaborative scenarios.

Method used

A monitoring system is adopted, including a data perception and acquisition module, a main controller, a multi-agent system, and a security response system. By collecting risk behavior signals in real time and combining them with the semantic information of user-initiated processes, the system performs self-monitoring, self-training, and policy self-updating, thereby achieving parallel analysis and dynamic protection of the AI ​​Agent.

Benefits of technology

Significantly reduces false alarm and false negative rates, shortens decision-making delays, improves blocking success rates, enables comprehensive real-time monitoring and adaptive protection of complex agent behavior, and continuously addresses emerging risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121902138A_ABST
    Figure CN121902138A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and particularly relates to a monitoring system for AI Agent risk behaviors. The system comprises a data sensing and collecting module used for collecting risk behavior signals and external data sources generated in the operation process of the AI Agent in real time; the main controller is used for receiving and distributing the risk behavior signal and an external data source and carrying out aggregation and risk assessment on a feedback analysis result; the multi-agent system comprises a plurality of sub-agents with independent functions; the sub-agents are configured to receive data from the main controller and perform parallel analysis and risk assessment on behaviors of the AI Agent from a plurality of dimensions; the safety response system is used for implementing dynamic protection and emergency response to the execution environment of the AI Agent according to the risk assessment result and the response instruction issued by the main controller; and the knowledge base is used for storing security rules, threat models and agent behavior parameters, and performing dynamic updating based on system operation feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a monitoring system for monitoring risky behavior of AI agents. Background Technology

[0002] In recent years, AI agents (agent-based artificial intelligence, hereinafter referred to as Agents) have been widely used in personal terminals and enterprise systems. Early Agents were mostly used for simple automation tasks, such as batch file organization, schedule reminders, or the execution of fixed scripts. Their behavior patterns were relatively fixed, and the security risks were low. With the maturity of natural language processing, reinforcement learning, large models, and multimodal interaction technologies, the capabilities of Agents have gradually expanded to the autonomous planning and execution of complex task chains. For example, in office scenarios, Agents can complete cross-document data aggregation and analysis based on natural language instructions; in development and operations scenarios, Agents can automatically call code repositories, execute build tests, and make configuration modifications; in network collaboration scenarios, Agents can complete cross-platform collaborative processing of emails, sending and receiving messages, and manipulating files.

[0003] However, these agents complete tasks by simulating user actions or making autonomous decisions, and the execution process is complex and opaque, making it more difficult to monitor than traditional user actions. For example, malicious agents can download and execute malicious programs in the background under the guise of system updates; agents in corporate email may actively tamper with email forwarding rules, redirecting internal company communications to external email addresses; and during system maintenance, agents may even modify scheduled tasks or startup items to establish persistent backdoor channels.

[0004] Therefore, real-time monitoring technology for agent operation behavior has become a key research direction for the security protection of personal terminals and enterprise-level systems. The market urgently needs solutions that can penetrate the agent's execution logic and identify abnormal operation intentions.

[0005] Currently, research and practice on agent behavior control mainly fall into three categories regarding monitoring and protection methods: 1. Static policies based on firewalls, access control, and other permission controls; 2. Post-event retrospective analysis based on logs and behavioral auditing; 3. Threat detection and discovery technologies based on feature detection or model recognition.

[0006] However, existing technologies still have significant shortcomings in agent behavior monitoring. These shortcomings primarily manifest in their inability to effectively identify agent dynamic behavior and operational intentions, their difficulty in timely blocking of abnormal operations, and their limited coverage, resulting in insufficient end-to-end risk protection against unknown threats and multi-model collaborative scenarios. Specifically: 1. Static access control policies cannot handle dynamic agent behavior. First, these methods cannot analyze agent task planning and multi-step operations in real time; they can only monitor individual operations, resulting in incomplete coverage of the entire task chain and making them vulnerable to malicious "step-by-step" operations. Second, static access control cannot associate agent historical operations with current tasks; for example, it cannot identify abnormal behavior such as deleting backup files using legitimate scripts. Furthermore, rule-based monitoring systems are prone to false alarms from normal automated operations, leading to inaccurate risk management.

[0007] 2. Delayed log and behavior audit interventions make real-time protection difficult. Log-based and behavior auditing technologies suffer from response lag, failing to intervene or block high-risk actions by agents in real time, thus hindering immediate protection. Furthermore, traditional log auditing lacks real-time monitoring and encryption protection for sensitive information acquired by agents, such as clipboard content or documents uploaded to the cloud, creating blind spots in content monitoring and making it difficult to prevent data leaks or unauthorized dissemination. These technologies also have limitations when handling multi-step or cross-system tasks. Due to the lack of correlation analysis between task context and historical operations, they cannot effectively identify the abuse of legitimate operations, such as malicious modification or deletion operations performed through normal scripts. Even when combined with anomaly pattern analysis or machine learning methods, the real-time protection capabilities of traditional log auditing remain limited, failing to meet the dynamic and complex needs of agent behavior monitoring.

[0008] 3. Feature detection and model recognition rely on limited coverage of known threat features. Threat detection methods based on signatures, signature codes, or model recognition rely on static rules or known attack characteristics to identify threats. While these methods react quickly to known attacks, they lack effective protection against unknown or novel threats, making proactive early warning and predictive protection difficult. Furthermore, signature databases require frequent updates to address new attacks, resulting in high maintenance costs. Although security measures for AGI models have been proposed, such as data encryption, rule verification, adversarial sample detection, and prevention of model reverse engineering, these methods primarily focus on the model itself and cannot comprehensively monitor the risks of multi-model collaboration and environmental operations during agent execution, thus limiting the overall security resilience of the system.

[0009] Therefore, it is very important to design a monitoring system for AI Agent risky behavior that can combine the specific semantic information of the user-initiated process with the real-time execution node of the Agent to achieve self-monitoring, self-training and self-updating of Agent operations. Summary of the Invention

[0010] This invention aims to overcome the shortcomings of existing technologies for monitoring agent behavior, which suffer from the inability to effectively identify the dynamic behavior and operational intentions of agents, the difficulty in timely blocking of abnormal operations, limited coverage, and insufficient full-process risk protection capabilities for unknown threats and multi-model collaborative scenarios. It provides a monitoring system for AI Agent risky behavior that combines the specific semantic information of the user-initiated process with the real-time execution nodes of the agent to achieve self-monitoring, self-training, and self-updating of agent operations.

[0011] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: Monitoring systems for detecting risky behavior by AI agents include: The data perception and acquisition module is used to collect risk behavior signals and external data sources generated during the operation of the agent-type artificial intelligence (AI) agent in real time. The main controller is used to receive and distribute the risk behavior signals and external data sources, and to aggregate and assess the risk of the feedback analysis results. A multi-agent system includes several functionally independent sub-agents; the sub-agents are configured to receive data from the main controller and perform parallel analysis and risk assessment of the AI ​​Agent's behavior from six dimensions: intent risk, object security, content security, privacy risk, program security, and log security. The security response system is used to implement dynamic protection and emergency response for the execution environment of the AI ​​Agent based on the risk assessment results and response instructions issued by the main controller. The knowledge base stores security rules, threat models, and agent behavior parameters, and is dynamically updated based on system operation feedback.

[0012] Preferably, the main controller includes: An input distributor is used to distribute received data to the corresponding sub-agents in the multi-agent system according to the corresponding type; The result aggregator is used to receive and aggregate risk assessment reports from each sub-agent in a multi-agent system to form structured information. The risk assessor is used to perform comprehensive analysis based on the structured information and the security rules in the knowledge base to generate an overall risk assessment report containing a unified risk level. The strategy executor is used to generate and issue standardized security response instructions to the security response system based on the overall risk assessment report and with reference to the response strategies defined in the knowledge base.

[0013] Preferably, the risk assessor employs a quantitative risk assessment model, specifically calculating the risk value using the following formula: Risk value = Probability of occurrence score × Vulnerability score × Impact score; The probability score, vulnerability score, and impact score are all within the range of 1 to 3 points; the risk assessor classifies the risk level into low risk, medium risk, or high risk based on the calculated risk value range.

[0014] Preferably, according to the risk value calculation formula, a final risk value of 1-6 points is defined as low risk; a final risk value of 7-14 points is defined as medium risk; and a final risk value of 15-27 points is defined as high risk. High-risk behaviors require immediate action to intercept; medium-risk behaviors require management attention and mitigation measures; low-risk behaviors can be tolerated and logged, without additional control.

[0015] Preferably, the multi-agent system includes: The Intent Risk Intelligent Agent module is responsible for monitoring and analyzing abnormal operational intents of users or AI Agents; The object security intelligent agent module is responsible for assessing the access permissions of the target objects operated by the AI ​​Agent and the compliance of the operational behavior; The content security intelligent agent module is used for security detection of data content processed or generated by the AI ​​Agent; The privacy risk intelligent agent module is used to assess the risk of leakage of personal privacy information during data processing; The program security intelligent agent module is used to analyze code vulnerabilities and malicious behaviors during the execution of the AI ​​Agent program; The Log Security Intelligent Agent module is used to analyze system security logs and discover abnormal events, attack traces, errors, or warning messages.

[0016] Preferably, the sub-agents in the multi-agent system interact with each other and engage in collaborative reasoning, sharing their respective analysis results.

[0017] Preferably, the security response system includes: The instruction parsing module is used to parse the security response instructions issued by the main controller and determine the execution priority; The strategy orchestration module, connected to the knowledge base, is used to generate specific handling strategies based on the parsed instruction parameters and the real-time status of the system. An execution engine cluster is used to execute the aforementioned processing strategy; The feedback loop module is used to verify whether the handling strategy is successful after the execution engine cluster completes the execution action, and at the same time feeds back the handling execution result and environmental status changes to the main controller and knowledge base.

[0018] Preferably, the execution engine cluster includes: The blocking and isolation engine is used to perform processes to terminate, disconnect, or isolate resources for high-risk behaviors. The downgrade and restriction engine is used to perform permission downgrades, access rate restrictions, or enable human verification mechanisms for medium-risk behaviors. The auditing and tracing engine is used to capture full data packets, record screens, and retain contextual information for low-risk behaviors, and then label them as suspicious.

[0019] Preferably, the knowledge base is a dynamic adaptive knowledge base, used to receive risk assessment results from the main controller and response effect feedback from the security response system, and dynamically update the security rules, threat models and agent behavior parameters stored in the knowledge base based on the risk assessment results and response effect feedback.

[0020] Compared with existing technologies, the beneficial effects of this invention are: (1) Comprehensive risk identification dimensions and significantly reduced false alarms and missed alarms: Unlike static strategies based on firewalls, access control and other permission controls, this invention expands the coverage of only one or a few dimensions to six dimensions of parallel scanning: intent, object, program, content, privacy and logs through a hierarchical monitoring architecture of "1 main controller + 6 sub-intelligent agents"; the actual test results show that under the same business system, the false alarm rate and missed alarm rate of complex Agent behavior are greatly reduced, and the risk visibility of "three-dimensional protection network" level is achieved; (2) Significantly shortened decision delay and more timely attack blocking: In response to the problem of slow post-event traceability based on log and behavior audit, the six intelligent agent sub-modules adopt parallel analysis and independent report generation, and the main controller The device only needs to perform lightweight aggregation to determine high / medium / low risk; compared with the traditional serial detection link, the parallel analysis mechanism in this invention significantly shortens the decision delay and can complete the interception before the attack payload reaches the business logic, thus improving the blocking success rate; (3) Adaptive closed loop enables the protection capability to continuously improve with the evolution of business: In view of the shortcomings of feature detection relying on static feature library and not being able to automatically improve the protection capability, this invention constructs a closed loop mechanism of "main controller → multi-agent collaboration → historical data feedback → model retraining → policy self-updating" to realize the dynamic optimization of content violation detection and risk judgment; this mechanism can improve the detection accuracy without manual re-labeling or manual parameter tuning, significantly reduce the operation and maintenance cost, and enable the system protection capability to automatically improve with the evolution of business and continuously cope with new risks. Attached Figure Description

[0021] Figure 1This is a schematic diagram of the principle architecture of a monitoring system for monitoring risky behavior of AI agents based on the present invention; Figure 2 This is a schematic diagram of the principle architecture of the safety response system in this invention. Detailed Implementation

[0022] To more clearly illustrate the embodiments of the present invention, specific implementation methods will be described below with reference to the accompanying drawings. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings and other implementation methods can be obtained based on these drawings without any creative effort.

[0023] like Figure 1 As shown, this invention proposes a monitoring system for monitoring risky behavior of AI agents, specifically including: The data perception and acquisition module is used to collect risk behavior signals and external data sources generated during the operation of the agent-type artificial intelligence (AI) agent in real time. The main controller is used to receive and distribute the risk behavior signals and external data sources, and to aggregate and assess the risk of the feedback analysis results. A multi-agent system includes several functionally independent sub-agents; the sub-agents are configured to receive data from the main controller and perform parallel analysis and risk assessment of the AI ​​Agent's behavior from six dimensions: intent risk, object security, content security, privacy risk, program security, and log security. The security response system is used to implement dynamic protection and emergency response for the execution environment of the AI ​​Agent based on the risk assessment results and response instructions issued by the main controller. The knowledge base stores security rules, threat models, and agent behavior parameters, and is dynamically updated based on system operation feedback.

[0024] The functions and roles of each module are as follows: 1. Data Sensing and Acquisition Module: When the AI ​​Agent generates risky behavior signals during its operation, the data perception and acquisition module designed in this invention will collect and record relevant external data sources in real time as input for risk analysis.

[0025] 2. Main controller system: The system includes: an input distributor, a result aggregator, a risk assessor, and a strategy executor.

[0026] As the core management unit of the system, the main controller can actively or passively receive external data sources and transmit the data to the input distributor.

[0027] The input distributor distributes inputs to the appropriate agents in the multi-agent system for processing based on the data type and analysis requirements. The processed results are then returned to the result aggregator.

[0028] Results aggregators are primarily used to aggregate, analyze, and summarize multiple data results to form structured information.

[0029] The risk assessor performs a comprehensive analysis of the information gathered by the results aggregator to evaluate the overall security risk level, threat types, and potential impact of the current system. This module can weight, correlate, and normalize the fine-grained risk indicators provided by the agent based on security rules in the knowledge base, generating a unified risk assessment report.

[0030] If the behavior could lead to system paralysis, core data leakage, or significant legal or financial losses, it is considered high-risk and requires immediate action to intercept it; if it causes partial service interruption or limited data loss, it is considered medium-risk and requires management attention, and is acceptable but should have mitigation measures; if it only causes minor impact and can be quickly recovered, it is considered low-risk, can be tolerated and logged, and usually does not require additional control.

[0031] Based on the assessment results from the risk assessment module and referencing the response strategies defined in the knowledge base, the policy executor generates specific security response instructions. These instructions are then sent to the security response module.

[0032] 3. Multi-agent systems: The system comprises six intelligent agent sub-modules: intent risk intelligent agent, object security intelligent agent, content security intelligent agent, privacy risk intelligent agent, program security intelligent agent, and log security intelligent agent.

[0033] The intent-based risk agent is responsible for monitoring unusual intents of users or the system. For example, it monitors risky behaviors such as repeatedly attempting to transfer money to high-risk countries. This agent ensures that the AI ​​agent's understanding and execution of user instructions comply with security policies.

[0034] The object security agent is responsible for assessing the security of various objects and detecting risks such as data leakage, tampering, and unauthorized access. For example, it assesses unauthorized access to a medical database by a doctor's account outside of working hours. This agent ensures that the AIAgent has legitimate permissions when operating on various objects (such as files, database records, API interfaces, system resources, etc.) and that its actions comply with preset security procedures.

[0035] The content security agent is responsible for analyzing data content and detecting risks such as sensitive information leakage, malicious code injection, and the spread of illegal content. For example, if it analyzes a seemingly normal landscape photo that contains embedded child abuse content, the agent ensures that the AI ​​agent does not include sensitive information, illegal content, or malicious payloads when processing, generating, or transmitting any content.

[0036] The privacy risk intelligence agent is responsible for assessing potential privacy risks during data processing and ensuring the compliance of data use. For example, it monitors the behavior of an e-commerce platform's employees who package and sell users' shipping addresses and phone numbers to third-party marketing companies. This intelligence agent assesses and handles any actions that could lead to the unauthorized collection, storage, transfer, inference, or leakage of personal privacy information, ensuring that operations comply with regulations and the organization's privacy policy.

[0037] The program security agent is responsible for analyzing the security of application or service code and runtime, detecting vulnerabilities and malicious program behavior. For example, it monitors the behavior of constructing special inputs to maliciously exploit algorithms to make customer service robots return other users' credit card numbers. This agent monitors the integrity and security of the AI ​​Agent's own running program to prevent the program from being tampered with, injected with malicious code, or having logical vulnerabilities.

[0038] The log security agent is responsible for analyzing various system security logs to detect abnormal events, attack traces, errors, or warning messages. For example, it monitors hackers deleting their login records from server logs after infiltrating the enterprise intranet. This agent performs real-time or periodic analysis on all logs generated during operation to identify abnormal patterns, security events, or potential risks.

[0039] The agents can exchange information and engage in collaborative reasoning, sharing their respective analysis results to improve the accuracy and comprehensiveness of the overall risk assessment. For example, when the intent-risk agent detects abnormal behavior, it can request the log security agent to provide relevant logs for further confirmation.

[0040] Subsequently, the multi-agent system sends the risk assessment report generated by the six sub-agents to the result aggregator in the main controller, which receives and performs preliminary integration.

[0041] 4. Security Response System: The security response system is the defense and handling component directly facing the AI ​​Agent's operating environment in this architecture. Its core function is to receive instructions from the main controller and implement dynamic protection and emergency response based on risk level, threat type, and real-time scenario. To achieve precise, tiered, and closed-loop risk control, the system's logical architecture specifically includes four core modules: The module includes an instruction parsing module, a strategy orchestration module, an execution engine cluster, and a feedback loop module. For example... Figure 2 As shown, the detailed process descriptions for each module are as follows: Instruction parsing module: This module is responsible for receiving and parsing the standardized response commands sent by the main controller, obtaining command parameters, which include at least the risk level, threat type, target process ID, and network session ID; obtaining real-time load status data of the system, including CPU utilization and network congestion level; and determining the execution priority of the standardized response commands based on the risk level and the real-time load status data. Specifically, when the risk level is high, the corresponding blocking command is set to the highest execution priority.

[0042] Strategy orchestration module: This module connects to the knowledge base in real time, serving as the core hub linking the decision-making and execution layers. Its main function is to transform the abstract risk instructions output by the instruction parsing module into a sequence of specific technical actions that can be recognized by the underlying system. Specifically, based on the parsed instruction parameters (risk level, threat type), this module retrieves a matching baseline response template from the knowledge base; combined with the real-time environmental context of the current system (such as load status, business priority), it parameterizes the baseline template to generate standardized handling scripts or configuration parameters (e.g., generating specific iptables rules or API rate limiting thresholds), and then distributes this instantiated strategy to the execution engine cluster.

[0043] Execution engine cluster: This is the execution unit that directly operates within the AI ​​Agent's runtime environment. It contains three dedicated sub-engines for different risk levels, enabling a tiered response from "tolerance of records" to "forced blocking." (1) Blocking and Isolation Engine (for high-risk situations): When the system detects high-risk behaviors such as malicious code injection or core data theft, the engine leverages operating system kernel-level privileges to immediately execute interception measures. Specific actions include forcibly terminating malicious processes, severing abnormal network sessions, and moving infected files or processes into an isolation sandbox, ensuring that threats are eliminated before causing substantial damage.

[0044] (2) Downgrading and limiting the engine (for medium risk): When the system detects medium-risk behaviors such as high-frequency abnormal access or unauthorized non-sensitive operations, the engine takes mitigation measures. Specific actions include dynamically adjusting the QPS (queries per second) threshold for API calls, revoking some sensitive permissions (such as downgrading read / write permissions to read-only), or enabling a "human-in-the-loop" mechanism to reduce risk without completely interrupting service.

[0045] (3) Audit and Tracking Engine (for low-risk applications): When the system detects low-risk behavior with unclear intent or minor violations, the engine adopts a tolerance and logging strategy. It activates "Shadow Mode," capturing all subsequent data packets, recording the screen, and retaining contextual information for all agent operations, then labeling them as suspicious. This provides comprehensive data support for subsequent risk escalation, location, and tracing.

[0046] Feedback closed-loop module: This module is responsible for the "closing down" and "evolution" of the response process. After the execution engine completes its actions, this module verifies whether the handling was successful (e.g., verifying whether the malicious process has indeed stopped) and sends the handling result (success / failure / partial success) and the new environment status data after handling back to the main controller and knowledge base. This feedback mechanism enables the system to automatically adjust subsequent risk assessment weights and optimize security strategies based on actual combat results, thereby forming a dynamic, traceable, closed-loop security management system.

[0047] 5. Knowledge Base: The security rules of the knowledge base are not static. Risk assessment results and response effectiveness generated throughout the process are fed back into the knowledge base to update security policies, threat models, and agent behavior parameters, enabling continuous learning and optimization of the system, thus forming a closed-loop intelligent security management system. We propose a quantifiable risk level scoring system, which is added to the knowledge base for the risk assessment module to use: Risk value = Probability of occurrence score × Vulnerability score × Impact score The three factors—probability of occurrence, vulnerability, and impact—are rated on a scale of 1 to 3. Likelihood of occurrence: 1 (low risk): Almost impossible, with few or no threat sources. For example, small intranet environments rarely suffer from external attacks.

[0048] Probability of occurrence: 2 (Medium risk): There is a certain probability of occurrence; known attack methods or moderate threat sources exist. Examples include network scanning and phishing email attacks commonly seen in medium-sized enterprises.

[0049] Likelihood of Occurrence: 3 (High Risk): Highly likely to occur; the threat source is active or has already targeted the system. For example, a known vulnerability has been publicly disclosed and has already been exploited by a large-scale group attack.

[0050] Vulnerability score 1 (low risk): Almost no weaknesses or they have been patched, making it difficult to exploit. For example, the system has all patches installed, strong passwords, and two-factor authentication.

[0051] Vulnerability score 2 (medium risk): There is a certain weakness, but certain technical conditions are required to exploit it. For example, the service has a medium-risk vulnerability, but a complex exploit chain is required to succeed.

[0052] Vulnerability 3 (High Risk): A major vulnerability or obvious flaw that is easily exploitable. For example, an unencrypted database or a publicly exposed management port.

[0053] Impact 1 (Low Risk): Limited impact, business operations are largely unaffected, and recovery is rapid. For example, a short-term interruption of non-core systems may result in log errors but not affect overall operation.

[0054] Impact 2 points (Medium Risk): The impact is significant, affecting some business operations and posing certain economic or compliance risks. For example, partial customer data leakage or order delays on e-commerce platforms.

[0055] Impact rating: 3 (High Risk): Severe impact, leading to business disruption and significant economic losses. Examples include the theft of a core database, resulting in a complete system crash and production stoppage, or fines for major privacy violations.

[0056] The final risk score is calculated using a formula. A score of 1-6 is defined as low risk, 7-14 as medium risk, and 15-27 as high risk. High-risk behaviors require immediate intervention; medium-risk behaviors require management attention and are acceptable but should have mitigation measures; low-risk behaviors are tolerable and can be logged, generally requiring no additional control.

[0057] Based on the technical solution of this invention, the implementation process of this invention in practical applications will be illustrated through the following case scenarios. The specific application implementation scheme is as follows: This invention can cover the two major categories and thirteen risk scenarios listed, enabling comprehensive real-time monitoring and intervention of agent risk behavior, and effectively improving the security and controllability of intelligent agents in practical applications.

[0058] (1) Environmental risk monitoring mechanism: 1. Blocking malicious pop-ups and ad clicks; 2. Real-time identification and blocking of agent access to phishing websites; 3. Identify or protect against phishing emails targeting agents; 4. Effectively address the behavior of agents automatically bypassing reCAPTCHA; 5. Detect actions that indicate fraudulent intent regarding account access; 6. Prevent the risk of agents being tricked into generating misleading text.

[0059] (2) User-initiated risk monitoring mechanism: 1. Monitor the Agent for illegal downloads or unauthorized access during web page operations; 2. Block the agent from posting sensitive or inappropriate content on social media platforms; 3. Prevent agents from leaking important documents or sharing them without authorization in office software; 4. Review file read / write operations for potentially sensitive information; 5. Prevent the agent from executing dangerous system-level commands; 6. Identify the programming behavior executed by the Agent to prevent private content from being pushed to public repositories; 7. Detect the risk of sensitive content transmission by the Agent in email operations.

[0060] Taking pop-up ads as an example, the most direct and core detection point lies in content security, which focuses on whether the content of the ad itself complies with regulations. However, a complete security detection system will utilize other intelligent agents to perform cross-verification from different perspectives. The working mechanism of this invention is as follows: (1) Initial triggering and initial content screening: When the system detects a new pop-up / ad request, the content security intelligence is first triggered to conduct a preliminary review of the ad's text, images, and video content. If the content is clearly in violation of regulations, it can be blocked immediately, and the result will be reported back to the decision center.

[0061] (2) Multi-dimensional cross-validation: If the content passes the initial screening or is only considered low-risk, the system will simultaneously distribute advertising-related information (such as URL, source, and display context) to other agents for parallel or sequential detection. The intent risk agent will analyze the timing of its appearance, interactive behavior, and redirect links to determine if there is malicious inducement. The privacy risk agent will monitor its data collection behavior and permission requests to assess the risk of privacy breaches. The program security agent will run the advertising code in a sandbox to detect any vulnerability exploits or malicious scripts. The object security agent will verify the legality of the advertising's delivery source and the compliance of permissions granted to the target object.

[0062] (3) Risk assessment and decision-making: The detection results and risk scores of all agents are aggregated in real time to the result aggregator, and then transmitted to the risk assessor. The risk assessor performs a comprehensive risk assessment of pop-ups / ads based on preset risk weights and aggregation rules for multi-agent detection results, generating a final risk level. For example, ordinary pop-up ads are primarily for commercial promotion and typically do not carry malicious code, only causing interface interference and a decline in user experience when users visit web pages. According to the risk level scoring system, in terms of probability of occurrence, ordinary ads are a common online promotion method with a high frequency of occurrence, but are not direct attacks, therefore they are classified as medium risk (2 points); in terms of vulnerability, even if the system lacks ad-blocking measures, ordinary ads usually will not lead to system breaches, therefore they are classified as medium risk (2 points); in terms of impact, the consequences of ordinary ads are limited to impaired user experience and will not cause data leaks or business interruptions, therefore they are classified as low risk (1 point). The comprehensive risk value is calculated as 2 × 2 × 1 = 4 points, corresponding to a low risk level. Malicious pop-up ads, on the other hand, exhibit significant security threats. Attackers often use advertising networks or malicious scripts to deliver pop-up ads containing phishing links, Trojan programs, or insecure software, enticing users to click and trigger attacks. According to the risk level scoring system, in terms of probability, malicious ads are widely distributed, and attackers are sophisticated in their methods; users only need to visit the relevant page to become a target, therefore it is classified as high risk (3 points). Regarding vulnerabilities, if the user's browser or system has vulnerabilities or lacks effective protection mechanisms, clicking the ad makes it extremely easy to exploit, also classified as high risk (3 points). In terms of impact, malicious ads may lead to the leakage of sensitive information, system infection with malicious programs, or even significant economic losses, therefore it is classified as high-level (3 points). The overall risk score is 3 × 3 × 3 = 27 points, corresponding to a high-risk level.

[0063] (4) Dynamic response and tracing: Based on the comprehensive risk assessment results, the decision-making center will trigger corresponding security response measures: Immediately block: For high-risk or confirmed malicious pop-ups / ads.

[0064] Warning: Issue a warning to users for advertisements that are of low to medium risk or questionable.

[0065] Restrict interaction: Limit certain user actions on the advertisement, such as disabling click-through redirection.

[0066] Reporting Analysis: Submit suspicious samples for manual review or further threat intelligence analysis.

[0067] The log security agent continuously collects and correlates all relevant logs throughout the process. Once an anomaly is detected or a response is triggered, it can immediately provide a detailed behavioral tracing path and chain of evidence, supporting subsequent investigations, evidence collection, and blacklist updates. For example, if the content security agent determines that an advertisement is unsafe, the log security agent can quickly provide all loading records, source IPs, and user interaction records of that advertisement, providing a basis for blocking the advertisement source.

[0068] The innovative aspects of this invention are as follows: 1. Multi-agent cooperative system: Based on the three stages of perception, decision-making, and execution in which AI Agents complete tasks, six types of intelligent agents are combined: "intent, object, content, privacy, program, and log".

[0069] At the user intent level, the intent risk intelligence agent uses an intent-step node matching algorithm and a behavior sequence comparison model to accurately analyze the semantics of the user-initiated process and determine in real time whether the Agent's "step nodes" match the original intent, thereby identifying process deviations or malicious intents at an early stage. At the object operation level, the object security intelligent agent ensures the legality of the target object operated by the user and the compliance of the operation permission through fine-grained, multi-dimensional access control model and object model recognition technology, effectively preventing unauthorized operations and illegal access. At the content interaction level, the content security agent, leveraging multimodal analysis capabilities covering text, images, and videos, deeply analyzes user input or system-generated information through a built-in sensitive information detection engine and context-sensitive content recognition technology. This agent can accurately identify and filter sensitive words, prohibited images, and inappropriate remarks in real-time interactions, thereby achieving intelligent security review of various forms of content.

[0070] At the data privacy level, the privacy risk agent utilizes sensitive privacy information identification and classification algorithms, as well as data flow tracking and leakage risk assessment models, to accurately detect and classify personally identifiable information (PII) and other sensitive privacy information in user data. Simultaneously, this agent can assess the data transmission and storage process, determining whether there are leakage risks or non-compliant uses, thereby ensuring the privacy compliance of the data processing process.

[0071] At the program execution level, the program security intelligent agent combines static and dynamic hybrid program behavior analysis technology and runtime integrity verification mechanism to deeply analyze the static characteristics and runtime behavior of program code, and promptly detect code vulnerabilities, malicious behaviors or abnormal calls.

[0072] At the system operation level, the log security intelligent agent collects and analyzes various log data generated by the system through intelligent log aggregation, parsing and standardization technologies, as well as a multi-source log correlation analysis engine. It intelligently identifies abnormal events, attack attempts or violations, and provides comprehensive audit tracking and threat tracing capabilities.

[0073] The intelligent agents can exchange information and engage in collaborative reasoning, sharing analysis results and making joint judgments. This forms a comprehensive security monitoring system for the risky behavior of AI agents.

[0074] 2. Risk assessment model: The risk assessment module is based on a three-factor quantitative model: Risk Value = Probability Score × Vulnerability Score × Impact Score. Each factor is scored from 1 to 3 points. The final risk value is calculated using the formula: 1-6 points are defined as low risk, 7-14 points as medium risk, and 15-27 points as high risk. For high-risk behaviors, the system will immediately take mandatory interception or isolation measures, such as blocking related operations, suspending affected services, and isolating abnormal processes or users, to ensure that the risk does not spread. For medium-risk behaviors, the system will trigger warnings and mitigation actions, such as sending alarm notifications to administrators, restricting access to certain functions, conducting security audits, or implementing temporary hardening measures, so that management can intervene in a timely manner and mitigate potential impacts. For low-risk behaviors, the system typically logs and monitors them to facilitate subsequent analysis and security policy optimization without immediate intervention, thereby achieving efficient resource utilization.

[0075] By using quantifiable risk values, the system can assess risks in real time during operation and drive different levels of security response strategies, improving the accuracy and automation of risk management. This model not only quantifies risks in real time but also automatically triggers tiered security response strategies based on different risk levels. For example, high-risk events will immediately trigger interception or isolation measures, medium-risk events will trigger warnings and mitigation actions, while low-risk events will be logged for subsequent analysis. In this way, the system achieves accurate assessment, automated control, and dynamic response to the risky behavior of AI agents, significantly improving security protection efficiency and decision-making accuracy.

[0076] 3. Dynamically adaptive knowledge base: During system operation, the risk assessor collects agent analysis data and response results, which are fed back to the knowledge base in real time. This data is used to adjust security rules and optimize agent behavior parameters. The knowledge base dynamically updates security policies based on this feedback, ensuring that rules remain consistent with the system's current threat environment. Simultaneously, the knowledge base continuously evolves its risk model by collecting the latest attack samples and abnormal behavior patterns, improving the accuracy of risk prediction. By combining historical events and feedback results, the knowledge base can optimize the analysis strategies and thresholds of each agent, achieving self-learning and adaptive analysis.

[0077] Furthermore, the knowledge base works in conjunction with the risk assessment module and policy executor to form a closed-loop management system, enabling the system to continuously optimize risk control and response efficiency. Through this dynamic adaptive mechanism, the knowledge base not only provides rule and policy references but also becomes the core support for the system's continuous learning and intelligent optimization, ensuring the real-time performance, flexibility, and reliability of the entire security management system.

[0078] The above description is merely a detailed explanation of preferred embodiments and principles of the present invention. For those skilled in the art, there may be changes in specific implementation methods based on the ideas provided by the present invention, and these changes should also be considered within the scope of protection of the present invention.

Claims

1. A monitoring system for monitoring risky behavior of AI agents, characterized in that, include: The data perception and acquisition module is used to collect risk behavior signals and external data sources generated during the operation of the agent-type artificial intelligence (AI) agent in real time. The main controller is used to receive and distribute the risk behavior signals and external data sources, and to aggregate and assess the risk of the feedback analysis results. A multi-agent system includes several functionally independent sub-agents; the sub-agents are configured to receive data from the main controller and perform parallel analysis and risk assessment of the AI ​​Agent's behavior from six dimensions: intent risk, object security, content security, privacy risk, program security, and log security. The security response system is used to implement dynamic protection and emergency response for the execution environment of the AI ​​Agent based on the risk assessment results and response instructions issued by the main controller. The knowledge base stores security rules, threat models, and agent behavior parameters, and is dynamically updated based on system operation feedback.

2. The monitoring system for monitoring risky behavior of AI agents according to claim 1, characterized in that, The main controller includes: An input distributor is used to distribute received data to the corresponding sub-agents in the multi-agent system according to the corresponding type; The result aggregator is used to receive and aggregate risk assessment reports from each sub-agent in a multi-agent system to form structured information. The risk assessor is used to perform comprehensive analysis based on the structured information and the security rules in the knowledge base to generate an overall risk assessment report containing a unified risk level. The strategy executor is used to generate and issue standardized security response instructions to the security response system based on the overall risk assessment report and with reference to the response strategies defined in the knowledge base.

3. The monitoring system for monitoring risky behavior of AI agents according to claim 2, characterized in that, The risk assessor employs a quantitative risk assessment model, specifically calculating the risk value using the following formula: Risk value = Probability of occurrence score × Vulnerability score × Impact score; The probability score, vulnerability score, and impact score are all within the range of 1 to 3 points; the risk assessor classifies the risk level into low risk, medium risk, or high risk based on the calculated risk value range.

4. The monitoring system for monitoring risky behavior of AI agents according to claim 3, characterized in that, According to the risk value calculation formula, a final risk value of 1-6 points is defined as low risk; a final risk value of 7-14 points is defined as medium risk; and a final risk value of 15-27 points is defined as high risk. High-risk behaviors require immediate action to intercept; medium-risk behaviors require management attention and mitigation measures; low-risk behaviors can be tolerated and logged, without additional control.

5. The monitoring system for monitoring risky behavior of AI agents according to claim 1, characterized in that, The multi-agent system includes: The Intent Risk Intelligent Agent module is responsible for monitoring and analyzing abnormal operational intents of users or AI Agents; The object security intelligent agent module is responsible for assessing the access permissions of the target objects operated by the AI ​​Agent and the compliance of the operational behavior; The content security intelligent agent module is used for security detection of data content processed or generated by the AI ​​Agent; The privacy risk intelligent agent module is used to assess the risk of leakage of personal privacy information during data processing; The program security intelligent agent module is used to analyze code vulnerabilities and malicious behaviors during the execution of the AI ​​Agent program; The Log Security Intelligent Agent module is used to analyze system security logs and discover abnormal events, attack traces, errors, or warning messages.

6. The monitoring system for monitoring risky behavior of AI agents according to claim 5, characterized in that, In the multi-agent system, the sub-agents interact with each other and engage in collaborative reasoning, sharing their respective analysis results.

7. The monitoring system for monitoring risky behavior of AI agents according to claim 1, characterized in that, The security response system includes: The instruction parsing module is used to parse the security response instructions issued by the main controller and determine the execution priority; The strategy orchestration module, connected to the knowledge base, is used to generate specific handling strategies based on the parsed instruction parameters and the real-time status of the system. An execution engine cluster is used to execute the aforementioned processing strategy; The feedback loop module is used to verify whether the handling strategy is successful after the execution engine cluster completes the execution action, and at the same time feeds back the handling execution result and environmental status changes to the main controller and knowledge base.

8. The monitoring system for monitoring risky behavior of AI agents according to claim 7, characterized in that, The execution engine cluster includes: The blocking and isolation engine is used to perform processes to terminate, disconnect, or isolate resources for high-risk behaviors. The downgrade and restriction engine is used to perform permission downgrades, access rate restrictions, or enable human verification mechanisms for medium-risk behaviors. The auditing and tracing engine is used to capture full data packets, record screens, and retain contextual information for low-risk behaviors, and then label them as suspicious.

9. The monitoring system for monitoring risky behavior of AI agents according to claim 1, characterized in that, The knowledge base is a dynamic adaptive knowledge base, used to receive risk assessment results from the main controller and response effect feedback from the security response system, and dynamically update the security rules, threat models and agent behavior parameters stored in the knowledge base based on the risk assessment results and response effect feedback.

Citation Information

Patent Citations

  • Parallel simulation processing method and device for multiple agents, terminal equipment and storage medium

    CN120597457A

  • Dynamic authority management system and method and multi-source heterogeneous message middleware

    CN121000532A

  • Security defense strategy method and system based on AI Agent dynamic optimization

    CN121077781A

  • Zero-trust security policy intelligent generation method and system for autonomous AI agent

    CN121567383A