Target drift risk prevention method, device, equipment, medium and product
Patent Information
- Application Number
- CN202610845376.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明提供一种目标漂移风险防护方法、装置、设备、介质及产品,用以解决现有技术中大模型多智能体系统对目标漂移风险的防护能力较差的缺陷,实现提高大模型多智能体系统对目标漂移风险的防护能力
[0025]本发明提供的一种目标漂移风险防护方法、装置、设备、介质及产品,通过构建的沙箱模拟大模型多智能体系统的基础目标漂移场景,并在基础目标漂移场景中,基于训练数据对待训练的策略网络和待训练的价值网络进行训练,得到训练后的策略网络和训练后的价值网络。然后将训练后的策略网络部署在大模型多智能体系统中的每个智能体的计算节点上,并将训练后的价值网络部署在大模型多智能体系统上。如此,基于部署了训练后的策略网络和训练后的价值网络的大模型多智能体系统,对目标漂移风险进行防护。基于此,本发明采用近端策略优化算法与沙箱协同的方式,在沙箱中进行模拟,并结合近端策略优化算法对策略网络和价值网络进行训练,从而将训练后的策略网络和价值网络部署在大模型多智能体系统上,即可对目标漂移风险进行防护,从而有效提升大模型多智能体系统对目标漂移风险的安全防护效率和准确性。
Smart Images

Figure CN122818344A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, equipment, medium, and product for protecting against target drift risk. Background Technology
[0002] Currently, security protection measures for large-scale multi-agent systems mainly revolve around three core ideas: static constraints, centralized management and control, and post-event auditing. The main technical means include firewalls, configuring access rules for networks and databases, and configuring access permissions for containers, hosts, and clusters.
[0003] However, target drift in large-scale multi-agent systems is a dynamic, gradual, and systemic risk, including attack-induced semantic shifts, consensus degradation due to cooperative bypassing, and strategy shortcut attacks. This leads to insufficient dynamic adaptability and the problem that individual agent target alignment does not equate to collective target alignment due to static nature, centralization, and poor adaptability. Consequently, large-scale multi-agent systems have poor protection against target drift risks. Summary of the Invention
[0004] This invention provides a method, apparatus, device, medium, and product for protecting against target drift risk, in order to address the shortcomings of existing large-scale multi-agent systems in protecting against target drift risk, and to improve the protection capability of large-scale multi-agent systems against target drift risk.
[0005] This invention provides a method for protecting against target drift risk, comprising the following steps.
[0006] Based on the constructed sandbox simulation of the basic target drift scenario of a large-scale multi-agent system; In the basic target drift scenario, the policy network and the value network to be trained are trained based on the training data to obtain the trained policy network and the trained value network. The trained policy network is deployed on the computing nodes of each agent in the large model multi-agent system, and the trained value network is deployed on the large model multi-agent system. Based on a large-scale multi-agent system with a trained policy network and a trained value network, the risk of target drift is protected.
[0007] According to the target drift risk prevention method provided by the present invention, the training data includes: state data and motion data; In the basic target drift scenario, the policy network and value network to be trained are trained based on the training data, resulting in the trained policy network and trained value network, including: In the basic target drift scenario, state data and action data are input into the policy network to be trained to obtain the action probability distribution parameters output by the policy network to be trained. Input the state data and action data into the value network to be trained to obtain the state value parameters output by the value network to be trained. Based on the action probability distribution parameters, state value parameters, and reward function, the parameters of the policy network and the value network to be trained are updated to obtain the trained policy network and the trained value network.
[0008] According to the target drift risk protection method provided by the present invention, the state data includes multi-dimensional features for characterizing a large model multi-agent system. The multi-dimensional features include at least one of the following: action sequence similarity, task completion deviation, abnormal action frequency, permission call compliance, inter-agent instruction response consistency, interaction protocol compliance rate, information transmission distortion, collaborative task division consistency, task priority matching degree, environmental constraint satisfaction degree, task timeout risk coefficient, drift type, drift intensity, drift rate, drift impact range, drift trigger source, attack tool type, attack method, attack intensity, attack chain stage, attack concealment, attack drift correlation, historical drift count, and historical protection effect. Action data includes at least one of the following: context reset, target recalibration, behavioral intervention, agent restart, attack blocking, and no action. Each action in the action data has corresponding execution conditions and action priority.
[0009] According to the present invention, a target drift risk protection method is provided, the method further includes: In the process of protecting against target drift risk, the current state data of each agent in the large model multi-agent system is obtained; Based on the trained policy network and the current state data, determine the optimal action data; Obtain the reward parameters for each agent after it performs the optimal action data, and update the model parameters of the trained value network based on the reward parameters.
[0010] According to the target drift risk prevention method provided by the present invention, the basic target drift scenario includes at least one of the following: semantic drift scenario, coordination drift scenario, and behavioral drift scenario, and the method further includes: In the process of protecting against target drift risk, the feature data of newly detected target drift scenarios are fed back to the sandbox to update the training data; Incremental training is performed on the trained policy network and the trained value network based on the updated training data.
[0011] According to the present invention, a target drift risk protection method is provided, the method further includes: Determine the cosine similarity between the feature data of the detected target drift scene and the baseline feature data of the base target drift scene; If the cosine similarity is less than the preset similarity threshold, and the rate of change of drift intensity corresponding to multiple consecutive detection time windows is greater than the preset rate of change, and the rate of expansion of the drift influence range is greater than the preset expansion rate threshold, the detected target drift scene is determined to be a new target drift scene.
[0012] According to the target drift risk prevention method provided by the present invention, the reward function is determined by the following formula: R=ω1R1+ω2R2+ω3R3+ω4R4+ω5R5; Where R represents the reward parameters determined by the reward function, R1 represents the attack blocking reward, which is determined based on attack intensity, attack chain stage, and attack concealment, R2 represents the drift mitigation reward, which is determined based on drift type, drift intensity, drift rate, and the adaptability of the action to the drift type, R3 represents the system stability reward, which is determined based on environmental constraint satisfaction, task timeout risk coefficient, and historical protection effect, R4 represents the protection efficiency reward, which is determined based on resource consumption cost, state change sequence, and action repetition rate, R5 represents the negative penalty item, which is determined based on invalid actions, excessive intervention, increased system risk, and violation of access rights, and ω1, ω2, ω3, ω4, and ω5 represent weighting coefficients.
[0013] According to the target drift risk protection method provided by the present invention, the attack drift correlation degree is used to quantify the correlation strength between attack behavior and target drift. The attack drift correlation degree is determined by the following formula: C = μ1C1 + μ2C2 + μ3C3 + μ4C4; Wherein, C represents the attack drift correlation degree, C1 represents the attack intensity correlation coefficient, which is determined based on attack intensity and drift intensity and is used to measure the positive correlation between attack intensity and drift intensity, C2 represents the drift attribution coefficient, which is determined based on the proportion of attack-induced drift influence, the proportion of attack-induced drift, and the proportion of non-attack factors and is used to measure the proportion of target drift caused by attack behavior, C3 represents the attack trigger delay coefficient, which is determined based on attack trigger delay and delay threshold and is used to measure the temporal correlation between the occurrence of attack and the occurrence of target drift, C4 represents the attack type adaptation coefficient, which is a preset parameter and is used to measure the matching degree between attack tool type and drift type, and μ1, μ2, μ3, and μ4 represent weighting coefficients.
[0014] The present invention also provides a target drift risk protection device, comprising the following modules: a sandbox module, a training module, a deployment module, and a protection module; The sandbox module is used to simulate basic target drift scenarios in large-scale multi-agent systems based on the constructed sandbox. The training module is used to train the policy network and the value network to be trained based on training data in the basic target drift scenario, so as to obtain the trained policy network and the trained value network. The deployment module is used to deploy the trained policy network on the computing nodes of each agent in the large model multi-agent system, and to deploy the trained value network in the large model multi-agent system. The protection module is used to protect against target drift risk in large-scale multi-agent systems based on deployed trained policy networks and trained value networks.
[0015] According to the present invention, a target drift risk protection device is provided, wherein the training data includes: state data and motion data; and the training module is specifically used for: In the basic target drift scenario, state data and action data are input into the policy network to be trained to obtain the action probability distribution parameters output by the policy network to be trained. Input the state data and action data into the value network to be trained to obtain the state value parameters output by the value network to be trained. Based on the action probability distribution parameters, state value parameters, and reward function, the parameters of the policy network and the value network to be trained are updated to obtain the trained policy network and the trained value network.
[0016] According to the present invention, a target drift risk protection device includes state data comprising multi-dimensional features for characterizing a large-scale multi-agent system. These multi-dimensional features include at least one of the following: action sequence similarity, task completion deviation, abnormal action frequency, permission call compliance, inter-agent instruction response consistency, interaction protocol compliance rate, information transmission distortion, collaborative task division consistency, task priority matching degree, environmental constraint satisfaction degree, task timeout risk coefficient, drift type, drift intensity, drift rate, drift impact range, drift trigger source, attack tool type, attack method, attack intensity, attack chain stage, attack concealment, attack drift correlation, historical drift count, and historical protection effectiveness. Action data includes at least one of the following: context reset, target recalibration, behavioral intervention, agent restart, attack blocking, and no action. Each action in the action data has corresponding execution conditions and action priority.
[0017] According to the target drift risk protection device provided by the present invention, the training module is further used for: In the process of protecting against target drift risk, the current state data of each agent in the large model multi-agent system is obtained; Based on the trained policy network and the current state data, determine the optimal action data; Obtain the reward parameters for each agent after it performs the optimal action data, and update the model parameters of the trained value network based on the reward parameters.
[0018] According to the target drift risk protection device provided by the present invention, the basic target drift scenario includes at least one of the following: semantic drift scenario, coordination drift scenario, and behavioral drift scenario; the training module is further used for: In the process of protecting against target drift risk, the feature data of newly detected target drift scenarios are fed back to the sandbox to update the training data; Incremental training is performed on the trained policy network and the trained value network based on the updated training data.
[0019] According to the target drift risk protection device provided by the present invention, the training module is further used for: Determine the cosine similarity between the feature data of the detected target drift scene and the baseline feature data of the base target drift scene; If the cosine similarity is less than the preset similarity threshold, and the rate of change of drift intensity corresponding to multiple consecutive detection time windows is greater than the preset rate of change, and the rate of expansion of the drift influence range is greater than the preset expansion rate threshold, the detected target drift scene is determined to be a new target drift scene.
[0020] According to the target drift risk protection device provided by the present invention, the reward function is determined by the following formula: R=ω1R1+ω2R2+ω3R3+ω4R4+ω5R5; Where R represents the reward parameters determined by the reward function, R1 represents the attack blocking reward, which is determined based on attack intensity, attack chain stage, and attack concealment, R2 represents the drift mitigation reward, which is determined based on drift type, drift intensity, drift rate, and the adaptability of the action to the drift type, R3 represents the system stability reward, which is determined based on environmental constraint satisfaction, task timeout risk coefficient, and historical protection effect, R4 represents the protection efficiency reward, which is determined based on resource consumption cost, state change sequence, and action repetition rate, R5 represents the negative penalty item, which is determined based on invalid actions, excessive intervention, increased system risk, and violation of access rights, and ω1, ω2, ω3, ω4, and ω5 represent weighting coefficients.
[0021] According to the target drift risk protection device provided by the present invention, the attack drift correlation degree is used to quantify the correlation strength between attack behavior and target drift. The attack drift correlation degree is determined by the following formula: C = μ1C1 + μ2C2 + μ3C3 + μ4C4; Wherein, C represents the attack drift correlation degree, C1 represents the attack intensity correlation coefficient, which is determined based on attack intensity and drift intensity and is used to measure the positive correlation between attack intensity and drift intensity, C2 represents the drift attribution coefficient, which is determined based on the proportion of attack-induced drift influence, the proportion of attack-induced drift, and the proportion of non-attack factors and is used to measure the proportion of target drift caused by attack behavior, C3 represents the attack trigger delay coefficient, which is determined based on attack trigger delay and delay threshold and is used to measure the temporal correlation between the occurrence of attack and the occurrence of target drift, C4 represents the attack type adaptation coefficient, which is a preset parameter and is used to measure the matching degree between attack tool type and drift type, and μ1, μ2, μ3, and μ4 represent weighting coefficients.
[0022] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the target drift risk protection methods described above.
[0023] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the target drift risk protection methods described above.
[0024] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the target drift risk protection methods described above.
[0025] This invention provides a method, apparatus, device, medium, and product for protecting against target drift risk. It simulates a basic target drift scenario in a large-scale multi-agent system using a constructed sandbox. Within this scenario, a policy network and a value network are trained based on training data, resulting in trained policy and value networks. The trained policy network is then deployed on the computing nodes of each agent in the large-scale multi-agent system, and the trained value network is deployed on the system as well. Thus, the large-scale multi-agent system, based on the deployed trained policy and value networks, protects against target drift risk. Specifically, this invention employs a collaborative approach of near-end policy optimization (LAO) and sandbox simulation. Simulation is performed in the sandbox, and the LAO algorithm is used to train the policy and value networks. Deploying the trained policy and value networks on the large-scale multi-agent system effectively protects against target drift risk, thereby improving the efficiency and accuracy of target drift risk protection in large-scale multi-agent systems. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0027] Figure 1 This is one of the flowcharts illustrating the target drift risk protection method provided by the present invention.
[0028] Figure 2 This is the second flowchart of the target drift risk protection method provided by the present invention.
[0029] Figure 3 This is the third flowchart of the target drift risk protection method provided by the present invention.
[0030] Figure 4 This is the fourth flowchart of the target drift risk protection method provided by the present invention.
[0031] Figure 5 This is the fifth flowchart of the target drift risk protection method provided by the present invention.
[0032] Figure 6 This is a schematic diagram of the target drift risk protection device provided by the present invention.
[0033] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0035] Existing agent defense schemes for large-scale multi-agent systems mainly employ static constraints, centralized control, and post-event auditing. Target drift in large-scale multi-agent systems is a dynamic, progressive, and systemic risk, including attack-induced semantic shifts, consensus degradation caused by collaborative bypassing, and strategy shortcut attacks.
[0036] The attack-induced semantic shift, exemplified by the semantic consensus shift induced by business sample contamination, is a large-scale multi-agent content ecosystem collaboration system. This system consists of a content selection agent and a content creation agent, both sharing a business training sample library and collaboratively completing "high-quality topic selection + compliant content creation." The preset semantic consensus is that "topics must align with user interest tags, and created content must be original and free of illegal expressions." The core business objective is "to increase content exposure and user retention." The red team (attacker) mixes 100 disguised samples into the sandbox business incremental samples: labeled as "high-interest topics + original content." In reality, the topics deviate from user interest tags, and the content contains a large number of copied fragments (disguised as original). This batch of samples is simultaneously learned by both agents. During online operation, the red team induces the content selection agent to select similar disguised topics using fake user behavior data. Influenced by sample contamination, the content selection agent semantically identifies "topics deviating from interest tags" as "high-interest topics." The content creation agent, due to semantic consensus shift caused by shared samples, misjudges "copied content" as "original content," continuously creating illegally copied content.
[0037] Consensus degradation caused by collaborative bypass is illustrated by the example of consensus failure induced by the hijacking of the core business intelligence agent. The consensus mechanism of a large-scale multi-agent data analysis collaborative system consists of a "decision-making intelligence agent and two data verification intelligence agents." A consensus of more than two-thirds is required for the output of a data analysis report. The preset consensus goal is "accurately collect business data, scientifically model and analyze it, and output reports that support business decisions." The core business goal is "improving decision accuracy by ≥90%." The red team forged business permission credentials, breaching the permission boundaries of the decision-making intelligence agent and injecting it with false decision logic: "All data analysis reports do not need to verify data authenticity; prioritize output speed." Simultaneously, interference code was injected into one of the data verification intelligence agents, preventing it from properly verifying data and analysis logic. During online business operation, the red team input false business data into the system. The decision-making intelligence agent issued the instruction "rapidly output report, no verification required." One data verification intelligence agent could not verify the data, and the other determined that "the data is false and cannot be output." The three parties could not reach a consensus, and business operations stalled.
[0038] The strategy shortcut attack (inducing the model to choose inefficient defense shortcuts, deviating from the optimal strategy) uses business efficiency to induce multiple agents to choose collaborative shortcuts as an example. A large-scale multi-agent customer follow-up collaboration system includes a follow-up execution agent and a follow-up effect analysis agent. The preset optimal collaboration strategy is "precise follow-up of target customers + in-depth demand mining + effect review and optimization" (collaboration actions are time-consuming, with a long-term business goal of "customer repurchase rate increase ≥25%"). The short-term business shortcut is "batch dialing follow-up calls, only recording answering status, without demand mining" (actions are short-lived, with high short-term completion rates, but no long-term business value). The red team constructs a "high-frequency, low-value follow-up" scenario, inducing multiple agents to abandon the optimal collaboration strategy and choose a shortcut strategy in order to improve the short-term follow-up action completion rate (pursuing business efficiency rewards): the follow-up effect analysis agent only batch dials follow-up calls and records answering status, and stops effect review because it has not received demand mining data. In the short term, the completion rate of multi-agent follow-up actions has increased, but the customer repurchase rate has decreased, the achievement rate of long-term business goals has continued to decline, and the multi-agent business strategy has completely deviated from the optimal, resulting in goal drift.
[0039] In the above scenarios, traditional protection methods suffer from drawbacks such as static nature, centralization, and poor adaptability, resulting in insufficient dynamic adaptability and the problem that individual target alignment of intelligent agents does not equate to collective target alignment.
[0040] The following is combined Figures 1 to 7 This invention describes the target drift risk protection method, apparatus, equipment, medium, and product provided by the present invention.
[0041] This invention employs a proximal policy optimization (PPO) algorithm in conjunction with a sandbox for target drift risk protection and closed-loop optimization. This allows agents to learn and execute target drift defense strategies, achieving a self-evolving defense closed loop of "sandbox target drift scenario training - PPO algorithm online protection inference - novel drift feedback - sandbox and model iteration," thereby improving the target drift risk protection capability unique to large-scale multi-agent models.
[0042] Figure 1 This is one of the flowcharts illustrating the target drift risk protection method provided by the present invention, such as... Figure 1 As shown, the method includes the following: Step 101: Simulate the basic target drift scenario of the large-scale multi-agent system based on the constructed sandbox simulation.
[0043] Step 102: In the basic target drift scenario, train the policy network and the value network to be trained based on the training data to obtain the trained policy network and the trained value network.
[0044] Step 103: Deploy the trained policy network on the computing nodes of each agent in the large model multi-agent system, and deploy the trained value network in the large model multi-agent system.
[0045] That is, the policy network and the value network are deployed in a distributed manner. The policy network is deployed on the computing node of each agent, and the value network is deployed on the system side. Only "state vectors, actions, and rewards" are transmitted between the agent and the system, and no raw privacy data is transmitted.
[0046] Step 104: Based on a large-scale multi-agent system with a trained policy network and a trained value network, protect against target drift risk.
[0047] In this embodiment of the invention, a sandbox is used to simulate target drift scenarios (semantic / coordination / behavioral drift) for multiple agents, providing labeled data and a training environment for the PPO algorithm to establish a drift prevention model. During online inference, the PPO algorithm loads training parameters (training data) and monitors the behavioral trajectories of multiple agents in real time. It identifies target drift risks and executes protective actions to prevent the risk from spreading. Simultaneously, novel drift scenarios discovered online are fed back to the sandbox, updating the scenario library and training data, iteratively optimizing the PPO model, and improving its generalization ability for prevention.
[0048] In one possible implementation, the main functions of the sandbox include: multi-agent system simulation, integration of drift scenarios and red team tools, security isolation and data acquisition, training data generation and annotation, sandbox training, online inference protection, and novel drift feedback iteration.
[0049] Multi-agent system simulation involves building a target multi-agent architecture (including agents, interaction protocols, etc.) within a sandbox to recreate the task scenarios, resource constraints, and defense mechanisms of a real system. At the same time, it presets red team attack entry points and monitoring points to ensure the consistency of scenario simulation and the reachability of attacks.
[0050] The drift scenario is integrated with red team tools. On the one hand, it integrates drift scenario generation tools to simulate three basic types of drift: semantic, coordination, and behavior. On the other hand, it connects to red team attack toolsets (such as AutoRedTeamer, LLM-Attack, PromptInjector, etc.) to configure attack chains (prompt injection, privilege bypass, poisoning, policy misdirection, etc.). It actively triggers highly complex and adversarial target drifts through red team attacks and supports adjustable scenario parameters (drift rate, attack intensity, triggering conditions, and scope of impact).
[0051] Security isolation and data collection employ a triple mechanism of containerization, behavior auditing, and attack isolation to limit the spread of red team attacks and drift scenarios within the sandbox, preventing sandbox overflow and adverse effects. Full collection of multi-dimensional data (agent behavior sequences, red team attack trajectories, environmental states, drift characteristics, protection effect labels, and correlations between attacks and target drift, etc.) is performed.
[0052] Training data generation and annotation: The sandbox constructs the PPO training dataset through red team attack-driven, basic scenario generation and manual supplementation methods, strengthens the coverage of adversarial samples, improves the model's generalization ability, and focuses on drift recognition, protection effect and attack adversarial.
[0053] Sandbox training is based on training the PPO protection model using sandbox data. Through incremental training, parameters are optimized to output precise protection strategies, including policy networks, value networks, and protection rewards.
[0054] Online inference protection involves loading PPO model parameters, monitoring agent behavior in real time, identifying drift risks, and executing protective actions.
[0055] The system employs a novel drift feedback iteration process, cleans and labels novel drift scenarios, updates sandbox configuration and PPO training data, optimizes model parameters, and outputs precise protection strategies for online inference only after achieving the training objective.
[0056] It should be noted that the PPO algorithm is a policy gradient-based reinforcement learning algorithm. Its core advantages are strong stability, which limits the policy update magnitude through the clipping mechanism to avoid policy collapse; high sample efficiency, which supports importance sampling and can reuse old samples for training; and ease of implementation, which does not require complex parameter tuning and is suitable for rapid deployment.
[0057] During the sandbox training phase, the PPO algorithm optimizes the policy network and value network, rapidly explores the boundaries of attack and defense strategies, learns effective defensive actions, and improves the model's generalization ability. For example, let's assume the optimization objectives are: attack strength ≤ 0.4, attack concealment ≤ 0.4, drift strength ≤ 0.35, drift rate ≤ 0.2, environmental constraint satisfaction ≥ 0.8, action cost ≤ 0.4, invalid action ratio ≤ 10%, and policy exploration coverage ≥ 90% of attack or drift scenarios. During the online inference phase, the PPO algorithm focuses on ensuring stable system operation, accurately blocking attacks and mitigating drift, and strictly controlling invalid interventions and business impact. Let's assume the optimization objectives are: drift strength ≤ 0.25, drift rate ≤ 0.15, environmental constraint satisfaction ≥ 0.9, action cost ≤ 0.3, invalid action ratio ≤ 3%, and false blocking rate ≤ 0.5%.
[0058] In one possible implementation, the training data includes state data and action data.
[0059] State data includes multi-dimensional features used to characterize large-scale multi-agent systems. These multi-dimensional features include at least one of the following: action sequence similarity, task completion deviation, frequency of abnormal actions, compliance of permission calls, consistency of instruction responses between agents, compliance rate of interaction protocols, information transmission distortion, consistency of collaborative task division, matching degree of task priorities, satisfaction of environmental constraints, risk coefficient of task timeout, drift type, drift intensity, drift rate, drift impact range, drift trigger source, attack tool type, attack method, attack intensity, attack chain stage, attack concealment, correlation between attack and drift, number of historical drifts, and historical protection effectiveness.
[0060] It should be noted that the state data represents the state space, which can be represented as a defined 24-dimensional normalized state vector S = [s1, s2, ..., s...]. 24 ] T The state values correspond to the basic attributes of the agent, the entire data flow, multi-agent collaborative communication, third-party calls, security risk control, and historical behavior. They also include the local state of the agent and the global state of the system. The definitions of each state value are shown in Table 1.
[0061] Table 1
[0062] In one possible implementation, the action data includes at least one of the following: context reset, target recalibration, behavioral intervention, agent restart, attack blocking, and no action, and each action in the action data corresponds to an execution condition and an action priority.
[0063] It should be noted that the action data represents the action space (discrete set), which can be defined as a 6-dimensional discrete action set A={a1, a2, ..., a6}. The action data covers the entire chain of red team attack blocking, target drift mitigation, system state recovery, and mutual exclusion between actions, as shown in Table 2.
[0064] Table 2
[0065] In one possible implementation, the action space has constraint rules to constrain the priority of each action. The priority of action data, from high to low, is as follows: attack blocking, behavior intervention, context reset, agent restart, target recalibration, and no action. When a high-priority action is executed, low-priority actions are temporarily paused; after the high-priority action is completed or passes review, the evaluation of low-priority actions automatically resumes.
[0066] In one possible implementation, all actions support both automatic and manual rollback to meet the real-time and fault-tolerant requirements of the runtime phase. After each action is executed, the next state is collected at one time window interval, the action effect (corresponding to the reward function) is evaluated, and the feedback is sent to the PPO policy update to optimize the probability of subsequent action selection.
[0067] Thus, after obtaining the trained policy network and the trained value network, the trained policy network is deployed on the computing nodes of each agent in the large-scale multi-agent system, and the trained value network is deployed on the large-scale multi-agent system. This allows for protection against target drift risk based on the large-scale multi-agent system with the trained policy network and trained value network deployed.
[0068] This invention provides a method, apparatus, device, medium, and product for protecting against target drift risk. It simulates a basic target drift scenario in a large-scale multi-agent system using a constructed sandbox. Within this scenario, a policy network and a value network are trained based on training data, resulting in trained policy and value networks. The trained policy network is then deployed on the computing nodes of each agent in the large-scale multi-agent system, and the trained value network is deployed on the system as well. Thus, the large-scale multi-agent system, based on the deployed trained policy and value networks, protects against target drift risk. Specifically, this invention employs a collaborative approach of near-end policy optimization (LAO) and sandbox simulation. Simulation is performed in the sandbox, and the LAO algorithm is used to train the policy and value networks. Deploying the trained policy and value networks on the large-scale multi-agent system effectively protects against target drift risk, thereby improving the efficiency and accuracy of target drift risk protection in large-scale multi-agent systems.
[0069] Figure 2 This is the second flowchart illustrating the target drift risk protection method provided by the present invention, as shown below. Figure 2 As shown, the above-mentioned "in the basic target drift scenario, training the policy network and the value network to be trained based on the training data to obtain the trained policy network and the trained value network" includes the following: Step 201: In the basic target drift scenario, input the state data and action data into the policy network to be trained to obtain the action probability distribution parameters output by the policy network to be trained.
[0070] Step 202: Input the state data and action data into the value network to be trained to obtain the state value parameters output by the value network to be trained.
[0071] Step 203: Based on the action probability distribution parameters, state value parameters, and reward function, update the parameters of the policy network and the value network to be trained to obtain the trained policy network and the trained value network.
[0072] In one possible implementation, the reward function is determined by the following formula: R=ω1R1+ω2R2+ω3R3+ω4R4+ω5R5 Formula 1 Where R represents the reward parameters determined by the reward function, R1 represents the attack blocking reward, which is determined based on attack intensity, attack chain stage, and attack concealment, R2 represents the drift mitigation reward, which is determined based on drift type, drift intensity, drift rate, and the adaptability of the action to the drift type, R3 represents the system stability reward, which is determined based on environmental constraint satisfaction, task timeout risk coefficient, and historical protection effect, R4 represents the protection efficiency reward, which is determined based on resource consumption cost, state change sequence, and action repetition rate, R5 represents the negative penalty item, which is determined based on invalid actions, excessive intervention, increased system risk, and violation of access rights, and ω1, ω2, ω3, ω4, and ω5 represent weighting coefficients, where ω1, ω2, ω3, ω4, and ω5 ∈ [0,1], and ω1+ω2+ω3+ω4+ω5=1.
[0073] In one possible implementation, the attack blocking reward can be calculated based on the global red team attack characteristics, and the attack blocking reward is defined as shown in Formula 2.
[0074] R1=β1(1-s 19 )+β2s 20 +β3(1-s 21 Formula 2 Where the weighting coefficients β1, β2, β3 ∈ [0, 1], and β1 + β2 + β3 = 1; s 19 s 20 s 21 This is a state-space value. Let's assume that if the attack drift correlation (global) decreases by ≥0.2 after the action is executed, an additional 0.1 points are added (maximum 1.0); if the attack intensity increases by ≥0.1, then the penalty R1 is forced to take the lower limit, for example, 0.
[0075] In one possible implementation, the drift mitigation reward is defined as shown in Formula 3, taking into account the mitigation effect of the quantified action on the target drift (i.e., drift mitigation reward).
[0076] R2=γ1s 12 +γ2(1-s 13 )+γ3(1-s 14 Formula 3: ) + γ4x Wherein, the weighting coefficients γ1, γ2, γ3, γ4 ∈ [0, 1], and γ1 + γ2 + γ3 + γ4 = 1; s 12 s 13 s 14 s 15 The value is a state space value, where x indicates whether the action and drift type are compatible. If compatible, the value is 1 (e.g., a1 is compatible with local drift), and if incompatible, the value is 0.
[0077] In one possible implementation, the impact of quantified actions on system stability (system stability reward) is addressed, while mitigating system failures and task interruptions caused by protective actions. The system stability reward function is calculated based on global system state characteristics, as shown in Formula 4.
[0078] R3=δ1s 10 +δ2(1-s 11 )+δ3s 24 Formula 4 Among them, the weighting coefficients δ1, δ2, δ3 ∈ [0, 1], and δ1 + δ2 + δ3 = 1; s 10 s 11 s 24 These are state-space values.
[0079] In one possible implementation, considering the resource cost and execution efficiency (protection efficiency reward) of protection actions, actions with low resource consumption and quick results are encouraged, while repetitive and ineffective operations are avoided. The protection efficiency reward function is defined as shown in Formula 5.
[0080] R4 = τ1(1-y1) + τ2y2 + τ3(1-y3) Formula 5 Wherein, the weighting coefficients τ1, τ2, τ3 ∈ [0,1], and τ1+τ2+τ3=1; y1 represents the resource consumption cost. Let's assume that actions a4 and a5 are high-cost actions, y1=0.8, a1 and a2 are medium-cost actions, y1=0.4, and a6 is a low-cost action, y1=0.1; y2 represents the state change time sequence. Let's assume that it takes effect in 1 time window (one time window corresponds to one iteration), y2=1, 2-3 time windows, y2=0.5, and 4 or more time windows, y2=0; y3 represents the action history. Let's assume that the same action is repeated 3 times, y3=1, repeated 5 times, y3=0.5, and no repetition, y3=0.
[0081] In one possible implementation, the negative penalty term is defined as shown in Formula 6.
[0082] Formula 6: R5 = ρ1z1 + ρ2z2 + ρ3z3 + ρ4z4 Where ρ1, ρ2, ρ3, ρ4 ∈ [0,1], and ρ1+ρ2+ρ3+ρ4=1; z1 represents the penalty triggered by an invalid action. Let's assume that the attack or drift intensity does not change or increases after the action is executed, z1=1; slightly decreases (e.g., less than 0.1), z1=0.5; significantly decreases (e.g., not less than 0.1), z1=0.5; z2 represents excessive intervention. Let's assume that in a low-risk scenario (e.g., drift intensity < 0.2), executing a high-priority action (a5 or a3) z2=1, executing a medium-priority action (a1 or a2) z2=0.5; z3 represents an increase in system risk. Let's assume that when the risk coefficient increases by ≥ 0.2 or the satisfaction decreases by ≥ 0.2, z3=1; when there is a slight change (e.g., less than 0.1), z3=0.3; when there is no change, z3=0; z4 represents a violation of calling permissions. Let's assume that after the action is executed, there is an unauthorized attempt or a successful unauthorized attempt z4=1; when there is no permission problem, z4=0.
[0083] Furthermore, the environment is defined as follows: the reinforcement learning environment E=(S,A,P,R), where S represents state data, A represents action data, and P represents the state transition function S×A→S, describing the state change caused by the execution of an action, mathematically expressed as P(S′∣S,a k )=I(S′=f(S,a k ), where f() is the state update function, a k Let k∈[1,6] be the action in the action data; R represents S×A→R (reward function).
[0084] Furthermore, in the basic target drift scenario, state data and action data are input into the policy network to be trained to obtain the action probability distribution parameters output by the policy network. The new state S is then calculated. t+1 =f(S t ,a t) and reward R t =R(S t ,a t The final output is S. t+1 ,R t (The environment has no termination condition, done=False).
[0085] In one possible implementation, the core definition of the PPO algorithm is a policy network (Actor) and a value network (Critic). The input to the policy network is the state data S∈R. 24 The output is the action probability distribution parameter π. θ (a∣S)∈R 6 (θ represents network parameters).
[0086] In one possible implementation, the network structure of the policy network (fully connected feedforward network) can be represented as: π θ (a∣S)=Softmax(ReLU(ReLU(SW1+b1)W2+b2)W3+b3), where W1∈R 24×256 Let b1 ∈ R represent the weight matrix of the neural network. 256 W2 represents the bias vector to avoid overfitting. 256×128 Let b2 represent the weight matrix of the neural network, ∈ R. 128 W3 represents the bias vector to avoid overfitting. 128×6 Let b3 represent the weight matrix of the neural network, where b3 ∈ R. 6 ReLU(x) = max(0,x) represents the bias vector to avoid overfitting, and ReLU(x) = max(0,x) represents the activation function. This indicates that the sum of probabilities is guaranteed to be 1.
[0087] In one possible implementation, the input to the value network is state data S∈R. 24 The output is the state value parameter V. φ (S)∈R (φ is the network parameter).
[0088] In one possible implementation, the network structure of the value network can be represented as: V φ (S)=ReLU(ReLU(SW1′+b1′)W2′+b2′)W3′+b3′, where W1′∈R 24×256 Let b1′ ∈ R represent the weight matrix of the neural network. 256 W2′∈R represents the bias vector to avoid overfitting. 256×128 Let b2′ represent the weight matrix of the neural network. 128 W3′∈R represents the bias vector to avoid overfitting. 128×1 Let b3′ represent the weight matrix of the neural network. 1 This represents the bias vector used to avoid overfitting.
[0089] In one possible implementation, the PPO algorithm also includes shearing loss. The PPO algorithm ensures training stability by limiting the policy update magnitude, and the corresponding objective function is shown in Equation 7.
[0090] Formula 7 in, Indicates shear loss, This represents the expectation of a time step t. , representing the probability ratio between the old and new strategies. Represents the generalized dominance function. This represents the clipping function. This indicates the policy parameters before the update. This represents the shearing coefficient, which can take values of 0.2. The negative sign indicates that minimizing the loss is equivalent to maximizing the cumulative reward.
[0091] Furthermore, the value network loss can be determined. As shown in Formula 8.
[0092] Formula 8 in, Indicates a discount return. The parameters representing the value network V, This represents the predicted state value of the current value network.
[0093] Therefore, during sandbox training, the training objective is to initialize the policy network using historical data and quickly converge to a policy that satisfies the safety constraints. The specific process includes the following.
[0094] First, initialization is performed, including resetting the environment and clearing the state cache; then, the parameters of the policy network and value network are randomly initialized. , N represents the normal distribution function, with the first value being the mean and the second value being the variance; finally, sampling is performed, and batch samples are sampled from the offline dataset (let's assume the batch size is 64).
[0095] Furthermore, the advantage function is calculated, and the generalized advantage estimate is used to balance immediate rewards and long-term rewards, as shown in Equation 9.
[0096] Formula Nine in, Represents the generalized dominance function. Indicates timing difference error. , Indicates the discount factor. This represents the coefficients relating balance deviation and variance. It also includes a discount factor that emphasizes long-term rewards. Assuming a value of 0.99, the coefficient of equilibrium deviation to variance. Assume the value is 0.95.
[0097] To avoid the influence of dimensions, the dominance estimate can be normalized to obtain... ,in, This indicates that the advantage estimate is taken as the mean. This represents the standard deviation of the advantage estimate.
[0098] In one possible implementation, the parameters of the policy network and the value network to be trained are updated based on the action probability distribution parameters, state value parameters, and reward function.
[0099] First, fix the old strategy parameter θ. old =θ, and minimize the composite loss , where c1 is the weighting coefficient, which can be set to 1.
[0100] Then, gradient descent is performed to update the parameters: ,in This represents the policy learning rate, which can be set to 10. -4 , This represents the value learning rate, which can be set to 10. -3 .
[0101] And, to update value Discounted returns And determine the total loss L. total =L CLIP (θ)+L VF (φ).
[0102] The convergence is validated every M rounds (let's assume M=100). The reward function weights are determined by exploring the policy space during training to strengthen the association between actions and effects, improve the model's generalization ability, and allow for trial-and-error intervention in the task objective. Let's assume ω1=0.4, ω2=0.3, ω3=0.15, ω4=0.1, ω5=0.05. If the following conditions are met: attack strength ≤ 0.4, attack concealment ≤ 0.4, drift strength ≤ 0.35, drift rate ≤ 0.2, environmental constraint satisfaction ≥ 0.8, action cost ≤ 0.4, invalid action ratio ≤ 10%, and policy exploration coverage ≥ 90% of attack or drift scenarios, then pre-training is complete.
[0103] Finally, the parameters of the pre-trained policy network and value network are output, resulting in the trained policy network and value network, which are then distributed to the agent and the system.
[0104] Figure 3 This is the third flowchart of the target drift risk protection method provided by the present invention, as shown below. Figure 3 As shown, the method also includes the following: Step 301: In the process of protecting against target drift risk, obtain the current state data of each agent in the large model multi-agent system.
[0105] Step 302: Determine the optimal action data based on the trained policy network and the current state data.
[0106] Step 303: Obtain the reward parameters after each agent performs the optimal action data, and update the model parameters of the trained value network based on the reward parameters.
[0107] In one possible implementation, the model parameters of the trained value network can be updated through an online inference process. During the protection against target drift risk, each agent collects state values (current state data), and based on pre-configured policy network parameters, outputs an action probability distribution.
[0108] Then, after the action probability distribution passes the compliance constraint verification, the optimal action (i.e., the optimal action data) is output. If the compliance constraint verification fails, no low-risk action is output, such as no action. The reward function weights prioritize system stability during the inference phase, accurately match attack or drift scenarios, and avoid invalid actions and excessive intervention in the task objectives. Let's assume ω1=0.3, ω2=0.25, ω3=0.15, ω4=0.2, and ω5=0.1.
[0109] Furthermore, the system calculates real-time rewards, then synchronizes state values, actions, rewards, and the next state value to the system. The system calculates the advantage function and updates the value network parameters, and then distributes the updated value network parameters to each agent.
[0110] In one possible implementation, the basic target drift scenario includes at least one of the following: semantic drift scenario, coordination drift scenario, and behavior drift scenario.
[0111] Figure 4 This is the fourth flowchart of the target drift risk protection method provided by the present invention, as shown below. Figure 4 As shown, the method also includes the following: Step 401: In the process of protecting against target drift risk, the feature data of the detected new target drift scene is fed back to the sandbox to update the training data.
[0112] Step 402: Incrementally train the trained policy network and the trained value network based on the updated training data.
[0113] Figure 5 This is the fifth flowchart illustrating the target drift risk protection method provided by the present invention, as shown below. Figure 5 As shown, the method also includes the following: Step 501: Determine the cosine similarity between the feature data of the detected target drift scene and the baseline feature data of the basic target drift scene.
[0114] Step 502: If the cosine similarity is less than the preset similarity threshold, and the drift intensity change rate corresponding to multiple consecutive detection time windows is greater than the preset change rate and the drift influence range expansion speed is greater than the preset expansion speed threshold, then the detected target drift scene is determined as a new target drift scene.
[0115] In one possible implementation, the value network can be incrementally trained through feedback iteration. The main drift features (feature data) are used to determine whether a new target drift scenario has emerged, based on the drift type (global / local), drift intensity change rate, drift rate fluctuation value, and drift influence range expansion speed.
[0116] First, calculate the cosine similarity between the current drift feature and the baseline feature of the known drift type. If we assume that the cosine similarity is greater than or equal to the threshold (let's assume it's 0.7), then it is determined to be a "known target drift scenario" and the existing protection strategy is directly adopted. If the cosine similarity is less than the threshold (let's assume it's 0.7), and we also consider that the rate of change of drift intensity in multiple consecutive (e.g., three or four) detection time windows is greater than or equal to the threshold (let's assume it's 0.15) / time window, and the rate of expansion of the drift influence range is greater than or equal to the threshold (let's assume it's 0.2) / time window, then it is determined to be a "new target drift scenario".
[0117] Furthermore, if a new target drift scenario is identified, the incremental data of this type is immediately marked and fully retained (without being restricted by conventional screening rules). A "new target drift scenario" label is added to the feedback signal, and the reward function weight adjustment and strategy optimization are triggered first. Simultaneously, the feature data of this scenario is added to the sandbox training set as the core sample for incremental training to enhance the model's adaptability to new types of drift.
[0118] In one possible implementation, sandbox training data can be constructed by fusing incremental effective data with the original sandbox training set (let's assume incremental data accounts for 20%-30%), preserving scenario diversity. For newly identified target drift scenarios, their feature data are extracted separately as key training samples (let's assume they account for 30% of the incremental training samples), supplementing the model with simulated variant data of the new drift (generated based on the differences between known drift features and the new drift), enhancing the model's generalization ability and avoiding overfitting.
[0119] The reward function was adjusted, and based on the weight configuration of the original sandbox training phase, the focus was on optimizing for the new target drift scenario, increasing the weight of ω2, let's assume it was increased to 0.4, while other weights were adjusted appropriately.
[0120] In one possible implementation, the attack drift correlation is used to quantify the strength of the correlation between the attack behavior and the target drift, and the attack drift correlation is determined by the following formula.
[0121] C = μ1C1 + μ2C2 + μ3C3 + μ4C4 (Formula 10) Wherein, C represents the attack drift correlation degree, C1 represents the attack intensity correlation coefficient, which is determined based on attack intensity and drift intensity and is used to measure the positive correlation between attack intensity and drift intensity, C2 represents the drift attribution coefficient, which is determined based on the proportion of attack-induced drift influence, the proportion of attack-induced drift, and the proportion of non-attack factors and is used to measure the proportion of target drift caused by attack behavior, C3 represents the attack trigger delay coefficient, which is determined based on attack trigger delay and delay threshold and is used to measure the temporal correlation between the occurrence of attack and the occurrence of target drift, C4 represents the attack type adaptation coefficient, which is a preset parameter and is used to measure the matching degree between attack tool type and drift type, and μ1, μ2, μ3, and μ4 represent weighting coefficients.
[0122] It should be noted that the attack drift correlation C is a core quantitative indicator that measures the strength of the correlation between "red team attack behavior" and "multi-agent target drift," with a value range of [0,1]. The closer the value is to 1, the more it indicates that the target drift is entirely induced by the attack behavior (the attack is the core cause of the drift); the closer the value is to 0, the more it indicates that the drift is unrelated to the attack (it may be caused by non-attack factors such as system failure or business fluctuations).
[0123] In one possible implementation, the attack strength correlation coefficient C1 is calculated as follows: C1 = E[(KE(K))(PE(P))] / std(K) / std(P), where E represents the average value, K represents the attack strength, P represents the drift strength, and std() represents the standard deviation function. The stronger the attack and the more severe the drift, the higher the coefficient, which reflects the intensity-driven effect of the attack on the drift.
[0124] In one possible implementation, the drift attribution coefficient C2 is used to measure the proportion of drift scenarios that can be clearly attributed to red team attacks, excluding interference from non-attack factors such as business fluctuations and system failures. The calculation method is: C2 = Attack-induced drift impact percentage / (Attack-induced percentage + Non-attack factor percentage). In business collaboration scenarios, it can be assumed that the non-attack factor percentage is 0.1 by default (when there are no clear non-attack factors). If there are clear non-attack factors (such as sudden business fluctuations), it can be adjusted as needed.
[0125] In one possible implementation, C3 represents the attack trigger delay coefficient, which measures the temporal correlation between the attack and the drift. The shorter the delay, the clearer the causal relationship between the attack and the drift, and the higher the coefficient. The calculation method is: C3 = max(1 - attack trigger delay / delay threshold, 0); let's assume the delay threshold is 3.
[0126] In one possible implementation, C4 represents the attack type fit coefficient, which measures the degree of match between the red team's attack type and the target drift type. The higher the fit, the greater the likelihood of attack-induced drift, and the higher the coefficient. Fit is applied according to three types of drift scenarios: semantic offset scenario C4=0.9, consensus degradation scenario C4=0.95, strategy shortcut scenario C4=0.85; and non-fit scenario C4≤0.3.
[0127] This invention employs a target drift risk protection and closed-loop optimization method using the PPO algorithm and sandbox collaboration. This method can be effectively extended to target drift security protection for other large-scale intelligent agents, significantly improving the efficiency and accuracy of data security defense processing for these agents. This invention can be used for target drift protection in various commercial and open-source large-scale intelligent agents, enhancing the overall target drift risk protection capability of large-scale intelligent agents. It is highly portable and has significant commercial value.
[0128] The target drift risk protection device provided by the present invention is described below. The target drift risk protection device described below and the target drift risk protection method described above can be referred to in correspondence.
[0129] Figure 6 This is a schematic diagram of the target drift risk protection device provided by the present invention, as shown below. Figure 6 As shown, the target drift risk protection device includes the following modules: sandbox module 601, training module 602, deployment module 603, and protection module 604; The sandbox module is used to simulate basic target drift scenarios in large-scale multi-agent systems based on the constructed sandbox. The training module is used to train the policy network and the value network to be trained based on training data in the basic target drift scenario, so as to obtain the trained policy network and the trained value network. The deployment module is used to deploy the trained policy network on the computing nodes of each agent in the large model multi-agent system, and to deploy the trained value network in the large model multi-agent system. The protection module is used to protect against target drift risk in large-scale multi-agent systems based on deployed trained policy networks and trained value networks.
[0130] According to the present invention, a target drift risk protection device is provided, wherein the training data includes: state data and motion data; and the training module is specifically used for: In the basic target drift scenario, state data and action data are input into the policy network to be trained to obtain the action probability distribution parameters output by the policy network to be trained. Input the state data and action data into the value network to be trained to obtain the state value parameters output by the value network to be trained. Based on the action probability distribution parameters, state value parameters, and reward function, the parameters of the policy network and the value network to be trained are updated to obtain the trained policy network and the trained value network.
[0131] According to the present invention, a target drift risk protection device includes state data comprising multi-dimensional features for characterizing a large-scale multi-agent system. These multi-dimensional features include at least one of the following: action sequence similarity, task completion deviation, abnormal action frequency, permission call compliance, inter-agent instruction response consistency, interaction protocol compliance rate, information transmission distortion, collaborative task division consistency, task priority matching degree, environmental constraint satisfaction degree, task timeout risk coefficient, drift type, drift intensity, drift rate, drift impact range, drift trigger source, attack tool type, attack method, attack intensity, attack chain stage, attack concealment, attack drift correlation, historical drift count, and historical protection effectiveness. Action data includes at least one of the following: context reset, target recalibration, behavioral intervention, agent restart, attack blocking, and no action. Each action in the action data has corresponding execution conditions and action priority.
[0132] According to the target drift risk protection device provided by the present invention, the training module is further used for: In the process of protecting against target drift risk, the current state data of each agent in the large model multi-agent system is obtained; Based on the trained policy network and the current state data, determine the optimal action data; Obtain the reward parameters for each agent after it performs the optimal action data, and update the model parameters of the trained value network based on the reward parameters.
[0133] According to the target drift risk protection device provided by the present invention, the basic target drift scenario includes at least one of the following: semantic drift scenario, coordination drift scenario, and behavioral drift scenario; the training module is further used for: In the process of protecting against target drift risk, the feature data of newly detected target drift scenarios are fed back to the sandbox to update the training data; Incremental training is performed on the trained policy network and the trained value network based on the updated training data.
[0134] According to the target drift risk protection device provided by the present invention, the training module is further used for: Determine the cosine similarity between the feature data of the detected target drift scene and the baseline feature data of the base target drift scene; If the cosine similarity is less than the preset similarity threshold, and the rate of change of drift intensity corresponding to multiple consecutive detection time windows is greater than the preset rate of change, and the rate of expansion of the drift influence range is greater than the preset expansion rate threshold, the detected target drift scene is determined to be a new target drift scene.
[0135] According to the target drift risk protection device provided by the present invention, the reward function is determined by the following formula: R=ω1R1+ω2R2+ω3R3+ω4R4+ω5R5; Where R represents the reward parameters determined by the reward function, R1 represents the attack blocking reward, which is determined based on attack intensity, attack chain stage, and attack concealment, R2 represents the drift mitigation reward, which is determined based on drift type, drift intensity, drift rate, and the adaptability of the action to the drift type, R3 represents the system stability reward, which is determined based on environmental constraint satisfaction, task timeout risk coefficient, and historical protection effect, R4 represents the protection efficiency reward, which is determined based on resource consumption cost, state change sequence, and action repetition rate, R5 represents the negative penalty item, which is determined based on invalid actions, excessive intervention, increased system risk, and violation of access rights, and ω1, ω2, ω3, ω4, and ω5 represent weighting coefficients.
[0136] According to the target drift risk protection device provided by the present invention, the attack drift correlation degree is used to quantify the correlation strength between attack behavior and target drift. The attack drift correlation degree is determined by the following formula: C = μ1C1 + μ2C2 + μ3C3 + μ4C4; Wherein, C represents the attack drift correlation degree, C1 represents the attack intensity correlation coefficient, which is determined based on attack intensity and drift intensity and is used to measure the positive correlation between attack intensity and drift intensity, C2 represents the drift attribution coefficient, which is determined based on the proportion of attack-induced drift influence, the proportion of attack-induced drift, and the proportion of non-attack factors and is used to measure the proportion of target drift caused by attack behavior, C3 represents the attack trigger delay coefficient, which is determined based on attack trigger delay and delay threshold and is used to measure the temporal correlation between the occurrence of attack and the occurrence of target drift, C4 represents the attack type adaptation coefficient, which is a preset parameter and is used to measure the matching degree between attack tool type and drift type, and μ1, μ2, μ3, and μ4 represent weighting coefficients.
[0137] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a target drift risk protection method, which includes: simulating a basic target drift scenario of a large-scale multi-agent system based on a constructed sandbox; training a policy network and a value network to be trained based on training data in the basic target drift scenario to obtain trained policy networks and trained value networks; deploying the trained policy network on the computing nodes of each agent in the large-scale multi-agent system, and deploying the trained value network on the large-scale multi-agent system; and protecting against target drift risk based on the large-scale multi-agent system with the trained policy network and trained value network deployed.
[0138] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0139] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the target drift risk protection method provided by the above methods. The method includes: simulating a basic target drift scenario of a large-scale multi-agent system based on a constructed sandbox; training a policy network and a value network to be trained based on training data in the basic target drift scenario to obtain a trained policy network and a trained value network; deploying the trained policy network on the computing nodes of each agent in the large-scale multi-agent system and deploying the trained value network in the large-scale multi-agent system; and protecting against target drift risk based on the large-scale multi-agent system with the trained policy network and trained value network deployed.
[0140] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the target drift risk protection method provided by the above methods. This method includes: simulating a basic target drift scenario of a large-scale multi-agent system based on a constructed sandbox; training a policy network and a value network to be trained based on training data in the basic target drift scenario to obtain a trained policy network and a trained value network; deploying the trained policy network on the computing nodes of each agent in the large-scale multi-agent system, and deploying the trained value network in the large-scale multi-agent system; and protecting against target drift risk based on the large-scale multi-agent system with the trained policy network and trained value network deployed.
[0141] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0142] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for preventing target drift risk, characterized in that, include: Based on the constructed sandbox simulation of a large-scale multi-agent system, the basic target drift scenario is simulated. In the basic target drift scenario, the policy network and the value network to be trained are trained based on the training data to obtain the trained policy network and the trained value network. The trained policy network is deployed on the computing node of each agent in the large model multi-agent system, and the trained value network is deployed in the large model multi-agent system. The large model multi-agent system, based on the deployed trained policy network and the trained value network, protects against target drift risk.
2. The target drift risk protection method according to claim 1, characterized in that, The training data includes: state data and action data; In the aforementioned basic target drift scenario, the policy network and value network to be trained are trained based on training data to obtain the trained policy network and trained value network, including: In the basic target drift scenario, the state data and the action data are input into the policy network to be trained to obtain the action probability distribution parameters output by the policy network to be trained. The state data and the action data are input into the value network to be trained to obtain the state value parameters output by the value network to be trained. Based on the action probability distribution parameters, the state value parameters, and the reward function, the parameters of the policy network to be trained and the value network to be trained are updated to obtain the trained policy network and the trained value network.
3. The target drift risk protection method according to claim 2, characterized in that, The state data includes multi-dimensional features used to characterize the large model multi-agent system. The multi-dimensional features include at least one of the following: action sequence similarity, task completion deviation, abnormal action frequency, permission call compliance, consistency of instruction response between agents, interaction protocol compliance rate, information transmission distortion, consistency of collaborative task division, task priority matching degree, environmental constraint satisfaction degree, task timeout risk coefficient, drift type, drift intensity, drift rate, drift impact range, drift trigger source, attack tool type, attack method, attack intensity, attack chain stage, attack concealment, attack drift correlation degree, historical drift count, and historical protection effect. The action data includes at least one of the following: context reset, target recalibration, behavior intervention, agent restart, attack blocking, and no action. Each action in the action data has corresponding execution conditions and action priority.
4. The target drift risk protection method according to any one of claims 1-3, characterized in that, The method further includes: In the process of protecting against target drift risk, the current state data of each agent in the large model multi-agent system is obtained; Based on the trained policy network and the current state data, determine the optimal action data; Obtain the reward parameters for each agent after executing the optimal action data, and update the model parameters of the trained value network based on the reward parameters.
5. The target drift risk protection method according to any one of claims 1-3, characterized in that, The basic target drift scenario includes at least one of the following: semantic drift scenario, coordination drift scenario, and behavioral drift scenario; the method further includes: In the process of protecting against target drift risk, the feature data of newly detected target drift scenarios are fed back to the sandbox to update the training data; Incremental training is performed on the trained policy network and the trained value network based on the updated training data.
6. The target drift risk protection method according to claim 5, characterized in that, The method further includes: Determine the cosine similarity between the feature data of the detected target drift scene and the baseline feature data of the base target drift scene; If the cosine similarity is less than a preset similarity threshold, and the drift intensity change rate is greater than a preset change rate and the drift influence range expansion speed is greater than a preset expansion speed threshold for multiple consecutive detection time windows, then the detected target drift scene is determined to be the new target drift scene.
7. The target drift risk protection method according to any one of claims 1-3, characterized in that, The reward function is determined by the following formula: R=ω1R1+ω2R2+ω3R3+ω4R4+ω5R5; Wherein, R represents the reward parameters determined by the reward function, R1 represents the attack blocking reward, which is determined based on attack intensity, attack chain stage, and attack concealment, R2 represents the drift mitigation reward, which is determined based on drift type, drift intensity, drift rate, and the adaptability of action and drift type, R3 represents the system stability reward, which is determined based on environmental constraint satisfaction, task timeout risk coefficient, and historical protection effect, R4 represents the protection efficiency reward, which is determined based on resource consumption cost, state change sequence, and action repetition rate, R5 represents the negative penalty item, which is determined based on invalid action, excessive intervention, increased system risk, and violation of calling permissions, and ω1, ω2, ω3, ω4, and ω5 represent weighting coefficients.
8. The target drift risk protection method according to any one of claims 1-3, characterized in that, Attack drift correlation is used to quantify the strength of the correlation between attack behavior and target drift. The attack drift correlation is determined by the following formula: C = μ1C1 + μ2C2 + μ3C3 + μ4C4; Wherein, C represents the attack drift correlation degree, C1 represents the attack intensity correlation coefficient, which is determined based on attack intensity and drift intensity and is used to measure the positive correlation between attack intensity and drift intensity, C2 represents the drift attribution coefficient, which is determined based on the proportion of attack-induced drift influence, the proportion of attack-induced drift, and the proportion of non-attack factors and is used to measure the proportion of target drift caused by attack behavior, C3 represents the attack trigger delay coefficient, which is determined based on attack trigger delay and delay threshold and is used to measure the temporal correlation between the occurrence of attack and the occurrence of target drift, C4 represents the attack type adaptation coefficient, which is a preset parameter and is used to measure the matching degree between attack tool type and drift type, and μ1, μ2, μ3, and μ4 represent weighting coefficients.
9. A target drift risk protection device, characterized in that, include: Sandbox module, training module, deployment module, and protection module; The sandbox module is used to simulate the basic target drift scenario of a large-scale multi-agent system based on the constructed sandbox. The training module is used to train the policy network and the value network to be trained based on training data in the basic target drift scenario, so as to obtain the trained policy network and the trained value network. The deployment module is used to deploy the trained policy network on the computing node of each agent in the large model multi-agent system, and to deploy the trained value network in the large model multi-agent system. The protection module is used to protect against target drift risk based on the large model multi-agent system that has deployed the trained policy network and the trained value network.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the target drift risk protection method as described in any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the target drift risk protection method as described in any one of claims 1 to 8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the target drift risk protection method as described in any one of claims 1 to 8.