Multi-agent cooperation system with double-layer credible guarantee

Through the multi-agent collaboration system with two-layer trustworthy guarantee, the multi-agent collaboration system solves the problems of malicious knowledge source identification and privacy leakage in the communication environment, achieving safe, reliable, private, secure and efficient collaboration effects.

CN120434004APending Publication Date: 2025-08-05GUANGDONG POLYTECHNIC NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510634373.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Multi-agent collaboration systems are difficult to effectively identify malicious knowledge sources in a confrontational communication environment, the risk of collaborative privacy leakage is high, and it is difficult to balance security and efficiency.

Method used

Design a multi-agent collaboration system with two-layer trustworthiness guarantees, including a trusted layer of knowledge source and a trusted layer of knowledge transmission. The trusted layer of knowledge source ensures the authenticity and integrity of the message source through the agent trust management, message identity authentication and malicious message detection modules; the trusted layer of knowledge transmission ensures privacy security and information accuracy through the privacy protection and collaborative message correction module.

Benefits of technology

It realizes effective identification and screening of malicious messages in a confrontation communication environment, reduces the penetration of forged and tampered messages, ensures information security and collaboration efficiency, and improves the security defense capabilities and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120434004A_ABST
    Figure CN120434004A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of agents, and discloses a double-layer credible guarantee multi-agent cooperation system, which comprises a knowledge source credible layer for dynamically evaluating the authenticity and integrity of the source of data transmitted by each agent, and comprises a knowledge source credible layer for dynamically evaluating the authenticity and integrity of the source of data transmitted by each agent based on message quality evaluation and a dynamic reputation updating mode, screening credible intelligent body co-authors; carrying out identity verification and data integrity verification on an intelligent agent transmission message; according to a constraint criterion based on decision consensus and uncertainty change, verifying the credibility of each collaborative message content, identifying malicious messages, and tracing and marking the identity of an intelligent agent; the knowledge transmission credible layer is used for guaranteeing privacy security and distortion compensation, and comprises the following steps: blocking a sensitive information leakage channel through noise injection and privacy budget control of intelligent agent observation data or coding characteristics based on a differential privacy mechanism; and carrying out local prediction and error correction on the cooperation message after noise disturbance by adopting a teammate modeling mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent agent technology, and in particular to a multi-agent collaboration system with double-layer trustworthiness assurance. Background Art

[0002] In recent years, with the rapid development of multi-agent reinforcement learning (MARL) technology and its widespread application in complex systems such as drone formations, intelligent connected vehicle scheduling, and distributed IoT control, collaboration based on inter-agent communication has become a key path to improving the distributed decision-making capabilities of these systems. Through efficient information exchange and joint learning between agents, we can effectively address challenges faced by individual agents, such as partial observability, environmental dynamics, and individual decision conflicts. This enables collaborative decision-making similar to "partial centralization," significantly improving overall task performance.

[0003] However, as multi-agent systems expand in scale, become more open, and communication scenarios become more complex, communication-dependent collaborative MARL faces increasingly severe security and privacy challenges, particularly in adversarial communication environments with interference, malicious attacks, or resource competition. Specifically, existing technologies face the following challenges:

[0004] 1. Untrustworthy knowledge sources and malicious injection: During the collaborative process, some agents may be controlled by attackers or engage in malicious behavior, intentionally sending forged, tampered, or poisoned information (such as false observation data, incorrect strategy recommendations, etc.). Existing security protection methods, such as blockchain-based trust mechanisms, dynamic reputation updates, or defenses against specific attack types (such as GAN-type stealth poisoning), can identify and isolate some malicious agents to a certain extent and block obvious attack messages. However, their main drawbacks are:

[0005] (1) Lack of ability to identify complex disguised attacks: For malicious messages that are concealed, intermittent, or highly similar to normal behavior, traditional reputation / detection methods based on thresholds or simple statistics are difficult to perform dynamic and reliable verification and identification.

[0006] (2) Failure to achieve dynamic verification and accurate tracking of the entire chain: Existing trust management is mostly concentrated at the node level, and fails to conduct dynamic strong verification at the content level of a single message, making it difficult to determine the authenticity of the message, and even more difficult to quickly and accurately trace the specific malicious message source when receiving an untrustworthy message.

[0007] (3) Insufficient integration with cryptographic technology: Most existing security mechanisms rarely integrate cryptographic methods such as message signatures into the collaborative communication process, which makes it difficult to strictly authenticate the identity of the intelligent agent and the data sent cannot prove its originality and non-tamperability, providing an opportunity for forged identities to send malicious messages.

[0008] 2. Risk of leakage of privacy-sensitive information: Multi-agent collaboration usually requires agents to share sensitive data such as their local observation states, strategy information, and belief distribution. In an adversarial environment, an adversary may be able to infer the key privacy states of the agents (for example, the precise location of the drone, flight intentions, type of observed targets, etc.) by eavesdropping and analyzing these communication messages, leading to significant data security and privacy leakage risks. Existing privacy protection methods, such as differential privacy (DP), homomorphic encryption (HE), or federated learning (FL), can provide a certain degree of information obfuscation or encryption, but their main drawbacks are:

[0009] (1) Noise perturbation significantly impairs collaborative efficiency: To meet strict privacy requirements (such as ε-delta in differential privacy), it is usually necessary to inject high-intensity noise into the transmitted data. This excessive noise perturbation can severely distort the usability of the original message, making it difficult for the receiver to accurately understand and utilize collaborative information, significantly reducing the performance and efficiency of collaborative learning, and even causing the algorithm to become unstable or divergent.

[0010] (2) It is difficult to effectively deal with poisoning attacks disguised as noise: Malicious agents may exploit the characteristics of the differential privacy mechanism itself (i.e., noise is its core component) and deliberately inject noise that appears random but has a specific attack purpose, disguised as legitimate privacy perturbations, to carry out covert poisoning attacks. Existing privacy protection mechanisms often have difficulty distinguishing between legitimate privacy noise and malicious poisoning noise.

[0011] (3) Efficiency issues of homomorphic encryption: Although homomorphic encryption can perform calculations in the ciphertext domain and process sensitive data without decryption, it usually requires decryption and re-encryption when processing nonlinear operations (such as activation functions in neural networks), which introduces significant computational overhead and delay, making it difficult to meet the real-time requirements of multi-agent collaboration scenarios.

[0012] 3. It is difficult to balance security and privacy: Most existing research and technical solutions tend to address security threats (such as malicious message detection and trust management) or privacy leaks (such as differential privacy noise and data encryption) in isolation. Even if there are individual research attempts to integrate the two, they often remain at a relatively preliminary and weakly coupled stage, only able to achieve simple mitigation of malicious influences, but unable to provide a systematic, end-to-end full-link trusted security guarantee. The fundamental reasons for this separation or weak integration are:

[0013] (1) Lack of a unified system architecture design: There is no overall framework that can organically integrate message source identification, identity authentication, content verification, privacy protection, and distortion correction. Existing methods usually only address one point or one link in "bridging security" (such as node trust) or "bridging privacy" (such as end-to-end encryption / noising), and lack the ability to simultaneously achieve the dual goals of "trusted knowledge source" and "trusted knowledge transmission" from an architectural perspective.

[0014] (2) Lack of intelligent processing capabilities under privacy disturbances: Existing mechanisms are unable to accurately determine the true intent or content of a message after the data has been disturbed by means such as differential privacy, and are even more unable to intelligently identify the cause of the distortion caused by the disturbance and make dynamic corrections. As a result, when applied to safety-critical scenarios such as drones in adversarial communication environments, the collaborative efficiency, system robustness, information security, and privacy protection cannot meet actual needs.

[0015] In summary, the core technical bottlenecks of current multi-agent collaboration in adversarial communication environments are the difficulty in effectively identifying malicious knowledge sources, the high risk of collaborative privacy leakage, and the difficulty in balancing security and efficiency. This field urgently needs a new multi-agent collaboration mechanism that can provide dual trustworthy guarantees for knowledge sources and knowledge transmission at the architectural level, while maintaining high collaboration efficiency and robustness in privacy-perturbing environments. Summary of the Invention

[0016] The purpose of this invention is to design a multi-agent collaboration system with double-layer trust assurance, which solves the problems of difficulty in effectively identifying malicious knowledge sources, high risk of collaborative privacy leakage, and difficulty in balancing security and efficiency.

[0017] The present invention provides a multi-agent collaboration system with dual-layer trustworthiness assurance, comprising:

[0018] The knowledge source trust layer is used to dynamically evaluate the authenticity and integrity of the data transmitted by each agent during the multi-agent collaboration process. It includes:

[0019] The agent trust management module selects trustworthy agent collaborators based on message quality assessment and dynamic reputation update;

[0020] The message identity authentication module uses elliptic curve encryption and hash signature to authenticate the identity of the agent's transmitted messages and verify the data integrity;

[0021] The malicious message detection and labeling module verifies the authenticity and credibility of the content of each collaborative message based on the constraint criteria based on decision consensus and uncertainty changes, identifies malicious messages and traces the identity of the intelligent agent;

[0022] The knowledge transmission trust layer is used to ensure privacy and distortion compensation during the message transmission process of intelligent agents. It includes:

[0023] The privacy protection processing module, based on the differential privacy mechanism, blocks the leakage of sensitive information by injecting noise into the agent's observation data or encoding features and controlling the privacy budget;

[0024] The collaborative message correction module uses teammate modeling to perform local intelligent prediction and error correction on collaborative messages after noise disturbance.

[0025] In the above solution, the core function of the knowledge source trust layer is to dynamically and reliably verify the authenticity and integrity of collaborative messages from various intelligent entities, ensuring that collaborative information originates from trusted entities and is not forged or tampered with, thereby fundamentally preventing malicious information from disrupting collaborative decision-making. This layer includes the following three main technical modules:

[0026] The agent trust management module establishes and adjusts trust in agents in real time through dynamic message quality assessment and reputation update strategies. Message quality assessment quantifies the effectiveness of messages by comparing the change in the recipient's decision uncertainty before and after a collaborative message is sent. Based on this, a message quality score is dynamically calculated. Furthermore, a dynamic reputation adjustment function is used to update each agent's reputation using a time-weighted cumulative evaluation of historical messages. High credibility and high-quality messages are combined to form an overall agent score. Agents below a preset threshold are excluded from the collaborative team, ensuring the reliability and stability of the collaborative entity.

[0027] The message authentication module utilizes the Edwards25519 elliptic curve cryptography algorithm and the SHA-512 hash function to establish a digital signature mechanism, ensuring that all collaborative messages are accompanied by an unforgeable signature from the agent. Identity authentication involves signing the message hash with the agent's private key at the message generator, and the receiving end can quickly verify the message using its public key, ensuring message integrity, consistency, and non-repudiation of the sender's identity. This module's design balances security and system overhead, supporting both single and batch signature verification, making it suitable for resource-constrained, large-scale multi-agent systems.

[0028] The malicious message detection and labeling module analyzes the impact of each collaborative message on the overall collaborative decision-making process, targeting potentially tampered or interfering information sent by malicious agents. Using a dual constraint based on decision consensus and changes in system uncertainty, the module analyzes the impact of each message on the overall collaborative decision-making process. For example, if a message causes a decrease in confidence in the optimal strategy or an increase in global system uncertainty, the message is deemed malicious. Combined with the message signature, the source agent of the malicious message is labeled and its reputation downgraded. This approach leverages the semantic impact of messages for security detection, breaking through the bottleneck of traditional detection based on thresholds and simple statistics, and possesses powerful capabilities against complex camouflage attacks.

[0029] When an intelligent agent sends a message, the message identity authentication module first ensures the authenticity of the identity and the integrity of the message; the receiver confirms the credibility of the message through the malicious message detection and marking module, combined with the measurement of decision uncertainty, and feeds back the conclusion to the intelligent agent trust management module to dynamically adjust the trust evaluation, forming a closed-loop feedback mechanism; at the same time, the list of trusted intelligent agents output by the trust management directly guides the screening of system collaborative entities, ensuring a high-credibility source of data in subsequent interaction processes.

[0030] The knowledge transfer trust layer focuses on protecting the privacy of collaborative messages between agents and accurately restoring messages in non-ideal channel environments such as noise disturbances, ensuring privacy security and collaboration efficiency among members. It includes two core modules:

[0031] The privacy-preserving processing module targets sensitive observation data and encoding features during the collaborative process. It employs a differential privacy mechanism by dynamically injecting Gaussian noise of controlled intensity into the evidence encoding end. This mechanism also imposes strict constraints on hidden state sensitivity (e.g., spectral norm normalization and gradient clipping). This full-process privacy budget control ensures mathematically quantified privacy protection. This module effectively blocks the direct mapping between sensitive data and transmitted messages, making it difficult for external observers to infer the privacy state of the agent by analyzing the communication content, thus ensuring data confidentiality during transmission.

[0032] The collaborative message correction module addresses the noise perturbations introduced by privacy protection and message distortion caused by external environmental noise. Based on teammate modeling technology, each agent locally constructs a dynamic behavior model of its neighboring agents (based on historical observations and identity identifiers), predicts the distribution characteristics of their original noise-free messages, and fuses this prediction with the received noisy messages. Using statistical methods such as weighted averaging, it performs error correction to restore more accurate message characteristics. This process fully leverages the collaborative nature of agents to intelligently correct distorted signals, significantly improving the system's robustness and collaborative performance in adversarial communication environments.

[0033] These two modules work together to achieve an organic combination of privacy protection and message distortion correction: the privacy protection processing module first blurs sensitive information to ensure information security; the collaborative message correction module then uses the built-in teammate prediction model to accurately correct the information bias caused by disturbances, so that the noise protection mechanism and collaborative accuracy are balanced, overcoming the irreconcilable contradiction between privacy and efficiency in traditional privacy protection mechanisms.

[0034] Preferably, in the agent trust management module, screening of trustworthy agent collaborators based on message quality evaluation and dynamic reputation update includes:

[0035] The quality of message transmission from agent i to agent j is defined as:

[0036]

[0037] Among them, u j represents the decision uncertainty of receiver j before communication, is the uncertainty remaining after communication; if v ij >0, indicating that the message effectively reduces the receiver's decision ambiguity; otherwise, it is considered low-quality or malicious information;

[0038] The time decay function is used to weight the quality of historical messages. The quality of the message transmitted by agent i at time t is: Its historical message quality V i Expressed as:

[0039]

[0040] Among them, γ∈(0,1) is the attenuation factor, which gives higher weight to recent communications, and T is the evaluation window length;

[0041] Introducing reputation value l i ∈[0,1] quantifies the credibility of agent i. Initially, all agents are set to have a neutral credibility l i =0.5. After each round of collaboration, the following rules are used to adjust the message based on whether it contains malicious content or significantly deviates from the true state:

[0042]

[0043] where Δ l is the reputation increment parameter;

[0044] The comprehensive score of agent i is defined as the product of message quality and reputation:

[0045] ε i =V i ·l i

[0046] Setting the threshold When not satisfied When , agent i is not selected as the cooperation partner of j, and the message transmitted by i to j is filtered.

[0047] Preferably, in the message identity authentication module, using elliptic curve encryption and hash signature to authenticate the identity and verify the data integrity of the agent transmission message includes:

[0048] Randomly generate a 256-bit random number d k , calculate its SHA-512 hash value and take the first 256 bits as the private key sk k ; Using the private key sk k With the elliptic curve base point G, generate the public key:

[0049] pk k =sk k ·G

[0050] Agent public binding ID k The public key pk k , private key is stored locally;

[0051] Based on the private key sk k and the hash value SHA-512(sk k ||m), generate a temporary random number r k mod l; by R k =r k G generates a temporary public key; for the temporary public key R k 、Public key pk k Hash with message m to get

[0052] e k =SHA-512(R k ||pk k ||m)mod l

[0053] Calculate signature s k =(r k +e k ·sk k ) mod l, and finally output the signature pair σ k =(R k ,s k ), the length is fixed at 64 bytes;

[0054] In the single signature verification phase, the verifier receives data m and signature σ k and public key pk k Afterwards, according to R k 、pk k 、m recalculate e k :

[0055]

[0056] Equality verification check, if it is true, the signature is valid.

[0057] Preferably, the message identity authentication module further includes:

[0058] Introduce a random number α for each signature k , to prevent attackers from bypassing verification by combining invalid signatures; k 、R k and e k ·pk k Linearly combine them to generate the aggregate value S agg 、R agg and Pagg ;verify:

[0059]

[0060] Complete several signature verifications at once, reducing the computational complexity from O(K) to O(1).

[0061] Preferably, in the malicious message detection and marking module, verifying the authenticity and credibility of the content of each collaborative message based on the constraint principle based on decision consensus and uncertainty change, identifying malicious messages and tracing the identity of the intelligent agent includes:

[0062] Malicious message characteristics are mathematically expressed as follows: If a message e i A message that violates any of the following cooperation principles during the integration process is defined as a malicious message:

[0063] Forward consensus principle: optimal action a * The combined confidence quality of coop (a * ,k) i (a * ,k);

[0064] Uncertainty reduction principle: global uncertainty of the system u coop Rise (u coop >u i ).

[0065] Preferably, message verification is performed based on the principles of positive consensus and uncertainty reduction, including:

[0066] Give each agent an initial reputation score, which can be the same initial value;

[0067] For each agent:

[0068] Calculate a comprehensive score based on the quality of the agent's past messages and its current reputation score;

[0069] Select agents with good credit and worthy of cooperation based on the scores and send them collaboration requests;

[0070] For each message received:

[0071] Verify the digital signature of the message to ensure that the message is indeed from the claimed intelligent agent and has not been tampered with;

[0072] Use the principles of positive consensus and uncertainty reduction to verify whether the message is true and credible;

[0073] Adjust the reputation value of the agent that sent the message based on the verification result:

[0074] ​If the message is reliable, the reputation value of the agent is increased;

[0075] If there is a problem with the message, reduce the reputation value;

[0076] Calculate the quality index of the message and update the historical message quality record of the agent;

[0077] Based on the received and verified collaborative knowledge, the agent selects actions and adjusts strategies based on credible information to improve the overall system performance;

[0078] Continuously select and accept collaboration requests based on reputation scores, complete the signing, sending and verification of messages, forming a dynamic adjustment and optimization cycle.

[0079] Preferably, in the privacy protection processing module, based on the differential privacy mechanism, by injecting noise into the agent's observation data or encoding features and controlling the privacy budget, blocking the sensitive information leakage channel includes:

[0080] The Gaussian noise mechanism is introduced, which is specifically defined as: Let the hidden state of the agent be

[0081] The output of the evidence encoder is:

[0082] e=ReLU(Wh+b)+η,

[0083] Where η~N(0,σ 2 I) is independent Gaussian noise, W and b are weight matrix and bias term respectively;

[0084] By constraining the sensitivity of the hidden state Δ h , that is, the maximum L2-norm change of adjacent hidden states; the spectral norm of the weight matrix ||W||2, calculates the sensitivity of the evidence encoding function:

[0085] Δ f =||W||2·Δ h ,

[0086] Ensure that the noise parameter σ satisfies Satisfies (∈,δ)-differential privacy;

[0087] In the implementation, gradient clipping is used to limit the update amplitude of the hidden state, and the spectral norm of the weight matrix is normalized W←W / ||W||2·c to further control the sensitivity;

[0088] For multiple rounds of training, the strong combination theorem is used to calculate the cumulative privacy loss:

[0089] and δ=Tδ0+δ′,

[0090] Ensure the overall privacy budget is under control.

[0091] Preferably, the privacy protection processing module further includes:

[0092] Input agent set U and message set M;

[0093] All agents build corresponding teammate models based on historical observations and teammate identity information;

[0094] For each agent u i :

[0095] For all messages received Perform the following processing:

[0096] c. Generate an estimated value for the message;

[0097] d. Correcting the received message based on the above estimate to improve the reliability of the message;

[0098] c. By minimizing the cross entropy between the estimated message and the true message, the estimation model parameters are optimized to improve the estimation accuracy;

[0099] Based on the corrected collaboration message, the agent performs action selection and strategy update;

[0100] When the agent receives a collaboration request, it fuzzifies its local observation information based on evidence encoding and then injects differential privacy noise to achieve privacy protection.

[0101] The agent sends a privacy-preserving message Complete privacy-protected message sharing.

[0102] Preferably, in the collaborative message correction module, using a teammate modeling method to perform local prediction and error correction on the collaborative message after noise disturbance includes:

[0103] Each agent i has a local observation history τ i and teammate identification d j Build a teammate model to infer the transmission message of teammate j. The specific process is as follows:

[0104] Input local observation history τ i , through the gated recurrent unit to i Encoded as hidden state

[0105] Will with d j Input the multi-layer perceptron to generate fusion features h fuse :

[0106]

[0107] Further generate the estimated confidence quality through linear mapping and normalization and uncertainty quality

[0108] in

[0109] Minimize the cross entropy between estimated evidence and true evidence to optimize the estimated model:

[0110]

[0111] Agent i receives a noisy message from teammate j;

[0112] The correction process dynamically integrates the received message with the teammate model estimate:

[0113]

[0114] in:

[0115]

[0116] The error variance of the teammate model estimation is calculated by L model optimization;

[0117]

[0118] The total noise variance reflects the interference intensity of the communication channel.

[0119] In the above scheme, trusted knowledge sources refer to ensuring the credibility of knowledge during the process of collaborative knowledge transfer between agents to achieve collaborative goals. Previous research has mostly used technologies such as malicious message detection and reputation mechanisms for countermeasures. Trusted knowledge transfer refers to ensuring that private data is not leaked during the process of collaborative knowledge transfer between agents. Key technologies used include differential privacy, homomorphic encryption, and federated learning. It is well known that DP noise can hinder the original learning process, leading to unstable or even divergent algorithms, especially for deep learning-based methods.

[0120] Compared with the prior art, the present invention has the following beneficial effects:

[0121] The present invention discloses a multi-agent collaboration system with dual-layer trustworthiness assurance. The knowledge source trustworthiness layer ensures the credibility and integrity of the message source, effectively identifies and screens reliable agents and their messages, and builds a trustworthy knowledge base. The knowledge transmission trustworthiness layer ensures the security of sensitive information through privacy protection and disturbance correction during message transmission, while also improving the impact of information distortion and ensuring the efficiency and accuracy of collaboration. The ultimate goal is to achieve safe, reliable, privacy-safe, robust and efficient operation of the multi-agent collaboration system, thereby meeting practical application scenarios with high security requirements, such as drone collaboration and intelligent transportation. Specifically:

[0122] (1) Adopting dynamic reputation management and message signing mechanisms to effectively screen out malicious or low-credibility agents, reduce the possibility of forged and tampered message penetration, and fundamentally improve the authenticity and integrity of messages;

[0123] (2) Based on the content verification rules of decision-making consensus and uncertainty changes, timely detect and mark untrustworthy message sources, prevent malicious information from disrupting the system at the source, and enhance the system's security defense capabilities;

[0124] (3) Through elliptic curve digital signatures, we can ensure that the source of the message is legitimate and the content has not been tampered with, thereby improving the reliability of identity authentication and strengthening the system's identity authentication capabilities;

[0125] (4) Introducing a differential privacy mechanism to perturb sensitive features, blocking the leakage channel of sensitive information, and providing mathematical privacy protection for multi-agent systems during information sharing;

[0126] (5) Through teammate modeling, message prediction and error correction, the information distortion caused by disturbances can be effectively alleviated, and the cooperation efficiency and decision-making quality of multi-agent systems in noisy disturbance environments can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0127] Figure 1 This is a schematic diagram of a multi-agent collaboration system module with double-layer trust assurance provided by an embodiment of the present invention;

[0128] Figure 2 This is a model diagram of a drone-assisted communication system based on multi-agent reinforcement learning provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0129] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0130] like Figure 1 As shown, this application provides a multi-agent collaboration system with dual-layer trust assurance, including:

[0131] The knowledge source trust layer is used to dynamically evaluate the authenticity and integrity of the data transmitted by each agent during the multi-agent collaboration process. It includes:

[0132] The agent trust management module selects trustworthy agent collaborators based on message quality assessment and dynamic reputation update;

[0133] The message identity authentication module uses elliptic curve encryption and hash signature to authenticate the identity of the agent's transmitted messages and verify the data integrity;

[0134] The malicious message detection and labeling module verifies the authenticity and credibility of the content of each collaborative message based on the constraint criteria based on decision consensus and uncertainty changes, identifies malicious messages and traces the identity of the intelligent agent;

[0135] The knowledge transmission trust layer is used to ensure privacy and distortion compensation during the message transmission process of intelligent agents. It includes:

[0136] The privacy protection processing module, based on the differential privacy mechanism, blocks the leakage of sensitive information by injecting noise into the agent's observation data or encoding features and controlling the privacy budget;

[0137] The collaborative message correction module uses teammate modeling to perform local intelligent prediction and error correction on collaborative messages after noise disturbance.

[0138] In one embodiment provided in this application, a UAV-assisted communication scenario is considered, where multiple UAVs U = {u1, ..., u N} As an airborne base station, it provides network services to ground targets (such as vehicles and users) in a collaborative manner. Figure 2 As shown, the base station and RSU are fixed, and their positions and coverage radius are known. The drone is equipped with a signal transceiver, and its coverage radius is Depends on the current flight altitude h u The drone controller works in conjunction with the base station to determine the drone's location based on the vehicle, user location, and drone status. Assuming that the drone has limited energy and needs to return to the charging station regularly for charging. Assuming that there is a drone charging station (CS) in the area, its location l charge Given: The vehicle locations are grouped at the cell level, and G represents all cells in the area.

[0139] A. Channel Model:

[0140] The channel model describes the signal transmission characteristics between infrastructure nodes (such as base stations, UAVs, or roadside units) and vehicles. The signal-to-noise ratio (SNR) is calculated as follows:

[0141] Ground-to-Ground (G2G) Channel:

[0142]

[0143] d i,v : Euclidean distance between node i and vehicle v. ζ: Path loss exponent, usually between 2 and 6. X0: Path loss at reference distance.

[0144] Air-to-Ground (A2G) Channel:

[0145]

[0146] ζ LoS ,ζ NLoS : are the path loss exponents of LoS and NLoS links respectively. are the shadow effects of LoS and NLoS links, respectively, modeled as zero-mean Gaussian random variables.

[0147] B. Communication Coverage Model

[0148] The communication coverage model defines the coverage areas of base stations, roadside units, and UAVs.

[0149] Coverage radius of base stations and roadside units:

[0150] Base station coverage radius R b and the coverage radius R of the roadside unit r is fixed.

[0151] UAV coverage radius:

[0152]

[0153] d u : The maximum distance between the UAV and the ground, defined as the path loss below a certain threshold The distance at h u : The flight altitude of the UAV.

[0154] Association between vehicles and infrastructure nodes:

[0155]

[0156] e v,i : Boolean indicator variable indicating whether vehicle v is within the coverage of infrastructure node i.

[0157] C. UAV energy consumption model

[0158] The energy consumption model of the drone takes into account the energy consumption during flight and hovering. The total energy consumption is calculated as follows:

[0159] Energy consumption per unit time:

[0160]

[0161] P0, P1, P2: coefficients of blade profile power, induced power, and parasitic power, respectively. υ0, υ1: rotor blade tip speed and average rotor induced speed in hover, respectively.

[0162] Total energy consumption:

[0163]

[0164] P c : Communication related power, in watts (W). V t : The set of vehicles at time t.

[0165] D. Vehicle Spatiotemporal Distribution Model

[0166] The spatiotemporal distribution model of vehicles describes the distribution characteristics of vehicles in space and time. The main indicators include:

[0167] Spatial Variation (SV):

[0168]

[0169] D: The set of vehicle counts in all cells. The standard deviation of the number of vehicles. μ(D): The mean of the number of vehicles.

[0170] Time change (TV):

[0171] Time Coefficient Variation (TCV):

[0172] ξ g : The vehicle proportion vector of cell g in the time period.

[0173] Step coefficient variation (SCV):

[0174] The scenario is modeled as a fully cooperative multi-agent system, where each drone makes action decisions based on its local observations (such as the location of the target and other drones) to maximize the joint reward (such as the number of targets covered). The fully cooperative multi-agent system is modeled as a decentralized partially observable Markov decision process G = {S,O i ,A i ,P,r i ,π i}, agents complete learning tasks such as target tracking and collaborative perception through mutual communication and collaboration. At each time step t, the learning agent i observes the environment to obtain local observations And based on current observations Strategy distribution Select task action Then receive rewards from the environment related to task performance (such as tracking accuracy) The environment is distributed through state transition P∈[0,1] |S| Transfer to new state s (t+1) ∈S. The goal of learning agent i is to optimize the policy π i , to maximize the expected long-term discounted reward The discount factor γ represents the relative importance of immediate rewards and future rewards.

[0175] state:

[0176] s t ={D t ,L t ,E t}

[0177] D t : The distribution of vehicles at time step t, expressed as the mean vector of the number of vehicles in all cells. L t ={l u,t |∈U}: The position vector of the UAV at time step t. E t ={E u,t |∈U}: The residual energy vector of the UAV at time step t.

[0178] action:

[0179] a t ={w u1,t ,…,w uN,t}

[0180] The normalized 3D coordinates of UAV u at time step t. The next position of UAV is calculated as:

[0181]

[0182] φ u,t : Boolean parameter indicating whether the UAVu is in service state (1) or charging state (0)

[0183] Reward function:

[0184]

[0185] The fairness value at time step t is defined as:

[0186]

[0187] X i,t : The vehicle access rate vector of the i-th arriving UAV service at time step t. ΔT i,t : The time interval for the i-th UAV to provide service after arriving at the target location. |V t |: The number of vehicles at time step t. The original fairness value when there is no UAV working.

[0188] Preferably, in the agent trust management module, screening of trustworthy agent collaborators based on message quality evaluation and dynamic reputation update includes:

[0189] The quality of message transmission from agent i to agent j is defined as:

[0190]

[0191] Among them, u j represents the decision uncertainty of receiver j before communication, is the uncertainty remaining after communication; if v ij >0, indicating that the message effectively reduces the receiver's decision ambiguity; otherwise, it is considered low-quality or malicious information;

[0192] The time decay function is used to weight the quality of historical messages. The quality of the message transmitted by agent i at time t is: Its historical message quality V i Expressed as:

[0193]

[0194] Among them, γ∈(0,1) is the attenuation factor, which gives higher weight to recent communications, and T is the evaluation window length;

[0195] Introducing reputation value l i ∈[0,1] quantifies the credibility of agent i. Initially, all agents are set to have a neutral credibility l i =0.5. After each round of collaboration, the following rules are used to adjust the message based on whether it contains malicious content or significantly deviates from the true state:

[0196]

[0197] where Δ l is the reputation increment parameter;

[0198] The comprehensive score of agent i is defined as the product of message quality and reputation:

[0199] εi =V i ·l i

[0200] Setting the threshold When not satisfied When , agent i is not selected as the cooperation partner of j, and the message transmitted by i to j is filtered.

[0201] In the above scheme, in the fully cooperative multi-agent reinforcement learning framework, the selection of trusted cooperation partners is a key link in ensuring the efficiency and security of system collaboration.

[0202] Preferably, in the message identity authentication module, using elliptic curve encryption and hash signature to authenticate the identity and verify the data integrity of the agent transmission message includes:

[0203] Randomly generate a 256-bit random number d k , calculate its SHA-512 hash value and take the first 256 bits as the private key sk k ; Using the private key sk k With the elliptic curve base point G, generate the public key:

[0204] pk k =sk k ·G

[0205] Agent public binding ID k The public key pk k , private key is stored locally;

[0206] Based on the private key sk k and the hash value SHA-512(sk k ||m), generate a temporary random number r k mod l; by R k =r k G generates a temporary public key; for the temporary public key R k 、Public key pk k Hash with message m to get

[0207] e k =SHA-512(R k ||pk k ||m)mod l

[0208] Calculate signature s k =(r k +e k ·sk k ) mod l, and finally output the signature pair σ k =(R k ,s k ), the length is fixed at 64 bytes;

[0209] In the single signature verification phase, the verifier receives data m and signature σ k and public key pk k Afterwards, according to R k 、pk k 、m recalculate e k :

[0210]

[0211] Equality verification check, if it is true, the signature is valid.

[0212] In this solution, a message signing mechanism is designed to ensure that received messages originate from collaborative agents, rather than data tampered or forged by malicious attackers. First, the Edwards25519 elliptic curve is used as the cryptographic foundation. This curve is the standard choice for the EdDSA algorithm, combining efficiency and security. SHA-512 is also specified as the hash function to ensure data integrity. Global parameters (such as the base point G and the curve equation) are uniformly published to ensure that all agents use the same standard.

[0213] Preferably, the message identity authentication module further includes:

[0214] Introduce a random number α for each signature k , to prevent attackers from bypassing verification by combining invalid signatures; k 、R k and e k ·pk k Linearly combine them to generate the aggregate value S agg 、R agg and P agg ;verify:

[0215]

[0216] Complete several signature verifications at once, reducing the computational complexity from O(K) to O(1).

[0217] Preferably, in the malicious message detection and marking module, verifying the authenticity and credibility of the content of each collaborative message based on the constraint principle based on decision consensus and uncertainty change, identifying malicious messages and tracing the identity of the intelligent agent includes:

[0218] Malicious message characteristics are mathematically expressed as follows: If a message e i A message that violates any of the following cooperation principles during the integration process is defined as a malicious message:

[0219] Forward consensus principle: optimal action a * The combined confidence quality of coop (a* ,k) i (a * ,k);

[0220] Uncertainty reduction principle: global uncertainty of the system u coop Rise (u coop >u i ).

[0221] Preferably, message verification is performed based on the principles of positive consensus and uncertainty reduction, including:

[0222] Give each agent an initial reputation score, which can be the same initial value;

[0223] For each agent:

[0224] Calculate a comprehensive score based on the quality of the agent's past messages and its current reputation score;

[0225] Select agents with good credit and worthy of cooperation based on the scores and send them collaboration requests;

[0226] For each message received:

[0227] Verify the digital signature of the message to ensure that the message is indeed from the claimed intelligent agent and has not been tampered with;

[0228] Use the principles of positive consensus and uncertainty reduction to verify whether the message is true and credible;

[0229] Adjust the reputation value of the agent that sent the message based on the verification result:

[0230] If the message is reliable, the reputation value of the agent is increased;

[0231] If there is a problem with the message, reduce the reputation value;

[0232] Calculate the quality index of the message and update the historical message quality record of the agent;

[0233] Based on the received and verified collaborative knowledge, the agent selects actions and adjusts strategies based on credible information to improve the overall system performance;

[0234] Continuously select and accept collaboration requests based on reputation scores, complete the signing, sending and verification of messages, forming a dynamic adjustment and optimization cycle.

[0235] ​In the above scenario, the attack may originate from an insider threat (i.e., an attacker controls some vulnerable agents). To address this issue, messages must be verified for malicious intent and the source of messages identified as malicious must be labeled as malicious agents. Regarding observational verification encoded as evidence: Intuitively, promoting agent collaboration through message passing can be quantified as an increase in the belief value of an action and a decrease in decision uncertainty. Malicious messages, on the other hand, can cause actions with low belief values to have higher belief values or increase decision uncertainty.

[0236] Preferably, in the privacy protection processing module, based on the differential privacy mechanism, by injecting noise into the agent's observation data or encoding features and controlling the privacy budget, blocking the sensitive information leakage channel includes:

[0237] The Gaussian noise mechanism is introduced, which is specifically defined as: Let the hidden state of the agent be

[0238] The output of the evidence encoder is:

[0239] e=ReLU(Wh+b)+η,

[0240] Where η~N(0,σ 2 I) is independent Gaussian noise, W and b are weight matrix and bias term respectively;

[0241] By constraining the sensitivity of the hidden state Δ h , that is, the maximum L2-norm change of adjacent hidden states; the spectral norm of the weight matrix ||W||2, calculates the sensitivity of the evidence encoding function:

[0242] Δ f =||W||2·Δ h ,

[0243] Ensure that the noise parameter σ satisfies Satisfies (∈,δ)-differential privacy;

[0244] In the implementation, gradient clipping is used to limit the update amplitude of the hidden state, and the spectral norm of the weight matrix is normalized W←W / ||W||2·c to further control the sensitivity;

[0245] For multiple rounds of training, the strong combination theorem is used to calculate the cumulative privacy loss:

[0246] and δ=Tδ0+δ′,

[0247] Ensure the overall privacy budget is under control.

[0248] Preferably, the privacy protection processing module further includes:

[0249] Input agent set U and message set M;

[0250] All agents build corresponding teammate models based on historical observations and teammate identity information;

[0251] For each agent u i :

[0252] For all messages received Perform the following processing:

[0253] e. Generate an estimated value for the message;

[0254] f. Correcting the received message based on the above estimate to improve the reliability of the message;

[0255] c. By minimizing the cross entropy between the estimated message and the true message, the estimation model parameters are optimized to improve the estimation accuracy;

[0256] Based on the corrected collaboration message, the agent performs action selection and strategy update;

[0257] When the agent receives a collaboration request, it fuzzifies its local observation information based on evidence encoding and then injects differential privacy noise to achieve privacy protection.

[0258] The agent sends a privacy-preserving message Complete privacy-protected message sharing.

[0259] In the aforementioned scheme, one of the core challenges in multi-agent collaborative tasks is how to efficiently transmit information while protecting individual privacy. As the communication content, the evidence vector contains richer observation information than traditional direct encoding, which can improve the efficiency of communication between multi-agents. In the communication-based multi-agent reinforcement learning collaboration process, malicious agents or attackers can reversely infer sensitive state information by intercepting and analyzing the communication data (such as action strategies and observation information) between agents. To enhance privacy protection, a Gaussian noise mechanism is introduced in the evidence encoding module.

[0260] Preferably, in the collaborative message correction module, using a teammate modeling method to perform local prediction and error correction on the collaborative message after noise disturbance includes:

[0261] Each agent i has a local observation history τ i and teammate identification d j Build a teammate model to infer the transmission message of teammate j. The specific process is as follows:

[0262] Input local observation history τ i , through the gated recurrent unit to i Encoded as hidden state

[0263] Will with d j Input the multi-layer perceptron to generate fusion features h fuse :

[0264]

[0265] Further generate the estimated confidence quality through linear mapping and normalization and uncertainty quality

[0266] in

[0267] Minimize the cross entropy between estimated evidence and true evidence to optimize the estimated model:

[0268]

[0269] Agent i receives a noisy message from teammate j;

[0270] The correction process dynamically integrates the received message with the teammate model estimate:

[0271]

[0272] in:

[0273]

[0274] The error variance of the teammate model estimation is calculated by L model optimization;

[0275]

[0276] The total noise variance reflects the interference intensity of the communication channel.

[0277] In the above scheme, noise can distort the sent evidence, leading to inaccuracies during fusion at the receiving end. For example, differential privacy noise can cause the confidence distribution of evidence to deviate from the true value, and environmental noise can introduce additional errors. Adversary modeling is a key technique in multi-agent systems. It involves reasoning about and modeling the opponent's behavior, intentions, beliefs, and so on, to assist in making better decisions. Through adversary modeling, agents can better understand and respond to opponents in complex interactive environments, thereby improving their performance and benefits in cooperative or competitive scenarios. Inspired by adversary modeling, by learning a target teammate model, the teammate's messages are locally predicted, and then the predicted messages are partially corrected to improve the efficiency of collaboration in adversarial communication environments.

[0278] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A multi-agent collaborative system with double-layer trust assurance, characterized by: include: The knowledge source trust layer is used to dynamically evaluate the authenticity and integrity of the data transmitted by each agent during the multi-agent collaboration process. It includes: The agent trust management module selects trustworthy agent collaborators based on message quality assessment and dynamic reputation update; The message identity authentication module uses elliptic curve encryption and hash signature to authenticate the identity of the agent's transmitted messages and verify the data integrity; The malicious message detection and labeling module verifies the authenticity and credibility of the content of each collaborative message based on the constraint criteria based on decision consensus and uncertainty changes, identifies malicious messages and traces the identity of the intelligent agent; The knowledge transmission trust layer is used to ensure privacy and distortion compensation during the message transmission process of intelligent agents. It includes: The privacy protection processing module, based on the differential privacy mechanism, blocks the leakage of sensitive information by injecting noise into the agent's observation data or encoding features and controlling the privacy budget; The collaborative message correction module uses teammate modeling to perform local prediction and error correction on collaborative messages after noise disturbance.

2. A multi-agent collaboration system with double-layer trust assurance according to claim 1, characterized in that: In the agent trust management module, based on message quality evaluation and dynamic reputation update, trusted agent collaborators are screened, including: The quality of message transmission from agent i to agent j is defined as: Among them, u j represents the decision uncertainty of receiver j before communication, is the uncertainty remaining after communication; if v ij >0, indicating that the message effectively reduces the receiver's decision ambiguity; otherwise it is considered low-quality or malicious information; The time decay function is used to weight the quality of historical messages. The quality of the message transmitted by agent i at time t is: Its historical message quality V i Expressed as: Among them, γ∈(0,1) is the attenuation factor, which gives higher weight to recent communications, and T is the evaluation window length; Introducing reputation value l i ∈[0,1] quantifies the credibility of agent i. Initially, all agents are set to have a neutral credibility l i =0.

5. After each round of collaboration, the following rules are used to adjust the message based on whether it contains malicious content or significantly deviates from the true state: where Δ l is the reputation increment parameter; The comprehensive score of agent i is defined as the product of message quality and reputation: e i =V i ·l i Setting the threshold When not satisfied When , agent i is not selected as the cooperation partner of j, and the message transmitted by i to j is filtered.

3. A multi-agent collaboration system with double-layer trust assurance according to claim 2, characterized in that: In the message authentication module, elliptic curve encryption and hash signature are used to authenticate the identity of the agent's transmitted message and verify the data integrity, including: Randomly generate a 256-bit random number d k , calculate its SHA-512 hash value and take the first 256 bits as the private key sk k ; Using the private key sk k With the elliptic curve base point G, generate the public key: pk k =en k ·G Agent public binding ID k The public key pk k , private key is stored locally; Based on the private key sk k and the hash value SHA-512(sk k ||m), generate a temporary random number r k mod l; by R k =r k G generates a temporary public key; for the temporary public key R k 、Public key pk k Hash with message m to get e k =SHA-512(R k ||pk k ||m)mod l Calculate signature s k =(r k +e k ·sk k ) mod l, and finally output the signature pair σ k =(R k ,s k ), the length is fixed at 64 bytes; In the single signature verification phase, the verifier receives data m and signature σ k and public key pk k Afterwards, according to R k 、pk k 、m recalculate e k : Equality verification check, if it is true, the signature is valid.

4. A multi-agent collaboration system with double-layer trust assurance according to claim 3, characterized in that: The message authentication module also includes: Introduce a random number α for each signature k , to prevent attackers from bypassing verification by combining invalid signatures; k 、R k and e k ·pk k Linearly combine them to generate the aggregate value S agg 、R agg and P agg ;verify: Complete several signature verifications at once, reducing the computational complexity from O(K) to O(1).

5. A multi-agent collaboration system with double-layer trust assurance according to claim 3, characterized in that: In the malicious message detection and tagging module, the authenticity and credibility of the content of each collaborative message is verified based on the constraint principle based on decision consensus and uncertainty changes. The malicious message is identified and the identity of the agent is traced and marked, including: Malicious message characteristics are mathematically expressed as: If a message e i A message that violates any of the following cooperation principles during the integration process is considered malicious: Forward consensus principle: optimal action a * The combined confidence quality of coop (a * ,k) i (a * ,k);​ Uncertainty reduction principle: global uncertainty of the system u coop Rise (u coop >u i ).

6. A multi-agent collaboration system with double-layer trust assurance according to claim 5, characterized in that: Message verification is based on the principles of positive consensus and uncertainty reduction, including: Give each agent an initial reputation score, which can be the same initial value; For each agent: Calculate a comprehensive score based on the quality of the agent's past messages and its current reputation score; Select agents with good credit and worthy of cooperation based on the scores and send them collaboration requests; For each message received: Verify the digital signature of the message to ensure that the message is indeed from the claimed intelligent agent and has not been tampered with; Use the principles of positive consensus and uncertainty reduction to verify whether the message is true and credible; Adjust the reputation value of the agent that sent the message based on the verification result: If the message is reliable, the reputation value of the agent is increased; If there is a problem with the message, reduce the reputation value; Calculate the quality index of the message and update the historical message quality record of the agent; Based on the received and verified collaborative knowledge, the agent selects actions and adjusts strategies based on credible information to improve the overall system performance; Continuously select and accept collaboration requests based on reputation scores, complete the signing, sending and verification of messages, forming a dynamic adjustment and optimization cycle.

7. A multi-agent collaboration system with double-layer trust assurance according to claim 5, characterized in that: In the privacy protection processing module, based on the differential privacy mechanism, noise injection into the agent's observation data or encoding features and privacy budget control are used to block sensitive information leakage channels, including: The Gaussian noise mechanism is introduced, which is specifically defined as: Let the hidden state of the agent be The output of the evidence encoder is: e=ReLU(Wh+b)+η, Where η~N(0,σ 2 I) is independent Gaussian noise, W and b are weight matrix and bias term respectively; By constraining the sensitivity of the hidden state Δ h , that is, the maximum L2-norm change of adjacent hidden states; the spectral norm of the weight matrix ||W||2, calculates the sensitivity of the evidence encoding function: D f =||W||2·D h , Ensure that the noise parameter σ satisfies Satisfies (∈,δ)-differential privacy; In the implementation, gradient clipping is used to limit the update amplitude of the hidden state, and the spectral norm of the weight matrix is normalized W←W / ||W||2·c to further control the sensitivity; For multiple rounds of training, the strong combination theorem is used to calculate the cumulative privacy loss: Ensure the overall privacy budget is under control.

8. A multi-agent collaboration system with double-layer trustworthiness assurance according to claim 7, characterized in that: The privacy protection processing module also includes: Input agent set U and message set M; All agents build corresponding teammate models based on historical observations and teammate identity information; For each agent u i : For all messages received Perform the following processing: a. Generate an estimated value for the message; b. Correcting the received message based on the above estimated value to improve the reliability of the message; c. By minimizing the cross entropy between the estimated message and the true message, the estimation model parameters are optimized to improve the estimation accuracy; Based on the corrected collaboration message, the agent performs action selection and strategy update; When the agent receives a collaboration request, it fuzzifies its local observation information based on evidence encoding and then injects differential privacy noise to achieve privacy protection. The agent sends a privacy-preserving message Complete privacy-protected message sharing.

9. A multi-agent collaboration system with double-layer trust assurance according to claim 8, characterized in that: In the collaborative message correction module, teammate modeling is used to perform local prediction and error correction on collaborative messages after noise disturbance, including: Each agent i has a local observation history τ i and teammate identification d j Build a teammate model to infer the transmission message of teammate j. The specific process is as follows: Input local observation history τ i , through the gated recurrent unit to i Encoded as hidden state Will with d j Input the multi-layer perceptron to generate fusion features h fuse : Further generate the estimated confidence quality through linear mapping and normalization and uncertainty quality in Minimize the cross entropy between estimated evidence and true evidence to optimize the estimated model: Agent i receives a noisy message from teammate j; The correction process dynamically integrates the received message with the teammate model estimate: in: The error variance of the teammate model estimation is calculated by L model optimization; The total noise variance reflects the interference intensity of the communication channel.

Citation Information

Cited By

  • Privacy protection collaborative decision-making system and method

    CN120956544A

  • Charging link network-oriented network link dynamic construction method, device and equipment

    CN121098902A