Multi-agent distributed security cooperation method and related apparatus

By leveraging the collaborative work of edge agents and local and cloud agents and optimizing parameters, the problems of insufficient perception and low learning efficiency in multi-agent systems are solved, resulting in more efficient anomaly detection and more stable network security management.

CN120378884BActive Publication Date: 2026-01-02BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510415880.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2026-01-02
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

In 5G networks, the insufficient perception, low learning efficiency, and unstable parameter optimization of multi-agent systems lead to insufficient robustness and flexibility in network security management.

Method used

The edge agent perceives the local environment, selects the optimal action, and outputs abnormal parameters using a pre-trained local security monitoring model. These parameters are then combined with the aggregated parameters from the cloud agent for optimization, generating a global anomaly detection result. The side agent receives and aggregates the local anomaly parameters, transmits them to the cloud agent for further optimization, and finally generates a global anomaly detection result, which is then fed back to the edge agent.

Benefits of technology

It improves the comprehensiveness of information perception in multi-agent systems, enhances the learning efficiency of edge agents, ensures that the model adapts to changes in the local environment, and improves detection accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378884B_ABST
    Figure CN120378884B_ABST
Patent Text Reader

Abstract

The present disclosure provides a multi-agent distributed security cooperation method and related device, which comprises: perceiving the local environment to obtain current state information and local parameters; selecting based on the current state information and the local parameters to obtain the optimal behavior action; inputting the optimal behavior action into the pre-trained local security monitoring model to output the local abnormal parameters, receiving the cloud agent aggregation parameters transmitted by the edge agent, optimizing based on the local abnormal parameters and the cloud agent aggregation parameters to obtain the local abnormal detection result; transmitting the local abnormal parameters to the edge agent to make the edge agent generate the global abnormal detection result; receiving the global abnormal detection result transmitted by the edge agent, judging the local abnormal detection result based on the global abnormal detection result to obtain the optimized local abnormal detection result. The present disclosure can effectively solve the problems of insufficient multi-agent information perception, low learning efficiency and unstable parameter optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of cybersecurity technology, and in particular to a multi-agent distributed security collaboration method and related apparatus. Background Technology

[0002] This section is intended to provide background or context for the embodiments of this disclosure as set forth in the claims. The description herein is not intended to be a prior art simply because it is included in this section.

[0003] Currently, 5G networks utilize a cloud-edge collaborative architecture and an MEC platform to achieve low-latency, high-bandwidth data processing and security management. The cloud is responsible for overall management and decision-making, while the edge is responsible for data collection and real-time processing. Building on this foundation, distributed security collaboration of multi-agent systems achieves secure coordination between the cloud, edge, and endpoints through autonomous decision-making and information exchange among agents. This enhances the robustness and flexibility of the network and is widely applied in the field of network security.

[0004] However, the related technologies suffer from problems such as insufficient perception, low learning efficiency, and unstable parameter optimization. Summary of the Invention

[0005] In view of this, the purpose of this disclosure is to propose a multi-agent distributed secure cooperation method and related apparatus, which at least to some extent solves one of the technical problems in the related technologies.

[0006] To achieve the above objectives, a first aspect of the exemplary embodiments of this disclosure provides a multi-agent distributed secure cooperation method, applied to end agents, the method comprising:

[0007] It can perceive the local environment and obtain current status information and local parameters;

[0008] The optimal action is obtained by selecting based on the current state information and the local parameters;

[0009] The optimal behavior is input into a pre-trained local security monitoring model, and local anomaly parameters are output. Cloud agent aggregation parameters transmitted by the edge agent are received. Based on the local anomaly parameters and the cloud agent aggregation parameters, optimization is performed to obtain local anomaly detection results.

[0010] The local anomaly parameters are transmitted to the edge agent so that the edge agent generates a global anomaly detection result; the global anomaly detection result transmitted by the edge agent is received, and the local anomaly detection result is judged based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0011] Based on the same inventive concept, a second aspect of the exemplary embodiments of this disclosure provides a multi-agent distributed secure cooperation method applied to edge agents, the method comprising:

[0012] The local anomaly parameters transmitted by the receiving agent are aggregated to obtain the edge agent aggregated parameters.

[0013] The edge agent aggregation parameters are transmitted to the cloud agent, the cloud agent aggregation parameters transmitted by the cloud agent are received, and the cloud agent aggregation parameters are transmitted to the end agent; optimization is performed based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and the global anomaly detection result is transmitted to the end agent; wherein, the cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters.

[0014] Based on the same inventive concept, a third aspect of the exemplary embodiments of this disclosure provides a multi-agent distributed secure cooperation device, applied to edge agents, comprising:

[0015] The perception information determination module is configured to perceive the local environment and obtain current state information and local parameters;

[0016] The optimal action determination module is configured to select the optimal action based on the current state information and the local parameters.

[0017] The local detection result determination module is configured to input the optimal behavior into a pre-trained local security monitoring model, output local anomaly parameters, receive cloud agent aggregation parameters transmitted by the edge agent, optimize based on the local anomaly parameters and the cloud agent aggregation parameters, and obtain local anomaly detection results.

[0018] The local anomaly detection result determination module is configured to transmit the local anomaly parameters to the edge agent so that the edge agent generates a global anomaly detection result; receive the global anomaly detection result transmitted by the edge agent; and judge the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0019] Based on the same inventive concept, a fourth aspect of the exemplary embodiments of this disclosure provides a multi-agent distributed secure cooperation device applied to edge agents, comprising:

[0020] The edge agent aggregation parameter determination module is configured to receive local anomaly parameters transmitted by the end agent, and aggregate the local anomaly parameters to obtain edge agent aggregation parameters.

[0021] The global anomaly detection result transmission module is configured to transmit the edge agent aggregation parameters to the cloud agent, receive the cloud agent aggregation parameters transmitted by the cloud agent, and transmit the cloud agent aggregation parameters to the edge agent; optimize based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and transmit the global anomaly detection result to the edge agent; wherein, the cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters.

[0022] Based on the same inventive concept, a fifth aspect of the exemplary embodiments of this disclosure provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method as described in the first aspect.

[0023] Based on the same inventive concept, a sixth aspect of the exemplary embodiments of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method as described in the first aspect.

[0024] Based on the same inventive concept, a seventh aspect of the exemplary embodiments of this disclosure provides a computer program product including computer program instructions that, when run on a computer, cause the computer to perform the method as described in the first aspect.

[0025] As can be seen from the above, the multi-agent distributed secure cooperation method provided in this disclosure embodiment is applied to end agents, and the method includes:

[0026] It can perceive the local environment and obtain current status information and local parameters;

[0027] The optimal action is obtained by selecting based on the current state information and the local parameters;

[0028] The optimal behavior is input into a pre-trained local security monitoring model, and local anomaly parameters are output. Cloud agent aggregation parameters transmitted by the edge agent are received. Based on the local anomaly parameters and the cloud agent aggregation parameters, optimization is performed to obtain local anomaly detection results.

[0029] The local anomaly parameters are transmitted to the edge agent so that the edge agent generates a global anomaly detection result; the global anomaly detection result transmitted by the edge agent is received, and the local anomaly detection result is judged based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0030] A multi-agent distributed secure cooperation method, applied to edge agents, the method comprising:

[0031] The local anomaly parameters transmitted by the receiving agent are aggregated to obtain the edge agent aggregated parameters.

[0032] The edge agent aggregation parameters are transmitted to the cloud agent, the cloud agent aggregation parameters are received from the cloud agent, and the cloud agent aggregation parameters are transmitted to the edge agent. Based on the edge agent aggregation parameters and the cloud agent aggregation parameters, optimization is performed to obtain a global anomaly detection result, which is then transmitted to the edge agent. The cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters. This disclosure effectively addresses the problems of insufficient information perception, low learning efficiency, and unstable parameter optimization among multi-agent systems. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 A schematic diagram illustrating an application scenario of the multi-agent distributed secure cooperation method provided as an exemplary embodiment of this disclosure;

[0035] Figure 2 A flowchart illustrating an application of a multi-agent distributed secure cooperation method for end agents, provided as an exemplary embodiment of this disclosure.

[0036] Figure 3 A schematic diagram of a system architecture for a multi-agent distributed secure cooperation method provided as an exemplary embodiment of the present disclosure;

[0037] Figure 4 A flowchart illustrating a multi-agent distributed secure cooperation method applied to an edge agent, as provided in an exemplary embodiment of this disclosure.

[0038] Figure 5 A schematic diagram of a multi-agent distributed security cooperation device applied to an end agent, provided as an exemplary embodiment of the present disclosure;

[0039] Figure 6 A schematic diagram of a multi-agent distributed security cooperation device applied to an edge agent, provided as an exemplary embodiment of the present disclosure;

[0040] Figure 7 A schematic diagram of the hardware structure of an electronic device provided for an exemplary embodiment of this disclosure. Detailed Implementation

[0041] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0042] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this application's technical solution, based on the prompt message.

[0043] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0044] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this application. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this application.

[0045] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0046] To make the objectives, technical solutions, and advantages of this disclosure clearer, the principles and spirit of this disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0047] In this article, it is important to understand that any number of elements in the accompanying figures is for illustrative purposes and not for limitation, and any naming is for distinction only and has no limiting meaning.

[0048] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar words used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly. The article "a" or "an" preceding an element does not exclude the existence of multiple such elements.

[0049] The principles and spirit of this disclosure will be explained in detail below with reference to several representative embodiments.

[0050] As described in the background section, related technologies suffer from problems such as insufficient perception, low learning efficiency, and unstable parameter optimization. Specifically, with the rapid development of mobile communication, cloud computing is being integrated into redundant 5G heterogeneous networks. Therefore, 5G networks not only solve the bandwidth shortage problem previously carried by 4G networks but also achieve the goal of rapidly improving network processing capabilities with AI support. Currently, operators' 5G networks are deployed in an edge-cloud collaborative architecture. Near the service end, base stations interact with the MEC platform via 5G communication to collect data and execute policies. MEC deployment reduces data processing and operation latency, improves the real-time performance and stability of services, and provides a more agile and flexible path for service terminal operations. MECs are typically deployed in edge data centers in cities and counties. The cloud platform implements control and management functions such as network status monitoring, remote module operation and maintenance, 5G LAN communication capabilities, traffic pool management, and security management. As the "central hub" of 5G network technology, it enables integrated management, analysis, decision-making, and handling of the entire network. However, cloud-edge collaborative network security management differs from traditional fragmented, single-point protection. It requires a "systematic" approach. Therefore, integrating security capabilities into a cloud-edge collaborative system addresses security issues from the perspective of business systems, ensuring business resilience to cope with the disappearance of traditional security boundaries and the continuous expansion of the network edge. However, adopting a business-centric cloud-edge collaborative approach to solve security problems requires the collaboration of security agents (generally referred to as intelligent agents in modern times). The key question is how security agents in the cloud can collaborate with those at the edge to maximize security.

[0051] First, agents in different locations perceive different environments. Second, due to differences in agent location (e.g., MEC or cloud), their learning capabilities (including learning efficiency and synchronous learning) also differ. Updating estimated network parameters solely through gradient descent will reduce the edge's response speed to attacks. Finally, improper parameter aggregation methods in the cloud or significant differences between the distributed parameters and the agent's own training parameters can lead to unstable edge model performance. Furthermore, each agent faces different data and has different locally trained models. Directly using the locally trained model as the security monitoring model results in incomplete samples and low monitoring accuracy; uploading locally trained model parameters to the edge agent for aggregation before distributing them to the edge agent may lead to a "one-size-fits-all" approach, making the local security monitoring model unable to adapt to changes in the local environment and reducing monitoring accuracy.

[0052] To address the aforementioned problems, this disclosure provides a multi-agent distributed secure cooperation method and related apparatus, specifically including:

[0053] As can be seen from the above, the multi-agent distributed secure cooperation method provided in this disclosure embodiment is applied to end agents, and the method includes:

[0054] It can perceive the local environment and obtain current status information and local parameters;

[0055] The optimal action is obtained by selecting based on the current state information and the local parameters;

[0056] The optimal behavior is input into a pre-trained local security monitoring model, and local anomaly parameters are output. Cloud agent aggregation parameters transmitted by the edge agent are received. Based on the local anomaly parameters and the cloud agent aggregation parameters, optimization is performed to obtain local anomaly detection results.

[0057] The local anomaly parameters are transmitted to the edge agent so that the edge agent generates a global anomaly detection result; the global anomaly detection result transmitted by the edge agent is received, and the local anomaly detection result is judged based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0058] A multi-agent distributed secure cooperation method, applied to edge agents, the method comprising:

[0059] The local anomaly parameters transmitted by the receiving agent are aggregated to obtain the edge agent aggregated parameters.

[0060] The edge agent aggregates parameters to the cloud agent, receives cloud agent aggregates parameters from the cloud agent, and transmits the cloud agent aggregates parameters to the endpoint agent. Based on the edge agent aggregates parameters and the cloud agent aggregates parameters, optimization is performed to obtain a global anomaly detection result, which is then transmitted to the endpoint agent. The cloud agent aggregates parameters are obtained by the cloud agent aggregating the edge agent aggregates parameters. This disclosure utilizes the endpoint agent's perception of the local environment and selection of optimal actions, outputs local anomaly parameters using a pre-trained local security monitoring model, and optimizes these parameters in conjunction with the aggregates transmitted from the cloud agent to obtain a local anomaly detection result. Simultaneously, the local anomaly parameters are transmitted to the edge agent to generate a global anomaly detection result, which is then used to judge the local anomaly detection result to obtain an optimized detection result. The edge agent receives the local anomaly parameters transmitted from the endpoint agent, aggregates them to obtain edge agent aggregates parameters, transmits these to the cloud agent, and then receives the cloud agent aggregates parameters from the edge agent. Based on both, optimization is performed to obtain a global anomaly detection result, which is then transmitted to the endpoint agent. Based on this, the present invention can effectively solve the problem of insufficient information perception among multiple agents through information interaction, enabling each agent to obtain more comprehensive environmental information; and improve the learning efficiency of edge agents through parameter sharing mechanism, enabling them to quickly learn the experience and knowledge of cloud agents; in addition, it can effectively solve the problem of unstable parameter optimization through parameter aggregation optimization and model adaptive adjustment, enabling the model to better adapt to changes in the local environment and improve detection accuracy.

[0061] After introducing the basic principles of this disclosure, various non-limiting embodiments of this disclosure will be described in detail below.

[0062] refer to Figure 1 This is a schematic diagram of an application scenario of the multi-agent distributed secure cooperation method provided in the exemplary embodiments of this disclosure.

[0063] This application scenario includes an edge agent 101, a cloud agent 102, and a terminal agent 103. The edge agent 101, the cloud agent 102, and the terminal agent 103 can all connect via wired or wireless communication networks to achieve data interaction.

[0064] The edge intelligent agent 101 can be an electronic device located close to the user side, possessing data transmission and multimedia input / output functions, including but not limited to desktop computers, mobile phones, mobile computers, tablet computers, media players, smart wearable devices, personal digital assistants (PDAs), or other electronic devices capable of performing the aforementioned functions. This electronic device may include a processor and a display screen with touch input functionality. The display screen is used to present a graphical user interface (GUI), which can display an application interface. The processor is used to process application data, generate the GUI, and control the display of the GUI on the screen.

[0065] Both edge agent 102 and cloud agent 103 can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0066] In some exemplary embodiments, the multi-agent distributed secure cooperation method can operate on end agent 101 or edge agent 102.

[0067] When the multi-agent distributed secure collaboration method is running on edge agent 102, edge agent 102 provides multi-agent distributed secure collaboration services to users of end agent 101.

[0068] End agent 101 perceives the local environment and obtains current state information and local parameters;

[0069] The endpoint agent 101 selects the optimal action based on the current state information and the local parameters;

[0070] The edge agent 101 inputs the optimal behavior into the pre-trained local security monitoring model and outputs local anomaly parameters. The edge agent 101 receives the cloud agent aggregation parameters transmitted by the edge agent 102. The edge agent 101 optimizes based on the local anomaly parameters and the cloud agent aggregation parameters to obtain local anomaly detection results.

[0071] End agent 101 transmits the local anomaly parameters to edge agent 102, so that edge agent 102 generates a global anomaly detection result; end agent 101 receives the global anomaly detection result transmitted by edge agent 102, and end agent 101 judges the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0072] Edge agent 102 receives local anomaly parameters transmitted by end agent 101, and aggregates the local anomaly parameters to obtain edge agent aggregated parameters.

[0073] Edge agent 102 transmits the edge agent aggregation parameters to cloud agent 103, receives the cloud agent aggregation parameters transmitted by cloud agent 103, and transmits the cloud agent aggregation parameters to end agent 101; edge agent 102 optimizes based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and transmits the global anomaly detection result to end agent 101; wherein, the cloud agent aggregation parameters are obtained by cloud agent 103 aggregating the edge agent aggregation parameters.

[0074] It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of this disclosure, and the implementation of this disclosure is not limited in any way. On the contrary, the implementation of this disclosure can be applied to any applicable scenario.

[0075] refer to Figure 2 A multi-agent distributed secure cooperation method, applied to end agents, the method includes the following steps:

[0076] Step S210: Perceive the local environment and obtain the current status information and local parameters.

[0077] In practice, the following methods are used to perceive the local environment and obtain current state information and local parameters:

[0078] refer to Figure 3This includes: x terminal agents (at least one terminal agent can be set in this scheme), edge agents, and cloud agents; each terminal agent includes an environment perception module, an experience pool, a local training model, a local security monitoring model, and a loss function generation module, and can upload parameters to and receive parameters transmitted by the edge agents; each edge agent includes a parameter filtering module, a parameter aggregation module, a local security monitoring model, and a sample pool, and can transmit parameters to and receive functions from both terminal and cloud agents; each cloud agent includes an edge agent parameter library, an edge agent parameter filtering module, and a parameter integration module, and can receive parameters from and transmit parameters to both terminal and cloud agents. In this exemplary embodiment, the terminal agents interact with the environment through the environment perception module, perceive the environment state, obtain current environment state information (i.e., current state information), and collect local training model parameters (i.e., local parameters).

[0079] Step S220: Select the optimal action based on the current state information and the local parameters.

[0080] In this exemplary embodiment, the optimal behavioral action is obtained based on the current state information and the local parameters, including:

[0081] The action space of the current state information is determined, and the optimal action is obtained by selecting actions based on the action space, the current state information and the local parameters according to the maximum action value function.

[0082] In specific implementation, the action space for determining the current state information is as follows:

[0083] The interaction between edge agents and their environment can be understood as a reinforcement learning process, which can be considered a Markov decision process, also known as an MDP. An MDP can be represented by a quintuple (S, A, P, R, γ).

[0084] S represents the set of states, that is, all possible states that the edge agent explores in the environment. t This represents the state of the endpoint agent at time t.

[0085] A represents the set of actions performed by the edge agent in response to the environment. This set of actions differs among different edge agents because the environments faced by the edge agents may be inconsistent. Therefore, there are cases where the set of environments and the set of actions of an edge agent are incomplete. t This represents the action taken by the end-user agent in response to the environmental state at time t.

[0086] In this exemplary embodiment, the optimal action is obtained by selecting actions based on the action space, the current state information, and the local parameters according to the maximum action value function.

[0087] refer to Figure 3 In a locally trained model, given the local training model parameters θ (i.e., local parameters), the optimal action policy a needs to be determined by selecting the maximum action value function. max (i.e., optimal behavior).

[0088] a max =argmax a∈A Q(s, a; θ)

[0089] Where, argmax a∈A This means finding the action a from the action set A that maximizes Q(s, a; θ). max ; s represents the current state information; a represents the executable action space (i.e., action space) in the current state information.

[0090] Step S230: Input the optimal behavior action into the pre-trained local security monitoring model, output local anomaly parameters, receive cloud agent aggregation parameters transmitted by the edge agent, optimize based on the local anomaly parameters and the cloud agent aggregation parameters, and obtain local anomaly detection results.

[0091] In this exemplary embodiment, the local security monitoring model is trained in the following manner:

[0092] Construct a sample set comprising several samples; wherein the samples include: sample data and label data; the sample data includes optimal behavior actions for training; the label data includes local anomaly parameters for training;

[0093] The sample data is input into the local security monitoring model to be trained to obtain the predicted data output by the model, wherein the predicted data includes the predicted local anomaly parameters output by the model.

[0094] Determine the sample differences between the predicted data and the label data;

[0095] Based on the sample differences, the parameters of the local security monitoring model to be trained are updated through a loss function until the sample differences between the predicted data and the labeled data are minimized, thereby obtaining the pre-trained local security monitoring model.

[0096] In specific implementation, a sample set is constructed, comprising several samples; wherein, the samples include: sample data and label data; the sample data includes optimal behavior actions for training; the label data includes local anomaly parameters for training, which refer to:

[0097] refer to Figure 3 When training a local security monitoring model on an edge agent, a sample set containing several samples needs to be constructed for training the model. Each sample consists of two parts, including the optimal action 'a' used for training. max And local anomaly parameters used for training.

[0098] In practice, the sample data is input into the local security monitoring model to be trained to obtain the predicted data output by the model. The predicted data includes the predicted local anomaly parameters output by the model, which refer to:

[0099] refer to Figure 3 The endpoint agent inputs sample data output from the locally trained model into the local security monitoring model to be trained. The local security monitoring model then uses the input sample data a... max The system outputs predicted data, including predicted local anomaly parameters from the model's output. Specifically, the maximum action is input into the local safety monitoring model to calculate the maximum action value function y. t (i.e., predicting local anomaly parameters):

[0100] y t =r t +γQ(s t+1 argmax a Q(s t+1 ,a;θ);θ′)

[0101] Where, r t s represents the immediate reward obtained by the agent after performing an action at time step t; t+1 Let θ represent the environment state at time step t+1; θ′ represent the target local anomaly parameter (i.e., the local anomaly parameter used for training); and γ represent the discount factor, specifically the proportion of the expected value of the future at the current moment. To better measure the reward value obtained by the action performed by the end agent, MDP typically uses the action value function (Q function) to measure the expected future reward obtained by the end agent after performing action a under the guidance of policy π in the environment state st.

[0102] Q π (s,a)=E π [G t |s,a]

[0103] Where Gt represents the discount reward, and the expected reward over the next n time steps is R. t+n If we want to calculate the expected reward R from the next time step t+1 The sum of expected rewards up to n future moments can be expressed as follows:

[0104] G t=R t+1 +γR t+2 +γ 2 R t+3 +...+γ n-1 R t+n

[0105] Therefore, the value discount factor indicates that future rewards have a smaller and smaller impact on current rewards.

[0106] Combination Formula Q can be used π (s,a)=E π [G t |s,a] should be changed to:

[0107] Q π (s,a)=E π [R(s t ,a t )]+γQ π (s t+1 ,a t+1 ).

[0108] In practice, the loss function can be generated in the following ways:

[0109] refer to Figure 3 The edge agent calculates the difference between the predicted data and the labeled data of the local security monitoring model during training, and obtains the loss function δ. Specifically:

[0110] δ=|Q π (s t ,a t )-y t+1 =|

[0111] =|Q(s) t ,a t ;θ)-(r t +γQ(s t+1 argmax a Q(s t+1 ,a;θ);θ′))|.

[0112] In the above exemplary embodiments, a method for training a local security monitoring model was introduced. Below, we specifically describe a method for updating the parameters of the local security monitoring model to be trained using a loss function based on the sample differences, until the sample differences between the predicted data and the labeled data are minimized, thereby obtaining the pre-trained local security monitoring model:

[0113] In this exemplary embodiment, based on the sample differences, the parameters of the local security monitoring model to be trained are updated through a loss function until the sample differences between the predicted data and the labeled data are minimized, thereby obtaining the pre-trained local security monitoring model, including:

[0114] The differences in the samples are sorted in ascending order to obtain the difference ranking results;

[0115] The ranking results of the difference sorting are calculated to obtain the sampling probability results;

[0116] The parameters of the local security monitoring model to be trained are updated using the sampling probability results and the loss function until the sample difference between the predicted data and the labeled data is minimized, thus obtaining the pre-trained local security monitoring model.

[0117] In practice, the sample differences are sorted in ascending order to obtain the difference ranking results as follows:

[0118] Through multiple rounds of iteration, given the initial θ and θ′, the samples from the experience pool are input into the two models mentioned above. Each sample receives a different error (i.e., sample difference). The errors are sorted in ascending order to obtain the ranking of each sample (i.e., the difference ranking result).

[0119] In specific implementation, the ranking results of the difference sorting are calculated to obtain the sampling probability results; the parameters of the local security monitoring model to be trained are updated using the sampling probability results and the loss function until the sample difference between the predicted data and the labeled data is minimized, thus obtaining the pre-trained local security monitoring model.

[0120] Prioritization of samples is implemented based on error. Assuming the rank of sample i is ranki, then the sampling probability of this sample is p. i (i.e., the sampling probability result):

[0121]

[0122] Where σ represents the offset, which is between 0 and 1.

[0123] After prioritizing the sampling probabilities using the experience pool, "excellent" samples are input into the local training model and the local security monitoring model to obtain the preliminarily optimized θ and θ′.

[0124] The local security monitoring model based on this edge agent has been trained using deep reinforcement learning on the edge's experience pool samples, and has obtained θ′ after preliminary optimization.

[0125] In specific implementation, the cloud agent aggregation parameters transmitted by the receiving edge agent are optimized based on the local anomaly parameters and the cloud agent aggregation parameters to obtain the local anomaly detection result.

[0126] refer to Figure 3 After the cloud agent completes parameter integration, it sends the aggregated parameters to the edge agent. Once the edge agent receives the aggregated parameters, it sends them back to the end agent, thus obtaining the aggregated parameters (denoted as θ′). e-cloud The terminal agent receives θ′ e-cloud Then, combined with the local abnormal parameter y t Optimize to obtain the updated model parameters θ new The endpoint agent uses the updated model parameters θ new Perform anomaly detection to obtain more accurate local anomaly detection results.

[0127] θ new =τθ′ e-cloud +(1-τ)θ′

[0128] Where τ represents the adjustment factor, which takes a value between 0 and 1. The adjustment factor measures the difference between the parameters θ′ issued by the cloud agent and the optimized parameters θ′ of the edge agent. e-cloud The similarity.

[0129]

[0130] Step S240: Transmit the local anomaly parameters to the edge agent so that the edge agent generates a global anomaly detection result; receive the global anomaly detection result transmitted by the edge agent, and judge the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0131] In specific implementation, the local anomaly parameters are transmitted to the edge agent so that the edge agent generates a global anomaly detection result; the method for receiving the global anomaly detection result transmitted by the edge agent is as follows:

[0132] refer to Figure 3 The edge agent uploads the local anomaly parameters output by the local security monitoring model to the edge agent so that the edge agent can generate global anomaly detection results, and the edge agent receives the global anomaly detection results transmitted by the edge agent.

[0133] In this exemplary embodiment, the optimized local anomaly detection result includes: local environment anomaly result or local environment normal result;

[0134] The step of judging the local anomaly detection result based on the global anomaly detection result to obtain the optimized local anomaly detection result includes:

[0135] The local anomaly detection result is judged based on the global anomaly detection result to obtain the local environment anomaly result;

[0136] or,

[0137] The local anomaly detection result is judged based on the global anomaly detection result to obtain a normal result for the local environment.

[0138] In specific implementation, the local anomaly detection result is judged based on the global anomaly detection result to obtain an abnormal result for the local environment; or, the local anomaly detection result is judged based on the global anomaly detection result to obtain a normal result for the local environment.

[0139] refer to Figure 3 The edge agent receives global anomaly detection results from the side agent and combines them with its own locally generated anomaly detection results to make a comprehensive judgment. If the global anomaly detection results show abnormal behavior that matches the abnormal characteristics or types detected locally, or if the global anomaly detection results indicate that the abnormal behavior is spreading and similar anomalies are detected locally, then the edge agent will confirm that there is an anomaly in the local environment and thus obtain an anomaly result for the local environment. Conversely, if the global anomaly detection results do not find any obvious anomalies, or if the anomalies detected locally are not supported by the global results, the edge agent will determine that the local environment is normal and thus obtain a normal result for the local environment.

[0140] In the above exemplary embodiment, the optimized local anomaly detection result was introduced. The following further describes a method for judging the local anomaly detection result based on the global anomaly detection result to obtain a normal result for the local environment:

[0141] In this exemplary embodiment, the local anomaly detection result is judged based on the global anomaly detection result to obtain a normal local environment result, including:

[0142] The detection difference between the global anomaly detection result and the local anomaly detection result is determined, and it is determined whether the detection difference is within the preset index value range. In response to the detection difference being within the preset index value range, the normal result of the local environment is obtained.

[0143] In specific implementation, the detection difference between the global anomaly detection result and the local anomaly detection result is determined, and it is judged whether the detection difference is within a preset index value range. If the detection difference is within the preset index value range, the normal result for the local environment is obtained as follows:

[0144] refer to Figure 3 The edge agent first receives the global anomaly detection results transmitted by the side agent and compares them with its own locally generated anomaly detection results to determine the detection differences between the two. These differences can be measured by calculating the difference in anomaly scores, the similarity of anomaly features, or other quantitative indicators. Subsequently, the edge agent compares the calculated detection differences with a preset range of indicator values. If the detection difference is within the preset range, it indicates a high degree of consistency between the local and global detection results, or that the locally detected anomaly is considered normal fluctuation or a false alarm in the global context. Therefore, the edge agent judges the local environment as normal, thus obtaining a normal result for the local environment. This process, by introducing the global detection results as a reference benchmark and combining them with the judgment logic of preset indicator ranges, effectively improves the accuracy and reliability of local anomaly detection, avoiding erroneous decisions caused by local false alarms or missed alarms.

[0145] refer to Figure 4 A multi-agent distributed secure cooperation method, applied to edge agents, the method includes the following steps:

[0146] Step S410: Receive the local anomaly parameters transmitted by the receiving end agent, and aggregate the local anomaly parameters to obtain the edge agent aggregated parameters.

[0147] In specific implementation, the local anomaly parameters transmitted by the receiving agent are aggregated to obtain the edge agent aggregated parameters in the following way:

[0148] refer to Figure 3 Edge agents utilize outlier detection. Before parameter aggregation, outlier detection is performed on the parameters uploaded by each edge agent to exclude those that significantly deviate from the normal range. By performing cluster analysis on each type of parameter, edge agents with large outliers can be eliminated, leaving only k edge agents.

[0149] By introducing a regularization term (L2 regularization), the size of the model parameters can be controlled, overfitting can be avoided, and the smoothness of the model parameters can be maintained. Each edge agent computes a local model update based on its local data. When these updates are sent to edge agents for aggregation, cross-entropy can be used during the aggregation process. The initial parameters of the edge agents are θ′. e Loss function: L(θ′) e For m reliable samples sent from the experience pool:

[0150]

[0151] Where y represents the label value. This represents the estimated value obtained by training samples using the initial parameters of the edge agent.

[0152] Based on the loss parameters mentioned above, the parameter gradient for each end agent is updated as follows: Then, based on the finer gradients, the parameter update gradients for each agent are obtained as: Δθ′1, Δθ′2, Δθ′3, ..., Δθ′ k After applying a weighted average to the gradients, the updated gradient is obtained as follows:

[0153]

[0154] After applying L2 regularization, the loss function becomes as follows:

[0155]

[0156] λ is the regularization parameter, ranging from 0 to 1. After regularization, the loss function is iterated through multiple iterations to obtain the set loss threshold, resulting in the optimized parameters θ′. e-new for:

[0157] θ′ e-new =θ′ e -η(Δθ′|+2λθ′ e )

[0158] Where η is the learning rate.

[0159] Step S420: Transmit the edge agent aggregation parameters to the cloud agent, receive the cloud agent aggregation parameters transmitted by the cloud agent, and transmit the cloud agent aggregation parameters to the end agent; optimize based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and transmit the global anomaly detection result to the end agent; wherein, the cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters.

[0160] In specific implementation, the method for transmitting the aggregated parameters of the edge agent to the cloud agent is as follows:

[0161] refer to Figure 3 The edge agent will optimize the parameters θ′ e-new After being transmitted to the local security monitoring model, the data is sent to one of the modules of the cloud agent—the edge agent parameter library—where the cloud agent enables autonomous aggregation of parameters based on different scenarios.

[0162] The cloud agent aggregation parameters are obtained by aggregating the edge agent aggregation parameters by the cloud agent:

[0163] refer to Figure 3The cloud-based intelligent agent uses outlier detection; before parameter aggregation, outlier detection can be performed on the parameters uploaded by each end agent to exclude those parameters that significantly deviate from the normal range. By performing cluster analysis on each type of parameter, end agents with large outliers can be excluded, leaving only k' agents.

[0164] Secondly, since there are certain differences in the parameters between edge agents, it is necessary to filter them according to the distribution of edge agent parameters. The filtering rule is to vectorize all parameters and then compare the similarity of each vector. If the similarity between the parameter vectors of two edge agents is less than a certain threshold, then these two edge agents can be clustered. This process is repeated iteratively until all edge agents are clustered, retaining the larger clusters and removing the smaller clusters.

[0165]

[0166] Assumption and These are the vectors of edge agent j1 and edge agent j2, respectively. If similar AB If the values ​​are less than the threshold "similar", then A and B are considered to be of the same class.

[0167] Once all clustering is complete, select clusters based on the number of clusters, retaining approximately 68% (samples within one standard deviation) of the data, assuming a sample size of P.

[0168] Based on this, aggregation is performed using a weighted average method of aggregation parameters.

[0169]

[0170] In specific implementation, the cloud agent aggregate parameters transmitted by the cloud agent are received and transmitted to the edge agent; optimization is performed based on the edge agent aggregate parameters and the cloud agent aggregate parameters to obtain a global anomaly detection result, and the global anomaly detection result is transmitted to the edge agent in the following manner:

[0171] After the cloud agent aggregates parameters, it begins to distribute them. However, due to the differences in the environments of the edge agents, each edge agent needs to autonomously optimize its own parameters based on these differences. (i.e., the global anomaly detection result) can be represented as:

[0172]

[0173] Where τ represents the adjustment factor, which takes a value between 0 and 1. The adjustment factor is a measure of the parameter θ′ issued by the cloud agent.e-cloud With the optimized parameters θ′ of the edge agent e-new The similarity.

[0174]

[0175] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.

[0176] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0177] Based on the same inventive concept, corresponding to any of the above embodiments, this disclosure also provides a multi-agent distributed secure collaboration device.

[0178] refer to Figure 5 The multi-agent distributed secure cooperation device, applied to end agents, includes:

[0179] The perception information determination module 510 is configured to perceive the local environment and obtain current status information and local parameters;

[0180] The optimal action determination module 520 is configured to select the optimal action based on the current state information and the local parameters.

[0181] The local detection result determination module 530 is configured to input the optimal behavior action into a pre-trained local security monitoring model, output local anomaly parameters, receive cloud agent aggregation parameters transmitted by the edge agent, optimize based on the local anomaly parameters and the cloud agent aggregation parameters, and obtain local anomaly detection results.

[0182] The local anomaly detection result determination module 540 is configured to transmit the local anomaly parameters to the edge agent so that the edge agent generates a global anomaly detection result; receive the global anomaly detection result transmitted by the edge agent; and judge the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0183] In this exemplary embodiment, the perception information determination module 510 is specifically configured as follows:

[0184] It can perceive the local environment and obtain current status information and local parameters.

[0185] In this exemplary embodiment, the optimal action determination module 520 is specifically configured as follows:

[0186] The action space of the current state information is determined, and the optimal action is obtained by selecting actions based on the action space, the current state information and the local parameters according to the maximum action value function.

[0187] In this exemplary embodiment, the local detection result determination module 530 is specifically configured as follows:

[0188] The optimal behavior is input into a pre-trained local security monitoring model, and local anomaly parameters are output. The local security monitoring model is trained using the following method:

[0189] A sample set is constructed, comprising several samples; wherein the samples include: sample data and label data; the sample data includes optimal actions for training; the label data includes local anomaly parameters for training; the sample data is input into a local security monitoring model to be trained to obtain predicted data output by the model, wherein the predicted data includes predicted local anomaly parameters output by the model; the sample differences between the predicted data and the label data are determined; the sample differences are sorted in ascending order to obtain a difference ranking result; the ranking result is calculated to obtain a sampling probability result; the parameters of the local security monitoring model to be trained are updated using the sampling probability result and the loss function until the sample differences between the predicted data and the label data are minimized, thereby obtaining the pre-trained local security monitoring model; cloud agent aggregation parameters transmitted by the edge agent are received, and optimization is performed based on the local anomaly parameters and the cloud agent aggregation parameters to obtain local anomaly detection results.

[0190] In this exemplary embodiment, the local detection result determination module 540 is specifically configured as follows:

[0191] The local anomaly parameters are transmitted to the edge agent to generate a global anomaly detection result; the global anomaly detection result transmitted by the edge agent is received, and the local anomaly detection result is judged based on the global anomaly detection result to obtain a local environment anomaly result; or, the detection difference between the global anomaly detection result and the local anomaly detection result is determined, and it is judged whether the detection difference is within a preset index value range. If the detection difference is within the preset index value range, a local environment normal result is obtained.

[0192] refer to Figure 6 The multi-agent distributed secure cooperation device, applied to edge agents, includes:

[0193] The edge agent aggregation parameter determination module 610 is configured to receive local abnormal parameters transmitted by the end agent, and aggregate the local abnormal parameters to obtain edge agent aggregation parameters.

[0194] The global anomaly detection result transmission module 620 is configured to transmit the edge agent aggregation parameters to the cloud agent, receive the cloud agent aggregation parameters transmitted by the cloud agent, and transmit the cloud agent aggregation parameters to the edge agent; optimize based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and transmit the global anomaly detection result to the edge agent; wherein, the cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters.

[0195] In this exemplary embodiment, the edge agent aggregation parameter determination module 610 is specifically configured as follows:

[0196] The local anomaly parameters transmitted by the receiving agent are aggregated to obtain the edge agent aggregated parameters.

[0197] In this exemplary embodiment, the global anomaly detection result transmission module 620 is specifically configured as follows:

[0198] The edge agent aggregation parameters are transmitted to the cloud agent, the cloud agent aggregation parameters transmitted by the cloud agent are received, and the cloud agent aggregation parameters are transmitted to the end agent; optimization is performed based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and the global anomaly detection result is transmitted to the end agent; wherein, the cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters.

[0199] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0200] The apparatus described above is used to implement the corresponding multi-agent distributed secure cooperation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0201] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multi-agent distributed secure cooperation method described in any of the above embodiments.

[0202] Figure 7 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0203] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0204] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0205] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0206] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0207] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0208] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0209] The electronic devices described above are used to implement the corresponding multi-agent distributed secure cooperation method in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0210] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the multi-agent distributed secure cooperation method as described in any of the above embodiments.

[0211] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0212] The aforementioned non-transitory computer-readable storage media can be any available medium or data storage device that a computer can access, including but not limited to magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MOs), etc.), optical storage (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs)).

[0213] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the multi-agent distributed secure cooperation method as described in any of the embodiments in the exemplary method section above, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0214] Based on the same inventive concept, corresponding to the multi-agent distributed secure cooperation method described in any of the above embodiments, this disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processors to execute the multi-agent distributed secure cooperation method. Corresponding to the execution entity for each step in each embodiment of the multi-agent distributed secure cooperation method, the processor executing the corresponding step may belong to the corresponding execution entity.

[0215] The computer program products of the above embodiments are used to cause the computer and / or the processor to execute the multi-agent distributed secure cooperation method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0216] Those skilled in the art will recognize that embodiments of this disclosure can be implemented as a system, method, or computer program product. Therefore, this disclosure can be implemented as entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this disclosure can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.

[0217] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example,, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (not exhaustive) of a computer-readable storage medium may include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0218] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0219] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0220] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0221] It should be understood that each block of a flowchart and / or block diagram, as well as combinations of blocks in a flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine that, when executed by a computer or other programmable data processing device, creates means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0222] These computer program instructions may also be stored in a computer-readable medium that enables a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce a product comprising an instruction apparatus that implements the functions / operations specified in the boxes of a flowchart and / or block diagram.

[0223] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions that execute on the computer or other programmable apparatus can provide a process for implementing the functions / operations specified in the boxes of a flowchart and / or block diagram.

[0224] Furthermore, although the operations of the methods of this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Rather, the steps depicted in the flowcharts may be executed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0225] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0226] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0227] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.

[0228] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0229] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0230] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.

[0231] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be interpreted in the broadest sense, thereby encompassing all such modifications and equivalent structures and functions.

Claims

1. A multi-agent distributed security cooperation method, characterized in that, The application is applied to an end intelligent agent, comprising: sensing a local environment to obtain current state information and local parameters; selecting based on the current state information and the local parameters to obtain an optimal behavior action; inputting the optimal behavior action into a pre-trained local security monitoring model to output a local abnormal parameter, receiving a cloud intelligent agent aggregation parameter transmitted by an edge intelligent agent, optimizing based on the local abnormal parameter and the cloud intelligent agent aggregation parameter to obtain a local abnormal detection result, wherein the cloud intelligent agent aggregation parameter is obtained by aggregating the local abnormal parameter to obtain an edge intelligent agent aggregation parameter, and then aggregating the edge intelligent agent aggregation parameter; transmitting the local abnormal parameter to the edge intelligent agent to enable the edge intelligent agent to generate a global abnormal detection result; receiving the global abnormal detection result transmitted by the edge intelligent agent, judging the local abnormal detection result based on the global abnormal detection result to obtain an optimized local abnormal detection result.

2. The method of claim 1, wherein, The selecting based on the current state information and the local parameters to obtain an optimal behavior action comprises: determining an action space of the current state information, selecting an action based on a maximum action value function on the action space, the current state information and the local parameters to obtain the optimal behavior action.

3. The method of claim 1, wherein, The method further comprises training the local security monitoring model by the following method: constructing a sample set comprising a plurality of samples; wherein the sample comprises: sample data and label data; the sample data comprises a training optimal behavior action; and the label data comprises a training local abnormal parameter; inputting the sample data into a to-be-trained local security monitoring model to obtain prediction data output by the model, wherein the prediction data comprises a predicted local abnormal parameter output by the model; determining a sample difference between the prediction data and the label data; updating parameters of the to-be-trained local security monitoring model based on the sample difference through a loss function until the sample difference between the prediction data and the label data is minimized to obtain the pre-trained local security monitoring model.

4. The method of claim 3, wherein, The updating parameters of the to-be-trained local security monitoring model based on the sample difference through a loss function until the sample difference between the prediction data and the label data is minimized to obtain the pre-trained local security monitoring model comprises: arranging the sample difference in ascending order to obtain a difference sorting result; performing ranking calculation on the difference sorting result to obtain a sampling probability result; updating the parameters of the to-be-trained local security monitoring model through the sampling probability result and the loss function until the sample difference between the prediction data and the label data is minimized to obtain the pre-trained local security monitoring model.

5. The method of claim 1, wherein, The optimized local abnormal detection result comprises a local environment abnormal result or a local environment normal result. The optimizing based on the global abnormal detection result to obtain an optimized local abnormal detection result comprises: judging the local anomaly detection result based on the global anomaly detection result to obtain a local environment anomaly result; or, judging the local anomaly detection result based on the global anomaly detection result to obtain a local environment normal result.

6. The method of claim 5, wherein, The judging the local anomaly detection result based on the global anomaly detection result to obtain a local environment normal result comprises: determining a detection difference between the global anomaly detection result and the local anomaly detection result, judging whether the detection difference is within a preset index value range, and obtaining the local environment normal result in response to the detection difference being within the preset index value range.

7. A multi-agent distributed security cooperation method, characterized in that, application to edge agent, comprising: receiving the local anomaly parameter transmitted by the end agent, aggregating the local anomaly parameter to obtain the edge agent aggregation parameter, wherein the local anomaly parameter is obtained by sensing the local environment to obtain the current state information and the local parameter, and the optimal behavior action is obtained based on the current state information and the local parameter, and finally input into the pre-trained local safety monitoring model; transmitting the edge agent aggregation parameter to the cloud agent, receiving the cloud agent aggregation parameter transmitted by the cloud agent, and transmitting the cloud agent aggregation parameter to the end agent to enable the end agent to optimize based on the local anomaly parameter and the cloud agent aggregation parameter to obtain the local anomaly detection result; based on the edge agent aggregation parameter and the cloud agent aggregation parameter, the global anomaly detection result is obtained, and the global anomaly detection result is transmitted to the end agent to enable the end agent to judge the local anomaly detection result based on the global anomaly detection result to obtain the optimized local anomaly detection result; wherein the cloud agent aggregation parameter is obtained by the cloud agent aggregating the edge agent aggregation parameter.

8. A multi-agent distributed security cooperation apparatus, characterized by, application to end agent, comprising: a sensing information determination module configured to sense the local environment to obtain the current state information and the local parameter; an optimal action determination module configured to select based on the current state information and the local parameter to obtain the optimal behavior action; a local detection result determination module configured to input the optimal behavior action into the pre-trained local safety monitoring model to output the local anomaly parameter, receive the cloud agent aggregation parameter transmitted by the edge agent, and optimize based on the local anomaly parameter and the cloud agent aggregation parameter to obtain the local anomaly detection result, wherein the cloud agent aggregation parameter is obtained by aggregating the edge agent aggregation parameter based on the local anomaly parameter; an optimized local detection result determination module configured to transmit the local anomaly parameter to the edge agent to enable the edge agent to generate the global anomaly detection result; receiving the global anomaly detection result transmitted by the edge agent, judging the local anomaly detection result based on the global anomaly detection result to obtain the optimized local anomaly detection result.

9. A multi-agent distributed security cooperation apparatus, characterized by, application to edge agent, comprising: The edge agent aggregation parameter determination module is configured to receive the local abnormal parameter transmitted by the end agent, aggregate the local abnormal parameter, and obtain the edge agent aggregation parameter. The local abnormal parameter is obtained by sensing the local environment to obtain current state information and a local parameter, and the optimal behavior action is obtained based on the current state information and the local parameter. Finally, the optimal behavior action is input into a pre-trained local safety monitoring model to obtain the edge agent aggregation parameter. The global abnormality detection result determination module is configured to transmit the edge agent aggregation parameter to the cloud agent, receive the cloud agent aggregation parameter transmitted by the cloud agent, and transmit the cloud agent aggregation parameter to the end agent to enable the end agent to optimize based on the local abnormal parameter and the cloud agent aggregation parameter to obtain a local abnormality detection result. The edge agent aggregation parameter and the cloud agent aggregation parameter are optimized to obtain a global abnormality detection result, and the global abnormality detection result is transmitted to the end agent to enable the end agent to judge the local abnormality detection result based on the global abnormality detection result to obtain an optimized local abnormality detection result. The cloud agent aggregation parameter is obtained by the cloud agent aggregating the edge agent aggregation parameter.

10. An electronic device, comprising: A computer program product includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the method of any one of claims 1 to 7.