Multi-agent distributed security cooperation method and related device

Through information interaction and parameter optimization between the end agent and the edge agent, the problems of insufficient perception and low learning efficiency in the multi-agent system are solved, more stable network security detection is achieved, and the robustness and flexibility of the 5G network are improved.

CN120378884AActive Publication Date: 2025-07-25BEIJING UNIV OF POSTS & TELECOMM
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510415880.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-25
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

In the prior art, multi-agent systems have problems such as insufficient perception, low learning efficiency, and unstable parameter optimization in 5G networks, resulting in insufficient robustness and flexibility of network security management.

Method used

Through the end agent, the local environment is perceived, the optimal behavioral action is selected, and the pre-trained local security monitoring model outputs abnormal parameters, combined with the cloud agent aggregation parameters for optimization, and generate global abnormal detection results; the edge agent receives and aggregates local abnormal parameters, transmits them to the cloud agent for further optimization, and finally generates global abnormal detection results and feeds them back to the end agent to realize parameter sharing and model adaptive adjustment.

Benefits of technology

It improves the comprehensiveness of information perception of multi-agent systems, enhances learning efficiency, ensures the stability of parameter optimization and model adaptability, and improves the accuracy and response speed of network security detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378884A_ABST
    Figure CN120378884A_ABST
Patent Text Reader

Abstract

The invention provides a multi-agent distributed security cooperation method and a related device, and the method comprises the steps: sensing a local environment, and obtaining current state information and local parameters; selecting based on the current state information and the local parameters to obtain an optimal behavior action; inputting the optimal behavior action into a pre-trained local security monitoring model, outputting to obtain a local abnormal parameter, receiving a cloud agent aggregation parameter transmitted by the edge agent, and performing optimization based on the local abnormal parameter and the cloud agent aggregation parameter to obtain a local anomaly detection result; transmitting the local anomaly parameter to the edge agent, so that the edge agent generates a global anomaly detection result; and receiving a global anomaly detection result transmitted by the side agent, and judging the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result. According to the method and the device, the problems of insufficient multi-agent information perception, low learning efficiency and unstable parameter optimization can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of network security technologies, and in particular, to a multi-agent distributed security collaboration method and related devices. Background Art

[0002] This section aims to provide background or context for the embodiments of the present disclosure described in the claims. The description herein is not admitted to be prior art merely because it is included in this section.

[0003] Currently, through the cloud-edge collaboration architecture, the 5G network uses the MEC platform to achieve low-latency and high-bandwidth data processing and security control. The cloud side is responsible for overall management and decision-making, and the edge side is responsible for data collection and real-time processing. On this basis, the distributed security collaboration of the multi-agent system realizes the security collaboration of the cloud-edge-end through the autonomous decision-making and information interaction between agents, improving the robustness and flexibility of the network, and is widely used in the field of network security.

[0004] However, in related technologies, there are problems of insufficient perception, low learning efficiency, and unstable parameter optimization. Summary of the Invention

[0005] In view of this, an object of the present disclosure is to propose a multi-agent distributed security collaboration method and related devices, which at least solve one of the technical problems in related technologies to a certain extent.

[0006] Based on the above object, in the first aspect of an exemplary embodiment of the present disclosure, a multi-agent distributed security collaboration method is provided, which is applied to an end agent, and the method includes:

[0007] Perceive the local environment to obtain the current state information and local parameters;

[0008] Select based on the current state information and the local parameters to obtain the optimal action;

[0009] Input the optimal action into a pre-trained local security monitoring model, output the local abnormal parameters, receive the aggregated parameters of the cloud agent transmitted by the edge agent, and optimize based on the local abnormal parameters and the aggregated parameters of the cloud agent to obtain the local abnormal detection result;

[0010] Transmit the local abnormal parameters to the edge agent so that the edge agent generates a global abnormal detection result; receive the global abnormal detection result transmitted by the edge agent, and judge the local abnormal detection result based on the global abnormal detection result to obtain the optimized local abnormal detection result.

[0011] Based on the same inventive concept, a second aspect of the exemplary embodiments of the present disclosure provides a multi-agent distributed security collaboration method, which is applied to edge agents. The method includes:

[0012] Receiving the local anomaly parameters transmitted by the receiving agent, aggregating the local anomaly parameters to obtain edge agent aggregation parameters;

[0013] Transmitting the edge agent aggregation parameters to the cloud agent, receiving the cloud agent aggregation parameters transmitted by the cloud agent, and transmitting the cloud agent aggregation parameters to the end agent; optimizing based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and transmitting the global anomaly detection result to the end agent; wherein, the cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters.

[0014] Based on the same inventive concept, a third aspect of the exemplary embodiments of the present disclosure provides a multi-agent distributed security collaboration device, which is applied to edge agents and includes:

[0015] A perception information determination module configured to perceive the local environment to obtain the current state information and local parameters;

[0016] An optimal action determination module configured to select based on the current state information and the local parameters to obtain an optimal behavior action;

[0017] A local detection result determination module configured to input the optimal behavior action into a pre-trained local security monitoring model, output local anomaly parameters, receive the cloud agent aggregation parameters transmitted by the edge agent, and optimize based on the local anomaly parameters and the cloud agent aggregation parameters to obtain a local anomaly detection result;

[0018] An optimized local detection result determination module configured to transmit the local anomaly parameters to the edge agent so that the edge agent generates a global anomaly detection result; receive the global anomaly detection result transmitted by the edge agent, and judge the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0019] Based on the same inventive concept, a fourth aspect of the exemplary embodiments of the present disclosure provides a multi-agent distributed security collaboration device, which is applied to edge agents and includes:

[0020] An edge agent aggregation parameter determination module configured to receive the local anomaly parameters transmitted by the end agent, aggregate the local anomaly parameters to obtain edge agent aggregation parameters;

[0021] The global anomaly detection result transmission module is configured to transmit the edge agent aggregation parameters to the cloud agent, receive the cloud agent aggregation parameters transmitted by the cloud agent, and transmit the cloud agent aggregation parameters to the edge agent; optimize based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and transmit the global anomaly detection result to the edge agent; wherein, the cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters.

[0022] Based on the same inventive concept, a fifth aspect of the exemplary embodiments of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.

[0023] Based on the same inventive concept, a sixth aspect of the exemplary embodiments of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the first aspect.

[0024] Based on the same inventive concept, a seventh aspect of the exemplary embodiments of the present disclosure provides a computer program product including computer program instructions. When the computer program instructions run on a computer, the computer is caused to execute the method described in the first aspect.

[0025] As can be seen from the above, a multi-agent distributed security collaboration method provided by the embodiments of the present disclosure is applied to an edge agent, and the method includes:

[0026] Perceive the local environment to obtain the current state information and local parameters;

[0027] Select based on the current state information and the local parameters to obtain the optimal behavior action;

[0028] Input the optimal behavior action into a pre-trained local security monitoring model, output to obtain local anomaly parameters, receive the cloud agent aggregation parameters transmitted by the edge agent, and optimize based on the local anomaly parameters and the cloud agent aggregation parameters to obtain a local anomaly detection result;

[0029] Transmit the local anomaly parameters to the edge agent to enable the edge agent to generate a global anomaly detection result; receive the global anomaly detection result transmitted by the edge agent, and judge the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0030] A multi-agent distributed security collaboration method is applied to an edge agent, and the method includes:

[0031] Receive the local abnormal parameters transmitted by the receiving - end agent, aggregate the local abnormal parameters to obtain the edge - agent aggregation parameters;

[0032] Transmit the edge - agent aggregation parameters to the cloud agent, receive the cloud - agent aggregation parameters transmitted by the cloud agent, and transmit the cloud - agent aggregation parameters to the end agent; optimize based on the edge - agent aggregation parameters and the cloud - agent aggregation parameters to obtain the global abnormal detection result, and transmit the global abnormal detection result to the end agent; wherein, the cloud - agent aggregation parameters are obtained by the cloud agent aggregating the edge - agent aggregation parameters. The present disclosure can effectively address the problems of insufficient multi - agent information perception, low learning efficiency, and unstable parameter optimization. Brief Description of the Drawings

[0033] To more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related - technology descriptions. Obviously, the drawings in the following description are only the embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0034] Figure 1 FIG. is a schematic diagram of an application scenario of a multi - agent distributed security cooperation method provided by an exemplary embodiment of the present disclosure;

[0035] Figure 2 FIG. is a schematic flowchart of a multi - agent distributed security cooperation method applied to an end agent provided by an exemplary embodiment of the present disclosure;

[0036] Figure 3 FIG. is a schematic system architecture diagram of a multi - agent distributed security cooperation method provided by an exemplary embodiment of the present disclosure;

[0037] Figure 4 FIG. is a schematic flowchart of a multi - agent distributed security cooperation method applied to an edge agent provided by an exemplary embodiment of the present disclosure;

[0038] Figure 5 FIG. is a schematic structural diagram of a multi - agent distributed security cooperation device applied to an end agent provided by an exemplary embodiment of the present disclosure;

[0039] Figure 6 FIG. is a schematic structural diagram of a multi - agent distributed security cooperation device applied to an edge agent provided by an exemplary embodiment of the present disclosure;

[0040] Figure 7 FIG. is a schematic diagram of the hardware structure of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners

[0041] It can be understood that before using the technical solutions disclosed in the embodiments of the present application, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present application should be informed to the user and the user's authorization should be obtained through appropriate means in accordance with relevant laws and regulations.

[0042] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be executed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operations of the technical solutions of the present application according to the prompt message.

[0043] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0044] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present application. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present application.

[0045] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0046] To make the purpose, technical solutions and advantages of the present disclosure clearer and more understandable, the principles and spirit of the present disclosure will be described below with reference to several exemplary implementation manners. It should be understood that these implementation manners are only provided to enable those skilled in the art to better understand and then implement the present disclosure, rather than limiting the scope of the present disclosure in any way. On the contrary, these implementation manners are provided to make the present disclosure more thorough and complete, and to be able to convey the scope of the present disclosure completely to those skilled in the art.

[0047] In this article, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0048] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not represent any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing in front of the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connecting" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly. The article "one" or "a" before an element does not exclude the existence of multiple such elements.

[0049] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure.

[0050] As described in the background technology, there are problems of insufficient perception, low learning efficiency, and unstable parameter optimization in the related technologies. Specifically, with the rapid development of mobile communications, cloud computing is being integrated into redundant 5G heterogeneous networks. Therefore, the 5G network not only solves the problem of insufficient bandwidth originally carried by the 4G network, but also achieves the goal of rapidly improving network processing capabilities with AI support. The current operator's 5G network is an architecture with edge-cloud collaborative deployment. It interacts with the MEC platform through the base station close to the business end to implement data collection and policy execution. Based on the deployment of MEC, the delay of data processing and operation processing is reduced, the real-time and stability of the business are improved, and a more agile and flexible path is provided for the work execution of business terminals. MEC is generally deployed in edge computer rooms in cities, counties and districts. The cloud platform realizes control and management functions such as network status monitoring, remote operation and maintenance of modules, communication capabilities of 5GLAN, traffic pool management, and security management. As the "hub" of 5G network technology, it realizes integrated management, integrated analysis, integrated decision-making and integrated disposal for the entire network. However, network security management based on cloud-edge collaboration is different from traditional fragmented single-point protection and requires a "systematic" approach. Therefore, security capabilities are integrated into the cloud-edge collaboration system to solve security problems from the perspective of the business system and ensure business resilience to cope with the disappearance of traditional security boundaries and the continuous expansion of the network edge. However, using the business perspective of cloud-edge collaboration to solve security problems requires the collaboration of security agents (generally referred to as intelligent agents in modern times). How can the security agent on the cloud collaborate with the security agent on the edge to maximize security?

[0051] First, the environments perceived by agents at different locations are different. Second, due to the different positions of agents (such as MEC or the cloud), their learning abilities (including learning efficiency and synchronous learning) are also different. If only the gradient descent method is used to update the estimated network parameters, the response speed of the edge side to attack behaviors will be reduced. Finally, improper parameter aggregation methods of cloud agents or large differences between the issued parameters and their own training parameters may lead to unstable performance of the edge model. In addition, the data faced by each agent is different, and the local training models are also different. If the local training model is directly used as the security monitoring model, incomplete samples will result in low monitoring accuracy; if the parameters of the local training model are uploaded to the edge agent for aggregation and then issued to the end agent, there may be a "one-size-fits-all" phenomenon, making it difficult for the local security monitoring model to adapt to local environmental changes and resulting in a decline in monitoring accuracy.

[0052] To solve the above problems, the present disclosure provides a multi-agent distributed security collaboration method and related device solutions, specifically including:

[0053] As can be seen from the above, a multi-agent distributed security collaboration method provided by an embodiment of the present disclosure is applied to an end agent, and the method includes:

[0054] Perceive the local environment to obtain the current state information and local parameters;

[0055] Select based on the current state information and the local parameters to obtain the optimal behavioral action;

[0056] Input the optimal behavioral action into a pre-trained local security monitoring model, output the local abnormal parameters, receive the aggregated parameters of the cloud agent transmitted by the edge agent, and optimize based on the local abnormal parameters and the aggregated parameters of the cloud agent to obtain the local anomaly detection result;

[0057] Transmit the local abnormal parameters to the edge agent to enable the edge agent to generate a global anomaly detection result; receive the global anomaly detection result transmitted by the edge agent, and judge the local anomaly detection result based on the global anomaly detection result to obtain the optimized local anomaly detection result.

[0058] A multi-agent distributed security collaboration method is applied to an edge agent, and the method includes:

[0059] Receive the local abnormal parameters transmitted by the end agent, aggregate the local abnormal parameters to obtain the aggregated parameters of the edge agent;

[0060] Transmit the edge agent aggregation parameter to the cloud agent, receive the cloud agent aggregation parameter transmitted by the cloud agent, and transmit the cloud agent aggregation parameter to the terminal agent; optimize based on the edge agent aggregation parameter and the cloud agent aggregation parameter to obtain a global anomaly detection result, and transmit the global anomaly detection result to the terminal agent; wherein, the cloud agent aggregation parameter is obtained by the cloud agent aggregating the edge agent aggregation parameter. Through the perception of the local environment by the terminal agent and the selection of the optimal behavior action, the present disclosure uses a pre-trained local security monitoring model to output local anomaly parameters, and combines the aggregation parameters transmitted by the cloud agent for optimization to obtain a local anomaly detection result. At the same time, transmit the local anomaly parameters to the edge agent to generate a global anomaly detection result, and use this to judge the local anomaly detection result to obtain an optimized detection result. The edge agent receives the local anomaly parameters transmitted by the terminal agent for aggregation to obtain an edge agent aggregation parameter, transmits it to the cloud agent, then receives the cloud agent aggregation parameter transmitted by the edge agent, optimizes based on the two to obtain a global anomaly detection result and transmits it to the terminal agent. Based on this, the present disclosure can effectively solve the problem of insufficient information perception of multiple agents through information interaction, enabling each agent to obtain more comprehensive environmental information; and through the parameter sharing mechanism, improving the learning efficiency of the edge agent, enabling it to quickly learn the experience and knowledge of the cloud agent; in addition, through parameter aggregation optimization and model adaptive adjustment, the problem of unstable parameter optimization is effectively solved, enabling the model to better adapt to the changes in the local environment and improving the detection accuracy.

[0061] After introducing the basic principle of the present disclosure, the following specifically introduces various non-limiting implementation manners of the present disclosure.

[0062] Refer to Figure 1 , which is a schematic diagram of an application scenario of a multi-agent distributed security cooperation method provided by an exemplary embodiment of the present disclosure.

[0063] In this application scenario, it includes a terminal agent 101, an edge agent 102, and a cloud agent 103. Among them, the terminal agent 101, the edge agent 102, and the cloud agent 103 can all be connected through a wired or wireless communication network to achieve data interaction.

[0064] The edge agent 101 can be an electronic device near the user side with data transmission, multimedia input / output functions, including but not limited to desktop computers, mobile phones, mobile computers, tablet computers, media players, smart wearable devices, personal digital assistants (PDAs), or other electronic devices capable of implementing the above functions. The electronic device may include a processor and a display screen with touch input function, the display screen is used to present a graphical user interface, the graphical user interface can display an application interface, and the processor is used to process application data, generate a graphical user interface, and control the display of the graphical user interface on the display screen.

[0065] Both the edge agent 102 and the cloud agent 103 can be independent physical servers, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0066] In some exemplary embodiments, the multi-agent distributed security collaboration method can run on the edge agent 101 or the edge agent 102.

[0067] When the multi-agent distributed security collaboration method runs on the edge agent 102, the edge agent 102 provides a multi-agent distributed security collaboration service for the users of the edge agent 101.

[0068] The edge agent 101 perceives the local environment to obtain the current state information and local parameters;

[0069] The edge agent 101 makes a selection based on the current state information and the local parameters to obtain the optimal action;

[0070] The edge agent 101 inputs the optimal action into a pre-trained local security monitoring model to output local anomaly parameters. The edge agent 101 receives the cloud agent aggregation parameters transmitted by the edge agent 102, and the edge agent 101 optimizes based on the local anomaly parameters and the cloud agent aggregation parameters to obtain a local anomaly detection result;

[0071] The edge agent 101 transmits the local anomaly parameters to the edge agent 102 so that the edge agent 102 generates a global anomaly detection result; the edge agent 101 receives the global anomaly detection result transmitted by the edge agent 102, and the edge agent 101 judges the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0072] The edge agent 102 receives the local anomaly parameters transmitted by the edge agent 101, and the edge agent 102 aggregates the local anomaly parameters to obtain edge agent aggregation parameters;

[0073] The edge agent 102 transmits the edge agent aggregation parameters to the cloud agent 103, the edge agent 102 receives the cloud agent aggregation parameters transmitted by the cloud agent 103, and the edge agent 102 transmits the cloud agent aggregation parameters to the edge agent 101; the edge agent 102 optimizes based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and the edge agent 102 transmits the global anomaly detection result to the edge agent 101; wherein, the cloud agent aggregation parameters are obtained by the cloud agent 103 aggregating the edge agent aggregation parameters.

[0074] It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0075] Reference Figure 2 , a multi-agent distributed security cooperation method, applied to an edge agent, the method comprising the following steps:

[0076] Step S210, sense the local environment to obtain the current state information and local parameters.

[0077] Specifically, when implementing, the method for sensing the local environment to obtain the current state information and local parameters:

[0078] Reference Figure 3, including: x edge agents (at least one edge agent can be set in this solution), edge agents, and cloud agents; the edge agent includes an environment perception module, an experience pool, a local training model, a local security monitoring model, and a loss function generation module, and the edge agent can upload parameters to the edge agent and can also receive parameters transmitted by the edge agent; the edge agent includes a parameter filtering module, a parameter aggregation module, a local security monitoring model, and a sample pool, where the edge agent can transmit parameters to the edge agent and the cloud agent, and can also receive functions transmitted from the edge agent and the cloud agent; the cloud agent includes an edge agent parameter library, an edge agent parameter filtering module, and a parameter integration module, where the intelligent agent can receive parameters transmitted from the edge agent and can also transmit parameters to the edge agent and the edge agent. In this exemplary embodiment, the edge agent interacts with the environment through the environment perception module, perceives the environmental state, obtains the current environmental state information (i.e., the current state information), and collects the local training model parameters (i.e., the local parameters).

[0079] Step S220: Select based on the current state information and the local parameters to obtain the optimal behavior action.

[0080] In this exemplary embodiment, selecting based on the current state information and the local parameters to obtain the optimal behavior action includes:

[0081] Determine the action space of the current state information, and select an action based on the maximum action value function for the action space, the current state information, and the local parameters to obtain the optimal behavior action.

[0082] Specifically, when implementing, the method for determining the action space of the current state information:

[0083] The interaction between the edge agent and the environment can be understood as a process of reinforcement learning, which can be considered as a Markov decision process, also known as MDP. MDP can be represented by a five-tuple (S, A, P, R, γ).

[0084] S represents the state set, that is, all possible states explored by the edge agent in the environment. s t represents the state of the edge agent at time t.

[0085] A represents the set of actions taken by the edge agent in the face of the environment. This set of actions is different in different edge agents because the environments faced by the edge agents may be inconsistent. Therefore, there are cases where the environment set and the action set of the edge agent are incomplete. a t represents the action taken by the edge agent at time t when facing the environmental state.

[0086] In this exemplary embodiment, an optimal behavior action is obtained by selecting an action from the action space, the current state information, and the local parameters based on the maximum action value function.

[0087] Reference Figure 3 , in the local training model, under the established local training model parameters θ (i.e., local parameters), the maximum action value function to be selected is used to solve the optimal behavior strategy a max (i.e., the optimal behavior action).

[0088] a max = argmax a∈A Q(s, a; θ)

[0089] where argmax a∈A represents finding the action a that maximizes Q(s, a; θ) from the action set A max ; s represents the current state information; a represents the executable action space in the current state information (i.e., the action space).

[0090] Step S230: Input the optimal behavior action into a pre-trained local security monitoring model, output the local anomaly parameter, receive the cloud agent aggregation parameter transmitted by the edge agent, and optimize based on the local anomaly parameter and the cloud agent aggregation parameter to obtain the local anomaly detection result.

[0091] In this exemplary embodiment, the local security monitoring model is trained in the following manner:

[0092] Construct a sample set including a number of samples; wherein, the samples include: sample data and label data; the sample data includes the optimal behavior action for training; the label data includes the local anomaly parameter for training;

[0093] Input the sample data into the local security monitoring model to be trained to obtain the predicted data output by the model, wherein the predicted data includes the predicted local anomaly parameter output by the model;

[0094] Determine the sample difference between the predicted data and the label data;

[0095] Based on the sample difference, update the parameters of the local security monitoring model to be trained through a loss function until the sample difference between the predicted data and the label data is minimized, and obtain the pre-trained local security monitoring model.

[0096] Specifically, when implementing, constructing a sample set including a number of samples; wherein, the samples include: sample data and label data; the sample data includes the optimal behavior action for training; the label data includes the local anomaly parameter for training means:

[0097] Reference Figure 3 , when training the local security monitoring model with the edge agent, a sample set containing several samples needs to be constructed for training the local security monitoring model. Each sample includes two parts, including the optimal action a for training max and the local anomaly parameter for training.

[0098] During specific implementation, the sample data is input into the local security monitoring model to be trained, and the predicted data output by the model is obtained. Among them, the predicted local anomaly parameter output by the model refers to:[[]]

[0099] Reference Figure 3 , the edge agent inputs the sample data output from the local training model into the local security monitoring model to be trained. The local security monitoring model outputs the predicted data according to the input sample data a max , and the predicted data includes the predicted local anomaly parameter output by the model. Specifically, the maximum action is input into the local security monitoring model to calculate the maximum action value function y t (i.e., the predicted local anomaly parameter):

[0100] y t = r t + γQ(s t+1 , argmax a Q(s t+1 , a; θ); θ')[[]]

[0101] Among them, r t represents the immediate reward obtained by the edge agent after executing the action at time step t; s t+1 represents the environmental state at time step t + 1; θ' represents the target local anomaly parameter (i.e., the local anomaly parameter for training); γ represents the discount factor, specifically the value proportion of future expectations at the current moment. To better measure the reward value obtained by the actions executed by the edge agent, MDP usually uses the action value function (Q function) to measure the future expected return obtained by the edge agent after executing the action a under the guidance of the policy π when facing the environmental state st.[[]]

[0102] Q π (s, a) = E π [G t |s, a][[]]

[0103] Among them, Gt represents the discounted reward, and the expected reward for the next n moments is R t+n , if you want to calculate the sum of the expected rewards from the expected reward R t+1 at the next moment to the expected rewards for the next n moments, then it can be expressed as follows:[[]]

[0104] G t= R t+1 + γR t+2 + γ 2 R t+3 +... + γ n-1 R t+n

[0105] It can be seen from this that the value discount factor indicates that the influence of future rewards on the current reward is getting smaller and smaller.

[0106] Combined with The formula Q π (s, a) = E π [G t |s, a] can be changed to:

[0107] Q π (s, a) = E π [R(s t , a t )] + γQ π (s t+1 , a t+1 ).

[0108] In specific implementation, the generation method of the loss function includes:

[0109] Referring to Figure 3 , the edge agent calculates the difference between the prediction data and the label data of the local security monitoring model during training to obtain the specific loss function δ:

[0110] δ = |Q π (s t , a t ) - y t+1 = |

[0111] = |Q(s t , a t ; θ) - (r t + γQ(s t+1 , argmax a Q(s t+1 , a; θ); θ'))|.

[0112] In the above exemplary embodiment, the method of training the local security monitoring model is introduced. Next, the method of updating the parameters of the to-be-trained local security monitoring model based on the sample difference through the loss function until the sample difference between the prediction data and the label data is minimized to obtain the pre-trained local security monitoring model is specifically introduced:

[0113] In this exemplary embodiment, based on the sample differences, the parameters of the to-be-trained local security monitoring model are updated through a loss function until the sample differences between the predicted data and the label data are minimized, and the pre-trained local security monitoring model is obtained, including:

[0114] Arrange the sample differences in ascending order to obtain a difference sorting result;

[0115] Perform ranking calculation on the difference sorting result to obtain a sampling probability result;

[0116] Update the parameters of the to-be-trained local security monitoring model through the sampling probability result and the loss function until the sample differences between the predicted data and the label data are minimized, and the pre-trained local security monitoring model is obtained.

[0117] Specifically, when implementing, the method for arranging the sample differences in ascending order to obtain a difference sorting result:

[0118] Through multiple rounds of iteration, when the initial θ and θ′ are determined, the samples in the experience pool are input into the above two models, and each sample obtains different errors (i.e., sample differences). Sort the above errors in ascending order to obtain the ranking of each sample (i.e., the difference sorting result).

[0119] Specifically, when implementing, the method for performing ranking calculation on the difference sorting result to obtain a sampling probability result; and updating the parameters of the to-be-trained local security monitoring model through the sampling probability result and the loss function until the sample differences between the predicted data and the label data are minimized, and the pre-trained local security monitoring model is obtained:

[0120] Priority processing of samples is achieved based on errors. Assume that the ranking of sample i is ranki, then the sampling probability of this sample is p i (i.e., the sampling probability result):

[0121]

[0122] where σ represents the offset and is between 0 and 1.

[0123] After the experience pool evaluates the priority of the sampling probability, the "excellent" samples are input into the local training model and the local security monitoring model to obtain the preliminarily optimized θ and θ′.

[0124] The local security monitoring model based on this edge agent has been trained using deep reinforcement learning under the samples in the edge-side experience pool, and the preliminarily optimized θ′ has been obtained.

[0125] During specific implementation, the method for receiving the cloud agent aggregation parameters transmitted by the edge agent, optimizing based on the local anomaly parameters and the cloud agent aggregation parameters, and obtaining the local anomaly detection result is as follows:

[0126] Reference Figure 3 , after the cloud agent completes parameter integration, the cloud agent aggregation parameters are sent to the edge agent, and after the edge agent receives the cloud agent aggregation parameters, the cloud agent aggregation parameters are sent to the terminal agent, so that the terminal agent obtains the cloud agent aggregation parameters (which can be denoted as θ′ e-cloud ), after the terminal agent receives θ′ e-cloud , it is combined with the local anomaly parameter y t for optimization to obtain the updated model parameter θ new , and the terminal agent uses the updated model parameter θ new for anomaly detection to obtain a more accurate local anomaly detection result.

[0127] θ new =τθ′ e-cloud +(1 - τ)θ′

[0128] where τ represents a regulation factor, with a value between 0 and 1. The regulation factor measures the similarity between the parameter θ′ sent by the cloud agent and the parameter θ′ optimized by the edge agent e-cloud .

[0129]

[0130] Step S240: Transmit the local anomaly parameters to the edge agent to enable the edge agent to generate a global anomaly detection result; receive the global anomaly detection result transmitted by the edge agent, and judge the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0131] During specific implementation, the method for transmitting the local anomaly parameters to the edge agent to enable the edge agent to generate a global anomaly detection result; receiving the global anomaly detection result transmitted by the edge agent is as follows:

[0132] Reference Figure 3 , the terminal agent uploads the local anomaly parameters output by the local security monitoring model to the edge agent to enable the edge agent to generate a global anomaly detection result, and the terminal agent receives the global anomaly detection result transmitted by the edge agent.

[0133] In this exemplary embodiment, the optimized local anomaly detection result includes: a local environment anomaly result or a local environment normal result;

[0134] Judging the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result, including:

[0135] Judging the local anomaly detection result based on the global anomaly detection result to obtain a local environment anomaly result;

[0136] Or,

[0137] Judging the local anomaly detection result based on the global anomaly detection result to obtain a local environment normal result.

[0138] In specific implementation, judging the local anomaly detection result based on the global anomaly detection result to obtain a local environment anomaly result; or, the method of judging the local anomaly detection result based on the global anomaly detection result to obtain a local environment normal result:

[0139] Referring to Figure 3 , the edge agent receives the global anomaly detection result transmitted by the edge agent and makes a comprehensive judgment in combination with the local anomaly detection result generated by itself. If the global anomaly detection result shows that there is an abnormal behavior and it matches the abnormal features or types detected locally, or the global anomaly detection result indicates that the abnormal behavior is spreading and similar anomalies are detected locally, then the edge agent will confirm that there is an anomaly in the local environment, thereby obtaining a local environment anomaly result. On the contrary, if the global anomaly detection result does not find obvious anomalies, or the anomalies detected locally are not supported in the global result, the edge agent will judge that the local environment is normal, thereby obtaining a local environment normal result.

[0140] In the above exemplary embodiment, the optimized local anomaly detection result is introduced. Next, the method of judging the local anomaly detection result based on the global anomaly detection result to obtain a local environment normal result is further introduced:

[0141] In this exemplary embodiment, judging the local anomaly detection result based on the global anomaly detection result to obtain a local environment normal result, including:

[0142] Determine the detection difference between the global anomaly detection result and the local anomaly detection result, judge whether the detection difference is within a preset index value range, and in response to the detection difference being within the preset index value range, obtain the local environment normal result.

[0143] In specific implementation, determine the detection difference between the global anomaly detection result and the local anomaly detection result, judge whether the detection difference is within a preset index value range, and in response to the detection difference being within the preset index value range, the method of obtaining the local environment normal result:

[0144] Reference Figure 3 The edge agent first receives the global anomaly detection result transmitted by the edge agent and compares it with the local anomaly detection result generated by itself to determine the detection difference between the two. This difference can be measured by calculating the difference in anomaly scores, the similarity of anomaly features, or other quantitative metrics. Subsequently, the edge agent compares the calculated detection difference with the preset metric value range. If the detection difference is within the preset metric value range, it indicates that the consistency between the local detection result and the global detection result is high, or the anomalies detected locally are considered normal fluctuations or false alarms in the global context. Therefore, the edge agent determines that the local environment is in a normal state, thus obtaining the local environment normal result. This process effectively improves the accuracy and reliability of local anomaly detection by introducing the global detection result as a reference benchmark and combining the judgment logic of the preset metric value range, avoiding incorrect decisions caused by local false alarms or missed detections.

[0145] Reference Figure 4 A multi-agent distributed security collaboration method is applied to the edge agent. The method includes the following steps:

[0146] Step S410: Receive the local anomaly parameters transmitted by the edge agent, aggregate the local anomaly parameters, and obtain the edge agent aggregation parameters.

[0147] Specifically, when implementing, the method of receiving the local anomaly parameters transmitted by the edge agent and aggregating the local anomaly parameters to obtain the edge agent aggregation parameters:

[0148] Reference Figure 3 The edge agent performs outlier monitoring. Before parameter aggregation, outlier detection can be performed on the parameters uploaded by each edge agent to exclude those parameters that are significantly deviated from the normal range. By performing clustering analysis on each type of parameter, edge agents with large outliers can be excluded, leaving only k edge agents.

[0149] By introducing a regularization term (L2 regularization), the size of the model parameters can be controlled to avoid overfitting, and it also helps to maintain the smoothness of the model parameters. Each edge agent will calculate a local model update based on its own local data. When these updates are sent to the edge agent for aggregation, during the aggregation process, the cross-entropy loss function of the edge agent's initial parameter is θ′ e Loss function: L(θ′ e ) For the m trusted samples sent from the experience pool:

[0150]

[0151] where y represents the label value, Denotes the estimated value obtained by training samples using the initial parameters of the edge agent.

[0152] Based on the above loss parameters, the parameter gradient of each edge agent is updated to Then, based on the finer gradient, the parameter update gradients of each agent are: Δθ′1, Δθ′2, Δθ′3..., Δθ′ k , After weighted averaging the gradients, the weighted average updated gradient is:

[0153]

[0154] After using the L2 regularization term in the loss function, the loss function is as follows:

[0155]

[0156] λ is the regularization parameter, taking values between 0 and 1. After regularization, by iterating multiple times to make the loss function satisfy the set loss threshold, the optimized parameter θ′ is obtained e-new is:

[0157] θ′ e-new = θ′ e - η(Δθ′| + 2λθ′ e )

[0158] where η is the learning rate.

[0159] Step S420: Transmit the aggregated parameters of the edge agent to the cloud agent, receive the aggregated parameters of the cloud agent transmitted by the cloud agent, and transmit the aggregated parameters of the cloud agent to the edge agent; optimize based on the aggregated parameters of the edge agent and the aggregated parameters of the cloud agent to obtain a global anomaly detection result, and transmit the global anomaly detection result to the edge agent; wherein, the aggregated parameters of the cloud agent are obtained by the cloud agent aggregating the aggregated parameters of the edge agent.

[0160] Specifically, when implemented, the method of transmitting the aggregated parameters of the edge agent to the cloud agent:

[0161] Refer to Figure 3 , After the edge agent transmits the optimized parameter θ′ e-new to the local security monitoring model, it is sent to one of the modules of the cloud agent - the edge agent parameter library, and parameter autonomous aggregation based on different scenarios is realized in the cloud agent.

[0162] The way the aggregated parameters of the cloud agent are obtained by the cloud agent aggregating the aggregated parameters of the edge agent:

[0163] Refer to Figure 3, the cloud intelligent agent performs outlier monitoring; before parameter aggregation, it can detect outliers in the parameters uploaded by each edge intelligent agent to exclude those parameters that deviate significantly from the normal range. By performing clustering analysis on each type of parameter, edge intelligent agents with large outliers can be excluded, leaving only k' intelligent agents.

[0164] Secondly, due to certain differences in the parameters between edge intelligent agents, it is necessary to filter according to the distribution of the parameters of edge intelligent agents. The filtering rule is to vectorize all the parameters and then compare the similarity of each vector. If the similarity between the parameter vectors of two edge intelligent agents is less than a certain threshold, then these two edge intelligent agents can perform clustering operations. Repeat this iteration until all edge intelligent agents are clustered, retain the cluster with a larger number, and eliminate the cluster with a smaller number.

[0165]

[0166] Suppose and are the vectors of edge intelligent agent j1 and edge intelligent agent j2 respectively. If similar AB is less than the threshold similar, then A and B are regarded as the same class.

[0167] When all clustering is completed, select the clusters according to the number of clusters, and retain about 68% (the samples within one standard deviation) of the data samples. Suppose the sample size is P.

[0168] On this basis, the method of weighted average of aggregation parameters is used for aggregation.

[0169]

[0170] Specifically, when implementing, receive the cloud intelligent agent aggregation parameters transmitted by the cloud intelligent agent, and transmit the cloud intelligent agent aggregation parameters to the edge intelligent agent; optimize based on the edge intelligent agent aggregation parameters and the cloud intelligent agent aggregation parameters to obtain the global anomaly detection result, and transmit the global anomaly detection result to the edge intelligent agent in the following way:

[0171] When the cloud intelligent agent completes parameter aggregation, it starts to distribute the parameters. However, due to certain differences in the environments where the edge intelligent agents are located, the edge intelligent agents need to perform autonomous optimization according to the differences in their own parameters for the parameters distributed by the cloud intelligent agent. The autonomous optimization of the parameters of the edge intelligent agent (i.e., the global anomaly detection result) can be expressed as:

[0172]

[0173] Among them, τ represents the adjustment factor, and its value ranges from 0 to 1. The adjustment factor is a measure of the parameter θ' distributed by the cloud intelligent agente-cloud The similarity with the optimized parameters θ′ of the edge agent e-new .

[0174]

[0175] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0176] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0177] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure further provides a multi-agent distributed security cooperation device.

[0178] Referring to Figure 5 , the multi-agent distributed security cooperation device, applied to the edge agent, includes:

[0179] A perception information determination module 510, configured to perceive the local environment to obtain the current state information and local parameters;

[0180] An optimal action determination module 520, configured to select based on the current state information and the local parameters to obtain the optimal behavioral action;

[0181] A local detection result determination module 530, configured to input the optimal behavioral action into a pre-trained local security monitoring model, output and obtain local abnormal parameters, receive the aggregated parameters of the cloud agent transmitted by the edge agent, and optimize based on the local abnormal parameters and the aggregated parameters of the cloud agent to obtain the local abnormal detection result;

[0182] Optimize the local detection result determination module 540, which is configured to transmit the local anomaly parameter to the edge agent so that the edge agent generates a global anomaly detection result; receive the global anomaly detection result transmitted by the edge agent, and judge the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result.

[0183] In this exemplary embodiment, the perception information determination module 510 is specifically configured to:

[0184] Perceive the local environment to obtain the current state information and local parameters.

[0185] In this exemplary embodiment, the optimal action determination module 520 is specifically configured to:

[0186] Determine the action space of the current state information, and select an action based on the maximum action value function for the action space, the current state information, and the local parameters to obtain the optimal behavior action.

[0187] In this exemplary embodiment, the local detection result determination module 530 is specifically configured to:

[0188] Input the optimal behavior action into a pre-trained local security monitoring model, and output a local anomaly parameter. The local security monitoring model is trained by the following method:

[0189] Construct a sample set including a number of samples; wherein, the sample includes: sample data and label data; the sample data includes the optimal behavior action for training; the label data includes the local anomaly parameter for training; input the sample data into the local security monitoring model to be trained, and obtain the predicted data output by the model. The predicted data includes the predicted local anomaly parameter output by the model; determine the sample difference between the predicted data and the label data; sort the sample differences in ascending order to obtain a difference sorting result; calculate the ranking of the difference sorting result to obtain a sampling probability result; update the parameters of the local security monitoring model to be trained through the sampling probability result and the loss function until the sample difference between the predicted data and the label data is minimized to obtain the pre-trained local security monitoring model; receive the cloud agent aggregation parameter transmitted by the edge agent, and optimize based on the local anomaly parameter and the cloud agent aggregation parameter to obtain a local anomaly detection result.

[0190] In this exemplary embodiment, the local detection result determination module 540 is specifically configured to:

[0191] Transmit the local anomaly parameters to the edge agent so that the edge agent generates a global anomaly detection result; receive the global anomaly detection result transmitted by the edge agent, and judge the local anomaly detection result based on the global anomaly detection result to obtain a local environment anomaly result; or, determine the detection difference between the global anomaly detection result and the local anomaly detection result, and judge whether the detection difference is within a preset index value range. In response to the detection difference being within the preset index value range, obtain a local environment normal result.

[0192] Reference Figure 6 , the multi-agent distributed security cooperation device, which is applied to an edge agent, includes:

[0193] An edge agent aggregation parameter determination module 610, configured to receive the local anomaly parameters transmitted by the end agent, aggregate the local anomaly parameters, and obtain edge agent aggregation parameters;

[0194] A global anomaly detection result transmission module 620, configured to transmit the edge agent aggregation parameters to the cloud agent, receive the cloud agent aggregation parameters transmitted by the cloud agent, and transmit the cloud agent aggregation parameters to the end agent; optimize based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and transmit the global anomaly detection result to the end agent; wherein, the cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters.

[0195] In this exemplary embodiment, the edge agent aggregation parameter determination module 610 is specifically configured to:

[0196] Receive the local anomaly parameters transmitted by the end agent, aggregate the local anomaly parameters, and obtain edge agent aggregation parameters.

[0197] In this exemplary embodiment, the global anomaly detection result transmission module 620 is specifically configured to:

[0198] Transmit the edge agent aggregation parameters to the cloud agent, receive the cloud agent aggregation parameters transmitted by the cloud agent, and transmit the cloud agent aggregation parameters to the end agent; optimize based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and transmit the global anomaly detection result to the end agent; wherein, the cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters.

[0199] For the convenience of description, the above device is described by function as various modules respectively. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0200] The device of the above embodiments is used to implement the corresponding multi-agent distributed security cooperation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0201] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the multi-agent distributed security cooperation method described in any of the above embodiments.

[0202] Figure 7 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0203] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0204] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0205] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0206] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. The communication module can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0207] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0208] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0209] The electronic device of the above embodiment is used to implement the corresponding multi-agent distributed security cooperation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0210] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the multi-agent distributed security cooperation method described in any of the foregoing embodiments.

[0211] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0212] The above non-transitory computer-readable storage medium can be any available medium or data storage device accessible by a computer, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical discs (MO), etc.), optical memories (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memories (such as ROMs, EPROMs, EEPROMs, non-volatile memories (NAND FLASH), solid-state drives (SSD)), etc.

[0213] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the multi-agent distributed security cooperation method described in any one of the above exemplary method embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0214] Based on the same inventive concept, corresponding to the multi-agent distributed security cooperation method described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of the computer to cause the computer and / or the processor to execute the multi-agent distributed security cooperation method. Corresponding to the execution subjects corresponding to the steps in the respective embodiments of the multi-agent distributed security cooperation method, the processor executing the corresponding steps can belong to the corresponding execution subject.

[0215] The computer program product of the above embodiments is used to cause the computer and / or the processor to execute the multi-agent distributed security cooperation method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0216] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to as "circuit", "module", or "system" in this article. In addition, in some embodiments, the present disclosure can also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.

[0217] Any combination of one or more computer-readable media may be employed. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium may include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device.

[0218] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0219] The program code embodied on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0220] The computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0221] It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program instructions, when executed by the computer or other programmable data processing apparatus, create means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0222] These computer program instructions can also be stored in a computer-readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instruction means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0223] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer or other programmable apparatus provide processes for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0224] In addition, although the operations of the methods of the present disclosure are depicted in the figures in a particular order, this is not required or implied to perform the operations in that particular order, or to perform all of the illustrated operations to achieve the desired result. On the contrary, the steps depicted in the flowchart may be executed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step and executed, and / or one step may be decomposed into multiple steps and executed.

[0225] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion thereof, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the figures. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams or flowcharts, and combinations of blocks in the block diagrams or flowcharts, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0226] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0227] Those of ordinary skill in the art should understand that: The discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; Under the concept of the present application, the technical features between the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.

[0228] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the device may be shown in block diagram form to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0229] Although the present application has been described in connection with specific embodiments of the present application, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) can be used with the embodiments discussed.

[0230] The embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application shall be included within the protection scope of the present application.

[0231] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. Such division is only for convenience of expression. The present disclosure aims to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures and functions.

Claims

1. A multi-agent distributed security cooperation method, characterized in that Applied to the edge agent, including: Perceive the local environment to obtain the current state information and local parameters; Select based on the current state information and the local parameters to obtain the optimal behavior action; Input the optimal behavior action into a pre-trained local security monitoring model, output the local abnormal parameters, receive the cloud agent aggregation parameters transmitted by the edge agent, and optimize based on the local abnormal parameters and the cloud agent aggregation parameters to obtain the local anomaly detection result; Transmit the local abnormal parameters to the edge agent to enable the edge agent to generate a global anomaly detection result; receive the global anomaly detection result transmitted by the edge agent, and judge the local anomaly detection result based on the global anomaly detection result to obtain the optimized local anomaly detection result.

2. The method according to claim 1, wherein The selecting based on the current state information and the local parameters to obtain the optimal behavior action includes: Determine the action space of the current state information, and select an action based on the maximum action value function for the action space, the current state information, and the local parameters to obtain the optimal behavior action.

3. The method according to claim 1, wherein The method further includes training the local security monitoring model by the following method: Construct a sample set including a number of samples; wherein, the sample includes: sample data and label data; the sample data includes the optimal behavior action for training; the label data includes the local abnormal parameters for training; Input the sample data into the local security monitoring model to be trained to obtain the predicted data output by the model, wherein the predicted data includes the predicted local abnormal parameters output by the model; Determine the sample difference between the predicted data and the label data; Based on the sample difference, update the parameters of the local security monitoring model to be trained through a loss function until the sample difference between the predicted data and the label data is minimized to obtain the pre-trained local security monitoring model.

4. The method according to claim 3, characterized in that The updating the parameters of the local security monitoring model to be trained through a loss function based on the sample difference until the sample difference between the predicted data and the label data is minimized to obtain the pre-trained local security monitoring model includes: Arrange the sample differences in ascending order to obtain a difference sorting result; Calculate the ranking of the difference sorting result to obtain a sampling probability result; Update the parameters of the local security monitoring model to be trained through the sampling probability result and the loss function until the sample difference between the predicted data and the label data is minimized to obtain the pre-trained local security monitoring model.

5. The method according to claim 1, wherein The optimized local anomaly detection result includes: local environment anomaly result or local environment normal result; The judging the local anomaly detection result based on the global anomaly detection result to obtain the optimized local anomaly detection result includes: Judge the local anomaly detection result based on the global anomaly detection result to obtain a local environment anomaly result; Or, Judge the local anomaly detection result based on the global anomaly detection result to obtain a local environment normal result.

6. The method according to claim 5, wherein Judging the local anomaly detection result based on the global anomaly detection result to obtain a normal local environment result, including: Determining the detection difference between the global anomaly detection result and the local anomaly detection result, judging whether the detection difference is within a preset index value range, and in response to the detection difference being within the preset index value range, obtaining the normal local environment result.

7. A multi-agent distributed security cooperation method, characterized in that Applied to an edge agent, including: Receiving local anomaly parameters transmitted by an end agent, aggregating the local anomaly parameters to obtain edge agent aggregation parameters; Transmitting the edge agent aggregation parameters to a cloud agent, receiving the cloud agent aggregation parameters transmitted by the cloud agent, and transmitting the cloud agent aggregation parameters to the end agent; optimizing based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and transmitting the global anomaly detection result to the end agent; wherein, the cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters.

8. A multi-agent distributed security collaboration device, characterized in that, Applied to an end agent, including: A perception information determination module configured to perceive the local environment to obtain current state information and local parameters; An optimal action determination module configured to select based on the current state information and the local parameters to obtain an optimal behavior action; A local detection result determination module configured to input the optimal behavior action into a pre-trained local security monitoring model, output local anomaly parameters, receive the cloud agent aggregation parameters transmitted by an edge agent, and optimize based on the local anomaly parameters and the cloud agent aggregation parameters to obtain a local anomaly detection result; An optimized local detection result determination module configured to transmit the local anomaly parameters to the edge agent so that the edge agent generates a global anomaly detection result; receiving the global anomaly detection result transmitted by the edge agent, and judging the local anomaly detection result based on the global anomaly detection result to obtain an optimized local anomaly detection result.

9. A multi-agent distributed security collaboration device, characterized in that, Applied to an edge agent, including: An edge agent aggregation parameter determination module configured to receive local anomaly parameters transmitted by an end agent, aggregate the local anomaly parameters to obtain edge agent aggregation parameters; A global anomaly detection result transmission module configured to transmit the edge agent aggregation parameters to a cloud agent, receive the cloud agent aggregation parameters transmitted by the cloud agent, and transmit the cloud agent aggregation parameters to the end agent; optimizing based on the edge agent aggregation parameters and the cloud agent aggregation parameters to obtain a global anomaly detection result, and transmitting the global anomaly detection result to the end agent; wherein, the cloud agent aggregation parameters are obtained by the cloud agent aggregating the edge agent aggregation parameters.

10. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Distributed multi-agent cooperative fault detection method, storage medium and equipment

    CN111290277A

  • Intelligent agent control method based on end-side cloud cooperation

    CN112099510A

  • Global perception model construction method and device based on adaptive task scheduling

    CN114398160A

  • Multi-intelligence hybrid collaborative optimization method based on reinforcement learning

    CN114528766A

  • Point cloud-based multi-agent beyond-visual-range networking cooperative sensing dynamic decision-making method

    CN114815832A