Malicious traffic defense method, system and equipment based on data processing unit
By deploying the target policy network and the target value network on the network controller and executing multi-agent reinforcement learning algorithms in the distributed DPU cluster, the problem of difficult to dynamically balance defense effects and resource consumption in existing DDoS defense technologies is solved, and efficient defense efficiency and resource utilization are achieved.
Patent Information
- Application Number
- CN202510326204.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-13
AI Technical Summary
The existing DDoS defense technology is difficult to achieve dynamic balance between defense effects and resource consumption, resulting in a sharp decline in the service quality of normal traffic under high load scenarios, making it difficult to achieve real-time response.
By deploying the trained target policy network and target value network on the network controller, traffic defense policies are generated, and load balancing, packet detection and traffic filtering functions are performed in the distributed DPU cluster, and the multi-agent reinforcement learning algorithm is used to achieve dynamic resource optimization and real-time adjustment of defense policies.
It improves defense efficiency and resource utilization, achieves the service quality of normal traffic under high load scenarios to remain stable, and can quickly respond to DDoS attacks.
Smart Images

Figure CN120151044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network attack defense, and particularly to a malicious traffic defense method, system and device based on a data processing unit. Background Art
[0002] With the evolution of network attack technologies, Distributed Denial of Service (DDoS) attacks have become one of the core challenges threatening the availability of network services. Such attacks flood the target system with a massive amount of malicious traffic, resulting in the interruption of normal services. Especially in latency-sensitive scenarios such as finance, industrial control, and cloud computing, traditional defense solutions face severe bottlenecks. Existing technologies mostly rely on external traffic cleaning centers or high-defense servers, and need to redirect traffic to remote nodes for filtering. Although it can relieve the attack pressure, the additional transmission latency and control plane load introduced are difficult to meet the real-time requirements. At the same time, although a data processing unit (DPU) has the ability of near-source processing, its limited storage and computing resources make it difficult for a single node to independently cope with large-scale attack traffic. It is urgent to achieve efficient defense through multi-DPU cooperation and dynamic resource allocation. Under this background, how to combine the line-speed processing advantage of DPU and design a hierarchical dynamic defense system with low latency and high resource utilization has become the core requirement for ensuring the quality of service of critical services.
[0003] Existing DDoS traffic defense technologies are mainly divided into two categories, both of which have significant limitations. One is the localized defense system based on static rules, such as a firewall or an Intrusion Detection System (IDS) triggered by a threshold. Such solutions rely on predefined rules and are difficult to adapt to the dynamic changes of DDoS attacks (such as pulsed traffic, multi-vector attack combinations), and the lag in rule updates is likely to cause misjudgment or missed detection. The other is the centralized solution based on a traffic cleaning center, which filters traffic by redirecting it to a dedicated cleaning node. Although such methods can relieve the pressure on the target server, the detour of traffic increases the latency of control services, and the cleaning center is likely to become a single point of failure or a target of new attacks.
[0004] After summarization, existing DDoS defense function deployment solutions generally lack fine-grained modeling of DPU data plane resources and cannot achieve dynamic balance between defense effects and resource consumption, resulting in a sharp decline in the quality of service of normal traffic in high-load scenarios. For example, the optimization method of single-agent reinforcement learning abstracts the whole network defense decision into the action space of a single agent. However, in the scenario of multi-DPU cooperation, the deployment of functional chains involves multi-dimensional coupling problems such as resource allocation, topological dependence, and timing constraints. The single-agent model faces the problems of state space explosion and policy convergence difficulty and is difficult to achieve real-time response. Summary of the Invention
[0005] The present invention provides a malicious traffic defense method, system and device based on a data processing unit to improve the defense efficiency and resource utilization rate.
[0006] According to an aspect of the present invention, there is provided a malicious traffic defense method based on a data processing unit, the method comprising:
[0007] When the distributed DPU cluster receives the traffic to be processed, the network controller generates a traffic defense strategy according to the traffic to be processed and sends the traffic defense strategy to the distributed DPU cluster; wherein, a trained target policy network and a target value network are deployed on the network controller;
[0008] The distributed DPU cluster performs at least one of a load balancing function, a packet detection function and a traffic filtering function on the traffic to be processed based on the traffic defense strategy.
[0009] In a possible implementation manner, the training process of the target policy network and the target value network includes:
[0010] Based on the network structure information of the distributed DPU cluster, the distributed DPU cluster is divided into multiple agents according to the network topology level to obtain an agent set;
[0011] Define a state space, an observation space, an action space and a reward function;
[0012] According to the agent set, the state space, the observation space, the action space and the reward function, use a multi-agent reinforcement learning algorithm to train an initial policy network and an initial value network;
[0013] When the training of the initial policy network and the initial value network is completed, the target policy network and the target value network are obtained.
[0014] In a possible implementation manner, the action space includes an action estimation stage and an action placement stage corresponding to each agent in the agent set, wherein,
[0015] The action estimation stage includes generating a function execution scale based on the initial policy network;
[0016] The action placement stage includes placing primitives in the distributed DPU cluster based on the function execution scale.
[0017] In a possible implementation manner, the reward function includes a local reward function, a global reward function and a penalty term, wherein,
[0018] The local reward function is determined based on the packet loss rates before and after the agent executes the function; wherein, the agent executes the function based on the function execution scale.
[0019] The global reward function is determined based on the sum of the packet loss rates corresponding to all the agents in the agent set.
[0020] The penalty term is determined based on the placement success rate when placing the primitive.
[0021] In a possible implementation manner, training the initial policy network and the initial value network by using the multi-agent reinforcement learning algorithm includes:
[0022] Constructing a state pool as a training data set.
[0023] Generating an initial traffic defense policy based on the training data set and the initial policy network.
[0024] Generating a target reward based on the initial traffic defense policy and the initial value network.
[0025] Based on the target reward, adjusting the neural network parameters corresponding to the initial policy network and the initial value network to obtain the target policy network and the target value network.
[0026] In a possible implementation manner, adjusting the neural network parameters corresponding to the initial policy network and the initial value network based on the target reward includes:
[0027] Updating the neural network parameters corresponding to the initial value network according to the target reward and the mean squared error loss function through backpropagation of gradients.
[0028] Updating the neural network parameters corresponding to the initial policy network through policy gradients.
[0029] In a possible implementation manner, the network controller generates a traffic defense policy according to the traffic to be processed, including:
[0030] Generating the traffic defense policy through the target policy network in the network controller according to the traffic to be processed and the current network state corresponding to the distributed DPU cluster.
[0031] According to another aspect of the present invention, there is provided a malicious traffic defense system based on a data processing unit, and the system includes:
[0032] A network controller, configured to generate a traffic defense policy according to the traffic to be processed when a distributed DPU cluster receives the traffic to be processed, and send the traffic defense policy to the distributed DPU cluster; wherein, a trained target policy network and a target value network are deployed on the network controller;
[0033] The distributed DPU cluster is configured to perform at least one of a load balancing function, a packet detection function, and a traffic filtering function on the traffic to be processed based on the traffic defense policy.
[0034] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0035] At least one processor;
[0036] And a memory communicatively connected to the at least one processor; wherein,
[0037] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the malicious traffic defense method based on a data processing unit according to any embodiment of the present invention.
[0038] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the malicious traffic defense method based on a data processing unit according to any embodiment of the present invention when executed.
[0039] In the technical solution of the embodiment of the present invention, when the distributed DPU cluster receives the traffic to be processed, the network controller generates a traffic defense policy according to the traffic to be processed and sends it to the distributed DPU cluster; wherein, a trained target policy network and a target value network are deployed on the network controller; the distributed DPU cluster performs at least one of a load balancing function, a packet detection function, and a traffic filtering function on the traffic to be processed based on the traffic defense policy. In the technical solution of the present invention, a trained target policy network and a target value network are deployed on the network controller to generate a traffic defense policy, solving the problem that the existing DDoS defense function deployment solution cannot achieve a dynamic balance between the defense effect and resource consumption, resulting in a sharp decline in the service quality of normal traffic in high-load scenarios and making it difficult to achieve real-time response, and improving the defense efficiency and resource utilization rate.
[0040] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0042] Figure 1 It is a flowchart of a malicious traffic defense method based on a data processing unit provided by an embodiment of the present invention;
[0043] Figure 2 It is a framework diagram applicable to a malicious traffic defense method based on a data processing unit provided by an embodiment of the present invention;
[0044] Figure 3 It is a flowchart of another malicious traffic defense method based on a data processing unit provided by an embodiment of the present invention;
[0045] Figure 4 It is a schematic structural diagram of a malicious traffic defense system based on a data processing unit provided by an embodiment of the present invention;
[0046] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0047] In order to enable those skilled in the art to better understand the solution of the present invention, the following clearly and completely describes the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0049] Figure 1 The figure is a flowchart of a malicious traffic defense method based on a data processing unit provided by an embodiment of the present invention. This embodiment is applicable to the case of DDoS defense. This method can be executed by a malicious traffic defense system, which can be implemented in the form of hardware and / or software, and can include a distributed DPU cluster and a network controller. As Figure 1 shown, the method specifically includes the following steps:
[0050] S110. When the distributed DPU cluster receives the traffic to be processed, the network controller generates a traffic defense strategy according to the traffic to be processed and sends the traffic defense strategy to the distributed DPU cluster.
[0051] Among them, the traffic to be processed refers to network traffic, and the traffic to be processed may include malicious traffic. When the distributed DPU cluster receives the traffic to be processed, the network controller can generate a traffic defense strategy for the traffic to be processed and send the traffic defense strategy to the distributed DTU cluster in real time.
[0052] In this embodiment, a trained target policy network and a target value network are deployed on the network controller.
[0053] In a possible implementation, the network controller generates a traffic defense strategy according to the traffic to be processed, which may include: generating the traffic defense strategy through the target policy network in the network controller according to the traffic to be processed and the current network state corresponding to the distributed DPU cluster.
[0054] Among them, the current network state may be the available resource state of the distributed DPU cluster; the target policy network may be a pre-trained machine learning model for generating a defense strategy corresponding to the traffic to be processed.
[0055] Specifically, the target policy network can infer a reasonable traffic defense strategy based on the traffic to be processed and the current network state.
[0056] S120. The distributed DPU cluster performs at least one of a load balancing function, a packet detection function, and a traffic filtering function on the traffic to be processed based on the traffic defense strategy.
[0057] To further clarify the implementation process of the present invention, in combination with Figure 2 the present invention is introduced. Figure 2 The figure is a framework diagram applicable to a malicious traffic defense method based on a DPU provided by an embodiment of the present invention.
[0058] In response to the potential network layer DDoS attack threat to the business system, the present invention proposes a malicious traffic defense method, which combines three functions of load balancing (Load Balancer), packet detection (Packets Inpector), and traffic filtering (Filter) running on the DPU programmable data plane, aiming to make full use of data plane resources to maximize service quality.
[0059] In a DDoS attack scenario, multiple DPUs together form a traffic cleaning array, located between end users and business servers, and each DPU deploys a portion of the data plane functional primitives. The DPU has relevant management interfaces and is uniformly controlled by the network controller. The network controller aggregates information in real time, understands the broader context and network behavior, and issues DDoS traffic defense strategies to the distributed DPU clusters under its jurisdiction.
[0060] DPU supports protocol-independent forwarding architecture to ensure line-speed processing during data forwarding. The advantage of this approach is that it does not go through additional traffic cleaning centers, high-defense servers, etc., and can better guarantee the control service latency requirements during DDoS attacks.
[0061] It should also be noted that the DPU programmable data plane resource distribution is divided into multiple stages (stages), including matching action tables, ALUs, registers, etc., which constitute programs in the form of primitives. There are deployment constraints in terms of resource limits and primitive dependencies.
[0062] The technical solution of the embodiment of the present invention is that when the distributed DPU cluster receives the traffic to be processed, the network controller generates a traffic defense strategy according to the traffic to be processed, and sends it to the distributed DPU cluster; wherein, the network controller is deployed with a trained target policy network and a target value network; the distributed DPU cluster performs at least one of the load balancing function, the data packet detection function and the traffic filtering function on the traffic to be processed based on the traffic defense strategy. In the technical solution of the present invention, a trained target policy network and a target value network are deployed on the network controller to generate the traffic defense strategy, which solves the problem that the existing DDoS defense function deployment scheme cannot achieve a dynamic balance between the defense effect and resource consumption, resulting in a sharp decline in the service quality of normal traffic in high-load scenarios, and it is difficult to achieve real-time response, thereby improving the defense efficiency and resource utilization.
[0063] The present invention aims to solve the three core problems in traditional DDoS defense solutions, namely poor adaptability to latency-sensitive scenarios, low collaborative efficiency of multi-node resources, and insufficient ability to cope with dynamic attacks. By constructing a hierarchical defense system based on multi-agent reinforcement learning, near-source traffic cleaning and dynamic resource optimization are achieved. Specifically, by embedding the load balancing, detection, and filtering function chains into the DPU data plane, traffic detouring through external cleaning nodes is avoided, ensuring the low-latency requirements of control-type services; a hierarchical agent architecture is used to logically layer multiple DPU clusters, combined with a dynamic function chain deployment mechanism, to achieve a global optimal balance between resource utilization and defense effect; further, through a multi-agent reinforcement learning algorithm, the system can autonomously perceive changes in traffic characteristics and adjust defense strategies in real time, thereby effectively coping with the diversity and dynamicity of DDoS attacks, and finally minimizing the packet loss rate of normal traffic in a complex network environment, improving the service quality of critical services and the system's survivability.
[0064] Figure 3 It is a flowchart of another malicious traffic defense method based on a data processing unit provided by an embodiment of the present invention. This embodiment can perform problem modeling and train an initial policy network and an initial value network using a multi-agent reinforcement learning algorithm, and then generate a traffic defense strategy through the trained initial policy network.
[0065] Before introducing the specific steps of the embodiments of the present invention, the problem modeling will be described first. It can be understood that the DDoS defense function of the DPU directly acts on traffic, including changing the traffic flow direction ratio, distinguishing malicious traffic from normal traffic, and filtering detected malicious traffic. Different types of functions will have different dimensional impacts on traffic and occupy DPU resources according to the traffic scale being processed. During specific execution, the Balancer (load balancing), Inspector (packet detection), and Filter (firewall) functions have a logical execution order, that is, first, for the current network traffic and resource distribution, traffic is rate-limited and shunted to adjust the network traffic view, and then malicious DDoS traffic contained in the network traffic is statistically identified and finally filtered. The data plane resources required for the deployment of the three types of functional programs are all positively correlated with the traffic scale being processed and are compiled in a modular manner.
[0066] Let Tr * be the proportional component in the traffic Tr to be processed. SCALE(Tr * ) = k·SCALE(Tr), k ∈ [0, 1], which is used to measure the execution scale of the three types of functions. Based on this, the formal expressions of the three types of functions are given:
[0067] (1) Load balancing (B)
[0068] Functional definition: Change the original forwarding rules, perform HASH calculation based on the packet header information, and shunt or limit the traffic, that is, adjust the flow direction and traffic scale of the traffic next hop.
[0069] Tr → {Tr 0 , Tr 1 , …, Tr b-1} means dividing the original traffic Tr into multiple parts, that is, {Tr0, Tr1, ..., Trb-1}, where b is the number of possible next-hop nodes of the node traffic, and by default, the flow direction of Tr 0 is the same as that of Tr. It satisfies ∑ j=0,1,…,b-1 SCALE(Tr j ) ≤ SCALE(Tr), SCALE(Tr j ) ≥ 0, that is, the traffic scale SCALE(Trj) of each shunted Trj (where j ranges from 0 to b-1) is less than or equal to the scale SCALE(Tr) of the original traffic Tr and greater than or equal to 0, and the total traffic scale of all shunts does not exceed the scale of the original traffic Tr.
[0070] The execution scale of the load balancing function is the traffic scale it acts on, that is, SCALE(Tr * ) = k·SCALE(Tr), which has nothing to do with the adopted strategy and execution effect. The execution scale is proportional to the storage resource requirements. When this function is executed, malicious traffic and normal traffic in the network are not distinguished.
[0071] (2) Packet Detection (I)
[0072] Functional definition: Statistically analyze the traffic information on the DPU side and perform in-network analysis, distinguish the real-time network traffic of a given scale input to the DPU into normal traffic and malicious traffic that can be directly filtered, and share the traffic status and analysis results.
[0073] Tr * → {Tr *η , Tr *ζ}
[0074] Satisfy Tr * = Tr *η ∪Tr *ζ . The execution scale of the traffic detection function is the traffic scale to be detected, that is, SCALE(Tr * ) = k·SCALE(Tr), which is proportional to the storage resource requirements.
[0075] (3) Traffic Filtering (F)
[0076] Functional scope: Similar to the firewall function, filter out malicious packets passing through the DPU to avoid subsequent link congestion. At the same time, proportionally reduce the traffic bandwidth and block the malicious links contained therein.
[0077] SCALE(Tr *ζ ) → (1 - z)·SCALE(Tr *ζ ), z ∈ [0, 1]
[0078] The execution scale of the traffic filtering function is z·SCALE(Tr *ζ ), which is proportional to the storage - class resource requirement.
[0079] The present invention aims to optimize by reducing the average link reliability of normal traffic in the system during a DDoS attack. First, boolean variables are defined to represent primitive placement indicating that the primitive Pr n in the function chain fc p is placed on the DPU node V i , otherwise it means not placed.
[0080] Objective: Based on the traffic and resource status view at a certain moment, generate a set of optimal function chains After deployment under the data - plane resource constraints, obtain an expected new traffic state such that the expected packet - loss rate of the original service traffic in the network area is the lowest:
[0081]
[0082] where PL is the number of lost packets of the traffic, and PK is the total number of data packets of the traffic. Three function - chain placement constraints need to be satisfied:
[0083] a. Resource constraint: Each DPU node has a limited number of different types of resources R(V j ), and the total resource consumption of all nodes placed on the same DPU shall not exceed R(V j ):
[0084]
[0085] b. Stage constraint: For each DPU node, the number of occupied stages should be less than the maximum number of stages.
[0086] c. Order constraint:
[0087]
[0088] where the boolean matrices rsr, rsc indicate the primitive placement constraints within the function chain, respectively representing the primitive pr n in fci ,pr j Execution dependency before and after Equal to 1 indicates a dependency, otherwise none; whether cross-DPU placement is allowed Equal to 1 indicates allowed, otherwise not allowed.
[0089] The deployment of the above DDoS defense function chain is not a simple discrete space decision problem under resource constraints. Similar problems are often solved by Integer Linear Programming (ILP). In the scenario of the present invention, the execution scales of the three types of functions are restricted by the available resources, presenting as definite capacity intervals. At the same time, the execution of the functions will predictably change the traffic view, thus changing the functional requirements of the rest of the network. Therefore, conventional ILP cannot achieve the solution of the present invention.
[0090] For the above problem model, the present invention proposes a multi-agent reinforcement learning algorithm. Combining the characteristics of DDoS attacks, a hierarchical one-dimensional defense agent array is constructed to eliminate the complex graph attributes with a single DPU as an agent, and a multi-agent reinforcement learning algorithm based on hierarchical structure policy gradient is proposed.
[0091] Assume that the global network topology structure and network link model parameters are given in advance. The multi-agent decision-making is defined as a cooperative POMDP problem, represented by the agent set Environment state set Local observation set Action set denotes. Subsequently, the environment transitions to a new state where is a set of deterministic transition functions, determined by network topology, traffic model, and performance model parameters; represents the joint reward function.
[0092] Assume that all decisions are made within a certain response cycle t when a DDoS attack occurs. Additionally, within a response cycle, the original network traffic state does not change spontaneously, and the traffic state is only affected by the functions deployed by the DPU. Then, the present invention will specifically define the components of the state, observation, action, and reward, including the following steps:
[0093] S210. Based on the network structure information of the distributed DPU cluster, divide the distributed DPU cluster into multiple agents according to the network topology hierarchy to obtain an agent set.
[0094] Specifically, the agent set The determination process can be as follows: The DPU in the network area is stratified, starting from the entrance node, and each layer is used as a group of agents. Denote the DPU group of the entrance node as Agent0, and the group of DPU with the minimum hop distance of 1 from the entrance DPU as Agent0. And so on, continue to divide Agent1, 2, 3... according to the distance from the entrance DPU until reaching the cloud-side access node connected to the business system. In addition, when dividing the levels, multiple connected DPUs can be combined into a single logical DPU, and the logical DPU maintains all external topological connection relationships, and the resources are the sum of the internal DPUs.
[0095]
[0096] Each Agent contains 3 actors, representing three types of functions respectively. The actor of the DPU group at the k-th level is denoted as
[0097] S220. Define the state space, observation space, action space, and reward function.
[0098] Specifically, the state space characterizes the network traffic distribution state within the response period t, the available resources within the DPU, and the relevant definitions have been given above. The global state is formed by connecting all the observations from the DPU.
[0099]
[0100] Among them, represents the original state, and s(k) represents the network state after the k-th Agent executes the action, which conforms to the Markov property.
[0101] Specifically, the observation space The state that Agent k can observe from the environment within the response period consists of two parts, namely the traffic state flowing through this level and the remaining available resources of the DPU at this level. All the DPUs at this level share this observation information.
[0102]
[0103] o k ={Tr k ,R k}
[0104] Among them, Tr kIndicates the traffic status in the network links that have a topological connection with the DPU at level k, including the links from level k-1 to level k, the links within level k, and the links from level k to level k+1. Tr k It includes both the traffic flow relationship between nodes and the link connection relationship, and the traffic carried by the links may be 0. R k Represents the set of available resources for each DPU at level k.
[0105] In a possible implementation, the action space includes the action presumption stage and the action placement stage for each agent in the set of agents. Among them, the action presumption stage includes executing a scale based on the initial policy network generation function; the action placement stage includes placing primitives in the distributed DPU cluster based on the function execution scale.
[0106] Specifically, the action space
[0107]
[0108] Agents execute actions in hierarchical order, and each within an agent Executes actions in the order of B, I, F. The specific actions Are divided into a presumption phase and a placement phase.
[0109] a) In the presumption phase, the agent obtains a policy from the output of k From (where Refers to the initial policy network), and the DPU group at this level needs to jointly initially select actions from the continuous action space set. The action is defined as the function execution scale for each link on each DPU, which determines the function chain requirements for the three types of functions on each DPU. Then, the program primitives corresponding to the function chains are placed in the resource view.
[0110] b) In the placement phase: Taking itself as the action execution node, first place the nodes without dependencies in sequence from back to front, and then place the necessary pre-dependency nodes in sequence. When placing a new primitive, update the processing resources and storage resource status corresponding to the DPU. If there is the same primitive at the current node, only update the storage resource status. If the current node does not meet the resource requirements, look for the next node. If no node can be placed, the placement fails. Under this strategy, there is no repeated deployment of the same type of function chain in each node.
[0111] The placement strategy is based on the greedy idea. First, determine the final execution primitive pr for placement in this DPU. (end) Then, starting from this level k, place the dependent primitives level by level forward, and finally place the independent primitives level by level backward starting from level 0. If the placement is successful, update the traffic view. If the placement of the function chain fails, recycle the deployed primitives and impose a penalty in the reward function.
[0112] In a possible implementation, the reward function includes a local reward function, a global reward function, and a penalty term. Among them, the local reward function is determined based on the packet loss rate before and after the agent executes the function; among them, the agent executes the function based on the function execution scale; the global reward function is determined based on the sum of the packet loss rates corresponding to all the agents in the agent set; the penalty term is determined based on the placement success rate when placing the primitive.
[0113] Specifically, for the reward function: after each Agent takes an action in turn, obtain the reward for this round from the environment. The positive reward obtained by Agent k from the action is the local reward. Among them, the local reward is defined as the improvement of the objective function value after the action of this Agent compared to before the execution, and is obtained after the action is executed.
[0114]
[0115] Define the global reward as the improvement of the objective function value compared to the original network state after all Agent actions are executed. Therefore, we have:
[0116]
[0117] Obviously, this is a cooperative multi-agent mode. It is easy to prove that maximizing the sum of the positive rewards of all Agents is equivalent to minimizing the global objective function.
[0118] In addition, add a small negative penalty term to the reward function: r cost represents the resource cost deployed by the Agent. When the network resources are sufficient, the effects achieved by DDoS defense are basically similar. At this time, the goal is to reduce the use of resources. r fail represents the penalty for the failure of the function chain placement of this Agent, which depends on the number of failures:
[0119]
[0120] Among them, c is a given coefficient vector. |mtype| is the number of types of storage class resources; |ftype| is the number of types of function class resources.
[0121] Finally, the reward obtained by Agent k is defined as:
[0122] S230. According to the agent set, state space, observation space, action space, and reward function, use the multi-agent reinforcement learning algorithm to train the initial policy network and the initial value network.
[0123] In a possible implementation, the training of the initial policy network and the initial value network using the multi-agent reinforcement learning algorithm includes: constructing a state pool as the training data set; generating an initial traffic defense policy based on the training data set and the initial policy network; generating a target reward based on the initial traffic defense policy and the initial value network; and adjusting the neural network parameters corresponding to the initial policy network and the initial value network based on the target reward to obtain the target policy network and the target value network.
[0124] S240. When the training of the initial policy network and the initial value network is completed, obtain the target policy network and the target value network.
[0125] In a possible implementation, step S230 can be optimized into sub-steps S231 - S239, and step S240 can be optimized into sub-step S241.
[0126] S231. Construct the state pool and the training data set.
[0127] Generate a number of initial environments by randomly generating or replicating real attacks, and construct a state pool as the training data set (TrainDataSet). Use the initial traffic state to simulate DDoS attacks of different intensities, and the initial resource state remains fixed during the model training process.
[0128] S232. Initialize the neural network parameters
[0129] Initialize the parameters of all Actor group networks and Critic networks. The Actor network is responsible for generating actions, and the Critic network is responsible for evaluating the state change after the action is executed.
[0130] S233. The training round starts.
[0131] Start a new training round, and the maximum number of training rounds is preset.
[0132] S234. Sampling and state observation.
[0133] Sample the current network state from the state pool, including the network traffic distribution state and the available resources inside the DPU.
[0134] Each agent observes the traffic status flowing through its own level and the remaining available resources of the DPU at this level according to its level in the network.
[0135] S235, Action Estimation Phase.
[0136] Each agent obtains a policy from the corresponding Actor network according to the local observation vector. The three Actor networks within the agent (corresponding to load balancing, detection, and filtering functions respectively) jointly and initially select actions to determine the functional chain requirements of the three types of functions on each DPU.
[0137] S236, Action Placement Phase.
[0138] The agent places the program primitives corresponding to the functional chain in the DPU data plane according to the estimated actions.
[0139] The placement strategy is based on the greedy idea. First, place the final execution primitive on this DPU, then place the dependent primitives level by level forward starting from this level, and finally place the independent primitives level by level backward starting from the entry level.
[0140] If there are the same primitives at the current node, only update the storage resource status; if the current node does not meet the resource requirements, look for the next node; if no placeable node can be found, the placement fails.
[0141] S237, Obtain Reward and Update State.
[0142] Each agent obtains the reward for this round from the environment. The reward includes the local reward (the improvement of the objective function value after the agent's action is executed) and the global reward (the overall improvement of the objective function value after all agents' actions are executed).
[0143] If the placement of the functional chain fails, a penalty term is added to the reward function.
[0144] Update the network state according to the placement result, including the traffic view and the DPU resource status.
[0145] S238, Neural Network Parameter Update.
[0146] In a possible implementation, the neural network parameters corresponding to the initial policy network and the initial value network based on the target reward include: updating the neural network parameters corresponding to the initial value network according to the target reward and minimizing the mean square error loss function through backpropagation of gradients; updating the neural network parameters corresponding to the initial policy network through policy gradients.
[0147] Specifically, the parameters of the Critic network are updated by using the mean squared error loss function and backpropagation of gradients. The parameters of the Actor network are updated using policy gradients to guide the Actor network to descend in the direction of faster gradients.
[0148] S239. Judgment on the end of the training round.
[0149] Judge whether the current training round has reached the maximum number of times. If it has reached, end the training; if not, return to step S233 to start the next training round.
[0150] S241: System deployment and operation.
[0151] Deploy the trained neural network parameters to the network controller and the DPU cluster. The system starts to run, and the network controller issues DDoS traffic defense policies to the DPU cluster according to the real-time network status. This facilitates the DPU cluster to perform load balancing, packet detection, and traffic filtering functions according to the policies, realizing near-source traffic cleaning and dynamic resource optimization.
[0152] In practical applications, for the hierarchical decision-making problem described before step S210, K Agents will adopt an actor-critic architecture with distributed decision-making and centralized training. Specifically, each agent executes actions in hierarchical order. The DPU managed by each Agent jointly decides the scale estimation and specific placement of the function chain. There are three actor networks corresponding to Balaner, Inspector, and Filter respectively. There are a total of K critic neural networks, and their evaluation scope expands hierarchically. Each critic is used to overall evaluate the state change after the action execution and guide the actor to select actions in a globally better direction.
[0153] The actor and critic neural networks of all Agents are deployed on the controller for centralized training.
[0154] Next, the DRL state representation and training process based on MAHSPG will be introduced in detail.
[0155] 1) actor network
[0156] Based on the observation information, the action sequentially executed by Agent k under the guidance of actor B k, actor I k, actor F k is defined as:
[0157]
[0158] There are three types of actor parameters respectively, and the model structures are the same. a k The output is defined as E k -dimensional vector, where E k is the number of links involved in level k:
[0159]
[0160] The function execution scale factor corresponding to each link determines the execution strategy of the three types of functions on each DPU. The larger the factor value, the larger the function execution scale and the more resources are required. The output value of the factor that is too small is set to 0 according to the threshold ∈.
[0161] After Agent k finishes executing a k , it obtains the local reward according to the placement result When all Agents have finished executing the actions, the environment returns the rewards r of each Agent in this round k .
[0162] 2) Critic network
[0163] The Critic network performs an aggregated evaluation on the state value after the actions of each Agent are executed. For Critic k, its evaluation range is defined as Agents 0 to k, that is, the network state G from level 0 to level k after Agent k finishes executing the action {0~k} . Each Critic essentially evaluates the {0~k} rationality of the deployed functions in s globally, predicts the potential global reward size through the current local state, rather than limiting the perspective to the traffic state at this level, to avoid the policy falling into a local optimum.
[0164] 3) Training method
[0165] The actor and critic networks on all Agents are deployed on the network controller. A state pool {s (i)} is constructed using the given preset or real network state, and experience samples ex(s (i) ) are sampled by interacting with the states in it for centralized training. A single experience sample for s (i) can be expressed as ex(s (i) ) = {s (i) , (o 0 , a 0 , r 0 ), …, (o K , a K , r K )}. In one episode, for the experience set {ex(s(1) ), ex(s (2) ), ex(s (i) ), …, ex(s (N=Batch Size) )} is trained as follows:
[0166] The purpose of the critic is to predict the globally obtainable reward value through the current process state and evaluate the regional network s from level 0 to level k {0~k} That is, logically Observed long-term benefits. The output of the k-th critic neural network is denoted as Use this function to replace the action set after the given initial state s long-term benefits. That is, let:
[0167]
[0168] Parameter Update by minimizing the following loss function:
[0169]
[0170] where represents the sum of the Agent reward values in this empirical sample.
[0171] For the actor group of Agent k, The parameter training will utilize the evaluation results of critic k to guide the training of the actor group to descend in a faster gradient direction. The policy gradient corresponding to the actor parameter update can be calculated as:
[0172]
[0173] where is used to evaluate the current regional state s caused by the actions selected by the actor groups of Agent 0 to k {0~k} , and approximate the gradient of this function in the a k direction as the action gradient of the actor of Agent k based on local observations.
[0174] Based on the above analysis, the steps of the DDoS defense function chain deployment algorithm based on hierarchical agents proposed by the present invention are as follows:
[0175] + Initialization:
[0176] + Define network structure information and related network performance model parameters.
[0177] +Generate a number of initial environments s by randomly generating or replicating real attacks, and construct a state pool as the training data set TrainDataSet. Use the initial traffic state to simulate DDoS attacks of different intensities, and the initial resource state remains fixed during the model training process.
[0178] +Initialize the parameters of all actor group networks and critic networks
[0179] +For episode = 1, 2, …, Episode max do: # In each training round, the maximum is Episode max
[0180] +sample {s} |Batch Size| ∈TrainDataSet
[0181] +for s ∈ {s} |Batch Size| do:
[0182] +for Agent k in # Generate ex(s)
[0183] +Obtain the current observation o from the environment k
[0184] +Execute the action estimation stage,
[0185] +Execute the placement stage, and update the current network state to s(k) according to the predicted or real environment feedback
[0186] +Obtain the reward r from the environment k
[0187] +Calculate the comprehensive reward of the sample in this training round
[0188] +for k = 0, 1, 2,..., K do:
[0189] +Use the mean squared error loss function to update the critic k parameters through backpropagation of gradients:
[0190] +Update actor k using the policy gradient of this round:
[0191] S250. When the distributed DPU cluster receives the traffic to be processed, the network controller generates a traffic defense policy based on the traffic to be processed and distributes the traffic defense policy to the distributed DPU cluster.
[0192] Among them, a trained target policy network and a target value network are deployed on the network controller;
[0193] S260. The distributed DPU cluster performs at least one of a load balancing function, a packet detection function, and a traffic filtering function on the traffic to be processed based on the traffic defense policy.
[0194] The core innovation of the present invention is to divide the distributed DPU cluster into multiple groups of agents according to the network topology level, and construct a hierarchical collaborative defense system. Specifically, starting from the traffic entrance DPU, the DPU is divided into multiple levels (such as Agent 0, Agent 1, etc.) according to the hop distance from the business server. Each level of DPU group shares local observation information (including the traffic status and remaining resources of this level), and independently decides the execution scale of the three functions of load balancing, detection, and filtering. This design decomposes the complex global optimization problem into hierarchical local decisions by reducing the state-action space dimension, which not only adapts to the characteristics of DDoS attack traffic "penetrating from the outside to the inside gradually", but also avoids the communication overhead and policy conflict problems in the traditional multi-agent model. For example, the agent close to the edge can preferentially perform coarse-grained traffic diversion to reduce the spread of attack traffic to the core layer; while the last-level agent close to the server focuses on high-precision filtering to maximize the service quality of critical business traffic. The levels are centrally coordinated through the network controller, and finally a deep defense system from the edge to the core is formed to improve the defense efficiency and resource utilization rate.
[0195] To meet the real-time requirements of function chain deployment in dynamic attack scenarios, the present invention proposes a hierarchical policy gradient algorithm (MAHSPG) to achieve multi-DPU collaborative optimization through distributed decision-making and centralized training. Each level of agent contains three independent Actor networks, which respectively correspond to the load balancing (B), detection (I), and filtering (F) functions, and generate continuous actions (function execution scale factors) in the logical order (B→I→F) to determine the resource allocation ratio; at the same time, a Critic network with hierarchical extension is designed to evaluate the impact of the joint actions from the current level to the entrance level on the global reward layer by layer. For example, Critic k predicts the long-term benefit of its defense policy on the overall packet loss rate by aggregating the actions and states of Agent 0 to k, so as to guide the Actor network to update the parameters in the direction of the global optimum. This architecture effectively solves the "credit assignment" problem in traditional multi-agent reinforcement learning, retains the independence of hierarchical decision-making, and implicitly realizes cross-level policy coordination through the hierarchical value evaluation of Critic.
[0196] The technical solution of the present invention embeds three types of function chains, namely load balancing, detection, and filtering, into the programmable data plane of the DPU, and directly processes traffic in real time at the network edge, avoiding the additional transmission delay caused by traffic detouring through the cleaning center or high-defense server in the traditional solution. Especially in latency-sensitive scenarios (such as power dispatching and industrial control), this design can process malicious traffic at line speed to ensure that the end-to-end latency of normal service traffic meets strict quality of service requirements. At the same time, the traffic cleaning array composed of multiple DPUs can achieve near-source shunting and current limiting at the initial stage of the penetration of attack traffic through hierarchical division of labor, significantly reducing the impact on core business servers and enhancing the robustness of the overall system.
[0197] Figure 4 It is a schematic structural diagram of a malicious traffic defense system based on a data processing unit provided by an embodiment of the present invention. As Figure 4 shown, the system includes:
[0198] A network controller 310, configured to generate a traffic defense strategy according to the traffic to be processed when the distributed DPU cluster receives the traffic to be processed, and send the traffic defense strategy to the distributed DPU cluster; wherein, a trained target policy network and a target value network are deployed on the network controller;
[0199] A distributed DPU cluster 320, configured to perform at least one of a load balancing function, a packet detection function, and a traffic filtering function on the traffic to be processed based on the traffic defense strategy.
[0200] The malicious traffic defense system based on a data processing unit provided by an embodiment of the present invention can execute the malicious traffic defense method based on a data processing unit provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0201] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0202] As Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0203] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0204] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the malicious traffic defense method based on the DPU.
[0205] In some embodiments, the malicious traffic defense method based on the data processing unit can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the malicious traffic defense method based on the data processing unit described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the malicious traffic defense method based on the data processing unit in any other appropriate manner (e.g., by means of firmware).
[0206] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0207] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0208] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0209] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0210] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0211] A computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0212] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0213] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A malicious traffic defense method based on a data processing unit, characterized in that: include: When the distributed DPU cluster receives the traffic to be processed, the network controller generates a traffic defense strategy according to the traffic to be processed, and sends the traffic defense strategy to the distributed DPU cluster; wherein the network controller is deployed with a trained target strategy network and a target value network; The distributed DPU cluster performs at least one of a load balancing function, a data packet detection function and a traffic filtering function on the traffic to be processed based on the traffic defense strategy.
2. The method according to claim 1, characterized in that The training process of the target strategy network and the target value network includes: Based on the network structure information of the distributed DPU cluster, the distributed DPU cluster is divided into a plurality of intelligent agents according to the network topology level to obtain an intelligent agent set; Define state space, observation space, action space and reward function; Training an initial strategy network and an initial value network using a multi-agent reinforcement learning algorithm according to the agent set, the state space, the observation space, the action space, and the reward function; When the training of the initial policy network and the initial value network is completed, the target policy network and the target value network are obtained.
3. The method according to claim 2, characterized in that The action space includes an action estimation phase and an action placement phase corresponding to each agent in the agent set, wherein: The action inference phase includes generating a function execution scale based on the initial policy network; The action placement phase includes placing primitives in the distributed DPU cluster based on the functional execution metric.
4. The method according to claim 3, characterized in that The reward function includes a local reward function, a global reward function and a penalty term, wherein: The local reward function is determined based on the packet loss rate before and after the agent performs a function; wherein the agent performs the function based on the function execution scale; The global reward function is determined based on the sum of the packet loss rates corresponding to all the agents in the agent set; The penalty term is determined based on a placement success rate when placing the primitive.
5. The method according to claim 2, characterized in that: The method of using a multi-agent reinforcement learning algorithm to train an initial strategy network and an initial value network includes: By constructing a state pool as a training dataset; Generate an initial traffic defense strategy based on the training data set and the initial strategy network; Generate a target reward based on the initial traffic defense strategy and the initial value network; The neural network parameters corresponding to the initial policy network and the initial value network are adjusted based on the target reward to obtain the target policy network and the target value network.
6. The method according to claim 5, characterized in that The neural network parameters corresponding to the initial strategy network and the initial value network based on the target reward include: According to the target reward and the minimized mean square error loss function, the neural network parameters corresponding to the initial value network are updated by reverse gradient propagation; The corresponding neural network parameters of the initial policy network are updated through policy gradient.
7. The method according to claim 1, characterized in that The network controller generates a traffic defense strategy according to the traffic to be processed, including: The traffic defense strategy is generated through a target strategy network in the network controller according to the current network status corresponding to the traffic to be processed and the distributed DPU cluster.
8. A malicious traffic defense system based on a data processing unit, characterized in that: include: A network controller is used to generate a traffic defense strategy according to the traffic to be processed when the distributed DPU cluster receives the traffic to be processed, and send the traffic defense strategy to the distributed DPU cluster; wherein the network controller is deployed with a trained target strategy network and a target value network; The distributed DPU cluster is used to perform at least one of a load balancing function, a data packet detection function and a traffic filtering function on the traffic to be processed based on the traffic defense strategy.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the malicious traffic defense method based on a data processing unit as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the malicious traffic defense method based on a data processing unit according to any one of claims 1 to 7 when executed.
Citation Information
Cited By
Hardware-level network isolation and security protection method based on data processing unit
CN121984793A