Distributed consistent fusion estimation denial of service sequence screening method and system
By constructing a communication topology and Markov decision model under the influence of denial-of-service (DoS) attacks, and using Q-learning to select the optimal DoS sequence, the problem of dynamic topology switching in a distributed consistent fusion estimation system is solved, thereby improving the system's defense capability and convergence resilience.
Patent Information
- Application Number
- CN202511561790.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies are insufficient to effectively address denial-of-service attacks that cause dynamic topology switching in distributed consistent fusion estimation systems, affecting system convergence speed and connectivity, and lack dynamic defense mechanisms.
A graph theory approach is used to construct the communication topology, quantify the impact of denial-of-service (DoS) attacks, build a Markov decision model, use Q-learning reinforcement learning to select the optimal DoS sequence, and dynamically adjust the strategy to delay system convergence.
It significantly improves the resilience and defense flexibility of distributed systems, quantifies network connectivity by minimum non-zero eigenvalues, optimizes denial-of-service strategies, maximizes the delay of system convergence speed, and provides a theoretical basis for defense under low energy constraints.
Smart Images

Figure CN121508922A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary fields of networked control and information security, specifically to a denial-of-service sequence filtering method and system based on Q-learning-based distributed consistent fusion estimation. Background Technology
[0002] Distributed consensus fusion estimation and optimal denial-of-service strategies are core research areas for secure collaborative control of networked multi-agent systems, widely applied in critical scenarios such as UAV swarms, industrial IoT, and smart grids. Current distributed estimation methods achieve collaborative state estimation among nodes through consensus protocols, offering significant advantages over centralized methods, including high fault tolerance and strong scalability. However, the openness of distributed networks also exposes them to security threats such as denial-of-service (DoS) attacks. Malicious actors can disrupt topological connectivity by blocking communication channels, potentially leading to decreased estimation performance or even system failure. Therefore, researching optimal defense strategies and resilience enhancement mechanisms for distributed estimation systems is crucial for ensuring the secure operation of critical infrastructure.
[0003] The core of distributed consensus fusion estimation lies in achieving global state convergence through information exchange between neighboring nodes. Distributed algorithms based on consensus protocols use the Laplace matrix to describe the network topology, and its smallest non-zero eigenvalue (algebraic connectivity) directly determines the system's convergence speed. Existing research uses semidefinite programming (SDP) to optimize communication weights to accelerate convergence or designs event-triggered mechanisms to reduce communication load. However, these methods typically assume that the network topology is static or ideally controllable, while in real systems, channel packet loss, node failures, or denial-of-service interference can lead to dynamic topology switching, significantly affecting estimation performance. How to model the topology dynamics under the influence of denial-of-service and quantify its impact on consistency has become a key issue in the reliability of distributed estimation.
[0004] Denial-of-service (DoS) attacks disrupt network connectivity by affecting channel quality, posing a major security challenge to distributed systems. However, existing research often assumes that DoS behavior is unidirectional. For example, Chinese invention patent application CN120151030A, "A Distributed DoS Attack and Defense Method Based on Multi-Agent Reinforcement Learning," abstracts the DoS attack and defense scenario into a stochastic game model, clarifies the strategies of both parties and the optimal strategy, and progressively solves the Nash equilibrium of the stochastic game to optimize the attack and defense strategies. However, this method neglects the dynamic responses of the system (such as power adjustments or channel switching). How to construct a secure game model and solve for the equilibrium strategy is a key direction for future research.
[0005] The collaborative analysis of distributed estimation and denial-of-service (DoS) defense requires the support of multidisciplinary technologies. Algebraic connectivity in graph theory provides an indicator for quantifying the impact of DoS attacks, consensus protocols and event-triggered mechanisms in control theory lay the foundation for designing robust estimation methods, while reinforcement learning and game theory provide tools for optimizing security strategies. Current challenges lie in real-time defense adjustment under dynamic topologies, the accuracy of DoS detection, and the design of multi-agent collaborative fault-tolerance mechanisms. With the popularization of the Industrial Internet of Things (IIoT) and smart grids, the secure collaborative control of distributed systems will become a common focus of academia and engineering. Future research needs to further explore secure game theory under local information constraints, lightweight fault-tolerance algorithm design, and experimental verification in real-world systems to promote the transformation of theory into application. For example, in UAV swarms, the system needs to design adaptive topology reconfiguration strategies to address potential DoS interference. Modeling and optimizing such dynamic security scenarios will deepen the understanding of distributed system reliability and provide new ideas for collaborative control in complex environments. Summary of the Invention
[0006] The technical problem to be solved by this invention is how to consider network topology switching, and how to select the denial-of-service sequence that maximizes the delay of the convergence speed of the distributed consistent fusion estimation system under the minimum energy allocation, thereby providing a theoretical guarantee for improving the resilience of critical infrastructure.
[0007] The present invention solves the above-mentioned technical problems through the following technical means: This invention provides a method for filtering denial-of-service sequences based on distributed consistent fusion estimation, comprising the following steps: S1. Construct the communication topology of the estimation system based on the distributed consensus protocol using graph theory methods; quantify its consensus convergence state, communication topology switching under the influence of denial-of-service, and network connectivity. S2. Construct an objective function to measure the impact of denial-of-service influencer energy allocation strategies on the consistency convergence performance of distributed systems. S3. Construct a Markov decision model for network topology switching caused by denial of service, resulting in decreased connectivity and slowed convergence speed. S4, based on Q-learning Reinforcement learning is used to iteratively solve the problem and select the optimal denial-of-service sequence under the constraint of finite energy.
[0008] Furthermore, the communication topology of the estimation system for constructing the distributed consensus protocol described in step S1 includes the following steps: (1) Represent the network topology as G ,sensor i scalar state at each time step Then in The time represents the initial state of the sensor network; the first...i The dynamic model of a first-order continuous system with one sensor is: ; in, For the first i The status values of each sensor, Representing the i Control input from one sensor, V The number of sensor network nodes. ; (2) The first in the system i The control input for each sensor is:
[0009] in, Representative and the i A set of neighboring nodes for which sensors exchange data. It is a sensor j To the sensor i The weight of the corresponding communication edge; A continuous system is represented as:
[0010] And transcribed as:
[0011] in, L For the image The Laplacian matrix, and the corresponding eigenvectors are .
[0012] Further, the estimation of the system consensus convergence state of the quantization distributed consensus protocol described in step S1 includes the following steps: (1) The states of all sensors in the system converge to a uniform state, which is represented as:
[0013] in, This is the only equilibrium state of the system. The average value for each agent, i.e. ; Distributed networks reach consensus according to average consensus protocols. Continuous systems The discrete form is:
[0014] And transcribed as:
[0015] in, For uniformity gain and , P The state transition matrix is and , I It is the identity matrix. Therefore, in the discrete state, the discrete final convergent state satisfies the average consensus protocol as follows:
[0016] in, for Time of the first i The status of a sensor network.
[0017] Furthermore, the communication topology switching of the estimation system for the quantized distributed consensus protocol described in step S1 under the influence of denial-of-service includes the following steps: (1) The relationship between denial of service and channel packet loss rate is modeled using the signal-to-interference-plus-noise ratio (SINR) model. The signal error transmission probability (SER) is calculated as follows:
[0018] in, H It is the right-tailed function of the standard normal distribution, i.e. ; Channel success transmission probability The calculation is as follows:
[0019] in, The blocking variable introduced under the influence of network denial-of-service (DoS) represents whether the channel is blocked, and is expressed as follows:
[0020] (2) The constraint condition for ensuring that the final state of the sensor network remains consistent even when the topology is switched is:
[0021] From the formula It can be seen that there is A unique equilibrium state is reached when the states of all sensors converge uniformly. ,in Therefore, the limit of the product of the state transition matrices is .
[0022] Furthermore, the estimated network connectivity of the distributed consensus protocol estimation system described in step S1 includes the following steps: Pick The smallest non-zero eigenvalue of the Laplace matrix L, and at the same time express The network connectivity, and the lower bound of the average uniform convergence rate in the case of a strongly connected undirected network are: ,Right now:
[0023] in, For the image Medium sensor i The degree is obtained through the locality weighting method.
[0024] Furthermore, the specific operation method of step S2 is as follows: Minimum eigenvalue The lower bound representing the average uniform convergence rate, taking the reward The objective function is given by the output energy vector as Average expected return at each time point r The sum is given as follows:
[0025] in Indicates the expectation. This refers to time points; it is assumed that at each moment, the influencer has an energy limit. P The objective function for measuring the impact of denial-of-service influencer energy allocation strategies on the consistency convergence performance of a distributed system is as follows:
[0026]
[0027] That is, at the upper limit of energy P Under the constraints, the solution makes Parameters for obtaining the maximum value .
[0028] Further, step S3 includes the following steps: S31. Define the set of states at a given moment. This indicates the current topology connectivity, where, Indicates channel i Whether connectivity is maintained, when the channel is interrupted due to a denial-of-service condition. Otherwise, the channel will function normally. ; Define the set of states as:
[0029] Define the set of states of a Markov chain as:
[0030] The action set is:
[0031] express k The energy output of each channel is influenced by the constant time. Indicating influencers' views on the first i Energy injected by each channel; energy injected by influencers is categorized as follows: L Energy levels, energy levels It is equivalent to the Laplace matrix.
[0032] Define the Markov action set as:
[0033] S32. For the influencer, the desired topological connectivity is reduced, and they want to use as little energy as possible. Then at time k The reward is expressed as ,in, It is a weight parameter; k Indicates time; Influencers in Always choose actions to obtain the highest possible reward. Then, it enters the next state, where the influencer injects energy into a certain channel. By calculating the signal transmission error rate SER , k Time Channel i of SER Represented as Taking action is represented as Then the state transition probability matrix can be defined as follows:
[0034] The objective function is given by the output energy vector. The average expected total return r Give Since the influencer's energy has an upper limit, the optimal selection strategy is as follows:
[0035]
[0036] in, The upper limit of the energy of the influencer.
[0037] Furthermore, step S4 includes the following steps: S41. The objective function can be transformed into:
[0038] Solution of the objective function It satisfies the Bellman equation, as shown below:
[0039] in s It is the defined initial state. This represents the possible states under the new action; therefore, we can obtain... Q-learning Under this method, mapping Satisfies Bellman's equation. The updated expression is as follows:
[0040] in, This indicates the energy that can be taken under the new condition; S42. Use a model-free reinforcement learning method to obtain the optimal solution, and define the action value function:
[0041] And based on the following formula:
[0042] Transform the problem into a computation .
[0043] Furthermore, the aforementioned The specific calculation method is as follows: (1) First initialize All values are random, and then at each time step... k The status is s Based on the formula:
[0044] Select Action And by formula Receive relevant rewards ; (2) From the formula:
[0045] Calculate the state at the next time step; (3) Iterative calculation, and use the equation after each iteration.
[0046] renew Q Value, of which, This is the learning rate.
[0047] This invention also provides a distributed consistent fusion estimation denial-of-service attack control system. The system operates using the above-described method and includes the following modules: The basic model building module uses graph theory to construct the communication topology of the estimation system based on the distributed consensus protocol; it quantifies the consistency convergence state, the communication topology switching under the influence of denial-of-service, and the network connectivity. The objective function building module is used to construct an objective function that measures the impact of denial-of-service influencer energy allocation strategies on the consistency convergence performance of a distributed system. The decision model building module is used to build Markov decision models that suffer from network topology switching, decreased connectivity, and slowed convergence speed due to denial of service. The strategy output module is used for... Q-learning Reinforcement learning is used to iteratively solve the problem and select the optimal denial-of-service sequence under the constraint of finite energy.
[0048] The advantages of this invention are: (1) The network topology switching of this invention screens out the optimal denial-of-service impact sequence in the time domain for the distributed consistent fusion estimation system. By dynamically adjusting the strategy through the Markov decision model, the flexibility and effectiveness of the screening method are significantly improved. At the same time, by quantifying the network connectivity through the minimum non-zero eigenvalue, the influencer can prioritize destroying the key channel and maximize the delay of the system convergence speed. This provides a theoretical basis for studying denial-of-service defense against the system under low energy constraints.
[0049] (2) Adopt Q-learning The algorithm iteratively solves for the optimal strategy, avoiding the limitation of traditional methods that require global information. It achieves efficient decision-making under conditions of limited computing resources and information, while accelerating the convergence of the solution process through dynamic learning rate. Attached Figure Description
[0050] Figure 1 This is a flowchart illustrating the denial-of-service sequence filtering method based on distributed consistent fusion estimation according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the network graph structure and optimal weights under the influence of denial of service in an embodiment of the present invention; Figure 3 This is a schematic diagram of network topology switching under the influence of denial of service in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the convergence time of an embodiment of the present invention without any impact. Figure 5 This is a schematic diagram illustrating the convergence time under the influence of denial-of-service in an embodiment of the present invention; Figure 6 This is a schematic diagram of the learning process for the optimal strategy in an embodiment of the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0052] Example 1 This embodiment provides a method for filtering denial-of-service sequences using distributed consistent fusion estimation. The specific implementation process is as follows: Figure 1 As shown, it includes the following steps: S1. Construct the communication topology of the estimation system based on the distributed consensus protocol using graph theory methods; quantify its consensus convergence state, communication topology switching under the influence of denial-of-service, and network connectivity. The communication topology of the estimation system for constructing the distributed consensus protocol includes the following steps: (1) Represent the network topology as G ,sensor i scalar state at each time step Then in The time represents the initial state of the sensor network; the first... i The dynamic model of a first-order continuous system with one sensor is: ; in, For the first i The status values of each sensor, Representing the i Control input from one sensor, V The number of sensor network nodes. ; (2) The first in the system i The control input for each sensor is:
[0053] in, Representative and the i A set of neighboring nodes for which sensors exchange data. It is a sensor j To the sensor i The weight of the corresponding communication edge; Since the current state update of each sensor depends only on the state values of its own and neighboring sensors, the continuous system is represented as:
[0054] And transcribed as:
[0055] in, L For the image The Laplace matrix, by definition, always has zero eigenvalues, and the corresponding eigenvectors are... .
[0056] The estimation of the system consensus convergence state of the quantitative distributed consensus protocol includes the following steps: (1) The states of all sensors in the system converge to a uniform state, which is represented as:
[0057] in, This is the only equilibrium state of the system. The average value for each agent, i.e. ; Distributed networks reach consensus according to average consensus protocols. Continuous systems The discrete form is:
[0058] And transcribed as:
[0059] in, For uniformity gain and , P The state transition matrix is and , I It is the identity matrix. Therefore, in the discrete state, the discrete final convergent state satisfies the average consensus protocol as follows:
[0060] in, for Time of the first i The status of a sensor network.
[0061] The communication topology switching of the estimation system for the quantized distributed consensus protocol under the influence of denial-of-service includes the following steps: (1) The relationship between denial of service and channel packet loss rate is modeled using the signal-to-interference-plus-noise ratio (SINR) model. The signal error transmission probability (SER) is calculated as follows:
[0062] in, H It is the right-tailed function of the standard normal distribution, i.e. ; Channel success transmission probability The calculation is as follows:
[0063] in, The blocking variable introduced under the influence of network denial-of-service (DoS) represents whether the channel is blocked, and is expressed as follows:
[0064] (2) The constraint condition for ensuring that the final state of the sensor network remains consistent even when the topology is switched is:
[0065] From the formula It can be seen that there is A unique equilibrium state is reached when the states of all sensors converge uniformly. ,in Therefore, the limit of the product of the state transition matrices is .
[0066] In this embodiment, when the network topology and optimal weights are affected by denial of service, as follows: Figure 2 As shown, if the influencer affects edge 7-8 and edge 3-7 sequentially, the following will occur: Figure 3 The topology switching sequence is shown.
[0067] The method for estimating the network connectivity of a distributed consensus protocol includes the following steps: Pick The smallest non-zero eigenvalue of the Laplace matrix L, and at the same time express The network connectivity, and the lower bound of the average uniform convergence rate in the case of a strongly connected undirected network are: ,Right now:
[0068] in, For the image Medium sensor i The degree is obtained through the locality weighting method.
[0069] S2. Construct an objective function to measure the impact of denial-of-service influencer energy allocation strategies on the consensus convergence performance of a distributed system. In a distributed estimation system, the ability of each sensor state to reach consensus within a finite time is a key indicator. From the influencer's perspective, the influencer wants the sensor network to reach consensus for as long as possible. However, the influencer's own energy is finite. Therefore, the optimal denial-of-service strategy design problem can be transformed into a constrained optimization problem. The specific operation method is as follows: Minimum eigenvalue The lower bound representing the average uniform convergence rate, taking the reward The objective function is given by the output energy vector as Average expected return at each time point r The sum is given as follows:
[0070] in Indicates the expectation. This refers to time points; it is assumed that at each moment, the influencer has an energy limit. P The objective function for measuring the impact of denial-of-service influencer energy allocation strategies on the consistency convergence performance of a distributed system is as follows:
[0071]
[0072] That is, at the upper limit of energy P Under the constraints, the solution makes Parameters for obtaining the maximum value .
[0073] S3. Construct a Markov decision model for network topology switching caused by denial of service, resulting in decreased connectivity and slower convergence speed; the specific implementation includes the following steps: S31. Define the set of states at a given moment. This indicates the current topology connectivity, where, Indicates channel i Whether connectivity is maintained, when the channel is interrupted due to a denial-of-service condition. Otherwise, the channel will function normally. ; Define the set of states as:
[0074] Define the set of states of a Markov chain as:
[0075] The action set is:
[0076] express k The energy output of each channel is influenced by the constant time. Indicating influencers' views on the first i The energy injected by each channel; in practical wireless network systems, coarsely quantized energy is often transmitted, rather than energy of arbitrary values. In this model, to ensure that the Markov model's action set is finite, the energy injected by the influencer is classified into... LAt a certain energy level, when the difference between theoretical and practical results can be ignored, the energy level... It is equivalent to the Laplace matrix.
[0077] Define the Markov action set as:
[0078] S32. For the influencer, the desired topological connectivity is reduced, and they want to use as little energy as possible. Then at time k The reward is expressed as ,in, It is a weight parameter; k Indicates time; Influencers in Always choose actions to obtain the highest possible reward. Then, it enters the next state, where the influencer injects energy into a certain channel. By calculating the signal transmission error rate SER , k Time Channel i of SER Represented as Taking action is represented as Then the state transition probability matrix can be defined as follows:
[0079] The objective function is given by the output energy vector. The average expected total return r Give Since the influencer's energy has an upper limit, the optimal selection strategy is as follows:
[0080]
[0081] in, The upper limit of the energy of the influencer.
[0082] S4, based on Q-learning Reinforcement learning iteratively solves the problem, selecting the optimal denial-of-service sequence under finite energy constraints. The specific implementation includes the following steps: S41. The objective function can be transformed into:
[0083] Solution of the objective function It satisfies the Bellman equation, as shown below:
[0084] ins It is the defined initial state. This represents the possible states under the new action; therefore, we can obtain... Q-learning Under this method, mapping Satisfies Bellman's equation. The updated expression is as follows:
[0085] in, This represents the energy that can be taken in a new state. Directly solving the above equation to obtain the optimal solution is relatively difficult. In reinforcement learning, there are some well-known methods to solve this problem, such as value iteration and policy iteration. These methods are effective, but they require all the transition probability equations in the environment and the rewards of all states. However, the influencer may not be able to obtain all this information, and it requires a large amount of computation. Therefore, in order to deal with the limitations of information and computational power, model-free reinforcement learning methods are used to obtain the optimal solution.
[0086] S42. Use a model-free reinforcement learning method to obtain the optimal solution, and define the action value function:
[0087] And based on the following formula:
[0088] Transform the problem into a computation .
[0089] The The specific calculation method is as follows: (1) First initialize All values are random values, and then at each time step... k The status is s Based on the formula:
[0090] Select Action , and by formula Receive relevant rewards ; (2) From the formula:
[0091] Calculate the state at the next time step; (3) Iterative calculation, and use the equation after each iteration.
[0092] renew Q Value, of which, This is the learning rate.
[0093] In this embodiment, the sensor can learn the optimal policy online or offline, requiring a balance between time and decision-making. On one hand, learning can gradually approach the optimal policy through iteration, thereby improving policy performance. On the other hand, excessively long iteration steps can slow down the system execution speed. Therefore, Q-learning The algorithm's core idea is to try different methods to update knowledge, using only a small amount of information at each time step. As the learning process iterates, the influencer gains sufficient information and gradually converges to the optimal solution (the slowing-down target). Subsequent iterations are then assigned less weight. Therefore, it's necessary to design the learning rate. The influence of subsequent information on the decision-makers gradually decreases until it converges to the optimal value. The simulation will provide a specific formula for calculating the learning rate. ,in Indicates the influencer's state s Execute energy allocation strategy Number of times, z and b This is a configured value.
[0094] This embodiment also provides a simulation experiment, as detailed below: First set network noise With a system of 3 nodes, the topology and weights are shown in the Laplace matrix. L。
[0095]
[0096] The initial values of each node are Then the set of states is The action set is set to First, let's look at the impact on the consensus convergence of multi-agent networks, such as... Figure 4 Figure 5 shows the convergence speed of the system under two different conditions: no influence (convergence time of 10 steps) and influence (convergence time of 20 steps). This demonstrates that network influence slows down the convergence speed of a distributed system. Then, the Q-Learning algorithm is applied to construct a Markov decision model, assuming a decay factor... And derive the learning rate As the number of iterations increases, the system gathers more information about the environment, reducing random exploration and increasing the exploration rate. The calculation method is set as follows: By setting the learning rate and exploration rate as described above, states and actions that are accessed less frequently will be given higher weight in the learning process, and the learning process that optimally influences the policy will proceed as follows: Figure 6 As shown.
[0097] The initial value is set to 0, at time... k Corresponding action Q The value is based on the formula
[0098] Other Q The value remains unchanged after 1000 Monte Carlo simulation iterations. Q The values converged to the optimal value, as shown in the table below, where the bolded values represent the optimal values. Q value .
[0099]
[0100] As shown in the table above, in states (1,1,1), (0,1,1), and (1,0,1), the influencers tend to prioritize affecting channel 3 because channel 3 has the largest weight. In state (1,1,0), i.e., when channel 3 is broken, they prioritize affecting channel 2 because channel 2 has a greater weight than channel 1. Thus, it can be seen that the influencers will ultimately prioritize affecting the channel with the greater weight.
[0101] Example 2 It should be further explained that, based on the same inventive concept, this embodiment also provides a distributed consistent fusion estimation denial-of-service sequence filtering system. The system operates using the method described in Embodiment 1, and includes the following modules: The basic model building module uses graph theory to construct the communication topology of the estimation system based on the distributed consensus protocol; it quantifies the consistency convergence state, the communication topology switching under the influence of denial-of-service, and the network connectivity. The objective function building module is used to construct an objective function that measures the impact of denial-of-service influencer energy allocation strategies on the consistency convergence performance of a distributed system. The decision model building module is used to build Markov decision models that suffer from network topology switching, decreased connectivity, and slowed convergence speed due to denial of service. The strategy output module is used for... Q-learning Reinforcement learning is used to iteratively solve the problem and select the optimal denial-of-service sequence under the constraint of finite energy.
[0102] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for filtering denial-of-service sequences based on distributed consistent fusion estimation, characterized in that, Includes the following steps: S1. Construct the communication topology of the estimation system based on the distributed consensus protocol using graph theory methods; quantify its consensus convergence state, communication topology switching under the influence of denial-of-service, and network connectivity. S2. Construct an objective function to measure the impact of denial-of-service influencer energy allocation strategies on the consistency convergence performance of distributed systems. S3. Construct a Markov decision model for network topology switching caused by denial of service, resulting in decreased connectivity and slowed convergence speed. S4, based on Q-learning Reinforcement learning is used to iteratively solve the problem and select the optimal denial-of-service sequence under the constraint of finite energy.
2. The method for filtering denial-of-service sequences based on distributed consistent fusion estimation according to claim 1, characterized in that, The communication topology of the estimated system for constructing the distributed consensus protocol described in step S1 includes the following steps: (1) Represent the network topology as G ,sensor i scalar state at each time step Then in The time interval represents the initial state of the sensor network; the first... i The dynamic model of a first-order continuous system with one sensor is: ; in, For the first i The status values of each sensor, Representing the i Control input from one sensor, V The number of sensor network nodes. ; (2) The first in the system i The control input for each sensor is: in, Representative and the i A set of neighboring nodes for which sensors exchange data. It is a sensor j To the sensor i The weight of the corresponding communication edge; A continuous system is represented as: And transcribed as: in, L For the image The Laplacian matrix, and the corresponding eigenvectors are .
3. The method for filtering denial-of-service sequences based on distributed consistent fusion estimation according to claim 2, characterized in that, Step S1, which involves estimating the system consensus convergence state of the quantized distributed consensus protocol, includes the following steps: (1) The states of all sensors in the system converge to a uniform state, which is represented as: in, This is the only equilibrium state of the system. The average value for each agent, i.e. ; Distributed networks reach consensus according to average consensus protocols. Continuous systems The discrete form is: And transcribed as: in, For uniformity gain and , P The state transition matrix is and , I It is the identity matrix. Therefore, in the discrete state, the discrete final convergent state satisfies the average consensus protocol as follows: in, for Time of the first i The status of a sensor network.
4. The method for filtering denial-of-service sequences based on distributed consistent fusion estimation according to claim 3, characterized in that, The communication topology switching of the quantization distributed consensus protocol estimation system under the influence of denial-of-service, as described in step S1, includes the following steps: (1) The relationship between denial of service and channel packet loss rate is modeled using the signal-to-interference-plus-noise ratio (SINR) model. The signal error transmission probability (SER) is calculated as follows: in, H It is the right-tailed function of the standard normal distribution, i.e. ; Channel success transmission probability The calculation is as follows: in, The blocking variable introduced under the influence of network denial-of-service (DoS) represents whether the channel is blocked, and is expressed as follows: (2) The constraint condition for ensuring that the final state of the sensor network remains consistent even when the topology is switched is: From the formula It can be seen that there is A unique equilibrium state is reached when the states of all sensors converge uniformly. ,in Therefore, the limit of the product of the state transition matrices is .
5. The method for filtering denial-of-service sequences based on distributed consistent fusion estimation according to claim 4, characterized in that, Step S1, which involves estimating the network connectivity of a distributed consensus protocol system, includes the following steps: Pick The smallest non-zero eigenvalue of the Laplace matrix L, and at the same time express The network connectivity, and the lower bound of the average uniform convergence rate in the case of a strongly connected undirected network are: ,Right now: in, For the image Medium sensor i The degree is obtained through the locality weighting method.
6. The method for filtering denial-of-service sequences based on distributed consistent fusion estimation according to claim 5, characterized in that, The specific operation method of step S2 is as follows: Minimum eigenvalue The lower bound representing the average uniform convergence rate, taking the reward The objective function is given by the output energy vector as Average expected return at each time point r The sum is given as follows: in Indicates the expectation. This refers to time points; it is assumed that at each moment, the influencer has an energy limit. P The objective function for measuring the impact of denial-of-service influencer energy allocation strategies on the consistency convergence performance of a distributed system is as follows: That is, at the upper limit of energy P Under the constraints, the solution makes Parameters for obtaining the maximum value .
7. The method for filtering denial-of-service sequences based on distributed consistent fusion estimation according to claim 6, characterized in that, Step S3 includes the following steps: S31. Define the set of states at a given moment. This indicates the current topology connectivity, where, Indicates channel i Whether connectivity is maintained, when the channel is interrupted due to a denial-of-service condition. Otherwise, the channel will function normally. ; Define the set of states as: Define the set of states of a Markov chain as: The action set is: express k The energy output of each channel is influenced by the constant time. Indicating influencers' views on the first i Energy injected by each channel; energy injected by influencers is categorized as follows: L Energy levels, energy levels Equivalent to the Laplace matrix; Define the Markov action set as: S32. For the influencer, the desired topological connectivity is reduced, and they want to use as little energy as possible. Then at time k The reward is expressed as ,in, It is a weight parameter; k Indicates time; Influencers in Always choose actions to obtain the highest possible reward. Then, it enters the next state, where the influencer injects energy into a certain channel. By calculating the signal transmission error rate SER , k Time Channel i of SER Represented as Taking action is represented as Then the state transition probability matrix can be defined as follows: The objective function is given by the output energy vector. The average expected total return r Give Since the influencer's energy has an upper limit, the optimal selection strategy is as follows: in, The upper limit of the influencer's energy.
8. The method for filtering denial-of-service sequences based on distributed consistent fusion estimation according to claim 7, characterized in that, Step S4 includes the following steps: S41. The objective function can be transformed into: Solution of the objective function It satisfies the Bellman equation, as shown below: in s It is the defined initial state. This represents the possible states under the new action; therefore, we can obtain... Q-learning Under this method, mapping Satisfies Bellman's equation. The updated expression is as follows: in, This indicates the energy that can be taken under the new condition; S42. Use a model-free reinforcement learning method to obtain the optimal solution, and define the action value function: And based on the following formula: Transform the problem into a computation .
9. The method for filtering denial-of-service sequences based on distributed consistent fusion estimation according to claim 8, characterized in that, The The specific calculation method is as follows: (1) First initialize All values are random, and then at each time step... k The status is s Based on the formula: Select Action And by formula Receive relevant rewards ; (2) From the formula: Calculate the state at the next time step; (3) Iterative calculation, and use the equation after each iteration. renew Q Value, of which, This is the learning rate.
10. A denial-of-service attack control system based on distributed consistent fusion estimation, characterized in that, Includes the following modules: The basic model building module uses graph theory to construct the communication topology of the estimation system based on the distributed consensus protocol; it quantifies the consistency convergence state, the communication topology switching under the influence of denial-of-service, and the network connectivity. The objective function building module is used to construct an objective function that measures the impact of denial-of-service influencer energy allocation strategies on the consistency convergence performance of a distributed system. The decision model building module is used to build Markov decision models that suffer from network topology switching, decreased connectivity, and slowed convergence speed due to denial of service. The strategy output module is used for... Q-learning Reinforcement learning is used to iteratively solve the problem and select the optimal denial-of-service sequence under the constraint of finite energy.
Citation Information
Patent Citations
Distributed denial of service attack and defense method based on multi-agent reinforcement learning
CN120151030A