Wireless router complete machine one-stop test method and system, and electronic equipment

By constructing a one-stop testing method for wireless routers, utilizing Markov decision processes and deep reinforcement learning to generate traffic patterns, and combining generative adversarial networks and multi-agent cooperative attack systems, the problem of simulating real network environments and unknown attacks in wireless router testing is solved, enabling efficient and comprehensive evaluation of router performance and security capabilities.

CN121078437APending Publication Date: 2025-12-05SHENZHEN TWOWING TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511060714.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing wireless router testing methods cannot effectively simulate real network environments and attack scenarios, are difficult to detect long-term performance degradation and resource leaks, lack the ability to defend against unknown attack patterns, and are inefficient in testing and miss critical fault scenarios.

Method used

We employ Markov decision processes to generate typical traffic patterns, combine deep reinforcement learning algorithms and generative adversarial networks to create attack patterns, construct a multi-agent collaborative attack system, achieve collaborative decision-making based on a shared value network, dynamically adjust test intensity, and build a multi-dimensional evaluation index system.

Benefits of technology

Significantly improves test coverage, identifies potential stability issues and security risks in advance, improves product quality and safety, and reduces user complaint rates and equipment return rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121078437A_ABST
    Figure CN121078437A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wireless network testing, and discloses a wireless router complete machine one-stop testing method and system and electronic device.The wireless router complete machine one-stop testing method comprises the steps that a typical network application characteristic flow mode library is constructed, and different flow modes are dynamically combined through the Markov decision process; creating a novel attack mode by using the generative adversarial network; constructing a multi-agent cooperative attack system; dynamically adjusting the test intensity according to the real-time response of the router; constructing a multi-dimensional evaluation index system and analyzing a test result; according to the invention, by combining artificial intelligence and network security technologies, the performance and security protection capability of the wireless router in a complex and changeable network environment and attack scene are comprehensively tested, potential stability problems and security holes are found in advance, and the product quality and the user experience are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of wireless network testing, more particularly, it relates to a wireless router whole machine one-stop testing method, system and electronic equipment. BACKGROUND

[0002] As the core device of modern networks, the performance and security of wireless routers directly affect the network experience and information security of users. With the diversification of network applications and the complexity of network threats, higher requirements are placed on the testing of wireless routers.

[0003] However, the traditional wireless router testing method has many technical limitations. The existing testing method mainly uses a fixed traffic model and a predefined attack mode. This static testing method cannot effectively simulate the dynamic and variable traffic characteristics and the constantly evolving attack methods in the actual network environment. The traditional testing method cannot detect the performance degradation and resource leakage of the router in the long-term running state. The existing testing usually focuses on short-term performance, and lacks effective evaluation means for the long-term stability of the device. The existing testing system is insufficient in evaluating the defense capability of the router against unknown attack patterns. Artificially designed test cases are often limited to known attack methods and cannot predict and simulate new attack patterns emerging in the field of network security.

[0004] The current testing method cannot effectively simulate the distributed attack mode of multiple points and multiple roles cooperating. In actual network threat scenarios, attackers usually launch attacks from multiple network locations in coordination, and the traditional single-source traffic attack test cannot reflect this complex attack scenario, resulting in a gap between the test results and the security threats in the actual network environment. At the same time, under the constraints of limited testing resources and time, the existing testing method cannot efficiently find the key problems that are most likely to cause router failure. The test often uses exhaustive or random methods, lacks intelligent testing strategy guidance, resulting in low testing efficiency and possible omission of critical failure scenarios.

[0005] Therefore, an innovative whole machine one-stop testing method is needed to comprehensively and efficiently evaluate the performance and security protection capability of wireless routers in various complex scenarios. SUMMARY

[0006] The present application provides a wireless router whole machine one-stop testing method, system and electronic equipment, which solves the technical problem of being unable to effectively simulate real network environment and attack scenarios in related technologies.

[0007] The present application provides a wireless router whole machine one-stop testing method, which includes the following steps: A library of traffic patterns characteristic of typical network applications is constructed. Different traffic patterns are dynamically combined using Markov decision process. Attack traffic patterns are generated by deep reinforcement learning algorithm and injected through software-defined networking technology. Based on traffic patterns, novel attack patterns are created using adversarial generative networks, variational autoencoders are applied to generate variant attack samples, and knowledge graphs are combined to construct composite attack chains. Based on the generated attack patterns, a multi-agent collaborative attack system is constructed to realize a collaborative decision-making mechanism based on a shared value network and automatically determine the optimal attack location distribution according to the test target. During the coordinated attack execution, the test intensity is dynamically adjusted according to the router's real-time response to realize the attack and defense game model and build a resource monitoring probe to collect performance data. Based on the collected performance data, a multi-dimensional evaluation index system is constructed to achieve automatic performance bottleneck analysis and build a performance prediction model.

[0008] In a preferred embodiment, the step of constructing a typical network application characteristic traffic pattern library specifically includes: By analyzing traffic data in real network environments, characteristic parameters of different network applications are extracted; Statistical modeling is performed on the traffic characteristics of each network application to obtain the probability distribution function of the characteristic parameters; Build a traffic feature library that includes video streaming, online games, web browsing, file transfer, VoIP calls, and common network applications.

[0009] In a preferred embodiment, the step of generating attack traffic patterns using a deep reinforcement learning algorithm adopts a dual-deep Q-network structure. The input of the dual-deep Q-network is the current state of the router, including performance indicators, and the output is the attack action, including attack type, attack parameters, and attack strength. It also employs priority experience replay technology to improve learning efficiency.

[0010] In a preferred embodiment, the step of generating mutation attack samples using a variational autoencoder specifically includes: The encoder network maps the input attack samples to the distribution parameters of the latent space; Latent variables are sampled from the distribution; the decoder network reconstructs the latent variables into new attack samples; The attack pattern difference assessment module is designed to calculate the difference between the generated sample and the known attack sample, and only the generated sample with a difference exceeding a preset threshold is retained.

[0011] In a preferred embodiment, the step of constructing a multi-agent cooperative attack system includes at least four types of agents: Probe agent: responsible for detecting the vulnerabilities of the network and the router; Resource consumption agent: responsible for consuming the computing and network resources of the router; Service denial attack agent: responsible for performing various denial of service attacks; Coordination agent: responsible for coordinating the actions of various attack agents.

[0012] In a preferred embodiment, the step of dynamically adjusting the test intensity according to the real-time response of the router specifically comprises: Defining a router state vector containing multiple performance indicators, while defining a normal state interval, a warning state interval and a dangerous state interval; Defining a target state within the warning state interval; Calculating the gap between the current state and the target state; Based on the proportional-integral-derivative control algorithm, calculating the test intensity adjustment amount; Updating the test intensity and introducing a security protection mechanism.

[0013] In a preferred embodiment, the step of implementing the attack and defense game model specifically comprises: Building a two-person zero-sum game model, including the strategy space of the attacker, the strategy space of the defender and their respective utility functions; using the Bayesian game framework, introducing the belief distribution of the defense strategy; Building an attack strategy mutation algorithm, including four links of defense detection, defense type identification, attack strategy mutation and mutation effect evaluation.

[0014] In a preferred embodiment, the step of constructing a multi-dimensional evaluation index system specifically comprises: Building an evaluation index system from performance stability, security protection, functional correctness and multiple dimensions; Using a radar chart to represent the comprehensive evaluation results of the router; Calculating the comprehensive score and evaluating the performance boundary under different working conditions.

[0015] In a preferred embodiment, the step of constructing a performance prediction model specifically comprises: Performing feature engineering to extract feature vectors and define target vectors; Evaluating multiple prediction models and performing hyperparameter tuning; Calculating the prediction accuracy index and analyzing the residual distribution; Quantifying the uncertainty of the prediction results and using active learning technology to guide the test.

[0016] In a preferred embodiment, a wireless router whole machine one-stop test system for performing a wireless router whole machine one-stop test method, comprising: a traffic pattern generation module for generating traffic patterns simulating real network environments, including a traffic pattern library construction unit, a Markov decision unit, a reinforcement learning unit, and a traffic injection unit; an attack pattern creation module for creating unknown attack patterns, including a generative adversarial network unit, a variational autoencoder unit, and a cross-protocol layer attack combination unit; a coordinated attack simulation module for simulating distributed coordinated attacks, including a multi-agent coordination framework unit, a shared value network unit, and an attack location optimization unit; a test strategy adjustment module for adaptively adjusting test strategies, including a strength adjustment unit, a game model unit, and a resource monitoring unit; a result analysis and evaluation module for analyzing test results and evaluating router performance, including an attack classification unit, an effect analysis unit, a defense evaluation unit, and a report generation unit.

[0017] The present application has the following advantages: Improved test coverage: Compared with traditional fixed traffic testing, the present application can discover more potential stability problems, significantly improve test coverage, and can discover serious performance degradation problems in advance, providing more comprehensive protection for product quality.

[0018] Prospective safety evaluation: By testing the defense capabilities of unknown attacks on the device, the present application can discover potential safety risks in advance, advance the discovery time of safety vulnerabilities to the product development stage, and significantly reduce the risk of market safety incidents.

[0019] Real environment simulation: The present application uses dynamic adjustable traffic generation technology, which can adaptively adjust the test strategy according to the router state, better meet the unpredictability in real use environment, and discover problems that the product may face in actual application scenarios in advance.

[0020] Test efficiency improvement: Through intelligent testing scheme, the present application can efficiently discover key problems that are most likely to cause router failure under limited test resources and time constraints, saving test time and resource cost.

[0021] Product quality assurance: The present application comprehensively evaluates the performance, resource management, and security protection capabilities of wireless routers when facing various normal and malicious traffic, significantly improves the reliability, stability, and security of the product, and reduces user complaint rate and equipment repair rate. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is a flowchart of a one-stop test method for a wireless router of the present application. DETAILED DESCRIPTION

[0023] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that discussions of these implementations are merely provided to enable those skilled in the art to better understand so as to best use the subject matter described herein, and variations of elements, functions, and arrangements of elements can be substituted for those illustrated and described herein without departing from the scope of the present specification. Various examples can omit, substitute, or add various procedures or components as appropriate, and the descriptions and representations of various examples are used merely to provide illustrative examples for the principles described herein. Additionally, some of the examples described herein are directed to what could be deemed to be fabricated implementations.

[0024] A wireless router whole machine one station test method is disclosed in at least one embodiment of the present application, as shown in Figure 1 The method comprises the following steps: Step 1, construct a typical network application characteristic traffic pattern library, adopt Markov decision process to dynamically combine different traffic patterns, apply deep reinforcement learning algorithm to generate attack traffic patterns, and inject variable traffic through software defined network technology; The method mainly comprises the following sub-steps: Step 1.1, construct a typical network application characteristic traffic pattern library; By analyzing a large amount of traffic data in a real network environment, the characteristic parameters of different network applications are extracted, and a typical network application characteristic traffic pattern library is constructed.

[0025] The pattern library contains traffic characteristics of common network applications such as video streaming, online games, web browsing, file transfer, VoIP calls, etc. The traffic characteristics of each application include but are not limited to: packet size distribution, inter-packet interval distribution, traffic burstiness parameters, protocol type distribution, TCP / UDP port distribution, etc.

[0026] To ensure the authenticity of the traffic patterns, statistical modeling is performed on the traffic characteristics of each network application to obtain the probability distribution function of the characteristic parameters. For example, for the packet size distribution, the following mathematical model can be established: Where P(s) represents the probability of packet size s; n represents the number of components of the mixed distribution; w i represents the weight of the i-th distribution component, which satisfies f i (s,μ i ,σ i ) represents the i-th probability distribution function; μ i ,σ i respectively represent the mean and standard deviation of the i-th distribution.

[0027] The finally constructed traffic pattern library contains m typical network application traffic patterns, denoted as F={F1,F2,...,F m}, where m represents the total number of typical network application types included in the traffic pattern library; F represents the entire set of the traffic pattern library; F1, F2, F m These represent the traffic patterns of network applications type 1, 2, and m, respectively; each traffic pattern F j It is represented by a set of feature parameter vectors, which include network traffic characteristics such as packet size distribution, inter-packet interval distribution, traffic burstiness parameters, protocol type distribution, and TCP / UDP port distribution.

[0028] Step 1.2: Employ a Markov decision process-based dynamic combination flow model; A Markov decision process is defined as a quintuple (S, A, P, R, γ), where: S is the state space, representing the set of possible states in the network environment; A is the action space, representing the combination of flow patterns that can be selected in each state; P:S×A×S→[0,1] is the state transition probability function, which represents the probability of transitioning to state s′ after performing action a from the current state s; It is a reward function that provides an immediate reward based on the current state and the action performed. γ∈[0,1] is a discount factor used to balance immediate rewards and future rewards.

[0029] In practical implementation, state s∈S represents the current load and traffic characteristics of the network environment; action a∈A represents which traffic patterns and their intensity are selected; transition probability P is set according to the traffic change pattern in the actual network environment; reward function R is designed as a measure of traffic authenticity and diversity.

[0030] The optimal combination strategy for traffic patterns can be obtained by solving the following optimal strategy equation: π * (s)=argmax a∈A [R(s,a)+γ∑ s′∈S P(s′|s,a)V * (s′)]; Where, π * (s) represents the optimal strategy in state s, i.e., which combination of traffic patterns is optimal; argmax a∈A This represents selecting the action 'a' from all possible actions 'a' that maximizes the subsequent expression; R(s,a) represents the immediate reward obtained by performing action 'a' in state s, reflecting the authenticity and diversity of the current flow combination; γ is a discount factor, ranging from [0,1], used to balance the importance of immediate rewards and future rewards; P(s′|s,a) represents the probability of transitioning to state s′ after performing action 'a' in state s; V *(s') is the optimal value function value of state s';∑ s′∈S represents the summation over all possible next states s'.

[0031] where V * (s) is the optimal value function, representing the maximum expected return that can be achieved starting from state s and following the optimal policy: V * (s) = max a∈A [R(s, a) + γ∑ s′∈S P(s'|s, a)V * (s')]; where R(s, a) is the immediate reward received after taking action a in state s. where V * (s) represents the optimal value of state s; max a∈A represents the selection of the action a that maximizes the following expression among all possible actions; other parameters have the same meaning as above.

[0032] This optimal policy ensures that the generated traffic pattern combination not only conforms to the characteristics of the real network environment, but also fully covers various traffic scenarios.

[0033] Step 1.3, applying a deep reinforcement learning algorithm to generate attack traffic patterns; On the basis of generating normal traffic patterns, according to another embodiment of the present application, a deep reinforcement learning (DRL) algorithm is further applied to generate attack traffic patterns. This algorithm takes the degree of router performance degradation as a reward signal, and continuously learns and optimizes the attack strategy through interaction with the environment.

[0034] The deep reinforcement learning model adopts a DoubleDQN structure, consisting of a target network and an evaluation network. The input of the model is the current state s t of the router, including performance indicators such as CPU utilization, memory occupation, and processing delay; the output is the attack action a t , including attack type, attack parameter, and attack strength, etc. The network parameters are optimized by minimizing the following loss function: where θ represents the parameters of the evaluation network, used to calculate the Q value of the current state-action pair; θ - represents the parameters of the target network, which has a lower update frequency and is used to stabilize the learning process; D represents the experience replay buffer, which stores state transition samples (s t , a t , r t , s t+1 ) in the interaction process; s t represents the current state, containing performance indicators such as CPU utilization and memory occupation of the router; a trepresents the attack action performed at the current time, including the attack type, parameters and intensity; r t represents the attack action a t obtained reward after, positively correlated with the degree of router performance degradation; s t+1 represents the next state after performing the action; γ is the discount factor, the value range is [0, 1], used to balance the importance of immediate reward and future reward; Q(s t ,a t ; θ) represents the value estimate of the current state-action pair calculated using the evaluation network; Q(s t+1 ,a; θ - ) represents the value estimate of each action in the next state calculated using the target network; argmax a represents the action that maximizes the Q value; represents the mathematical expectation, that is, the average of the samples in the experience replay buffer.

[0035] In order to improve the learning efficiency, the application adopts the Prioritized Experience Replay (Prioritized Experience Replay) technology, and gives different sampling weights according to the time difference (TD) error of the sample: Where, p i represents the sampling probability of sample i, that is, the probability of selecting the sample for training from the experience replay buffer; δ i represents the TD error of sample i (time difference error), calculated as the actual obtained reward plus the estimated value of the next state, minus the estimated value of the current state-action pair, reflecting the gap between the predicted value and the target value; |δ i | represents the absolute value of the TD error, the greater the error, the greater the information quantity of the sample, and the more valuable it is to learn; α is a hyperparameter that controls the priority degree, the value range is usually [0, 1], when α = 0, it is equivalent to uniform sampling, when α = 1, it is completely sampled according to the priority of the TD error; ∑ j |δ j | α represents the sum of the αth power of the absolute value of the TD error of all samples, used to normalize the sampling probability, ensuring that the sum of the sampling probabilities of all samples is 1.

[0036] Through the above deep reinforcement learning algorithm, the system can automatically discover and generate the traffic pattern that has the most attack effect on a specific router.

[0037] Step 1.4, inject variable traffic through software defined network technology; According to one embodiment of the present application, the step utilizes software-defined network (SDN) technology to inject the generated traffic pattern into the test network. The SDN controller controls the network devices through the OpenFlow protocol to achieve precise traffic injection and control.

[0038] It should be understood that the software-defined network architecture includes three layers: the application layer, the control layer, and the data layer. The application layer implements traffic generation logic; the control layer is responsible for converting traffic rules and issuing them to the data layer; and the data layer performs specific traffic forwarding and processing.

[0039] The traffic injection process is achieved through the following steps: Convert the generated traffic pattern into SDN flow table rules, including MatchFields and ActionSet; the controller issues the flow table rules to the data forwarding device through the OpenFlow protocol; The data forwarding device generates and forwards network traffic that meets the specific pattern according to the flow table rules; Real-time monitoring of traffic generation to ensure that traffic characteristics meet the expected pattern.

[0040] Through the flexible control capability of SDN technology, the system can adjust the traffic injection strategy in real time, simulate the dynamic changes in the network environment, and provide a real and complex traffic environment for router testing.

[0041] Step 2, based on the traffic pattern, use the generative adversarial network to create new attack patterns, use the variational autoencoder to generate mutated attack samples, and combine the knowledge graph to build a composite attack chain; The main steps include the following sub-steps: Step 2.1, use the generative adversarial network to create new attack patterns; The present application uses the generative adversarial network (GAN) technology to create new attack patterns. The GAN model consists of two neural networks: the generator (Generator) and the discriminator (Discriminator). The generator is responsible for creating new attack patterns, and the discriminator is responsible for evaluating the attack effect.

[0042] The generator network G receives a random noise vector z ~ p z (z) as input and outputs the attack pattern G(z); the discriminator network D receives the attack pattern as input and outputs a probability value D(x), representing the probability that the input x is a real attack (rather than a generated attack). The training process of GAN is a game of opposition between the generator and the discriminator, and the objective function is: where min G max DThe generator G tries to minimize the objective function, while the discriminator D tries to maximize the objective function, forming an adversarial game relationship; V(D, G) represents the value function between the generator and the discriminator; p data (x) represents the distribution of the real attack mode; p z (z) represents the prior distribution of the noise vector, which is usually uniform distribution or Gaussian distribution; represents the expectation of the real attack mode distribution; represents the expectation of the noise vector distribution; logD(x) represents the logarithmic probability output of the discriminator for the real attack mode; log(1-D(G(z))) represents the logarithmic probability output of the discriminator for the generated attack mode.

[0043] In order to improve the quality and diversity of the generated attack mode, the improved (WassersteinGAN with Gradient Penalty, WGANGP) algorithm is adopted in the present application, and the objective function is modified as: Among them, represents the average score of the discriminator for the real attack mode; represents the average score of the discriminator for the generated attack mode; λ is the gradient penalty coefficient, used to control the weight of the gradient penalty term; represents the uniform sampling distribution between the real sample and the generated sample; represents the gradient of the discriminator for the sampling point ; represents the L2 norm of the gradient; represents the square of the difference between the gradient norm and 1, used to punish the case where the gradient deviates from 1.

[0044] In addition, in order to ensure that the generated attack mode conforms to the network protocol specification, the present application introduces a constraint condition specific to the network protocol, which is realized by adding a protocol compliance checking module in the discriminator: D(x)=D adv (x)·D protocol (x); Among them, D(x) represents the overall score of the discriminator for the input x; D adv (x) evaluates the attack effect and outputs the score of attack effectiveness; D protocol (x) evaluates the protocol compliance and outputs the score of protocol compliance, ensuring that the generated attack mode conforms to the network protocol specification.

[0045] Step 2.2, apply the variational autoencoder to generate the mutation attack sample; According to another embodiment of the present application, the step generates an adversarial sample mutant using a variational autoencoder (VAE) technique to realize automatic mutation of the attack mode. The VAE is composed of an encoder and a decoder, and can learn the latent representation of data and generate new samples.

[0046] The mathematical model of the variational autoencoder is defined as follows: Encoder network q φ (z|x) maps the input attack sample x to the distribution parameters μ and σ of the latent space; The latent variable z is sampled from the distribution . Decoder network p θ (x|z) reconstructs the latent variable z into a new attack sample

[0047] The optimization objective function of the VAE is: where, is the loss function of the VAE, which depends on the encoder parameters φ, the decoder parameters θ, and the input sample x; represents the expectation of z on the distribution q φ (z|x) generated by the encoder; logp θ (x|z) is the log-likelihood of the decoder reconstructing the original input, which measures the reconstruction quality; D KL (q φ (z|x)||p(z)) is the KL divergence between the encoder distribution and the prior distribution, which serves as a regularization term; β is a weighting factor that controls the proportion of reconstruction accuracy and distribution regularization, and a larger β value will make the model pay more attention to generating latent variables that conform to the prior distribution; p(z) is the prior distribution of the latent variable, which is usually set as a standard normal distribution q φ (z|x) is the posterior distribution output by the encoder network, which is usually a Gaussian distribution p θ (x|z) is the conditional distribution defined by the decoder network, which represents the probability of generating a sample x given the latent variable z.

[0048] By interpolating or adding directional perturbations in the latent space, the VAE can generate new attack samples with specific attributes. To enhance the directionality of attack mutation, the present application uses a conditional variational autoencoder (CVAE) and introduces an attack effect label y as a condition: where y is the attack effect label, which is used to guide the generation of samples with specific attack effects; q φ(z|x,y) is a conditional encoder network that maps the input sample x and label y together to the latent space; p θ (x|z,y) is a conditional decoder network that generates sample x based on latent variable z and label y; p(z|y) is the conditional prior distribution of latent variable z given label y; the meanings of other parameters are the same as those of the basic VAE.

[0049] This application also includes an attack pattern difference assessment module for calculating the difference between generated samples and known attack samples: Diff(x gen ,x known )=α·d struct (x gen ,x known )+(1-α)·d func (x gen ,x known ); Wherein, Diff(x) gen ,x known ) is the generated sample x gen Compared with known attack sample x known The overall degree of difference between them; d struct (x gen ,x known ) Measure the structural similarity of attacks by calculating the distance between two attack samples in terms of structural features; d func (x gen ,x known The similarity of attack functions is measured to evaluate the difference in functional effect between two attack samples; α is a weight coefficient that balances structural similarity and functional similarity, and its value ranges from [0,1]. When α is close to 1, the difference evaluation focuses more on structural differences; when α is close to 0, it focuses more on functional differences.

[0050] Only generated samples with a difference exceeding a preset threshold are retained, ensuring the diversity of the attack sample library.

[0051] Step 2.3: Construct a composite attack chain by combining the knowledge graph; According to another embodiment of this application, this step constructs a complex attack chain through network security knowledge graph technology, forming a multi-stage, multi-dimensional complex attack scenario.

[0052] The entities and relations in a knowledge graph are defined as follows: The entity set E includes attack types, attack tools, vulnerability types, etc. The set of relationships R describes the associations between entities, such as "utilize", "trigger", "bypass", etc. The set of triples T represents a complete set of knowledge facts.

[0053] The composite attack chain generation algorithm is based on random walk and heuristic search strategy, and mainly includes the following steps: From the initial attack node e start Start; According to the relationship weight w ij Calculate the transition probability According to the transition probability, the next attack node is selected, and the partial attack path path = [e start ,...,e current ] is constructed. Apply heuristic function: h(path) = a severity(path) + b diversity(path) + g concealment(path) evaluate the current path, wherein a, b, g are weight coefficients. Repeat step 2.4 until the target node or the maximum path length is reached.

[0054] In order to improve the practicability of the composite attack chain, the timing relationship constraint is introduced to ensure the timing rationality of the attack steps: Wherein, valid(e i →e j ) indicates whether the transition from attack step e i To e j Is valid; t(e) indicates the time stamp of attack step e, which ensures that the attack is executed in time sequence; pre(e j ) indicates the precondition set of attack step e j , that is, the environment state required to execute the step; effects(path[0:i]) indicates the effect set after executing the first i steps of the attack path; path[0:i] indicates all attack steps in the attack path from index 0 to index i. Indicates that the set contains the relationship, and requires that all preconditions of attack step e j All preconditions of attack step e j Have been met by the previous attack steps.

[0055] Step 3, according to the generated attack mode, build a multi-agent cooperative attack system, realize the cooperative decision mechanism based on shared value network, and automatically determine the optimal attack position distribution according to the test target; Mainly includes the following sub-steps: Step 3.1, build a multi-agent cooperative attack system; According to one embodiment of the present application, this step constructs a distributed attack coordination framework based on multi-agent reinforcement learning, including intelligent agents with different functional roles such as detectors, resource consumers, denial-of-service attackers, etc.

[0056] Each agent i is composed of the following components: Observation space O i : the set of environment states that the agent can observe; Action space A i : the set of actions that the agent can perform; State transition function T i : O i × A i → O i : describes the probability of state transition after action execution; Reward function evaluates the effect of action execution; Policy function π i : O i → A i : decides what action to perform under a specific observation state.

[0057] In a multi-agent system, state transition and rewards are not only dependent on the actions of a single agent, but also influenced by the behaviors of other agents, forming a complex interdependent relationship. The system state transition can be expressed as: P(o′1,o′2,...,o′ n | o1,o2,...,o n ,a1,a2,...,a n ); where o i represents the current observation state of agent i, i.e. the environment information that agent i can perceive; o′ i represents the next observation state of agent i, i.e. the new environment state that agent i can observe after executing the action; a i represents the action performed by agent i, i.e. the operation selected by agent i based on the current observation state; P(o′1,o′2,...,o′ n | o1,o2,...,o n ,a1,a2,...,a n ) represents the probability of the system transitioning to the new observation state (o′1,o′2,...,o′ n ) after all agents perform actions (a1,a2,...,a n ) under the current observation states (o1,o2,...,o n ).

[0058] In a specific implementation, different types of agents are defined according to functional roles: Probe agent: responsible for probing the network and router vulnerabilities, observation space includes network topology information, open ports, service types, etc.; action space includes various probing techniques such as port scanning, service identification, etc.

[0059] Resource consumption agent: responsible for consuming router computing and network resources, observation space includes router resource utilization, response time, etc.; action space includes sending specific types of resource-intensive requests.

[0060] Service denial attack agent: responsible for performing various denial of service attacks, observation space includes network traffic state, router response state, etc.; action space includes various DOS attack techniques.

[0061] Coordination agent: responsible for coordinating the actions of various attack agents, observation space is the global state; action space is the action instruction to other agents.

[0062] Through the above multi-agent framework, the system can simulate complex collaborative attack scenarios and test the defense capabilities of routers when facing multi-point, multi-role attacks.

[0063] Step 3.2, implement a collaborative decision-making mechanism based on a shared value network; According to another embodiment of the present application, this step implements a collaborative decision-making mechanism based on a shared value network (Shared Value Network), enabling each attack point to coordinate action strategies based on global state.

[0064] The shared value network architecture includes the following components: Local observation encoder E i : encodes the local observation o i of agent i into a feature vector h i ; Communication channel C: aggregates feature vectors of each agent into a global state representation g; Value network V: evaluates the value of the global state, outputs the global value estimate v; Policy network P i : generates a policy π i for agent i based on local observation and global state.

[0065] The decision-making process of agent i is represented as: π i = P i (o i , g), where g = C(h1, h2,..., h n ), h i = Ei (o i ); where π i represents the policy function of agent i, which determines the action P i that the agent will take i represents the local observation of agent i, i.e., the environment information that the agent can perceive g represents the global state representation, which is aggregated from the feature vectors of all agents C represents the communication channel function, which is used to aggregate the feature vectors of all agents h i represents the feature vector of agent i, which is the encoding result of the local observation E i represents the local observation encoder of agent i, which converts the original observation into a feature vector The system adopts the paradigm of centralized training and decentralized execution. In the training phase, the global value network is updated by temporal difference (TD) learning to minimize the following loss function: where L V represents the loss function of the value network represents the mathematical expectation; r t represents the immediate reward obtained at the current time t γ represents the discount factor, which is used to balance the importance of immediate rewards and future rewards, and its value range is [0, 1]; V(g t ) represents the value estimate of the current global state g t ; V(g t+1 ) represents the value estimate of the next global state g t+1 .

[0066] The policy network is updated by the policy gradient method to maximize the expected cumulative reward: where represents the gradient of the policy objective function J with respect to the parameter θ i ; θ i represents the policy network parameter of agent i represents the gradient of the log probability of the policy function with respect to the parameter θ i ; π i (a i |o i , g) represents the probability of agent i selecting action a i under observation o i and global state g; (r t + γV(g t+1 ) - V(g t )) represents the advantage function, which is used to evaluate the relative value of the action.

[0067] To enhance the collaboration ability between agents, the application introduces a differentiated reward mechanism, including: Global reward r g : based on the overall attack effect, shared by all agents; Local reward r i : based on the individual contribution of agent i; Collaboration reward r c : based on the degree of collaboration between agents.

[0068] The total reward of agent i is calculated as: Where, represents the total reward obtained by agent i, α represents the weight coefficient of global reward, used to adjust the importance of global goal; β represents the weight coefficient of local reward, used to adjust the importance of individual contribution; γ represents the weight coefficient of collaboration reward, used to adjust the importance of collaboration behavior.

[0069] Step 3.3, automatically determine the optimal attack position distribution according to the test target; According to another embodiment of the application, this step constructs an attack position distribution algorithm according to the test target and network topology, and automatically determines the optimal deployment position of the attack agent.

[0070] First, perform topology analysis on the test network to construct a network topology graph G=(V,E), where the vertex set V represents the network nodes and the edge set E represents the network connections. The node importance evaluation is based on the following indicators: Degree Centrality: the number of connections of node v, denoted as deg(v) represents the degree of node v, i.e. the number of edges directly connected to node v; |V| represents the total number of nodes in the network; |V|-1 represents the number of other nodes excluding node v; C D (v) value range is [0,1], the larger the value, the more connected the node is, and the more important it is in the network.

[0071] Betweenness Centrality: the number of times node v is located on the shortest path between other node pairs, denoted as Where, σ st represents the total number of shortest paths from node s to node t; σ st (v) represents the number of shortest paths from node s to node t passing through node v; ∑ s≠v≠t represents the sum of all node pairs (s,t) that do not contain node v; C B (v) value, the stronger the control ability of the node in the network information flow.

[0072] Closeness Centrality: the average reciprocal of the distance from node v to all other nodes, denoted as where d(v, u) represents the shortest distance (hop count) from node v to node u; ∑ u≠v d(v, u) represents the sum of distances from node v to all other nodes; |V| - 1 represents the number of other nodes excluding node v; C C (v) value, the shorter the average distance from the node to other nodes in the network, the higher the information propagation efficiency.

[0073] Considering the above indicators, the comprehensive importance score of the node is calculated: I(v) = w1·C D (v) + w2·C B (v) + w3·C C (v); where w1, w2, w3 are weight coefficients of degree centrality, betweenness centrality and closeness centrality, respectively, used to adjust the importance of each indicator in the comprehensive score; w1 + w2 + w3 = 1 to ensure weight normalization; I(v) represents the comprehensive importance score of node v, and the higher the value, the more critical the node in the network According to the attack type and test target, the application defines the position fitness function F i (v) of different types of attack agents. For example, for probing agents, it is more suitable to be deployed near the target; for distributed denial of service attacks, it needs to be deployed dispersedly at the edge of the network.

[0074] Finally, by solving the following optimization problem, the optimal deployment position of n attack agents is determined: where P represents the selected attack position set, which is a subset of the vertex set V; |P| = n represents selecting exactly n positions to deploy attack agents; v i represents the deployment position of the i-th attack agent; F i (v i ) represents the fitness value of the i-th attack agent at position v i ; represents the total fitness value of all attack agent position distributions; max represents finding the position combination that maximizes the total fitness.

[0075] Meanwhile, certain constraints such as minimum distance between attack points, coverage range, etc. are met.

[0076] Through the above algorithm, the system can automatically determine the optimal attack position distribution according to the test target, maximize the attack effect, and thus comprehensively test the defense capability of the router.

[0077] Step 4, during the cooperative attack execution process, the test intensity is dynamically adjusted according to the real-time response of the router, an attack and defense game model is realized, and a resource monitoring probe is constructed to collect performance data; The method mainly includes the following sub-steps: Step 4.1, dynamically adjusting the test intensity according to the real-time response of the router; According to an embodiment of the present application, this step constructs an adaptive test intensity adjustment algorithm, dynamically adjusts the test intensity according to the real-time response state of the router, and gradually advances to the limit state but does not exceed the collapse point.

[0078] Define the router state vector Wherein, , respectively represent the 1st, 2nd, mth performance indicators, and m represents the number of performance indicators, such as CPU utilization, memory occupation, throughput, delay, etc.

[0079] At the same time, define the normal state interval [s min ,s safe ], the alert state interval [s safe ,s critical ] and the dangerous state interval [s critical ,s max ]. Wherein, s t represents the router state vector at time point t; represents the value of the ith performance indicator at time point t; m represents the total number of monitored performance indicators; s min represents the minimum normal value of each indicator; s safe represents the safety threshold of each indicator, which exceeds the value to enter the alert state; s critical represents the critical threshold of each indicator, which exceeds the value to enter the dangerous state; s max represents the maximum acceptable value of each indicator, which exceeds the value that may cause the router to crash.

[0080] The test intensity adjustment is based on the following adaptive control algorithm: Define the target state s target , which is located in the alert state interval, indicating the stress state that the router is expected to reach; Calculate the gap Δs t between the current state and the target state = s target -s t ; Based on the proportional-integral-derivative (PID) control algorithm, calculate the test intensity adjustment amount: Where, ΔI t represents the test intensity adjustment amount at time point t; K prepresents the proportional coefficient, controlling the response strength to the current state deviation; K i represents the integral coefficient, controlling the response strength to the historical accumulated deviation; K d represents the differential coefficient, controlling the response strength to the rate of change of the deviation; represents the state deviation accumulation from the start to the time point t; Δs t -Δs t-1 represents the change amount of the state deviation between the current time point and the previous time point.

[0081] update test strength: I t+1 =I t +ΔI t , and ensure that I t+1 is within the effective range. Wherein, I t represents the test strength at time point t; I t+1 represents the test strength at time point t+1; In order to prevent the router from collapsing suddenly, the application introduces a security protection mechanism: define the state change rate threshold Δs threshold , when detecting rapid deterioration of the state (|Δs t |>Δs threshold ), immediately reduce the test strength; wherein, Δs threshold represents the safety threshold of state change; |Δs t | represents the absolute value of the state change.

[0082] Implement the recovery period mode, after the router experiences high pressure test, give a certain recovery time, and then perform the next round of test; Record the response characteristics of the router under different test strengths, establish a response model, and use it to predict the reaction of the router to the change of the test strength.

[0083] Through the above adaptive adjustment algorithm, the system can accurately control the test strength, and explore the performance limit of the router to the greatest extent without causing the router to completely collapse.

[0084] Step 4.2, implement the attack and defense game model; According to another embodiment of the application, this step builds an attack and defense game model, introduces the defense measures of the router as environmental feedback into the test process, automatically mutates the attack strategy when detecting the defense measures, and realizes the continuous evolution of attack and defense confrontation.

[0085] The attack and defense game model is defined as a two-person zero-sum game G={A, D, U A , U D}, wherein A is the strategy space of the attacker, including various attack strategies; D is the strategy space of the defender (router), including various defense strategies; UA is the attacker's utility function, mapping the combination of attack strategy and defense strategy to a real-valued utility; U D is the defender's utility function, and U D = -U A , indicating that the defender's payoff is opposite to the attacker.

[0086] In actual tests, the defender's strategy is usually not directly observable, but can only be inferred through router responses. Therefore, this application adopts a Bayesian game framework and introduces the belief distribution μ(d) of the defense strategy, representing the probability of the defender adopting strategy d ∈ D.

[0087] The attacker's goal is to choose the optimal attack strategy to maximize the expected utility: where a * represents the optimal attack strategy; argmax a∈A represents finding the strategy that maximizes the objective function among all possible attack strategies; represents the expected utility of attack strategy a under the defense strategy distribution μ; ∑ d∈D μ(d)·U A (a,d) represents the weighted sum over all possible defense strategies.

[0088] To counter the defense measures of the router, this application constructs an attack strategy mutation algorithm: Defense detection: By monitoring the response characteristics of the router (such as response time changes, packet loss rate patterns, connection reset behaviors, etc.), determine whether the router has enabled specific defense measures; Defense type identification: Based on the response characteristics, use multiple classifiers to identify the specific defense type d ∈ D and update the belief distribution μ(d); Attack strategy mutation: Generate a mutated attack strategy for the detected defense measures. Mutation operations include: Parameter adjustment: Modify attack parameters such as rate, duration, packet size, etc. Feature confusion: Change the statistical characteristics of attack traffic to avoid feature detection; Protocol transformation: Switch to a different protocol or attack carrier while maintaining attack effectiveness; Timing pattern mutation: Change the timing pattern of the attack to avoid timing-based detection.

[0089] Mutation effect evaluation: Evaluate the effect of the mutated attack strategy and update the attack strategy utility function U A (a,d).

[0090] Through the game model and mutation algorithm, the system can dynamically adjust the attack strategy according to the defense measures of the router, and continuously explore the boundary and weakness of the router defense mechanism.

[0091] Step 4.3, build resource monitoring probe; According to another embodiment of the application, this step builds a special memory and CPU resource monitoring probe, which captures resource usage anomalies in real time and provides a basis for test strategy adjustment.

[0092] The resource monitoring probe system includes the following components: Data acquisition module: through SNMP, SSH or device API interface, collect resource usage data of the router, including: CPU utilization (overall and each core); Memory usage (total memory, used memory, buffer, cache, etc.); Network interface statistics (throughput, packet count, error count, etc.); Connection table status (active connection number, connection establishment rate, etc.); System log information; Abnormality detection module: use multiple abnormality detection algorithms to identify resource usage anomalies in real time: Statistical threshold detection: set dynamic threshold based on statistical distribution; Time series analysis: detect resource usage trends and mutations; Correlation analysis: detect abnormal correlations between multiple resource indicators; Periodicity analysis: identify abnormal periodic changes in resource usage.

[0093] Root cause analysis module: when resource anomalies are detected, analyze possible causes: Build resource indicator dependency graph to track abnormal propagation path; Correlate test activities with resource anomalies to establish causal relationships; Apply decision tree or Bayesian network for probability inference to determine the most likely root cause.

[0094] Warning system: graded warning according to the severity of the anomaly: Mild anomaly: record but do not interrupt the test; Moderate anomaly: adjust test parameters and reduce related test intensity; Serious anomaly: suspend the test and enter the recovery period.

[0095] Resource monitoring data is stored in a time series database, supporting high-frequency writing and efficient querying, facilitating subsequent analysis. At the same time, a real-time visualization interface is provided to display resource usage and abnormal events.

[0096] Through the above resource monitoring probe system, the abnormal resource use of the router in the test process can be captured in real time, accurate basis is provided for dynamic adjustment of the test strategy, and the test process is prevented from causing irrecoverable damage to the equipment.

[0097] Step 5, based on the collected performance data, a multi-dimensional evaluation index system is constructed, automatic performance bottleneck analysis is realized, and a performance prediction model is constructed; The method mainly includes the following sub-steps: Step 5.1, constructing a multi-dimensional evaluation index system; This step constructs a multi-dimensional evaluation index system, which evaluates the router overall performance from multiple dimensions such as performance stability, security protection, and function correctness, and forms a complete evaluation report.

[0098] The multi-dimensional evaluation index system includes the following main dimensions: Performance stability dimension: Long-time high-load stability: the time during which the system can stably run under continuous high load (≥90% of maximum processing capacity); resource utilization balance: the balance degree of resource utilization such as CPU, memory, and network interface, quantified by the coefficient of variation CV=σ / μ; performance fluctuation degree: the variance or standard deviation of performance indicators under stable load; Recovery capability: the time required to recover from an extreme stress state to a normal working state.

[0099] Security protection dimension: Attack detection rate: the percentage of the number of attacks successfully detected to the total number of attacks; False positive rate: the percentage of normal traffic incorrectly identified as attacks to the total normal traffic; Defense effectiveness: the percentage of the number of attacks successfully defended to the total number of attacks; Defense response time: the average time from the start of the attack to the effectiveness of the defense measures.

[0100] Function correctness dimension: Protocol compliance: the degree of compliance of the router to network protocol standards; Quality of service guarantee: the quality of service guarantee level of key services under stress conditions; Boundary condition handling: the degree of maintaining function correctness under extreme conditions (such as interface rate mutation, connection number explosion, etc.).

[0101] Based on the above dimensions, the application uses a radar chart to represent the comprehensive evaluation results of the router, and intuitively displays the performance of each dimension. At the same time, the comprehensive score is calculated: Wherein, w iis the weight of dimension i, indicating the importance of this dimension in the overall evaluation, the weight value ranges from 0 to 1, and the sum of all dimension weights is equal to 1; Score i is the score of dimension i, indicating the performance of the router in this dimension, usually adopting a standardized score of 0100; n is the total number of evaluation dimensions, that is, the number of performance indicator dimensions considered, including the aforementioned performance stability, security protection, functional correctness and other dimensions; Score is the final comprehensive score, reflecting the overall performance level of the router.

[0102] In addition, the performance boundary under different working conditions should be evaluated, and the performance boundary curve should be drawn to reveal the performance bottleneck. For example, by fixing other parameters, the influence of the change of a single parameter (such as the number of concurrent connections) on performance is studied to determine the performance inflection point.

[0103] The evaluation report finally contains the following contents: quantitative performance indicators, quality level assessment, performance boundary analysis, performance bottleneck identification, improvement suggestions, etc., providing a scientific basis for further optimization of the router.

[0104] Step 5.2, automatic performance bottleneck analysis is realized; This step realizes automatic performance bottleneck analysis, identifies the performance bottleneck of the router through correlation analysis, locates the root cause of performance degradation, and provides direction for performance optimization.

[0105] The realization of automatic performance bottleneck analysis includes the following key steps: Data collection and preprocessing: Collect comprehensive performance data of the router during the test process, including system-level indicators (CPU, memory, interrupts, etc.) and application-level indicators (throughput, delay, etc.); Perform data cleaning and standardization to handle outliers and missing values; Apply time window smoothing to reduce the impact of transient fluctuations.

[0106] Bottleneck detection algorithm: Resource saturation analysis: detect the time points and duration when resource utilization rate approaches saturation (such as CPU utilization > 95%); Performance mutation detection: identify the mutation points of performance indicators, using CUSUM (Cumulative Sum) or PELT (Pruned Exact Linear Time) algorithm; Correlation heat map analysis: calculate the correlation coefficient matrix between performance indicators, and draw a heat map to identify a group of strongly correlated indicators; Principal component analysis (PCA): dimensionality reduction analysis to identify the factor set that contributes most to performance variation.

[0107] Root cause positioning: Construct performance causal diagram: based on domain knowledge and data correlation, construct the causal relationship between indicators; Automatically discover causal relationships using causal inference algorithms (such as PC algorithm or FCI algorithm); Apply decision tree or random forest algorithm to identify the decision path that causes performance anomalies; Use Bayesian networks to calculate the probability distribution of each factor as a root cause.

[0108] Bottleneck verification: Design a control experiment to control the suspected bottleneck factor and verify its impact on performance; Simulate bottleneck scenarios and compare system behavior with real bottleneck scenarios; Calculate the confidence level to represent the degree of certainty of the identified bottleneck.

[0109] The application also implements a bottleneck visualization system, including: Performance topology map: show system components and their performance status; Bottleneck hotspot marking: mark the bottleneck location and severity on the topology map; Time correlation diagram: show the time correlation between the bottleneck and performance decline; Root cause probability distribution diagram: display the probability of each factor as a root cause.

[0110] Through the above automatic bottleneck analysis system, the router performance bottleneck can be quickly and accurately identified, providing a scientific basis for targeted optimization.

[0111] Step 5.3, construct performance prediction model; According to another embodiment of the application, this step constructs a performance prediction model to predict the performance of the router in untested scenarios based on historical test data, expands the test coverage, and reduces the cost of full-scale testing.

[0112] The construction process of the performance prediction model is as follows: Feature engineering: Extract feature vector X, including input parameters (such as number of concurrent connections, packet size, traffic type, etc.); Define target vector Y, including performance indicators (such as throughput, delay, packet loss rate, etc.); Feature transformation and normalization to ensure comparability of features with different dimensions; Feature selection, using regularization methods such as Lasso, Ridge, or principal component analysis to reduce feature dimensionality.

[0113] Model selection and training: Evaluate multiple prediction models, including linear regression, support vector regression (SVR), random forest, gradient boosting tree (GBDT), etc.; Select the most suitable model type for different performance indicators; Use grid search or Bayesian optimization for hyperparameter tuning; Use K-fold cross-validation to evaluate model performance and avoid overfitting.

[0114] Model evaluation: Calculate mean squared error (MSE), mean absolute error (MAE), coefficient of determination (R 2 ) and other indicators to evaluate prediction accuracy; Analyze residual distribution and test model assumptions; Evaluate model performance in different working intervals to ensure the model is effective in the entire prediction range.

[0115] Uncertainty quantification: Calculate the confidence interval of the prediction result using Bootstrap or Bayesian method; Analyze model sensitivity and evaluate the impact of input parameter changes on prediction results; Identify high uncertainty areas of the model to guide further testing.

[0116] For high-dimensional nonlinear relationships, this application uses deep learning models such as: LSTM network: captures long-term dependencies in time-series performance data; Autoencoder: used for feature dimensionality reduction and anomaly detection; Mixed expert model (MoE): divides the prediction space into multiple subspaces and uses specialized predictors for different subspaces.

[0117] In addition, active learning techniques should be used to guide testing: based on the uncertainty of the prediction model, automatically select the most valuable next test point, maximize information gain, and optimize the allocation of test resources.

[0118] Through the above performance prediction model, while maintaining prediction accuracy, the test cost and time are significantly reduced, achieving efficient and comprehensive router performance evaluation and analysis.

[0119] Device embodiment: Based on the above method implementation, a wireless router whole machine one-stop testing device is provided, which includes the following modules: traffic pattern generation module: for generating traffic patterns simulating real network environment, including traffic pattern library construction unit, Markov decision unit, reinforcement learning unit and traffic injection unit; Attack pattern creation module: for creating unknown attack patterns, including generative adversarial network unit, variational autoencoder unit and cross-protocol layer attack combination unit; Coordinated attack simulation module: for simulating distributed coordinated attacks, including multi-agent coordination framework unit, shared value network unit and attack location optimization unit; Test strategy adjustment module: for adaptive adjustment of test strategy, including intensity adjustment unit, game model unit and resource monitoring unit; Result analysis and evaluation module: for analyzing test results and evaluating router performance, including attack classification unit, effect analysis unit, defense evaluation unit and report generation unit.

[0120] The modules communicate through a unified data exchange interface, and the data generated during the test is stored in a central database for sharing and analysis by each module.

[0121] Computer device embodiment: The method of the present application can be implemented on a general-purpose computer device, which includes: Processor: at least one central processing unit (CPU) or graphics processing unit (GPU) for performing various computing tasks; Memory: including random access memory (RAM) and read-only memory (ROM) for storing operating systems, applications and data; Communication interface: including Ethernet interface, WiFi interface, etc., for communication with the tested router and other test nodes; Storage device: including hard disk drive or solid state drive, for storing test data and results; Input and output device: for user interaction and result display.

[0122] The software system running on the computer device includes operating system, database management system, test control system and result analysis system, etc., which work together to realize each step of the above test method.

[0123] The above describes the embodiments of the present application, but the embodiments are not limited to the specific implementation described above, which is only illustrative and not limiting, and those skilled in the art can make more forms of equivalent embodiments under the inspiration of the embodiments, which are all within the protection of the embodiments.

Claims

1. A wireless router system one station test method, characterized in that, The method comprises the following steps: a typical network application feature traffic pattern library is constructed, different traffic patterns are dynamically combined using Markov decision process, attack traffic patterns are generated using deep reinforcement learning algorithm, and variable traffic is injected through software defined network technology; based on the traffic patterns, new attack patterns are created using generative adversarial network, and variant attack samples are generated using variational autoencoder, and a composite attack chain is constructed combining knowledge graph; according to the generated attack patterns, a multi-agent collaborative attack system is constructed, a collaborative decision mechanism based on shared value network is realized, and the optimal attack position distribution is automatically determined according to the test target; during the execution of the collaborative attack, the test intensity is dynamically adjusted according to the real-time response of the router, an attack and defense game model is realized, and a resource monitoring probe is constructed to collect performance data; based on the collected performance data, a multi-dimensional evaluation index system is constructed, automatic performance bottleneck analysis is realized, and a performance prediction model is constructed.

2. The one station test method for a wireless router complete machine according to claim 1, wherein, The step of constructing the typical network application feature traffic pattern library specifically comprises: by analyzing the traffic data in the real network environment, the characteristic parameters of different network applications are extracted; statistical modeling is performed on the traffic characteristics of each network application to obtain the probability distribution function of the characteristic parameters; a traffic feature library containing video streaming media, online games, web browsing, file transfer, VoIP calls and common network applications is constructed.

3. The one station test method for a wireless router complete machine according to claim 1, wherein, The step of generating attack traffic patterns using deep reinforcement learning algorithm adopts a double deep Q network structure, the input of the double deep Q network is the current state of the router, including performance indicators, and the output is attack actions, including attack type, attack parameter and attack intensity; and priority experience replay technology is used to improve learning efficiency.

4. The one station test method for a wireless router complete machine according to claim 1, wherein, The step of generating variant attack samples using variational autoencoder specifically comprises: the encoder network maps the input attack sample to the distribution parameters of the latent space; latent variables are sampled from the distribution; the decoder network reconstructs the latent variables into new attack samples; an attack pattern difference evaluation module is designed to calculate the difference degree of the generated samples and the known attack samples, and only the generated samples with a difference degree exceeding a preset threshold are retained.

5. The one station test method for a wireless router complete machine according to claim 1, wherein, The step of constructing the multi-agent collaborative attack system includes at least four types of agents: probe agent: responsible for detecting the vulnerability of the network and the router; resource consumption agent: responsible for consuming the computing and network resources of the router; denial of service attack agent: responsible for performing various denial of service attacks; coordination agent: responsible for coordinating the actions of various attack agents.

6. The one station test method for a wireless router complete machine according to claim 1, wherein, The step of dynamically adjusting the test intensity according to the real-time response of the router specifically comprises: define a router state vector containing multiple performance indicators, and define normal state interval, alert state interval and dangerous state interval; define the target state in the alert state interval; calculate the gap between the current state and the target state; based on the proportional-integral-derivative control algorithm, calculate the test intensity adjustment amount; update the test intensity and introduce a security protection mechanism.

7. The one station test method for a wireless router complete machine according to claim 1, wherein, The step of realizing the attack and defense game model specifically comprises: construct a two-person zero-sum game model, including the strategy space of the attacker, the strategy space of the defender and their respective utility functions; A Bayesian game framework is adopted, and belief distribution of defense strategy is introduced. An attack strategy mutation algorithm is constructed, including four links of defense detection, defense type identification, attack strategy mutation and mutation effect evaluation.

8. The one station test method for a wireless router complete machine according to claim 1, wherein, The step of constructing the multi-dimensional evaluation index system specifically includes: An evaluation index system is constructed from performance stability, security protection, functional correctness and multiple dimensions; A radar chart is used to represent the comprehensive evaluation result of the router; The comprehensive score is calculated and the performance boundary under different working conditions is evaluated.

9. The one station test method for a wireless router complete machine according to claim 1, wherein, The step of constructing the performance prediction model specifically includes: Feature engineering is performed to extract feature vectors and define target vectors; Multiple prediction models are evaluated and hyperparameters are optimized; Prediction accuracy indicators are calculated and residual distribution is analyzed; The uncertainty of the prediction result is quantified and active learning technology is used to guide the test.

10. A wireless router system one-stop test system for performing the wireless router system one-stop test method of any one of claims 1-9, wherein It includes: A traffic pattern generation module for generating traffic patterns simulating real network environments, including a traffic pattern library construction unit, a Markov decision unit, a reinforcement learning unit and a traffic injection unit; An attack pattern creation module for creating unknown attack patterns, including a generative adversarial network unit, a variational autoencoder unit and a cross-protocol layer attack combination unit; A collaborative attack simulation module for simulating distributed collaborative attacks, including a multi-agent collaboration framework unit, a shared value network unit and an attack location optimization unit; A test strategy adjustment module for adaptively adjusting the test strategy, including a strength adjustment unit, a game model unit and a resource monitoring unit; A result analysis and evaluation module for analyzing test results and evaluating router performance, including an attack classification unit, an effect analysis unit, a defense evaluation unit and a report generation unit.

Citation Information

Cited By

  • Software performance test system, method and equipment under high-concurrency scene

    CN121542147A