Multi-probe collaborative network detection method and system
Through the multi-probe collaborative network detection method, combined with linear regression model and MADDPG algorithm to optimize probe deployment, the resource limitation and redundancy problems of traditional detection systems are solved, and efficient IoT device detection and vulnerability warning are achieved.
Patent Information
- Application Number
- CN202510839135.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
Traditional single-node detection systems are difficult to ensure detection accuracy and real-time detection under limited resource conditions, while distributed detection systems have problems such as resource utilization imbalance, high redundancy rate and deviation in fingerprint database timeliness.
Multi-probe collaboration network detection method is adopted, combined with boundary gateway protocol and IoT device information, and optimize probe deployment through linear regression model and integer linear planning, the MADDPG algorithm is used to realize intelligent collaboration of probe groups, dynamically adjust the detection strategy, and introduce real-time reward function to optimize resource utilization.
Under limited resources, it maximizes coverage probability, ensures detection accuracy and real-timeness, reduces redundancy, improves resource utilization, and ensures the timeliness of fingerprint databases and the effectiveness of vulnerability warnings.
Smart Images

Figure CN120358177A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet of Things security technologies, and in particular, to a network detection method and system with multi-probe collaboration. Background Art
[0002] With the number of global Internet of Things (IoT) devices exceeding 30 billion and continuing to surge, network space mapping technology has become the core pillar of network security situation awareness. Its core task is to dynamically capture and accurately identify device fingerprint features through a detection system, providing real-time data support for vulnerability early warning and security protection.
[0003] In the prior art, traditional single-node detection systems and distributed detection systems are usually used to dynamically capture and accurately identify device fingerprint features. However, the traditional single-node architecture is restricted by the rigid constraints of local computing power and network bandwidth. A single probe can only maintain a detection rate of about 2,000 devices per second. Currently, the device access density of mainstream 5G base stations has reached the scale of tens of thousands. Therefore, in the face of the high-frequency changes in device status and the massive protocol interaction data streams in a dynamic network environment, the contradiction between the computationally intensive fingerprint matching algorithm and the hardware resources of the traditional single-node architecture will intensify, resulting in the device identification delay exceeding a reasonable threshold. Moreover, the traditional single-node architecture lacks a dynamic resource allocation mechanism, resulting in the coexistence of CPU overload and network bandwidth idleness during the detection period. At the same time, in order to avoid resource overrun, a downsampling strategy needs to be adopted, which causes the missed detection of key device statuses and directly affects the detection quality. Therefore, in summary, the detection system with the traditional single-node architecture is difficult to ensure detection accuracy and detection real-time performance under limited resource conditions.
[0004] The distributed architecture of the distributed detection system can theoretically break through the single-node resource limit through horizontal expansion. However, in actual deployment, the static task sharding mechanism of the distributed architecture may lead to unbalanced resource utilization. In a large-scale detection system, using a static task sharding mechanism may cause some nodes to be overloaded and some nodes to be idle. Moreover, the detection targets between adjacent probes may have significant overlap, resulting in a relatively high proportion of redundant data generated by repeated detection. At the same time, the existing polling heartbeat synchronization mechanism of the distributed detection system usually has a fixed interval time. However, in scenarios where device status changes rapidly, this fixed interval may be difficult to adapt to, resulting in a timeliness deviation of the target device fingerprint database and directly affecting the effectiveness of vulnerability early warning. Therefore, in summary, the distributed detection system is prone to technical problems such as low efficiency of distributed node collaboration, high redundancy rate, timeliness deviation of the fingerprint database, and decline in the effectiveness of vulnerability early warning.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] In view of the above problems, the present application provides a network detection method and system with multi-probe collaboration, which can hardly ensure detection accuracy and detection real-time performance under limited resource conditions, reduce redundancy and improve resource utilization rate, and at the same time ensure the timeliness of the fingerprint database and the effectiveness of vulnerability warning.
[0007] To achieve the objectives of the present application, the following technical solutions are provided in the present application:
[0008] In a first aspect, the present application provides a network detection method with multi-probe collaboration, including:
[0009] Combining the Border Gateway Protocol with existing Internet of Things device information to obtain multiple Autonomous System (AS) subsystems, calculating the Internet of Things coverage probability of each autonomous system subsystem using a linear regression model, and determining the deployment location of the probe group through an integer linear programming method;
[0010] Defining the local state of each probe, obtaining the probability distribution of each subnet through the MADDPG algorithm, and obtaining the subnet index to be scanned by the probe;
[0011] After the probe executes an action, update the local state, calculate the global reward through an immediate reward function, store the experience tuple in the replay buffer, and update the Critic network and Actor network of the MADDPG algorithm.
[0012] In a possible implementation manner, the step of combining the Border Gateway Protocol with existing Internet of Things device information to obtain multiple autonomous system subsystems, calculating the Internet of Things coverage probability of each autonomous system subsystem using a linear regression model, and determining the deployment location of the probe group through an integer linear programming method includes:
[0013] Extract autonomous system features according to the Border Gateway Protocol and existing Internet of Things device information, construct a graph according to the autonomous system features to obtain an autonomous system, and split the autonomous system into x subsystems;
[0014] Construct a linear regression model to predict the Internet of Things device coverage density probability of each autonomous system subsystem, and establish a mean square error loss function to train the linear regression model;
[0015] Determine the deployment location of the probe group through an integer linear programming method according to the Internet of Things device coverage density probability of each autonomous system subsystem.
[0016] In a possible implementation, the autonomous system feature vector is , is the number of IP addresses managed, is the regional location, represents the traffic pattern, represents the known IoT device distribution;
[0017] The graph is constructed as , where is the graph construction, is the set of autonomous systems, is the peer relationship; The autonomous system is represented as .
[0018] In a possible implementation, the linear regression model is ; where is the predicted IoT device density, is the autonomous system feature vector, is the transpose of the vector, is the first optimization parameter, is the second optimization parameter;
[0019] The mean squared error loss function is ; where is the mean squared error loss function, is the actual IoT device density, is the number of autonomous system subsystems;
[0020] The specific training formula for training the linear regression model is and ; where represents the learning rate.
[0021] In a possible implementation, the steps of defining the local state of each probe, obtaining the probability distribution of each subnet through the MADDPG algorithm, and obtaining the subnet index that the probe is about to execute a scan include:
[0022] Define the local state of the probe and construct an Actor network based on the MADDPG algorithm;
[0023] Input the local state of the probe into the Actor network to obtain the scan probability distribution of each subnet;
[0024] According to the scan probability distribution of each subnet, use -greedy strategy to select actions in stages to obtain the subnet index that the probe is about to execute a scan.
[0025] In a possible implementation, the local state of the probe is:
[0026] ;
[0027] where is the local state of the i-th probe, is the scanned subnets represented by a binary vector, is the response feature dimension, is the global statistic dimension;
[0028] The scanning probability distribution of each subnet is:
[0029] ;
[0030] where is the scanning probability distribution of each subnet, represents the Actor network, and softmax is the action score for each subnet.
[0031] In a possible implementation, the method of using -greedy strategy to select actions in stages is:
[0032] In the training stage, with probability sample the subnet index from the scanning probability distribution of each subnet through the multinomial distribution, and with probability randomly select the subnet index from the uniform distribution; in the evaluation stage, select the subnet index with the highest probability.
[0033] In a possible implementation, the steps for the probe to update the local state after performing an action, calculate the global reward through the immediate reward function, and store the experience tuple in the replay buffer to update the Critic network and Actor network of the MADDPG algorithm include:
[0034] The probe executes the selected action according to the subnet index to be scanned, receives the environmental feedback, updates the state and experience data, and obtains the global reward value through the immediate reward function;
[0035] After updating the local state and global state, recycle the experience tuple, and the experience tuple includes: the current state, action, global reward value, and the next state;
[0036] Update the Critic and Actor networks in the MADDPG algorithm according to the recycled experience tuple.
[0037] In a possible implementation, the immediate reward function is:
[0038] ;
[0039] Among them, is the immediate reward function, is the indicator function, with boolean value characteristics, is the total number of probes, is the reward factor when a new device is detected, is the penalty factor for taking each probing action, is the penalty factor for a probe being detected or blocked, is the newly discovered device in the i-th scanning action, is the event that a probe is detected or blocked in the i-th scanning action. If the scan is detected or blocked, the value of I(detected_i) is 1, otherwise it is 0;
[0040] The local state update method is:
[0041] ;
[0042] ;
[0043] Among them, is a vector with a length of 1000, is the bitwise OR operation;
[0044] The global state update method is:
[0045] ;
[0046] The update strategy of the Actor network is:
[0047] ;
[0048] The update strategy of the Critic network is:
[0049] ;
[0050] Among them, is the parameter gradient of the Actor network, is the expectation of experience replay, is the gradient of the Actor with respect to its own parameters, is the gradient of the Critic with respect to the action, represents the action generated by the policy , is the Q value of the Critic network, , refers to a batch of samples randomly sampled from the experience replay pool, is the network parameter, is the target network parameter, is the current state, is the next state, is the current action, is the next action, is the discount factor, is the mean squared error loss of the Critic network, is from sample state transition tuples from is the target Critic network for all possible Q-value maximum of
[0051] In a second aspect, the present application also provides a multi-probe collaborative network detection system for performing the above multi-probe collaborative network detection method. The system includes:
[0052] A probe deployment module, which is used to combine the Border Gateway Protocol with existing Internet of Things device information to obtain multiple autonomous system subsystems, calculate the Internet of Things coverage probability of each autonomous system subsystem by using a linear regression model, and determine the deployment location of the probe group through an integer linear programming method;
[0053] An action acquisition module, which is used to define the local state of each probe, obtain the probability distribution of each subnet through the MADDPG algorithm, and obtain the subnet index that the probe is about to execute a scan on;
[0054] An algorithm update module, which is used to update the local state after the probe executes an action, calculate the global reward through an immediate reward function, store the experience tuple in the replay buffer, and update the Critic network and Actor network of the MADDPG algorithm.
[0055] The technical solution provided by the present application may include the following beneficial effects:
[0056] Through a multi-probe collaborative network detection method and system provided by the present application, it is possible to optimize the probe deployment through a linear regression model and integer linear programming, achieve the maximum coverage probability under limited resources, and ensure accuracy and real-time performance; moreover, the MADDPG algorithm is used to achieve multi-probe intelligent collaboration, dynamically adjust the detection strategy, reduce redundancy and improve resource utilization; at the same time, dynamic resource perception and an immediate reward function are introduced, comprehensively considering the detection rate, the number of actions and concealment, enhancing the system adaptability, reducing redundancy and improving resource utilization, while ensuring the timeliness of the fingerprint database and the effectiveness of vulnerability early warning.
[0057] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0059] Figure 1 It is a schematic flowchart of a network detection method with multi-probe collaboration provided by an embodiment of the present application;
[0060] Figure 2 It is a schematic flowchart of step S100 of a network detection method with multi-probe collaboration provided by an embodiment of the present application;
[0061] Figure 3 It is a schematic flowchart of step S200 of a network detection method with multi-probe collaboration provided by an embodiment of the present application;
[0062] Figure 4 It is a schematic flowchart of step S300 of a network detection method with multi-probe collaboration provided by an embodiment of the present application;
[0063] Figure 5 It is a schematic diagram of the MADDPG algorithm framework of a network detection method with multi-probe collaboration provided by an embodiment of the present application;
[0064] Figure 6 It is a schematic structural diagram of a network detection system with multi-probe collaboration provided by an embodiment of the present application. Detailed implementation manners
[0065] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments.
[0066] In this example embodiment, a network detection method with multi-probe collaboration is first provided. Referring to Figure 1 shown in, the network detection method with multi-probe collaboration may include the following steps:
[0067] Step S100: Combine the Border Gateway Protocol with existing Internet of Things device information to obtain multiple autonomous system subsystems, calculate the Internet of Things coverage probability of each autonomous system subsystem using a linear regression model, and determine the deployment location of the probe group through an integer linear programming method.
[0068] Step S200: Define the local state of each probe, and obtain the subnet index where the probe is about to perform a scan through the probability distribution of each subnet.
[0069] Step S300: Update the local state after the probe performs an action, calculate the global reward through the immediate reward function, store the experience tuple in the replay buffer, and update the Critic network and Actor network of the MADDPG algorithm.
[0070] Through the above network detection method with multi-probe collaboration, it is possible to combine the Border Gateway Protocol (BGP) and existing Internet of Things device information, construct the Internet of Things coverage probability of the autonomous system subsystem using a linear regression model, and reasonably select the deployment location of the probe group through the integer linear programming method to maximize the coverage probability under limited resources and optimize the probe deployment cost; moreover, using the MADDPG algorithm, adopting a centralized training and distributed execution architecture, through the collaborative work of the joint state observation, policy network, and value network, optimize the cooperation relationship between multiple probes, adjust the detection strategy of the probe in real time, improve the detection efficiency, reduce redundant scans, and improve the intelligence of the system; at the same time, by constructing a collaborative strategy for multiple probes, it is possible to dynamically respond to changes in the Internet of Things environment, improve the flexibility and intelligence of Internet of Things device detection. The immediate reward function adopted takes into account the detection rate, the total number of detection actions, the detection efficiency, and concealment, further promoting the optimization of the system. By optimizing the probe deployment location and dynamic detection strategy, it is possible to maximize the coverage rate of Internet of Things devices under limited resources, improving the detection efficiency and detection ability of Internet of Things devices.
[0071] Next, reference will be made to Figures 2 to 5 to describe each step of the above network detection method with multi-probe collaboration in the present exemplary embodiment in more detail.
[0072] In step S100, multiple autonomous system subsystems are obtained by combining the Border Gateway Protocol and existing Internet of Things device information, the Internet of Things coverage probability of each autonomous system subsystem is calculated using a linear regression model, and the deployment location of the probe group is determined through the integer linear programming method.
[0073] In a possible implementation manner, step S100 may further include the following sub-steps:
[0074] In step S110, extract the autonomous system features according to the Border Gateway Protocol and existing Internet of Things device information, construct a graph according to the autonomous system features to obtain the autonomous system, and split the autonomous system into x subsystems.
[0075] Among them, the autonomous system feature vector is , is the number of IP addresses managed, is the regional location, represents the traffic pattern, represents the known IoT device distribution; the graph is constructed as , where is the set of autonomous systems, is the peer relationship; the autonomous system is represented as .
[0076] It can be understood that the encoding of the regional location is a one - hot vector; the traffic pattern can be quantified as a scalar, such as the HTTP traffic ratio, which can be estimated through traffic sniffing or public datasets; the known IoT device distribution can be the number of devices recorded using Shodan; in terms of mathematical representation of the graph construction, the adjacency matrix of the graph is , where is the number of autonomous systems, represents and are connected by an edge, , is the degree of autonomous system , that is, the number of other autonomous systems directly connected to , reflecting the connectivity of . Using the autonomous system topology information can assist in inference. For example, a highly connected autonomous system (such as a tier - 1 ISP) may serve more end - users, indirectly affecting the IoT device density.
[0077] In step S120, a linear regression model is constructed to predict the probability of the IoT device coverage density of each autonomous system subsystem, and a mean - squared error loss function is established to train the linear regression model.
[0078] Among them, the linear regression model is ; where is the predicted IoT device density, is the autonomous system feature vector, is the transpose of the vector, is the first optimization parameter, is the second optimization parameter;
[0079] The mean - squared error loss function is ; where is the mean - squared error loss function, is the actual IoT device density, is the number of autonomous system subsystems;
[0080] The specific training formula for training the linear regression model is and ; where Denotes the learning rate.
[0081] It can be understood that after obtaining multiple autonomous system subsystems, each The density of IoT devices is estimated as , assuming that the upper limit of probe deployment is k at this time, the optimization objective is to select several subsets of autonomous systems to maximize the objective function:
[0082] ;
[0083] At this time, the estimation of is particularly crucial, so linear regression prediction is used , that is, for the feature vector , find a linear function corresponding to it for probability prediction, that is, through to represent , and at the same time, the gradient descent algorithm is used to train and , and the learning rate in the specific training formula can be modified according to the number of autonomous system subsystems.
[0084] Optionally, the range of the trained may be negative or greater than 1, and it needs to be normalized through the normalization formula, and the normalization formula is .
[0085] In step S130, according to the IoT device coverage density probability of each autonomous system subsystem, the deployment location of the probe group is determined by the integer linear programming method.
[0086] It can be understood that after obtaining the IoT device density probability of the th autonomous system subsystem, next, the integer linear programming method is selected to solve the selection of the autonomous system subsystem, and its calculation method is:
[0087] ;
[0088] This method can select the subsystems with as high a density as possible in the case of the highest k subsystems to maximize the coverage rate.
[0089] In step S200, the local state of each probe is defined, and the subnet index to be scanned by the probe is obtained through the probability distribution of each subnet.
[0090] In a possible implementation manner, step S200 may further include the following sub-steps:
[0091] In step S210, the local state of the probe is defined, and the Actor network based on the MADDPG algorithm is constructed.
[0092] Among them, the local state of the probe is:
[0093] ;
[0094] Among them, is the local state of the i-th probe, is the scanned subnets represented by a binary vector, is the response feature dimension, is the global statistic dimension.
[0095] It should be noted that , among them, is the number of subnets, 1 indicates scanned, 0 indicates not scanned; , among them, ; , among them, .
[0096] It can be understood that, as Figure 5 shown, it is the framework of the MADDPG algorithm of this application, including a training layer and a decision layer, composed of multiple intelligent agent probes, and each intelligent agent probe is composed of a Critic network and an Actor network.
[0097] In step S220, the local state of the probe is input into the Actor network to obtain the scanning probability distribution of each subnet.
[0098] Among them, the scanning probability distribution of each subnet is:
[0099] ;
[0100] Among them, is the scanning probability distribution of each subnet, represents the Actor network, and softmax is the action score of each subnet.
[0101] It should be noted that the Actor network has two fully connected layers, each with 128 units; the softmax function selects ReLU, and its specific calculation method is , and the output dimensional vector, and the specific mathematical calculation is:
[0102] ;
[0103] ;
[0104] ;
[0105] Among them, is the weight matrix of the first fully connected layer, , is the weight matrix of the second fully connected layer, , is the weight matrix of the third fully connected layer, , is the bias vector of the first layer, is the bias vector of the second layer, is the bias vector of the third layer, is the output of the first hidden layer, is the output of the second hidden layer, is the score vector;
[0106] Its probability distribution .
[0107] In step S230, according to the scanning probability distribution of each subnet, the -greedy strategy is used to select actions in stages to obtain the subnet index that the probe is about to perform a scan on.
[0108] Among them, the method of using the -greedy strategy to select actions in stages is as follows:
[0109] In the training stage, with probability sample the subnet index from the scanning probability distribution of each subnet through the multinomial distribution, and with probability randomly select the subnet index from the uniform distribution; in the evaluation stage, select the subnet index with the highest probability.
[0110] It can be understood that the probe selects an action from the probability distribution , and decides which subnet to detect from the discrete set , . The exploration strategy is the -greedy strategy, which combines randomness during training and selects the optimal action during evaluation. The implementation method is , and the decay step is 1000.
[0111] In the training stage, sample from : , with probability randomly select an action: , that is, during training, use a random number generator, and according to judge whether to sample. When sampling, draw a subnet index from the 1000-dimensional probability distribution, for example, use the multinomial distribution for sampling.
[0112] During the evaluation, the subnet index with the highest probability can be directly selected: .
[0113] In step S300, after the probe performs an action, it updates the local state, calculates the global reward through the immediate reward function, stores the experience tuple in the replay buffer, and updates the Critic network and Actor network of the MADDPG algorithm.
[0114] In a possible implementation, step S300 may include the following sub-steps:
[0115] In step S310, the probe performs the selected action according to the subnet index to be scanned, receives the environmental feedback, updates the state and experience data, and obtains the global reward value through the immediate reward function.
[0116] Among them, the immediate reward function is:
[0117] ;
[0118] Among them, is the immediate reward function, is the indicator function, with boolean value characteristics, is the total number of probes, is the reward factor when a new device is detected, is the penalty factor for taking each detection action, is the penalty factor for the probe being detected or blocked, is the newly discovered device in the i-th scan action, is the event that the probe is detected or blocked in the i-th scan action. If the scan is detected or blocked, the value of I(detected_i) is 1, otherwise it is 0.
[0119] It should be noted that the indicator function is used to judge the authenticity of an event, represents that a new device is detected, otherwise it is 0, similar to if detected, otherwise 0; and, the penalty factor for taking each detection action is used to encourage the detection efficiency, and both the reward factor when a new device is detected and the penalty factor for taking each detection action are taken as 1; the penalty factor for the probe being detected or blocked is taken as 10 to encourage the concealment of the probe.
[0120] In step S320, after updating the local state and the global state, the experience tuple is recycled. The experience tuple includes: the current state, action, global reward value, and the next state.
[0121] Among them, the local state update method is:
[0122] ;
[0123] ;
[0124] Among them, is a vector with a length of 1000, is a bitwise OR operation, is the response returned by the environment.
[0125] It should be noted that the bit in is 1 and the rest are 0;
[0126] The global state update method is as follows:
[0127] .
[0128] It should be noted that the global state update is completed by the central server, and the central server broadcasts through communication , and the probe receives and splices it into the complete state.
[0129] In step S330, according to the recycled experience tuples, the Critic and Actor networks in the MADDPG algorithm are updated.
[0130] Among them, the update strategy of the Actor network is:
[0131] ;
[0132] The update strategy of the Critic network is:
[0133] ;
[0134] Among them, is the parameter gradient of the Actor network, is the expectation of experience replay, is the gradient of the Actor with respect to its own parameters, is the gradient of the Critic with respect to the action, represents the action generated by the policy , is the Q value of the Critic network, , refers to a batch of samples randomly sampled from the experience replay pool, is the network parameter, is the target network parameter, is the current state, is the next state, is the current action, is the next action, is the discount factor, is the mean squared error loss of the Critic network, is to sample state transition tuples from , is the maximum Q value of all possible under the target Critic network for .
[0135] It should be noted that the update of the Critic network determines the scanning efficiency and intelligence between probe groups. The input of the Critic network combines the state and action, which brings the update rules of the Actor network and the Critic network. For the update strategies of the Actor network and the Critic network, in this application, the MADDPG parameters used in this invention are as follows:
[0136] The input state of the Actor network is , the hidden layer is a two-layer 128 neural network, and the output is dimensional softmax. The input of the Critic network is all states and actions. The hidden layer is a three-layer structure, with each layer dimension being 512, 256, and 128 respectively, and the output is 1 dimension. The discount factor is , the buffer size is 10000, the network update frequency is every 100 steps, the soft update coefficient is 0.01, and the -greedy strategy is adopted, decaying from 1.0 to 0.1 in 1000 steps.
[0137] Furthermore, in this exemplary embodiment, a network detection system with multi-probe collaboration is also provided for performing the above-mentioned network detection method with multi-probe collaboration. Referring to Figure 6 shown in, the system may include: a probe deployment module, an action acquisition module, and an algorithm update module.
[0138] The probe deployment module is used to combine the Border Gateway Protocol and the existing Internet of Things device information to obtain multiple autonomous system subsystems, calculate the Internet of Things coverage probability of each autonomous system subsystem using a linear regression model, and determine the deployment location of the probe group through an integer linear programming method.
[0139] The action acquisition module is used to define the local state of each probe, obtain the probability distribution of each subnet through the MADDPG algorithm, and obtain the subnet index that the probe is about to execute the scan.
[0140] The algorithm update module is used to update the local state after the probe executes the action, calculate the global reward through an immediate reward function, store the experience tuple in the replay cache, and update the Critic network and the Actor network of the MADDPG algorithm.
[0141] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the art not disclosed herein. The specification and examples are only illustrative, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
[0142] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. The present application is not limited to the exact structures described above and illustrated in the drawings, and it cannot be considered that the specific implementation of the present application is only limited to these descriptions. For those of ordinary skill in the technical field to which the present application belongs, without departing from the concept of the present application, various changes and modifications made should be regarded as falling within the protection scope of the present application.
Claims
1. A network detection method with multi-probe collaboration, characterized in that Including: Combining the Border Gateway Protocol with existing Internet of Things device information to obtain multiple autonomous system subsystems, calculating the Internet of Things coverage probability of each autonomous system subsystem using a linear regression model, and determining the deployment location of the probe group through an integer linear programming method; Defining the local state of each probe, obtaining the probability distribution of each subnet through the MADDPG algorithm, and obtaining the subnet index that the probe is about to execute a scan on; After the probe executes an action, update the local state, calculate the global reward through an immediate reward function, store the experience tuple in the replay buffer, and update the Critic network and Actor network of the MADDPG algorithm.
2. The network detection method with multi-probe collaboration according to claim 1, characterized in that The step of combining the Border Gateway Protocol with existing Internet of Things device information to obtain multiple autonomous system subsystems, calculating the Internet of Things coverage probability of each autonomous system subsystem using a linear regression model, and determining the deployment location of the probe group includes: Performing autonomous system feature extraction based on the Border Gateway Protocol and existing Internet of Things device information, constructing a graph according to the autonomous system features to obtain an autonomous system, and splitting the autonomous system into x subsystems; Constructing a linear regression model to predict the Internet of Things device coverage density probability of each autonomous system subsystem, and establishing a mean squared error loss function to train the linear regression model; Determining the deployment location of the probe group through an integer linear programming method according to the Internet of Things device coverage density probability of each autonomous system subsystem.
3. The network detection method with multi-probe collaboration according to claim 2, wherein The autonomous system eigenvector is , is the autonomous system eigenvector, is the number of IP addresses managed, is the regional location, represents the traffic pattern, represents the known IoT device distribution; The said graph is constructed as , where is the graph structure, is the set of autonomous systems, is the peer relationship; the autonomous system is represented as .
4. The network detection method with multi-probe collaboration according to claim 3, characterized in that The linear regression model is ; where is to predict the density of Internet of Things devices, is the eigenvector of the autonomous system, is the transpose of the vector, is the first optimization parameter, is the second optimization parameter; The mean squared error loss function is ; where is the mean squared error loss function,[[]] is the actual IoT device density,[[]] is the number of autonomous system subsystems; The specific training formula for training the linear regression model is and ; where represents the learning rate.
5. The network detection method with multi-probe collaboration according to claim 1, wherein The step of defining the local state of each probe, obtaining the probability distribution of each subnet through the MADDPG algorithm, and obtaining the subnet index that the probe is about to execute a scan on includes: Defining the local state of the probe and constructing an Actor network based on the MADDPG algorithm; Inputting the local state of the probe into the Actor network to obtain the scan probability distribution of each subnet; According to the scanning probability distribution of each subnet, the -greedy strategy is used to perform action selection in stages to obtain the subnet index where the probe is about to perform a scan.
6. The network detection method with multi-probe collaboration according to claim 5, characterized in that, The local state of the probe is: wherein, is the local state of the i-th probe, uses a binary vector to represent the scanned subnets, is the response feature dimension, is the global statistical dimension; The scan probability distribution of each subnet is: Among them, is the scanning probability distribution of each subnet, represents the Actor network, and softmax is the action score of each subnet.
7. The network detection method with multi-probe collaboration according to claim 5, characterized in that The method of using - the greedy strategy to select actions in stages is as follows: During the training phase, with probability sample the subnet index from the scanning probability distribution of each subnet through multinomial distribution, and with probability randomly select the subnet index from the uniform distribution; during the evaluation phase, select the subnet index with the highest probability.
8. The network detection method with multi-probe collaboration according to claim 1, characterized in that, The step of updating the local state after the probe executes an action, calculating the global reward through an immediate reward function, storing the experience tuple in the replay buffer, and updating the Critic network and Actor network of the MADDPG algorithm includes: The probe executes the selected action according to the subnet index that is about to execute a scan, receives environmental feedback, updates the state and experience data, and obtains the global reward value through an immediate reward function; After updating the local state and the global state, recycle the experience tuple, and the experience tuple includes: the current state, action, global reward value, and next state; Update the Critic and Actor networks in the MADDPG algorithm according to the recycled experience tuple.
9. The multi-probe collaborative network detection method according to claim 8, wherein, The immediate reward function is: ; Among them, is the immediate reward function, is the indicator function, with boolean value characteristics, is the total number of probes, is the reward factor when a new device is detected, is the penalty factor for taking each probing action, is the penalty factor for the probe being detected or blocked, is the newly discovered device in the i-th scanning action, is the event that the probe is detected or blocked in the i-th scanning action. If the scan is detected or blocked, the value of I(detected_i) is 1, otherwise it is 0; The local state update method is: ; ; Among them, is a vector with a length of 1000, is a bitwise OR operation; The global state update method is: ; The update strategy of the Actor network is: ; The update strategy of the Critic network is: ; Among them, is the parameter gradient of the Actor network, is the expectation of experience replay, is the gradient of the Actor with respect to its own parameters, is the gradient of the Critic with respect to the action, represents the action generated by the policy is the Q value of the Critic network, , refers to a batch of samples randomly sampled from the experience replay pool, are the network parameters, are the target network parameters, is the current state, is the next state, is the current action, is the next action, is the discount factor, is the mean squared error loss of the Critic network, is sampled from the state transition tuples, is the target Critic network for all possible maximum Q value of 10. A network detection system with multi-probe collaboration, characterized in that, The system is used to execute the multi-probe collaborative network detection method described in any one of claims 1 to 9, and the system includes: A probe deployment module, which is used to combine the Border Gateway Protocol with existing Internet of Things device information to obtain multiple autonomous system subsystems, calculate the Internet of Things coverage probability of each autonomous system subsystem using a linear regression model, and determine the deployment location of the probe group through an integer linear programming method; An action acquisition module, which is used to define the local state of each probe, obtain the probability distribution of each subnet through the MADDPG algorithm, and obtain the subnet index that the probe is about to perform a scan on; An algorithm update module, which is used to update the local state after the probe executes an action, calculate the global reward through an immediate reward function, store the experience tuple in the replay buffer, and update the Critic network and Actor network of the MADDPG algorithm.
Citation Information
Patent Citations
Method for detecting P2P network search hot spots based on multi-probe nodes
CN104009891A
Modeling method for discrete counting data based on multi-probe locality sensitive hash negative binomial regression model
CN114297582A
Agent policy learning method with privacy protection in mobile edge computing
WO2024254892A1