Internet of Things equipment detection method and system

By modeling the network scanning process as partially observable Markov decision-making process and clustering algorithm, the scanning strategy is optimized, and the problem of high redundant detection rate in network scanning is solved, and the IP-device mapping changes are efficiently detected under limited resources, which improves the timeliness of detection and resource utilization.

CN120358176AActive Publication Date: 2025-07-22XI AN JIAOTONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510847694.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The prior art lacks intelligent perception of the dynamic characteristics of IP devices in network scanning, resulting in high redundant detection rates, unable to make full use of the limited scanning budget, and difficult to adapt to the dynamic changes in IP-device mapping in the network.

Method used

The mapping change process of scanning IP-device is modeled as part of the observable Markov decision-making process, and the IP pool is divided through the clustering algorithm, a hierarchical action selection mechanism is constructed, and the scanning strategy is optimized based on information gain and time-weighted strategies, and a dynamic management mechanism of the IP pool is established to adapt to network state changes.

Benefits of technology

Maximize the capture of IP-device mapping mutations under a limited scanning budget, reduce the redundant detection rate, improve the timeliness of detection and resource utilization, and adapt to dynamic network changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358176A_ABST
    Figure CN120358176A_ABST
Patent Text Reader

Abstract

The invention relates to an Internet of Things equipment detection method and system, and the method comprises the steps: modeling a mapping change process of scanning IP-equipment into a partially observable Markov decision process, defining a state, an action, observation, state transition, an observation probability and an award, and optimizing a scanning strategy; according to IP static or semi-static characteristics, an IP address space is initially divided by adopting a clustering algorithm to form a manageable IP pool; a hierarchical action selection mechanism is constructed, IP pool detection priority is scheduled through an upper-layer element scheduler based on information gain and a time weighting strategy, and a lower-layer action decision is combined with a historical sequence and a reward function to select a detection action; and an IP pool dynamic management mechanism is established, reclustering is triggered in combination with network state change, and IP pool division is adjusted by updating static and dynamic characteristics. By adopting the technical scheme of the invention, the number of IP-device mapping abrupt changes can be captured to the maximum extent under a limited scanning budget, and the timeliness, effectiveness and resource utilization rate of detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet of Things (IoT) security technologies, and particularly to a method and system for detecting IoT devices. Background Art

[0002] With the popularization of Internet of Things (IoT) devices, the number of devices on the Internet has grown exponentially, and these devices communicate with the network through IP addresses. As an important tool for Internet information retrieval, one of the core tasks of a network search engine is to discover and track the mapping relationship between IP devices and physical devices through active scanning and fingerprint recognition technologies. However, the Internet address space is extremely large (the IPv4 address space is , and the IPv6 address space is ), and the IP device mapping relationship is highly dynamic and variable. For example, a device may change its IP address due to network configuration changes, device migration, or failures, or some devices may only be online during specific time periods. Therefore, in actual network monitoring scenarios, how to efficiently track the evolution of IP device mappings with a limited scanning budget has become a key point in building a timely and accurate network search engine.

[0003] In the prior art, traditional network scanning methods include a sequential scanning strategy that sequentially probes IP addresses in order and a random scanning strategy that randomly selects IP addresses for probing. However, both of the above two traditional network scanning methods have obvious limitations in complex network scenarios. First, traditional network scanning methods lack intelligent perception of the dynamic characteristics of IP devices and are prone to repeatedly probing invalid or long-term static IP addresses, resulting in a high redundancy detection rate and wasting computing resources and bandwidth. Second, traditional network scanning methods usually divide network protocols into different categories and manually define the scanning frequencies of various protocols, failing to fully consider the dynamic change characteristics of IP-device mappings in different networks and unable to optimize the scanning strategy according to the real-time state of the network, thus making it difficult to fully utilize the limited scanning budget.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] To address the above problems, this application provides a method and system for detecting IoT devices, which can maximize the number of captured IP-device mapping mutations with a limited scanning budget and improve the timeliness, effectiveness, and resource utilization rate of detection.

[0006] To achieve the objectives of this application, the following technical solutions are provided in this application:

[0007] In a first aspect, the present application provides an Internet of Things device detection method, including:

[0008] Model the mapping change process of scanning IP-devices as a partially observable Markov decision process (POMDP), define states, actions, observations, state transitions, observation probabilities, and rewards, and optimize the scanning strategy;

[0009] According to the static or semi-static characteristics of IPs, use a clustering algorithm to initially partition the IP address space to form a manageable IP pool;

[0010] Construct a hierarchical action selection mechanism. The upper-level meta-scheduler schedules the detection priorities of the IP pool based on information gain and time-weighted strategies, and the lower-level action decision combines historical sequences and reward functions to select detection actions;

[0011] Establish an IP pool dynamic management mechanism, trigger reclustering in combination with network state changes, and adjust the IP pool partition by updating static and dynamic characteristics.

[0012] In a possible implementation manner, the step of modeling the mapping change process of scanning IP-devices as a partially observable Markov decision process, defining states, actions, observations, state transitions, observation probabilities, and rewards, and optimizing the scanning strategy includes:

[0013] Define the state as the potential true state of the IP pool, including IP activity, device type distribution, and mapping stability;

[0014] Define the action as a collection of discrete actions executable on the IP pool;

[0015] Define the observation as the observation result obtained after executing the action;

[0016] Define the state transition as the probability of the true state evolution;

[0017] Define the observation probability as the probability of obtaining a predetermined observation result after executing a specified action under the target true state;

[0018] Define the reward as an immediate feedback signal based on the action and the observation result;

[0019] Update the belief state through the observation result, calculate the expected reward of each action, and select the action sequence that can maximize the long-term cumulative reward to optimize the scanning strategy.

[0020] In a possible implementation manner, the step of updating the belief state through the observation result, calculating the expected reward of each action, and selecting the action sequence that can maximize the long-term cumulative reward to optimize the scanning strategy includes:

[0021] Iteratively correct the belief estimate of the IP pool state based on the observation results and actions through the observation probability and state transition probability;

[0022] Calculate the expected reward of each action, and select the action sequence that can maximize the long-term cumulative reward as the scanning strategy.

[0023] In a possible implementation, the step of initially partitioning the IP address space using a clustering algorithm according to IP static or semi-static characteristics to form a manageable IP pool includes:

[0024] Select IP address features and represent them as a feature vector containing the selected features. The IP address features include IP geographical location, autonomous system number, and subnet prefix;

[0025] Partition the IP pool through the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm, divide the IP addresses that are density-connected into the same IP pool, and handle the noise points separately or merge them into the nearest cluster to obtain multiple IP pools;

[0026] Assign a unique ID to each IP pool to complete the initial static clustering and form a manageable IP pool.

[0027] In a possible implementation, before the step of partitioning the IP pool through the DBSCAN algorithm, dividing the IP addresses that are density-connected into the same IP pool, handling the noise points separately or merging them into the nearest cluster to obtain multiple IP pools, it further includes:

[0028] Represent each IP address as a feature vector;

[0029] Set the key parameters of the DBSCAN algorithm, and the key parameters are the neighborhood radius and the minimum number of neighbors.

[0030] In a possible implementation, the step of constructing a hierarchical action selection mechanism, where the upper-level meta-scheduler schedules the IP pool probing priority based on information gain and time-weighted strategy, and the lower-level action decision selects the probing action by combining the historical sequence and the reward function, includes:

[0031] Calculate the uncertainty of the belief state before IP pool probing and the expected uncertainty of the posterior belief state after probing through Shannon entropy to obtain the information gain metric;

[0032] Record the last probing time of each IP pool and calculate the time interval from the current time;

[0033] Fuse the information gain and time interval through the priority formula, and select the IP pool with the highest priority for the next step of detection;

[0034] Extract the fixed-length historical sequence of observations and actions included in the IP pool;

[0035] Set a reward function for the selected action, reward the expected behavior, and punish ineffective detection and high costs, and iteratively update the agent;

[0036] Input the historical sequence into the agent to obtain the hidden state at the last time step, and input the hidden state into the subsequent fully connected layer to obtain the Q-values of each action, and select the optimal action to execute through the Q-values of each action.

[0037] In a possible implementation, the calculation formula of the information gain metric is:

[0038] ;

[0039] Among them, is the information gain, is the IP address pool divided by clustering, is the Shannon entropy, is the simplified belief state, is the updated simplified belief state, is the mathematical expectation;

[0040] The calculation formula of the time interval is:

[0041] ;

[0042] Among them, is the time interval, is the current detection time, is the last detection time of the IP pool;

[0043] The calculation formula of the priority formula is:

[0044] ;

[0045] Among them, is the priority, is the third weight, is the information gain, is the fourth weight, is the time interval.

[0046] In a possible implementation, the historical sequence is:

[0047] ;

[0048] Among them, is the observation result, a is the action taken in the previous step, is the current time step, is the length of the historical window;

[0049] The reward function is: ;

[0050] where, is the reward function, assigns different reward values according to the observed change type , is the first weight coefficient, is the second weight coefficient, is the cost associated with executing the action a.

[0051] In a possible implementation manner, the steps of establishing the IP pool dynamic management mechanism, triggering reclustering in combination with network state changes, and adjusting the IP pool division by updating static and dynamic features include:

[0052] Set dynamic reclustering conditions, and perform reclustering when the dynamic reclustering conditions are met. The dynamic reclustering conditions include periodic triggering and event triggering;

[0053] Introduce dynamic features on the basis of the initial static features to form an updated feature vector;

[0054] Use a clustering algorithm to recluster the IP addresses containing dynamic features and adjust the division of the IP pool accordingly.

[0055] In a second aspect, the present application also provides an Internet of Things device detection system for executing the above-mentioned Internet of Things device detection method. The system includes:

[0056] A scanning and modeling module, which is used to model the mapping change process of scanning IP-devices as a partially observable Markov decision process, define states, actions, observations, state transitions, observation probabilities, and rewards, and optimize the scanning strategy;

[0057] An initial clustering module, which is used to initially divide the IP address space using a clustering algorithm according to IP static or semi-static features to form a manageable IP pool;

[0058] A hierarchical selection module, which is used to construct a hierarchical action selection mechanism. The upper-level meta-scheduler schedules the IP pool detection priorities based on information gain and time-weighted strategies, and the lower-level action decision selects detection actions in combination with historical sequences and reward functions;

[0059] A dynamic reclustering module, which is used to establish an IP pool dynamic management mechanism, trigger reclustering in combination with network state changes, and adjust the IP pool division by updating static and dynamic features.

[0060] The technical solution provided by this application may include the following beneficial effects:

[0061] Through an Internet of Things device detection method and system provided by this application, it is possible to efficiently detect and identify IP-device mapping changes of Internet of Things devices by dynamically adjusting the scanning strategy, reduce the redundant detection rate, reduce waste of computing resources and bandwidth, and be able to adapt to network dynamic changes, so as to maximize the number of captured mutations under a limited scanning budget, and improve the timeliness, effectiveness and resource utilization rate of detection.

[0062] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this disclosure. Brief Description of the Drawings

[0063] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. Obviously, the drawings in the following description are only some embodiments of this disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0064] Figure 1 It is a schematic flowchart of an Internet of Things device detection method provided by an embodiment of this application;

[0065] Figure 2 It is a schematic flowchart of step S100 of an Internet of Things device detection method provided by an embodiment of this application;

[0066] Figure 3 It is a schematic flowchart of step S200 of an Internet of Things device detection method provided by an embodiment of this application;

[0067] Figure 4 It is a schematic flowchart of step S300 of an Internet of Things device detection method provided by an embodiment of this application;

[0068] Figure 5 It is a schematic flowchart of step S400 of an Internet of Things device detection method provided by an embodiment of this application;

[0069] Figure 6 It is a schematic structural diagram of an Internet of Things device detection system provided by an embodiment of this application. Detailed Embodiments

[0070] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0071] In this example embodiment, a method for detecting Internet of Things (IoT) devices under feedback sparsity and observation loss conditions for dynamic interference power management is first provided. Referring to Figure 1 as shown, the method for detecting IoT devices may include the following steps:

[0072] Step S100: Model the mapping change process of scanning IP-devices as a partially observable Markov decision process, define states, actions, observations, state transitions, observation probabilities, and rewards, and optimize the scanning strategy.

[0073] Step S200: According to the static or semi-static characteristics of IPs, use a clustering algorithm to initially partition the IP address space to form a manageable IP pool.

[0074] Step S300: Construct a hierarchical action selection mechanism. The upper-level meta-scheduler schedules the detection priorities of the IP pool based on information gain and time-weighted strategies, and the lower-level action decision combines historical sequences and the reward function to select detection actions.

[0075] Step S400: Establish an IP pool dynamic management mechanism, trigger reclustering in combination with network state changes, and adjust the IP pool partition by updating static and dynamic characteristics.

[0076] Through the above method for detecting IoT devices, the scanning process is modeled as a partially observable Markov decision process, the decision complexity is reduced by the IP address range clustering algorithm, a hierarchical action selection mechanism including information gain measurement and time-weighted strategies is constructed to achieve efficient detection scheduling, and an IP dynamic management mechanism is established to adapt to the evolution of the network environment; thereby intelligently scheduling scanning tasks under a limited scanning budget, dynamically adjusting the scanning frequency and target selection, significantly reducing the redundant detection rate, reducing waste of computing resources and bandwidth, adapting to network dynamic changes in real time, maximizing the number of captured IP device mapping mutations, and effectively improving the timeliness, accuracy, and resource utilization efficiency of network search engine detection.

[0077] Next, each step of the above method for detecting IoT devices in this example embodiment will be described in more detail with reference to Figures 2 to 5 .

[0078] In step S100, the process of mapping changes in the scanned IP-devices is modeled as a partially observable Markov decision process, defining the state, action, observation, state transition, observation probability, and reward to optimize the scanning strategy.

[0079] It can be understood that the process of mapping changes in the scanned IP-devices is regarded as a continuous decision-making process and modeled as a partially observable Markov decision process. The goal is to optimize the scanning strategy to capture as many IP address changes as possible.

[0080] In a possible implementation, step S100 may further include the following sub-steps:

[0081] In step S110, the state is defined as the potential true state of the IP pool, including IP activity, device type distribution, and mapping stability.

[0082] It can be understood that the state S is defined as the potential true state of each IP pool. This state is partially observable and contains information such as the activity of IP addresses in the pool, device type distribution, stability of the IP-device mapping, and change frequency.

[0083] In step S120, the action is defined as the set of discrete actions executable on the IP pool.

[0084] It should be noted that the action mainly involves adjusting the detection frequency:

[0085]

[0086] Among them, is the action, is to increase the detection frequency by a preset level, is to decrease by a preset level, is to maintain the current frequency.

[0087] In step S130, the observation is defined as the observation result obtained after executing the action.

[0088] It should be noted that after defining the execution of the detection action, the set of observation information obtained by the agent from the environment ; these observations are partial reflections of the true state. Observation can be a combination or a single value of the following discrete values:

[0089] : A high-priority change is detected.

[0090] : A general change is detected.

[0091] : No significant change is found compared with the previous detection.

[0092] : The probing request for the target IP timed out.

[0093] : The target IP rejected the connection (e.g., received a TCP RST packet).

[0094] : Other probing errors occurred (e.g., the network was unreachable).

[0095] : According to the sampling results of this probing, it is estimated that the proportion of active IP addresses in the pool exceeds or is lower than the threshold.

[0096] In step S140, the state transition is defined as the probability of the true state evolution.

[0097] It should be noted that the probability of the true state evolution is .

[0098] In step S150, the observation probability is defined as the probability of obtaining a predetermined observation result after performing a specified action under the target true state.

[0099] It can be understood that, exemplarily, the observation probability O represents the probability of observing after performing the action under a certain true state . .

[0100] In step S160, the reward is defined as an immediate feedback signal based on the action and the observation result.

[0101] It should be noted that the immediate feedback signal is , and the expression of the reward can be ;

[0102] Among them, assigns different reward values according to the observed change type ω, is the first weight coefficient, is the second weight coefficient, is the cost associated with performing the action a;

[0103] The weight coefficients and need to be adjusted according to the actual scanning budget and the target, and the initial values are = 1, = 1.

[0104] In step S170, the belief state is updated through observation results, the expected reward of each action is calculated, and the action sequence that can maximize the long-term cumulative reward is selected to optimize the scanning strategy.

[0105] It is understandable that since the true state cannot be fully known, by maintaining a belief state , indicating the real state The probability distribution of is estimated by using the hidden state of the agent.

[0106] Furthermore, the step of updating the belief state through observation results, calculating the expected reward of each action, and selecting the action sequence that can maximize the long-term cumulative reward to optimize the scanning strategy includes:

[0107] According to the observation results and actions, the belief estimate of the IP pool state is iteratively corrected through observation probability and state transition probability;

[0108] Calculate the expected reward for each action and select the action sequence that maximizes the long-term cumulative reward as the scanning strategy.

[0109] It is understandable that by maintaining a belief state (Real status of IP pool The probability distribution estimate of , which is implicitly represented by the agent's hidden state), combined with the observations after each detection , based on the Bayesian principle, the belief state is updated, and then the expected reward of each action is calculated, and the action sequence that can maximize the long-term cumulative reward is selected. The benefits and costs of actions are quantified, and the intelligent agent is driven to give priority to high-benefit, low-cost detection actions.

[0110] In step S200, the IP address space is initially divided using a clustering algorithm according to the static or semi-static features of the IP addresses to form a manageable IP pool.

[0111] It is understandable that step S200 divides the massive IP address space into manageable units to reduce the complexity of decision making. The IP address range clustering algorithm is used to estimate the IP address pool to reduce the complexity of subsequent decisions. The distribution of different device types is calculated by scanning records, and the DBSCAN method is used for initial static clustering.

[0112] In a possible implementation, step S200 may further include the following sub-steps:

[0113] In step S210, IP address features are selected and represented as feature vectors containing the selected features, wherein the IP address features include the IP geographic location, the autonomous system number to which it belongs, and the subnet prefix.

[0114] It is understandable that the initial static clustering mainly performs initial clustering based on static or semi-static features such as IP geographical location, autonomous system number (ASN) to which it belongs, subnet prefix, etc.

[0115] In step S220, the IP pool is divided by the DBSCAN algorithm, and IP addresses that are density-connected are divided into the same IP pool. Noise points are processed separately or merged into the nearest cluster, resulting in multiple IP pools.

[0116] It is understandable that since the DBSCAN algorithm does not require specifying the number of clusters in advance, initial clustering is performed using the DBSCAN method by scanning the IP distributions of different device types.

[0117] In step S230, a unique ID is assigned to each IP pool to complete the initial static clustering and form manageable IP pools.

[0118] It should be noted that the result of clustering is to generate a series of IP pools, and each pool contains a group of IP addresses with similar features or close geographical / network locations.

[0119] Furthermore, before step S220, it also includes:

[0120] In step S2201, each IP address is represented as a feature vector.

[0121] In step S2202, the key parameters of the DBSCAN algorithm are set, and the key parameters are the neighborhood radius and the minimum number of neighbors.

[0122] It is understandable that step S2201 is to process the selected IP feature data for data preparation; step S2202 is to set the DBSCAN algorithm parameters, the neighborhood radius eps and the minimum number of neighbors .

[0123] In step S300, a hierarchical action selection mechanism is constructed. The upper-layer meta-scheduler schedules the IP pool detection priorities based on information gain and time-weighted strategies, and the lower-layer action decision combines the historical sequence and the reward function to select detection actions.

[0124] It should be noted that the action selection mechanism consists of two layers, the upper-layer meta-scheduler is responsible for macro-scheduling among multiple IP pools, determining which IP pool to preferentially select for detection or policy adjustment at the next decision time point, and the lower-layer action decision-making is for the IP pool selected by the upper-layer meta-scheduler and is responsible for selecting specific detection actions; the decision-making basis of the upper-layer meta-scheduler includes the information gain metric of the IP pool and the time-weighted strategy. In order to estimate which IP pool to detect can minimize the uncertainty about the state of the pool, thereby helping the meta-scheduler determine the next detection target, information gain calculation is introduced; at the same time, by maintaining a simplified Bayesian belief model, it is used to approximately track the state uncertainty of the IP pool.

[0125] In a possible implementation manner, step S300 may include the following sub-steps:

[0126] In step S310, calculate the belief state uncertainty before IP pool detection and the expected uncertainty of the posterior belief state after detection through Shannon entropy to obtain the information gain metric.

[0127] Furthermore, the calculation formula of the information gain metric is:

[0128] ;

[0129] where is the information gain, is the IP address pool divided by clustering, is the Shannon entropy, is the simplified belief state, is the updated simplified belief state, is the mathematical expectation.

[0130] It should be noted that the Shannon entropy represents the uncertainty contained in the current belief state of the IP pool before detection; the mathematical expectation represents the posterior belief state updated after obtaining a specific observation ω after detection. The information gain represents how much uncertainty is expected to be reduced through detection. After each detection, the belief state is updated according to the observation ω .

[0131] In step S320, record the last detection time of each IP pool and calculate the time interval from the current time.

[0132] Furthermore, the calculation formula of the time interval is:

[0133] ;

[0134] where is the time interval, is the current detection time, It is the last detection time of the IP pool.

[0135] It should be noted that the time-weighted strategy gives priority to the IP pool with the longest interval since the last detection to ensure the coverage of detection.

[0136] In step S330, the information gain and time interval are fused through the priority formula, and the IP pool with the highest priority is selected for the next detection.

[0137] Furthermore, the calculation formula of the priority formula is:

[0138] ;

[0139] Among them, is the priority, is the third weight, is the information gain, is the fourth weight, is the time interval.

[0140] It should be noted that when making decision fusion, the information gain and time weighting are combined to select the next IP pool to be detected.

[0141] It can be understood that steps S310 - S330 are for the upper-layer meta-scheduler to schedule the detection priority of the IP pool based on the information gain and time-weighted strategy.

[0142] In step S340, a fixed-length historical sequence containing observations and actions of the IP pool is extracted.

[0143] Furthermore, the historical sequence is:

[0144] ;

[0145] Among them, is the observation result, a is the action taken in the previous step, is the current time step, is the length of the historical window.

[0146] It can be understood that the lower-layer action decision is for the IP pool selected by the upper-layer meta-scheduler and is responsible for selecting specific detection actions .

[0147] In step S350, a reward function is set for the selected action to reward the expected behavior and punish ineffective detection and high costs, and the agent is iteratively updated.

[0148] Furthermore, the reward function is:

[0149] ;

[0150] Among them, is the reward function, assigns different reward values according to the observed change type ω, is the first weight coefficient, is the second weight coefficient, is the cost associated with performing the action a.

[0151] It can be understood that is the first weight coefficient, the second weight coefficient is adjusted according to the actual scanning budget and the target, and the initial value is = 1, = 1.

[0152] In step S360, the historical sequence is input into the agent to obtain the hidden state at the last time step, and the hidden state is input into the subsequent fully connected layer to obtain the Q-values of each action, and the optimal action is selected and executed through the Q-values of each action.

[0153] Among them, the Q-value is an index output by the decision network for measuring the value of each possible action.

[0154] It can be understood that the constructed sequence is input into the agent, and the agent outputs the hidden state at the last time step ; this hidden state is input into the subsequent fully connected layer, and finally the Q-values corresponding to each possible action are output ; the historical sequence of IP pool i is input into the decision network, introducing the exploration and exploitation strategy, and finally the action to be executed is selected from it .

[0155] It can be understood that steps S340 - S360 are for the lower-level action decision to select the detection action by combining the historical sequence and the reward function.

[0156] In step S400, an IP pool dynamic management mechanism is established, and re-clustering is triggered in combination with the network state change, and the IP pool division is adjusted by updating the static and dynamic features.

[0157] It should be noted that when establishing the IP pool dynamic management mechanism, regularly or when significant network topology changes are detected, in combination with the different mutation frequencies of the IP static features and the IP-device mapping, the re-clustering process is triggered to adapt to the evolution of the network environment. The number and scale of the IP pools are optional configuration parameters.

[0158] In a possible implementation manner, step S400 may include the following sub-steps:

[0159] In step S410, dynamic reclustering conditions are set, and reclustering is performed when the dynamic reclustering conditions are met. The dynamic reclustering conditions include periodic triggering and event triggering.

[0160] It should be noted that periodic triggering can set a fixed time interval for reclustering, such as once a week or once a month; event triggering can be triggered when a major event that may affect the IP address distribution or behavior is detected, such as a significant increase or decrease in the mutation frequency of the IP-device mapping in a certain IP pool, exceeding a preset threshold. Among them, the mutation frequency is estimated by counting the proportion in recent observations and ; detecting a large-scale network topology change; a drastic change in the overall activity of a certain IP pool.

[0161] In step S420, dynamic features are introduced on the basis of the initial static features to form an updated feature vector.

[0162] In step S430, a clustering algorithm is used to recluster the IP addresses containing dynamic features, and the division of the IP pool is adjusted accordingly.

[0163] It can be understood that finally, the division of the IP pool is adjusted according to the new clustering result, and the IP pool list and its related status information managed by the meta-scheduler are updated.

[0164] Furthermore, in this exemplary embodiment, an Internet of Things device detection system is also provided for performing the above-mentioned Internet of Things device detection method. Referring to Figure 6 shown in, the system may include a scanning and modeling module, an initial clustering module, a hierarchical selection module, and a dynamic reclustering module.

[0165] The scanning and modeling module is used to model the mapping change process of scanning IP-devices as a partially observable Markov decision process, define states, actions, observations, state transitions, observation probabilities, and rewards, and optimize the scanning strategy.

[0166] The initial clustering module is used to initially divide the IP address space according to IP static or semi-static features by using a clustering algorithm to form manageable IP pools.

[0167] The hierarchical selection module is used to construct a hierarchical action selection mechanism. The upper-level meta-scheduler schedules the IP pool detection priority based on the information gain and time-weighted strategy, and the lower-level action decision selects the detection action by combining the historical sequence and the reward function.

[0168] The dynamic reclustering module is used to establish an IP pool dynamic management mechanism, trigger reclustering in combination with network state changes, and adjust the IP pool division by updating static and dynamic features.

[0169] Other embodiments of the present disclosure will be readily apparent to those skilled in the art in view of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only illustrative, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

[0170] The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit it. The present application is not limited to the exact structures described above and illustrated in the drawings, and it cannot be considered that the specific implementation of the present application is only limited to these descriptions. For those of ordinary skill in the technical field to which the present application belongs, various changes and modifications made without departing from the concept of the present application should be regarded as falling within the protection scope of the present application.

Claims

1. A method for detecting Internet of Things devices, characterized in that, Including: Model the mapping change process of scanning IP-devices as a partially observable Markov decision process, define states, actions, observations, state transitions, observation probabilities, and rewards, and optimize the scanning strategy. According to the static or semi-static characteristics of IPs, use a clustering algorithm to initially partition the IP address space to form a manageable IP pool. Construct a hierarchical action selection mechanism. The upper-layer meta-scheduler schedules the probing priorities of the IP pool based on information gain and time-weighted strategies, and the lower-layer action decision combines historical sequences and the reward function to select probing actions. Establish a dynamic management mechanism for the IP pool, trigger reclustering in combination with network state changes, and adjust the IP pool partition by updating static and dynamic characteristics.

2. The method for detecting an Internet of Things device according to claim 1, wherein The step of modeling the mapping change process of scanning IP-devices as a partially observable Markov decision process, defining states, actions, observations, state transitions, observation probabilities, and rewards, and optimizing the scanning strategy includes: Define the state as the potential true state of the IP pool, including IP activity, device type distribution, and mapping stability. Define the action as the set of discrete actions executable on the IP pool. Define the observation as the observation result obtained after executing an action. Define the state transition as the probability of the true state evolution. Define the observation probability as the probability of obtaining a predetermined observation result after executing a specified action under the target true state. Define the reward as an immediate feedback signal based on the action and the observation result. Update the belief state through the observation result, calculate the expected reward of each action, and select the action sequence that can maximize the long-term cumulative reward to optimize the scanning strategy.

3. The method for detecting an Internet of Things device according to claim 2, wherein The step of updating the belief state through the observation result, calculating the expected reward of each action, and selecting the action sequence that can maximize the long-term cumulative reward to optimize the scanning strategy includes: According to the observation result and the action, iteratively correct the belief estimate of the IP pool state through the observation probability and the state transition probability. Calculate the expected reward of each action, and select the action sequence that can maximize the long-term cumulative reward as the scanning strategy.

4. The method for detecting an Internet of Things device according to claim 1, wherein The step of initially partitioning the IP address space using a clustering algorithm according to the static or semi-static characteristics of IPs to form a manageable IP pool includes: Select IP address features and represent them as feature vectors containing the selected features. The IP address features include IP geographical location, autonomous system number, and subnet prefix. Partition the IP pool using the DBSCAN algorithm, divide IP addresses that are density-connected into the same IP pool, and handle noise points separately or merge them into the nearest cluster to obtain multiple IP pools. Assign a unique ID to each IP pool to complete the initial static clustering and form a manageable IP pool.

5. The method for detecting an Internet of Things device according to claim 4, wherein Before the step of partitioning the IP pool using the DBSCAN algorithm, dividing IP addresses that are density-connected into the same IP pool, handling noise points separately or merging them into the nearest cluster to obtain multiple IP pools, it also includes: Represent each IP address as a feature vector. Set the key parameters of the DBSCAN algorithm. The key parameters are the neighborhood radius and the minimum number of neighbors.

6. The method for detecting an Internet of Things device according to claim 1, wherein The steps of constructing the hierarchical action selection mechanism, where the upper-layer meta-scheduler schedules the probing priorities of the IP pools based on the information gain and time-weighting strategy, and the lower-layer action decision selects the probing actions by combining the historical sequence and the reward function, include: Calculate the uncertainty of the belief state before IP pool probing and the expected uncertainty of the posterior belief state after probing through Shannon entropy to obtain the information gain metric; Record the last probing time of each IP pool and calculate the time interval from the current time; Fuse the information gain and the time interval through the priority formula to select the IP pool with the highest priority for the next probing; Extract the fixed-length historical sequence of observations and actions contained in the IP pool; Set a reward function for the selected actions, reward the expected behaviors, and punish the ineffective probing and high costs, and iteratively update the agent; Input the historical sequence into the agent to obtain the hidden state at the last time step, and input the hidden state into the subsequent fully connected layer to obtain the Q-values of each action, and select the optimal action to execute through the Q-values of each action.

7. The method for detecting an Internet of Things device according to claim 6, wherein The calculation formula of the information gain metric is: ; Among them, is the information gain, is the IP address pool divided by clustering, is the Shannon entropy, is the simplified belief state, is the updated simplified belief state, is the mathematical expectation; The calculation formula of the time interval is: ; Among them, is the time interval, is the current detection time, is the last detection time of the IP pool; The calculation formula of the priority formula is: ; Among them, is the priority, is the third weight, is the information gain, is the fourth weight, is the time interval.

8. The method for detecting an Internet of Things device according to claim 6, wherein, The historical sequence is: ; where, is the observation result, a is the action taken in the previous step, is the current time step, is the length of the historical window; The reward function is as follows: ; Among them, is the reward function, assigns different reward values according to the observed change type , is the first weight coefficient, is the second weight coefficient, is the cost associated with performing action a.

9. The method for detecting an Internet of Things device according to claim 1, wherein The steps of establishing the IP pool dynamic management mechanism, triggering reclustering in combination with network state changes, and adjusting the IP pool division by updating the static and dynamic features, include: Set the dynamic reclustering conditions, and perform reclustering when the dynamic reclustering conditions are met. The dynamic reclustering conditions include periodic triggering and event triggering; Introduce dynamic features on the basis of the initial static features to form an updated feature vector; Use the clustering algorithm to recluster the IP addresses containing dynamic features and adjust the division of the IP pool accordingly.

10. An Internet of Things device detection system, characterized in that, The system is used to execute the Internet of Things device probing method described in any one of claims 1 to 9. The system includes: A scanning modeling module, which is used to model the mapping change process of scanning IP-devices as a partially observable Markov decision process, define states, actions, observations, state transitions, observation probabilities, and rewards, and optimize the scanning strategy; An initial clustering module, which is used to initially divide the IP address space by using a clustering algorithm according to the IP static or semi-static features to form a manageable IP pool; A hierarchical selection module, which is used to construct a hierarchical action selection mechanism, where the upper-layer meta-scheduler schedules the probing priorities of the IP pools based on the information gain and time-weighting strategy, and the lower-layer action decision selects the probing actions by combining the historical sequence and the reward function; A dynamic reclustering module, which is used to establish an IP pool dynamic management mechanism, trigger reclustering in combination with network state changes, and adjust the IP pool division by updating the static and dynamic features.

Citation Information

Patent Citations

  • Industrial control network automatic defense decision-making method oriented to partially unknown security state

    CN116582330A

  • Rail transit station-level intelligent agent implementation method and system

    CN120017481A

  • Information security management method based on data processing

    CN120128361A

  • Intelligent distribution of data for robotic and autonomous systems

    US20210011461A1