An AI-based real-time detection system and method for Internet of Things IP address status
Through the combination of AI-based deep neural network, principal component analysis and high-order entropy model, the device status recognition problem under the shared IP address of multiple devices in the Internet of Things is solved, and the accurate identification of the online status of the device and the efficient management of the IP address are achieved.
Patent Information
- Application Number
- CN202510667516.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-05-22
AI Technical Summary
In the case where multiple devices share the same export IP address in the Internet of Things devices, the existing IP detection methods cannot accurately distinguish the real online status of each device, resulting in misjudgment of IP address recycling and waste of resources.
Using an AI-based method, the communication data and historical connection records of the device are analyzed through the deep neural network model, combined with principal component analysis and higher-order entropy model, the dynamic changes of data flow and control flow and the uncertainty of traffic distribution are evaluated, so as to achieve accurate identification of suspected offline devices and the recovery of IP addresses.
It improves the accuracy and robustness of device status determination, avoids incorrect operations of IP address recycling, and improves the management efficiency and resource utilization of the Internet of Things network.
Smart Images

Figure CN120186134B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of IP address detection, and more specifically, to an AI-based real-time detection system and method for the IP address status of the Internet of Things. Background Art
[0002] In existing technologies, IoT devices typically communicate with external networks through edge proxies or gateway devices. In such scenarios, the communication traffic of multiple devices may share the same egress IP address, which complicates device status determination. Existing IP detection methods typically rely on IP addresses as unique identifiers, but in the case of the same source IP, they cannot accurately distinguish the true online status of each device. For example, when certain devices have actually been offline for a long time, the response of the proxy device or gateway may cause the detection system to mistakenly determine that the device is still online. In addition, the operating status of the proxy device itself, such as load fluctuations or communication delays, may also distort the detection results.
[0003] Therefore, if the true online status of a single device cannot be accurately identified in a complex scenario where multiple devices have the same source IP, it will lead to misjudgment of IP address recycling, and it will be impossible to effectively distinguish between long-term idle devices and normally online devices. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide an AI-based real-time detection system and method for the Internet of Things IP address status to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] An AI-based real-time detection method for IoT IP address status includes the following steps:
[0007] Collect communication data and historical connection records of each device, perform field extraction and vectorization processing through feature selection algorithms, and generate input feature vectors for deep neural networks;
[0008] Use the trained deep neural network model to perform online inference on the input feature vector, combine the preset condition screening rules to output the initial offline judgment result of the device, and add suspected offline devices to the verification list;
[0009] Analyze the control flow and data flow distribution characteristics of suspected offline devices through principal component analysis to evaluate the dynamic changes in the ratio of data flow to control flow;
[0010] The communication traffic distribution characteristics of suspected offline devices are analyzed using a high-order entropy model to evaluate the uncertainty variation characteristics of traffic over multiple time periods.
[0011] Based on the dynamic changes in the ratio of data flow to control flow and the uncertain changes in traffic in multiple time periods, the suspected offline devices in the verification list are checked to see if they are continuously offline. If so, the IP address reclaim operation is performed.
[0012] In a preferred embodiment, the communication data and historical connection records of each device are collected, and field extraction and vectorization processing are performed using a feature selection algorithm to generate an input feature vector for a deep neural network, including:
[0013] Collect communication data of each device in the IoT network, including transmission protocol type, destination address, source address, port information and packet size;
[0014] The communication data is associated with the historical connection records of the corresponding device and stored. The historical connection records include the device's connection frequency, average connection duration, and protocol usage distribution;
[0015] Extract fields from communication data and historical connection records, including device identification fields, connection timing fields, and traffic distribution fields;
[0016] The extracted fields are vectorized and converted into high-dimensional feature vectors to generate input feature vectors suitable for the deep neural network model.
[0017] In a preferred embodiment, the trained deep neural network model is used to perform online inference on the input feature vector, and the initial offline determination result of the device is output in combination with the preset condition screening rules, and the suspected offline device is included in the verification list, including:
[0018] Loading an offline trained deep neural network model. The deep neural network model is trained based on the device's communication data and historical connection records. It has a multi-layer structure, with each layer consisting of a preset number of nodes.
[0019] The input feature vector generated by vectorization is input into the deep neural network model, and the activation value of the input feature vector at each node is calculated layer by layer;
[0020] Generate online or offline initial judgment results for the device based on the matching results of the feature values output by the deep neural network model and the preset condition screening rules;
[0021] Devices suspected of being offline are recorded in a verification list, which is used to further verify the device's online status in subsequent steps.
[0022] In a preferred embodiment, principal component analysis is used to analyze the control flow and data flow distribution characteristics of suspected offline devices and evaluate the dynamic changes in the ratio of data flow to control flow, including:
[0023] Extract control flow data and data flow data for each suspected offline device in the verification list. Control flow data includes the number of request messages, protocol type, and average message duration of the device. Data flow data includes the number of transmitted packets, packet size, and destination address distribution of the device.
[0024] Construct a feature matrix of the control flow and data flow of suspected offline devices, where each row corresponds to a device and each column is the extracted feature field;
[0025] Based on the feature matrix, the principal component analysis method is used to calculate the linear combination of features and generate the principal component vector to represent the main change trends of control flow and data flow;
[0026] The projection value of each device in the principal component space is calculated to generate the dynamic change characteristics of the ratio of control flow to data flow.
[0027] In a preferred embodiment, the projection value of each device in the principal component space is calculated to generate a dynamic change feature of the ratio of control flow to data flow, specifically:
[0028] Based on the generated principal component vector, the projection value of each device in the principal component space is calculated to generate the dynamic change characteristics of the ratio of control flow to data flow;
[0029] The projection value of the principal component of each device is calculated as follows: ;in, Indicates the The device is in The projection value of the principal components, Indicates the The feature in The linear combination coefficients of the principal components, Indicates the Standardized characteristic values of each device, Indicates the number of features.
[0030] In a preferred embodiment, the communication traffic distribution characteristics of suspected offline devices are analyzed using a high-order entropy model to evaluate the uncertainty variation characteristics of the traffic in multiple time periods, including:
[0031] Extract communication traffic data for each suspected offline device in the verification list over different time periods. The communication traffic data includes the device's total traffic, traffic distribution by destination address, and traffic distribution by protocol type.
[0032] The extracted communication traffic data is grouped by time period to generate a traffic distribution feature matrix for multiple time periods, where each row represents a time period and each column represents a traffic feature field within the corresponding time period;
[0033] The high-order entropy value in each time period is calculated based on the traffic distribution characteristic matrix to describe the uncertainty characteristics of the device traffic distribution;
[0034] The high-order entropy change rate is calculated for the high-order entropy values in all time periods to evaluate the uncertainty variation characteristics of flow in multiple time periods.
[0035] In a preferred embodiment, the high-order entropy value in each time period is calculated based on the traffic distribution characteristic matrix to describe the uncertainty characteristics of the device traffic distribution, specifically:
[0036] The formula for calculating the high-order entropy value of each time period is: ;in, Indicates the The high-order entropy value of the time period, is the order parameter of the higher-order entropy, Indicates the Time period The normalized probability distribution value of the traffic characteristics, Indicates the total number of traffic feature fields, The index of the time period. Indicates the index of the traffic feature field. Indicates the In the time period Normalized probability distribution value of traffic characteristics of power;
[0037] The calculated high-order entropy value for each time period is stored as a time series.
[0038] In a preferred embodiment, based on the dynamic changes in the ratio of data flow to control flow and the uncertainty of traffic changes in multiple time periods, the suspected offline device in the verification list is checked to see if it is continuously offline. If so, the IP address reclaiming operation is performed, including:
[0039] Obtain the projection value and high-order entropy change rate of the principal component, and set the projection threshold corresponding to the projection value of the principal component and the entropy change rate threshold corresponding to the high-order entropy change rate;
[0040] When the projection value of the principal component is less than its corresponding projection threshold, and the high-order entropy change rate is less than its corresponding entropy change rate threshold, it is determined that the suspected offline device in the verification list does not meet the continuous offline condition; otherwise, the suspected offline device in the verification list is determined to be continuously offline;
[0041] If it is determined to be in a continuous offline state, the Internet Protocol address corresponding to the suspected offline device will be recycled and the address usage status record will be updated.
[0042] On the other hand, the present invention provides an AI-based real-time detection system for the status of an Internet of Things IP address, comprising a data acquisition module, an initial determination module, a dynamic analysis module, an entropy value evaluation module, and a status review module;
[0043] Data acquisition module: collects communication data and historical connection records of each device, and performs field extraction and vectorization processing through feature selection algorithms to generate input feature vectors for deep neural networks;
[0044] Initial judgment module: This module uses the trained deep neural network model to perform online inference on the input feature vector, combines it with preset condition screening rules to output the initial offline judgment results of the device, and adds suspected offline devices to the verification list;
[0045] Dynamic Analysis Module: This module analyzes the control flow and data flow distribution characteristics of suspected offline devices through principal component analysis, and evaluates the dynamic changes in the ratio of data flow to control flow.
[0046] Entropy evaluation module: This module uses a high-order entropy model to analyze the traffic distribution characteristics of suspected offline devices and evaluate the uncertainty of traffic changes over multiple time periods.
[0047] Status review module: Based on the dynamic changes in the ratio of data flow to control flow and the uncertain changes in traffic flow in multiple time periods, it verifies whether the suspected offline devices in the verification list are continuously offline. If so, the IP address reclaim operation is performed.
[0048] The technical effects and advantages of the present invention's AI-based real-time detection system and method for Internet of Things IP address status are as follows:
[0049] 1. Compared with existing detection methods that only rely on the unique identification of IP addresses, the present invention can comprehensively analyze the communication data of devices, the dynamic changes of control flows and data flows, and the uncertainty characteristics of traffic distribution in complex scenarios with multiple devices having the same source IP, thereby effectively avoiding misjudgments caused by proxy device responses; by reviewing the multi-dimensional dynamic characteristics of suspected offline devices, it can accurately identify the true online status of a single device, improve the accuracy and robustness of device status judgment, and avoid erroneous operations of IP address recovery.
[0050] 2. While achieving real-time detection, it fully utilizes a variety of artificial intelligence technologies to conduct in-depth learning and analysis of the behavioral characteristics of devices, and can promptly detect abnormal communication patterns and potential offline states of devices. By combining principal component projection values and high-order entropy change rates, it can not only capture fluctuations in the communication behavior of devices, but also accurately quantify the dynamic changes in traffic distribution, making the allocation and recovery of IP address resources more efficient and reliable, avoiding the waste of long-term idle addresses, and improving the management efficiency and resource utilization of the IoT network. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a schematic diagram of a method for real-time detection of IoT IP address status based on AI in the present invention;
[0052] Figure 2 This is a structural diagram of an AI-based real-time detection system for Internet of Things IP address status in the present invention. DETAILED DESCRIPTION
[0053] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0054] Example 1: Figure 1 The present invention provides an AI-based real-time detection method for the status of an Internet of Things IP address, which includes the following steps:
[0055] The communication data and historical connection records of each device are collected, and field extraction and vectorization processing are performed through feature selection algorithms to generate input feature vectors for deep neural networks.
[0056] The trained deep neural network model is used to perform online inference on the input feature vector, and the preset condition screening rules are combined to output the initial offline judgment results of the device, and suspected offline devices are included in the verification list.
[0057] The control flow and data flow distribution characteristics of suspected offline devices are analyzed through principal component analysis to evaluate the dynamic changes in the ratio of data flow to control flow.
[0058] The communication traffic distribution characteristics of suspected offline devices are analyzed through a high-order entropy model to evaluate the uncertainty change characteristics of traffic in multiple time periods.
[0059] Based on the dynamic changes in the ratio of data flow to control flow and the uncertain changes in traffic in multiple time periods, the suspected offline devices in the verification list are checked to see if they are continuously offline. If so, the IP address reclaim operation is performed.
[0060] The communication data and historical connection records of each device are collected, and field extraction and vectorization processing are performed through feature selection algorithms to generate input feature vectors for deep neural networks, including:
[0061] Collect communication data of each device in the IoT network, including transmission protocol type, destination address, source address, port information and packet size:
[0062] In a distributed IoT network, in order to accurately identify the online status of devices, it is first necessary to comprehensively collect the communication data of each device. Communication data contains key interaction information of devices in the network, including but not limited to the following:
[0063] Transmission protocol type: The protocol type used by each device during communication, such as Transmission Control Protocol, User Datagram Protocol, etc. By collecting the transmission protocol type, you can determine the technical protocol characteristics of device communication.
[0064] Destination address: The destination address of each data packet, used to record the destination of the device data flow.
[0065] Source address: The source address of each data packet, identifying the device that originates the data packet.
[0066] Port information: The source port number and destination port number used by the device, which is used to determine the application layer characteristics of the device.
[0067] Data packet size: The number of bytes in each data packet transmitted by the device, reflecting the communication load of the device.
[0068] The above data is collected by monitoring network traffic, ensuring that complete data fields are captured during collection and generating independent data record files for each device for subsequent analysis and processing.
[0069] The communication data is stored in association with the historical connection records of the corresponding device. The historical connection records include the device's connection frequency, average connection duration, and protocol usage distribution:
[0070] To improve the accuracy of device status analysis, the collected communication data needs to be associated and stored with the device's historical connection records to generate a time-series comprehensive data set. The historical connection records include the following:
[0071] Connection frequency: Counts the number of connections made by a device within a specific time range to reflect the device's activity level.
[0072] Average Connection Duration: Calculates the average duration of each device connection to assess whether the device has abnormal connection patterns.
[0073] Protocol Usage Distribution: Analyzes the proportion of various protocols used by devices in different time periods, such as the proportion of Transmission Control Protocol and User Datagram Protocol, to reflect the communication behavior characteristics of the devices.
[0074] When storing data in a linked manner, the communication data should be sorted by timestamp and matched with the corresponding historical connection records. Each data record should retain a clear device identifier to ensure the traceability of device data in subsequent processing.
[0075] Extract fields from communication data and historical connection records, including device identification fields, connection timing fields, and traffic distribution fields:
[0076] A feature selection algorithm is used to extract fields from communication data and historical connection records, retaining key fields for subsequent deep neural network input feature vector generation. The extracted fields include:
[0077] Device identification field: An attribute used to uniquely identify a device, such as the device's physical address or virtual address.
[0078] Connection timing field: describes the device's connection activity in different time periods, including statistical values of the connection start time, end time, and connection duration.
[0079] Traffic distribution field: records the data traffic distribution of the device within different time ranges, including the traffic proportion of the transmission protocol type and the total traffic size.
[0080] During the field extraction process, feature selection algorithms are used to filter and sort data fields, removing redundant fields and retaining fields that significantly contribute to device status determination. For example, by analyzing the distribution of device protocol usage, it is possible to identify whether a device with abnormal behavior has been using only a certain type of protocol for a long time, thereby assisting in determining whether the device is online.
[0081] Perform vectorization on the extracted fields, convert them into high-dimensional feature vectors, and generate input feature vectors suitable for the deep neural network model:
[0082] Perform data normalization on the device's connection timing field and traffic distribution field, and map data fields of different magnitudes to a unified numerical range. For example, limit the value range of the traffic distribution field to between 0 and 1 to facilitate feature learning of the deep learning model.
[0083] Perform one-hot encoding on the device identification field to map the discrete attributes of the device into a high-dimensional sparse vector. For example, for the physical address field, encode each device address into a fixed-length binary vector.
[0084] A time series embedding algorithm is used for the connection timing field to convert the characteristics of the device's connection behavior at different time points into time series vectors, preserving the dynamic change characteristics of the device behavior.
[0085] The normalized, encoded, and embedded fields are concatenated to generate a high-dimensional feature vector for the device.
[0086] The generated feature vector has integrity and structured characteristics, can adapt to the input requirements of the deep neural network model, and serve as the basic data for subsequent equipment status judgment and analysis.
[0087] The trained deep neural network model is used to perform online inference on the input feature vector, and the initial offline judgment result of the device is output based on the preset condition screening rules. The suspected offline devices are then included in the verification list, including:
[0088] Load the offline trained deep neural network model. The deep neural network model is trained based on the device's communication data and historical connection records. It has a multi-layer structure, with each layer consisting of a preset number of nodes:
[0089] Each node contains weight parameters and activation functions, which are used to learn the behavioral characteristics of the device. Specifically, the training goal of the deep neural network is to minimize the loss function, which is used to measure the difference between the model's prediction results and the actual device state label. The optimization formula during the training process is as follows: ;in, Represents the value of the loss function; Indicates the number of training samples; Indicates the The actual state label of the training sample (1 for online and 0 for offline); Indicates the The input feature vector of training samples; Represents the model prediction result, which is determined by the weight parameters and bias parameters; represents the weight parameter; Represents the bias parameter.
[0090] After model training is completed, it is stored in a loadable file format and reloaded during the online inference phase.
[0091] The input feature vector generated by vectorization is input into the deep neural network model, and the activation value of the input feature vector at each node is calculated layer by layer:
[0092] After loading the deep neural network model, the input feature vector generated by the previous steps is fed into the model for layer-by-layer calculation to obtain the prediction results of the device.
[0093] The specific reasoning process is as follows:
[0094] The input feature vector represents the communication behavior characteristics of the device and the extracted values of the historical connection records.
[0095] For each hidden layer The output value calculation formula of each node is as follows: ;in, Indicates the hidden layer The activation value of a node reflects the activation state of the input feature at the node; Represents an activation function, such as a ReLU function or a Sigmoid function; Indicates the first The feature pair The weight of each node is used to describe the relationship between the input features and the nodes; Represents the feature value of a dimension in the input feature vector, which is extracted from the device's communication behavior data and historical connection records; Indicates the The bias parameter of each node is used to adjust the calculation result of the node activation value; represents the dimension of the input feature vector, that is, the number of input features; is an index variable used to identify the first Feature positions, ranging from 1 to ,in is the dimension of the feature vector; The index variable representing the hidden layer node is used to identify the node whose activation value is currently being calculated.
[0096] The calculation results are passed to the next layer of the model in sequence until the output layer generates the final result.
[0097] In the output layer, the deep neural network model performs a binary classification prediction of the device's online status based on the activation values calculated by the hidden layer. The specific process is as follows: the output layer receives all node activation values transmitted by the hidden layer and generates a score for each category through weighted calculation. The classification scores are then processed using a normalization function to obtain a probability value for each category, indicating the likelihood that the device belongs to a certain status (for example, "online" or "offline"). By comparing the probability values of the two categories, the output layer ultimately determines the device's online status. If the probability value of a category exceeds a preset classification threshold (for example, 0.5), the device is classified as that status. The binary classification threshold can be adjusted according to actual scenario requirements to optimize the accuracy and sensitivity of the prediction. The device's online status prediction result is then used for further judgment and verification processes.
[0098] Generate the initial online or offline judgment result of the device based on the matching results of the feature values output by the deep neural network model and the preset condition screening rules:
[0099] Filtering rules include but are not limited to the following conditions:
[0100] Communication behavior pattern threshold range: whether the device's communication behavior characteristics exceed the reasonable range of historical behavior;
[0101] Degree of deviation from historical connection records: whether the current connection behavior of the device differs significantly from the historical records;
[0102] Abnormal protocol usage distribution characteristics: whether the device's communication ratio using different protocol types is abnormal.
[0103] The device status probability values output by the deep neural network are matched with the screening rules one by one. If the device meets multiple abnormal conditions, it is determined to be a suspected offline device.
[0104] Based on the matching results of the filtering rules, the device is marked as "online" or "suspected offline".
[0105] Devices suspected of being offline are recorded in the verification list. The verification list is used for further review of the device's online status in subsequent steps:
[0106] The validation list is stored in a table format, and each record includes the following fields:
[0107] Device ID: Attributes that uniquely identify a device, such as its physical address;
[0108] Preliminary determination result: indicates the current status of the device, such as suspected offline;
[0109] Abnormal Feature Description: Lists the main criteria for determining that the device is suspected to be offline, such as abnormal protocol distribution and connection timing fluctuations.
[0110] Verify that the list is updated in real time, adding new records or modifying device status tags to ensure data integrity.
[0111] Principal component analysis is used to analyze the control flow and data flow distribution characteristics of suspected offline devices and evaluate the dynamic changes in the ratio of data flow to control flow, including:
[0112] Extract the control flow data and data flow data of each suspected offline device in the verification list. The control flow data includes the number of request messages, protocol type, and average message duration of the device. The data flow data includes the number of transmitted data packets, packet size, and destination address distribution of the device:
[0113] Number of request messages: This counts the total number of request messages sent by a device within a certain period of time, reflecting the control flow communication strength of the device.
[0114] Protocol Type: records the network protocol type involved in the device control flow, such as Transmission Control Protocol and User Datagram Protocol, to identify the communication mode.
[0115] Average packet duration: calculates the average duration of each control flow packet on the device, reflecting the control flow behavior characteristics of the device.
[0116] Number of transmitted data packets: Counts the total number of data packets transmitted by the device in the data flow, describing the load situation of the device.
[0117] Packet size: records the size of each packet and calculates the average and variance to reflect the distribution characteristics of the data flow.
[0118] Destination address distribution: Statistics on the types and proportions of destination addresses of device data flow transmission, used to analyze the data flow direction characteristics of the device.
[0119] Construct a feature matrix of the control flow and data flow of suspected offline devices, where each row corresponds to a device and each column is the extracted feature field:
[0120] Each row of the feature matrix represents the data of a device; each column represents the extracted feature fields, including the number of request messages, protocol type, average message duration, number of transmitted data packets, data packet size, and destination address distribution.
[0121] For example, a feature matrix contains three devices and five feature fields. The matrix dimensions are 3×5, with each row representing a device and each column representing a feature field. For example, the value in row 1, column 2 represents the data in feature field 2 for device 1 (e.g., the number of request packets is 50). Similarly, the value in row 3, column 4 represents the data in feature field 4 for device 3 (e.g., the packet size is 1024 bytes).
[0122] To eliminate the influence of the eigenvalue dimension, each column of the feature matrix is normalized. The normalized feature matrix is used in the subsequent principal component analysis to ensure that each feature field contributes evenly to the analysis.
[0123] Based on the feature matrix, the principal component analysis method is used to calculate the linear combination of features and generate the principal component vector to represent the main change trends of control flow and data flow:
[0124] The principal component is a linear combination of the feature matrix that is used to maximize the variance explained by the data. The calculation formula for the principal component is as follows: ;in, Indicates the principal components, Indicates the The feature in The linear combination coefficients of the principal components (principal component vectors), represents the standardized eigenvalue; Indicates the number of features.
[0125] The explained variance of each principal component was calculated, and the first few principal components whose cumulative variance contribution rate reached 95% were selected as representatives of the main change trend.
[0126] The generated principal component vector is stored for subsequent analysis of dynamic change characteristics.
[0127] Calculate the projection value of each device in the principal component space to generate the dynamic change characteristics of the ratio of control flow to data flow:
[0128] Based on the generated principal component vector, the projection value of each device in the principal component space is calculated to generate the dynamic change characteristics of the ratio of control flow to data flow.
[0129] The projection value of the principal component of each device is calculated as follows: ;in, Indicates the The device is in The projection value of the principal components, Indicates the The feature in The linear combination coefficients of the principal components, Indicates the Standardized characteristic values of each device.
[0130] The projection values of the principal components represent the changing characteristics of the control flow and data flow ratios of a device in the principal component space and serve as the primary basis for dynamic change characteristics. The projection values for each device are stored as a numerical vector, providing input data for subsequent review steps.
[0131] It's worth noting that the formula for calculating the principal component and the formula for its projection value are mathematically identical: both are linear combinations of features. However, the principal component represents the global representation of all device feature data, used to extract key trends in feature fields and perform dimensionality reduction. The projection value, on the other hand, is the projection of the principal component onto a single device. It focuses on describing the device's dynamic characteristics along the principal component's direction, quantifying changes in device behavior and enabling device status analysis.
[0132] A larger principal component projection value indicates a more significant dynamic change in the ratio of data flow to control flow, indicating unusual fluctuations in the device's communication behavior and possible deviations from normal operation. For suspected offline devices, a larger principal component projection value indicates more pronounced abnormal changes and a higher likelihood of persistent offline status. Therefore, the principal component projection value is an important indicator of offline status.
[0133] The high-order entropy model is used to analyze the traffic distribution characteristics of suspected offline devices and evaluate the uncertainty variation characteristics of traffic over multiple time periods, including:
[0134] Extract the communication traffic data of each suspected offline device in the verification list in different time periods. The communication traffic data includes the total traffic of the device, traffic distribution by destination address, and traffic distribution by protocol type:
[0135] Total traffic: Statistics on the total traffic value of the device in each time period, used to reflect the overall load of device communication.
[0136] Destination address traffic distribution: records the amount of traffic sent by the device to different destination addresses in each time period and calculates the traffic proportion of the destination address.
[0137] Traffic distribution by protocol type: This statistics displays the traffic proportions of different communication protocols (such as Transmission Control Protocol and User Datagram Protocol) used by the device in each time period.
[0138] The extracted communication traffic data is grouped by time period to generate a traffic distribution feature matrix for multiple time periods, where each row represents a time period and each column represents the traffic feature field within the corresponding time period:
[0139] Each row of the feature matrix corresponds to a time period, and each column represents a traffic feature field (including total traffic, traffic distribution by destination address, traffic distribution by protocol type, etc.).
[0140] Assuming a device has eight time periods, each of which contains four traffic feature fields, the dimension of the feature matrix is 8 × 4. Each element in the matrix represents a specific traffic feature value in a certain time period.
[0141] In order to avoid the influence of eigenvalues on the analysis results due to the difference in order of magnitude, the eigenvalues of each column in the matrix are normalized, and the characteristic matrix generated after normalization is used for the subsequent calculation of high-order entropy values.
[0142] The high-order entropy value in each time period is calculated based on the traffic distribution characteristic matrix to describe the uncertainty characteristics of the device traffic distribution:
[0143] The formula for calculating the high-order entropy value of each time period is: ;in, Indicates the The high-order entropy value of each time period; is the order parameter of the higher-order entropy, which is used to adjust the sensitivity of the entropy value to the traffic distribution; Indicates the Time period Normalized probability distribution value of traffic characteristics; Indicates the total number of traffic feature fields; The index of the time period, which is used to identify the time range for calculating the current high-order entropy value; Indicates the index of the traffic feature field; Indicates the In the time period Normalized probability distribution value of traffic characteristics of The power is used to adjust the weight of different probability values to the overall uncertainty in entropy calculation.
[0144] Among them, the order parameter of the high-order entropy is used to adjust the sensitivity of the entropy value to the traffic distribution. When the order value is greater than 1, the influence of the high-probability component is emphasized more; when the order value is less than 1 but greater than 0, more attention is paid to the changes in the low-probability component, thereby adapting to the analysis needs of different distribution characteristics.
[0145] The calculation formula is: ;in, represents the normalized eigenvalue.
[0146] The larger the high-order entropy value, the greater the uncertainty of the traffic distribution in that time period; the smaller the entropy value, the more concentrated the traffic distribution. The calculated high-order entropy value for each time period is stored as a time series for subsequent evaluation of dynamic change characteristics.
[0147] The high-order entropy change rate is calculated for the high-order entropy values in all time periods to evaluate the uncertainty change characteristics of the flow in multiple time periods:
[0148] Calculate the high-order entropy change rate. The high-order entropy change rate is the absolute value of the high-order entropy value change rate. The absolute value of the high-order entropy value change rate refers to the amplitude of the high-order entropy value change in adjacent time periods. It is used to quantify the severity of traffic distribution uncertainty in different time periods.
[0149] A high rate of change in high-order entropy indicates that the device's traffic distribution has significantly changed between time periods, reflecting the dramatic fluctuations in the uncertainty of traffic distribution. Drastic changes in high-order entropy values indicate unstable communication behavior, potentially intermittent anomalies or complete interruption of communication, further increasing the likelihood that the device will be identified as persistently offline.
[0150] Based on the dynamic changes in the ratio of data flow to control flow and the uncertainty of traffic changes over multiple time periods, the system verifies whether the suspected offline devices in the verification list are continuously offline. If so, it performs IP address reclaiming operations, including:
[0151] Set the projection threshold corresponding to the projection value of the principal component, and set the entropy change rate threshold corresponding to the high-order entropy change rate.
[0152] Among them, the projection threshold is set according to the distribution of principal component projection values of normal and abnormal communication behaviors of devices in historical data. The projection value range of normal devices is selected as a reference, and an upper limit value of more than 90% of the projection value of normal devices is set as the threshold to ensure a high anomaly detection sensitivity; the entropy change rate threshold is set according to the historical data of the change rate of high-order entropy values of normal devices in multiple time periods, and the 95% quantile of the absolute value of the change rate is taken as the threshold to cover the fluctuation range of the vast majority of normal devices and capture the violent fluctuation characteristics of abnormal devices.
[0153] When the projection value of the principal component is less than its corresponding projection threshold and the high-order entropy change rate is less than its corresponding entropy change rate threshold, it is determined that the suspected offline device in the verification list does not meet the continuous offline condition; otherwise, it is determined that the suspected offline device in the verification list is continuously offline.
[0154] If the device is determined to be in a persistent offline state, the Internet Protocol address corresponding to the suspected offline device will be recycled and the address usage status record will be updated:
[0155] When a device is determined to be persistently offline, the system sends a reclaim request to the address management module, marking the device's corresponding Internet Protocol address as "pending reclaim" and removing it from the existing allocation pool. Subsequently, the address usage record is updated, changing the device's online status to "offline," and recording the time and reason for the reclaim. The reclaimed Internet Protocol address is stored in the address reclaim pool, awaiting reallocation. Simultaneously, relevant systems are notified to synchronize address status updates to ensure efficient and accurate network resource allocation.
[0156] It is worth noting that the present invention uses a deep neural network model for online reasoning to achieve intelligent preliminary judgment of the online status of the device; at the same time, combined with principal component analysis and high-order entropy models, artificial intelligence technology is used to evaluate the dynamic changes in the uncertainty of the data flow and control flow ratio and traffic distribution, thereby improving the accuracy and efficiency of detection. For the communication data and historical connection records of devices in the Internet of Things environment, including the total traffic, target address and protocol type traffic of the device, by analyzing these data features related to the Internet Protocol address, it is possible to identify whether the device is in an offline state, thereby ensuring the effective use of Internet Protocol address resources. By performing field extraction, online reasoning and dynamic analysis on the collected data, and combining preset conditions to determine and review suspected offline devices, real-time monitoring of the device's online status and timely recovery of Internet Protocol addresses are achieved. With artificial intelligence as the core, for the Internet of Things scenario, intelligent and real-time detection and management of Internet Protocol address status are achieved.
[0157] Example 2: The difference between Example 2 of the present invention and Example 1 is that this example introduces an AI-based real-time detection system for the IP address status of the Internet of Things.
[0158] Figure 2 A structural schematic diagram of an AI-based real-time detection system for the status of an Internet of Things IP address is given in the present invention. The AI-based real-time detection system for the status of an Internet of Things IP address includes a data acquisition module, an initial judgment module, a dynamic analysis module, an entropy value evaluation module, and a status review module.
[0159] Data acquisition module: collects communication data and historical connection records of each device, and performs field extraction and vectorization processing through feature selection algorithm to generate input feature vectors for deep neural network.
[0160] Initial judgment module: Use the trained deep neural network model to perform online inference on the input feature vector, combine the preset condition screening rules to output the device's offline initial judgment results, and include suspected offline devices in the verification list.
[0161] Dynamic Analysis Module: Analyzes the control flow and data flow distribution characteristics of suspected offline devices through principal component analysis, and evaluates the dynamic changes in the ratio of data flow to control flow.
[0162] Entropy evaluation module: Analyzes the distribution characteristics of communication traffic of suspected offline devices through a high-order entropy model and evaluates the uncertainty change characteristics of traffic in multiple time periods.
[0163] Status review module: Based on the dynamic changes in the ratio of data flow to control flow and the uncertain changes in traffic flow in multiple time periods, it verifies whether the suspected offline devices in the verification list are continuously offline. If so, the IP address reclaim operation is performed.
[0164] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.
[0165] The above embodiments can be implemented in whole or in part via software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in the embodiments of this application are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0166] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0167] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0168] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0169] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0170] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0171] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0172] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
[0173] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A real-time detection method for Internet of Things IP address status based on AI, characterized in that: The steps include: Collect communication data and historical connection records of each device, perform field extraction and vectorization processing through feature selection algorithms, and generate input feature vectors for deep neural networks; Use the trained deep neural network model to perform online inference on the input feature vector, combine the preset condition screening rules to output the initial offline judgment result of the device, and add suspected offline devices to the verification list; Principal component analysis is used to analyze the control flow and data flow distribution characteristics of suspected offline devices and evaluate the dynamic changes in the ratio of data flow to control flow, including: Extract the control flow data and data flow data of each suspected offline device in the verification list; Construct a feature matrix of the control flow and data flow of suspected offline devices, where each row corresponds to a device and each column is the extracted feature field; Based on the feature matrix, the principal component analysis method is used to calculate the linear combination of features and generate the principal component vector to represent the main change trends of control flow and data flow; Calculate the projection value of each device in the principal component space to generate the dynamic change characteristics of the ratio of control flow to data flow, specifically: Based on the generated principal component vector, the projection value of each device in the principal component space is calculated to generate the dynamic change characteristics of the ratio of control flow to data flow; The projection value of the principal component of each device is calculated as follows: ;in, Indicates the The device is in The projection value of the principal components, Indicates the The feature in The linear combination coefficients of the principal components, Indicates the Standardized characteristic values of each device, Indicates the number of features; The communication traffic distribution characteristics of suspected offline devices are analyzed using a high-order entropy model to evaluate the uncertainty variation characteristics of traffic over multiple time periods. Based on the dynamic changes in the ratio of data flow to control flow and the uncertain changes in traffic in multiple time periods, the suspected offline devices in the verification list are checked to see if they are continuously offline. If so, the IP address reclaim operation is performed.
2. The method for real-time detection of Internet of Things IP address status based on AI according to claim 1 is characterized in that: The communication data and historical connection records of each device are collected, and field extraction and vectorization processing are performed through feature selection algorithms to generate input feature vectors for deep neural networks, including: Collect communication data of each device in the IoT network, including transmission protocol type, destination address, source address, port information and packet size; The communication data is associated with the historical connection records of the corresponding device and stored. The historical connection records include the device's connection frequency, average connection duration, and protocol usage distribution; Extract fields from communication data and historical connection records, including device identification fields, connection timing fields, and traffic distribution fields; The extracted fields are vectorized and converted into high-dimensional feature vectors to generate input feature vectors suitable for the deep neural network model.
3. The method for real-time detection of Internet of Things IP address status based on AI according to claim 1 is characterized in that: The trained deep neural network model is used to perform online inference on the input feature vector, and the initial offline judgment result of the device is output based on the preset condition screening rules. The suspected offline devices are then included in the verification list, including: Loading an offline trained deep neural network model. The deep neural network model is trained based on the device's communication data and historical connection records. It has a multi-layer structure, with each layer consisting of a preset number of nodes. The input feature vector generated by vectorization is input into the deep neural network model, and the activation value of the input feature vector at each node is calculated layer by layer; Generate online or offline initial judgment results for the device based on the matching results of the feature values output by the deep neural network model and the preset condition screening rules; Devices suspected of being offline are recorded in a verification list, which is used to further verify the device's online status in subsequent steps.
4. The method for real-time detection of Internet of Things IP address status based on AI according to claim 1 is characterized in that: Control flow data includes the number of request messages, protocol type, and average message duration of the device, and data flow data includes the number of transmitted data packets, packet size, and destination address distribution of the device.
5. The method for real-time detection of Internet of Things IP address status based on AI according to claim 1 is characterized in that: The high-order entropy model is used to analyze the traffic distribution characteristics of suspected offline devices and evaluate the uncertainty variation characteristics of traffic over multiple time periods, including: Extract communication traffic data for each suspected offline device in the verification list over different time periods. The communication traffic data includes the device's total traffic, traffic distribution by destination address, and traffic distribution by protocol type. The extracted communication traffic data is grouped by time period to generate a traffic distribution feature matrix for multiple time periods, where each row represents a time period and each column represents a traffic feature field within the corresponding time period; The high-order entropy value in each time period is calculated based on the traffic distribution characteristic matrix to describe the uncertainty characteristics of the device traffic distribution; The high-order entropy change rate is calculated for the high-order entropy values in all time periods to evaluate the uncertainty variation characteristics of flow in multiple time periods.
6. The method for real-time detection of Internet of Things IP address status based on AI according to claim 5 is characterized in that: The high-order entropy value in each time period is calculated based on the traffic distribution characteristic matrix to describe the uncertainty characteristics of the device traffic distribution. Specifically: The formula for calculating the high-order entropy value of each time period is: ;in, Indicates the The high-order entropy value of the time period, is the order parameter of the higher-order entropy, Indicates the Time period The normalized probability distribution value of the traffic characteristics, Indicates the total number of traffic feature fields, The index of the time period. Indicates the index of the traffic feature field. Indicates the In the time period Normalized probability distribution value of traffic characteristics of power; The calculated high-order entropy value for each time period is stored as a time series.
7. The method for real-time detection of Internet of Things IP address status based on AI according to claim 1, characterized in that: Based on the dynamic changes in the ratio of data flow to control flow and the uncertainty of traffic changes over multiple time periods, the system verifies whether the suspected offline devices in the verification list are continuously offline. If so, it performs IP address reclaiming operations, including: Obtain the projection value and high-order entropy change rate of the principal component, and set the projection threshold corresponding to the projection value of the principal component and the entropy change rate threshold corresponding to the high-order entropy change rate; When the projection value of the principal component is less than its corresponding projection threshold, and the high-order entropy change rate is less than its corresponding entropy change rate threshold, it is determined that the suspected offline device in the verification list does not meet the continuous offline condition; otherwise, the suspected offline device in the verification list is determined to be continuously offline; If it is determined to be in a continuous offline state, the Internet Protocol address corresponding to the suspected offline device will be recycled and the address usage status record will be updated.
8. An AI-based real-time detection system for the status of an Internet of Things IP address, used to implement the AI-based real-time detection method for the status of an Internet of Things IP address according to any one of claims 1 to 7, characterized in that: It includes data acquisition module, initial judgment module, dynamic analysis module, entropy value evaluation module and status review module; Data acquisition module: collects communication data and historical connection records of each device, and performs field extraction and vectorization processing through feature selection algorithms to generate input feature vectors for deep neural networks; Initial judgment module: This module uses the trained deep neural network model to perform online inference on the input feature vector, combines it with preset condition screening rules to output the initial offline judgment results of the device, and adds suspected offline devices to the verification list; Dynamic Analysis Module: This module analyzes the control flow and data flow distribution characteristics of suspected offline devices through principal component analysis, and evaluates the dynamic changes in the ratio of data flow to control flow. Entropy evaluation module: This module uses a high-order entropy model to analyze the traffic distribution characteristics of suspected offline devices and evaluate the uncertainty of traffic changes over multiple time periods. Status review module: Based on the dynamic changes in the ratio of data flow to control flow and the uncertain changes in traffic flow in multiple time periods, it verifies whether the suspected offline devices in the verification list are continuously offline. If so, the IP address reclaim operation is performed.
Citation Information
Patent Citations
IP address checking method and device based on port state data
CN119383171A