Method for identifying and detecting data of Internet of Things
By generating comparison time-series curves in the IoT system and combining multi-scale spurious event feature vector detection and window multi-source data verification, the problem of spurious detection of IoT sensor data is solved, enabling accurate identification and differentiated processing of spurious data, and improving the security and stability of the system.
Patent Information
- Application Number
- CN202511483963.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-17
AI Technical Summary
IoT sensor data is vulnerable to spoofing attacks during transmission and acquisition, leading to system misrepresentation and security incidents. Existing technologies are insufficient to effectively detect and respond to spoofing data in the complex and ever-changing IoT environment.
By generating comparison time-series curves through feature matching of operating data, and combining multi-scale fake event feature vector detection and window multi-source data verification mechanism, we can achieve accurate identification and differentiated processing of fake events in IoT data.
It improves the accuracy and reliability of fake data detection, effectively identifies various types of fake data attacks, reduces false alarms, saves computing resources, and ensures the security and stability of the system.
Smart Images

Figure CN120951003A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data recognition technology, and in particular to a method for identifying and detecting Internet of Things (IoT) data. Background Technology
[0002] In the Internet of Things (IoT), false data refers to data manipulated or falsified to mislead or deceive system users and other stakeholders. IoT sensor data is susceptible to various forms of passive manipulation during transmission and acquisition, leading to data distortion or the generation of false data. False data not only affects the normal operation of the system and the accuracy of decision-making but can also cause security incidents and economic losses.
[0003] The impact of fake data on IoT systems can be catastrophic. It can cause control systems to respond incorrectly to certain situations, leading to safety incidents and economic losses. Because the IoT environment is inherently interconnected and universal, fake data in the IoT is difficult to detect. Besides improving the security of IoT data during transmission and collection, it's possible to build fake data feature identification models and thresholds based on empirical values. However, these methods have poor adaptability and struggle to cope with the complex and ever-changing IoT environment.
[0004] Therefore, how to effectively detect spoofed data in IoT sensor data has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides a method for identifying and detecting IoT data. It generates comparison time-series curves by matching operating condition data features, and combines multi-scale fake event feature vector detection and window multi-source data verification mechanism to achieve accurate identification and differentiated processing of fake events in IoT data.
[0006] This application provides a method for identifying and detecting Internet of Things (IoT) data, including: S101, obtain the target data time-series curve generated by fitting the target data sequence collected by the target sensor in the current detection cycle, and match the target data comparison time-series curve according to the working condition data characteristics of the target sensor. S102: Extract the multi-scale false event feature vector of each sampling point in the time series curve of the target data, input it into the pre-trained data false identification model, and output the false event result associated with the target data of the corresponding sampling point; S103, using a preset window multi-source data verification mechanism, obtains decision indicators for false event results at each sampling point in the target data time series curve; S104, based on the decision index of false event results of sampling points in the target data time series curve, detects the target data of sampling points in the current detection cycle of the target sensor.
[0007] Preferably, the operating condition data features are represented as follows: , Here, B represents the operating condition data characteristics of the target sensor during the current detection cycle, L represents the CPU load, and R represents the range of the target data time-series curve. The mean of the time series curve of the target data. The standard deviation of the target data time series curve. It is a frequency domain feature.
[0008] Preferably, the step of matching the target data with the comparison time-series curve based on the operating condition data characteristics of the target sensor specifically includes: S201, based on the operating mode of the device to which the current target sensor belongs, obtain the target data time-series curves for all detection cycles in the history of the target sensor under the same operating mode, and filter out the curves that are genuine and free from false attack injections, as a set of historical time-series curves. , For the j-th historical time series curve, record the historical operating condition data features corresponding to each historical time series curve to obtain the historical operating condition data feature set. , The historical operating condition data characteristics of the j-th historical time series curve; S202, based on the feature set of historical operating condition data, uses a preset clustering algorithm to group the historical time series curve set into N groups. And set working condition labels for each class group; S203, obtain the operating condition data characteristics of the target sensor in the current detection cycle, calculate its similarity with the operating condition label of each class group, and take the class group with the highest similarity as the target class group; S204. Based on all historical time-series curves in the target class group, obtain the high-fidelity standard curve under the label of this working condition, and match it as the comparison time-series curve of the target data.
[0009] Preferably, the multi-scale false event feature vector includes single-point deviation features, window pattern features, and frequency domain anomaly features, represented as follows: , This is the feature vector of the multi-scale spoof event at the k-th sampling point. For the single-point deviation feature of the k-th sampling point, For the window pattern features of the k-th sampling point, This represents the frequency domain anomaly feature of the k-th sampling point.
[0010] Preferably, the single-point deviation feature is determined based on the target data time series curve and the target data at the kth sampling point in the comparison time series curve; The window mode feature is determined based on the target data time series curve and the target data within a preset window length in the comparison time series curve; The frequency domain anomaly features are determined based on the frequency domain features of the target data within a preset window length in the target data time series curve and the comparison time series curve.
[0011] Preferably, the window mode feature specifically includes:
[0012] This refers to the window mode feature, where w is the preset window length. For the target data time series curve window The target data of the z-th sampling point. To compare timing curves in the window The mean of the target data.
[0013] Preferably, the pre-trained fake data identification model is obtained in the following way: S301, obtain the time-series curves of historical target data within a large number of historical detection cycles of the target sensor and the identification type of target data at all historical sampling points; S302, generate multi-scale spoof event feature vectors for all historical sampling points, and label them with the label content set as spoof event results; S303 uses all labeled multi-scale fake event feature vectors as the training set to train a pre-selected neural network structure, thus obtaining the final data fake detection model.
[0014] Preferably, in step S302, the labeled false event result is set as: quadruple label. When the data is real, When it is a single false event, When it is a series of false events, When it is a pattern attack event, .
[0015] Preferably, the preset window multi-source data verification mechanism specifically includes: S401, based on the window corresponding to the kth sampling point in the current detection period. Extract the associated data sequence of the target sensor within the window; S402, based on the target data sequence of the target sensor within the window and the associated data sequence of the associated sensor within the window, obtain the window cooperative deviation index between the target sensor and each associated sensor; S403, based on the window collaborative deviation index between the target sensor and each associated sensor, uses a preset false event result verification algorithm to output the corresponding decision index.
[0016] Preferably, the preset false event result verification algorithm is set as follows:
[0017] in, This is the decision metric for the false event result at the k-th sampling point of the target sensor within the current detection period. To provide a comprehensive coordination deviation index, Q represents the total number of sensors currently associated with the target sensor. This refers to the window collaborative deviation index corresponding to the k-th sampling point of the target sensor within the current detection period. The preset sensitivity coefficient controls the attenuation of the deviation on the decision indicator.
[0018] One or more technical solutions provided in this application have at least the following technical effects or advantages: By capturing the "static" and "dynamic" attributes of device status, similar operating conditions are grouped and labeled, enabling the comparison time-series curves to automatically adapt to different device states, such as battery depletion or load fluctuations. This avoids baseline drift caused by device malfunctions or environmental interference, ensuring the accuracy of the comparison benchmark. The comparison time-series curves are generated based on historical real curves, representing the ideal state free from spoofing attacks, thereby improving the detection accuracy of continuous or periodic forgery patterns. Simultaneously, by quantifying the similarity between the current operating condition and various operating condition labels, the comparison time-series curves are generated on demand, avoiding blind comparisons with a large number of historical time-series curves, saving computational resources, and increasing the matching rate. Since the characteristics of the operating condition data reflect the true operating status of the equipment, false alarms caused by noise or interference can be reduced, improving the accuracy of the monitoring system.
[0019] Multi-scale fake event feature vectors include single-point deviation features, window pattern features, and frequency domain anomaly features. By analyzing the change patterns of data at different scales, the features of fake events can be comprehensively captured, and various types of fake data attacks can be effectively detected, such as single-point fake events, continuous fake events, and pattern attack events.
[0020] By extracting the associated data sequences of the target sensor and its associated sensors within a window, the window coordination deviation index between the target sensor and each associated sensor is calculated. A pre-defined false event result verification algorithm is then used to output a decision index. This multi-source data verification mechanism comprehensively considers data from multiple sensors, enhancing the reliability of verification and reducing misjudgments caused by abnormal data from a single sensor. The decision index comprehensively considers the coordination deviation between the target sensor and its associated sensors, providing a more reliable basis for subsequent detection and processing of target data from sampling points within the current detection cycle of the target sensor. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the IoT data identification and detection method according to an embodiment of the present invention. Detailed Implementation
[0022] To facilitate understanding of the present invention, a more complete description of this application will be given below with reference to the accompanying drawings, which illustrate preferred embodiments of the invention. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to enable a more thorough and complete understanding of the disclosure of the present invention.
[0023] It should be noted that the terms "vertical," "horizontal," "up," "down," "left," "right," and similar expressions used in this article are for illustrative purposes only and do not represent the only possible implementation.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention; the term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0025] Example 1: Figure 1 This is a flowchart illustrating the IoT data identification and detection method according to an embodiment of the present invention.
[0026] like Figure 1 As shown, a method for identifying and detecting IoT data includes the following steps: S101, obtain the target data time-series curve generated by fitting the target data sequence collected by the target sensor in the current detection cycle, and match the target data comparison time-series curve according to the working condition data characteristics of the target sensor.
[0027] It should be noted that since IoT sensor data are all time-series data, commonly used sensor data in industrial production include humidity, temperature, pressure, and voltage. Target data is set as the data type collected by the target sensor, which can be any of humidity, temperature, pressure, and voltage. This invention does not limit this; target data is used to refer to a certain data type.
[0028] In some embodiments, the operating condition data characteristics of the target sensor include state parameter information of the operating equipment (directly from the IoT device connected to the sensor, quantifying the health status of the equipment itself, such as power supply stability and computing resources) and target data time-series curve characteristic information. For example, the state parameter information of the operating equipment includes, but is not limited to, battery level and CPU load at the beginning of the detection cycle, and the target data time-series curve characteristic information includes, but is not limited to, range, mean, standard deviation, and frequency domain characteristics. Specifically, the operating condition data characteristics are represented as follows: , The target sensor's current detection cycle operating data characteristics are represented by B, where B is the battery level, L is the CPU load, and R is the range of the target data time-series curve (the difference between the maximum and minimum values of the curve). The mean of the time series curve of the target data. The standard deviation of the target data time series curve. It is a frequency domain feature.
[0029] The frequency domain features are defined as follows:
[0030] Frequency domain features are used to capture the periodic patterns of the target data time series curve within the detection period (e.g., data oscillations caused by vibration, avoiding mean smoothing of true fluctuations).
[0031] In some embodiments, matching the comparison time-series curve of the target data based on the operating condition data characteristics of the target sensor specifically includes: S201, based on the operating mode of the device to which the current target sensor belongs (the operating modes of different IoT devices in an industrial environment, not limited to different operating modes such as start-up and shutdown), obtain the target data time-series curves for all detection cycles in the history of the target sensor under the same operating mode, and filter out the curves that are genuine and free from false attack injections, as the historical time-series curve set. , For the j-th historical time series curve, record the historical operating condition data features corresponding to each historical time series curve to obtain the historical operating condition data feature set. , The historical operating condition data features are those of the j-th historical time series curve.
[0032] S202, based on the historical operating condition data feature set, uses a preset clustering algorithm (e.g., DBSCAN density clustering or K-means clustering algorithm, which can automatically identify multimodal distributions and handle noise and non-spherical clusters; the working principle of related clustering algorithms will not be elaborated in this invention, but can be referred to relevant existing content) to group the historical time series curve set, obtaining N groups. , Create the Nth class group and set a working condition label for each class group.
[0033] It should be noted that the clustering process uses the Euclidean distance between the features of historical operating conditions corresponding to the historical time series curves as a similarity measure, minimizing the differences in operating condition features within each cluster.
[0034] The method for obtaining the working condition labels for each category group is as follows: Calculate the mean of the historical operating condition data characteristics of all historical time series curves in this category group, and use it as the operating condition label for this category group. , Let be the working condition label for the i-th class group.
[0035] S203, obtain the operating condition data characteristics of the target sensor in the current detection cycle, calculate the similarity between it and the operating condition label of each class group, and take the class group with the highest similarity as the target class group.
[0036] Specifically, the similarity between the condition label and the work condition label of each class group is calculated according to the following algorithm:
[0037] in, To measure the similarity with the working condition label of the i-th class group, The target sensor's operating condition data characteristics for the current detection cycle. For the working condition label of the i-th class group, Represents Euclidean distance. The higher the value, the higher the similarity.
[0038] S204. Based on all historical time-series curves in the target class group, obtain the high-fidelity standard curve under the label of this working condition, and match it as the comparison time-series curve of the target data.
[0039] In some embodiments, the high-fidelity standard curve under the working condition label can be determined by: calculating the similarity between all historical time-series curves in the working condition label and the target data time-series curve in the current detection cycle, and taking the historical time-series curve with the highest similarity as the high standard curve.
[0040] In other embodiments, the high-fidelity standard curve under the operating condition label can also be determined as follows:
[0041] For the high-fidelity standard curve under the working condition label of the i-th class group, Represents the i-th class group, This represents the m-th historical time series curve in the i-th class group. The weight value of the m-th historical time series curve is based on the current target data time series curve. With the m-th historical time series curve The dynamic feature distance (Euclidean distance) between curves is calculated, and curves with high similarity contribute more, reducing the impact of differences in the shape of target data changes.
[0042] Therefore, the operating condition data features aim to capture the "static" attributes (equipment status) and "dynamic" attributes (curve statistics) of operating conditions. Clustering groups similar operating conditions to ensure that high-fidelity standard curves are generated based on homogeneous historical data, avoiding baseline drift caused by equipment failure or environmental interference. The comparison time-series curves automatically adapt to different equipment states (such as battery depletion or load fluctuations), preventing baseline drift due to equipment failure. The comparison time-series curves are generated based on historical real curves, ensuring they represent an ideal state free from spoofing attacks, improving the detection accuracy of continuous or periodic forgery patterns. Similarity calculation uses a normalized distance metric, reducing matching complexity and making it suitable for real-time IoT edge devices. Integrating equipment status and curve statistical features (range / mean / standard deviation) constructs multi-dimensional operating condition labels, solving the problem of static baselines failing to respond to changes in operating conditions. By quantifying the similarity between the current operating condition and various operating condition labels, on-demand generation of comparison time-series curves is achieved, avoiding the computational waste of blindly comparing with a large number of historical time-series curves and improving the matching rate.
[0043] By extracting features from operating data, the operating status of equipment can be described more accurately, thereby improving the accuracy of the monitoring system. Since the features of operating data can reflect the true operating status of the equipment, false alarms caused by noise or interference can be reduced.
[0044] S102: Extract the multi-scale false event feature vector of each sampling point in the time series curve of the target data, input it into the pre-trained data false identification model, and output the false event result associated with the target data of the corresponding sampling point.
[0045] Specifically, the feature vector of multi-scale fake events includes single-point deviation features, window pattern features, and frequency domain anomaly features, represented as follows: , This is the feature vector of the multi-scale spoof event at the k-th sampling point. For the single-point deviation feature of the k-th sampling point, For the window pattern features of the k-th sampling point, This represents the frequency domain anomaly feature of the k-th sampling point.
[0046] Among them, the single-point deviation feature is determined based on the target data time series curve and the target data at the k-th sampling point in the comparison time series curve, specifically as follows:
[0047] This is a single-point deviation feature. The target data is the target data at the k-th sampling point of the target data time series curve. To compare the target data at the k-th sampling point of the time series curve, To obtain the standard deviation of the target data corresponding to real and authentic historical target data, the standard deviation can be calculated from the target data at the kth sampling point in the historical time series curve of the corresponding target group. This standard deviation serves as the benchmark for target data fluctuation under the same working conditions. The instantaneous offset is normalized to eliminate the influence of dimensions and obtain the relative deviation.
[0048] The window pattern feature is determined based on the target data time series curve and the target data within a preset window length in the comparison time series curve, specifically:
[0049] For window mode features, w is the preset window length, determined based on the average duration of typical false data attacks occurring within the detection cycle in industrial scenarios (e.g., ramp attacks require a duration greater than or equal to 10 sampling points) (this can be set using historical data and expert experience), and is less than the detection cycle length. For the target data time series curve window The target data of the z-th sampling point. To compare timing curves in the window The mean of the target data within the window is used as the ideal benchmark. By comparing the actual local mean with the ideal benchmark, the timeliness of local pattern analysis is ensured. By introducing window pattern features, it is possible to effectively avoid the omission of false events caused by single-point misjudgment. Continuous attacks that are missed by single-point features are captured by window pattern features.
[0050] It should be noted that the window is constructed using the current k-th sampling point as the base point. This creates a scale progression from "single point to local," addressing continuous forgery patterns (such as ramp attacks and staircase attacks) that single-point features cannot identify. For example, maliciously injecting slowly varying temperature data from 10 consecutive sampling points (simulating equipment overload), the deviation at a single point may not exceed a threshold, but the window pattern feature can be triggered due to a significant increase in accumulated deviation within the window. Additionally, it's important to note that if... If the sampling point exceeds the detection cycle range, the sampling point at the end of the detection cycle will be replaced. Therefore, this invention will not elaborate on this special case.
[0051] Among them, the frequency domain anomaly features are determined based on the frequency domain features of the target data within a preset window length in the time-series curve of the target data and the comparison time-series curve, specifically as follows:
[0052] For the frequency domain anomaly features of the k-th sampling point, The L2 norm, which will not be elaborated upon in this invention, is used to calculate the magnitude of the frequency domain vector after the FFT transform, characterizing the total energy in the frequency domain. The numerator represents the time series curve of the target data within the window. The frequency domain energy, with the denominator representing the window corresponding to the alignment time-series curve. Frequency domain energy reference, This is used to represent the frequency domain energy ratio between the target data time-series curve and the comparison time-series curve. For example, if... A value greater than 1 indicates that the target data time-series curve has higher frequency domain energy, implying the existence of abnormal high-frequency oscillations. A value less than 1 indicates that the frequency domain energy of the target data time series curve is lower, implying abnormal smoothness.
[0053] In some embodiments, the pre-trained fake data detection model is obtained in the following specific ways: S301, obtain the time-series curves of historical target data within a large number of historical detection cycles of the target sensor and the identification type of target data at all historical sampling points (real data, fake data, which can be quantified as 0 and 1, where 0 represents real data and 1 represents fake data). If the data is fake data, it also includes the type of fake event corresponding to the fake data (e.g., single-point fake event, continuous fake event, pattern attack event).
[0054] S302, generate multi-scale false event feature vectors for all historical sampling points and label them. The label content is set to false event results, including real data and false data. When it is false data, it also includes the false event type.
[0055] The labeled false event results can also be set as: quadruple labels. When the data is real, When it is a single false event, When it is a series of false events, When it is a pattern attack event, Where 1 indicates that it is 100% the corresponding identification type.
[0056] It should be noted that if more training data is needed, a large amount of simulated data can be generated by simulating fake data, and the simulated data can be labeled. This invention will not elaborate on this.
[0057] S303 uses the labeled feature vectors of all multi-scale fake events as the training set to train a pre-selected neural network structure, obtaining the final fake event identification model, which is used to output the results of fake event identification. , The probability of real data. The probability of a single false event. The probability of a series of false events. This represents the probability of a pattern attack event.
[0058] Therefore, the multi-scale fake event feature vector is a feature vector designed specifically for fake events. By analyzing the changing patterns of data at different scales, it can capture the characteristics of fake events and achieve effective detection of fake events.
[0059] S103 utilizes a preset window multi-source data verification mechanism to obtain decision indicators for false event results at each sampling point in the target data time series curve.
[0060] In some embodiments, the preset window multi-source data verification mechanism is specifically set according to the associated sensors in the current working environment of the target sensor, and specifically includes: S401, based on the window corresponding to the kth sampling point in the current detection period. Extract the associated data sequence of the target sensor (at least one) within the window.
[0061] It should be noted that the determination of the associated sensor is based on the data type that has a physical correlation or fluctuation coordination with the target data of the target sensor. This needs to be determined based on the working environment of different IoT devices in the industrial scenario, and is usually based on expert experience. For example, temperature sensors and voltage sensors are associated, but this invention will not elaborate on this.
[0062] S402, based on the target data sequence X of the target sensor within the window and the associated data sequence Y of the associated sensor within the window, obtain the window cooperative deviation index between the target sensor and each associated sensor.
[0063] Specifically, the window co-location deviation index is determined based on the DTW distance and standard DTW distance between the target data sequence and the associated data sequence within the window, and is set as follows:
[0064] in, Let X be the window cooperative deviation index corresponding to the k-th sampling point of the target sensor within the current detection period, and let X be the target data sequence of the target sensor within the window. Let q be the associated data sequence of the target sensor's associated sensor within the window. The window collaboration features between the target sensor and the qth associated sensor under the current operating condition data characteristics are determined as follows: Within the target group, all historical time-series curves with a similarity greater than 0.8 to the comparison time-series curves are selected, and the associated time-series curves of the corresponding associated sensors (associated time-series curves collected within the same detection period and without false data) are obtained. Calculate the DTW distance between each selected historical time series curve and its corresponding associated time series curve within the window of the kth sampling point. Take the average DTW distance of all selected historical time series curves as the window collaborative feature.
[0065] S403, based on the window collaborative deviation index between the target sensor and each associated sensor, uses a preset false event result verification algorithm to output the corresponding decision index.
[0066] Specifically, the preset algorithm for verifying false event results is set as follows:
[0067] in, This is the decision metric for the false event result at the k-th sampling point of the target sensor within the current detection period. To provide a comprehensive coordination deviation index, Q represents the total number of sensors currently associated with the target sensor. Let the window coordination deviation index be the k-th sampling point of the target sensor within the current detection period, and take the maximum value as the comprehensive coordination deviation index. A preset sensitivity coefficient (e.g., set to 0.5, adjusted based on expert experience and actual conditions) controls the attenuation of the deviation on the decision indicator. The higher the value, the stronger the consensus and the more reliable the results of false events.
[0068] S104, based on the decision index of false event results of sampling points in the target data time series curve, detects and processes the target data of sampling points in the current detection cycle of the target sensor.
[0069] In some embodiments, step S104 specifically includes: S501 combines the results of false events with decision indicators to obtain the probability of the target false event.
[0070] S501, based on the probability of target false events, executes a graded response strategy.
[0071] For example, when a fake event result is represented as a four-tuple label At that time, the probability of a false event is set to the target quadruple label: The maximum value in the target quadruple label is taken as the false probability of the target data at the sampling point (the false probability threshold is preset based on expert experience and historical data, for example, set to 0.6): When the maximum value corresponds to the real data, if the false probability is greater than 0.6, no action is taken; otherwise, a single-point replacement is performed (the target data of the corresponding sampling point in the comparison time-series curve is replaced with the sampling point of the current detection period). When the maximum value corresponds to a single false event, if the false probability is greater than 0.6, then the single point is replaced; otherwise, the system runs normally and the sampling point is marked for subsequent review. When the maximum value corresponds to consecutive false events, if the false probability is greater than 0.6, then window replacement is performed (the target data in the window of the corresponding sampling point in the comparison time series curve is replaced in the window of the sampling point in the current detection period); otherwise, the system runs normally and the sampling point window is marked for subsequent review. When the maximum value corresponds to a pattern attack event, if the false probability is greater than 0.6, the system will immediately freeze and issue an alarm; otherwise, the data from the target sensor will be isolated and the administrator will be immediately notified for review.
[0072] It is important to note that if decision indicators are detected... If the value is less than the preset coordination threshold (e.g., 0.4), it indicates a serious deviation, and further review of the single-point data or window data of the associated sensor corresponding to the target sensor is required.
[0073] It should be noted that different false probability thresholds can be set for different data types (real data, single point, continuous, pattern attacks), and can be adjusted according to actual scenarios and needs. The above is only an example, and this invention will not elaborate further.
[0074] In summary, a target data time-series curve is generated by fitting the data sequence collected by the target sensor within the current detection cycle. This curve reflects the change of target data (such as temperature and humidity) over time within the current detection cycle. Based on the operating condition data characteristics of the target sensor (including equipment status parameters and curve statistical characteristics), similar historical operating condition data curves are matched as comparison time-series curves. This ensures that the environmental conditions and equipment status of the comparison curve are similar to those of the current detection curve, thereby improving the accuracy of subsequent false data detection.
[0075] Multi-scale fake event feature vectors are extracted, including single-point deviation features, window pattern features, and frequency domain anomaly features, reflecting the potential for data falsification from different perspectives. The extracted feature vectors are input into a pre-trained data falsification detection model, which outputs the fake event results associated with the target data at corresponding sampling points. The model automatically identifies potential fake data points. Specifically, single-point deviation features reflect anomalies in individual data points, window pattern features reflect pattern anomalies within a data segment, and frequency domain anomaly features reflect anomalies in the frequency domain. The data falsification detection model is based on a neural network model trained on a large amount of historical data and can automatically identify and classify fake data.
[0076] A multi-source data verification mechanism using windows is employed to obtain decision indicators: Based on the sampling point window within the current detection period, associated sensor data of the target sensor is extracted, leveraging the correlation between sensors in the Internet of Things (IoT) to improve the reliability of false data detection. By calculating metrics such as the DTW distance between the data sequences of the target sensor and associated sensors within the window, a window collaborative deviation index is obtained, reflecting the consistency between the target sensor data and the associated sensor data. Based on the window collaborative deviation index, a pre-defined false event result verification algorithm outputs decision indicators to evaluate the reliability of false event results. Specifically, the false event result verification algorithm outputs decision indicators based on the window collaborative deviation index and the sensitivity coefficient to assess the reliability of false event results.
[0077] Based on decision indicators, the target data at the sampling points is detected and processed. By combining the results of false events with the decision indicators, the probability of false events is obtained, which is used to assess the likelihood that the target data at the sampling points is false data. A hierarchical response strategy is executed, and different response strategies are implemented according to the probability of false events, such as single-point replacement, window replacement, freezing and alarm, to ensure that the system can take appropriate measures according to different false data situations.
[0078] The technical solutions described in the embodiments of this application have at least the following technical effects or advantages: By capturing the "static" attributes (such as battery level and CPU load) and "dynamic" attributes (such as curve statistical characteristics like range, mean, and standard deviation), similar operating conditions are grouped and labeled. This allows the comparison time-series curves to automatically adapt to different equipment states, such as battery depletion or load fluctuations, avoiding baseline drift caused by equipment failure or environmental interference and ensuring the accuracy of the comparison benchmark. The comparison time-series curves are generated based on historical real curves, representing the ideal state free from spoofing attacks, thereby improving the detection accuracy of continuous or periodic forgery patterns. For example, it can more accurately identify attacks that mimic continuous, slowly changing temperature data simulating equipment overload. Simultaneously, by quantifying the similarity between the current operating condition and various operating condition labels, the comparison time-series curves can be generated on demand, avoiding blind comparisons with a large number of historical time-series curves, saving computational resources, and improving the matching rate. Since the operating condition data characteristics reflect the true operating state of the equipment, false alarms caused by noise or interference can be reduced, improving the accuracy of the monitoring system.
[0079] Multi-scale spoofing event feature vectors include single-point deviation features, window pattern features, and frequency domain anomaly features. These analyze data variation patterns at different scales. Single-point deviation features focus on the deviation of a single sampling point, window pattern features analyze local patterns using a preset window length, and frequency domain anomaly features capture abnormal energy changes from a frequency domain perspective. This comprehensive approach captures the characteristics of spoofing events and effectively detects various types of spoofing attacks, such as single-point spoofing events, continuous spoofing events, and pattern attacks. The introduction of window pattern features effectively avoids missing spoofing events due to single-point misjudgments. For example, in attacks that maliciously inject slowly varying data across multiple consecutive sampling points, a single-point deviation might not exceed a threshold, but the window pattern feature will be triggered due to a significant increase in accumulated deviation within the window, thus improving detection reliability.
[0080] By extracting the associated data sequences of the target sensor and its associated sensors within a window, the window coordination deviation index between the target sensor and each associated sensor is calculated. A pre-defined false event result verification algorithm is then used to output a decision index. This multi-source data verification mechanism comprehensively considers data from multiple sensors, enhancing the reliability of verification and reducing misjudgments caused by abnormal data from a single sensor. The decision index comprehensively considers the coordination deviation between the target sensor and its associated sensors, providing a more reliable basis for subsequent detection and processing of target data from sampling points within the current detection cycle of the target sensor.
[0081] Based on the results of false events (real data, single false events, continuous false events, and pattern attack events) and the probability of the target false event, a tiered response strategy is implemented. Different handling methods are adopted for different types and probabilities of false events, such as single-point replacement, window replacement, freezing and alarm, data isolation and reminder to administrator for review, etc. Differentiated handling can more reasonably deal with various false events and improve the security and stability of the system.
[0082] In summary, by obtaining the target data time-series curve and comparing it with the time-series curve, extracting multi-scale spurious event feature vectors and inputting them into the model, utilizing a windowed multi-source data verification mechanism to obtain decision indicators, and detecting and processing the target data at sampling points based on these decision indicators, spurious data detection of IoT sensor data was achieved. By combining multi-source information such as device status parameters, curve statistical characteristics, and correlated sensor data, the accuracy and reliability of spurious data detection were improved. Furthermore, the hierarchical response strategy ensured that the system could take appropriate countermeasures according to different spurious data situations, effectively preventing the impact of spurious data on the IoT system.
[0083] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for identifying and detecting Internet of Things (IoT) data, characterized in that, include: S101, obtain the target data time-series curve generated by fitting the target data sequence collected by the target sensor in the current detection cycle, and match the target data comparison time-series curve according to the working condition data characteristics of the target sensor. S102: Extract the multi-scale false event feature vector of each sampling point in the time series curve of the target data, input it into the pre-trained data false identification model, and output the false event result associated with the target data of the corresponding sampling point; S103, using a preset window multi-source data verification mechanism, obtains decision indicators for false event results at each sampling point in the target data time series curve; S104, based on the decision index of false event results of sampling points in the target data time series curve, detects the target data of sampling points in the current detection cycle of the target sensor.
2. The method for identifying and detecting IoT data as described in claim 1, characterized in that, The operating condition data features are represented as follows: , Here, B represents the operating condition data characteristics of the target sensor during the current detection cycle, L represents the CPU load, and R represents the range of the target data time-series curve. The mean of the time series curve of the target data. The standard deviation of the target data time series curve. It is a frequency domain feature.
3. The method for identifying and detecting IoT data as described in claim 2, characterized in that, The step of matching the target data with the comparison time-series curve based on the operating condition data characteristics of the target sensor specifically includes: S201, based on the operating mode of the device to which the current target sensor belongs, obtain the target data time-series curves for all detection cycles in the history of the target sensor under the same operating mode, and filter out the curves that are genuine and free from false attack injections, as a set of historical time-series curves. , For the j-th historical time series curve, record the historical operating condition data features corresponding to each historical time series curve to obtain the historical operating condition data feature set. , The historical operating condition data characteristics of the j-th historical time series curve; S202, based on the feature set of historical operating condition data, uses a preset clustering algorithm to group the historical time series curve set into N groups. And set working condition labels for each class group; S203, obtain the operating condition data characteristics of the target sensor in the current detection cycle, calculate its similarity with the operating condition label of each class group, and take the class group with the highest similarity as the target class group; S204. Based on all historical time-series curves in the target class group, obtain the high-fidelity standard curve under the label of this working condition, and match it as the comparison time-series curve of the target data.
4. The method for identifying and detecting IoT data as described in claim 3, characterized in that, The multi-scale fake event feature vector includes single-point deviation features, window pattern features, and frequency domain anomaly features, represented as follows: , This is the feature vector of the multi-scale spoof event at the k-th sampling point. For the single-point deviation feature of the k-th sampling point, For the window pattern features of the k-th sampling point, This represents the frequency domain anomaly feature of the k-th sampling point.
5. The method for identifying and detecting IoT data as described in claim 4, characterized in that, The single-point deviation feature is determined based on the target data time series curve and the target data at the kth sampling point in the comparison time series curve; The window mode feature is determined based on the target data time series curve and the target data within a preset window length in the comparison time series curve; The frequency domain anomaly features are determined based on the frequency domain features of the target data within a preset window length in the target data time series curve and the comparison time series curve.
6. The method for identifying and detecting IoT data as described in claim 5, characterized in that, The specific features of the window mode are: This refers to the window mode feature, where w is the preset window length. For the target data time series curve window The target data of the z-th sampling point. To compare timing curves in the window The mean of the target data.
7. The method for identifying and detecting IoT data as described in claim 4, characterized in that, The pre-trained data fraud detection model is obtained in the following specific way: S301, obtain the time-series curves of historical target data within a large number of historical detection cycles of the target sensor and the identification type of target data at all historical sampling points; S302, generate multi-scale spoof event feature vectors for all historical sampling points, and label them with the label content set as spoof event results; S303 uses all labeled multi-scale fake event feature vectors as the training set to train a pre-selected neural network structure, thus obtaining the final data fake detection model.
8. The method for identifying and detecting IoT data as described in claim 7, characterized in that, In step S302, the labeled false event result is set as: quadruple label. When the data is real, When it is a single false event, When it is a series of false events, When it is a pattern attack event, .
9. The method for identifying and detecting IoT data as described in claim 6, characterized in that, The preset window multi-source data verification mechanism specifically includes: S401, based on the window corresponding to the kth sampling point in the current detection period. Extract the associated data sequence of the target sensor within the window; S402, based on the target data sequence of the target sensor within the window and the associated data sequence of the associated sensor within the window, obtain the window cooperative deviation index between the target sensor and each associated sensor; S403, based on the window collaborative deviation index between the target sensor and each associated sensor, uses a preset false event result verification algorithm to output the corresponding decision index.
10. The method for identifying and detecting IoT data as described in claim 9, characterized in that, The preset algorithm for verifying false event results is set as follows: in, This is the decision metric for the false event result at the k-th sampling point of the target sensor within the current detection period. To provide a comprehensive coordination deviation index, Q represents the total number of sensors currently associated with the target sensor. This refers to the window collaborative deviation index corresponding to the k-th sampling point of the target sensor within the current detection period. The preset sensitivity coefficient controls the attenuation of the deviation on the decision indicator.
Citation Information
Patent Citations
Multi-dimensional periodicity detection of IoT device behavior
CN113424157A
Sensor false data detection method and device, computer equipment and storage medium
CN115081524A
Smart home management method and system based on Internet of Things
CN120762296A
System and method for sensor outlier detection using ensemble algorithm considering characteristic environmental variables
KR102752526B1