An identification detection method for internet of things data

By generating comparison time-series curves in the Internet of Things system and combining multi-scale fake event feature vectors and multi-source data verification, the accuracy and adaptability issues of fake data detection are solved, enabling accurate identification and differentiated processing of fake events, and improving the security and stability of the system.

CN120951003BActive Publication Date: 2025-12-09SHANGHAI RENWEI ELECTRONIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511483963.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-12-09
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

The detection of fake data in IoT systems is difficult to identify effectively, especially in complex and ever-changing environments where the system has poor adaptability, affecting the normal operation and security of the system.

Method used

By generating comparison time-series curves through feature matching of operating data, and combining multi-scale fake event feature vector detection and window multi-source data verification mechanism, we can achieve accurate identification and differentiated processing of fake events in IoT data.

Benefits of technology

It improves the accuracy and reliability of fake data detection, reduces false alarms caused by equipment failure or environmental interference, and can effectively detect various types of fake data attacks, ensuring the security and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951003B_ABST
    Figure CN120951003B_ABST
Patent Text Reader

Abstract

The application discloses a kind of identification detection methods of Internet of Things data, the method includes: obtaining target data time series curve that target data sequence fitted generation that target sensor is collected in current detection period, according to the comparison time series curve of matching target data feature of working condition data;Extract the multi-scale false event feature vector of each sampling point, input into data false identification model, and output the false event result associated with the corresponding sampling point target data;Using the preset window multi-source data verification mechanism, obtain the decision index of the false event result of each sampling point in the target data time series curve;Based on the decision index of the false event result of sampling point in the target data time series curve, the sampling point target data in the current detection period of target sensor is detected.Therefore, the accuracy of Internet of Things data identification detection is significantly improved, and the occurrence of missed report and false report is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of data recognition, and in particular to a recognition and detection method for Internet of Things data. BACKGROUND

[0002] False data in the Internet of Things refers to misleading or deceiving system users and other related parties by manipulating or falsifying sensor data. Internet of Things sensor data is vulnerable to various passive manipulations during transmission and collection, resulting in distorted data or false data. False data not only affects the normal operation of the system and the accuracy of decision-making, but also may cause safety accidents and economic losses.

[0003] False data can have a disastrous impact on Internet of Things systems. It can cause control systems to respond incorrectly to certain situations, leading to safety accidents and economic losses. Due to the nature of the Internet of Things environment, which is interconnected and ubiquitous, false data in the Internet of Things is hidden. In addition to improving the security of Internet of Things data during transmission and collection, false data feature recognition models and threshold values can be constructed based on experience values to identify false data, but their adaptive ability is poor and it is difficult to cope with complex and changing Internet of Things environments.

[0004] Therefore, how to effectively detect false data in Internet of Things sensor data has become a problem to be solved. SUMMARY

[0005] The application provides a recognition and detection method for Internet of Things data, which generates a comparison time series curve by matching the working condition data features, combines a multi-scale false event feature vector detection and a window multi-source data verification mechanism, and realizes accurate recognition and differential processing of false events in Internet of Things data.

[0006] The application provides a recognition and detection method for Internet of Things data, which includes:

[0007] S101, obtaining a target data time series curve fitted and generated by a target data sequence collected by a target sensor in a current detection period, and matching a comparison time series curve of the target data according to the working condition data features of the target sensor;

[0008] S102, extracting a multi-scale false event feature vector of each sampling point in the target data time series curve, inputting the multi-scale false event feature vector into a pre-trained data false recognition model, and outputting a false event result associated with the target data of the corresponding sampling point;

[0009] S103, obtaining a decision index of the false event result of each sampling point in the target data time series curve by using a preset window multi-source data verification mechanism;

[0010] S104, detecting the sampling point target data in the current detection period of the target sensor based on the decision index of the false event result of the sampling point in the target data time series curve.

[0011] Preferably, the working condition data feature is represented as , is the working condition data feature of the target sensor in the current detection period, B is the battery level, L is the CPU load, R is the range value of the target data time series curve, is the mean value of the target data time series curve, is the standard deviation value of the target data time series curve, is the frequency domain feature.

[0012] Preferably, the comparison time series curve of the target data matched according to the working condition data feature of the target sensor specifically comprises:

[0013] S201, according to the working mode of the device to which the current target sensor belongs, obtaining the target data time series curve in all detection periods of the target sensor in the same working mode in history, filtering out the curve without false attack injection to obtain the historical time series curve set , is the jth historical time series curve, recording the historical working condition data feature corresponding to each historical time series curve to obtain the historical working condition data feature set , is the historical working condition data feature of the jth historical time series curve;

[0014] S202, based on the historical working condition data feature set, using a preset clustering algorithm to group the historical time series curve set to obtain N class groups , and setting a working condition label for each class group;

[0015] S203, obtaining the working condition data feature of the target sensor in the current detection period, calculating the similarity between the working condition data feature and the working condition label of each class group, and taking the class group with the largest similarity as the target class group;

[0016] S204, based on all historical time series curves in the target class group, obtaining the high-fidelity standard curve under the working condition label, and matching as the comparison time series curve of the target data.

[0017] Preferably, the multi-scale false event feature vector includes single-point deviation feature, window mode feature and frequency domain abnormal feature, and is represented as , is the multi-scale false event feature vector of the kth sampling point, is the single-point deviation feature of the kth sampling point, is the window mode feature of the kth sampling point, The frequency domain anomaly feature of the kth sampling point.

[0018] Preferably, the single-point deviation feature is determined according to target data of the kth sampling point in the target data time series curve and the comparison time series curve.

[0019] The window mode feature is determined according to target data in a preset window length in the target data time series curve and the comparison time series curve.

[0020] The frequency domain anomaly feature is determined according to frequency domain features of target data in a preset window length in the target data time series curve and the comparison time series curve.

[0021] Preferably, the window mode feature is specifically:

[0022]

[0023] For the window mode feature, w is a preset window length, The target data of the zth sampling point in the window The mean value of the target data in the window The comparison time series curve.

[0024] Preferably, the pre-trained data false identification model is obtained in the following manner:

[0025] S301, obtaining historical target data time series curves and identification types of target data of all historical sampling points in a large number of historical detection periods of a target sensor;

[0026] S302, generating multi-scale false event feature vectors of all historical sampling points and labeling them, and the labeling content is set as a false event result;

[0027] S303, taking all labeled multi-scale false event feature vectors as a training set, training a preselected neural network structure, and obtaining a final data false identification model.

[0028] Preferably, in S302, the false event result is set as a four-tuple label When it is real data, When it is a single-point false event, When it is a continuous false event, When it is a mode attack event, .

[0029] Preferably, the preset window multi-source data verification mechanism specifically includes:

[0030] ​S401, obtaining a window corresponding to the kth sampling point in the current detection period based on the target sensor extracting a correlation data sequence corresponding to the associated sensor of the target sensor in the window;

[0031] S402, obtaining a window collaborative deviation index of the target sensor and each associated sensor based on the target data sequence of the target sensor in the window and the correlation data sequence of the associated sensor in the window;

[0032] S403, based on the window collaborative deviation index of the target sensor and each associated sensor, using a preset false event result verification algorithm to output a corresponding decision index.

[0033] Preferably, the preset false event result verification algorithm is set as:

[0034]

[0035] wherein, is the decision index of the false event result of the target sensor at the kth sampling point in the current detection period, is a comprehensive collaborative deviation index, Q is the total number of current associated sensors of the target sensor, is the window collaborative deviation index of the target sensor corresponding to the kth sampling point in the current detection period, is a preset sensitivity coefficient, controlling the attenuation intensity of the deviation degree on the decision index.

[0036] One or more technical solutions provided in the present application have at least the following technical effects or advantages:

[0037] By capturing the "static" and "dynamic" properties of the device state, grouping similar working conditions and setting working condition labels, the comparison time curve can automatically adapt to different device states, such as battery depletion or load fluctuation, avoiding baseline drift caused by device failure or environmental interference, and ensuring the accuracy of the comparison reference; based on the historical real curve, the comparison time curve is generated, which can represent the ideal state without false attack, thereby improving the detection accuracy of subsequent continuous or periodic fake patterns. At the same time, by quantifying the similarity of the current working condition and each working condition label, the comparison time curve is generated on demand, avoiding blind comparison with a large number of historical time curves, saving computing resources, and improving the matching rate. Since the working condition data features can reflect the real running state of the device, false positives caused by noise or interference can be reduced, and the accuracy of the monitoring system can be improved.

[0038] The multi-scale false event feature vector includes single-point deviation features, window mode features and frequency domain anomaly features, which can comprehensively capture the features of false events and effectively detect various types of false data attacks, such as single-point false events, continuous false events and mode attack events.

[0039] By extracting the associated data sequence of the associated sensor of the target sensor in the window, the window cooperative deviation index of the target sensor and each associated sensor is calculated, and the decision index is output by using the preset false event result verification algorithm, the multi-source data verification mechanism can comprehensively consider the data of multiple sensors, enhance the reliability of verification, and reduce the false judgment caused by the abnormal data of a single sensor; the decision index comprehensively considers the cooperative deviation of the target sensor and the associated sensor, and provides a more reliable decision basis for subsequent detection and processing of the sampling point target data of the target sensor in the current detection period. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 A flowchart of the identification and detection method of the Internet of Things data of the embodiment of the present application. DETAILED DESCRIPTION

[0041] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the related drawings; the preferred embodiments of the present application are shown in the drawings, but the present application can be realized in many different forms, and is not limited to the embodiments described herein; on the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0042] It should be noted that the terms "vertical", "horizontal", "up", "down", "left", "right" and similar expressions used herein are only for the purpose of illustration, and do not represent the only embodiment.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs; the terms used herein in the specification of the present application are only for the purpose of describing the specific embodiments, and are not intended to limit the present application; the term "and / or" used herein includes any and all combinations of one or more related listed items.

[0044] Embodiment one: Figure 1 A flowchart of the identification and detection method of the Internet of Things data of the embodiment of the present application.

[0045] As shown in Figure 1 A method for identifying and detecting Internet of Things data, comprising the following steps:

[0046] S101, obtain a target data time series curve fitted and generated by a target data sequence collected by a target sensor in a current detection period, and match a comparison time series curve of the target data according to a working condition data feature of the target sensor.

[0047] It should be noted that, since the Internet of Things sensor data are all time series data, the commonly used sensor data in industrial production are humidity, temperature, pressure, voltage, etc., the target data is set as the data type collected by the target sensor, which can be any one of humidity, temperature, pressure, and voltage, and the present application does not limit this, and the target data is used to refer to a certain data type.

[0048] In some embodiments, the working condition data feature of the target sensor includes state parameter information of the running equipment to which it belongs (directly from the Internet of Things equipment connected to the sensor, quantifying the health status of the equipment itself, such as power supply stability, computing power resources) and target data time series curve feature information, for example, the state parameter information of the running equipment includes but is not limited to the battery level at the starting point of the detection period and the CPU load, and the target data time series curve feature information includes but is not limited to the range value, the mean value, the standard deviation value, and the frequency domain feature. Specifically, the working condition data feature is represented as , is the working condition data feature of the target sensor in the current detection period, B is the battery level, L is the CPU load, R is the range value of the target data time series curve (the difference between the maximum value and the minimum value of the curve), is the mean value of the target data time series curve, is the standard deviation value of the target data time series curve, is the frequency domain feature.

[0049] Among them, the frequency domain feature is defined as:

[0050]

[0051] The frequency domain feature is used to capture the periodic pattern of the target data time series curve in the detection period (such as data oscillation caused by vibration, to avoid smoothing of real fluctuations by the mean value).

[0052] In some embodiments, the comparison time series curve of the target data is matched according to the working condition data feature of the target sensor, specifically including:

[0053] S201, according to the working mode of the equipment to which the current target sensor belongs (the working mode of different Internet of Things equipment in an industrial environment, not limited to different working modes such as startup and shutdown), obtain the target data time series curve in all detection periods of the target sensor in the history under the same working mode, and filter out the curves without false attack injection as a historical time series curve set , record the historical working condition data features corresponding to each historical time series curve to obtain a historical working condition data feature set , is the historical working condition data feature of the jth historical time series curve.

[0054] S202, based on the historical working condition data feature set, a preset clustering algorithm (for example, DBSCAN density clustering or K-means clustering algorithm, which can automatically identify multimodal distribution, process noise and non-spherical clusters, and the working principle of related clustering algorithms will not be repeated here. It can refer to related existing content) is used to group the historical time series curve set to obtain N class groups , is the Nth class group, and a working condition label is set for each class group.

[0055] It should be noted that the clustering process takes the Euclidean distance between the historical working condition data features corresponding to the historical time series curves as the similarity measure, and minimizes the intra-class working condition feature difference.

[0056] wherein the working condition label of each class group is obtained in the following manner:

[0057] The mean value of the historical working condition data features of all historical time series curves in the class group is calculated as the working condition label of the class group , is the working condition label of the ith class group.

[0058] S203, obtain the working condition data feature of the target sensor in the current detection period, calculate the similarity between it and the working condition label of each class group, and take the class group with the highest similarity as the target class group.

[0059] Specifically, the similarity between it and the working condition label of each class group is calculated according to the following algorithm:

[0060]

[0061] wherein, is the similarity with the working condition label of the ith class group, is the working condition data feature of the target sensor in the current detection period, is the working condition label of the ith class group, indicates the Euclidean distance, , the greater the value, the higher the similarity.

[0062] S204, based on all historical time series curves in the target class group, obtain a high-fidelity standard curve under the working condition label, and match it as the comparison time series curve of the target data.

[0063] In some embodiments, the determination of the high-fidelity standard curve under the working condition label can be: similarity calculation is performed between all historical time series curves in the working condition label and the target data time series curve in the current detection period, and the historical time series curve with the highest similarity is taken as the high-standard curve.

[0064] In other embodiments, the determination of the high-fidelity standard curve under the working condition label can also be:

[0065]

[0066] the high-fidelity standard curve under the working condition label of the i-th class group, denotes the i-th class group, denotes the m-th historical time series curve in the i-th class group, is a weight value of the m-th historical time series curve, which is calculated based on the current target data time series curve and the dynamic feature distance (Euclidean distance) between the m-th historical time series curve The curve with high similarity contributes more, and the influence of the difference in the target data change form is reduced.

[0067] Therefore, the working condition data feature aims to capture the "static" attribute (device state) and "dynamic" attribute (curve statistics) of the working condition, group similar working conditions through clustering, ensure that the high-fidelity standard curve is generated based on homogeneous historical data, avoid baseline drift caused by device failure or environmental interference, automatically adapt to different device states (such as battery depletion or load fluctuation) when comparing time series curves, avoid baseline drift caused by device failure; generate comparison time series curves based on historical real curves, ensure that they represent the ideal state without false attacks, and improve the detection accuracy of subsequent continuous or periodic fake patterns; similarity calculation uses normalized distance measurement to reduce matching complexity and is suitable for real-time Internet of Things edge devices; integrate device state and curve statistical features (range / mean / standard deviation) to construct a multi-dimensional working condition label, solve the problem that the static baseline cannot respond to working condition changes; by quantifying the similarity of the current working condition and each working condition label, the comparison time series curve is generated on demand, avoiding the waste of computing resources caused by blind comparison with a large number of historical time series curves, and improving the matching rate.

[0068] By extracting the working condition data feature, the running state of the device can be more accurately described, thereby improving the accuracy of the monitoring system. Since the working condition data feature can reflect the real running state of the device, false positives caused by noise or interference can be reduced.

[0069] S102, extract the multi-scale false event feature vector of each sampling point in the target data time series curve, input it into the pre-trained data false identification model, and output the false event result associated with the target data of the corresponding sampling point.

[0070] Specifically, the multi-scale false event feature vector includes a single-point deviation feature, a window mode feature, and a frequency domain anomaly feature, and is represented as , is a multi-scale false event feature vector of the kth sampling point, is a single-point deviation feature of the kth sampling point, is a window mode feature of the kth sampling point, is a frequency domain anomaly feature of the kth sampling point.

[0071] The single-point deviation feature is determined according to the target data of the kth sampling point in the target data time series curve and the comparison time series curve, and is specifically:

[0072]

[0073] is a single-point deviation feature, is target data of the kth sampling point in the target data time series curve, is target data of the kth sampling point in the comparison time series curve, is a standard deviation of the corresponding real and false historical target data, which can be calculated from the target data of the kth sampling point in the historical time series curve in the corresponding target class group, as a target data fluctuation benchmark under the same working condition data feature condition, to normalize the instantaneous offset and eliminate the dimension effect to obtain a relative deviation degree.

[0074] The window mode feature is determined according to the target data in a preset window length in the target data time series curve and the comparison time series curve, and is specifically:

[0075]

[0076] is a window mode feature, w is a preset window length, and is determined according to the mean value of the duration of a typical false data attack (such as a ramp attack that needs to last more than or equal to 10 sampling points) appearing in a detection period in an industrial scene (which can be set by historical data and expert experience), and is less than the length of the detection period, is target data of the zth sampling point in the window in the target data time series curve, is a mean value of target data of the comparison time series curve in the window , which is used as an ideal benchmark for the window, and the actual local mean value is compared with the ideal benchmark to ensure the timeliness of the local mode analysis. By introducing the window mode feature, false events caused by single-point misjudgment can be effectively avoided, and continuous attacks caused by single-point feature missing can be captured by the window mode feature.

[0077] It should be noted that the current kth sampling point is taken as the base point to construct the window , forming a scale progression of "single point → local", solving the continuous counterfeit mode (such as ramp attack, step attack) that single point features cannot identify, for example, maliciously injecting 10 continuous sampling points of slowly changing temperature data (imitating device overload), the single point deviation may not exceed the threshold, but the window mode feature can be triggered due to the significant increase in the cumulative deviation within the window; in addition, it should be noted that if The tail sampling points of the detection period are replaced , and the present application does not elaborate on this special case.

[0078] Wherein, the frequency domain abnormal feature is determined according to the frequency domain feature of the target data in the preset window length in the target data time sequence curve and the comparison time sequence curve, specifically:

[0079]

[0080] is the frequency domain abnormal feature of the kth sampling point, L2 norm, the present application does not elaborate on this, which is used to calculate the module length of the frequency domain vector after FFT transformation, representing the total energy in the frequency domain, the numerator is used to represent the frequency domain energy of the target data time sequence curve in the window , and the denominator is used to represent the frequency domain energy benchmark of the comparison time sequence curve corresponding to the window , is used to represent the frequency domain energy ratio of the target data time sequence curve to the comparison time sequence curve. Exemplarily, if is greater than 1, it indicates that the frequency domain energy of the target data time sequence curve is higher, implying abnormal high-frequency oscillation, and if is less than 1, it indicates that the frequency domain energy of the target data time sequence curve is lower, implying abnormal smoothness.

[0081] In some embodiments, the pre-trained data false identification model is obtained in the following manner:

[0082] S301, obtaining a large number of historical target data time sequence curves in a historical detection period of a target sensor and the identification type (real data, false data, which can be quantified as 0 and 1, 0 representing real data and 1 representing false data) of all historical sampling points of the target data, if the data is false data, it also includes the false event type (such as single point false event, continuous false event, mode attack event) corresponding to the false data.

[0083] S302, generate a multi-scale false event feature vector of all historical sampling points, and label it, and the label content is set to a false event result, including real data and false data, and when it is false data, it also includes a false event type.

[0084] The labeled false event result can also be set as a four-tuple label When it is real data, When it is a single-point false event, When it is a continuous false event, When it is a pattern attack event, Wherein, 1 represents 100% of the corresponding identification type.

[0085] It should be noted that if more training data is needed, a large amount of simulation data can be generated by simulating false data, and the simulation data can be labeled, and the present application does not repeat the description.

[0086] S303, using all labeled multi-scale false event feature vectors as a training set, training a pre-selected neural network structure to obtain a final data false identification model for outputting an identified false event result , The probability of being real data, The probability of being a single-point false event, The probability of being a continuous false event, The probability of being a pattern attack event.

[0087] Thus, the multi-scale false event feature vector is a feature vector designed for false events, which can capture the characteristics of false events by analyzing the change pattern of data at different scales, and can effectively detect false events.

[0088] S103, using a preset window multi-source data verification mechanism, obtaining a decision indicator of a false event result of each sampling point in a target data time series curve.

[0089] In some embodiments, the preset window multi-source data verification mechanism is specifically set according to an associated sensor in a current working environment of the target sensor, and specifically includes:

[0090] S401, based on a window corresponding to the kth sampling point in the current detection period, extracting an associated data sequence of the associated sensor (at least one) of the target sensor in the window.

[0091] It should be noted that the determination of the associated sensor is based on the determination of the data type that is physically associated with the target data of the target sensor or fluctuates cooperatively, which needs to be determined based on the working environment of different Internet of Things devices in the industrial scene, and is usually determined based on expert experience, for example, the temperature sensor is associated with the voltage sensor, and the present application does not repeat the description.

[0092] S402, based on the target data sequence X of the target sensor in the window and the associated data sequence Y of the associated sensor in the window, obtaining the window cooperative deviation index of the target sensor and each associated sensor.

[0093] Specifically, the window cooperative deviation index is determined according to the DTW distance between the target data sequence and the associated data sequence in the window, and the standard DTW distance, and is set as:

[0094]

[0095] Wherein, is the window cooperative deviation index corresponding to the kth sampling point of the target sensor in the current detection period, X is the target data sequence of the target sensor in the window, is the associated data sequence of the qth associated sensor of the target sensor in the window, is the window cooperative feature of the target sensor and the qth associated sensor under the current working condition data feature, and the determination method is:

[0096] In the target class group, all historical time series curves with a similarity greater than 0.8 to the comparison time series curve are filtered out, and the associated time series curve of the corresponding associated sensor is obtained (the associated time series curve collected in the same detection period and without false data);

[0097] The DTW distance between each filtered historical time series curve and its corresponding associated time series curve in the kth sampling point window is calculated, and the mean value of the DTW distance corresponding to all filtered historical time series curves is taken as the window cooperative feature.

[0098] S403, based on the window cooperative deviation index of the target sensor and each associated sensor, using a preset false event result verification algorithm to output a corresponding decision index.

[0099] Specifically, the preset false event result verification algorithm is set as:

[0100]

[0101] Wherein, is the decision index of the false event result of the kth sampling point of the target sensor in the current detection period, Q is a total number of current associated sensors of the target sensor, Qk is a window collaborative deviation index corresponding to the kth sampling point of the target sensor in the current detection period, and the maximum value is taken as the comprehensive collaborative deviation index, Q is a preset sensitivity coefficient (for example, set to 0.5, which is adjusted according to expert experience and actual situation), and controls the attenuation intensity of the deviation degree on the decision index, The higher the value is, the stronger the collaborative consistency is, and the more reliable the false event result is.

[0102] S104, based on the decision index of the false event result of the sampling point in the target data time sequence curve, detecting and processing the sampling point target data in the current detection period of the target sensor.

[0103] In some embodiments, step S104 specifically includes:

[0104] S501, combining the false event result with the decision index to obtain a target false event probability.

[0105] S501, based on the target false event probability, performing the execution of the hierarchical response strategy.

[0106] Exemplarily, when the false event result is represented as a four-tuple label , the target false event probability is set as a target four-tuple label: , and the maximum value in the target four-tuple label is taken as a false probability of the sampling point target data (the false probability threshold is set in advance, which is set according to expert experience and historical data, for example, set to 0.6):

[0107] When the maximum value corresponds to real data, if the false probability is greater than 0.6, no processing is performed, otherwise, single-point replacement is performed (the target data of the corresponding sampling point in the comparison time sequence curve is replaced to the sampling point in the current detection period);

[0108] When the maximum value corresponds to a single-point false event, if the false probability is greater than 0.6, single-point replacement is performed, otherwise, normal operation is performed and the sampling point is marked for subsequent auditing;

[0109] When the maximum value corresponds to a continuous false event, if the false probability is greater than 0.6, window replacement is performed (the target data in the window of the corresponding sampling point in the comparison time sequence curve is replaced to the window of the sampling point in the current detection period); otherwise, normal operation is performed and the sampling point window is marked for subsequent auditing;

[0110] When the maximum value corresponds to a mode attack event, if the false probability is greater than 0.6, it is immediately frozen and alarmed; otherwise, the data of the target sensor is isolated and the administrator is immediately reminded to audit.

[0111] It should be noted that if the decision index Less than the preset synergy threshold (for example, 0.4), it means serious deviation, and further auditing instruction information is needed for the single-point data or window data of the target sensor corresponding to the associated sensor.

[0112] It should be noted that for different data types (real data, single point, continuous, mode attack), differentiated false probability threshold settings can be made, which can be adjusted according to actual scenes and needs. The above is only an example, and the present application does not elaborate on this.

[0113] In summary, the target data time series curve is generated by fitting the data sequence collected by the target sensor in the current detection period. This curve reflects the change of the target data (such as temperature, humidity, etc.) over time in the current detection period. According to the working condition data characteristics of the target sensor (including device state parameters and curve statistical characteristics), the similar working condition data curve in history is matched as the comparison time curve, which ensures that the comparison curve and the current detection curve have similar environmental conditions and device states, and improves the accuracy of subsequent false data detection.

[0114] The multi-scale false event feature vector is extracted, including single-point deviation feature, window mode feature and frequency domain abnormal feature, which reflects the possible falsity of the data from different angles. The extracted feature vector is input into the pre-trained data false identification model, and the false event result corresponding to the sampling point target data is output, and the possible false data points are automatically identified by the model. The single-point deviation feature reflects the abnormality of a single data point, the window mode feature reflects the mode abnormality in a section of data, and the frequency domain abnormal feature reflects the abnormality of the data in the frequency domain. The data false identification model is a neural network model trained based on a large amount of historical data, which can automatically identify and classify false data.

[0115] The decision index is obtained by using the window multi-source data verification mechanism: based on the sampling point window in the current detection period, the associated sensor data of the target sensor is extracted, the correlation between sensors in the Internet of Things is utilized, and the reliability of false data detection is improved; by calculating the DTW distance and other indicators of the data sequence of the target sensor and the associated sensor in the window, the window collaborative deviation index is obtained, which reflects the consistency between the target sensor data and the associated sensor data; based on the window collaborative deviation index, the preset false event result verification algorithm is used to output the decision index, which is used to evaluate the reliability of the false event result. The false event result verification algorithm outputs the decision index based on the window collaborative deviation index and the sensitivity coefficient, and evaluates the reliability of the false event result.

[0116] The target data of the sampling point is detected and processed based on the decision index, and the target false event probability is obtained by combining the false event result and the decision index, which is used to evaluate the possibility of the target data of the sampling point being false data; and a hierarchical response strategy is executed, different response strategies are executed according to the target false event probability, such as single-point replacement, window replacement, freezing and alarm, etc., so as to ensure that the system can take appropriate measures according to different false data situations.

[0117] The technical solutions in the embodiments of the application have at least the following technical effects or advantages:

[0118] By capturing the "static" attributes (such as battery level, CPU load) and "dynamic" attributes (such as curve statistical feature range, mean value, standard deviation) of the device state, similar working conditions are grouped and working condition labels are set, so that the comparison time curve can automatically adapt to different device states, such as battery depletion or load fluctuation, avoiding baseline drift caused by device failure or environmental interference, and ensuring the accuracy of the comparison reference; the comparison time curve is generated based on the historical real curve, so that it can represent the ideal state without false attack, thereby improving the detection accuracy of subsequent continuous or periodic fake patterns, for example, for continuous and slowly increasing temperature data attack simulating device overload, more accurate identification can be achieved. At the same time, by quantifying the similarity of the current working condition and each working condition label, the comparison time curve is generated on demand, avoiding blind comparison with a large number of historical time curves, saving computing resources and improving matching speed. Since the working condition data features can reflect the real running state of the device, false positives caused by noise or interference can be reduced, and the accuracy of the monitoring system can be improved.

[0119] The multi-scale false event feature vector includes single-point deviation feature, window mode feature, and frequency domain anomaly feature, which analyzes the change pattern of the data from different scales. The single-point deviation feature focuses on the deviation of a single sampling point, the window mode feature analyzes the local mode through a preset window length, and the frequency domain anomaly feature captures the abnormal energy change from the frequency domain. The multi-scale false event feature vector can comprehensively capture the features of false events and effectively detect various types of false data attacks, such as single-point false events, continuous false events, and pattern attack events. The introduction of the window mode feature can effectively avoid the omission of false events caused by single-point misjudgment. For example, for the attack of slowly increasing data injected into continuous multiple sampling points, the single-point deviation may not exceed the threshold, but the window mode feature will be triggered due to the significant increase in accumulated deviation within the window, thereby improving the reliability of detection.

[0120] By extracting the associated data sequence of the associated sensor of the target sensor in the window, the window collaborative deviation index of the target sensor and each associated sensor is calculated, and the preset false event result verification algorithm is used to output the decision index, the multi-source data verification mechanism can comprehensively consider the data of multiple sensors, enhance the reliability of verification, and reduce the misjudgment caused by the abnormal data of a single sensor; The decision index comprehensively considers the collaborative deviation of the target sensor and the associated sensor, and provides a more reliable decision basis for subsequent detection and processing of the sampling point target data of the target sensor in the current detection period.

[0121] According to the false event result (real data, single-point false event, continuous false event, mode attack event) and the target false event probability, a hierarchical response strategy is executed, different processing methods are adopted for different types and probabilities of false events, such as single-point replacement, window replacement, freezing and alarm, data isolation and reminding administrator to review, etc., and differential processing can more reasonably cope with various false events and improve the security and stability of the system.

[0122] In summary, by obtaining the target data time curve and comparing the time curve, extracting the multi-scale false event feature vector and inputting the model, using the window multi-source data verification mechanism to obtain the decision index, and detecting and processing the sampling point target data based on the decision index, the false data detection of the Internet of Things sensor data is realized, which combines device state parameters, curve statistical features and associated sensor data and other multi-source information, improves the accuracy and reliability of false data detection. At the same time, through the hierarchical response strategy, the system can take appropriate measures according to different false data situations, effectively preventing the influence of false data on the Internet of Things system.

[0123] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for detecting identification of Internet of Things data, characterized in that, The method comprises the following steps: S101, obtaining a target data time series curve fitted by a target data sequence collected by a target sensor in a current detection period, and matching a comparison time series curve of the target data according to a working condition data feature of the target sensor; S102, extracting a multi-scale false event feature vector of each sampling point in the target data time series curve, inputting the multi-scale false event feature vector into a pre-trained data false identification model, and outputting a false event result associated with the target data of the corresponding sampling point; S103, obtaining a decision index of the false event result of each sampling point in the target data time series curve by using a preset window multi-source data verification mechanism; S104, detecting the sampling point target data in the current detection period of the target sensor based on the decision index of the false event result of the sampling point in the target data time series curve. 2.The method of claim 1, wherein, The working condition data features are expressed as , is the working condition data feature of the current detection period of the target sensor, B is the battery level, L is the CPU load, R is the range value of the target data time series curve, is the mean value of the target data time series curve, is the standard deviation value of the target data time series curve, is the frequency domain feature. 3.The method of claim 2, wherein, The matching of the comparison time series curve of the target data according to the working condition data feature of the target sensor specifically comprises: S201, according to the working mode of the device to which the target sensor belongs, obtaining the target data time curve in all detection periods in the history of the target sensor under the same working mode, screening out the curve without false attack injection as the historical time curve set , is the jth historical time curve, records the historical working condition data characteristics corresponding to each historical time curve, and obtains the historical working condition data characteristic set , is the historical working condition data characteristic of the jth historical time curve. S202, based on the historical working condition data feature set, using a preset clustering algorithm to group the historical time series curve set to obtain N class groups , and setting a working condition label for each class group; S203, obtaining a working condition data feature of the target sensor in the current detection period, calculating the similarity of the working condition data feature to the working condition labels of each class group, and taking the class group with the largest similarity as a target class group; S204, obtaining a high-fidelity standard curve under the working condition label based on all historical time series curves in the target class group, and matching the high-fidelity standard curve as the comparison time series curve of the target data. 4.The method of claim 3, wherein, The multi-scale spurious event feature vector includes a single-point deviation feature, a window mode feature, and a frequency domain anomaly feature, denoted as , is a multi-scale spurious event feature vector of the kth sampling point, is a single-point deviation feature of the kth sampling point, is a window mode feature of the kth sampling point, is a frequency domain anomaly feature of the kth sampling point. 5.The method of claim 4, wherein, The single-point deviation feature is determined according to the target data of the kth sampling point in the target data time series curve and the comparison time series curve; The window mode feature is determined according to the target data in a preset window length in the target data time series curve and the comparison time series curve; The frequency domain abnormal feature is determined according to the frequency domain features of the target data in a preset window length in the target data time series curve and the comparison time series curve. 6.The method of claim 5, wherein, The window mode feature specifically comprises: w is a preset window length, is the target data in the window of the zth sampling point, is the mean value of the target data in the window of the comparison time series. 7.The method of claim 4, wherein, The pre-trained data false identification model is obtained in the following way: S301, obtaining historical target data time series curves in a large number of historical detection periods of the target sensor and identification types of the target data of all historical sampling points; S302, generating multi-scale false event feature vectors of all historical sampling points and labeling the multi-scale false event feature vectors, and the labeling content is set as a false event result; S303, taking all labeled multi-scale false event feature vectors as a training set, training a preselected neural network structure, and obtaining a final data false identification model. 8.The method of claim 7, wherein, In the S302, the labeled false event result is set as a four-tuple label When it is real data, When it is a single-point false event, When it is a continuous false event, When it is a pattern attack event, . 9.The method of claim 6, wherein, The preset window multi-source data verification mechanism specifically comprises: In S401, a window corresponding to a kth sampling point in a current detection period is determined In S403, a correlation data sequence corresponding to the window is extracted from the target sensor. S402, obtaining window collaborative deviation indexes of the target sensor and each associated sensor based on a target data sequence of the target sensor in a window and an associated data sequence of an associated sensor in the window; S403, outputting a corresponding decision index by using a preset false event result verification algorithm based on the window collaborative deviation indexes of the target sensor and each associated sensor. 10.The method of claim 9, wherein, The preset false event result verification algorithm is set as: wherein, is a decision indicator of the false event result of the target sensor at the kth sampling point in the current detection period, is a comprehensive collaborative deviation indicator, Q is the total number of the current associated sensors of the target sensor, is a window collaborative deviation indicator corresponding to the kth sampling point in the current detection period of the target sensor, is a preset sensitivity coefficient, controlling the attenuation intensity of the deviation degree on the decision indicator.

Citation Information

Patent Citations

  • Sensor false data detection method and device, computer equipment and storage medium

    CN115081524A

  • Smart home management method and system based on Internet of Things

    CN120762296A