Data processing method based on Internet of Things multi-source information
By preprocessing and quality evaluation of IoT multi-source data, the problems of data heterogeneity and quality instability are solved, adaptive data processing is achieved, the accuracy and efficiency of data fusion are improved, and the real-time and reliability of IoT systems are ensured.
Patent Information
- Application Number
- CN202510961548.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
There are problems in the multi-source data processing of IoT multi-source data, high data heterogeneity, unstable data quality and low multi-source data fusion efficiency, especially when data abnormalities cannot be adaptively adjusted.
By pre-processing multiple heterogeneous data, eliminating outliers, synchronizing timestamps and converting protocol formats, and establishing a data quality evaluation mechanism, and starting a data completion process when the data credibility is lower than the threshold to realize adaptive processing.
It realizes adaptive adjustment and processing when data abnormalities are abnormal, improves the accuracy and efficiency of data fusion, and ensures the real-time and reliability of the Internet of Things system.
Smart Images

Figure CN120492815A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electronic digital data processing, and in particular relates to a data processing method based on multi-source information of the Internet of Things. Background Art
[0002] With the rapid development of Internet of Things (IoT) technology, the data generated by various terminal devices (such as sensors, smart terminals, and industrial equipment) are multi-source and heterogeneous, including information with different protocols, different sampling frequencies, and different data formats. This information usually involves multiple fields such as environmental monitoring, industrial control, and smart cities, and is characterized by strong real-time performance, large data volume, and diverse dimensions.
[0003] Currently, IoT multi-source data processing faces the following technical challenges:
[0004] High data heterogeneity: Different devices use different communication protocols (such as MQTT, CoAP, and Modbus), and the data formats are not unified, making data fusion difficult.
[0005] Unstable data quality: Due to network delays, equipment failures, or environmental interference, the collected data may contain missing data, noise, or outliers, affecting the accuracy of subsequent analysis.
[0006] Low efficiency of multi-source data fusion: Traditional data processing methods usually use single data source analysis or simple weighted fusion, which makes it difficult to effectively mine the correlation features of cross-source data, resulting in insufficient decision-making accuracy.
[0007] Among existing technologies, some solutions use edge computing for data preprocessing, but this still suffers from poor adaptability of data fusion algorithms. Other solutions use cloud computing for centralized processing, but this results in high network transmission overhead and difficulty ensuring real-time performance. Furthermore, traditional methods often lack dynamic data quality assessment mechanisms, making it impossible to adaptively adjust processing when data anomalies occur. Summary of the Invention
[0008] The technical problem to be solved by the present invention is the technical problem that adaptive adjustment and processing cannot be carried out when data is abnormal. The present invention provides a data processing method based on multi-source information of the Internet of Things, which performs data preprocessing on the collected multi-dimensional heterogeneous data. The data preprocessing includes removing data outliers, synchronizing timestamps of multi-source data, and converting data of different protocols into a unified intermediate format. By establishing a data quality assessment mechanism, the data completion process is started when the credibility of the data after data preprocessing is lower than a threshold, thereby realizing adaptive adjustment and processing when data is abnormal.
[0009] In order to achieve the above object, the present invention is implemented by the following technical solutions:
[0010] A data processing method based on multi-source information of the Internet of Things comprises the following steps:
[0011] Step S1. Collecting multi-dimensional heterogeneous data from IoT terminal devices;
[0012] Step S2. Preprocess the collected multivariate heterogeneous data; data preprocessing includes: removing data outliers, synchronizing timestamps on multi-source data, and converting data from different protocols into a unified intermediate format;
[0013] Step S3. Establish a data quality assessment mechanism and start the data completion process when the credibility of the data after data preprocessing is lower than the threshold;
[0014] Step S4. Extract cross-source correlation features from the preprocessed data through feature-level fusion, and generate comprehensive evaluation results through decision-level fusion;
[0015] Step S5. Trigger the preset IoT control strategy based on the comprehensive evaluation results and output the processing results.
[0016] Optionally, in step S1, the collection of multi-dimensional heterogeneous data at least includes the collection of sensor data, device status data, and environmental parameter data.
[0017] Optionally, in step S2, based on embodiment 1, for step S2, the collected multivariate heterogeneous data is preprocessed as follows:
[0018] For removing data outliers, it is to identify and process noise data beyond the normal range, and apply it to non-normally distributed data without strict assumptions about the data distribution;
[0019] Timestamp synchronization of multi-source data is to align data from different sources and at different time granularities onto a unified timeline;
[0020] Converting data of different protocols into a unified intermediate format is to convert different communication protocols or data formats into a unified format.
[0021] Optionally, remove data outliers:
[0022] Elimination: Directly delete abnormal data that is confirmed to be invalid;
[0023] Correction: Replace outliers with interpolation or statistical values;
[0024] Marking: Mark the abnormal values that cannot be confirmed and retain the original data for subsequent verification.
[0025] Optionally, timestamp synchronization of multi-source data is performed to interpolate the data. Perform interpolation:
[0026] for but Points, using adjacent known points and calculate:
[0027] ;
[0028] At a certain point in time When there is no observation value in the original data source, the value of the point is estimated by the adjacent points. The time value is , The time value is ,but hour: .
[0029] Optionally, the following steps are performed to convert data of different protocols into a unified intermediate format:
[0030] Step a: Define a unified intermediate format;
[0031] Step b: parse the source data protocol;
[0032] Step c: field mapping and type conversion;
[0033] Step d: Complete missing fields and default values;
[0034] Step e: data verification and cleaning;
[0035] Step f: Output unified intermediate format data.
[0036] Optionally, in step S3, a data quality assessment mechanism is established by weighted calculation of multi-dimensional indicators based on data credibility, and the formula is as follows: ;
[0037] For the The weight of the evaluation indicators, For the The score of each indicator; Indicates the number of evaluation indicators involved in calculating credibility.
[0038] Optionally, in step S3, when the credibility of the data after data preprocessing is lower than a threshold, the data completion process is started as follows:
[0039] Step A: Define evaluation metrics and thresholds;
[0040] Step B: Data preprocessing and quality assessment;
[0041] Step C: Determine whether to start data completion;
[0042] Step D: Data completion process;
[0043] Step E: Re-evaluate after completion; perform quality assessment again on the completed data.
[0044] If the credibility meets the requirements, the subsequent analysis process will be entered; if not, the process will be completed in a loop until the requirements are met.
[0045] Optionally, in step S4, feature-level fusion extraction of cross-source correlation features is: feature splicing, and the specific operation is to directly splice the feature vectors of different data sources by dimension to form a high-dimensional feature vector.
[0046] Optionally, in step S5, the comprehensive evaluation result is to construct a comprehensive evaluation function:
[0047] ;
[0048] There are , For the Raw data collected by IoT sensors;
[0049] There are , For the The weight of the data; To evaluate the model function; For the comprehensive evaluation results.
[0050] Beneficial effects of the present invention:
[0051] The present invention performs data preprocessing on the collected multivariate heterogeneous data. The data preprocessing includes eliminating data outliers, synchronizing timestamps on multi-source data, and converting data of different protocols into a unified intermediate format. By establishing a data quality assessment mechanism, the data completion process is initiated when the credibility of the data after data preprocessing is lower than a threshold, thereby realizing adaptive adjustment and processing when data anomalies occur. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0053] Figure 1 Schematic diagram of the workflow of the present invention. DETAILED DESCRIPTION
[0054] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0055] Example 1:
[0056] like Figure 1 As shown, this embodiment provides a data processing method based on multi-source information of the Internet of Things, including the following steps:
[0057] Step S1. Collect multivariate heterogeneous data from IoT terminal devices, where the multivariate heterogeneous data includes at least sensor data, device status data, and environmental parameters;
[0058] Step S2. Preprocess the collected multivariate heterogeneous data, where the data preprocessing includes: removing data outliers, synchronizing timestamps on multi-source data, and converting data of different protocols into a unified intermediate format;
[0059] Step S3. Establish a data quality assessment mechanism and start the data completion process when the credibility of the data after data preprocessing is lower than the threshold;
[0060] Step S4. Extract cross-source correlation features from the preprocessed data through feature-level fusion, and generate comprehensive evaluation results through decision-level fusion;
[0061] Step S5. Trigger the preset IoT control strategy based on the comprehensive evaluation results and output the processing results.
[0062] Example 2:
[0063] Based on Example 1, step S1 includes the following steps for collecting multi-dimensional heterogeneous data from IoT terminal devices:
[0064] The peripheral sensing interface of the IoT terminal device is used to connect a variety of sensor devices, including but not limited to digital sensors (for obtaining sensor data), infrared sensors (for detecting device status), and environmental sensors (for obtaining environmental parameters). The multi-dimensional heterogeneous data collected by the sensor devices are read through the peripheral sensing interface.
[0065] Example 3:
[0066] Based on Example 1, for step S2, the collected multivariate heterogeneous data is preprocessed as follows:
[0067] Eliminating data outliers is to identify and process noise data that exceeds the normal range. It is applied to non-normally distributed data (such as skewed data) without strict assumptions about the data distribution. The following takes the IQR method as an example:
[0068] The quartiles Q1, Q3 and interquartile range IQR = Q3-Q1 were calculated, and the outlier range was defined as below Q1-1.5IQR or above Q3+1.5IQR;
[0069] Among them, the first quartile (Q1): the value of the 25% position after the data is sorted, ; The third quartile (Q3): the value of the 75% position after the data is sorted, ;
[0070] The abnormal value range is Q1-1.5IQR or less. ; The abnormal value range is Q3+1.5IQR or above, .
[0071] in, =1.5 identifies mild abnormality, =3 Identify extreme anomalies. Identify outliers if the value is less than the Lower or greater than the Upper value.
[0072] Outlier handling strategy:
[0073] Elimination: Directly delete abnormal data that has been confirmed to be invalid (to avoid excessive elimination that affects the sample size).
[0074] Correction: Replace outliers with interpolation methods (e.g., linear interpolation, polynomial interpolation) or statistical values (mean, median).
[0075] Marking: Mark the abnormal values that cannot be confirmed and retain the original data for subsequent verification.
[0076] Synchronizing timestamps of multi-source data is to align data from different sources and at different time granularities to a unified time axis. Specifically, this involves interpolating the data. Perform interpolation:
[0077] for but Points, using adjacent known points and calculate:
[0078] ;
[0079] At a certain point in time When there is no observation value in the original data source, the value of the point is estimated by the adjacent points. The time value is , The time value is ,but hour: .
[0080] For between and between , pass and It is obtained by linear combination according to the relative position ratio, so as to estimate the function value of the middle point based on the known discrete data points.
[0081] In experimental data processing, when only data at discrete time points are collected, linear interpolation can be used to estimate the data at intermediate moments for more detailed analysis.
[0082] Convert data from different protocols into a unified intermediate format, that is, convert different communication protocols (such as Modbus, OPC UA and MQTT) or data formats (such as JSON, XML and CSV) into a unified format.
[0083] The specific steps are:
[0084] Step a: Define a unified intermediate format (target format);
[0085] Determine a unified intermediate format (e.g., JSON Schema, Protocol Buffers), clarify field names, data types, and constraints (e.g., required fields, length limits). If the intermediate format is JSON, define the structure of the device data.
[0086] Step b: parse the source data protocol;
[0087] Write parsing rules to extract raw data based on the protocol type of the source data (such as Modbus, MQTT, HTTP REST, and custom binary protocols);
[0088] The source data is , the protocol parsing function is ,in, Output parsed structured data for source protocol rules (e.g. JSON parser, binary parsing template) : ;
[0089] If the source data is binary data of the Modbus protocol (for example, register address 0x0001 stores the temperature value), the parsing function needs to extract the register value according to the Modbus protocol rules and convert it to decimal (for example: ).
[0090] If the source data is in CSV format, the parsing function needs to extract fields based on comma delimiters (e.g., device_id, value, and timestamp are converted to objects).
[0091] Step c: field mapping and type conversion;
[0092] The parsed structured data The fields are mapped to the fields in the intermediate format and data type conversion is handled (such as string to value, time format conversion).
[0093] Defining a mapping function ,in, Output temporary data in an intermediate format for the mapping rule table (source field → target field, including type conversion logic) .
[0094] Specifically, the fields are renamed: Source data The dev_id in is mapped to device_id.
[0095] Type conversion: The string time 2023-10-01, 12:00:00 in the source data is converted to the Unix timestamp 1696161600.
[0096] Logical processing: The original voltage value in the source data needs to be converted to the actual value using the formula: voltage = original value × 0.1.
[0097] Step d: Complete missing fields and default values;
[0098] For mandatory fields in the intermediate format that are missing in the source data, default values must be added or obtained through other means (e.g., extracted from metadata or device configuration).
[0099] If the source data does not provide the metadata field, it is completed with metadata:{} (empty object).
[0100] Step e: data verification and cleaning;
[0101] Verify whether the converted data meets the constraints of the intermediate format (such as whether the field exists and whether the value is within a reasonable range), eliminate invalid data, or trigger error handling procedures.
[0102] Define the validation function , where Schema is the validation rule in the intermediate format (e.g. JSON Schema). To output the completed data, the verification function outputs valid data Or error: .
[0103] Verification rule example: The length of device_id cannot exceed 32 characters.
[0104] Step f: outputting unified intermediate format data;
[0105] Format the verified data into the final intermediate format (such as JSON string, Protobuf binary data) for subsequent processing (such as storage, analysis); the formatting function is , output the result : .
[0106] Convert the input data into the specified format. Indicates the target format type, which is different data format types such as JSON, XML and Protobuf. The parameter formatting function should convert valid data Convert to a specific format, the formula is to call the formatting function , the valid data According to the specified Format is processed.
[0107] Further clarified In JSON format, Indicates the target data obtained after formatting, that is, the final output data that meets the JSON format requirements. Calling the formatting function , the valid data Convert to JSON format.
[0108] It is to convert valid data in the form of objects into JavaScript environment (similar principles in other programming environments that support JSON processing) Convert to JSON string. In many scenarios, the final data storage and transmission operations need to process JSON data in the form of strings. For example, in Web development, when the front-end and back-end interact with each other, the data returned by the back-end to the front-end is usually transmitted in the form of JSON strings. So through this operation, the previously processed valid data is converted to JSON string. Convert it into target JSON data that can be used in actual application scenarios (such as storing it in a file or transmitting it over the network) .
[0109] Example 4:
[0110] Based on Example 1, for step S3, a data quality assessment mechanism is established by weighted calculation of multi-dimensional indicators based on data credibility (confidence), and the formula is as follows: ;
[0111] For the The weight of the evaluation indicators (multi-dimensional indicators) The weight needs to be set according to business needs, such as the integrity weight may be higher). The weight is a value between 0 and 1, and the sum of all indicator weights is 1, that is, , the weight reflects the relative importance of the indicator in evaluating data credibility. For example, in data quality assessment, the weight of the integrity indicator , which shows that integrity accounts for a higher proportion in the overall credibility assessment and is more concerned.
[0112] For the The score of each indicator (usually standardized to 0-1 or 0-100 points) is usually obtained according to the calculation rules of the specific indicator and is generally standardized to a range of 0-1 or 0-100. For example, the score of the integrity indicator , indicating that the data performed well in terms of completeness, with a score of 80% (converted to a range of 0−1).
[0113] Indicates the number of evaluation indicators involved in calculating credibility. For example, if the three indicators of completeness, accuracy and consistency are involved in the calculation, then =3;
[0114] This formula, used to calculate data credibility, uses a weighted summation approach. It combines multiple factors that influence data credibility, assigning each factor a corresponding weight, and then summing them to arrive at a composite credibility value. Higher credibility values generally indicate better data quality and, therefore, higher credibility. Credibility is the final calculated result, representing the overall credibility of the data and serving as a comprehensive indicator.
[0115] When the credibility of the data after data preprocessing is lower than the threshold, the data completion process is started as follows:
[0116] Step A: Define evaluation metrics and thresholds;
[0117] Determine core indicators: Filter key indicators based on needs (for example, financial scenarios focus more on accuracy, while IoT scenarios focus more on timeliness).
[0118] Set weights: completeness 40%, accuracy 30%, consistency 20%, timeliness 10%.
[0119] Set threshold: The confidence threshold is set to 0.7 (i.e. 70 points). The completion process starts when the confidence threshold is lower than this value.
[0120] Step B: Data preprocessing and quality assessment;
[0121] Data cleaning: remove duplicate values, handle outliers and standardize the format.
[0122] Indicator calculation: Calculate the scores of each dimension and the overall credibility.
[0123] If the threshold is 0.7, the credibility is met; if the threshold is 0.75, completion is triggered.
[0124] Step C: Determine whether to start data completion;
[0125] Condition trigger: If the comprehensive credibility is less than the threshold, the completion process will begin.
[0126] Identify the problem: Analyze the scores of each dimension and identify the weak links (e.g., insufficient completeness → many missing values, insufficient accuracy → many data errors).
[0127] Step D: Data completion process;
[0128] Choose a completion strategy based on the question type:
[0129] Missing value completion:
[0130] Numerical: mean / median filling, regression model prediction, time series model (such as ARIMA) prediction.
[0131] Classification: majority filling, hot encoding + clustering algorithm filling, dictionary mapping (such as: matching through business rules).
[0132] Error value correction:
[0133] Compare and correct with authoritative data sources (such as third-party APIs).
[0134] Verified by rule engine (such as verification rules, logical consistency rules).
[0135] Timely completion:
[0136] Collect incremental data in real time (e.g., get the latest data through a message queue).
[0137] Supplement historical data (e.g., extract from data archives).
[0138] Consistent completion:
[0139] Unify field definitions across multiple data sources (e.g., standardize fields through a data dictionary).
[0140] Arbitrate conflicting values (e.g., based on the primary data source).
[0141] Step E: Re-evaluate after completion;
[0142] Perform quality assessment again on the completed data.
[0143] If the credibility meets the requirements, the subsequent analysis process will be entered; if not, the process will be completed in a loop until the requirements are met or it is marked as "unavailable data".
[0144] Example 5:
[0145] Based on Example 1, for step S4, the feature-level fusion algorithm for extracting cross-source correlation features is as follows:
[0146] Feature concatenation is a process that directly concatenates feature vectors from different data sources by dimension to form a feature vector with a higher dimension.
[0147] The specific steps are:
[0148] Feature preprocessing: Normalization / standardization: The numerical ranges of different feature sources vary greatly (for example, image pixel values are [0, 255], and sensor data are [-10, 10]). Normalization (such as Min-Max scaling) or standardization (such as Z-Score) is needed to unify the scale to prevent features with large values from dominating model training.
[0149] Missing value processing: If there are missing features, they need to be supplemented by interpolation (such as mean, median, and KNN interpolation) or filling (such as constant filling).
[0150] Feature alignment: Ensure that samples from different feature sources correspond one to one (e.g., observation data of the same object at the same time point) to avoid sample misalignment during stitching.
[0151] Dimension splicing: Splicing by row (same sample dimension): If the number of samples of each feature source is the same (both are samples), then directly concatenate by feature dimension. For example:
[0152] Feature source A: ; samples, dimensional features); is a matrix whose elements are all real numbers, Row (corresponding to samples) and Columns (corresponding to each sample features), belonging to the set of real matrices , is a real number.
[0153] Feature source B: ; ( samples, dimensional features); is a matrix whose elements are all real numbers, Row (corresponding to samples) and Columns (corresponding to each sample features), belonging to the set of real matrices , is a real number.
[0154] After splicing: ,in, Indicates column-wise concatenation.
[0155] Column-wise concatenation (with consistent feature dimensions): This is used to concatenate different subsets of the same feature source (e.g., different feature groups from the same dataset).
[0156] Feature concatenation is a common method for feature-level fusion. It combines feature vectors from different sources or types by dimension to form a new feature vector containing multi-source information. This directly stacks features, preserving the original structure of each source feature. This ensures that the concatenated features are physically related (for example, different attributes of the same object) and avoids meaningless feature stacking. Feature concatenation directly integrates raw information from multiple sources, providing a foundation for subsequent feature extraction (for example, automatically learning cross-source correlation features through neural networks) or decision-level fusion.
[0157] The comprehensive evaluation results of decision-level integration are as follows:
[0158] Decision-level fusion is an integration at the model prediction result layer, combining the decisions of multiple single-source models or feature-level fusion models to generate a final comprehensive evaluation. It uses a meta-model (such as logistic regression and neural network) to learn the mapping relationship between the prediction results of the base model and the true label.
[0159] The steps of decision-level fusion are:
[0160] The base model predicts the training set and generates a new feature matrix (Each row is the predicted value of each base model for the sample);
[0161] Select multiple base models of different types or based on different feature subsets (e.g., decision trees, support vector machines, and neural networks). Train each base model independently on the training dataset. For example, select a decision tree, linear regression, and random forest as base models, and perform parameter optimization and model fitting on each of them on the training dataset.
[0162] Use the trained base model to predict the training set and test set respectively. For the training set, each base model will output a prediction value vector; for the test set, the corresponding prediction value vector is also obtained. Assume that the training set has samples, there are base models, then each base model will get a length of After predicting the test set, a prediction value vector with a length of the test set samples (set as ) is a vector of predicted values.
[0163] The metamodel is As input, training outputs the final result ;
[0164] The prediction results of the base model for the training set are used as new features to construct a new training dataset. Specifically, the prediction values of the base model are used as feature columns, and the labels of the original training set are used as target columns. If there is base model, then the feature dimension of the new training dataset is , the number of samples is still After predicting the training set with the three base models of decision tree, linear regression and random forest, the predicted values of these three models are used as new features to form new training data.
[0165] For the test set, the base model predicts , the meta-model predicts .
[0166] Select a meta-model (common simple models include linear regression and logistic regression, because the meta-model mainly learns the relationship between the base model prediction results and the true label, and does not need to be too complicated), and use the newly constructed training dataset to train the meta-model.
[0167] The predictions of the base model on the test set are used as input to the meta-model, which then outputs the final predictions. For example, the predictions of the decision tree, linear regression, and random forest on the test set are fed into the trained logistic regression meta-model to produce the final predictions.
[0168] Example 6:
[0169] Based on Example 1, for step S5, a comprehensive evaluation function is constructed:
[0170] ;
[0171] There are , For the Raw data collected by IoT sensors (such as temperature, humidity, voltage, etc.);
[0172] There are , For the The weight of each data point (reflecting its importance to the evaluation results, which can be determined through hierarchical analysis and machine learning training);
[0173] To evaluate model functions (e.g. weighted sum, neural networks, and fuzzy logic);
[0174] It is a comprehensive assessment result (i.e., value or status, such as risk level).
[0175] A comprehensive evaluation function is a tool that integrates multiple evaluation indicators (qualitative or quantitative data) into a comprehensive evaluation value. Its core functions are:
[0176] Multi-dimensional quantitative analysis: Decompose complex evaluation objects (such as equipment status, environmental parameters, and system performance) into measurable indicators, and then achieve cross-dimensional quantitative integration through calculations.
[0177] Decision support: Provides a unified assessment benchmark for IoT systems, smart devices, or management platforms to trigger control strategies (such as threshold alarms, automatic adjustments, and equipment start and stop).
[0178] Dynamic monitoring and optimization: Through real-time data input, the evaluation results are continuously updated to support dynamic adjustment of the system.
[0179] In the Internet of Things (IoT) system, control strategy trigger conditions refer to the rules or thresholds that determine whether to activate a preset control strategy after logical judgment or algorithmic calculation based on input information such as sensor data, status parameters, or external commands. This establishes a mapping between "input conditions" and "control actions" to ensure that the system automatically executes the corresponding action when specific conditions are met.
[0180] Specifically, the triggering conditions of the IoT control strategy are:
[0181] ;
[0182] in, For the The trigger threshold range of each policy (such as: ≥80 triggers alarm strategy);
[0183] It is a preset control strategy (such as equipment start and stop, threshold adjustment).
[0184] Send instructions to actuators (such as relays, control valves, and alarms) through the IoT platform;
[0185] Record the metadata of the policy execution time and parameters (e.g., execution time is , the parameter metadata is (i.e. starting voltage)).
[0186] Design principles for trigger conditions:
[0187] Accuracy: Ensure that the data source is reliable (e.g., sensors are calibrated regularly) to avoid false triggering or missed triggering due to data errors.
[0188] Real-time performance: For scenarios with strong timeliness (such as industrial equipment failure warnings), it is necessary to ensure that the delay in data collection and judgment is sufficiently low.
[0189] Flexibility: Supports dynamic adjustment of thresholds or rules (e.g., remote modification of trigger conditions through a cloud platform) to adapt to different scenario requirements.
[0190] Trigger conditions are the core logic for automated control in the IoT. Specific design must integrate business needs, technical feasibility, and scenario constraints to ensure the system responds to environmental changes at the right time and in the right way. By properly configuring data sources, decision logic, and threshold rules, the intelligence and reliability of IoT applications can be significantly improved.
[0191] Output:
[0192] Collect feedback data after execution (For example, the device status changes to "off", or the parameter stabilizes at the target value). Generate processing results :
[0193] Text report: " implement , the device has been shut down and is currently in normal status."
[0194] Visual output: Mark the policy execution effect on the monitoring interface (for example, use a red icon to indicate device shutdown).
[0195] Data storage: 、 and Stored in the database for historical tracing and model optimization.
[0196] Directly intervene in physical equipment or the environment through actuators (such as valves and switches) to meet pre-defined control objectives. Prevent equipment failures or safety incidents through early warning and protection mechanisms, ensuring system operation within safe thresholds. Dynamically adjust resource allocation based on real-time data to reduce energy consumption and improve efficiency.
[0197] By recording output results and execution effects, a closed loop of "data collection-strategy execution-effect feedback-strategy optimization" is formed to improve the intelligence level of the system.
[0198] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope of the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A data processing method based on multi-source information of the Internet of Things, characterized in that: The steps include: Step S1. Collecting multi-dimensional heterogeneous data from IoT terminal devices; Step S2. Preprocess the collected multivariate heterogeneous data; data preprocessing includes: removing data outliers, synchronizing timestamps on multi-source data, and converting data from different protocols into a unified intermediate format; Step S3. Establish a data quality assessment mechanism and start the data completion process when the credibility of the data after data preprocessing is lower than the threshold; Step S4. Extract cross-source correlation features from the preprocessed data through feature-level fusion, and generate comprehensive evaluation results through decision-level fusion; Step S5. Trigger the preset IoT control strategy based on the comprehensive evaluation results and output the processing results.
2. The data processing method based on multi-source information of the Internet of Things according to claim 1 is characterized in that: In step S1, the collection of the multi-dimensional heterogeneous data at least includes the collection of sensor data, device status data, and environmental parameter data.
3. The data processing method based on multi-source information of the Internet of Things according to claim 1 is characterized in that: In step S2, based on embodiment 1, for step S2, the collected multivariate heterogeneous data is preprocessed as follows: For removing the data outliers, the method is to identify and process the noise data beyond the normal range, and to apply it to non-normally distributed data without strict assumptions about the data distribution; The timestamp synchronization of multi-source data is to align data from different sources and at different time granularities to a unified time axis; Converting data of different protocols into the unified intermediate format is converting different communication protocols or data formats into a unified format.
4. The data processing method based on multi-source information of the Internet of Things according to claim 3 is characterized in that: Eliminating the data outliers is: Elimination: Directly delete abnormal data that is confirmed to be invalid; Correction: Replace outliers with interpolation or statistical values; Marking: Mark the abnormal values that cannot be confirmed and retain the original data for subsequent verification.
5. The data processing method based on multi-source information of the Internet of Things according to claim 3 is characterized in that: The timestamp synchronization of the multi-source data is to interpolate the data, that is, to Perform interpolation: for but Points, using adjacent known points and calculate: ; At a certain point in time When there is no observation value in the original data source, the value of the point is estimated by the adjacent points. The time value is , The time value is ,but hour: .
6. The data processing method based on multi-source information of the Internet of Things according to claim 3 is characterized in that: The steps for converting the data of different protocols into the unified intermediate format are as follows: Step a: Define a unified intermediate format; Step b: parse the source data protocol; Step c: field mapping and type conversion; Step d: Complete missing fields and default values; Step e: data verification and cleaning; Step f: Output unified intermediate format data.
7. The data processing method based on multi-source information of the Internet of Things according to claim 1 is characterized in that: In step S3, the data quality assessment mechanism is established by weighted calculation of multi-dimensional indicators based on data credibility, and the formula is as follows: ; For the The weight of the evaluation indicators, For the The score of each indicator; Indicates the number of evaluation indicators involved in calculating credibility.
8. The data processing method based on multi-source information of the Internet of Things according to claim 1 is characterized in that: In step S3, when the credibility of the data after the data preprocessing is lower than the threshold, the data completion process is started as follows: Step A: Define evaluation metrics and thresholds; Step B: Data preprocessing and quality assessment; Step C: Determine whether to start data completion; Step D: Data completion process; Step E: Re-evaluate after completion; perform quality assessment on the completed data again; If the credibility meets the standard, enter the subsequent analysis process; If the requirements are not met, complete the process repeatedly until they are met.
9. The data processing method based on multi-source information of the Internet of Things according to claim 1 is characterized in that: In step S4, the feature-level fusion extraction of cross-source correlation features is feature splicing, and the specific operation is to directly splice the feature vectors of different data sources according to the dimensions to form a high-dimensional feature vector.
10. The data processing method based on multi-source information of the Internet of Things according to claim 1, characterized in that: In step S5, the comprehensive evaluation result is to construct a comprehensive evaluation function: ; There are , For the Raw data collected by IoT sensors; There are , For the The weight of the data; To evaluate the model function; For the comprehensive evaluation results.
Citation Information
Patent Citations
Multi-source data investigation and real-time analysis method
CN116894152A
Building safety alert level dynamic adjustment method and system based on artificial intelligence driving
CN119204699A
Automatic DCMM evaluation system based on multi-source heterogeneous data integration mechanism
CN119557297A
Efficient processing method for real-time acquisition and synchronization of multi-source heterogeneous data
CN120234530A
Massive multi-source and multi-modal data fusion method
CN120277619A