An iot edge data quality joint repair method and system

CN122679031APending Publication Date: 2026-09-01GANZHOU LINGFANGE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610999955.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

然而在实际物理系统中,上述异常往往同时发生且强耦合(例如传感器故障可能同时导致数据漂移和间歇性缺失)

Benefits of technology

[0025]通过多异常联合感知与协同修复模块,在单次前向推理中同步输出异常分类概率和协同修复值,使修复值同时最小化多种异常类型的损失函数,有助于缓解线性流水线模式下容易出现的策略冲突问题;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122679031A_ABST
    Figure CN122679031A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for joint repair of data quality at the IoT edge, relating to the fields of IoT and edge computing. The method includes: acquiring a standardized data sequence; synchronously outputting anomaly classification probabilities and collaborative repair values ​​through a multi-task learning neural network with a shared encoder, achieving joint perception and collaborative repair of multiple anomalies; generating a dynamic adjustment factor based on the data stream variance change rate to adjust the anomaly detection threshold, achieving zero-overhead online adaptiveness; generating a confidence score through a computational complexity index, and deciding on immediate edge repair, on-demand offloading to the cloud, or triggering an alarm based on the confidence score. This invention aims to achieve millisecond-level, adaptive joint repair of multiple anomalies at resource-constrained edge devices and establish a dynamic edge-cloud collaborative mechanism for on-demand offloading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of the Internet of Things (IoT) and edge computing, specifically to a method and system for improving the data quality of IoT in resource-constrained scenarios. Background Technology

[0002] In IoT systems, sensor data is the foundation for intelligent decision-making and control. However, due to limitations in sensor accuracy, communication link stability, environmental interference, and equipment aging, the raw data streams collected generally suffer from quality issues such as missing data, drift, spikes, and duplication, severely misleading subsequent data analysis and status monitoring. Therefore, efficient and accurate data quality repair at or near the source of data generation has become a key challenge for the reliable operation of IoT systems.

[0003] Currently, IoT data quality restoration technologies are mainly divided into two categories: centralized cloud-based solutions and lightweight edge-based solutions. Cloud-based solutions (such as Generative Adversarial Imputation Network (GAIN) and Bidirectional Recurrent Imputation for TimeSeries (BRITS)) utilize deep learning models to perform restoration in cloud data centers, offering high accuracy but significant latency—data upload, cloud inference, and result feedback typically take hundreds of milliseconds or even seconds, failing to meet the millisecond-level real-time requirements of scenarios such as industrial control and autonomous driving. Lightweight edge-based solutions (such as linear interpolation and Kalman filtering) achieve rapid response on edge devices, but suffer from low restoration accuracy and fixed model parameters, making them unable to adapt to dynamic changes in the statistical characteristics of data streams.

[0004] A common characteristic of the aforementioned solutions is that they typically treat anomalies such as missing data, drift, and spikes as independent events, handling them separately using a linear pipeline approach of "detect first, repair later." However, in real-world physical systems, these anomalies often occur simultaneously and are strongly coupled (for example, sensor failure may simultaneously cause data drift and intermittent missing data). In such strongly coupled scenarios, the aforementioned solutions are prone to policy conflicts—for instance, interpolation operations performed earlier may "solidify" spikes into normal data points, causing subsequent spike detection to fail; while smoothing filtering performed earlier may smooth out short-term missing features, making it impossible for the interpolation algorithm to accurately identify the missing location. Furthermore, the aforementioned solutions typically adopt a "static layering" model for edge-cloud collaboration, meaning all data is fixed to be processed at the edge or uploaded to the cloud, making it difficult to flexibly adjust processing strategies according to the real-time complexity of the data stream, and posing risks of wasted network bandwidth or failure to meet real-time requirements. Summary of the Invention

[0005] (a) Purpose of the invention

[0006] The purpose of this invention is to overcome at least some of the technical problems in the background art mentioned above, and to provide a method and system for joint repair of IoT edge data quality. It aims to achieve millisecond-level, adaptive joint repair of multiple anomalies at resource-constrained edge devices, and to establish a dynamic edge-cloud collaboration mechanism that can be unloaded on demand, thereby achieving a dynamic balance between real-time performance and accuracy.

[0007] (II) Technical Solution

[0008] like Figure 1 As shown, to achieve the above objectives, this invention provides a joint repair method for IoT edge data quality, comprising the following steps:

[0009] The data acquisition step involves acquiring the raw data stream collected from the edge, and performing timestamp alignment, unit unification, and format standardization on the raw data stream to obtain a standardized data sequence X.

[0010] In the joint anomaly detection and collaborative repair step, the standardized data sequence X is input into the joint anomaly detection and collaborative repair module. For example... Figure 2 As shown, the joint anomaly perception and collaborative repair module employs a multi-task learning neural network with a shared encoder to simultaneously output anomaly classification probability vector P and collaborative repair value y_t. The anomaly classification probability vector P includes missing probability, drift probability, spike probability, and repetition probability. When multiple anomaly probabilities reach or exceed preset thresholds, the calculation of the collaborative repair value y_t simultaneously minimizes the weighted sum of the loss functions corresponding to multiple anomaly types, resulting in a global compromise solution that considers the coupling relationships between multiple anomalies.

[0011] The online adaptive step involves inputting the standardized data sequence X into a zero-overhead online adaptive engine. For example... Figure 3 As shown, the zero-overhead online adaptive engine calculates the variance change rate Δσ² based on the sliding window statistics of the data stream, and generates a dynamic adjustment factor α through a mapping function. This dynamic adjustment factor α is used to adjust the anomaly detection threshold of the joint anomaly perception and collaborative repair module. The entire adaptive process does not involve online model training or gradient backpropagation.

[0012] The complexity assessment and decision-making steps involve inputting the standardized data sequence X into the data flow complexity estimator. For example... Figure 4As shown, the data flow complexity estimator calculates the complexity index of the current data window and generates a confidence score C based on the complexity index. Different repair paths are executed according to the confidence score C: when C is below a first threshold, the collaborative repair value y_t output by the joint anomaly detection and collaborative repair module is used as the final repair result; when C is between the first and second thresholds, the data is offloaded to the cloud for deep repair; when C is above the second threshold, the current data is marked as untrustworthy and an alarm is triggered.

[0013] The output integration step integrates the repair results and outputs them in chronological order to ensure the temporal consistency of the output data stream.

[0014] Preferably, the shared encoder consists of convolutional layers and gated recurrent unit (GRU) layers; the joint anomaly perception and collaborative repair module further includes an anomaly classification head and a collaborative repair head, wherein the anomaly classification head is a fully connected layer with a softmax activation function, and the collaborative repair head is a fully connected layer with a linear activation function.

[0015] Preferably, in the weighted sum of the loss function, for the case where both drift and spike are detected simultaneously, the loss function is L = λ_1·L_smooth(y_t) + λ_2·L_spike(y_t), where L_smooth is the second-order difference smoothing loss, L_spike is the median-based suppression loss, and λ_1 and λ_2 are preset weight coefficients; the co-repair value y_t is iteratively solved during inference using the gradient descent method.

[0016] Preferably, the zero-overhead online adaptive engine maintains two sliding windows of length N, calculates the mean μ_t and variance σ_t² of the current window, and uses the mean μ_t to calculate the variance σ_t²; the variance change rate Δσ² = (σ_t² − σ_{tN}²) / (σ_{tN}² + ε), where ε is a very small positive number to prevent division by zero; the mapping function is α = 1.0 + 0.5·tanh(β·Δσ²), where β is the sensitivity coefficient; the dynamic adjustment factor α is restricted to a preset interval.

[0017] Preferably, the complexity index includes energy spectrum entropy and sample entropy; the confidence score C = w_1·E_psd_normalized + w_2·SampEn_normalized, where E_psd_normalized is the normalized value of energy spectrum entropy, SampEn_normalized is the normalized value of sample entropy, and w_1 and w_2 are preset weight coefficients.

[0018] Preferably, in the complexity assessment and decision-making step, the first threshold is 0.7 and the second threshold is 0.9.

[0019] Preferably, in the complexity assessment and decision-making step, when C is between the first threshold and the second threshold, only the 10 data points with the highest abnormal probability in the current data window and their 5 neighboring points before and after, for a total of 20 data points, are unloaded; if no response is received from the cloud within the preset timeout period, the collaborative repair value y_t is used as the repair result.

[0020] The present invention also provides an IoT edge data quality joint repair system, including an edge processing subsystem and a cloud enhancement subsystem.

[0021] The edge processing subsystem is deployed on IoT edge devices and includes: a data stream preprocessing module for preprocessing the raw data stream and outputting a standardized data sequence X; a joint anomaly perception and collaborative repair module, employing a multi-task learning neural network with a shared encoder to simultaneously output anomaly classification probability vector P and collaborative repair value y_t; a zero-overhead online adaptive engine for calculating the variance change rate based on the sliding window statistics of the data stream and generating a dynamic adjustment factor α through a mapping function; a data stream complexity estimator for calculating the complexity index of the current data window and generating a confidence score C; a task unloading decision-maker for deciding the repair path based on the confidence score C; and an output integrator for integrating and outputting the repair results in chronological order.

[0022] The cloud-based enhancement subsystem is deployed on a cloud server to receive and repair data offloaded from the edge, and to periodically update the edge model. Preferably, the cloud-based enhancement subsystem includes a high-precision deep repair model and a model update and distillation module; the high-precision deep repair model is based on a Transformer model; the model update and distillation module is used to periodically fine-tune the high-precision model using historical data, and to compress knowledge into a lightweight version through knowledge distillation and distribute it to the edge.

[0023] (III) Beneficial Effects

[0024] The technical solution of this invention has the following beneficial effects:

[0025] Through the multi-anomaly joint perception and collaborative repair module, the anomaly classification probability and collaborative repair value are output simultaneously in a single forward inference, so that the repair value minimizes the loss function of multiple anomaly types at the same time, which helps to alleviate the policy conflict problem that is easy to occur in the linear pipeline mode.

[0026] The zero-overhead online adaptive engine generates a dynamic adjustment factor to adjust the detection threshold using only the variance change rate of the data stream, without the need for online model training or gradient backpropagation, which helps to achieve dynamic adaptation with low computational overhead on resource-constrained edge devices.

[0027] By using a dynamic task offloading mechanism based on data flow complexity, intelligent decisions can be made based on confidence scores to either perform immediate edge repair or offload tasks to the cloud as needed, which helps to achieve a dynamic balance between real-time performance and accuracy. Attached Figure Description

[0028] Figure 1 This is a diagram of the overall system architecture of the present invention;

[0029] Figure 2 This is a structural diagram of the combined anomaly detection and collaborative repair module of the present invention;

[0030] Figure 3 This is the control loop diagram of the zero-overhead online adaptive engine of the present invention;

[0031] Figure 4 This is a flowchart of the data flow complexity assessment and dynamic unloading decision-making process of the present invention. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0033] Example 1: Vibration Monitoring of Industrial Motors

[0034] This embodiment applies the present invention to an industrial motor vibration monitoring scenario.

[0035] The system edge is deployed as an industrial IoT gateway near the motor, using an STM32F407 microcontroller (ARM Cortex-M4 core, 168MHz clock speed, 192KB memory). The cloud-based enhancement subsystem is deployed on the factory's private cloud server.

[0036] A piezoelectric accelerometer acquires vibration signals from the motor housing at a sampling rate of 1 kHz. After the data stream enters the edge, the preprocessing module performs timestamp alignment and unit normalization, extracting data from the most recent 64 time steps to form a normalized data sequence X.

[0037] The zero-overhead adaptive engine maintains two sliding windows of length 100, calculating the mean μ_t and variance σ_t² in real time. When the motor is loaded from no-load to rated load, the variance σ_t² surges from 0.5 to 2.0 in a short period of time. The engine calculates the variance change rate Δσ²≈3.0, and generates an adjustment factor through the mapping function α = 1.0 + 0.5·tanh(0.5×3.0) ≈ 1.453 to amplify the spike detection threshold, which helps to avoid misjudging the normal increase in vibration amplitude caused by load changes as abnormal spikes.

[0038] The data flow complexity estimator calculates the energy spectral entropy and sample entropy for the current window. When the motor is running stably, both are low, and the confidence score C is far below the first threshold of 0.7. The task unloading decision-maker selects the edge-end immediate repair path.

[0039] The joint sensing module model consists of a shared encoder comprised of two 1D convolutional layers and two GRU layers, an anomaly classification head, and a collaborative repair head. The anomaly classification head is a fully connected layer with a Softmax activation function, while the collaborative repair head is a fully connected layer with a linear activation function. Assuming that bearing wear causes slow drift (p_{drift}=0.85) within the current window, and there are also spikes caused by metal particle impact (p_{spike}=0.92), the loss function of the collaborative repair head includes both smoothing and suppression losses. The output repair value can reduce the spike amplitude while preserving the drift trend of the vibration signal. The entire inference process takes approximately 800μs, meeting the 1ms real-time requirement.

[0040] When a loose sensor cable causes intermittent signal interruptions, the confidence score C spikes to 0.85 (between the first threshold of 0.7 and the second threshold of 0.9). The task offloading decision-maker selects a medium-confidence path, offloading the 10 data points with the highest anomaly probability and their five nearest neighbors (a total of 20 points, approximately 800 bytes) to the cloud. The cloud-based BRITS model uses global context to infer missing values ​​and returns the repair results within a 500ms timeout.

[0041] The cloud performs weekly fine-tuning of the Transformer model based on accumulated high-quality vibration data, compresses it into a lightweight version through knowledge distillation, and distributes it to the edge to continuously improve repair accuracy.

[0042] Example 2: GPS positioning for autonomous driving

[0043] This embodiment applies the invention to an autonomous driving scenario, with the on-board edge computing unit being an NVIDIA Jetson TX2 and the cloud being a roadside edge computing node.

[0044] GPS and IMU data streams enter the Jetson TX2. When a vehicle enters a densely populated area with tall buildings, GPS signals are blocked and lost, while multipath effects cause drift. The joint sensing module uses GPS and IMU data as joint inputs, the anomaly classification head simultaneously identifies the loss and drift, and the collaborative repair head uses short-term motion information from the IMU to output a position repair value, which can correct drift while filling in the loss.

[0045] The zero-overhead adaptive engine monitors changes in the missing rate and dynamically adjusts the interpolation window size. The task offloading decision-maker selects edges for immediate repair in open road sections; when entering complex intersections, the C value reaches 0.75 (between the first threshold of 0.7 and the second threshold of 0.9), and the current window data is offloaded to the roadside node to request high-precision map matching model correction.

[0046] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for joint repair of IoT edge data quality, characterized in that, include: The data acquisition step involves acquiring the raw data stream collected from the edge terminal, preprocessing the raw data stream, and obtaining a standardized data sequence X. The joint anomaly perception and collaborative repair step involves inputting the standardized data sequence X into the joint anomaly perception and collaborative repair module. This module employs a multi-task learning neural network with a shared encoder, simultaneously outputting anomaly classification probability vector P and collaborative repair value y_t. The anomaly classification probability vector P includes missing probability, drift probability, spike probability, and duplication probability. When multiple anomaly probabilities reach or exceed preset thresholds, the calculation of the collaborative repair value y_t simultaneously minimizes the weighted sum of loss functions corresponding to multiple anomaly types. In the online adaptive step, the standardized data sequence X is input into the zero-overhead online adaptive engine. The zero-overhead online adaptive engine calculates the variance change rate Δσ² based on the sliding window statistics of the data stream and generates a dynamic adjustment factor α through a mapping function. The dynamic adjustment factor α is used to adjust the anomaly detection threshold of the joint anomaly perception and collaborative repair module. The complexity assessment and decision-making steps involve inputting the standardized data sequence X into a data flow complexity estimator, which calculates the complexity index of the current data window and generates a confidence score C; different repair paths are executed based on the confidence score C: when C is lower than a first threshold, the collaborative repair value y_t is used as the final repair result; When C is between the first and second thresholds, the data is offloaded to the cloud for repair; when C is above the second threshold, the current data is marked as untrusted and an alarm is triggered. The output integration steps integrate and output the repair results in chronological order.

2. The method according to claim 1, characterized in that, The shared encoder consists of convolutional layers and gated recurrent unit layers; the joint anomaly perception and collaborative repair module further includes an anomaly classification head and a collaborative repair head, wherein the anomaly classification head is a fully connected layer with a Softmax activation function, and the collaborative repair head is a fully connected layer with a linear activation function.

3. The method according to claim 1, characterized in that, In the weighted sum of the loss function, for the case where both drift and spike are detected simultaneously, the loss function is L = λ_1·L_smooth(y_t) + λ_2·L_spike(y_t), where L_smooth is the second-order difference smoothing loss, L_spike is the median-based suppression loss, and λ_1 and λ_2 are preset weight coefficients; the co-repair value y_t is iteratively solved during inference using the gradient descent method.

4. The method according to claim 1, characterized in that, The zero-overhead online adaptive engine maintains two sliding windows of length N, and calculates the mean μ_t and variance σ_t² of the current window; the variance change rate Δσ² = (σ_t² − σ_{tN}²) / (σ_{tN}² + ε), where ε is a very small positive number to prevent division by zero; the mapping function is α = 1.0 + 0.5·tanh(β·Δσ²), where β is the sensitivity coefficient; the dynamic adjustment factor α is restricted to a preset interval.

5. The method according to claim 1, characterized in that, The complexity index includes energy spectral entropy and sample entropy; the confidence score C = w_1·E_psd_normalized + w_2·SampEn_normalized, where E_psd_normalized and SampEn_normalized are normalized values, and w_1 and w_2 are preset weight coefficients.

6. The method according to claim 1, characterized in that, In the complexity assessment and decision-making step, when C is between the first threshold and the second threshold, only the predetermined number of data points with the highest abnormal probability in the current data window and their predetermined number of neighboring points before and after it are unloaded. If no response is received from the cloud within the preset timeout period, the collaborative repair value y_t will be used as the repair result.

7. A joint data quality repair system for IoT edge devices, characterized in that, include: An edge processing subsystem, deployed on IoT edge devices, includes: The data stream preprocessing module is used to preprocess the raw data stream and output a standardized data sequence X; The joint anomaly perception and collaborative repair module employs a multi-task learning neural network with a shared encoder to synchronously output the anomaly classification probability vector P and the collaborative repair value y_t. A zero-overhead online adaptive engine is used to calculate the variance change rate based on the sliding window statistics of the data stream, and to generate a dynamic adjustment factor α through a mapping function to adjust the anomaly detection threshold of the joint anomaly perception and collaborative repair module. A data flow complexity estimator is used to calculate the complexity index of the current data window and generate a confidence score C; A task unloading decision-maker is used to determine the repair path based on the confidence score C; Output integrator, used to integrate and output the repair results in chronological order; The cloud-based enhancement subsystem, deployed on a cloud server, is used to receive and repair data offloaded from the edge, as well as to periodically update the edge model.

8. The system according to claim 7, characterized in that, The model size of the joint anomaly perception and collaborative repair module does not exceed 150KB, and the single inference latency is less than 0.5ms; the data flow complexity estimator takes less than 100μs to calculate the confidence score C in a single calculation.

9. The system according to claim 7, characterized in that, The cloud-based enhancement subsystem includes a high-precision deep repair model and a model update and distillation module. The model update and distillation module is used to periodically fine-tune the high-precision model using historical data and compress knowledge into a lightweight version through knowledge distillation and distribute it to the edge.