An industrial internet-based data fusion method and processing system
Through streaming segmented timing alignment and the tensor decomposition model under the federated learning framework, the timing misalignment problem of high-concurrency industrial equipment data streams is solved, millisecond-level alignment and noise suppression are achieved, meeting the real-time control requirements of flexible production lines.
Patent Information
- Application Number
- CN202511014003.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-07-23
AI Technical Summary
In high-concurrency scenarios, the data streams of industrial equipment suffer from timing misalignment due to sampling rate differences and network jitter. The existing dynamic time warping algorithm has high computational latency and cannot adapt to sudden fluctuations, making it impossible to achieve millisecond-level dynamic adaptive fusion.
A streaming segmented timing alignment method is adopted to generate a dynamic time offset by calculating the local mutual information between the latest data segment and the historical data in the buffer. Resampling alignment is performed based on the device sampling rate, communication protocol, and type weight. Data fusion is performed using the tensor decomposition model under the federated learning framework, and a noise detection module is deployed to suppress impulse noise.
It achieves millisecond-level alignment of hundreds of data streams, adaptively adjusts compensation strength, avoids alignment distortion, and accurately marks pulse noise points, meeting the real-time control needs of flexible production lines and reducing computing resource consumption and manual maintenance costs.
Smart Images

Figure CN120524437B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of industrial internet and edge intelligence, and in particular to a data fusion method and processing system based on industrial internet. BACKGROUND
[0002] In discrete manufacturing lines, such as automobile welding and 3C electronic assembly, tens to hundreds of industrial devices concurrently upload time series data at millisecond level frequency; the sampling period, communication protocol and network jitter of different devices result in time offset of multi-source data streams; in such scenarios, device state monitoring needs to fuse multi-dimensional data such as vibration and current, but millisecond-level misalignment can mask the correlation of key events, for example, the phase misalignment of mechanical arm motion trajectory and welding gun current may misjudge the virtual welding fault.
[0003] Current solutions mostly use dynamic time warping (DTW) optimization algorithm combined with sliding window mechanism, the specific process includes: segmenting the data stream with a fixed time length, such as 200ms, and performing improved DTW, such as FastDTW, to align multi-device data in the window, and finally inputting the aligned segments into the fusion model; such solutions can alleviate global offset, but the complexity of FastDTW is still O(n), and the time-consuming of 100-way data stream alignment is more than 50ms, so the fixed window cannot adapt to sudden network fluctuations.
[0004] When the number of devices increases sharply or the network suddenly fluctuates in high-concurrency scenarios, the window truncation of existing solutions causes cross-window event breakage, and the alignment algorithm relies on a pre-set similarity function, which cannot adapt to sudden changes in device sampling rate, and the contradiction between real-time requirements and computing resources is intensified, such problems are particularly prominent in flexible production line dynamic reconfiguration, and the existing method has not realized millisecond-level dynamic adaptive fusion. SUMMARY
[0005] In view of the above existing problems, the present application is proposed.
[0006] The present application provides a data fusion method and processing system based on industrial internet to solve the problem of time offset of high-concurrency device data streams due to sampling rate difference and network jitter, and the high computational delay of existing dynamic time warping algorithm and the inability to adapt to sudden fluctuations.
[0007] To solve the above technical problems, the present application provides the following technical solutions:
[0008] In a first aspect, the present application provides a data fusion method based on industrial internet, which includes: step S1, receiving concurrent data streams from multiple industrial devices, the data streams containing timestamps and device identifiers;
[0009] Step S2, buffering the data streams to independent buffer areas according to device identifiers;
[0010] Step S3, performing stream segmentation time alignment on the data of each buffer:
[0011] a) calculating the local mutual information of the latest data segment and the historical data of the buffer;
[0012] b) generating a dynamic time offset according to the mutual information value;
[0013] c) resampling and aligning the data segment based on the time offset;
[0014] Step S4, inputting the aligned multi-device data into a fusion model to generate joint features.
[0015] As a preferred scheme of the data fusion method based on the industrial internet, the dynamic time offset of step S3 satisfies:
[0016] The offset calculation formula is: offset = reference offset x (1-mutual information value), wherein the reference offset is a preset maximum tolerance offset value.
[0017] As a preferred scheme of the data fusion method based on the industrial internet, in step S3, the dynamic time offset calculation in the stream segmentation time alignment includes: in the sliding window , the latest data segment is regarded as a random variable , the corresponding historical segment is regarded as , the normalized discrete mutual information is calculated, and the formula is:
[0018] ,
[0019] Wherein, represents the mutual information of the device at time , the device index is , the time index is , the number of histogram bins is , the joint probability of the device falling into the th bin is , and the respective edge probabilities are ,
[0020] For the device sampling rate and the protocol correction factor , the reference offset coefficient is given as follows:
[0021] ,
[0022] Wherein, is the reference offset coefficient of the device , Profinet takes protocol modifier, other protocols take , the sampling rate of the device ;
[0023] According to the inherent delay characteristics of different functional units, type weights are introduced:
[0024] ,
[0025] wherein, is the device type weight, the value is determined according to empirical test;
[0026] Then, the dynamic time offset and the skip threshold are determined, and the final time compensation amount is positioned as:
[0027] ,
[0028] wherein, is the time compensation required to be applied to the latest segment ;
[0029] If , , the data segment is directly transmitted, otherwise the sparse resampling window is entered; in the formula, is the skip interpolation threshold.
[0030] As a preferred scheme of the data fusion method based on the industrial internet, in the step S3, the resampling alignment adopts the cubic spline interpolation method, and only the data in the offset time window is interpolated.
[0031] As a preferred scheme of the data fusion method based on the industrial internet, in the step S4, the fusion model is a tensor decomposition model under the federal learning framework, and a noise detection module is arranged at the input end of the model:
[0032] The noise detection module identifies the impulse noise point through gradient mutation detection;
[0033] The noise point is input into the tensor decomposition model after a weight penalty factor is applied thereto.
[0034] As a preferred scheme of the data fusion method based on the industrial internet, in the step S4, in the process of inputting the noise point into the tensor decomposition model after a weight penalty factor is applied thereto, the aligned data sequence of each device is sequentially executed in the following flow, is the number of aligned sample points of the device in the current batch, and is unitless:
[0035] with a uniform sampling period Compute the gradient magnitude of adjacent samples:
[0036] ,
[0037] where, denotes the device gradient magnitude at index , is the sample index, is the th sample value, is the previous sample value, is the aligned uniform sampling period;
[0038] Robust statistics are then performed on to obtain:
[0039] , ;
[0040] where, is the gradient median, is the absolute deviation median of the gradient;
[0041] The adaptive threshold is set as:
[0042] ,
[0043] where, is the gradient mutation threshold, is the threshold amplification coefficient;
[0044] Define the noise indicator function:
[0045] ,
[0046] where, is the gradient mutation marker, 1 indicates a pulse noise point;
[0047] Apply a weight penalty to the detected noise points:
[0048] , ,
[0049] where, is the sample weight, is the penalty coefficient, is the weighted sample value, and is input to the tensor decomposition model.
[0050] In a second aspect, the present application provides an industrial internet-based data processing system, comprising,
[0051] The device access layer is configured to connect to multi-protocol industrial devices and parse raw data streams;
[0052] Stream processing engine, including:
[0053] Buffer management module, which isolates and stores data according to device identification;
[0054] Data alignment module, which performs streaming segment timing alignment operations;
[0055] Integrate the computing layer and deploy the federated learning framework and noise suppression tensor decomposition model.
[0056] As a preferred solution of the data processing system based on the industrial Internet described in the present invention, the data alignment module includes:
[0057] Mutual information calculation unit, which calculates the correlation of data segments in real time;
[0058] A dynamic offset compensation unit outputs a time offset instruction based on mutual information;
[0059] Sparse resampling unit, performs interpolation alignment according to offset instructions.
[0060] As a preferred solution of the data processing system based on the industrial Internet described in the present invention, the tensor decomposition model of the fusion computing layer includes:
[0061] Noise marking unit, which locates impulse noise through real-time gradient analysis;
[0062] The weighted decomposition unit assigns preset weight coefficients to noise points.
[0063] As a preferred solution of the data processing system based on the Industrial Internet described in the present invention, the device access layer supports protocols including OPC UA, Modbus, and Profinet, and performs semantic conflict resolution during protocol conversion:
[0064] Calculate KL divergence based on the distribution of historical data of the device;
[0065] When the KL divergence is greater than the preset threshold, the fusion weight is dynamically allocated according to the device confidence;
[0066] The confidence allocation in semantic conflict resolution satisfies:
[0067] Device confidence = 1 - the false alarm rate of the device in the past 24 hours. The fusion weight is normalized and allocated according to the confidence.
[0068] The beneficial effects of the present invention are as follows: the present invention replaces global calculation with streaming segmentation processing, only performs mutual information analysis and resampling on the latest data segment, compresses the time consumption of aligning hundreds of data streams to milliseconds, and meets the real-time control requirements of flexible production lines; dynamic offset generation combines the device sampling rate, communication protocol and type weight, adaptively adjusts the compensation strength, and avoids alignment distortion caused by device heterogeneity; introduces a noise perception mechanism to use the gradient median and absolute deviation to construct an adaptive threshold, accurately marks pulse noise points and applies weight penalties, while suppressing interference, retains the integrity of fault features, and prevents signal smoothing caused by traditional filtering; tensor decomposition under the federated learning framework supports distributed data privacy protection, and the noise suppression module is designed in advance to avoid contaminating the core feature extraction process.
[0069] The dynamic resolution of semantic conflicts in the present invention is based on KL divergence to quantify protocol differences, and combines the historical false alarm rate of the equipment to allocate fusion weights, eliminating implicit conflicts between protocols such as OPC UA and Modbus, and reducing manual maintenance costs. In addition, buffer isolation management and sparse resampling are used to reduce redundant calculations, and the stream processing engine is lightweight and deployed on edge nodes to alleviate cloud load pressure. Interpolation thresholds are skipped to avoid invalid operations under micro-shifts, thereby improving resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0071] Figure 1 This is a flow chart of the data fusion method based on the industrial Internet in Example 1.
[0072] Figure 2 This is a schematic diagram of the framework of the data processing system based on the industrial Internet in Example 1. DETAILED DESCRIPTION
[0073] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0074] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0075] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0076] Example 1, reference Figure 1 and Figure 2 , this embodiment provides a data fusion method based on the industrial Internet, comprising the following steps:
[0077] Step S1, receiving concurrent data streams from multiple industrial devices, the data streams including timestamps and device identifiers;
[0078] Step S2, caching the data stream into an independent buffer according to the device identification;
[0079] Step S3: Perform streaming segment timing alignment on the data in each buffer:
[0080] a) Calculate the local mutual information between the latest data segment and the historical data in the buffer;
[0081] b) generating a dynamic time offset according to the mutual information value;
[0082] c) resampling and aligning the data segments based on the time offset;
[0083] The dynamic time offset generation in step S3 satisfies:
[0084] The offset calculation formula is: offset = reference offset × (1-mutual information value), where the reference offset is the preset maximum tolerance offset value;
[0085] The rules for determining the base offset value include:
[0086] For devices with a sampling rate ≥ 50 Hz, the reference offset = 20 ms;
[0087] For devices with sampling rates < 50 Hz, the baseline offset = 50 ms;
[0088] If the device communication protocol is Profinet, the baseline offset value is reduced by 40%;
[0089] In step S3, the calculation of the dynamic time offset in the streaming segment timing alignment includes:
[0090] In the sliding window Within, the latest data segment is considered as a random variable , the corresponding historical segment is regarded as , calculate the normalized discrete mutual information, the formula is:
[0091] ,
[0092] in, Representation device At the moment The mutual information of For device index, is the time index, is the number of histogram bins, for Fall into The joint probability of the bins, 、 are their respective marginal probabilities;
[0093] For device sampling rate and protocol correction factors , giving the reference offset coefficient as follows:
[0094] ,
[0095] in, For equipment The reference offset coefficient, is the protocol correction factor, Profinet takes , other protocols take For equipment Sampling rate;
[0096] According to the inherent delay characteristics of different functional units, type weights are introduced:
[0097] ,
[0098] in, is the equipment type weight, the value is determined based on empirical tests;
[0099] Then the dynamic time offset and skip threshold are determined, and the final time compensation is located as follows:
[0100] ,
[0101] in, To be applied to the latest segment Time compensation;
[0102] like , , then directly pass the data segment, otherwise enter the sparse resampling window; where , is the skip interpolation threshold.
[0103] Specifically, mutual information measures the correlation between the latest segment and historical segments using binned probabilities, avoiding any assumptions about signal amplitude and thus adapting to both analog and discrete scenarios. The reference offset coefficient introduces sampling rate and protocol factors to unify the measurement of physical layer jitter and network layer jitter. The type weight further distinguishes actuator hysteresis from controller feedback differences, making the compensation strategy robust to scenario changes.
[0104] The resampling alignment in step S3 uses cubic spline interpolation, and only interpolates the data within the offset time window;
[0105] Step S4, inputting the aligned multi-device data into the fusion model to generate joint features;
[0106] The fusion model in step S4 is a tensor decomposition model under the federated learning framework, and a noise detection module is deployed at the model input:
[0107] The noise detection module identifies impulse noise points through gradient mutation detection;
[0108] Apply weight penalty factors to noise points and input them into tensor decomposition model;
[0109] In step S4, when the noise points are input into the tensor decomposition model after applying the weight penalty factor, in order to suppress the impulse noise before tensor decomposition, each device Aligned data sequence Follow the following steps in sequence: For equipment The number of aligned sample points in the current batch, unitless:
[0110] With a unified sampling period Calculate the gradient magnitude of adjacent samples:
[0111] ,
[0112] in, Representation device In the index The gradient amplitude at , is the sample index, For the Sample values, is the previous sample value, is the unified sampling period after alignment;
[0113] Then to Performing robust statistics, we get:
[0114] , ;
[0115] in, is the median gradient, is the median absolute deviation of the gradient;
[0116] The adaptive threshold is set as:
[0117] ,
[0118] in, is the gradient mutation threshold, is the threshold amplification factor;
[0119] Define the noise indicator function:
[0120] ,
[0121] in, is the gradient mutation mark, 1 represents the impulse noise point;
[0122] Apply weight penalty to detected noise points:
[0123] , ,
[0124] in, is the sample weight, is the penalty coefficient, is the weighted sample value, As tensor decomposition model input.
[0125] Specifically, gradient mutation detection uses the median and the median of absolute deviation to construct a threshold, without making any prior assumptions about the signal distribution, and remains robust to industrial process data containing spike interference; the threshold amplification factor and the penalty coefficient are designed separately, and the noise suppression strength can be independently adjusted without recalibrating the gradient statistics; weighted samples avoid destroying the time series integrity through dense noise suppression rather than hard deletion.
[0126] This embodiment also provides a data processing system based on the Industrial Internet, including:
[0127] The device access layer is configured to connect to multi-protocol industrial devices and parse raw data streams;
[0128] Stream processing engine, including:
[0129] Buffer management module, which isolates and stores data according to device identification;
[0130] Data alignment module, which performs streaming segment timing alignment operations;
[0131] The data alignment module includes:
[0132] Mutual information calculation unit, which calculates the correlation of data segments in real time;
[0133] A dynamic offset compensation unit outputs a time offset instruction based on mutual information;
[0134] Sparse resampling unit, performs interpolation alignment according to offset instructions;
[0135] Integrate the computing layer and deploy the federated learning framework and noise suppression tensor decomposition model;
[0136] The tensor decomposition model of the fusion computing layer includes:
[0137] Noise marking unit, which locates impulse noise through real-time gradient analysis;
[0138] A weighted decomposition unit assigns a preset weight coefficient of less than 0.1 to noise points;
[0139] The device access layer supports protocols including OPC UA, Modbus, and Profinet, and performs semantic conflict resolution during protocol conversion:
[0140] Calculate KL divergence based on the distribution of historical data of the device;
[0141] When the KL divergence is greater than the preset threshold, the fusion weight is dynamically allocated according to the device confidence;
[0142] The confidence allocation in semantic conflict resolution satisfies:
[0143] Device confidence = 1 - the device's false alarm rate in the past 24 hours. The fusion weight is normalized and allocated according to the confidence level.
[0144] The above threshold parameters (e.g., 20 ms, 50 ms, 0.6, 1.0, 2 ms, etc.) are calibrated based on large-sample experimental statistics and industry experience on typical discrete manufacturing production lines, and can be equivalently adjusted by those skilled in the art according to actual scenarios through experiments or automated parameter adjustment methods (e.g., grid search, Bayesian optimization). The present invention is not limited to specific numerical values.
[0145] This embodiment addresses the problem of data stream timing misalignment caused by the concurrent uploading of data at millisecond frequencies by dozens to hundreds of industrial devices in discrete manufacturing production lines. A dynamic adaptive data fusion method and system based on the Industrial Internet is proposed. This method first isolates cached data by device identifier, calculates the local mutual information between the latest and historical segments within a sliding window, and combines the sampling rate, protocol correction factor, and device type weight to generate a time compensation that is inversely proportional to the correlation. When the compensation amount exceeds the skip threshold, cubic spline interpolation is used only for samples within the offset window, avoiding the O(n) complexity brought by global DTW.
[0146] The aligned multi-source data is fed into a tensor decomposition model within the federated learning framework. A gradient-MAD-based impulse noise detection module is deployed at the model entry point, assigning low-weight penalties to outliers before participating in feature decomposition, thereby balancing privacy, security, and robustness. The supporting system comprises a device access layer, a stream processing engine, and a fusion computing layer, encompassing functional units such as protocol parsing, mutual information calculation, dynamic compensation, sparse interpolation, and weighted tensor decomposition. Key threshold parameters are derived from large-sample statistics and can be adjusted through automated parameter tuning.
[0147] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A data fusion method based on industrial Internet, characterized in that: The following steps are involved: Step S1, receiving concurrent data streams from multiple industrial devices, wherein the data streams include timestamps and device identifiers; Step S2, caching the data stream into an independent buffer according to the device identification; Step S3: Perform streaming segment timing alignment on the data in each buffer: a) Calculate the local mutual information between the latest data segment and the historical data in the buffer; b) generating a dynamic time offset according to the mutual information value; c) resampling and aligning the data segments based on the time offset; Step S4, inputting the aligned multi-device data into the fusion model to generate joint features; The dynamic time offset is defined as: , in, It is a dynamic time offset, which means it needs to be applied to the latest segment. Time compensation, is the device type weight, For devices The reference offset coefficient is based on the device Sampling rate and protocol correction factors Calculated, Representation device At the moment mutual information.
2. The data fusion method based on the industrial Internet according to claim 1, characterized in that: In step S3, the calculation of dynamic time offset in the streaming segment timing alignment includes: Within, the latest data segment is considered as a random variable , the corresponding historical segment is regarded as , calculate the normalized discrete mutual information, the formula is: , in, Representation device At the moment The mutual information of For device index, is the time index, is the number of histogram bins, for Fall into The joint probability of the bins, 、 are their respective marginal probabilities; For device sampling rate and protocol correction factors , giving the reference offset coefficient as follows: , in, For devices The reference offset coefficient, is the protocol correction factor, Profinet takes , other protocols take For devices Sampling rate; According to the inherent delay characteristics of different functional units, type weights are introduced: , in, is the equipment type weight, the value is determined based on empirical tests; Then make dynamic time offset and skip threshold determination, like , , then directly pass the data segment, otherwise enter the sparse resampling window; where , is the skip interpolation threshold.
3. The data fusion method based on the industrial Internet according to claim 1, characterized in that: The resampling alignment in step S3 uses the cubic spline interpolation method, and only interpolates the data within the offset time window.
4. The data fusion method based on the industrial Internet according to claim 1, characterized in that: The fusion model in step S4 is a tensor decomposition model under the federated learning framework, and a noise detection module is deployed at the model input: The noise detection module identifies impulse noise points by detecting gradient mutations; The noise points are subjected to weight penalty factors and then input into the tensor decomposition model.
5. The data fusion method based on the industrial Internet according to claim 4, characterized in that: In step S4, when the noise points are input into the tensor decomposition model after applying the weight penalty factor, each device Aligned data sequence Follow the following steps in sequence: For devices The number of aligned sample points in the current batch, unitless: With a unified sampling period Calculate the gradient magnitude of adjacent samples: , in, Representation device In the index The gradient amplitude at , is the sample index, For the Sample values, is the previous sample value, is the unified sampling period after alignment; Then to Performing robust statistics, we get: , ; in, is the median gradient, is the median absolute deviation of the gradient; The adaptive threshold is set as: , in, is the gradient mutation threshold, is the threshold amplification factor; Define the noise indicator function: , in, is the gradient mutation mark, 1 represents the impulse noise point; Apply weight penalty to detected noise points: , , in, is the sample weight, is the penalty coefficient, is the weighted sample value, As tensor decomposition model input.
6. A data processing system based on the industrial Internet, based on the data fusion method based on the industrial Internet according to any one of claims 1 to 5, characterized in that: include: The device access layer is configured to connect to multi-protocol industrial devices and parse raw data streams; Stream processing engine, including: Buffer management module, which isolates and stores data according to device identification; Data alignment module, which performs streaming segment timing alignment operations; Integrate the computing layer and deploy the federated learning framework and noise suppression tensor decomposition model.
7. The data processing system based on the Industrial Internet according to claim 6, characterized in that: The data alignment module includes: Mutual information calculation unit, which calculates the correlation of data segments in real time; A dynamic offset compensation unit outputs a time offset instruction based on mutual information; Sparse resampling unit, performs interpolation alignment according to offset instructions.
8. The data processing system based on the Industrial Internet according to claim 6, characterized in that: The tensor decomposition model of the fusion computing layer includes: Noise marking unit, which locates impulse noise through real-time gradient analysis; The weighted decomposition unit assigns preset weight coefficients to noise points.
9. The data processing system based on the Industrial Internet according to claim 6, characterized in that: The device access layer supports protocols including OPC UA, Modbus, and Profinet, and performs semantic conflict resolution during protocol conversion: Calculate KL divergence based on the distribution of historical data of the device; When the KL divergence is greater than the preset threshold, the fusion weight is dynamically allocated according to the device confidence; The confidence allocation in semantic conflict resolution satisfies: Device confidence = 1 - the false alarm rate of the device in the past 24 hours. The fusion weight is normalized and allocated according to the confidence.
Citation Information
Patent Citations
Data center intelligent operation and maintenance method and system based on industrial Internet of Things
CN120110939A
Multi-source heterogeneous data-based industry return on investment real-time acquisition method
CN120295993A