Cross-platform multi-source data integration analysis method and system based on industrial internet

Through technologies such as adaptive data adapter clusters and knowledge graphs, the compatibility, security and accuracy issues of multi-source data integration and analysis in the Industrial Internet have been solved, efficient data collection, integration and analysis have been achieved, and the level of intelligence in industrial production has been improved.

CN120596561AInactive Publication Date: 2025-09-05NANTONG WUXI INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510769438.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing industrial Internet data integration and analysis has problems such as inconsistent data formats, large differences in communication protocols, uneven data quality, insufficient security, and poor scalability and compatibility, which make it difficult to effectively integrate and analyze data, affecting industrial production efficiency and decision-making accuracy.

Method used

It adopts adaptive data adapter clusters, knowledge graph-based data quality assessment networks, edge computing distributed middleware architecture, federated blockchain technology, hybrid storage architecture, reinforcement learning-driven intelligent analysis framework, and in-depth defense system to achieve efficient collection, integration, analysis, and security protection of multi-source data.

Benefits of technology

The system's compatibility and scalability have been improved, enabling it to quickly adapt to different industrial scenarios and changes in data sources, ensuring data security and accuracy, providing a reliable data foundation and accurate analysis results, and improving the level of intelligence in industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596561A_ABST
    Figure CN120596561A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial internet data processing, in particular to a cross-platform multi-source data integration analysis method and system based on the industrial internet, and the system comprises a data collection layer, a data preprocessing layer, a data integration layer, a data storage layer, an intelligent analysis framework, an application layer and a data security layer. Through the data adapter and middleware technology, efficient collection and integration of multi-source heterogeneous data are achieved, multiple data sources and communication protocols are supported, the compatibility and expansibility of the system are improved, and the system can rapidly adapt to different industrial scenes and data source changes; according to the method, advanced big data and artificial intelligence technologies are applied, a specific analysis model for industrial production is combined, deep analysis can be performed on the data, potential values in the data are mined, an accurate analysis result is provided for equipment fault prediction and production process optimization, and the intelligent level of industrial production is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial Internet data processing technology, and specifically to a cross-platform multi-source data integration and analysis method and system based on the industrial Internet. Background Art

[0002] In the era of industrial Internet, a large amount of multi-source heterogeneous data is generated in the industrial production process. These data come from different devices, systems and platforms, such as sensors, industrial robots, and business management systems.

[0003] In the prior art, in an intelligent multi-source heterogeneous data analysis platform with application number 202010653255.3, the application includes a sales information data collection module, a cost information data collection module, a data information cache module, a risk information data collection module, a data information screening module, a data information analysis module, a platform information display module, and a risk assessment module. The output end of the sales information data collection module is electrically connected to the input end of the data information cache module. The sales information data is collected through the sales information data collection module, and the cost information is collected through the cost information data collection module. Then, according to the amount of information in the risk information data collection module and the coordination between the data screening module and the data information analysis module, the net profit is calculated, and at the same time, the risk level of the enterprise is evaluated.

[0004] Existing industrial internet data integration and analysis suffers from numerous issues. Inconsistent data formats and widely varying communication protocols hinder direct data interaction and integration. Data quality varies widely, containing significant amounts of noise, missing values, and outliers, impacting the accuracy of analytical results. Furthermore, insufficient data security presents the risk of data leakage and tampering. Furthermore, existing systems suffer from poor scalability and compatibility, making them difficult to adapt to evolving industrial scenarios and data sources. These issues hinder effective data integration and analysis, severely impacting industrial production efficiency and decision-making accuracy. Summary of the Invention

[0005] The purpose of the present invention is to provide a cross-platform multi-source data integration and analysis method and system based on the Industrial Internet to solve the problems raised in the above background technology.

[0006] To achieve the above object, the present invention provides the following technical solutions: A cross-platform, multi-source data integration and analysis system based on the Industrial Internet, including: Data collection layer: deploys a cluster of adaptive data adapters, each with a built-in dynamic protocol parsing engine and intelligent format conversion module; Data preprocessing layer: Build a data quality assessment network based on the knowledge graph, and evaluate from five dimensions: completeness, accuracy, consistency, timeliness, and relevance; Data integration layer: adopts a distributed middleware architecture based on edge computing and combines it with federated blockchain technology; Data storage layer: Build a hybrid storage architecture, including a real-time time series database based on LSM trees, a distributed data warehouse based on column storage, and an unstructured data lake based on object storage; Data analysis layer: Integrates a reinforcement learning-driven intelligent analysis framework, including an industry-specific model library and a dynamic optimization engine; Application layer: Provides a low-code application development platform that supports the rapid construction of data visualization, intelligent decision-making, equipment health management, and production collaborative scheduling applications. The data visualization module supports 3D digital twin visualization based on WebGL. The intelligent decision-making module integrates explainable AI technology and uses the SHAP value method to visually explain decision results. It also opens a full-lifecycle API interface that complies with the OpenAPI standard. Data security layer: Build a defense-in-depth system, use homomorphic encryption technology at the data collection end to achieve data encryption collection, use quantum key distribution QKD and national secret SM4 hybrid encryption in the transmission process, use attribute encryption ABE-based access control in the storage stage, and introduce secure multi-party computing MPC technology in the processing link to ensure data privacy and security.

[0007] Preferably, the dynamic protocol parsing engine is based on NLP semantic analysis technology, supports adaptive parsing of Modbus, OPCUA, MQTT standard protocols and custom protocols, and the intelligent format conversion module adopts a multimodal mapping algorithm to realize the conversion of multi-source data to JSON-LD semantic intermediate format, and supports three modes of real-time stream acquisition, batch timing acquisition and event-triggered acquisition; The adaptive data adapter cluster supports hot-plug expansion. Through the dynamic service discovery mechanism and service orchestration engine, it can automatically identify and adapt to newly connected data sources and realize online updating of protocol parsing rules and format conversion templates.

[0008] Preferably, the data preprocessing layer integrates an intelligent cleaning module, a semantic standardization module and a causal reasoning and error correction module. The causal reasoning and error correction module uses a Bayesian network model, combined with a knowledge graph of the equipment operation mechanism, to achieve intelligent correction of abnormal data through causal association analysis. The knowledge graph construction module of the data preprocessing layer converts the metadata of multi-source heterogeneous data into a unified semantic model through ontology mapping and semantic alignment technology, thereby realizing semantic association analysis of data quality assessment indicators.

[0009] Preferably, the middleware supports cross-regional and cross-platform data collaborative interaction. The federated blockchain nodes are composed of a consortium chain network consisting of data providers, analysts, and auditors. Each data integration operation generates an encrypted blockchain transaction containing a data source hash fingerprint, a zero-knowledge proof of the processing process, and a timestamp. Automated auditing and traceability of the entire data integration process are achieved through smart contracts. The federal blockchain of the data integration layer adopts a practical Byzantine fault-tolerant consensus algorithm combined with threshold signature technology to ensure the efficiency of data integration operations while ensuring that data cannot be tampered with and operations are traceable.

[0010] Preferably, the data storage layer adopts an intelligent storage strategy engine to dynamically schedule storage media according to the access frequency, timeliness and structured degree of the data, and is equipped with a data redundancy backup and rapid recovery mechanism based on erasure coding technology; The data analysis layer's industry-specific model library includes a Transformer-based equipment failure prediction model and a graph neural network-based production process optimization model. The dynamic optimization engine automatically adjusts model parameters and architecture based on real-time data feedback using a reinforcement learning algorithm, triggering incremental retraining when prediction errors exceed a threshold. The Transformer-based equipment fault prediction model of the data analysis layer adopts a multi-head attention mechanism to perform feature fusion and long sequence dependency modeling on multimodal time series data of equipment vibration signals, temperature curves, and current waveforms.

[0011] Preferably, the low-code development platform of the application layer has a built-in industrial application template library, which includes equipment predictive maintenance, energy management, and quality traceability scenario templates, and supports the rapid generation of customized applications through visual drag and drop and parameter configuration.

[0012] The cross-platform multi-source data integration and analysis method based on the Industrial Internet includes the following steps: Intelligent access to heterogeneous data: Through adaptive data adapter clusters, NLP semantic analysis and multimodal mapping algorithms are used to parse and convert multi-source data from PLCs, sensors, and ERP systems into a semantically-defined JSON-LD intermediate format. Dynamic expansion of new protocol parsing rules and format conversion strategies is supported through configuration files or automatic machine learning generation. Intelligent data quality governance: Build a data quality assessment network based on the knowledge graph to conduct a five-dimensional data assessment. Use interpolation, anomaly detection algorithms, and Bayesian causal reasoning to cleanse, standardize, and correct data errors. At the same time, achieve unified management of data metadata through semantic alignment. Trusted Distributed Integration: Utilizes distributed middleware based on edge computing to achieve cross-platform data collaboration, recording metadata, processing logs, and result hash values ​​of the data integration process in the form of encrypted transactions to the federated blockchain. Through the consortium chain consensus mechanism and smart contracts, automated verification and auditing of data integration operations are achieved. Intelligent storage and management: Based on data characteristics and access patterns, real-time data is stored in a time-series database, historical data and integrated data are stored in a data warehouse, and unstructured data is stored in a data lake. An intelligent storage policy engine is used to dynamically schedule storage media, and data hot and cold migration and redundant backup are performed regularly. Deep Intelligent Analysis: Based on a reinforcement learning-driven intelligent analysis framework, it uses algorithm models from an industry-specific model library to analyze data. It uses a dynamic optimization engine to monitor model performance in real time. When the prediction error exceeds the threshold, it uses incremental learning or transfer learning technology to optimize the model. Value application and sharing: Quickly build data visualization and intelligent decision-making applications through a low-code development platform; output analysis results in the form of API interfaces that comply with the OpenAPI standard, supporting data sharing and business collaboration with third-party systems; use explainable AI technology to visually explain analysis and decision-making results.

[0013] Preferably, in the step of intelligently accessing heterogeneous data, an adaptive data adapter cluster is used to implement access and conversion of multi-source heterogeneous data. The Transformer-based semantic parsing model in NLP semantic analysis technology is used in combination with a multi-head attention mechanism to perform semantic understanding of Modbus and OPCUA protocol data. For custom protocol data, an active learning algorithm is used based on a small amount of manually annotated data to optimize the loss function: in, is the true label, Automatically optimize the protocol parsing model to predict labels.

[0014] The multimodal mapping algorithm converts PLC time series data, sensor numerical data, and ERP system structured data into a JSON-LD semantic intermediate format. The algorithm uniformly maps different modal data by constructing a multimodal feature mapping matrix M. The formula is:

[0015] in, is the input multimodal data matrix, It is a converted semantic data matrix. At the same time, it supports manually adding new protocol parsing rules through configuration files, or automatically generating rules using machine learning algorithms to achieve rapid adaptation to new data sources.

[0016] Preferably, in the trusted distributed integration step, zero-knowledge proof technology is introduced to achieve legitimacy verification of data integration operations and trusted transmission of results without leaking original data.

[0017] Preferably, the anomaly detection uses the IsolationForest algorithm to calculate the anomaly score S of the data point by constructing an isolation tree: Where E(h(x)) is the path length of data point x, and c(n) is the average path length of the tree. Using Bayesian causal reasoning to correct data errors, building a Bayesian network based on the knowledge graph of the equipment operation mechanism, and using the Bayesian formula Calculate the posterior probability of the data, identify and correct abnormal data points. At the same time, through semantic alignment technology, the ontology mapping algorithm is used to transform the metadata of multi-source heterogeneous data into a unified semantic model to achieve unified management of data metadata.

[0018] Compared with the prior art, the present invention has the following beneficial effects: The present invention uses data adapter and middleware technology to achieve efficient collection and integration of multi-source heterogeneous data, supports multiple data sources and communication protocols, improves the compatibility and scalability of the system, and can quickly adapt to different industrial scenarios and data source changes. This invention uses advanced big data and artificial intelligence technologies, combined with specific analysis models for industrial production, to conduct in-depth analysis of data, explore the potential value in the data, provide accurate analysis results for equipment failure prediction and production process optimization, and improve the level of intelligence in industrial production. The present invention ensures the security of data throughout the entire processing process through multiple security technologies in the data security layer, prevents data leakage, tampering and illegal use, and improves the credibility and security of data.

[0019] The layered architecture design and standardized interface of the present invention make each layer of the system independent of each other, facilitate system maintenance and upgrade, and facilitate the access and integration of third-party application systems, thereby expanding the application scope of the system. The present invention effectively improves the accuracy and integrity of data through the quality assessment and error correction functions of the data preprocessing layer, provides a reliable data foundation for subsequent data analysis and application, and ensures the accuracy and reliability of the analysis results. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of the framework of the present invention. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. Example

[0022] See also Figure 1 , a cross-platform multi-source data integration and analysis method and system based on the Industrial Internet, including: Data collection layer: Deploys an adaptive data adapter cluster, each with a built-in dynamic protocol parsing engine and intelligent format conversion module. The dynamic protocol parsing engine, based on NLP semantic analysis technology, supports adaptive parsing of Modbus, OPCUA, MQTT standard protocols and custom protocols. The intelligent format conversion module uses a multimodal mapping algorithm to convert multi-source data into a semantic intermediate format of JSON-LD, supporting three modes: real-time stream collection, batch timed collection, and event-triggered collection. The adaptive data adapter cluster supports hot-swappable expansion. Through the dynamic service discovery mechanism and service orchestration engine, it can automatically identify and adapt to newly connected data sources, and implement online updates of protocol parsing rules and format conversion templates. Data preprocessing layer: Build a data quality assessment network based on the knowledge graph, and evaluate from five dimensions: completeness, accuracy, consistency, timeliness, and relevance; The data preprocessing layer integrates an intelligent cleaning module, a semantic standardization module, and a causal reasoning and error correction module. The causal reasoning and error correction module uses a Bayesian network model, combined with a knowledge graph of equipment operation mechanisms, to achieve intelligent correction of abnormal data through causal association analysis. The knowledge graph construction module of the data preprocessing layer transforms the metadata of multi-source heterogeneous data into a unified semantic model through ontology mapping and semantic alignment technology, thus realizing semantic correlation analysis of data quality assessment indicators. Data integration layer: This layer utilizes a distributed middleware architecture based on edge computing, combined with federated blockchain technology. The middleware supports cross-regional and cross-platform data collaboration and interaction. The federated blockchain nodes form a consortium chain network comprised of data providers, analysts, and auditors. Each data integration operation generates an encrypted blockchain transaction containing a data source hash fingerprint, zero-knowledge proof of the processing process, and a timestamp. Smart contracts enable automated auditing and traceability of the entire data integration process. The federated blockchain at the data integration layer uses a practical Byzantine fault-tolerant consensus algorithm, combined with threshold signature technology, to ensure data integrity and traceability while ensuring efficient data integration operations. Data storage layer: Build a hybrid storage architecture, including a real-time time series database based on LSM trees, a distributed data warehouse based on column storage, and an unstructured data lake based on object storage; The data storage layer uses an intelligent storage strategy engine to dynamically schedule storage media based on data access frequency, timeliness, and structuredness, and is equipped with a data redundancy backup and rapid recovery mechanism based on erasure coding technology. Data analysis layer: Integrates a reinforcement learning-driven intelligent analysis framework, including an industry-specific model library and a dynamic optimization engine; The data analysis layer's industry-specific model library includes Transformer-based equipment failure prediction models and graph neural network-based production process optimization models. The dynamic optimization engine automatically adjusts model parameters and architecture based on real-time data feedback using reinforcement learning algorithms, triggering incremental retraining when prediction errors exceed a threshold. The Transformer-based equipment fault prediction model in the data analysis layer uses a multi-head attention mechanism to perform feature fusion and long-sequence dependency modeling on multimodal time series data such as equipment vibration signals, temperature curves, and current waveforms. Application layer: Provides a low-code application development platform that supports the rapid construction of data visualization, intelligent decision-making, equipment health management, and production collaborative scheduling applications. The data visualization module supports 3D digital twin visualization based on WebGL. The intelligent decision-making module integrates explainable AI technology and uses the SHAP value method to visually explain decision results. It also opens a full-lifecycle API interface that complies with the OpenAPI standard. The low-code development platform at the application layer has a built-in industrial application template library, including scenario templates for equipment predictive maintenance, energy management, and quality traceability. It supports the rapid generation of customized applications through visual drag-and-drop and parameter configuration. Data security layer: Build a defense-in-depth system, use homomorphic encryption technology at the data collection end to achieve data encryption collection, use quantum key distribution QKD and national secret SM4 hybrid encryption in the transmission process, use attribute encryption ABE-based access control in the storage stage, and introduce secure multi-party computing MPC technology in the processing link to ensure data privacy and security. Example

[0023] The cross-platform multi-source data integration and analysis method based on the Industrial Internet includes the following steps: Intelligent access to heterogeneous data: Through adaptive data adapter clusters, NLP semantic analysis and multimodal mapping algorithms are used to parse and convert multi-source data from PLCs, sensors, and ERP systems into a semantically-defined JSON-LD intermediate format. Dynamic expansion of new protocol parsing rules and format conversion strategies is supported through configuration files or automatic machine learning generation. Intelligent access to heterogeneous data enables access and conversion of multi-source heterogeneous data through an adaptive data adapter cluster. It utilizes the Transformer-based semantic parsing model in NLP semantic analysis technology, combined with a multi-head attention mechanism, to perform semantic understanding of Modbus and OPCUA protocol data. For custom protocol data, an active learning algorithm is used based on a small amount of manually annotated data to optimize the loss function: in, is the true label, Automatically optimize the protocol parsing model to predict labels. The multimodal mapping algorithm converts PLC time series data, sensor numerical data, and ERP system structured data into a JSON-LD semantic intermediate format. The algorithm uniformly maps different modal data by constructing a multimodal feature mapping matrix M. The formula is: in, is the input multimodal data matrix, The converted semantic data matrix is ​​also supported. New protocol parsing rules can be manually added through configuration files, or automatically generated using machine learning algorithms, enabling rapid adaptation to new data sources.

[0024] Intelligent data quality governance: Build a data quality assessment network based on the knowledge graph to conduct a five-dimensional assessment of the data; use interpolation methods, anomaly detection algorithms, and Bayesian causal reasoning to clean, standardize, and correct data, while achieving unified management of data metadata through semantic alignment.

[0025] Anomaly detection uses the IsolationForest algorithm to calculate the anomaly score S of the data point by building an isolation tree: Where E(h(x)) is the path length of data point x, and c(n) is the average path length of the tree. Using Bayesian causal reasoning to correct data errors, building a Bayesian network based on the knowledge graph of the equipment operation mechanism, and using the Bayesian formula Calculate the posterior probability of the data, identify and correct abnormal data points. At the same time, through semantic alignment technology, the ontology mapping algorithm is used to transform the metadata of multi-source heterogeneous data into a unified semantic model to achieve unified management of data metadata.

[0026] Intelligent data quality governance builds a data quality assessment network based on the knowledge graph, evaluates data from five dimensions: completeness, accuracy, consistency, timeliness, and relevance, and uses the Dempster-Shafer evidence theory to integrate the results of multiple quality assessment indicators to calculate the comprehensive data quality score. in, is the weight of the j-th indicator, is the evaluation evidence of the j-th indicator. In terms of data cleaning, the LSTM interpolation method based on time series is used for missing values. The LSTM network is used to learn the long-term dependencies of time series data and predict missing values. The forward propagation formula of the LSTM network is: in, 、 、 They are input gate, forget gate, and output gate respectively. is the cell state, used to store long-term dependency information of the time series; is the hidden state at time t, which is the output of the network and contains the feature information of the current moment; σ is the activation function, which generally uses the sigmoid function to map the input to the (0, 1) interval to control the degree of opening and closing of the gate; is the input data at time t; is the hidden state at time t−1; is the cell state at time t−1; 、 、 、 、 、 、 、 、 、 、 It is the weight matrix of the network, which is used to adjust the influence of different inputs on the gate, cell state, and hidden state; 、 、 、 is a bias term used to increase the fitting ability of the model. Through these formulas, the LSTM network can learn complex patterns and long-term dependencies in time series data, thereby making more accurate predictions and interpolations of missing values.

[0027] Trusted distributed integration: Utilize distributed middleware based on edge computing to achieve cross-platform data collaboration, and record the metadata, processing logs, and result hash values ​​of the data integration process in the form of encrypted transactions to the federated blockchain; through the alliance chain consensus mechanism and smart contracts, realize automated verification and auditing of data integration operations.

[0028] Cross-platform data collaboration is achieved using distributed middleware based on edge computing. The middleware adopts a microservice architecture and realizes dynamic management of each service node through service discovery and registration mechanisms. The metadata, processing logs and result hash values ​​of the data integration process are recorded in the form of encrypted transactions to the federal blockchain. The federated blockchain uses the Practical Byzantine Fault Tolerance (PBFT) consensus algorithm, combined with threshold signature technology. In the PBFT algorithm, nodes reach consensus through message passing. The consensus process is divided into three stages: pre-preparation, preparation, and submission. Assuming that there are n nodes in the network, of which f are faulty nodes, the maximum number of faulty nodes that the algorithm can tolerate satisfies .

[0029] Each data integration operation generates a hash fingerprint H(D) of the data source and a zero-knowledge proof of the processing process. ZKP, encrypted blockchain transaction TX with timestamp T:

[0030] Among them, K is the encryption key, Encrypt is the encryption function, and the automated verification and auditing of data integration operations are achieved through smart contracts. Smart contracts verify the legitimacy of transactions based on preset rules.

[0031] Intelligent storage and management: Based on data characteristics and access patterns, real-time data is stored in a time series database, historical data and integrated data are stored in a data warehouse, and unstructured data is stored in a data lake. An intelligent storage policy engine is used to dynamically schedule storage media, and data hot and cold migration and redundant backup are performed regularly.

[0032] Build a hybrid storage architecture, including a real-time time series database based on LSM tree, a distributed data warehouse based on column storage, and an unstructured database based on object storage. Use an intelligent storage strategy engine and realize dynamic scheduling of storage media based on reinforcement learning algorithm. Define the state space S as various attributes of data (such as access frequency, data size, and timeliness), the action space A as the choice of storage medium (such as time series database, data warehouse, and data lake), and the reward function R based on the performance indicators of data storage and access (such as storage cost and query response time). The intelligent storage strategy engine maximizes the long-term cumulative reward. ; in, is the discount factor, Learn the optimal storage strategy for the reward at time t+k.

[0033] Regularly perform hot and cold data migration. For cold data, use data compression algorithms (such as Snappy and Zstandard) to reduce storage usage. For data redundancy backup, use erasure code technology, such as Reed-Solomon code, to divide the original data into k data blocks and generate n redundant blocks (n>k) through encoding. Even if some data blocks are lost, the original data can still be restored.

[0034] Deep Intelligent Analysis: Based on a reinforcement learning-driven intelligent analysis framework, it uses algorithm models from an industry-specific model library to analyze data. It uses a dynamic optimization engine to monitor model performance in real time. When the prediction error exceeds the threshold, it uses incremental learning or transfer learning technology to optimize the model. Value application and sharing: Quickly build data visualization and intelligent decision-making applications through a low-code development platform; output analysis results in the form of API interfaces that comply with the OpenAPI standard, supporting data sharing and business collaboration with third-party systems; use explainable AI technology to visually explain analysis and decision-making results.

[0035] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A cross-platform multi-source data integration and analysis system based on the Industrial Internet, characterized by: include: Data collection layer: deploys a cluster of adaptive data adapters, each with a built-in dynamic protocol parsing engine and intelligent format conversion module; Data preprocessing layer: Build a data quality assessment network based on the knowledge graph, and evaluate from five dimensions: completeness, accuracy, consistency, timeliness, and relevance; Data integration layer: adopts a distributed middleware architecture based on edge computing and combines it with federated blockchain technology; Data storage layer: Build a hybrid storage architecture, including a real-time time series database based on LSM trees, a distributed data warehouse based on column storage, and an unstructured data lake based on object storage; Data analysis layer: Integrates a reinforcement learning-driven intelligent analysis framework, including an industry-specific model library and a dynamic optimization engine; Application layer: Provides a low-code application development platform to support the rapid construction of data visualization, intelligent decision-making, equipment health management, and production collaborative scheduling applications; The data visualization module supports 3D digital twin visualization based on WebGL. The intelligent decision-making module integrates explainable AI technology, uses the SHAP value method to visually explain decision results, and opens a full lifecycle API interface that complies with the OpenAPI standard. Data security layer: Build a defense-in-depth system, use homomorphic encryption technology at the data collection end to achieve data encryption collection, use quantum key distribution QKD and national secret SM4 hybrid encryption in the transmission process, use attribute encryption ABE-based access control in the storage stage, and introduce secure multi-party computing MPC technology in the processing link to ensure data privacy and security.

2. The cross-platform multi-source data integration and analysis system based on the Industrial Internet according to claim 1 is characterized in that: The dynamic protocol parsing engine is based on NLP semantic analysis technology and supports adaptive parsing of Modbus, OPCUA, MQTT standard protocols and custom protocols. The intelligent format conversion module adopts a multimodal mapping algorithm to realize the conversion of multi-source data into a semantic intermediate format of JSON-LD, supporting three modes of real-time stream collection, batch timed collection and event-triggered collection. The adaptive data adapter cluster supports hot-plug expansion. Through the dynamic service discovery mechanism and service orchestration engine, it can automatically identify and adapt to newly connected data sources and realize online updating of protocol parsing rules and format conversion templates.

3. The cross-platform multi-source data integration and analysis system based on the Industrial Internet according to claim 1 is characterized in that: The data preprocessing layer integrates an intelligent cleaning module, a semantic standardization module, and a causal reasoning and error correction module. The causal reasoning and error correction module uses a Bayesian network model, combined with a knowledge graph of equipment operation mechanisms, to achieve intelligent correction of abnormal data through causal association analysis. The knowledge graph construction module of the data preprocessing layer converts the metadata of multi-source heterogeneous data into a unified semantic model through ontology mapping and semantic alignment technology, thereby realizing semantic association analysis of data quality assessment indicators.

4. The cross-platform multi-source data integration and analysis system based on the Industrial Internet according to claim 1 is characterized in that: The middleware supports cross-regional and cross-platform data collaboration and interaction. The federated blockchain nodes are composed of data providers, analysts, and auditors to form a consortium chain network. Each data integration operation generates an encrypted blockchain transaction containing the data source hash fingerprint, processing process zero-knowledge proof, and timestamp. Smart contracts are used to achieve automated auditing and traceability of the entire data integration process. The federal blockchain of the data integration layer adopts a practical Byzantine fault-tolerant consensus algorithm combined with threshold signature technology to ensure the efficiency of data integration operations while ensuring that data cannot be tampered with and operations are traceable.

5. The cross-platform multi-source data integration and analysis system based on the Industrial Internet according to claim 1 is characterized in that: The data storage layer uses an intelligent storage strategy engine to dynamically schedule storage media based on data access frequency, timeliness, and structuredness, and is equipped with a data redundancy backup and rapid recovery mechanism based on erasure coding technology; The data analysis layer's industry-specific model library includes a Transformer-based equipment failure prediction model and a graph neural network-based production process optimization model. The dynamic optimization engine automatically adjusts model parameters and architecture based on real-time data feedback using a reinforcement learning algorithm, triggering incremental retraining when prediction errors exceed a threshold. The Transformer-based equipment fault prediction model of the data analysis layer adopts a multi-head attention mechanism to perform feature fusion and long sequence dependency modeling on multimodal time series data of equipment vibration signals, temperature curves, and current waveforms.

6. The cross-platform multi-source data integration and analysis system based on the Industrial Internet according to claim 1 is characterized in that: The low-code development platform of the application layer has a built-in industrial application template library, which includes equipment predictive maintenance, energy management, and quality traceability scenario templates, and supports the rapid generation of customized applications through visual drag and drop and parameter configuration.

7. A cross-platform multi-source data integration and analysis method based on the Industrial Internet, characterized by: The following steps are involved: Intelligent access to heterogeneous data: Through adaptive data adapter clusters, NLP semantic analysis and multimodal mapping algorithms are used to parse and convert multi-source data from PLCs, sensors, and ERP systems into a semantically-defined JSON-LD intermediate format. Dynamic expansion of new protocol parsing rules and format conversion strategies is supported through configuration files or automatic machine learning generation. Intelligent data quality governance: Build a data quality assessment network based on the knowledge graph to conduct a five-dimensional data assessment. Use interpolation, anomaly detection algorithms, and Bayesian causal reasoning to cleanse, standardize, and correct data errors. At the same time, achieve unified management of data metadata through semantic alignment. Trusted distributed integration: Utilizes distributed middleware based on edge computing to achieve cross-platform data collaboration, and records metadata, processing logs, and result hash values ​​of the data integration process in the form of encrypted transactions to the federated blockchain; Through the alliance chain consensus mechanism and smart contracts, automatic verification and auditing of data integration operations are achieved; Intelligent storage and management: Based on data characteristics and access patterns, real-time data is stored in a time series database, historical data and integrated data are stored in a data warehouse, and unstructured data is stored in a data lake. Adopting an intelligent storage strategy engine to achieve dynamic scheduling of storage media, and regularly perform hot and cold data migration and redundant backup; Deep Intelligent Analysis: Based on a reinforcement learning-driven intelligent analysis framework, it uses algorithm models from an industry-specific model library to analyze data. Use a dynamic optimization engine to monitor model performance in real time. When the prediction error exceeds the threshold, use incremental learning or transfer learning technology to optimize the model. Value application and sharing: Rapidly build data visualization and intelligent decision-making applications through a low-code development platform; output analysis results in the form of API interfaces that comply with the OpenAPI standard, supporting data sharing and business collaboration with third-party systems; Use explainable AI technology to provide visual explanations of analysis and decision-making results.

8. The cross-platform multi-source data integration and analysis method based on the Industrial Internet according to claim 7 is characterized in that: In the step of intelligent access to heterogeneous data, an adaptive data adapter cluster is used to access and convert multi-source heterogeneous data. The Transformer-based semantic parsing model in NLP semantic analysis technology is used in combination with a multi-head attention mechanism to perform semantic understanding of Modbus and OPCUA protocol data. For custom protocol data, an active learning algorithm is used based on a small amount of manually annotated data to optimize the loss function: in, is the true label, Automatically optimize the protocol parsing model to predict labels. The multimodal mapping algorithm converts PLC time series data, sensor numerical data, and ERP system structured data into a JSON-LD semantic intermediate format. The algorithm uniformly maps different modal data by constructing a multimodal feature mapping matrix M. The formula is: in, is the input multimodal data matrix, It is a converted semantic data matrix; at the same time, it supports manually adding new protocol parsing rules through configuration files, or automatically generating rules using machine learning algorithms to achieve rapid adaptation to new data sources.

9. The cross-platform multi-source data integration and analysis method based on the Industrial Internet according to claim 7 is characterized in that: In the trusted distributed integration step, zero-knowledge proof technology is introduced to achieve the legitimacy verification of data integration operations and the trusted transmission of results without leaking the original data.

10. The cross-platform multi-source data integration and analysis method based on the Industrial Internet according to claim 7, characterized in that: The anomaly detection adopts the IsolationForest algorithm to calculate the anomaly score S of the data point by constructing an isolation tree: Where E(h(x)) is the path length of data point x, and c(n) is the average path length of the tree. Using Bayesian causal reasoning to correct data errors, building a Bayesian network based on the knowledge graph of the equipment operation mechanism, and using the Bayesian formula Calculate the posterior probability of the data, identify and correct abnormal data points; at the same time, through semantic alignment technology, use the ontology mapping algorithm to transform the metadata of multi-source heterogeneous data into a unified semantic model to achieve unified management of data metadata.

Citation Information

Patent Citations

  • Intelligence-oriented multi-source heterogeneous data analysis platform

    CN111915147A

Cited By

  • Production line visual management and control system based on MDC data

    CN121277046A

  • Multi-modal data and dynamic knowledge graph fusion method

    CN121503629A

  • Thermal power plant data processing method based on industrial Internet of Things and deep reinforcement learning

    CN121635144A

  • Ultrahigh-voltage equipment state prediction system based on big data driving

    CN121840886A

  • Industrial data analysis platform based on multi-source heterogeneous data fusion

    CN122470952A