Industrial Internet of Things Data Security Collection System Based on Blockchain

The data collection system that combines blockchain technology with the Industrial Internet of Things solves the problem of easy tampering of Industrial Internet of Things data, achieves the immutability and high credibility of data, and improves data security and the accuracy of equipment management.

CN119996012BActive Publication Date: 2025-09-05WUXI JISHU NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510185921.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-09-05
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

Existing industrial Internet of Things data collection systems are susceptible to tampering and forgery, resulting in insufficient data reliability and credibility and a lack of effective credibility assessment mechanisms.

Method used

A blockchain-based industrial Internet of Things data security collection system is adopted, including an industrial data perception layer, an edge computing layer, and a blockchain trust layer. Through real-time data collection, encrypted transmission, traceability verification, and data enhancement processing, the data is ensured to be tamper-proof and highly reliable.

Benefits of technology

It achieves data immutability and high credibility, enhances data transparency, ensures data security, and improves the accuracy of equipment monitoring and management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996012B_ABST
    Figure CN119996012B_ABST
Patent Text Reader

Abstract

The present invention discloses a blockchain-based secure data collection system for the Industrial Internet of Things (IIoT), belonging to the field of electronic data processing technology. The system includes: a collection network establishment module for connecting to the IIoT; a perception industrial data source acquisition module for real-time interaction with the industrial data perception layer; a traceability inspection result acquisition module for encrypting and transmitting the perceived industrial data source to the edge computing layer; an enhanced industrial data source acquisition module for activating the data enhancement dual channels within the edge computing layer based on the multiple traceability inspection results to enhance the perceived industrial data source; and a trusted industrial data block establishment module for inputting the enhanced industrial data source into the blockchain trusted layer. This application solves the technical problem in the prior art that IIoT data is susceptible to tampering and forgery during the collection process, resulting in insufficient data reliability and credibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic data processing, and in particular to a blockchain-based industrial Internet of Things data security collection system. Background Art

[0002] With the rapid development of the Industrial Internet of Things (IIoT), the amount of data generated by industrial equipment, sensors, and systems is growing exponentially. This data is widely used in areas such as equipment monitoring, production optimization, and intelligent decision-making. However, as data volumes increase, data security, reliability, and privacy issues have become increasingly prominent, becoming major challenges facing IIoT systems.

[0003] Currently, Industrial IoT data collection relies heavily on traditional centralized control systems, making them vulnerable to attacks and data tampering, leading to data leaks or losses. Furthermore, traditional traceability and data verification methods lack effective credibility assessment mechanisms, making it difficult to ensure data quality. Existing technologies fail to fully integrate decentralized technologies like blockchain to ensure data security and credibility, and this urgently needs improvement. Summary of the Invention

[0004] This application provides a blockchain-based industrial Internet of Things data security collection system to solve the technical problem in the existing technology that industrial Internet of Things data is easily tampered with and forged during the collection process, resulting in insufficient data reliability and credibility.

[0005] In view of the above problems, this application provides an industrial Internet of Things data security collection system based on blockchain.

[0006] The present application provides a blockchain-based industrial Internet of Things data security collection system, which includes a collection network establishment module for connecting the industrial Internet of Things and establishing an industrial data collection network, wherein the industrial data collection network includes an industrial data perception layer, an edge computing layer and a blockchain trusted layer; a perception industrial data source acquisition module for interacting with the industrial data perception layer in real time to obtain a perception industrial data source; a traceability inspection result acquisition module for encrypting and transmitting the perception industrial data source to the edge computing layer, and performing traceability inspection on the perception industrial data source by the traceability inspection channel in the edge computing layer to obtain multiple traceability inspection results; an enhanced industrial data source acquisition module for activating the data enhancement dual channel in the edge computing layer based on the multiple traceability inspection results to enhance the perception industrial data source and obtain an enhanced industrial data source; a trusted industrial data block establishment module for inputting the enhanced industrial data source into the blockchain trusted layer, and performing on-chain and off-chain secure storage of the enhanced industrial data source by the blockchain trusted layer to establish a trusted industrial data block.

[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0008] By integrating blockchain technology with the Industrial Internet of Things (IIoT) for data collection, including the collaborative work of an industrial data perception layer, an edge computing layer, and a blockchain trust layer, this solution addresses the existing technical issues of IIoT data being susceptible to tampering and forgery during the collection process, resulting in insufficient data reliability and trustworthiness. Through real-time data collection and blockchain storage, data immutability and high reliability are ensured, achieving the technical benefits of enhanced data transparency, guaranteed data security, and improved accuracy in equipment monitoring and management.

[0009] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 A schematic diagram of the structure of the blockchain-based industrial Internet of Things data security collection system provided in an embodiment of the present application.

[0011] Figure 2 This is a structural diagram of the trusted industrial data block establishment module in the blockchain-based industrial Internet of Things data security collection system provided in an embodiment of the present application.

[0012] Explanation of the accompanying drawings: collection network establishment module 11, perception industrial data source acquisition module 12, traceability inspection result acquisition module 13, enhanced industrial data source acquisition module 14, trusted industrial data block establishment module 15, multiple industrial data shard acquisition unit 151, off-chain storage block establishment unit 152, on-chain evidence block establishment unit 153, on-chain and off-chain mapping model construction unit 154, storage block alignment unit 155. DETAILED DESCRIPTION

[0013] The overall idea of ​​the technical solution provided by this application is as follows:

[0014] This embodiment of the present application provides a blockchain-based secure data collection system for the Industrial Internet of Things (IIoT). By integrating blockchain technology with the IIoT, secure data collection and trusted storage are achieved. The solution includes real-time data collection through the industrial data perception layer, preliminary data processing and verification by the edge computing layer, and encrypted data storage through the blockchain trusted layer. This not only ensures data security during collection and transmission, but also effectively prevents data tampering, improves data reliability and transparency, and meets the high security requirements of industrial equipment monitoring and management.

[0015] After introducing the basic principles of the present application, various non-limiting implementation methods of the present application will be specifically introduced in conjunction with the drawings in the specification.

[0016] Examples, such as Figure 1 As shown, the embodiment of the present application provides an industrial Internet of Things data security collection system based on blockchain, which includes:

[0017] The collection network establishment module 11 is used to connect the industrial Internet of Things and establish an industrial data collection network, wherein the industrial data collection network includes an industrial data perception layer, an edge computing layer and a blockchain trust layer.

[0018] Specifically, an industrial data collection network refers to a network system used to acquire and transmit data from industrial equipment. Within this network, data is transmitted to different layers for processing, analysis, and storage. The industrial data perception layer, the lowest layer of the IoT, is responsible for sensing and collecting data from field devices and sensors. The edge computing layer, located between the perception layer and the blockchain layer, performs preliminary data processing, analysis, and screening close to the data source, reducing the burden of data transmission and improving response speed. The blockchain trust layer utilizes blockchain technology to decentrally store and encrypt data, ensuring immutability and security, and guaranteeing long-term data reliability.

[0019] First, the industrial IoT is connected, acquiring real-time data through various sensing devices. This data is first transmitted through the industrial data sensing layer to the next layer. Next, the data transmitted by the sensing layer is sent to the edge computing layer. This layer is responsible for preliminary data processing. For example, edge computing can filter, clean, or perform simple pre-processing on temperature data, thereby reducing unnecessary data transmission. The edge computing layer can utilize local computing resources to monitor the status of industrial equipment and provide timely feedback on abnormal data, ensuring real-time response. Finally, the data processed by the edge computing layer is transmitted to the blockchain trust layer. The introduction of blockchain technology ensures the immutability and security of data. Each piece of data is packaged into a block and stored in the blockchain, forming a chain.

[0020] By building an industrial data collection network, the entire system can achieve more efficient and secure data collection, processing, and storage. The industrial data perception layer can accurately obtain real-time data, the edge computing layer can reduce the pressure of data transmission and improve system response speed, and the blockchain trust layer ensures data security, integrity, and traceability.

[0021] The perceived industrial data source obtaining module 12 is used to interact with the industrial data perception layer in real time to obtain the perceived industrial data source.

[0022] Specifically, perception industrial data sources refer to data sources collected by industrial sensors and equipment, such as real-time temperature data collected by temperature sensors and real-time pressure values ​​collected by pressure sensors.

[0023] First, various sensing devices exchange data with the perception layer in real time. These devices, such as temperature sensors, humidity sensors, and gas detectors, are connected to the system through sensor interfaces to acquire data in real time. To ensure timely and accurate data transmission, wireless or wired communication technologies are typically used for data exchange. Wireless communication is particularly suitable for large-scale coverage or complex on-site installations. For example, in large factories, LoRa wireless technology can ensure long-distance transmission, reaching every corner of the workshop.

[0024] Through real-time interaction with the industrial data perception layer, the system can accurately collect real-time data of equipment and environment, thus laying the foundation for subsequent data processing and analysis.

[0025] The traceability inspection result obtaining module 13 is used to encrypt and transmit the perceived industrial data source to the edge computing layer, and the traceability inspection channel in the edge computing layer performs traceability inspection on the perceived industrial data source to obtain multiple traceability inspection results.

[0026] Specifically, the traceability verification channel is a processing unit in the edge computing layer that traces the source and verifies the credibility of perceived industrial data. It verifies the authenticity of data by analyzing its source, transmission process, and environment.

[0027] First, the perception data obtained from the industrial data perception layer is sent to the edge computing layer via encrypted transmission. Encrypted transmission aims to protect data from external tampering or theft during transmission, ensuring data security. For example, data encryption using AES (Advanced Encryption Standard) or TLS (Transport Layer Security) ensures confidentiality during network transmission.

[0028] Next, the provenance verification channel within the edge computing layer performs provenance verification on the received encrypted data. Provenance verification traces and verifies the data's origin, ensuring it has not been tampered with and originates from a reliable device or sensor. Provenance verification determines the data's credibility by analyzing multiple dimensions of the sensor's data source, such as the sensor's device status (e.g., whether it is functioning properly), the sensor's environmental information (e.g., the temperature and humidity at the sensor's location), and the data transmission path (e.g., whether the data is transmitted through an encrypted channel).

[0029] Through encrypted transmission and traceability verification, the system ensures the security of data during transmission and effectively avoids the risk of data tampering or forgery.

[0030] The enhanced industrial data source acquisition module 14 is used to activate the data enhancement dual channels in the edge computing layer to enhance the perceived industrial data source based on the multiple traceability inspection results to obtain an enhanced industrial data source.

[0031] Specifically, the dual-channel data enhancement refers to two processing channels used in the edge computing layer. These channels enhance data quality through different processing methods, including a data cleaning channel and an anomaly correction channel. The data cleaning channel processes trusted data, removing noise, errors, or inconsistencies to improve data quality. The provenance anomaly correction channel corrects untrusted data, correcting possible errors or anomalies based on provenance information. Enhanced industrial data sources refer to perceptual industrial data sources that have undergone data enhancement processing. The processed data is more accurate and reliable, making it suitable for subsequent data analysis or storage.

[0032] First, based on the traceability verification results from the previous step, the system analyzes the credibility of each piece of data. If the data is deemed credible (i.e., traceability is reliable), it enters the data cleaning channel. If the data is deemed untrustworthy, it enters the traceability anomaly correction channel. Trustworthy data is processed through the data cleaning channel. This channel is responsible for removing noise or inconsistencies in the data. For example, if certain data points in temperature data collected by a sensor contain errors, the cleaning channel removes these errors to ensure data consistency and accuracy. Untrustworthy data is corrected through the traceability anomaly correction channel. This channel corrects the data based on traceability information (such as device status and environmental conditions). For example, if a temperature sensor malfunctions and generates anomalous data, the traceability correction channel will repair the data based on the device's historical status and ambient temperature information.

[0033] After cleansing and correcting, trusted and untrusted data are processed through their respective channels to create two enhanced data areas: the trusted industrial data area: Data cleansing ensures error-free data and improves data quality. The corrected industrial data area: Untrusted data is corrected through traceability and anomaly correction, improving data quality. Finally, the data in the trusted and corrected industrial data areas are merged to create the final enhanced industrial data source. This step ensures unified processing of data from different sources, improving overall data credibility.

[0034] By utilizing dual data enhancement channels, the system effectively improves the quality and credibility of industrial data. The data cleansing channel removes noise from trusted data, ensuring only accurate data is used. The traceability and anomaly correction channel corrects unreliable data, preventing erroneous data from influencing decision-making. This significantly enhances the system's robustness and accuracy in the face of uncertain environments and fluctuating data quality.

[0035] The trusted industrial data block establishment module 15 is used to input the enhanced industrial data source into the blockchain trusted layer, and the blockchain trusted layer securely stores the enhanced industrial data source on and off the chain to establish a trusted industrial data block.

[0036] Specifically, on-chain and off-chain secure storage refers to storing part of the data on the blockchain (on-chain) to ensure the data’s immutability, while storing most of the data in an off-chain storage system (off-chain) to reduce the storage pressure on the blockchain.

[0037] First, the system inputs multi-processed, enhanced industrial data sources into the blockchain's trusted layer. This data undergoes cleansing, anomaly correction, and enhancement to ensure its quality and credibility. Within the blockchain's trusted layer, data is stored in two parts: Off-chain storage: Due to the limited storage space of the blockchain, the actual data content is stored in a distributed storage system external to the blockchain (such as IPFS or cloud storage). This off-chain storage can accommodate larger amounts of industrial data, such as raw sensor data and equipment logs. On-chain storage: To ensure data immutability and long-term traceability, key data features (such as digests, hash values, data source IDs, and timestamps) are stored on the blockchain. Storing this feature information on the blockchain ensures data uniqueness and prevents tampering.

[0038] By storing key features of enhanced industrial data sources on the blockchain, the system forms a trusted industrial data block. This block contains verified, reliable data and corresponding on-chain evidence. The decentralized nature of blockchain allows anyone to verify the authenticity of the data while also ensuring that the data cannot be tampered with during storage. Blockchain encryption technology ensures that data stored on the blockchain cannot be tampered with, and all data changes are traceable. The data's source, storage process, and modification history can be traced at any time, ensuring data transparency and security.

[0039] Through this step, the system achieves enhanced secure storage and verification of industrial data sources, using blockchain technology to ensure the data's immutability and long-term traceability.

[0040] Furthermore, the traceability inspection result acquisition module includes: an industrial data extraction unit, used to extract the nth industrial data based on the perceived industrial data source, where n is a positive integer; a traceability scenario information acquisition unit, used to connect to the industrial data perception layer, retrieve the perception device status information and perception device environment information corresponding to the nth industrial data, and obtain the nth traceability scenario information; a traceability credibility coefficient acquisition unit, used to perform credibility evaluation based on the nth traceability scenario information, and obtain the nth traceability credibility coefficient; a traceability inspection result output unit, used to input the nth traceability credibility coefficient into the traceability inspection channel, output the nth traceability inspection result, and add the nth traceability inspection result to the multiple traceability inspection results.

[0041] Specifically, "nth industrial data" refers to the nth piece of specific data selected from a perceived industrial data source, where n is a positive integer greater than or equal to 1. Perceived device status information refers to the operating status data of devices related to the industrial data source, such as whether the sensor is functioning properly or has any faults. Perceived device environmental information refers to parameters of the environment in which the perceived device resides, such as external factors like temperature and humidity, which may affect device operation. Traceability scenario information is data generated based on device status and environmental information. It describes the context of the data collection process and helps understand the data's source, collection conditions, and possible sources of error. Credibility evaluation analyzes traceability scenario information to assess the data's credibility. This is typically measured by calculating a credibility coefficient, with data with a higher credibility coefficient considered more credible. The traceability credibility coefficient is a value calculated based on the traceability scenario information and the credibility evaluation to indicate the credibility of a piece of data. The higher the value, the more credible the data.

[0042] First, the perception industrial data source is data collected from the perception layer of the Industrial Internet of Things. This data is encrypted and transmitted to the edge computing layer. The traceability verification channel in the edge computing layer is responsible for verifying the data's provenance to ensure the authenticity of the data source. A specific piece of industrial data is selected from multiple perception data sources, referred to as the nth piece of industrial data. This might be the temperature value captured by a temperature sensor at a specific moment, or a reading from a pressure sensor. The system then retrieves the perception device status and perception device environment information associated with this nth piece of industrial data. For example, if this data comes from a temperature sensor, the system checks the sensor's operating status (whether it is operating properly) as well as the ambient temperature, humidity, and other environmental conditions. This information forms the nth piece of traceability scenario information, providing the basis for subsequent credibility evaluation.

[0043] The system uses the collected traceability scenario information to perform a credibility evaluation. A credibility evaluation model, such as a machine learning model or a rule-based system, is usually used to analyze the device status and environmental information. For example, if the temperature sensor is in normal working condition and the ambient temperature is stable, the credibility of the data is high. On the contrary, if the device is in a faulty state or the environment is abnormal, the data credibility is low. Based on the credibility evaluation results, the system generates a traceability credibility coefficient, which reflects the reliability of the data. Finally, the nth traceability credibility coefficient is input into the traceability inspection channel, and the channel determines whether the data is credible by comparing it with the set predetermined credibility threshold. If the credibility coefficient is higher than the preset threshold, the data passes the inspection and is considered credible; otherwise, the data is marked as unreliable. The traceability inspection result will be added to multiple traceability inspection results.

[0044] Through this traceability verification process, the system can evaluate the credibility of each piece of industrial data in real time, thereby ensuring data reliability and validity. This effectively prevents data anomalies caused by sensor failure, data tampering, or environmental factors, reduces the impact of erroneous data on production decisions, and thus improves the reliability and security of the entire Industrial Internet of Things system.

[0045] Furthermore, the traceability credibility coefficient acquisition unit includes: a sample set loading unit, used to connect to the industrial Internet of Things, load the traceability scenario sample set and the traceability credibility evaluation sample set; an evaluation model acquisition unit, used to perform supervised learning on K traceability credibility evaluation learners based on the traceability scenario sample set and the traceability credibility evaluation sample set, and obtain K traceability credibility evaluation models that meet the predetermined traceability credibility evaluation accuracy, wherein K is a positive integer greater than 1; a scene information input unit, used to input the nth traceability scenario information into the K traceability credibility evaluation models, and obtain K traceability credibility evaluation coefficients; a mean calculation unit, used to perform mean calculation on the K traceability credibility evaluation coefficients, and output the nth traceability credibility coefficient.

[0046] Specifically, the provenance trust evaluation learner is a machine learning model used to assess data credibility. It builds an evaluation mechanism by learning from a provenance scenario sample set and a provenance trust evaluation sample set. The provenance trust evaluation model is a trained model that outputs a credibility evaluation of the data based on the input provenance scenario information.

[0047] First, the system loads traceability scenario sample sets and traceability trust evaluation sample sets from the Industrial Internet of Things. These sample sets contain previously acquired traceability scenario information and corresponding known trustworthiness labels for subsequent model training. The sample sets can include different types of sensing devices and working environments, encompassing a wide range of possible sensor states and environmental conditions. For example, the status and environmental information of a temperature sensor may include whether the device is functioning properly, the operating environment temperature, and humidity.

[0048] Next, the system inputs the traceability scenario sample set and the traceability trust evaluation sample set into K traceability trust evaluation learners for supervised learning. By analyzing these sample sets, the learners learn how to evaluate the credibility of the data based on the traceability scenario information. Specifically, representative traceability scenario sample sets and corresponding credibility labels are collected as training data. Then, a suitable machine learning algorithm, such as a decision tree, support vector machine (SVM), or deep learning model, is selected to train the sample set. During the training process, the model generates a credibility assessment model by learning the relationship between the characteristics of the traceability scenario (such as device status, environmental information, etc.) and credibility. Finally, the model's performance is evaluated using methods such as cross-validation to ensure that its credibility evaluation accuracy meets the predetermined requirements. Each model can focus on different characteristics of the traceability scenario, such as sensor status, environmental changes, etc.

[0049] After training is complete, the newly acquired n-th traceability scenario information is input into these K traceability credibility evaluation models. Each model will generate K traceability credibility evaluation coefficients, representing each model's credibility assessment value for the data. For example, the n-th traceability scenario information may include information such as the operating status, temperature, and humidity of a temperature sensor. Different learning models may assess its credibility based on different characteristics, resulting in different results. To synthesize the evaluation results of multiple models, the K traceability credibility evaluation coefficients are averaged to obtain a final n-th traceability credibility coefficient. By taking the average, the potential bias of individual models can be reduced, improving the accuracy and stability of the evaluation.

[0050] Finally, the calculated nth traceability credibility coefficient is output, which is used to represent the credibility of the nth data. This credibility coefficient will be used for subsequent traceability verification to ensure that the data can be stored and used correctly.

[0051] This step uses multiple machine learning models to evaluate traceability scenario information, significantly improving the accuracy of data credibility assessments. This multi-model comprehensive assessment method effectively mitigates risks associated with sensor failures or data tampering, thereby ensuring data accuracy and reliability across the entire IoT system.

[0052] Furthermore, the traceability verification channel includes a traceability verification operator, and the traceability verification operator includes: a traceability credibility judgment unit, if the nth traceability credibility coefficient is greater than or equal to a predetermined traceability credibility coefficient, the nth traceability verification result is traceability credibility; a traceability untrustworthy judgment unit, if the nth traceability credibility coefficient is less than the predetermined traceability credibility coefficient, the nth traceability verification result is traceability untrustworthy.

[0053] Specifically, the traceability verification operator is a core component in the traceability verification channel. It is responsible for judging the traceability credibility coefficient and outputting a conclusion on whether the data is trustworthy. It uses a predetermined credibility threshold to judge the data. The predetermined traceability credibility coefficient is a set credibility threshold that is used for comparison with the traceability credibility coefficient as a standard for determining whether the data is trustworthy.

[0054] First, the nth piece of data is extracted from the perception data source, and the nth traceability credibility coefficient of the data is calculated through the aforementioned steps. For example, if the temperature sensor is in normal condition and the environment is stable, the credibility of the data is high, resulting in a higher credibility coefficient. Then, the nth traceability credibility coefficient is compared with the preset predetermined traceability credibility coefficient. The predetermined credibility coefficient is a set standard value that represents the minimum credibility requirement that the data should meet. For example, a predetermined credibility coefficient of 0.8 means that only data with a credibility coefficient greater than or equal to 0.8 is considered credible. Based on the comparison result, the traceability verification operator outputs the corresponding test result. If the nth traceability credibility coefficient is greater than or equal to the predetermined credibility coefficient, the output is "traceability is credible"; otherwise, the output is "traceability is unreliable."

[0055] Through this judgment mechanism of the traceability verification operator, the system can efficiently assess the credibility of data, ensuring that only highly reliable data is processed or stored. The system can promptly filter out unreliable data, ensuring the safety and efficiency of the production process.

[0056] Furthermore, the enhanced industrial data source acquisition module 14 includes: a data enhancement dual-channel unit, the data enhancement dual-channel includes a data cleaning model and a traceability anomaly correction model; an industrial data source classification unit, used to classify the perceived industrial data source according to the multiple traceability inspection results, and establish a trusted industrial data area and an untrusted industrial data area; a first enhanced industrial data area acquisition unit, used to input the trusted industrial data area into the data cleaning model to obtain a first enhanced industrial data area; a corrected industrial data area acquisition unit, used to perform data correction on the untrusted industrial data area according to the traceability anomaly correction model to obtain a corrected industrial data area; a second enhanced industrial data area acquisition unit, used to perform enhancement processing on the corrected industrial data area according to the data cleaning model to obtain a second enhanced industrial data area; an enhanced industrial data area fusion unit, used to fuse the first enhanced industrial data area and the second enhanced industrial data area to obtain the enhanced industrial data source.

[0057] Specifically, the data cleansing model is used to cleanse and optimize industrial data models, aiming to remove noise, outliers, or erroneous information from the data, making it more consistent and accurate. The traceability anomaly correction model is used to correct untrustworthy data models, using traceability information (such as equipment status and environmental conditions) to correct anomalies or errors in the data. The trusted industrial data area contains all data that has passed traceability verification and been deemed trustworthy. After undergoing credibility assessment, data enters this area for further processing. The untrustworthy industrial data area contains all data deemed untrustworthy. This data requires correction using the traceability anomaly correction model to improve its credibility. The first enhanced industrial data area is the data area obtained after processing by the data cleansing model, and all cleaned data is contained in this area. The second enhanced industrial data area is the data area corrected by the traceability anomaly correction model. This data was originally untrustworthy but has been enhanced to be trusted after processing.

[0058] First, the system categorizes all perceived industrial data sources based on previous traceability verification results. Based on these results, data is divided into two categories: trusted and untrusted. For example, if a temperature sensor's readings are consistent with pre-set standards and pass traceability verification, it is assigned to the trusted industrial data zone. However, if the data readings are abnormal or the sensor's operating status is unstable, it is assigned to the untrusted industrial data zone.

[0059] Next, all data belonging to the trusted industrial data zone is input into the data cleansing model. The data cleansing model construction method typically includes the following steps: Collect and preprocess raw data to identify missing values, duplicate values, and outliers. Statistical analysis methods or rule-based algorithms (such as Z-score and IQR) are used to detect abnormal data. Missing and outliers are handled using strategies such as interpolation, filling, or deletion to ensure data consistency and integrity. To improve model performance, machine learning methods such as decision trees or clustering algorithms are incorporated to automatically identify and correct noise in the data. Finally, cross-validation and evaluation metrics are used to ensure that the model cleansing results achieve the expected accuracy. The data cleansing model identifies and removes noise, erroneous values, and inconsistencies in the data to ensure data consistency and accuracy. For example, if temperature data within a certain period exhibits sudden, sharp fluctuations (perhaps caused by a temporary sensor failure), the data cleansing model removes these anomalies, retaining the true temperature fluctuation data. The cleansed data forms the first enhanced industrial data zone.

[0060] Data belonging to the untrusted industrial data zone is first input into the traceability anomaly correction model. The method for constructing the traceability anomaly correction model includes the following steps: Raw data with traceability information is collected and annotated. Based on historical data and device status information, supervised learning algorithms (such as regression analysis, decision trees, or support vector machines) are used to model the data and learn the relationship between the data and traceability information. This model corrects errors in the data based on traceability information (such as device status and operating environment). For example, if sensor data doesn't match other devices or historical data, the correction model adjusts the data to a reasonable range based on factors such as the sensor's operating status and environmental changes. After correction, this data forms the corrected industrial data zone.

[0061] Finally, the system further enhances the data in the corrected industrial data area. This enhancement typically uses a data cleaning model to further remove noise from the data, generating a second enhanced industrial data area. The system then fuses the first enhanced industrial data area (cleaned, trusted data) with the second enhanced industrial data area (corrected, untrusted data) to create the final enhanced industrial data source.

[0062] Through this dual-channel data enhancement approach, the system significantly improves data quality and credibility. The data cleansing model ensures the purity and consistency of trusted data, avoiding analytical errors caused by noise or erroneous values. The traceability anomaly correction model repairs unreliable data, bringing it closer to reality and making it more useful. This avoids production delays or decision-making errors caused by missing data.

[0063] Furthermore, the correction industrial data area acquisition unit includes: an untrusted industrial data area traversal unit, used to traverse the untrusted industrial data area and extract the first untrusted industrial data; a traceability scenario information corresponding unit, used to use the traceability scenario information corresponding to the first untrusted industrial data as the untrusted traceability scenario information; an anomaly detection unit, used to perform anomaly detection based on the untrusted traceability scenario information and obtain a traceability scenario anomaly detection result; a first correction industrial data acquisition unit, used to input the first untrusted industrial data and the traceability scenario anomaly detection result into the traceability anomaly correction model, obtain the first correction industrial data, and add the first correction industrial data to the correction industrial data area.

[0064] Specifically, untrusted traceability scenario information refers to the equipment and environment information related to untrusted industrial data. This includes background data such as equipment status and working environment, which helps determine the credibility of the data.

[0065] First, the system traverses the untrusted industrial data area and extracts the first untrusted industrial data. For each piece of untrusted data, the system extracts relevant traceability scenario information, such as the device's operating status, ambient temperature, and humidity. This traceability scenario information enables the system to understand the data's context and possible sources of anomalies. For example, if a temperature sensor operates in an extreme environment, it may generate anomalous data, and the traceability scenario information provides relevant context. Based on this extracted untrusted traceability scenario information, the system performs anomaly detection. This process uses data analysis methods such as statistical tests and model-based detection methods (such as decision trees and cluster analysis) to detect anomalies in the traceability scenario.

[0066] The first untrusted industrial data, along with the anomaly detection results from the traceability scenario, are input into the traceability anomaly correction model. The correction model uses the traceability information to correct the data. For example, if sensor data is inaccurate due to a device failure, the correction model will correct the data value based on the normal operating state of the device or environmental conditions, generating the first corrected industrial data. The corrected data is then added to the corrected industrial data area, ensuring improved quality and reliability after anomaly correction. The corrected data can be used for subsequent storage, analysis, and decision-making.

[0067] Through this process of the traceability anomaly correction model, the system can significantly improve the credibility of untrusted data, thereby ensuring the accuracy and reliability of subsequent data. This step can accurately correct abnormal data by using traceability information of the device and environment, avoiding the errors that may be caused by relying on simple filtering or mean padding to process abnormal data.

[0068] Further, such as Figure 2As shown, the trusted industrial data block establishment module 15 includes: multiple industrial data shard acquisition units, used to perform sharding processing on the enhanced industrial data source to obtain multiple industrial data shards; an off-chain storage block establishment unit, used to perform off-chain encrypted storage of the multiple industrial data shards according to the blockchain trusted layer, and establish an off-chain storage block; an on-chain evidence block establishment unit, used to perform on-chain secure storage of key features of the multiple industrial data shards according to the blockchain trusted layer, and establish an on-chain evidence block; an on-chain off-chain mapping model construction unit, used to construct an on-chain off-chain mapping model according to the mapping relationship between the on-chain evidence block and the off-chain storage block; a storage block alignment unit, used to align the on-chain evidence block and the off-chain storage block according to the on-chain off-chain mapping model to generate the trusted industrial data block.

[0069] Specifically, sharding refers to splitting large datasets into multiple smaller parts (shards) for more efficient storage and processing. Each shard contains a subset of the original data, facilitating distributed storage and parallel processing. On-chain evidence blocks are data blocks stored on the blockchain, containing key data characteristics (such as data hash values, digests, and timestamps). These evidence blocks ensure the immutability and traceability of data on the blockchain. Off-chain storage blocks are encrypted data blocks stored outside the blockchain, containing detailed information that enhances industrial data sources. Off-chain storage can store large amounts of data, while blockchain storage space is limited. The on-chain-off-chain mapping model is a model used to establish a correspondence between on-chain evidence blocks and off-chain storage blocks. This model helps the system align on-chain evidence information with off-chain data storage, ensuring data integrity.

[0070] First, the system shards the augmented industrial data source, breaking the raw data into multiple smaller shards. For example, a set of sensor data such as temperature and pressure is broken down into multiple data fragments, each containing data from a specific time range or device. This processing method improves data storage and transmission efficiency and distributes storage pressure. The system then stores the sharded data in off-chain storage blocks. Due to the limited storage space of the blockchain, the complete data needs to be stored in an external distributed storage system (such as IPFS or cloud storage). Before storage, the system encrypts the data to ensure privacy and security. Furthermore, the system extracts key features from the sharded data, such as the data digest (hash value), timestamp, and data source ID, and stores these features on the blockchain.

[0071] Build an on-chain and off-chain mapping model that establishes a correspondence between on-chain evidence blocks and off-chain storage blocks. For example, the on-chain evidence block records the hash value of temperature data, while the off-chain storage block stores the specific temperature value. This mapping model ensures the correspondence between these data, thereby guaranteeing data accuracy and consistency. Specifically, key features of the off-chain storage block and the on-chain evidence block are extracted, such as the data digest (hash value), timestamp, and data source ID. These features are used to establish a correspondence, and by recording the mapping between each on-chain evidence block and its corresponding off-chain storage block, data consistency is ensured. This mapping relationship is stored using a database or distributed ledger technology for easy query and verification.

[0072] Finally, the system aligns the on-chain evidence block with the off-chain storage block based on the mapping model to generate a trusted industrial data block. This data block contains enhanced and encrypted industrial data, ensuring its security, integrity, and long-term traceability.

[0073] Through this process, the system achieves secure storage and efficient management of industrial data. Data is sharded and encrypted and stored outside the blockchain, ensuring efficient use of storage space, while key features are stored on the blockchain, ensuring data immutability and verifiability.

[0074] Furthermore, the on-chain evidence block establishment unit includes: a shard summary block acquisition unit, which is used to perform summary feature recognition on the multiple industrial data shards to obtain a shard summary block; a shard storage address block acquisition unit, which is used to collect the off-chain storage address parameters corresponding to the multiple industrial data shards to obtain a shard storage address block; a shard source block acquisition unit, which is used to collect the source feature parameters of the multiple industrial data shards to obtain a shard source block; a shard access control block acquisition unit, which is used to perform access control feature configuration on the multiple industrial data shards to obtain a shard access control block; an industrial data key feature block establishment unit, which is used to fuse the shard summary block, the shard storage address block, the shard source block and the shard access control block to establish an industrial data key feature block; an on-chain secure storage unit, which is used to perform on-chain secure storage of the industrial data key feature block according to the blockchain trusted layer to obtain the on-chain evidence block.

[0075] Specifically, summary feature recognition involves hashing data to generate a summary (or hash value). The hash value uniquely identifies the data and can be used to verify its integrity and immutability. A shard summary block is a data block generated by performing summary feature recognition on a data shard. An industrial data key feature block consists of multiple key features, including the shard's hash value, storage address, source information, and access control information, ensuring data integrity, source traceability, and security.

[0076] First, each industrial data shard is hashed to generate a shard summary block. This step ensures the uniqueness of each data shard by calculating a hash value (e.g., using the SHA-256 algorithm). For example, a temperature sensor data shard is hashed to generate a unique summary representing the characteristics of that data. The system then obtains the storage address of each data shard outside the blockchain, known as the shard storage address block. This data is stored in distributed storage systems (such as IPFS or distributed cloud storage), so the system needs to record the specific storage location of each shard. For example, a file storing temperature data may be stored on a specific node in the IPFS network, and this address information is part of the storage address block. Furthermore, the system collects source characteristic parameters related to each data shard, such as the source device information and acquisition time, to form a shard source block. For example, source characteristics for temperature data may include the device ID of the temperature collector, the sensor model, and the acquisition environment. The system then configures access permissions for each data shard, forming a shard access control block. This step ensures data security by setting access permissions.

[0077] The system then combines the four features (shard summary block, shard storage address block, shard source block, and shard access control block) to form the industrial data key feature block. This key feature block contains all key information for the data shard, ensuring its uniqueness and traceability on the blockchain. Finally, the industrial data key feature block is stored in the blockchain's trusted layer, forming an on-chain evidence block. This evidence block is permanently stored on the blockchain, ensuring the immutability and security of the data. The hash value and other feature information in the evidence block can be verified at any time to ensure that the data has not been tampered with.

[0078] Through this step, the system achieves secure storage and verification of industrial data, ensuring the security and traceability of data within the distributed storage system. Key information for each data shard is stored on the blockchain, ensuring that the data source can be verified, the storage location can be traced, and only authorized users can access the data.

[0079] In summary, the blockchain-based industrial IoT data security collection system provided by the embodiments of the present application has the following technical effects:

[0080] 1. By combining blockchain with the Industrial Internet of Things, the security and credibility of the data collection process are ensured. The collaborative work of the perception layer, edge computing layer, and blockchain layer improves data reliability and transparency, effectively prevents data tampering and forgery, and enhances the intelligence of industrial equipment monitoring and management.

[0081] 2. Data enhancement: Dual channels improve the quality of both trusted and untrusted data through cleaning and correction. The dual effects of the cleaning model and the anomaly correction model ensure the accuracy and integrity of industrial data and enhance the reliability of data analysis results.

[0082] 3. Blockchain's on-chain and off-chain storage solutions ensure data immutability and long-term traceability. By sharding and encrypting data, combining on-chain evidence blocks with off-chain storage blocks, data security and integrity are enhanced, meeting the storage needs of large-scale industrial data.

[0083] Any step of the method described above can be stored as a computer instruction or program in an unlimited computer memory, and can be called and recognized by an unlimited computer processor to implement any method in the embodiments of the present application, without any unnecessary restrictions.

[0084] Furthermore, the terms "first" or "second" as described above may not only represent an order relationship but may also represent a specific concept and / or refer to the selection of multiple elements individually or collectively. Obviously, those skilled in the art may make various modifications and variations to this application without departing from the scope of this application. Thus, if such modifications and variations fall within the scope of this application and its equivalents, this application is intended to include such modifications and variations.

Claims

1. The industrial Internet of Things data security collection system based on blockchain is characterized by: The system comprises: A collection network establishment module, used to connect to the industrial Internet of Things and establish an industrial data collection network, wherein the industrial data collection network includes an industrial data perception layer, an edge computing layer, and a blockchain trust layer; A perception industrial data source acquisition module, configured to interact with the industrial data perception layer in real time to obtain a perception industrial data source; a traceability inspection result obtaining module, configured to encrypt and transmit the perceived industrial data source to the edge computing layer, and perform traceability inspection on the perceived industrial data source through a traceability inspection channel within the edge computing layer to obtain a plurality of traceability inspection results; An enhanced industrial data source acquisition module is configured to activate the data enhancement dual channels in the edge computing layer to enhance the perceived industrial data source based on the multiple traceability inspection results to obtain an enhanced industrial data source. The enhanced industrial data source acquisition module includes: A data enhancement dual-channel unit, wherein the data enhancement dual-channel includes a data cleaning model and a traceability anomaly correction model; An industrial data source classification unit, configured to classify the perceived industrial data sources according to the plurality of traceability inspection results, and establish a trusted industrial data area and an untrusted industrial data area; A first enhanced industrial data area obtaining unit, configured to input the trusted industrial data area into the data cleaning model to obtain a first enhanced industrial data area; a corrected industrial data area obtaining unit, configured to perform data correction on the untrusted industrial data area according to the traceability anomaly correction model to obtain a corrected industrial data area; A second enhanced industrial data area obtaining unit is configured to perform enhancement processing on the corrected industrial data area according to the data cleaning model to obtain a second enhanced industrial data area; an enhanced industrial data area fusion unit, configured to fuse the first enhanced industrial data area and the second enhanced industrial data area to obtain the enhanced industrial data source; The correction industrial data area obtaining unit includes: an untrusted industrial data area traversal unit, configured to traverse the untrusted industrial data area and extract first untrusted industrial data; a traceability scenario information corresponding unit, configured to use the traceability scenario information corresponding to the first untrusted industrial data as the untrusted traceability scenario information; an anomaly detection unit, configured to perform anomaly detection based on the untrusted traceability scenario information and obtain an anomaly detection result of the traceability scenario; a first corrected industrial data obtaining unit, configured to input the first untrusted industrial data and the traceability scenario anomaly detection result into the traceability anomaly correction model to obtain first corrected industrial data, and add the first corrected industrial data to the corrected industrial data area; The trusted industrial data block establishment module is used to input the enhanced industrial data source into the blockchain trusted layer, and the blockchain trusted layer securely stores the enhanced industrial data source on and off the chain to establish a trusted industrial data block.

2. The blockchain-based industrial Internet of Things data security collection system according to claim 1, characterized in that: The traceability test result acquisition module includes: an industrial data extraction unit, configured to extract n-th industrial data according to the sensed industrial data source, wherein n is a positive integer; a traceability scenario information obtaining unit, configured to connect to the industrial data perception layer, retrieve the perception device status information and the perception device environment information corresponding to the nth industrial data, and obtain the nth traceability scenario information; a traceability credibility coefficient obtaining unit, configured to perform credibility evaluation based on the nth traceability scenario information to obtain an nth traceability credibility coefficient; The traceability inspection result output unit is used to input the nth traceability credibility coefficient into the traceability inspection channel, output the nth traceability inspection result, and add the nth traceability inspection result to the multiple traceability inspection results.

3. The blockchain-based industrial Internet of Things data security collection system according to claim 2, characterized in that: The traceability credibility coefficient obtaining unit includes: A sample set loading unit, used to connect to the industrial Internet of Things and load the traceability scenario sample set and the traceability trust evaluation sample set; An evaluation model obtaining unit, configured to perform supervised learning on K traceability trustworthy evaluation learners based on the traceability scenario sample set and the traceability trustworthy evaluation sample set, to obtain K traceability trustworthy evaluation models that meet a predetermined traceability trustworthy evaluation accuracy, where K is a positive integer greater than 1; A scenario information input unit, configured to input the nth traceability scenario information into the K traceability credibility evaluation models to obtain K traceability credibility evaluation coefficients; The mean calculation unit is used to perform mean calculation on the K traceability credibility evaluation coefficients and output the nth traceability credibility coefficient.

4. The blockchain-based industrial Internet of Things data security collection system according to claim 2, characterized in that: The traceability verification channel includes a traceability verification operator, which includes: a traceability credibility judgment unit, wherein if the nth traceability credibility coefficient is greater than or equal to a predetermined traceability credibility coefficient, the nth traceability test result is traceability credibility; The traceability unreliable judgment unit is configured to determine that the traceability is unreliable if the nth traceability credibility coefficient is less than the predetermined traceability credibility coefficient.

5. The blockchain-based industrial Internet of Things data security collection system according to claim 1, characterized in that: The trusted industrial data block establishment module includes: a plurality of industrial data shard obtaining units, configured to perform sharding processing on the enhanced industrial data source to obtain a plurality of industrial data shards; An off-chain storage block establishing unit, configured to perform off-chain encrypted storage on the plurality of industrial data shards according to the blockchain trusted layer, and establish an off-chain storage block; An on-chain evidence block establishment unit, configured to securely store key features of the plurality of industrial data shards on-chain according to the blockchain trusted layer, and establish an on-chain evidence block; An on-chain and off-chain mapping model construction unit, configured to construct an on-chain and off-chain mapping model based on a mapping relationship between the on-chain evidence block and the off-chain storage block; The storage block alignment unit is used to align the on-chain evidence block and the off-chain storage block according to the on-chain and off-chain mapping model to generate the trusted industrial data block.

6. The blockchain-based industrial Internet of Things data security collection system according to claim 5, characterized in that: The on-chain evidence block establishment unit includes: A fragment summary block obtaining unit, configured to perform summary feature recognition on the plurality of industrial data fragments to obtain fragment summary blocks; A shard storage address block obtaining unit, configured to collect off-chain storage address parameters corresponding to the plurality of industrial data shards and obtain a shard storage address block; A slice source block obtaining unit, configured to collect source characteristic parameters of the plurality of industrial data slices to obtain slice source blocks; A shard access control block obtaining unit, configured to perform access control feature configuration on the plurality of industrial data shards to obtain a shard access control block; an industrial data key feature block establishing unit, configured to fuse the slice summary block, the slice storage address block, the slice source block, and the slice access control block to establish an industrial data key feature block; The on-chain secure storage unit is used to securely store the key feature block of industrial data on the chain according to the blockchain trusted layer to obtain the on-chain evidence block.

Citation Information

Patent Citations

  • Physical certificate traceability system for edge computing service based on blockchain

    CN108769031A

  • Cloud data operation behavior-oriented trusted traceability method

    CN113886841A