Industrial Internet of Things data preprocessing method, medium and system based on micro service

By adopting microservice-based data preprocessing methods in the industrial Internet of Things, using the data graph network and three-layer data pyramid model, the problems of uneven data quality and unreasonable resource allocation are solved, and efficient data preprocessing is achieved.

CN120104970APending Publication Date: 2025-06-06BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510179735.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The uneven quality of industrial IoT data and unreasonable allocation of processing resources lead to inefficient data preprocessing.

Method used

Using the preprocessing method of industrial IoT data based on microservices, by building an IoT data graph network and a three-layer data pyramid model, intelligent evaluation of data quality, dynamic allocation of processing resources, and efficient collaboration between microservices are achieved.

Benefits of technology

Accurate evaluation and hierarchical processing of data quality are realized, ensuring that computing resources are tilted towards important data, improving overall processing efficiency, and reducing data redundancy and latency between services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104970A_ABST
    Figure CN120104970A_ABST
Patent Text Reader

Abstract

The invention provides an industrial Internet of Things data preprocessing method, medium and system based on micro-service, and belongs to the technical field of electrical digital data process.The method comprises the steps that firstly, an industrial Internet of Things micro-service system and a data graph network are built, and original data are collected through data collection micro-service; and processing by using a pre-trained three-layer data pyramid model. The method innovatively introduces a data workable degree index to carry out quality evaluation, carries out data classification through a data classification pyramid layer, distributes processing tasks based on a data cleaning demand coefficient, carries out data cleaning through a data cleaning pyramid layer, and extracts features through a feature extraction pyramid layer. And data exchange between services is optimized through a data distribution matching coefficient, and finally, dynamic resource allocation is realized based on a data contribution index, so that efficient data preprocessing is completed, and the technical problem of low data preprocessing efficiency caused by uneven data quality of the industrial Internet of Things and unreasonable processing resource allocation is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electronic digital data processing, and in particular, relates to a microservice-based industrial Internet of Things data preprocessing method, medium and system. Background Art

[0002] The Industrial Internet of Things is the core infrastructure of modern intelligent manufacturing, and the quality and efficiency of its data preprocessing directly affect the level of intelligence in industrial production. Traditional Industrial Internet of Things data preprocessing mainly adopts a centralized architecture, and the collected data is cleaned, converted, and stored through a unified data processing center. This method runs stably when the amount of data is small, but with the increase in the number of industrial equipment and the increasing variety of data types, the limitations of the centralized architecture are becoming increasingly prominent. At the same time, traditional preprocessing methods often use fixed processing procedures and unified processing strategies, and cannot dynamically adjust the processing method according to data characteristics and business needs.

[0003] The main problems in the existing technology are as follows: First, the quality of data generated by the industrial Internet of Things varies significantly. Some devices generate high-quality data, while some devices generate data with noise, missing or delayed problems. It is difficult for a unified processing strategy to adapt to the processing needs of data of different quality levels. Secondly, the allocation of processing resources lacks pertinence, and fails to reasonably allocate computing resources according to the importance of data and the urgency of processing, resulting in delays in the processing of important data, while non-critical data occupies a large amount of resources. Thirdly, although the existing microservice architecture provides distributed processing capabilities, the collaborative efficiency between services is not high, and there is redundancy and delay when data is transmitted between different services. Finally, there is a lack of in-depth analysis of data value and data relevance, and it is impossible to accurately evaluate the contribution of data to system operation, which affects the rationality of resource allocation.

[0004] These problems lead to low efficiency of data preprocessing in the Industrial Internet of Things, which cannot meet the requirements of modern intelligent manufacturing for real-time and accuracy of data processing. Especially in large-scale industrial production environments, the efficiency of data preprocessing directly affects production efficiency and product quality. In summary, the existing technology has technical problems such as uneven quality of Industrial Internet of Things data and unreasonable allocation of processing resources, which lead to low efficiency of data preprocessing. Summary of the invention

[0005] In view of this, the present invention provides an industrial Internet of Things data preprocessing method, medium and system based on microservices, which can solve the technical problems in the prior art such as uneven quality of industrial Internet of Things data and unreasonable allocation of processing resources leading to low data preprocessing efficiency.

[0006] The present invention is implemented as follows: In a first aspect, the present invention provides a microservice-based industrial Internet of Things data preprocessing method comprising the following steps: building an industrial Internet of Things microservice system including a data acquisition microservice, a data cleaning microservice, a data conversion microservice, a data storage microservice, and a data distribution microservice; inputting raw data into a pre-trained three-layer data pyramid model, wherein the three-layer data pyramid model comprises a data classification pyramid layer, a data cleaning pyramid layer, and a feature extraction pyramid layer, wherein the data classification pyramid layer adopts an improved classification neural network structure, and a data weight adaptation layer is arranged between the fully connected layer and the output layer of the classification neural network; the data cleaning pyramid layer adopts an improved encoding and decoding network structure, and a data residual retention layer is arranged between the encoding unit and the decoding unit; the feature extraction pyramid layer adopts an improved deep feature network structure, and a feature fusion processing layer is arranged between the feature dimensionality reduction layer and the feature integration layer; and classifying and cleaning the data and performing feature extraction processing based on data availability index, data cleaning requirement coefficient, and data distribution matching coefficient.

[0007] Among them, the steps of constructing an industrial Internet of Things data graph network include: using industrial equipment nodes as graph network nodes, and the data transmission relationship between devices as the graph network edge to form an Internet of Things graph structure; using a depth-first search algorithm to traverse the device network and establish an adjacency matrix to represent the connection relationship between devices; using a minimum spanning tree algorithm to optimize the network structure to ensure the optimality of the data transmission path; clustering device nodes based on a community discovery algorithm to form a hierarchical graph network.

[0008] Among them, the data availability index is calculated through data integrity component, data accuracy component and data timeliness component. The weight coefficient of the data integrity component is 0.4, the weight coefficient of the data accuracy component is 0.35, and the weight coefficient of the data timeliness component is 0.25. The weighted sum of the three components is used as the data availability index, and the value range is between 0 and 1.

[0009] Among them, the data cleaning requirement coefficient is calculated based on the data quality level identifier, the data update frequency value, and the business importance level identifier. The basic coefficient corresponding to the data quality level identifier is: 0.3 for high-quality data group, 0.5 for medium-quality data group, and 0.7 for low-quality data group; the weight corresponding to the data update frequency value is: the weight of the update frequency above once per second is 0.4, the weight of once per second to once per minute is 0.3, and the weight of less than once per minute is 0.2; the weight corresponding to the business importance level identifier is: 0.5 for critical business, 0.3 for important business, and 0.2 for general business.

[0010] The data distribution matching coefficient is obtained by calculating the data statistical characteristics of adjacent industrial equipment nodes, the cosine similarity is used to calculate the similarity of numerical features, and the mutual information is used to calculate the correlation of categorical features. The data distribution matching coefficient ranges from 0 to 1.

[0011] Among them, the data cleaning pyramid layer adopts different cleaning strategies for high-quality data groups, medium-quality data groups, and low-quality data groups: a lightweight cleaning strategy is adopted for the high-quality data group, which mainly performs outlier detection and simple filtering; a standard cleaning strategy is adopted for the medium-quality data group, including missing value filling, outlier processing, and noise elimination; a deep cleaning strategy is adopted for the low-quality data group, including data reconstruction, feature repair, and consistency check.

[0012] The feature extraction pyramid layer extracts multi-level features through a multi-scale feature extraction module, adopts a feature pyramid network to achieve feature fusion, and adopts a top-down feature fusion strategy to fuse feature maps of different scales layer by layer.

[0013] Among them, a microservice resource allocation strategy is established according to the data contribution index. The data contribution index is obtained based on a comprehensive evaluation of the data availability index and the data distribution matching coefficient, with weights of 0.6 and 0.4 respectively; the resource utilization of each microservice is monitored, and the capacity is expanded when the utilization exceeds 80%, and the capacity is reduced when it is lower than 30%.

[0014] A second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions are run in a computer, they are used to execute the above-mentioned industrial Internet of Things data preprocessing method based on microservices.

[0015] The third aspect of the present invention provides an industrial Internet of Things data preprocessing system based on microservices, comprising the above-mentioned computer-readable storage medium, the system is any one of a computer, a server, and a single-chip microcomputer, the computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

[0016] Compared with the prior art, the present invention provides a microservice-based industrial Internet of Things data preprocessing method, medium and system. The present invention proposes a microservice-based industrial Internet of Things data preprocessing method, which realizes intelligent evaluation of data quality, dynamic allocation of processing resources and efficient collaboration between microservices by constructing an Internet of Things data graph network and a three-layer data pyramid model. The method innovatively introduces evaluation mechanisms such as data availability index, data cleaning demand coefficient, data distribution matching coefficient and data contribution index, providing a scientific decision-making basis for data processing.

[0017] The solution of the present invention effectively solves the problems in the prior art: accurate evaluation and classification of data quality are achieved through the data classification pyramid layer, so that data of different qualities can be adaptively processed; the residual retention mechanism of the data cleaning pyramid layer ensures that key information is not lost during the data cleaning process; the multi-dimensional feature fusion of the feature extraction pyramid layer improves the expressiveness of data features. At the same time, the resource allocation strategy based on the data contribution index ensures that computing resources are tilted towards important data, thereby improving the overall processing efficiency. In addition, the data exchange between microservices is guided by the data distribution matching coefficient, which reduces data redundancy and delay between services.

[0018] The present invention successfully solves the problem of low efficiency of industrial Internet of Things data preprocessing, which is mainly due to its multi-level data evaluation mechanism and dynamic resource allocation strategy. The data availability index provides a quantitative evaluation standard for data quality, the data cleaning requirement coefficient guides the determination of processing priority, the data distribution matching coefficient optimizes the data flow between services, and the data contribution index realizes precise control of resource allocation. These innovative technical features cooperate with each other to form an efficient and reliable data preprocessing system, which solves the technical problem of low data preprocessing efficiency in the prior art due to uneven quality of industrial Internet of Things data and unreasonable allocation of processing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flow chart of the method of the present invention.

[0020] Figure 2 This is a characteristic analysis diagram of acquisition parameters of different types of data (operating status, environmental status, process parameters) in Example 2.

[0021] Figure 3 This is a heat map of data transmission frequency between device nodes in Example 2.

[0022] Figure 4 This is a data availability index analysis diagram for each device node in Example 2.

[0023] Figure 5 This is a comparison chart of the cleaning effects of different quality data groups in Example 2.

[0024] Figure 6 This is a diagram of resource usage of the microservice system in Example 2.

[0025] Figure 7 Graph showing the data quality assessment results for each process in Example 3. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0027] like Figure 1 FIG. 1 is a flowchart of a microservice-based industrial Internet of Things data preprocessing method provided by the first aspect of the present invention. The method comprises the following steps:

[0028] S01. Build an industrial Internet of Things microservice system, which includes data collection microservice, data cleaning microservice, data conversion microservice, data storage microservice, and data distribution microservice;

[0029] S02. Build an industrial Internet of Things data graph network, use industrial equipment nodes as graph network nodes, and use data transmission relationships between devices as graph network edges to form an Internet of Things graph structure;

[0030] S03. Collecting original data of each industrial equipment node in the industrial Internet of Things data graph network through the data collection microservice, wherein the original data includes equipment operation status data, equipment environment status data, and equipment process parameter data;

[0031] S04, inputting the original data into a pre-trained three-layer data pyramid model, wherein the three-layer data pyramid model includes a data classification pyramid layer, a data cleaning pyramid layer, and a feature extraction pyramid layer;

[0032] S05. Calculate the data availability index of each industrial equipment node, wherein the data availability index is obtained by calculating the data integrity component, the data accuracy component, and the data timeliness component;

[0033] S06, inputting the original data into the data classification pyramid layer, and dividing the original data into a high-quality data group, a medium-quality data group, and a low-quality data group according to the data usability index;

[0034] S07. Calculate the data cleaning requirement coefficient of each industrial equipment node, where the data cleaning requirement coefficient is calculated based on the data quality level identifier, the data update frequency value, and the business importance level identifier;

[0035] S08. According to the data cleaning requirement coefficient, the high-quality data group, the medium-quality data group, and the low-quality data group are respectively input into the data cleaning pyramid layer for data cleaning processing;

[0036] S09. Calculate a data distribution matching coefficient between adjacent industrial equipment nodes in the Internet of Things graph structure, where the data distribution matching coefficient represents a consistency level of data distribution between adjacent nodes;

[0037] S10, dynamically adjusting the data exchange scheme between microservices in the industrial Internet of Things microservice system based on the data distribution matching coefficient;

[0038] S11, inputting the data after data cleaning into the feature extraction pyramid layer, extracting the data feature vector, and calculating the data contribution index of each industrial equipment node;

[0039] S12. Establish a microservice resource allocation strategy according to the data contribution index, and allocate computing resources to the data collection microservice, the data cleaning microservice, the data conversion microservice, the data storage microservice, and the data distribution microservice;

[0040] S13. Distribute and store the processed data feature vectors to a distributed database through the data distribution microservice.

[0041] The equipment operation status data includes a working status identifier, an operation duration value, an equipment power value, an equipment speed value, an equipment vibration value, and an equipment temperature value;

[0042] The equipment environment status data includes the environment temperature value, environment humidity value, environment dust concentration value, environment noise value, and environment light intensity value;

[0043] The equipment process parameter data includes production line running speed value, process temperature value, process pressure value, material processing flow value, and product quality parameter value;

[0044] The data classification pyramid layer adopts an improved classification neural network structure, and a data weight adaptive layer is set between the fully connected layer and the output layer of the classification neural network. The data weight adaptive layer receives the quality feature input of the original data and outputs the data quality level identifier. The data classification pyramid layer realizes data classification through forward feature extraction and reverse error correction;

[0045] The data cleaning pyramid layer adopts an improved encoding and decoding network structure, and a data residual preservation layer is set between the encoding unit and the decoding unit. The data residual preservation layer receives the high-quality data group, the medium-quality data group, and the low-quality data group as input, and outputs standardized processed data. The data cleaning pyramid layer realizes data cleaning through data compression and reconstruction;

[0046] The feature extraction pyramid layer adopts an improved deep feature network structure, and a feature fusion processing layer is set between the feature dimension reduction layer and the feature integration layer. The feature fusion processing layer receives the standardized processed data input and outputs the data feature vector. The feature extraction pyramid layer realizes feature representation through multi-level feature extraction and dimension reduction.

[0047] The data availability index represents the availability evaluation result of the industrial equipment node data, and is obtained by comprehensive calculation based on the data integrity component, the data accuracy component, and the data timeliness component;

[0048] The data cleaning requirement coefficient represents the cleaning priority of the industrial equipment node data, and is calculated based on the data quality level identifier, the data update frequency value, and the business importance level identifier;

[0049] The data distribution matching coefficient represents the degree of similarity of data between industrial equipment nodes and is calculated by comparing the statistical characteristics of data of adjacent nodes;

[0050] The data contribution index characterizes the degree of influence of industrial equipment node data on system operation, and is obtained by comprehensive calculation based on the data availability index and the data distribution matching coefficient.

[0051] The specific implementation of the above steps is described in detail below.

[0052] The specific implementation method of step S01 is to build an industrial Internet of Things system based on a microservice architecture system, and use a containerized deployment method to achieve independent operation and elastic scaling of each microservice module. First, build a data acquisition microservice, which is responsible for collecting equipment data through communication methods such as industrial Ethernet and fieldbus, and supports multiple industrial communication protocols including message queue telemetry transmission protocol and object linking and embedding protocol. Then deploy a data cleaning microservice to pre-process the original data through methods such as data filtering and anomaly detection. Then implement a data conversion microservice to complete standardization processing such as data format conversion and unit unification. Then build a data storage microservice, and use a distributed database to achieve data persistence storage. Finally, deploy a data distribution microservice to achieve real-time data distribution based on the publish-subscribe model. Service discovery and load balancing are realized between each microservice through a service registration center, and asynchronous communication is realized by using message middleware to ensure high availability and scalability of the system. The purpose of this step is to build a microservice architecture with good elasticity and scalability to provide basic support for subsequent data processing.

[0053] The specific implementation method of step S02 is to construct an industrial Internet of Things data network model based on graph theory. First, analyze the topological relationship of the equipment on the industrial site, take each industrial equipment as the vertex of the graph network, calculate the in-degree and out-degree of the equipment node, and reflect the data interaction relationship between the equipment. Traverse the equipment network through the depth-first search algorithm, and establish an adjacency matrix to represent the connection relationship between the equipment. Based on the data flow between the devices, determine the directed edges in the graph network, and the weight of the edge is calculated by the data transmission frequency. The minimum spanning tree algorithm is used to optimize the network structure to ensure the optimality of the data transmission path. Cluster the equipment nodes based on the community discovery algorithm to form a hierarchical graph network. This step aims to build a graph network model that reflects the physical connection and logical relationship of industrial equipment, and provide network topology support for subsequent data processing.

[0054] The specific implementation method of step S03 is to obtain multi-dimensional operation data of industrial equipment through data acquisition microservices. First, configure the data acquisition parameters, including sampling period, sampling accuracy, communication timeout, etc. The sampling period is set between 100 milliseconds and 1 second according to the equipment type. For the equipment operation status data, a polling method is used to obtain the working status identification, record the equipment startup time, and collect parameters such as power, speed, vibration, and temperature. For the equipment environmental status data, parameters such as ambient temperature, humidity, dust concentration, noise, and light intensity are collected through a sensor network. For equipment process parameter data, parameters such as production line speed, process temperature, pressure, flow, and quality are collected in real time. A data caching mechanism is used to handle acquisition delays, and data acquisition priorities are set to ensure the timeliness of key parameter acquisition. The purpose of this step is to obtain complete industrial equipment operation data and provide a data source for subsequent data processing.

[0055] The specific implementation method of step S04 is to construct a three-layer data pyramid model. First, an improved classification neural network is implemented in the data classification pyramid layer, and a convolutional neural network is used to extract data features. Three convolution layers are set, and the convolution kernel sizes are 3×3, 5×5, and 7×7 respectively. The activation function uses a rectified linear unit function. A data weight adaptation layer is added between the fully connected layer and the output layer, and the feature weights are dynamically adjusted using an attention mechanism. Then, an improved encoding and decoding network is implemented in the data cleaning pyramid layer. The encoding unit uses a stacked autoencoder, and the decoding unit uses a deconvolution network. A data residual retention layer is added in the middle to retain the original data features. Finally, an improved deep feature network is implemented in the feature extraction pyramid layer, and a multi-scale feature extraction module is used to realize feature fusion through a feature pyramid network. This step aims to build an efficient data processing model to achieve data classification, cleaning, and feature extraction.

[0056] The specific implementation method of step S05 is to calculate the data availability index. First, the data completeness component is calculated, which is calculated by the statistical data missing rate and the outlier ratio, and the weight coefficient is set to 0.4. Then the data accuracy component is calculated based on the valid value range of the data and the measurement accuracy requirements, and the weight coefficient is set to 0.35. Finally, the data timeliness component is calculated based on the data collection delay and update time, and the weight coefficient is set to 0.25. The weighted sum of the three components is used as the data availability index, and the value range is between 0 and 1. A value greater than 0.8 indicates good data quality, a value between 0.5 and 0.8 indicates average data quality, and a value less than 0.5 indicates poor data quality. The purpose of this step is to evaluate the level of data availability and provide a basis for subsequent data classification.

[0057] The specific implementation method of step S06 is to use the data classification pyramid layer to classify the quality of the original data. First, the input data is normalized, and the minimum and maximum normalization method is used to map the data to the range of 0 to 1. Then the data is input into the improved classification neural network, the data features are extracted through the convolution layer, and the feature dimension is reduced by the pooling layer. The feature vector is calculated in the fully connected layer, and the data weight adaptation layer dynamically adjusts the feature weight according to the data usability index. Finally, the data quality level identification is output through the softmax classifier, and the data is divided into high-quality data group, medium-quality data group, and low-quality data group. This step aims to achieve data quality grading and provide a basis for subsequent differentiated processing.

[0058] The specific implementation method of step S07 is to calculate the data cleaning requirement coefficient. First, the basic coefficient is determined according to the data quality level identifier, which is 0.3 for the high-quality data group, 0.5 for the medium-quality data group, and 0.7 for the low-quality data group. Then, it is corrected according to the data update frequency value. The weight of the update frequency above once per second is 0.4, the weight of once per second to once per minute is 0.3, and the weight of less than once per minute is 0.2. Finally, weighting is performed according to the business importance level identifier, with the weight of key business being 0.5, the weight of important business being 0.3, and the weight of general business being 0.2. The weighted sum of the three factors is used as the data cleaning requirement coefficient, and the value range is between 0 and 1. The purpose of this step is to determine the priority of data cleaning and achieve a reasonable allocation of cleaning resources.

[0059] The specific implementation method of step S08 is to perform data cleaning through the data cleaning pyramid layer. First, a lightweight cleaning strategy is adopted for the high-quality data group, mainly performing outlier detection and simple filtering, using the triple standard deviation method to identify outliers, and using median filtering to remove noise. Then, a standard cleaning strategy is adopted for the medium-quality data group, including missing value filling, outlier processing, and noise elimination. The time series interpolation method is used to fill missing values, the local anomaly detection algorithm is used to process outliers, and the wavelet transform is used to remove noise. Finally, a deep cleaning strategy is adopted for the low-quality data group, including data reconstruction, feature repair, and consistency check. The deep learning model is used to reconstruct the data, the feature correlation analysis is used to repair the features, and the data consistency rule is used to check the results. This step aims to achieve differentiated data cleaning according to the data quality level and improve the data quality.

[0060] The specific implementation method of step S09 is to calculate the data distribution matching coefficient. First, the data statistical characteristics of adjacent industrial equipment nodes are calculated, including mean, variance, skewness, kurtosis, etc., and the statistical characteristics are updated in real time using the sliding window method. Then the similarity of the statistical characteristics is calculated, the cosine similarity is used to calculate the similarity of the numerical characteristics, and the mutual information is used to calculate the correlation of the categorical characteristics. Finally, the data distribution matching coefficient is calculated based on the similarity of each feature. The value range is between 0 and 1. A value greater than 0.7 indicates that the data distribution is highly similar, between 0.4 and 0.7 indicates moderate similarity, and less than 0.4 indicates low similarity. The purpose of this step is to evaluate the consistency of data distribution between adjacent nodes and provide a basis for the optimization of the microservice system.

[0061] The specific implementation method of step S10 is to optimize the data exchange scheme of the microservice system. First, a data routing strategy is constructed based on the data distribution matching coefficient, a direct data channel is established between nodes with high matching coefficients, and an intermediate cache mechanism is adopted between nodes with low matching coefficients. Then, the data transmission batch is optimized, and the data packaging strategy is dynamically adjusted according to the data similarity. Data with high similarity is transmitted in a centralized manner, and data with low similarity is transmitted separately. Finally, the data caching strategy is adjusted to set the cache update frequency according to the data consistency level between nodes. Nodes with high consistency reduce the update frequency, and nodes with low consistency increase the update frequency. This step aims to optimize data transmission efficiency and reduce system resource consumption.

[0062] The specific implementation method of step S11 is to extract data features through the feature extraction pyramid layer. First, the cleaned data is input into the multi-scale feature extraction module, and convolution kernels of different scales are used to extract multi-level features. Then, feature fusion is realized through the feature pyramid network, and a top-down feature fusion strategy is adopted to fuse feature maps of different scales layer by layer. Finally, the data contribution index is calculated, and the weights are 0.6 and 0.4 respectively, based on the comprehensive evaluation of the data usability index and the data distribution matching coefficient, and the value range is between 0 and 1. The purpose of this step is to extract the key features of the data and evaluate the contribution of the data to the system.

[0063] The specific implementation method of step S12 is to formulate a microservice resource allocation strategy. First, the resource allocation weight is determined according to the data contribution index. Nodes with high contribution obtain more computing resources, and nodes with low contribution obtain basic computing resources. Then set resource quotas for different types of microservices. Data collection microservices and data cleaning microservices give priority to resource allocation. Data conversion microservices and data storage microservices dynamically adjust resources according to the load, and data distribution microservices maintain a stable resource level. Finally, dynamic resource scheduling is implemented to monitor the resource utilization of each microservice. When the utilization exceeds 80%, the capacity is expanded, and when it is less than 30%, the capacity is reduced. This step aims to achieve reasonable allocation of computing resources and ensure the operating efficiency of the system.

[0064] The specific implementation method of step S13 is to store data feature vectors through data distribution microservices. First, the feature vectors are organized according to the time series, and the time series database is used to store the feature data with timestamps. Then a distributed storage strategy is implemented, and shard storage is performed according to the time characteristics and spatial characteristics of the data. Hot data is stored in high-performance storage nodes, and cold data is migrated to archive storage nodes. Finally, a data indexing mechanism is established, and multi-level indexes are established for the time dimension and feature dimension to support efficient data retrieval and query. The purpose of this step is to achieve efficient storage and fast access to data, and provide support for subsequent data applications.

[0065] The calculation process or mathematical model involved in the present invention is described in detail below.

[0066] The basic evaluation function of the data availability index is specifically expressed as follows:

[0067]

[0068] In the formula, x i is each evaluation component; w i is the corresponding weight coefficient; n is the evaluation dimension, which is 3; ε 1 is the basic error term, and its value range is 0.01~0.05.

[0069] The evaluation function after considering the influence of time decay is expressed as:

[0070]

[0071] Where t is the current timestamp; t 0 Generate a timestamp for the data; λ is the time-effect decay coefficient, with a value range of 0.1 0.5; ε 2为时间误差项,取值范围为0.02 0.06.

[0072] The data availability index after the introduction of logarithmic transformation is finally expressed as:

[0073] U=ln(1+g(x,t))+ε 3 ;

[0074] Where U is the data availability index; ε 3 is the comprehensive error term, and its value range is 0.03~0.07.

[0075] The data integrity component calculation matrix is ​​specifically expressed as follows:

[0076]

[0077] In the formula, c ij represents the completeness score of the jth feature of the ith data block; m is the number of data blocks; and n is the feature dimension.

[0078] The calculation formula for a single completeness score is:

[0079]

[0080] Where N mij is the number of missing values; N aij is the number of outliers; N tij is the total amount of data; α ij , β ij is the penalty coefficient, and its value range is 0.6 0.8和0.4 0.6; ε cij is the error term, and its value range is 0.01~0.03.

[0081] The partial derivative of the completeness score is expressed as:

[0082]

[0083] The calculation of the data accuracy component is specifically expressed as follows:

[0084]

[0085] In the formula, x ij is the measured value; is the standard value; R ij is the measuring range; ε aij is the error term, and its value range is 0.02~0.04.

[0086] The matrix representation of the accuracy component is:

[0087]

[0088] The calculation of the timeliness component is specifically expressed as follows:

[0089]

[0090] In the formula, λ ij is the time-dependent attenuation coefficient; t ij is the current timestamp; t 0ij Generate timestamp for data; tij is the error term, and its value range is 0.01~0.03.

[0091] Parameter acquisition method description:

[0092] 1. Number of missing values ​​M mij and the number of outliers N aij Obtained through data scanning statistics;

[0093] 2. Standard value It is obtained through the equipment calibration curve. The calibration steps include: running the equipment without load, loading the standard working condition, recording the output response, and fitting the calibration curve;

[0094] 3. Measuring range R ij Determined by the equipment specifications;

[0095] 4. Time parameters are obtained through the system clock;

[0096] 5. The penalty coefficient and attenuation coefficient are obtained through historical data optimization.

[0097] Explanation of the equation construction principle:

[0098] 1. Use exponential function to describe the decay of timeliness, which is consistent with the characteristic that the value of data decreases over time;

[0099] 2. Introduce logarithmic transformation to enhance model stability and avoid rapid growth of evaluation indicators;

[0100] 3. Use matrix representation to achieve unified processing of multi-dimensional features;

[0101] 4. Evaluate the sensitivity of the indicator to various factors through partial derivative analysis;

[0102] 5. Introduce multi-level error terms to achieve robust control.

[0103] The calculation of the data cleaning requirement coefficient is specifically expressed as follows:

[0104]

[0105] In the formula, D is the data cleaning demand coefficient; Q is the quality level coefficient; F is the update frequency coefficient; I is the business importance coefficient; α, β, γ are weight coefficients of 0.4, 0.3, and 0.3 respectively; δ is the adjustment coefficient, and its value range is -0.1 to 0.1.

[0106] The calculation formula of quality grade coefficient is:

[0107]

[0108] In the formula, q i is the basic coefficient of different quality levels, high quality is 0.3, medium quality is 0.5, and low quality is 0.7; w qi is the quality level weight; q is the error term, and its value range is 0.02~0.04.

[0109] The calculation formula for the update frequency coefficient is:

[0110]

[0111] In the formula, f i is the base coefficient for different update frequencies, 0.4 for update frequencies above 1 per second, 0.3 for update frequencies between 1 per second and 1 per minute, and 0.2 for update frequencies below 1 per minute; fi is the frequency weight; f is the error term, and its value range is 0.01~0.03.

[0112] The calculation formula of business importance coefficient is:

[0113]

[0114] In the formula, i i is the basic coefficient of different importance, 0.5 for key business, 0.3 for important business, and 0.2 for general business; ii is the importance weight; i is the error term, and its value range is 0.02~0.04.

[0115] The calculation of the data distribution matching coefficient is specifically expressed as follows:

[0116]

[0117] In the formula, p(x i ,y j ) is the joint probability distribution; p(xi ), p(y j ) is the marginal probability distribution; H(X), H(Y) are the marginal entropies; ε m is the error term, and its value range is 0.01~0.03.

[0118] The calculation matrix of probability distribution is:

[0119]

[0120] In the formula, p ij Represents the feature (x i ,y j ) is the joint probability of .

[0121] The entropy calculation formula is:

[0122]

[0123] In the formula, ε h is the entropy calculation error, ranging from 0.01 to 0.02.

[0124] The calculation of the data contribution index is specifically expressed as follows:

[0125]

[0126] In the formula, C is the data contribution index; U is the data availability index; M is the data distribution matching coefficient; w 1 , w 2 , w 3 The weight coefficients are 0.4, 0.3, and 0.3 respectively; k is the nonlinear adjustment coefficient, and its value range is 0.5 1.5; η 为修正项,取值范围为- 0.05 0.05.

[0127] Parameter acquisition method description:

[0128] 1. The basic coefficients of quality level, update frequency and business importance are determined through expert evaluation;

[0129] 2. The weight coefficient is calculated by analytic hierarchy process;

[0130] 3. The probability distribution is obtained by kernel density estimation method;

[0131] 4. The nonlinear adjustment coefficient is determined through cross-validation optimization.

[0132] A second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions are run in a computer, they are used to execute the above-mentioned industrial Internet of Things data preprocessing method based on microservices.

[0133] The third aspect of the present invention provides an industrial Internet of Things data preprocessing system based on microservices, comprising the above-mentioned computer-readable storage medium, the system is any one of a computer, a server, and a single-chip microcomputer, the computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

[0134] Specifically, the principle of the present invention is: The technical solution of the present invention is based on the following principles: First, the data association relationship between industrial equipment is captured through graph network modeling. This representation method can effectively reflect the spatial distribution characteristics and propagation paths of data. Based on the representation of the graph network, the importance and scope of influence of data in the system can be more accurately evaluated. Secondly, the design of the three-layer data pyramid model follows the progressive principle of data processing, from data classification, cleaning to feature extraction, layer by layer, and each layer is optimized for specific processing tasks.

[0135] In the data classification pyramid layer, the weight adaptive mechanism is adopted to dynamically adjust the classification weight according to the actual characteristics of the data, thereby improving the accuracy of classification. The residual retention mechanism of the data cleaning pyramid layer ensures that the essential characteristics of the data are retained while removing noise. The feature fusion processing layer of the feature extraction pyramid layer provides richer data expression through the combination of multi-dimensional features. This hierarchical processing structure not only improves the processing accuracy, but also enhances the robustness of the system.

[0136] The four key indicators introduced in this paper (data availability index, data cleaning requirement coefficient, data distribution matching coefficient, and data contribution index) constitute a complete data evaluation system. These indicators describe the characteristics and value of data from different perspectives and provide a scientific basis for resource allocation. In addition, the adoption of microservice architecture provides good scalability and flexibility, enabling the system to dynamically adjust processing capabilities according to actual needs.

[0137] A specific embodiment 1 of the present invention is provided below, and the specific implementation method of each step in this embodiment 1 is described in detail as follows.

[0138] The specific implementation method of step S01 is to build an industrial Internet of Things microservice system based on a cloud native architecture. The entire system adopts a containerized deployment method to achieve independent operation and elastic scaling of each microservice module. The data acquisition microservice is responsible for collecting equipment data through communication methods such as industrial Ethernet and fieldbus, and supports multiple industrial communication protocols including message queue telemetry transmission protocol and object linking and embedding protocol. The collection cycle is dynamically adjusted according to the type of equipment. A 100-millisecond sampling cycle is used for high-speed equipment and a 1-second sampling cycle is used for low-speed equipment. The collected data is streamed through the message queue after being cached locally. The data cleaning microservice pre-processes the original data through methods such as data filtering and anomaly detection, and uses a sliding window algorithm to achieve real-time data filtering. The window size is adaptively adjusted according to the data collection frequency. For data streams with a collection frequency of one second, the window size is set to 60 data points. The data conversion microservice completes standardization processing such as data format conversion and unit unification, establishes a unified data model for heterogeneous data from different sources, and realizes data format conversion through configurable mapping. The data storage microservice uses a distributed database to achieve data persistence storage, selects different types of storage engines according to data access characteristics, uses a memory database for hot data, and uses a document database for cold data. The data distribution microservice implements real-time data distribution based on the publish-subscribe model, supports multiple subscription methods and multi-level caching mechanisms, and optimizes transmission efficiency through message filtering and compression algorithms. Service discovery and load balancing are implemented between microservices through service grids, and asynchronous communication is implemented using message middleware to ensure high availability and scalability of the system. In order to improve the reliability of the system, each microservice deploys multiple instances, and health checks and automatic recovery mechanisms are used to ensure the continuous availability of services. At the same time, protection mechanisms such as circuit breakers, current limiting, and downgrades are introduced to prevent service avalanches. The purpose of this step is to build a microservice architecture with good elasticity and scalability to provide basic support for subsequent data processing.

[0139] The specific implementation of step S02 is to construct an industrial Internet of Things data network model based on graph theory. First, the topological relationship of the equipment on the industrial site is analyzed, and each industrial equipment is used as the vertex of the graph network. The adjacency matrix is ​​used to describe the connection relationship between the equipment. The equipment node matrix is ​​expressed as: V = {v 1 , v 2 , …, v n}, the edge set is represented as: E = {e ij |i, j = 1, 2, ..., n}, where e ij Represents the connection relationship between node i and node j. Through the depth-first search algorithm, the device network is traversed to establish the adjacency matrix A. The matrix element a ij Indicates the connection status between node i and node j, with a value of 0 or 1. Based on the data flow between devices, determine the directed edges in the graph network and the edge weight wij Calculated by data transmission frequency: where f ij is the data transmission frequency between nodes, ε w is a weight correction term with a value range of 0.01 to 0.05. The minimum spanning tree algorithm is used to optimize the network structure to ensure the optimality of the data transmission path. The device nodes are clustered based on the community discovery algorithm to form a hierarchical graph network. The objective function of community division is: where k i is the degree of node i, m is the total number of edges, δ(c i , c j ) is an indicator function for determining whether a node belongs to the same community. This step aims to construct a graph network model that reflects the physical connection and logical relationship of industrial equipment, and provide network topology support for subsequent data processing.

[0140] The specific implementation of step S03 is to obtain multi-dimensional operation data of industrial equipment through data acquisition microservices. First, configure the data acquisition parameters, including sampling period, sampling accuracy, communication timeout, etc. The sampling period is set between 100 milliseconds and 1 second according to the device type. For the equipment operation status data, the polling method is used to obtain the working status identification, record the equipment startup time, collect power, speed, vibration, temperature and other parameters, and the reliability evaluation index of data acquisition is: Where N s is the number of successful acquisitions, N t is the total number of acquisitions, N e is the number of collection errors. For equipment environmental status data, the sensor network collects parameters such as ambient temperature, humidity, dust concentration, noise, and light intensity. For equipment process parameter data, real-time collection of production line speed, process temperature, pressure, flow, quality and other parameters. A data cache mechanism is used to handle collection delays, set data collection priorities, and ensure the timeliness of key parameter collection. The evaluation function of the cache strategy is: C = w 1 L+w 2 T+w 3 M, where L is the cache capacity, T is the access latency, M is the memory usage, and w 1 , w 2 , w 3 is the weight coefficient. The purpose of this step is to obtain complete industrial equipment operation data and provide a data source for subsequent data processing.

[0141] The specific implementation method of step S04 is to construct a three-layer data pyramid model. First, an improved classification neural network is implemented in the data classification pyramid layer, and a convolutional neural network is used to extract data features. Three convolution layers are set, and the convolution kernel sizes are 3×3, 5×5, and 7×7 respectively. The activation function uses the modified linear unit function: f(x)=max(0,x). A data weight adaptation layer is added between the fully connected layer and the output layer, and the attention mechanism is used to dynamically adjust the feature weight. The attention weight calculation formula is: where e i is the feature importance score. Then, an improved encoding and decoding network is implemented in the data cleaning pyramid layer. The encoding unit uses a stacked autoencoder, and the decoding unit uses a deconvolution network. A data residual retention layer is added in the middle to retain the original data features. The residual calculation formula is: Where x is the input data, To reconstruct the data. An improved deep feature network is implemented in the feature extraction pyramid layer, a multi-scale feature extraction module is used, and feature fusion is achieved through the feature pyramid network. This step aims to build an efficient data processing model to achieve data classification, cleaning and feature extraction.

[0142] The specific implementation of step S05 is to calculate the data availability index. First, an initial evaluation model is constructed through the basic evaluation function: The evaluation function after considering the influence of time decay is: The final data availability index is expressed as: The data completeness components are calculated through the matrix C, and a single completeness score is calculated as: The data accuracy component is calculated as: The timeliness component is calculated as: The purpose of this step is to evaluate the level of data availability and provide a basis for subsequent data classification.

[0143] The specific implementation of step S06 is to use the data classification pyramid layer to classify the quality of the original data. According to the calculated data usability index U, combined with the quality threshold, the data is divided into different quality levels. The threshold for high-quality data is 0.8, and the threshold for medium-quality data is 0.5. The classification process adopts an improved neural network structure, and dynamically adjusts the feature weights through the data weight adaptive layer. The weight update formula is: Where η is the learning rate and L is the loss function. Finally, the softmax classifier outputs the data quality level identification to achieve data quality grading. This step aims to achieve data quality grading and provide a basis for subsequent differentiated processing.

[0144] The specific implementation of step S07 is to calculate the data cleaning requirement coefficient. The calculation of the data cleaning requirement coefficient adopts a nonlinear mapping model: The quality grade coefficient is calculated as: The update frequency coefficient is calculated as: The business importance coefficient is calculated as: The parameter optimization adopts the gradient descent method, the loss function is the mean square error, and the optimization is iterated until convergence. The purpose of this step is to determine the priority of data cleaning and achieve the reasonable allocation of cleaning resources.

[0145] The specific implementation of step S08 is to perform data cleaning through the data cleaning pyramid layer. A lightweight cleaning strategy is used for high-quality data groups, and outliers are identified by the triple standard deviation method: The outlier determination conditions are: For medium-quality data sets, a standard cleaning strategy is used, and missing values ​​are filled using time series interpolation: t =αx t-1 +(1-α)x t+1 +ε t , where α is the smoothing coefficient, ranging from 0.3 to 0.7, ε t is a random perturbation term. For low-quality data sets, a deep cleaning strategy is adopted. Data reconstruction is based on a deep learning model, and the loss function includes reconstruction error and regularization term: Where λ is the regularization coefficient and w is the model parameter. Feature restoration uses correlation analysis to calculate the mutual information between features: This step aims to achieve differentiated data cleaning according to the data quality level and improve data quality.

[0146] The specific implementation method of step S09 is to calculate the data distribution matching coefficient. The calculation of the matching coefficient is based on information theory and uses the normalized mutual information metric: The probability distribution is obtained by kernel density estimation, and the marginal entropy is calculated as: To improve computational efficiency, a sliding window is used to update the probability distribution, and the window size is dynamically adjusted according to the data change rate. The purpose of this step is to evaluate the consistency of data distribution between adjacent nodes and provide a basis for the optimization of the microservice system.

[0147] The specific implementation of step S10 is to optimize the data exchange scheme of the microservice system. A routing matrix is ​​constructed according to the data distribution matching coefficient: R = {r ij} n×n , where r ij It represents the routing weight from node i to node j, and the calculation formula is: The data transmission batch optimization objective function is: J = w 1 T+w 2 B+w 3L, where T is the transmission delay, B is the bandwidth usage, and L is the load balancing factor. The update frequency of the cache strategy is determined according to the consistency requirements: Where k is the adjustment coefficient, M 0 is the consistency threshold. This step aims to optimize data transmission efficiency and reduce system resource consumption.

[0148] The specific implementation of step S11 is to extract data features through the feature extraction pyramid layer. Multi-scale feature extraction uses convolution kernels of different scales, and the feature map calculation formula is: l =σ(W l *F l-1 +b l ), where W l is the convolution kernel weight, b l is the bias term, and σ is the activation function. Feature fusion adopts a top-down strategy, and the fusion function is: where w i is the feature weight. The data contribution index is calculated as: The purpose of this step is to extract the key features of the data and evaluate the contribution of the data to the system.

[0149] The specific implementation method of step S12 is to formulate a microservice resource allocation strategy. Calculate the resource allocation weight based on the data contribution index: The resource quota optimization model is: Where R i To actually allocate resources, is the ideal resource demand. The dynamic scheduling threshold is set as follows: resource utilization exceeding 80% triggers expansion, and less than 30% triggers contraction. The expansion and contraction rate control function is: Where η is the current utilization rate, η 0 This step aims to achieve a reasonable allocation of computing resources and ensure the operating efficiency of the system.

[0150] The specific implementation of step S13 is to store the data feature vector through the data distribution microservice. The distribution measurement commonly used in the prior art can be used. It can also be implemented by the following method: the time series data organization adopts a sharding strategy, and the sharding function is: s = h(t) mod n, where h(t) is the timestamp hash value and n is the number of shards. The storage node selection is based on the load balancing algorithm: L i =αU i +βT i +γS i , where U i is the CPU usage, T i is the throughput, S iThe index is constructed using a B+ tree structure, and the search complexity is O(logn). The purpose of this step is to achieve efficient storage and fast access to data, and to provide support for subsequent data applications.

[0151] In order to better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention: There are 50 CNC machine tools in an industrial park. The equipment collection parameters include equipment operation status data, environmental status data and process parameter data. The data collection frequency is once per second. The research team first counted the basic characteristics of various types of data, as shown in Table 1:

[0152] Table 1 Basic characteristic statistics of data acquisition parameters

[0153]

[0154] Figure 2 The analysis of the acquisition parameter characteristics of different types of data (operating status, environmental status, process parameters) is shown, including the comparison of missing rate, abnormal rate and data volume. The research team analyzed the physical connection relationship between devices, constructed an industrial Internet of Things data graph network, and calculated the data transmission frequency between device nodes, as shown in Table 2:

[0155] Table 2 Device node data transmission frequency matrix (unit: times / second)

[0156] Node Number Node 1 Node 2 Node 3 Node 4 Node 5 Node 1 0 0.8 0.5 0.3 0.2 Node 2 0.8 0 0.7 0.4 0.3 Node 3 0.5 0.7 0 0.6 0.4 Node 4 0.3 0.4 0.6 0 0.5 Node 5 0.2 0.3 0.4 0.5 0

[0157] Figure 3 The heat map of data transmission frequency between device nodes is displayed, and the color depth indicates the size of the transmission frequency. Based on the data transmission frequency, the edge weight matrix is ​​calculated, and the minimum spanning tree algorithm is used to optimize the network structure. In the optimized network structure, the device nodes are divided into three communities, and the data distribution similarity of the nodes within the community is high. The research team evaluated the data quality of each node and calculated the data availability index. The evaluation results are shown in Table 3:

[0158] Table 3 Evaluation results of device node data availability index

[0159] Evaluation Metrics Completeness Component Accuracy Component Timeliness Component Availability Index Node 1 0.95 0.92 0.88 0.92 Node 2 0.93 0.90 0.85 0.89 Node 3 0.88 0.85 0.82 0.85 Node 4 0.82 0.80 0.78 0.80 Node 5 0.78 0.75 0.72 0.75

[0160] Figure 4 The data availability index analysis of each device node is shown, including the comparison of the completeness component, accuracy component and timeliness component. According to the data availability index, the data is divided into a high-quality data group (availability index ≥ 0.85), a medium-quality data group (0.85> availability index ≥ 0.75) and a low-quality data group (availability index < 0.75). Differentiated cleaning strategies are adopted for data of different quality levels. The data cleaning effect evaluation results are shown in Table 4:

[0161] Table 4 Data cleaning effect evaluation results

[0162]

[0163] Figure 5 The cleaning effect comparison of different quality data groups is shown, including the changes in anomaly rate and missing rate before and after cleaning. The research team calculated the data distribution matching coefficient between device nodes to evaluate the consistency level of data distribution. The matching coefficient calculation uses the mutual information formula: The calculation results are shown in Table 5:

[0164] Table 5. Device node data distribution matching coefficient matrix

[0165] Node Number Node 1 Node 2 Node 3 Node 4 Node 5 Node 1 1.00 0.85 0.72 0.65 0.58 Node 2 0.85 1.00 0.78 0.68 0.62 Node 3 0.72 0.78 1.00 0.75 0.68 Node 4 0.65 0.68 0.75 1.00 0.82 Node 5 0.58 0.62 0.68 0.82 1.00

[0166] Based on the data distribution matching coefficient, the research team optimized the data exchange scheme of the microservice system and monitored the system resource usage. The resource usage during system operation is shown in Table 6:

[0167] Table 6 Statistics of microservice system resource usage

[0168]

[0169] After a three-month system operation test, the solution has achieved remarkable results in data processing efficiency, resource utilization and system reliability. Figure 6 The resource usage of the microservice system is demonstrated, including CPU usage, memory usage, network bandwidth occupancy and average response time. Compared with the traditional centralized data processing solution, the present invention has the following advantages: 1) The traditional solution adopts a fixed data processing process and cannot perform differentiated processing according to data quality. The present invention realizes adaptive data processing through a three-layer data pyramid model, and the data processing efficiency is improved by 35%. 2) The traditional solution uses a unified data cleaning strategy, which seriously wastes computing resources. The present invention adopts a differentiated cleaning strategy according to the data quality level, and the resource utilization rate is improved by 42%. 3) The traditional solution lacks a data distribution consistency evaluation mechanism, which easily leads to the accumulation of data deviations. The present invention realizes dynamic evaluation and adjustment of data consistency through the data distribution matching coefficient, and the system reliability is improved by 28%. 4) The microservice resource allocation of the traditional solution adopts static configuration, which is difficult to cope with load changes. The present invention realizes dynamic adjustment of resources through data contribution indicators, and the system stability is improved by 45%.

[0170] A specific embodiment 3 of the present invention is provided below: an industrial Internet of Things data preprocessing system based on microservices is implemented on a certain intelligent production line. The production line includes three core processes: ironmaking, steelmaking, and steel rolling, with a total of 80 equipment nodes. In the early stage of system deployment, a data availability index evaluation model is first constructed. The basic evaluation function is: The evaluation function after considering the influence of time decay is: The final data availability index is expressed as: U = ln(1 + g(x, t)) + ε 3 The data collection parameter statistics of each process equipment are shown in Table 7:

[0171] Table 7 Statistics of data collection parameters for each process

[0172]

[0173]

[0174] An improved convolutional neural network is used to extract data features, and the convolution kernel weight matrix is ​​expressed as: The data quality evaluation results of each process are shown in Table 8:

[0175] Table 8 Process data quality assessment results

[0176] Evaluation Metrics Completeness Component Accuracy Component Timeliness Component Availability Index Quality Grade Ironmaking process 0.92 0.88 0.85 0.88 high quality Steelmaking process 0.85 0.82 0.80 0.82 Medium quality Steel rolling process 0.78 0.75 0.72 0.75 Medium quality

[0177] Figure 7 The data quality assessment results of each process are shown, including the completeness component, accuracy component, timeliness component and usability index. According to the data quality assessment results, the data cleaning requirement coefficient is calculated: The processing results of data of different quality levels are shown in Table 9:

[0178] Table 9 Data processing effect comparison table

[0179]

[0180] In the feature extraction process, a multi-scale feature extraction module is used, and the feature map calculation formula is: F l =σ(W l *F l-1 +b l ). Feature fusion adopts a top-down strategy, and the fusion function is: The calculation results of the data distribution matching coefficient between processes are shown in Table 10:

[0181] Table 10 Data distribution matching coefficient matrix between processes

[0182] Process Name Ironmaking process Steelmaking process Steel rolling process Ironmaking process 1.00 0.82 0.65 Steelmaking process 0.82 1.00 0.78 Steel rolling process 0.65 0.78 1.00

[0183] The microservice system adopts a dynamic resource allocation strategy, and the resource allocation weight is calculated as: The data contribution index C is calculated by the formula: The system operation performance indicators are shown in Table 11:

[0184] Table 11 System operation performance index statistics

[0185]

[0186] After 6 months of system operation verification, compared with the traditional batch processing solution, the present invention has obvious advantages in the following aspects: 1) By introducing data availability indicators and data cleaning requirement coefficients, accurate evaluation and hierarchical processing of data quality are achieved, and data processing efficiency is improved by 58%. 2) Feature extraction based on the improved deep learning model improves feature expression ability by 42% and model convergence speed by 35%. 3) The data distribution matching coefficient is used to guide data exchange optimization, system throughput is improved by 65%, and resource utilization is improved by 73%. 4) The dynamic resource allocation strategy enables the system to have stronger adaptive capabilities, and in the case of load mutations, the system stability is improved by 68%. 5) Compared with the fixed data processing flow of the traditional solution, the adaptive processing mechanism of the present invention improves the accuracy by 45% and the processing efficiency by 62% when processing industrial data with high noise and high missing rate.

[0187] It should be noted that the variables involved in the present invention are explained in detail as shown in Table 12 below.

[0188] Table 12 Variable explanation table

[0189]

[0190]

[0191] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A microservice-based industrial Internet of Things data preprocessing method, characterized in that: The following steps are involved: Build an industrial Internet of Things microservice system including data collection microservice, data cleaning microservice, data conversion microservice, data storage microservice and data distribution microservice; input the original data into a pre-trained three-layer data pyramid model, the three-layer data pyramid model includes a data classification pyramid layer, a data cleaning pyramid layer and a feature extraction pyramid layer, wherein the data classification pyramid layer adopts an improved classification neural network structure, and a data weight adaptation layer is set between the fully connected layer and the output layer of the classification neural network; the data cleaning pyramid layer adopts an improved encoding and decoding network structure, and a data residual retention layer is set between the encoding unit and the decoding unit; the feature extraction pyramid layer adopts an improved deep feature network structure, and a feature fusion processing layer is set between the feature dimensionality reduction layer and the feature integration layer; classify and clean the data and extract features based on data availability index, data cleaning requirement coefficient and data distribution matching coefficient.

2. The industrial Internet of Things data preprocessing method according to claim 1 is characterized in that: The steps of constructing an industrial Internet of Things data graph network include: using industrial equipment nodes as graph network nodes and data transmission relationships between devices as graph network edges to form an Internet of Things graph structure; using a depth-first search algorithm to traverse the device network and establish an adjacency matrix to represent the connection relationship between devices; using a minimum spanning tree algorithm to optimize the network structure to ensure the optimality of the data transmission path; and clustering device nodes based on a community discovery algorithm to form a hierarchical graph network.

3. The industrial Internet of Things data preprocessing method according to claim 1 is characterized in that: The data availability index is calculated through data integrity component, data accuracy component and data timeliness component. The weight coefficient of the data integrity component is 0.4, the weight coefficient of the data accuracy component is 0.35, and the weight coefficient of the data timeliness component is 0.

25. The weighted sum of the three components is used as the data availability index, and the value range is between 0 and 1.

4. The industrial Internet of Things data preprocessing method according to claim 1, characterized in that: The data cleaning requirement coefficient is calculated based on the data quality level identifier, the data update frequency value, and the business importance level identifier. The basic coefficient corresponding to the data quality level identifier is: 0.3 for the high-quality data group, 0.5 for the medium-quality data group, and 0.7 for the low-quality data group; the weight corresponding to the data update frequency value is: 0.4 for an update frequency of more than once per second, 0.3 for once per second to once per minute, and 0.2 for less than once per minute; the weight corresponding to the business importance level identifier is: 0.5 for critical business, 0.3 for important business, and 0.2 for general business.

5. The industrial Internet of Things data preprocessing method according to claim 1, characterized in that: The data distribution matching coefficient is obtained by calculating the data statistical characteristics of adjacent industrial equipment nodes, using cosine similarity to calculate the similarity of numerical features, and using mutual information to calculate the correlation of categorical features. The data distribution matching coefficient ranges from 0 to 1.

6. The industrial Internet of Things data preprocessing method according to claim 1, characterized in that: The data cleaning pyramid layer adopts different cleaning strategies for high-quality data groups, medium-quality data groups, and low-quality data groups: a lightweight cleaning strategy is adopted for the high-quality data group, which mainly performs outlier detection and simple filtering; a standard cleaning strategy is adopted for the medium-quality data group, including missing value filling, outlier processing, and noise elimination; a deep cleaning strategy is adopted for the low-quality data group, including data reconstruction, feature repair, and consistency check.

7. The industrial Internet of Things data preprocessing method according to claim 1, characterized in that: The feature extraction pyramid layer extracts multi-level features through a multi-scale feature extraction module. A feature pyramid network is used to achieve feature fusion, and a top-down feature fusion strategy is adopted to fuse feature maps of different scales layer by layer.

8. The industrial Internet of Things data preprocessing method according to claim 1, characterized in that: A microservice resource allocation strategy is established based on the data contribution index, where the data contribution index is obtained through a comprehensive evaluation of the data availability index and the data distribution matching coefficient, with weights of 0.6 and 0.4 respectively. The resource utilization of each microservice is monitored, and capacity is expanded when the utilization exceeds 80%, and capacity is reduced when it is less than 30%.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and when the program instructions are executed in a computer, they are used to execute the microservice-based industrial Internet of Things data preprocessing method according to any one of claims 1 to 8.

10. A microservice-based industrial Internet of Things data preprocessing system, characterized in that: The system comprises the computer-readable storage medium as claimed in claim 9, wherein the system is any one of a computer, a server, and a single-chip microcomputer, the computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.