Micro-service-based industrial data quality evaluation method, medium and system
By building a multi-level evaluation system, including distributed data acquisition networks and multi-layer neural network evaluation models, the problem of difficulty in adapting to the dynamic changes of multi-source heterogeneous data in the existing technology is solved, and accurate dynamic evaluation of industrial data quality is achieved.
Patent Information
- Application Number
- CN202510241302.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to adapt to the dynamic changes of multi-source heterogeneous data, resulting in insufficient accuracy and reliability of industrial data quality assessment.
Using the industrial data quality evaluation method based on microservices, a multi-level evaluation system is built, including distributed data acquisition network, multi-layer neural network evaluation model and collaboration mechanism, to achieve dynamic evaluation of multi-source heterogeneous data quality.
Accurate and dynamic evaluation of industrial multi-source heterogeneous data quality is achieved, which enhances the adaptability to dynamic data changes and improves the accuracy and reliability of evaluation results.
Smart Images

Figure CN120145009A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electrical digital data processing, and more particularly, relates to an industrial data quality assessment method, medium, and system based on microservices. Background Art
[0002] The rapid development of the industrial Internet has promoted the in-depth application of intelligent manufacturing, and the multi-source heterogeneous data generated by industrial equipment has become the core asset of manufacturing enterprises. Traditional industrial data quality assessment methods mainly rely on statistical analysis and rule checking, including basic assessment means such as data integrity checking, outlier identification, and consistency verification. These methods usually adopt fixed assessment indicators and thresholds, and judge data quality through static analysis of data characteristics. In practical applications, traditional assessment methods have achieved certain results in the quality assessment of single data sources, data anomaly detection, data consistency verification, etc.
[0003] However, with the improvement of the intelligence level of industrial equipment, the types of data generated during the production process are becoming increasingly rich, the data scale is continuously expanding, and the data characteristics show obvious dynamic change characteristics. Traditional static assessment methods are difficult to accurately capture the dynamic change rules of data quality and cannot effectively identify the trend characteristics of data quality evolution over time. Especially in the scenario of multi-source heterogeneous data, there are complex correlation relationships between different data sources, and the assessment of data quality needs to comprehensively consider the mutual influence between data. In addition, the fixed assessment models adopted by traditional methods lack adaptability and are difficult to dynamically adjust assessment strategies according to changes in data characteristics, resulting in insufficient accuracy and reliability of assessment results.
[0004] Facing the complex and changeable data environment in the industrial field, how to construct a quality assessment method that can adapt to the dynamic characteristics of data and achieve accurate assessment of multi-source heterogeneous data quality has become an urgent technical problem to be solved. In the prior art, rule-based assessment methods are difficult to cope with the dynamic changes of data characteristics, statistical-based assessment methods have strict assumptions about data distribution, and machine learning-based assessment methods have limitations such as insufficient model generalization ability and inability to fully utilize data time series characteristics. These technical means are difficult to effectively solve the problem of dynamic assessment of multi-source heterogeneous data quality. In summary, there is a technical problem in the prior art that industrial data quality assessment methods are difficult to adapt to the dynamic changes of multi-source heterogeneous data. Summary of the Invention
[0005] In view of this, the present invention provides an industrial data quality assessment method, medium, and system based on microservices, which can solve the technical problem in the prior art that industrial data quality assessment methods are difficult to adapt to the dynamic changes of multi-source heterogeneous data.
[0006] The present invention is implemented as follows: In a first aspect of the present invention, an industrial data quality assessment method based on microservices includes the following steps: collecting multi-source heterogeneous data of industrial equipment, establishing a distributed data collection network, and obtaining industrial data streams; constructing a data quality assessment index system, and establishing a data quality matrix including data deviation rate index, data completeness rate index, and data timeliness index; establishing a collaborative mechanism between the data quality assessment units, constructing a horizontal assessment network and a vertical assessment network, where the horizontal assessment network calculates the information gain value between different data sources, and the vertical assessment network calculates the correlation intensity value of the data sources; establishing a multi-layer neural network assessment model according to the data quality consistency index, where the multi-layer neural network assessment model includes a data encoding layer, a feature extraction layer, a pyramid processing layer, a feature fusion layer, and a pattern recognition layer, the pyramid processing layer includes three pyramid structures, the first pyramid structure includes an improved residual network structure, the second pyramid structure includes an improved attention network structure, and the third pyramid structure includes an improved dense connection network structure.
[0007] Among them, the specific structure of the first pyramid structure is composed of 3 residual blocks connected in series, each residual block contains two layers of convolution, batch normalization, and activation function, the batch normalization module is used to normalize the feature distribution and improve the stability of network training, and the data flow transmitted inside the first pyramid structure is the fusion feature obtained by adding the original feature after residual mapping and the shortcut connection.
[0008] Among them, the specific structure of the second pyramid structure includes a dual attention mechanism of a spatial attention sub-network and a channel attention sub-network, the channel attention module is used to learn the dependence relationship between feature channels and dynamically adjust the channel weights, and the data flow transmitted inside the second pyramid structure is the multi-scale feature map weighted by spatial and channel attention.
[0009] Among them, the specific structure of the third pyramid structure is a multi-path network with cross-layer feature reuse, including 4 dense blocks, the feature selection module is used to screen and combine multi-layer features and reduce redundant information, and the data flow transmitted inside the third pyramid structure is the dense connection combination of different scale features.
[0010] Among them, the relationship between the parameters in the basic assessment equation is: the basic quality index is equal to the product of the data deviation rate index and the first weight coefficient, plus the product of the data completeness rate index and the second weight coefficient, plus the product of the data timeliness index and the third weight coefficient, plus the product of the data fluctuation coefficient matrix and the fourth weight coefficient, and the sum of the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient is equal to 1.
[0011] Among them, the relationships of the parameters in the association evaluation equation are as follows: the association quality index is equal to the product of the information gain value and the first association coefficient, plus the product of the association strength value and the second association coefficient, plus the product of the natural logarithm value of the data quality consistency index and the third association coefficient, and the sum of the first association coefficient, the second association coefficient, and the third association coefficient is equal to 1.
[0012] Among them, the relationships of the parameters in the comprehensive evaluation equation are as follows: the comprehensive quality index is equal to the product of the partial derivative of the basic quality index with respect to time and the first comprehensive coefficient, plus the product of the association quality index and the second comprehensive coefficient, and the sum of the first comprehensive coefficient and the second comprehensive coefficient is equal to 1.
[0013] Among them, a data quality scoring model is constructed based on the weighted fusion result, and the output result of the comprehensive evaluation equation is mapped to the interval of 0 to 100 points through a normalization function. The normalization function uses a logarithmic form for numerical mapping to ensure the smoothness and discrimination of the scoring result. According to the mapped score, the data quality is divided into the following levels: 90 to 100 points represents high-quality data, 80 to 89 points represents good data, 70 to 79 points represents qualified data, and below 70 points represents unqualified data.
[0014] The second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions run on a computer, they are used to execute the above-mentioned industrial data quality evaluation method based on microservices.
[0015] The third aspect of the present invention provides an industrial data quality evaluation system based on microservices, which includes the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.
[0016] Compared with the prior art, the present invention provides an industrial data quality evaluation method, medium, and system based on microservices. The present invention proposes an industrial data quality evaluation method based on microservices, which realizes the dynamic evaluation of industrial multi-source heterogeneous data quality by constructing a multi-level evaluation system. The method uses a distributed data acquisition network to obtain multi-source data streams, establishes an evaluation system including indicators such as data deviation rate, completeness rate, and timeliness, and introduces a deep learning model to extract data features, and quantifies the dynamic characteristics of data through indicators such as feature entropy and fluctuation coefficient.
[0017] The method of the present invention breaks through the limitations of traditional evaluation methods. Through the parallel processing mechanism of the deep learning model, it captures the temporal and spatial features of data simultaneously, enhancing the adaptability to the dynamic changes of data. Multiple evaluation units are deployed using a microservices architecture, and the correlation relationships between data sources are analyzed through horizontal and vertical evaluation networks, constructing a dynamic evaluation model based on state transition. In particular, an improved residual network, attention mechanism, and dense connection structure are introduced in the pyramid processing layer, improving the accuracy of feature extraction and the adaptive ability of the model. By introducing a data quality evaluation equation set, the mathematical basis for data quality evaluation is established, realizing the interpretability of evaluation results.
[0018] The present invention successfully solves the technical problems of dynamic evaluation of industrial data quality, which are mainly reflected in the following aspects: First, through the feature extraction ability of the deep learning model, the accurate capture of the dynamic features of data is realized; Second, through the collaborative mechanism of the multi-layer evaluation system, the processing ability of multi-source heterogeneous data is improved; Third, through the improved neural network structure, the adaptive ability of the model is enhanced. The method of the present invention can not only accurately evaluate the current state of data quality, but also predict the change trend of data quality, providing effective technical support for industrial data quality management and solving the technical problem that the existing industrial data quality evaluation methods are difficult to adapt to the dynamic changes of multi-source heterogeneous data. Brief Description of the Drawings
[0019] Figure 1 It is a flowchart of the method of the present invention.
[0020] Figure 2 It is a graph of the performance monitoring data of the data acquisition node in Example 2 within 24 hours.
[0021] Figure 3 It is a graph of data during the 100-round training process of the deep learning model in Example 2.
[0022] Figure 4 It is a graph of the temporal change of the quality score within 30 days in Example 2. Detailed Embodiments
[0023] To make the purpose, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0024] As Figure 1 shown, it is a flowchart of a microservices-based industrial data quality evaluation method provided by the first aspect of the present invention. This method includes the following steps:
[0025] S01. Collect multi-source heterogeneous data of industrial equipment, establish a distributed data collection network, and obtain industrial data streams. The industrial data streams include production equipment status data, process parameter data, quality inspection data, and environmental monitoring data. The data collection is performed at a standard sampling period of 1 second, generating a sampling frequency value;
[0026] S02. Construct a data quality evaluation index system, and establish a data quality matrix including data deviation rate index, data completeness rate index, and data timeliness index. The data deviation rate index is calculated by the dispersion between the measured value and the standard value. The data completeness rate index is calculated by the ratio of the effective data volume to the theoretical data volume. The data timeliness index is calculated by the ratio of the data transmission delay time to the standard response time;
[0027] S03. Use a deep learning model to extract features from the industrial data streams, generating a feature probability distribution matrix. The deep learning model includes parallel long short-term memory network units and convolutional neural network units. The long short-term memory network units are used to extract temporal features, and the convolutional neural network units are used to extract spatial features;
[0028] S04. Calculate a data feature entropy matrix and a data fluctuation coefficient matrix based on the feature probability distribution matrix. The feature entropy matrix is used to measure the uncertainty degree of the data. The value range of the feature entropy matrix is from 0 to 1, where 0 represents complete certainty and 1 represents maximum uncertainty. The data fluctuation coefficient matrix is used to characterize the stability of the data;
[0029] S05. Deploy multiple data quality evaluation units using a microservices architecture. The data quality evaluation units construct a state transition matrix based on the time series of the data feature entropy matrix. The state transition matrix is used to describe the transition law of the data quality state;
[0030] S06. Establish a cooperation mechanism between the data quality evaluation units, and construct a horizontal evaluation network and a vertical evaluation network. The horizontal evaluation network calculates the information gain value between different data sources, and the vertical evaluation network calculates the correlation strength value of the data sources;
[0031] S07. Calculate a data quality consistency index based on the information gain value and the correlation strength value. The value range of the data quality consistency index is from -1 to +1, which is used to characterize the correlation degree between the data;
[0032] S08. Establish a multi-layer neural network evaluation model according to the data quality consistency index. The multi-layer neural network evaluation model includes a data encoding layer, a feature extraction layer, a pyramid processing layer, a feature fusion layer, and a pattern recognition layer;
[0033] S09. Use the attention mechanism to perform weighted fusion on the outputs of the three levels of the pyramid processing layer, calculate the importance of each layer of features through the soft attention weight matrix, and achieve adaptive feature selection;
[0034] S10. Build a data quality scoring model based on the weighted fusion result, map the output result of the comprehensive evaluation equation to the interval of 0 to 100 points through a normalization function. The normalization function uses a logarithmic form for numerical mapping to ensure the smoothness and discrimination of the scoring result. Divide the data quality into the following levels according to the mapped score: 90 to 100 points represent high-quality data, 80 to 89 points represent good data, 70 to 79 points represent qualified data, and below 70 points represent unqualified data;
[0035] The first-layer pyramid structure adopts an improved residual network structure, and a batch normalization module is added between the convolutional layer and the activation layer of the original residual network structure. The specific structure of the first-layer pyramid structure consists of 3 residual blocks connected in series, and each residual block contains two layers of convolution, batch normalization, and ReLU activation function;
[0036] The batch normalization module is used to normalize the feature distribution and improve the stability of network training. The data flow inside the first-layer pyramid structure transfers data as the fused feature obtained by adding the residual mapping of the original feature and the shortcut connection;
[0037] The second-layer pyramid structure adopts an improved attention network structure, and a channel attention module is added between the feature extraction layer and the fusion layer of the original attention network structure. The specific structure of the second-layer pyramid structure includes the dual attention mechanism of the spatial attention sub-network and the channel attention sub-network;
[0038] The channel attention module is used to learn the dependency relationship between feature channels and dynamically adjust the channel weights. The data flow inside the second-layer pyramid structure transfers data as the multi-scale feature map weighted by spatial and channel attention;
[0039] The third-layer pyramid structure adopts an improved dense connection network structure, and a feature selection module is added between the feature connection layer and the output layer of the original dense connection network structure. The specific structure of the third-layer pyramid structure is a multi-path network with cross-layer feature reuse, including 4 dense blocks;
[0040] The feature selection module is used to screen and combine multi-layer features and reduce redundant information. The data flow inside the third-layer pyramid structure transfers data as the dense connection combination of different-scale features;
[0041] The relationships of the parameters in the basic evaluation equation are described as follows:
[0042] The basic quality index is equal to the product of the data deviation rate index and the first weight coefficient, plus the product of the data completeness rate index and the second weight coefficient, plus the product of the data timeliness index and the third weight coefficient, plus the product of the data fluctuation coefficient matrix and the fourth weight coefficient, and the sum of the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient is equal to 1;
[0043] The relationships of the parameters in the correlation evaluation equation are described as follows:
[0044] The correlation quality index is equal to the product of the information gain value and the first correlation coefficient, plus the product of the correlation strength value and the second correlation coefficient, plus the natural logarithm value of the data quality consistency index and the third correlation coefficient, and the sum of the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient is equal to 1;
[0045] The relationships of the parameters in the comprehensive evaluation equation are described as follows:
[0046] The comprehensive quality index is equal to the product of the partial derivative of the basic quality index with respect to time and the first comprehensive coefficient, plus the product of the correlation quality index and the second comprehensive coefficient, and the sum of the first comprehensive coefficient and the second comprehensive coefficient is equal to 1;
[0047] The data deviation rate index characterizes the degree of dispersion of the measured value relative to the standard value and is used to evaluate the data accuracy;
[0048] The data completeness rate index characterizes the proportion of valid data in the total data and is used to evaluate the data integrity;
[0049] The data timeliness index characterizes the timeliness of data transmission and processing and is used to evaluate the data real-time performance;
[0050] The feature entropy matrix characterizes the degree of uncertainty of data distribution and is used to evaluate the fluctuation characteristics of data;
[0051] The data fluctuation coefficient matrix characterizes the degree of dispersion of data and is used to evaluate the data stability;
[0052] The information gain value characterizes the information correlation degree between data sources and is used to evaluate the correlation characteristics of data;
[0053] The correlation strength value characterizes the degree of dependence between data sources and is used to evaluate the coupling characteristics of data.
[0054] The specific implementation manners of the above steps are described in detail below. The specific implementation manner of step S01 is to construct an industrial data acquisition network based on the OPC UA protocol. By deploying data acquisition nodes on each production device, multi-source heterogeneous data including device status, process parameters, quality inspection, and environmental monitoring are collected. The data acquisition nodes collect data with a standard sampling period of 1 second, and at the same time, a data caching mechanism is set to ensure that data is not lost when the network fluctuates. The data acquisition nodes are connected to the edge server through the industrial Ethernet, and the edge server preprocesses the collected raw data, including data cleaning, format conversion, and timestamp synchronization. To ensure the reliability of data acquisition, the system also needs to configure a heartbeat detection mechanism. When a certain data acquisition node fails to upload data on time for 3 consecutive times, a fault alarm is automatically triggered. The purpose of this step is to establish a stable and reliable data acquisition infrastructure and provide a data source for subsequent data quality evaluation.
[0055] The specific implementation manner of step S02 is to construct a data quality evaluation system based on multi-dimensional metrics. First, the data deviation rate metric is calculated, which is obtained by calculating the root mean square error between the actual measurement value and the standard value. The maximum acceptable deviation rate is set to 5%. The data completeness rate metric is calculated by statistically analyzing the ratio of the number of valid data points actually obtained within a specified time window to the number of data points that should be obtained theoretically. The minimum completeness rate requirement is set to 95%. The data timeliness metric is calculated by measuring the ratio of the time delay from data generation to availability to the preset standard response time. The standard response time is set to 100 milliseconds. When the actual delay exceeds 2 times the standard response time, the timeliness metric will be significantly reduced. The purpose of this step is to establish a comprehensive data quality measurement standard and provide a basic metric system for quality evaluation.
[0056] The specific implementation manner of step S03 is to use a deep learning model to extract features from industrial data streams. This model adopts a parallel long short-term memory network and convolutional neural network structure. The long short-term memory network consists of 4 layers, with each layer containing 128 neurons, which is used to capture the temporal features of the data. The convolutional neural network uses 3 convolutional layers, and the sizes of the convolutional kernels are 3×3, 5×5, and 7×7 respectively, which are used to extract the spatial features of the data. The outputs of the two networks are merged through a feature fusion layer to generate a feature probability distribution matrix. The model is trained using the stochastic gradient descent algorithm, with the learning rate set to 0.001, the batch size set to 64, and the number of training epochs set to 100. The purpose of this step is to extract the deep features of the data through deep learning technology and provide feature support for subsequent quality evaluation.
[0057] The specific implementation of step S04 is to calculate the feature entropy matrix and the data fluctuation coefficient matrix based on the feature probability distribution matrix. The calculation of the feature entropy matrix uses the Shannon entropy formula, substituting each element in the feature probability distribution matrix for calculation to obtain an entropy value between 0 and 1. The data fluctuation coefficient matrix is obtained by calculating the ratio of the standard deviation to the average value of the data. The warning threshold for the fluctuation coefficient is set to 0.3, and exceeding this threshold indicates large data fluctuations. The calculation of the feature entropy matrix and the fluctuation coefficient matrix uses the sliding window method, with the window size set to 60 seconds and the step size set to 10 seconds. This step aims to quantify the uncertainty and stability characteristics of the data, providing an important basis for quality assessment.
[0058] The specific implementation of step S05 is to deploy the data quality assessment unit using a microservices architecture, with each assessment unit responsible for the quality assessment of a specific type of data. Based on the time series of the feature entropy matrix, the assessment unit constructs a state transition matrix using the Markov chain model, and the state space is defined as four levels: excellent, good, qualified, and unqualified. The state transition probability is calculated by statistically counting the frequency of state transitions at adjacent moments in historical data, and the time window is set to 24 hours. The assessment unit is deployed in a containerized manner, supporting horizontal scaling, and automatically triggering the scaling mechanism when the system load exceeds 80%. The purpose of this step is to establish a distributed quality assessment architecture, improving the scalability and reliability of the system.
[0059] The specific implementation of step S06 is to construct a collaborative assessment network among the assessment units. The horizontal assessment network uses the mutual information algorithm to calculate the information gain value between different data sources, and the assessment window is set to 5 minutes. The vertical assessment network uses the Pearson correlation coefficient to calculate the correlation strength value between data sources, and the correlation coefficient threshold is set to 0.7. Data sources with a correlation coefficient higher than this threshold are considered to have a strong correlation. The assessment network adopts a hierarchical structure. The first layer conducts the quality assessment within the data source, the second layer conducts the correlation analysis between data sources, and the third layer conducts global collaborative optimization. This step aims to achieve multi-dimensional collaborative assessment of data quality, improving the accuracy and reliability of the assessment results.
[0060] The specific implementation of step S07 is to calculate the data quality consistency index based on the information gain value and the correlation strength value. The calculation process first normalizes the information gain value and the correlation strength value, and then obtains the preliminary consistency index through weighted summation. The weight coefficient is determined by the grey relational analysis method, considering the importance and reliability of the data source. The value range of the consistency index is from -1 to +1, where a value greater than 0.8 indicates high consistency, 0.5 to 0.8 indicates medium consistency, and less than 0.5 indicates low consistency. The purpose of this step is to quantify the correlation degree between different data sources, providing a basis for subsequent comprehensive assessment.
[0061] The specific implementation of step S08 is to construct a multi-layer neural network evaluation model. The data encoding layer adopts an autoencoder structure to compress the input data into a 128-dimensional feature space. The feature extraction layer adopts a residual network structure, which contains 3 residual blocks, and each residual block consists of two convolutional layers and a skip connection. The pyramid processing layer adopts a 3-layer pyramid structure, and the size of the feature map of each layer is halved in turn. The feature fusion layer uses an attention mechanism to perform weighted fusion on features at different levels. The pattern recognition layer adopts a fully connected network structure, and the output layer uses a linear activation function. This step aims to establish an end-to-end data quality evaluation model to achieve automated quality evaluation.
[0062] The specific implementation of step S09 is to implement the feature fusion of the pyramid processing layer and use a soft attention mechanism to calculate the feature weights. The attention weight matrix is generated by a three-layer fully connected network, and the input is the statistical descriptors of the features of each layer, including mean, variance, and kurtosis. The weight matrix is linearly normalized to ensure that the sum of the feature weights of each layer is 1. Feature fusion is performed by weighted summation, and the fused features are added to the original features through a residual connection to improve the expression ability of the features. The purpose of this step is to achieve the adaptive selection and fusion of features and improve the feature extraction ability of the model.
[0063] The specific implementation of step S10 is to construct a data quality scoring model, which uses a logarithmic function to map the comprehensive evaluation result to the range of 0 to 100 points. The mapping function adopts a segmented processing method, and different mapping curvatures are used in different score intervals to ensure the discrimination of the scoring results. The scoring model also includes an anomaly detection module, which triggers a review mechanism when the scoring result shows a mutation. The scoring results are classified as high-quality data above 90 points, good data from 80 to 89 points, qualified data from 70 to 79 points, and unqualified data below 70 points. This step aims to convert the complex quality evaluation result into an intuitive score for easy understanding and use.
[0064] The equations or mathematical models involved in the present invention will be described in detail below.
[0065] The calculation equations of the data deviation rate index and the data completeness rate index are specifically expressed as follows:
[0066]
[0067] In the formula, DR is the data deviation rate index; M i is the i-th measured value; S i is the i-th standard value; n is the number of samples; α is the correction coefficient; σ is the standard deviation; CR is the data completeness rate index; N valid is the number of valid data points; N toal is the theoretical number of data points; β is the time decay coefficient; Δt is the sampling period deviation.
[0068] The calculation equation of the data timeliness index is specifically expressed as follows:
[0069]
[0070] In the formula, TR is the timeliness index; γ is the delay sensitivity coefficient; T delay is the data transmission delay time; T std is the standard response time; λ is the processing time weight coefficient; T process is the data processing time.
[0071] The calculation equation of the feature entropy matrix is specifically expressed as follows:
[0072]
[0073] In the formula, H ij is an element of the feature entropy matrix; P ijk is the probability value of the k-th feature in the i-th row and j-th column; m is the number of features; η is the information diffusion coefficient; D ij is the local density.
[0074] The calculation equation of the data fluctuation coefficient matrix is specifically expressed as follows:
[0075]
[0076] In the formula, V ij is an element of the fluctuation coefficient matrix; X ijt is the time series data; is the mean value; T is the time window length; μ is the fluctuation correction coefficient; R ij is the relative change rate.
[0077] The calculation equation of the information gain value is specifically expressed as follows:
[0078] IG ab = H(A) + H(B) - H(A, B) + θ·MI ab ;
[0079] In the formula, IG ab is the information gain value between data sources a and b; H(A) and H(B) are the entropy values of the data sources respectively; H(A, B) is the joint entropy; θ is the mutual information weight coefficient; MI ab is the mutual information value.
[0080] The calculation equation of the correlation strength value is specifically expressed as follows:
[0081]
[0082] In the formula, CS abis the association strength value between data sources a and b; X at and X bt are time series data; and are the means; ω is the synchronization weight coefficient; S ab is the synchronization coefficient.
[0083] The calculation equation of the data quality consistency index is specifically expressed as follows:
[0084]
[0085] In the formula, CI is the consistency index; w i and w j are the weight coefficients; φ is the global consistency adjustment coefficient; E is the system entropy.
[0086] The construction principles of these equations are as follows:
[0087] 1. The data deviation rate index equation uses the root mean square error as the basis, introduces the standard deviation term to consider the data distribution characteristics, and adjusts the error sensitivity through the correction coefficient;
[0088] 2. The data completeness rate index equation is based on the ratio of the number of data points, and introduces the time decay term to reflect the integrity of the data time series;
[0089] 3. The timeliness index equation combines the exponential decay and the inverse function to respectively characterize the effects of transmission delay and processing time;
[0090] 4. The feature entropy matrix equation is based on Shannon entropy, and adds the local density term to describe the feature distribution characteristics;
[0091] 5. The fluctuation coefficient matrix equation uses the coefficient of variation principle, and introduces the relative change rate to reflect the dynamic characteristics;
[0092] 6. The information gain value equation is based on the mutual information theory, and adds the mutual information weight term to enhance the description of feature correlation;
[0093] 7. The association strength value equation is based on the Pearson correlation coefficient, and introduces the synchronization coefficient to characterize the time series correlation;
[0094] 8. The consistency index equation uses the weighted average principle, and introduces the system entropy term to characterize the overall consistency level.
[0095] Description of the parameter acquisition method:
[0096] 1. α, β, γ, λ, η, μ, θ, ω, φ are system parameters, and the optimal values are determined through cross-validation;
[0097] 2. M i 、S i 、N valid, N total Obtained directly through the data acquisition system;
[0098] 3. T delay , T process Obtained by calculating the timestamp;
[0099] 4. P ijk Obtained by calculating through the feature probability distribution matrix;
[0100] 5. D ij , R ij , S ab , E Obtained by statistical analysis calculation;
[0101] 6. w i , w j Optimized by the grey relational analysis method.
[0102] The following details the specific calculation processes of multiple evaluation equations and feature fusion
[0103] Among them, the basic evaluation equation is specifically expressed as follows:
[0104]
[0105] In the formula, BQ is the basic quality index; w 1 , w 2 , w 3 , w 4 are weight coefficients and satisfy DR is the data deviation rate index; CR is the data completeness rate index; TR is the timeliness index; V ij is an element of the fluctuation coefficient matrix; ξ is the basic quality correction coefficient; Q b is the basic quality compensation term.
[0106] The correlation evaluation equation is specifically expressed as follows:
[0107]
[0108] In the formula, RQ is the correlation quality index; c 1 , c 2 , c 3 are correlation coefficients and satisfy IG ab is the information gain value; CS ab is the correlation strength value; CI is the consistency index; ψ is the correlation quality correction coefficient; Q r is the correlation quality compensation term.
[0109] The comprehensive evaluation equation is specifically expressed as follows:
[0110]
[0111] In the formula, CQ is the comprehensive quality index; k 1 , k 2 is the comprehensive coefficient and satisfies k 1 +k 2 = 1; is the partial derivative of the basic quality index with respect to time; RQ is the associated quality index; δ is the time-varying adjustment coefficient; f(t) is the time decay function; Q c is the comprehensive quality compensation term.
[0112] The calculation of the attention weight matrix is specifically expressed as follows:
[0113]
[0114] In the formula, A ij is the element of the attention weight matrix; e ij is the element of the energy matrix; Q i is the query vector; K j is the key vector; d is the feature dimension; ρ is the attention adjustment coefficient; S ij is the element of the similarity matrix; b ij is the bias term.
[0115] The calculation of the feature fusion is specifically expressed as follows:
[0116]
[0117] In the formula, F is the fused feature; w l is the hierarchical weight; F l is the feature of the l-th layer; L is the number of pyramid layers; τ is the feature interaction coefficient; α l is the feature importance index.
[0118] The construction principles of these equations are explained as follows:
[0119] 1. The basic evaluation equation adopts the weighted summation form, balances the importance of each index through the weight coefficient, and introduces the compensation term to handle special cases;
[0120] 2. The associated evaluation equation combines linear combination and logarithmic transformation to enhance the ability to characterize the associated characteristics;
[0121] 3. The comprehensive evaluation equation introduces the time partial derivative term to reflect the dynamic change characteristics of the quality index;
[0122] 4. The attention weight matrix is based on the dot product attention mechanism, and improves the feature selection ability through the scaled dot product and the similarity matrix;
[0123] 5. The feature fusion calculation adopts a form that combines weighted summation and geometric mean to enhance the expression of feature interaction.
[0124] Parameter acquisition method:
[0125] 1. ξ, ψ, δ, ρ, τ are system adjustment parameters, and the optimal values are determined through grid search;
[0126] 2. Q b 、Q r 、Q c are obtained through statistical analysis of historical data;
[0127] 3. f(t) adopts the form of an exponential decay function;
[0128] 4. Q i 、K j are obtained through feature transformation;
[0129] 5. b ij is obtained through random initialization and then training;
[0130] 6. α l is obtained through backpropagation optimization.
[0131] Among them, all the above parameters and variables are real numbers, and the domain is the set of non-negative real numbers. The optimal values of the parameter ranges are determined through experiments according to the actual application scenarios. All matrix operations satisfy the basic rules of matrix algebra to ensure the mathematical rationality of the calculation process. When the system runs, the processing efficiency is improved through parallel computing, and an adaptive adjustment mechanism is adopted to achieve dynamic optimization of the parameters.
[0132] Among them, the derivation process of the data deviation rate index equation:
[0133] First, the root mean square error is adopted as the basic form:
[0134] Considering the data distribution characteristics, the standard deviation term is introduced:
[0135] Finally, the complete deviation rate index is formed:
[0136] Among them, the standard deviation term is weighted and adjusted by the coefficient α, and the value range of α is from 0.1 to 0.5, which is determined through cross-validation.
[0137] The specific representation of the feature entropy matrix is:
[0138]
[0139] Among them, the matrix element h ij is calculated through the feature probability distribution:
[0140] Local density D ij is calculated using the kernel density estimation method:
[0141] where K is the kernel function and h is the bandwidth parameter, which is determined by regular distribution estimation.
[0142] The fluctuation coefficient matrix is expressed as:
[0143]
[0144] where the relative change rate R ij is calculated from time series data:
[0145] The calculation process of the attention weight matrix:
[0146] First, construct the query matrix Q and the key matrix K:
[0147]
[0148] Calculate the energy matrix through matrix multiplication: E = QK T ;
[0149] Apply the softmax function to obtain the attention weights:
[0150] The optimization process of the comprehensive evaluation equation:
[0151] 1. The time derivative of the basic quality index is approximated by difference:
[0152] 2. The time decay function is in exponential form: f(t) = e -λt ;
[0153] 3. Optimize the comprehensive coefficient by the gradient descent method:
[0154] where L is the loss function and η is the learning rate.
[0155] The optimization process of feature fusion includes:
[0156] 1. The hierarchical weights are calculated through the attention mechanism:
[0157] 2. The feature importance index is updated through backpropagation:
[0158] 3. The feature interaction coefficient is determined by the validation set: τ = argmin τ {L val (τ)}.
[0159] In the second aspect of the present invention, a computer-readable storage medium is provided. Program instructions are stored in the computer-readable storage medium. When the program instructions run on a computer, they are used to execute the above-mentioned industrial data quality assessment method based on microservices.
[0160] In the third aspect of the present invention, an industrial data quality assessment system based on microservices is provided, which includes the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is arranged inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is arranged inside the system.
[0161] Specifically, the principle of the present invention is as follows: The technical principle of the present invention is based on the organic combination of deep learning and microservice architecture, and a complete dynamic data quality assessment system is constructed. At the level of data feature extraction, a parallel long short-term memory network and a convolutional neural network are adopted. The long short-term memory network captures the long-term dependence relationship of data through a gating mechanism, while the convolutional neural network extracts the spatial features of data through a local receptive field. The combination of the two realizes the comprehensive extraction of the dynamic features of data.
[0162] In terms of the construction of the evaluation model, the present invention adopts a pyramid-style multi-layer processing structure, solves the problem of gradient disappearance in deep networks through an improved residual network, realizes the adaptive selection of features through an attention mechanism, and enhances the feature reuse ability through a densely connected network. This multi-level processing mechanism can adaptively adjust the feature extraction strategy to ensure the sensitivity of the evaluation model to changes in data features. The adoption of the microservice architecture provides the flexible deployment ability of evaluation units, and through the collaboration of horizontal and vertical evaluation networks, the comprehensive evaluation of multi-source data is realized.
[0163] In terms of the mathematical basis, the present invention establishes a complete set of evaluation equations, including three levels: basic evaluation, correlation evaluation, and comprehensive evaluation. By introducing indicators such as feature entropy and fluctuation coefficient, a quantitative evaluation system for data quality is constructed. In particular, the concepts of information gain and correlation strength are introduced in the correlation evaluation, and the dynamic characteristics of data quality are described through partial derivatives and logarithmic relationships, providing a reliable mathematical support for the evaluation results. This evaluation method based on a mathematical model ensures the scientificity and interpretability of the evaluation results.
[0164] A specific embodiment 1 of the present invention is provided below. The specific implementation manners of each step in this embodiment 1 are described in detail as follows.
[0165] To better understand and implement the present invention, the following provides an embodiment 2 of a specific application scenario of the present invention: When an automobile manufacturing enterprise implemented an industrial data quality assessment system, in-depth application research was conducted on the intelligent engine assembly production line. This production line includes 100 data collection points, covering multiple dimensions such as process parameters, equipment status, quality inspection, and environmental monitoring. First, a distributed data collection network was established, and the specific configuration data of the collection nodes is shown in Table 1:
[0166] Table 1 Configuration Parameter Table of Data Collection Nodes
[0167] Parameter type Parameter value Description Sampling period 1 second Standard sampling period Local cache 256 MB Can cache 2 hours of data when the network is disconnected Network bandwidth 100 Mbps Industrial Ethernet Heartbeat period 3 seconds Fault detection period Data compression ratio 85% Difference compression algorithm Backup period 30 minutes Data backup interval
[0168] Figure 2 Shows the performance monitoring data of the data collection nodes within 24 hours, including the real-time changes in CPU usage, memory usage, and network bandwidth usage. During the establishment of the data quality assessment index system, by analyzing one month of historical data, the benchmark values of various parameters were determined, as shown in Table 2:
[0169] Table 2 Benchmark Value Table of Quality Assessment Parameters
[0170] Parameter name Value Range Correction coefficient α 0.25 0.1~0.5 Time decay coefficient β 0.05 0.01~0.1 Delay sensitivity coefficient γ 0.5 0.1~1.0 Processing time weight coefficient λ 0.3 0.1~0.5 Information diffusion coefficient η 0.2 0.1~0.3 Fluctuation correction coefficient μ 0.3 0.1~0.5
[0171] The training of the deep learning model uses three months of historical data, totaling 2,592,000 data samples. Figure 3 Shows the change trends of training loss, validation loss, training accuracy, and validation accuracy during 100 rounds of training of the deep learning model. The training parameters of the long short-term memory network are shown in Table 3:
[0172] Table 3 Training Parameter Table of Long Short-Term Memory Network
[0173] Parameter name Value Number of network layers 4 layers Number of neurons 128 / layer Batch size 64 Learning rate 0.001 Number of training rounds 100 Validation ratio 0.2
[0174] Based on the trained model, the test data set is evaluated, and the data quality score distribution is shown in Table 4:
[0175] Table 4 Data Quality Score Distribution Table
[0176]
[0177]
[0178] Figure 4 Shows the time series change of the quality score within 30 days, including real-time quality score, benchmark score line, normal fluctuation range, and abnormal point identification. During the specific evaluation process, the following parameter values are used for the calculation of basic quality indicators: w 1 = 0.3, w 2 = 0.25, w3 = 0.25, w 4 = 0.2, compensation term Q b = 0.1. The associated quality index calculation uses the parameters: c 1 = 0.4, c 2 = 0.35, c 3 = 0.25, compensation term Q r = 0.15. The comprehensive evaluation uses the parameters: k 1 = 0.4, k 2 = 0.6, and the time-varying adjustment coefficient δ = 0.2.
[0179] The performance indicators of the system after running for one week are shown in Table 5 as follows:
[0180] Table 5 System Operating Performance Indicator Table
[0181] Metric name Actual value Target value Data acquisition success rate 99.8% 99.5% Average response time 85 ms 100 ms Data processing delay 150 ms 200 ms System availability 99.99% 99.9% Accurate warning rate 92.5% 90% False alarm rate 2.1% 3%
[0182] Traditional data quality assessment methods mainly rely on manual experience judgment and simple statistical analysis, suffering from problems such as strong subjectivity, poor real-time performance, and inability to handle massive data. Specifically, it is manifested as: the assessment indicators are single, usually only considering data integrity and accuracy; the assessment process lacks adaptability and cannot dynamically adjust the assessment strategy according to data characteristics; the interpretability of the assessment results is poor and it is difficult to provide effective guidance for quality improvement. The present invention adopts deep learning and multi-dimensional assessment methods and has made significant progress in the following aspects: First, a comprehensive assessment index system is constructed, covering multiple dimensions such as data deviation rate, completeness rate, timeliness, etc., which can comprehensively reflect the data quality status. Second, through the deep learning model, automatic feature extraction and optimization are realized, avoiding the limitations of manual feature engineering. Third, an improved pyramid structure and attention mechanism are adopted to improve the model's feature extraction and fusion capabilities. Fourth, a horizontal and vertical collaborative assessment network is established to achieve unified quality assessment of multi-source heterogeneous data. Fifth, an adaptive weight adjustment mechanism is introduced, which can dynamically optimize the assessment strategy according to data characteristics. In practical applications, compared with traditional methods, the accurate warning rate of the method of the present invention has increased by 15 percentage points, the false alarm rate has decreased by 5 percentage points, the processing efficiency has increased by 3 times, and the system scalability and maintainability have also been significantly improved.
[0183] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Table 6 below.
[0184] Table 6 Variable Explanation Table
[0185]
[0186]
[0187] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.
Claims
1. A microservice-based industrial data quality assessment method, characterized in that: The following steps are included: collecting multi-source heterogeneous data of industrial equipment, establishing a distributed data collection network, and obtaining industrial data streams; Construct a data quality assessment index system, and establish a data quality matrix including a data deviation rate index, a data completeness rate index, and a data timeliness index; establish a coordination mechanism between the data quality assessment units, and construct a horizontal assessment network and a vertical assessment network, the horizontal assessment network calculates the information gain value between different data sources, and the vertical assessment network calculates the association strength value of the data source; establish a multi-layer neural network assessment model based on the data quality consistency index, the multi-layer neural network assessment model includes a data encoding layer, a feature extraction layer, a pyramid processing layer, a feature fusion layer, and a pattern recognition layer, the pyramid processing layer includes a three-layer pyramid structure, the first layer pyramid structure includes an improved residual network structure, the second layer pyramid structure includes an improved attention network structure, and the third layer pyramid structure includes an improved densely connected network structure.
2. The microservice-based industrial data quality assessment method according to claim 1, characterized in that: The specific structure of the first-layer pyramid structure is composed of three residual blocks connected in series, each residual block contains two layers of convolution, batch normalization and activation function, the batch normalization module is used to normalize the feature distribution and improve the stability of network training. The data flow transmitted within the first-layer pyramid structure is the fusion feature of the original feature after residual mapping and short-circuit connection addition.
3. The microservice-based industrial data quality assessment method according to claim 1, characterized in that: The specific structure of the second-layer pyramid structure includes a dual attention mechanism of a spatial attention subnetwork and a channel attention subnetwork. The channel attention module is used to learn the dependencies between feature channels and dynamically adjust channel weights. The data flow inside the second-layer pyramid structure transmits data that is a multi-scale feature map weighted by spatial and channel attention.
4. The microservice-based industrial data quality assessment method according to claim 1, characterized in that: The specific structure of the third-layer pyramid structure is a multi-path network with cross-layer feature reuse, which includes 4 dense blocks. The feature selection module is used to screen and combine multi-layer features to reduce redundant information. The data flow transmission data inside the third-layer pyramid structure is a densely connected combination of features of different scales.
5. The microservice-based industrial data quality assessment method according to claim 1, characterized in that: The relationship between the parameters in the basic evaluation equation is: the basic quality index is equal to the product of the data deviation rate index and the first weight coefficient, plus the product of the data completeness rate index and the second weight coefficient, plus the product of the data timeliness index and the third weight coefficient, plus the product of the data volatility coefficient matrix and the fourth weight coefficient, and the sum of the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient is equal to 1.
6. The microservice-based industrial data quality assessment method according to claim 1, characterized in that: The relationship between the parameters in the association evaluation equation is: the association quality index is equal to the product of the information gain value and the first association coefficient, plus the product of the association strength value and the second association coefficient, plus the product of the natural logarithm of the data quality consistency index and the third association coefficient, and the sum of the first association coefficient, the second association coefficient, and the third association coefficient is equal to 1.
7. The microservice-based industrial data quality assessment method according to claim 1, characterized in that: The relationship between the parameters in the comprehensive evaluation equation is: the comprehensive quality index is equal to the product of the partial derivative of the basic quality index with respect to time and the first comprehensive coefficient, plus the product of the associated quality index and the second comprehensive coefficient, and the sum of the first comprehensive coefficient and the second comprehensive coefficient is equal to 1.
8. The microservice-based industrial data quality assessment method according to claim 1, characterized in that: A data quality scoring model is constructed based on the weighted fusion result, and the output result of the comprehensive evaluation equation is mapped to a range of 0 to 100 points through a normalization function. The normalization function uses a logarithmic form for numerical mapping to ensure the smoothness and discrimination of the scoring result. The data quality is divided into the following levels according to the mapped scores: 90 to 100 points represent high-quality data, 80 to 89 points represent good data, 70 to 79 points represent qualified data, and below 70 points represent unqualified data.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and when the program instructions are executed in a computer, they are used to execute the microservice-based industrial data quality assessment method according to any one of claims 1 to 8.
10. A microservice-based industrial data quality assessment system, characterized in that: The system comprises the computer-readable storage medium as claimed in claim 9, wherein the system is any one of a computer, a server, and a single-chip microcomputer, the computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.
Citation Information
Patent Citations
Business quality analysis method and system under micro-service architecture
CN109961204A
Quality evaluation method and system for multi-source heterogeneous data
CN111639850A
Single remote sensing image height information estimation method based on deep learning algorithm
CN114972989A
Heterogeneous multi-source multi-purpose data synchronization method and device, computer and storage medium
CN117149906A
Land resource management method and system based on big data analysis
CN119323327A
Cited By
AI-based data quality intelligent evaluation optimization system
CN120410574A
Real-time acquisition and quality evaluation method and system based on multi-source heterogeneous data
CN121542253A
Method and system for real-time collection and quality evaluation based on multi-source heterogeneous data
CN121542253B