A manufacturing big data quality identification method based on aging correlation rule

By constructing a manufacturing big data quality assessment method based on timeliness association rules, the problem of assessing the timeliness of sensor data has been solved, enabling effective assessment of sensor data quality and anomaly detection, and improving the reliability of data analysis in the manufacturing process.

CN116701379BActive Publication Date: 2026-01-16SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310821554.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-05
Publication Date
2026-01-16
Estimated Expiration
2043-07-05

AI Technical Summary

Technical Problem

The lack of effective methods in the current technology to assess the timeliness of sensor data makes it difficult to detect sensor anomalies in the manufacturing process, affecting data quality and the effectiveness of analysis results.

Method used

By constructing a manufacturing big data quality assessment method based on time-dependent correlation rules, including data preprocessing, timestamp alignment, calculation of correlation coefficients, and construction of numerical correlation graphs and time-delay cross-correlation graphs, the timeliness of sensor data is evaluated.

Benefits of technology

It enables effective identification of the timeliness of sensor data, ensures data quality and the reliability of analysis results, detects sensor anomalies, and improves the monitoring and management capabilities of the manufacturing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116701379B_ABST
    Figure CN116701379B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a manufacturing big data quality identification method based on time-effect correlation rules, relates to the technical field of data quality identification, and the correlation coefficient between the data of sensors is calculated through preprocessing and time stamp alignment of original data in a sensor group, so that the numerical correlation between the sensor data is determined. Through the numerical correlation between the preprocessed original data and each data, the time delay cross correlation between the data is determined, and finally the time effectiveness of the sensor data is effectively identified according to the numerical correlation and the time delay cross correlation between the data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data quality identification, and particularly relates to a manufacturing big data quality identification method based on time-effectiveness correlation rules. BACKGROUND

[0002] In recent years, with the rapid development of manufacturing industry, equipment sensors, intelligent instruments and other equipment are widely applied to the production process of manufacturing industry, and manufacturing data with large data scale are generated. These data are usually structured and have explicit time stamps, and can meet the application scenarios of monitoring and recording, sensing, analysis and the like. Using these data can better understand the state and fault conditions of the production line, which is helpful for the operation of the factory. How to ensure the quality of the data and thus ensure the effectiveness, reliability and usability of the data analysis results is a key factor affecting the development and operation of various enterprises.

[0003] Time-effectiveness is an important dimension affecting data quality. In different fields, too low data time-effectiveness sometimes brings many troubles and even huge losses. In the manufacturing industry, how to evaluate the time-effectiveness of sensor data and identify the data quality of sensors to find sensor abnormalities in the manufacturing process is very important. However, there is currently few evaluation method of time-effectiveness combined with data correlation.

[0004] Therefore, there is an urgent need for a manufacturing big data quality identification method based on time-effectiveness correlation rules. SUMMARY

[0005] In view of the above problems, the embodiments of the present application provide a manufacturing big data quality identification method based on time-effectiveness correlation rules, so as to overcome the above problems or at least partially solve the above problems.

[0006] In a first aspect, the embodiments of the present application provide a manufacturing big data quality identification method based on time-effectiveness correlation rules, and the method comprises the following steps:

[0007] obtaining original data recorded by each sensor in a sensor group, and performing data preprocessing on the original data to obtain first data;

[0008] aligning the first data with time stamps to obtain second data, and retaining the first data;

[0009] calculating respective correlation coefficients between data of all sensors in the sensor group according to the second data, to obtain a numerical correlation matrix of the sensor group;

[0010] constructing a numerical correlation graph by taking sensors as nodes and taking target correlation coefficients as edges, wherein the target correlation coefficients are correlation coefficients with absolute values greater than or equal to a correlation coefficient threshold in the correlation coefficients.

[0011] According to the numerical correlation graph and the first data, time offsets corresponding to each pair of sensors with adjacent nodes in the numerical correlation graph are calculated, and a time-delay cross-correlation graph is constructed by taking the sensors with adjacent nodes in the numerical correlation graph as nodes and taking the time offsets as edges.

[0012] According to the time offsets corresponding to each pair of sensors in the time-delay cross-correlation graph, a timeliness evaluation value of data of each sensor in the sensor group is determined.

[0013] According to the timeliness evaluation value of data of each sensor in the sensor group, the sensor data in the sensor group is subjected to timeliness identification.

[0014] Optionally, the determining of the timeliness evaluation value of data of each sensor in the sensor group according to the time offsets corresponding to each pair of sensors in the time-delay cross-correlation graph comprises:

[0015] The actual update time of a post-sampling sensor for which the timeliness evaluation value needs to be determined and the actual update time of a pre-sampling sensor that is adjacent to the post-sampling sensor are obtained, the post-sampling sensor representing a sensor with data updated later, and the pre-sampling sensor representing a sensor with data updated earlier;

[0016] The theoretical update time of the post-sampling sensor is determined according to the actual update time of the pre-sampling sensor and the time offset between the post-sampling sensor and the pre-sampling sensor.

[0017] The timeliness evaluation value of data of the post-sampling sensor is determined according to the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor.

[0018] Optionally, the determining of the timeliness evaluation value of data of the post-sampling sensor according to the actual update time of the post-sampling sensor and the theoretical update time comprises:

[0019] In a case where the difference between the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor is greater than or equal to 0 and less than or equal to the average sampling time interval of the sensor, the difference between the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor is substituted into a timeliness score equation to obtain the timeliness evaluation value of data of the post-sampling sensor, the timeliness score equation being shown in the following formula:

[0020]

[0021] wherein Q curr_tmpdiff(t, t') is the difference between the actual update time of the sensor and the theoretical update time of the sensor, t is the actual update time of the sensor, t' is the theoretical update time of the sensor, e is a constant, and T is the average sampling time interval of the sensor. avg The average sampling time interval of the sensor.

[0022] Optionally, in the case where the difference between the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor is less than 0, the timeliness evaluation value of the post-sampling sensor data is 1, and in the case where the difference between the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor is greater than the average sampling time interval of the sensor, the timeliness evaluation value of the post-sampling sensor data is 0.

[0023] Optionally, according to the numerical correlation graph and the first data, the time offset between each pair of sensors containing adjacent nodes in the numerical correlation graph is calculated, including:

[0024] The sensors containing adjacent nodes are determined from the numerical correlation graph.

[0025] According to the sensors containing adjacent nodes in the numerical correlation graph, the third data corresponding to each of the sensors containing adjacent nodes without time stamp alignment is read from the first data.

[0026] According to the third data, the sampling frequency of each sensor corresponding to each data in the third data is calculated.

[0027] According to the calculated sampling frequency of each sensor corresponding to each data, the time lag cross-correlation coefficient between any two sensors with the same sampling frequency is determined.

[0028] According to the time lag cross-correlation coefficient between any two sensors with the same sampling frequency, the time offset between any two sensors with the same sampling frequency is determined.

[0029] Optionally, the determination of the time offset between any two sensors with the same sampling frequency according to the time lag cross-correlation coefficient between any two sensors with the same sampling frequency includes:

[0030] According to a plurality of offset values in a preset offset interval, the second sensor in the any two sensors with the same sampling frequency is shifted as a whole with respect to the first sensor in the any two sensors with the same sampling frequency, to obtain the time lag cross-correlation coefficient between the first sensor and the second sensor under different offset values.

[0031] determine the time offset of the first sensor and the second sensor with the same sampling frequency according to the offset value corresponding to the time lag cross-correlation coefficient with the maximum value at different offset values.

[0032] determine the time offset of the first sensor and the second sensor with the same sampling frequency according to the offset value corresponding to the time lag cross-correlation coefficient with the maximum value at different offset values.

[0033] Optionally, the determining the time offset of the first sensor and the second sensor with the same sampling frequency according to the offset value corresponding to the time lag cross-correlation coefficient with the maximum value at different offset values comprises:

[0034] obtaining the time stamp of the first sensor at each recording point and the time stamp of the second sensor at each recording point at the offset value with the maximum value;

[0035] calculating the average value of the difference between the time stamp of the first sensor at each recording point and the time stamp of the second sensor at each recording point, and determining the average value of the difference between the time stamps as the time offset of the first sensor and the second sensor.

[0036] Optionally, the calculating the sampling frequency of the sensor corresponding to each data in the third data according to the third data comprises:

[0037] obtaining the sampling time interval of the data point corresponding to each data in the third data;

[0038] calculating the average sampling time interval according to the sampling time interval of the data point corresponding to each data in the third data;

[0039] calculating the sampling frequency of the sensor corresponding to each data in the third data according to the average sampling time interval.

[0040] Optionally, the calculating the correlation coefficient between each pair of data of all sensors in the sensor group according to the second data to obtain the numerical correlation matrix of the sensor group comprises:

[0041] dividing the second data into multiple sub-sequences, and calculating the correlation coefficient between the data in each sub-sequence;

[0042] obtaining the numerical correlation matrix of the sensor group by using the covariance matrix according to the correlation coefficient between the data in each sub-sequence.

[0043] Optionally, the timeliness identification of the sensor data in the sensor group according to the timeliness evaluation value of each sensor data in the sensor group comprises:

[0044] Comparing the timeliness evaluation value of each sensor data in the sensor group with a timeliness threshold value respectively;

[0045] In the case that the timeliness evaluation value is greater than or equal to the timeliness threshold value, it is determined that the timeliness of the sensor data is not abnormal;

[0046] In the case that the timeliness evaluation value is less than the timeliness threshold value, it is determined that the timeliness of the sensor data is abnormal.

[0047] The beneficial effects of the embodiments of the present application are as follows:

[0048] The embodiments of the present application provide a manufacturing big data quality identification method based on timeliness correlation rules, which comprises: obtaining original data recorded by each sensor in a sensor group, and performing data preprocessing on the original data to obtain first data; aligning the first data with time stamps to obtain second data, and retaining the first data; calculating respective corresponding correlation coefficients between data of all sensors in the sensor group according to the second data to obtain a numerical correlation matrix of the sensor group; constructing a numerical correlation graph with sensors as nodes and target correlation coefficients as edges according to the numerical correlation matrix; the target correlation coefficient is a correlation coefficient in the correlation coefficient whose absolute value is greater than or equal to a correlation coefficient threshold; calculating respective corresponding time offsets between sensors with adjacent nodes in the numerical correlation graph according to the numerical correlation graph and the first data, and constructing a time-lag cross-correlation graph with sensors with adjacent nodes in the numerical correlation graph as nodes and time offsets as edges; determining a timeliness evaluation value of each sensor data in the sensor group according to respective corresponding time offsets between two sensors in the time-lag cross-correlation graph; and performing timeliness identification of the sensor data in the sensor group according to the timeliness evaluation value of each sensor data in the sensor group. The present application calculates the correlation coefficient between the data of the sensors by preprocessing and time stamp alignment of the original data in the sensor group, so as to determine the numerical correlation between the sensor data. The time-lag cross-correlation between the data is determined through the preprocessed original data and the numerical correlation between the data, and finally the timeliness of the sensor data is effectively identified according to the numerical correlation and the time-lag cross-correlation between the data. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0050] Figure 1 is a step flow diagram of a manufacturing big data quality identification method based on time-effect correlation rules provided by the embodiments of the present application;

[0051] Figure 2 is a flow chart of a manufacturing big data quality identification method based on time-effect correlation rules provided by the embodiments of the present application;

[0052] Figure 3 is a numerical correlation graph provided by the embodiments of the present application;

[0053] Figure 4 is a time-lag cross-correlation graph provided by the embodiments of the present application;

[0054] Figure 5 is a function relationship diagram of a time-effectiveness scoring equation provided by the embodiments of the present application. DETAILED DESCRIPTION

[0055] The exemplary embodiments of the present application will be described in detail below with reference to the drawings of the embodiments of the present application. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0056] In a first aspect of the embodiments of the present application, a manufacturing big data quality identification method based on time-effect correlation rules is provided, referring to Figure 1 and Figure 2 shown, Figure 1 is a step flow diagram of a manufacturing big data quality identification method based on time-effect correlation rules provided by the embodiments of the present application, Figure 2 is a flow chart of a manufacturing big data quality identification method based on time-effect correlation rules provided by the embodiments of the present application, the method comprising:

[0057] Step S101, obtaining the original data recorded by each sensor in the sensor group, and performing data preprocessing on the original data to obtain first data;

[0058] Specifically, in the embodiments of the present application, in the sensor group, each sensor undertakes a key task, recording various data perceived in industrial production. These raw data can be various information such as temperature, humidity, pressure, light intensity and the like in industrial production. However, these raw data are not always perfect, and usually need to be pre-processed to become more accurate and usable.

[0059] In practical application, in the data preprocessing process, first, the raw data need to be cleaned. This can include detecting and repairing sensor faults or abnormalities, such as sensor disconnection or measurement errors. Next, some incomplete data may need to be processed, such as handling missing values. Sometimes sensors may not be able to record certain data for various reasons, at which time interpolation or other techniques are used to fill in these missing values to maintain the integrity of the data. Another important preprocessing step is to remove noise. Sensor data can be affected by environmental interference or equipment failure, resulting in inaccurate fluctuations in the data. Through techniques such as filtering and smoothing, the interference of noise on the data can be effectively reduced, so that more stable and reliable data is obtained. Normalization or standardization of data is also a common step in the preprocessing process. Since different sensors may have different measurement ranges and units, converting them to a uniform scale can facilitate subsequent data comparison and analysis. After preprocessing, the raw data becomes the first data, which can also be referred to as "once-processed data".

[0060] The embodiments of the present application can clean and convert raw data, eliminate possible outliers and noise, fill in missing data, and normalize or standardize data by preprocessing raw data, so that subsequent data analysis and application can be more reliable and effective.

[0061] Step S102, time stamp alignment is performed on the first data to obtain second data, and the first data is retained;

[0062] Specifically, after preprocessing the raw data recorded by the sensor group to obtain the first data, the next important task is to perform time stamp alignment on the first data to obtain the second data. Time stamp alignment is to synchronize the data recorded by different sensors according to their exact time points of occurrence, so that the data can be accurately compared and correlated in subsequent analysis and application.

[0063] In practical application, the process of time stamp alignment involves time sorting and matching of data recorded by each sensor. Since the sampling frequency and triggering mechanism of different sensors can be different, the time interval between data can be different. Therefore, interpolation, alignment algorithm or other technical means can be used to correct these time deviations, so that the data can be aligned at the same time point.

[0064] Step S103, according to the second data, calculating the respective corresponding correlation coefficients between the data of all sensors in the sensor group, obtaining the numerical correlation matrix of the sensor group;

[0065] Specifically, the second data contains the data recorded by all sensors in the sensor group, and the respective corresponding correlation coefficients between the data of all sensors in the sensor group are calculated, thereby obtaining the numerical correlation matrix of the sensor group.

[0066] In a preferred embodiment of the present application, the second data is divided into multiple sub-sequences, and the correlation coefficients between the data in each sub-sequence are calculated.

[0067] According to the correlation coefficients between the data in each sub-sequence, the covariance matrix is used to obtain the numerical correlation matrix of the sensor group.

[0068] Specifically, in this embodiment, in order to calculate the numerical correlation between the data of different sensors in the sensor group, the data recorded by all sensors in the second data are divided into multiple sub-sequences. Each sub-sequence represents a set of data in a certain time range.

[0069] In order to evaluate the correlation between the data in each sub-sequence, we can calculate the correlation coefficients between them. The correlation coefficient measures the degree of correlation between two variables. In practical applications, commonly used correlation coefficients include Pearson correlation coefficient, Spearman rank correlation coefficient, etc.

[0070] By calculating the correlation coefficients of the data in each sub-sequence, we can obtain a matrix about the correlation between the data. This matrix is called the numerical correlation matrix, and each element in the matrix represents the correlation coefficient between two sub-sequence data. By analyzing the numerical correlation matrix, we can obtain the correlation degree and interaction mode between different sensors in the sensor group.

[0071] Further, in order to obtain the numerical correlation matrix, we can use the covariance matrix. The covariance matrix is a symmetric matrix, and the elements in the matrix represent the covariance between different variables. Covariance measures whether the overall trend of two variables is consistent, but does not consider their scales. By normalizing the covariance matrix, we can obtain the numerical correlation matrix.

[0072] In this embodiment, the qth sub-sequence is taken as an example: all sensor data in the sensor group are divided into several sub-sequences with length k, and the sensor sub-sequence of the qth sub-sequence is taken as an example. Each of the sensor sub-sequences comprises k data points, i.e.

[0073] Further, the sensor sub-sequence D q is subjected to Z-score standardization processing to convert the sensor data into dimensionless data with a mean value of 0 and a variable deviation of 1, thereby eliminating the interference of the unit dimension of the sensor data on the value correlation analysis. From the sensors, two sensors, i.e., a first sensor S i and a second sensor S j , are taken, and the covariance between the data of the first sensor S i and the second sensor S j under the qth sub-sequence is calculated, and the calculated covariance is determined as the correlation coefficient between the data of the first sensor S i and the second sensor S j under the qth sub-sequence. The calculation formula of the correlation coefficient is shown in the following formula (1):

[0074]

[0075] wherein, is the correlation coefficient between the data of the first sensor S i and the second sensor S j under the qth sub-sequence, k is a constant, and respectively represent the values of the data of the first sensor S i and the second sensor S j at the gth time point within the qth time period. and represent the mean values of all data points within the time period.

[0076] Further, within the qth sub-sequence, the correlation coefficients between the data of all sensors in the sensor group two by two are calculated by the above formula (1), and the covariance matrix is used to obtain the value correlation matrix of all sensor data of the sensor group within the qth sub-sequence, which is shown in the following formula (2):

[0077]

[0078] wherein, NCM q is the value correlation matrix within the qth sub-sequence, is the correlation coefficient of any two sensors within the qth sub-sequence.

[0079] ​​​​Further, the numerical correlation matrix of the sensor group is obtained by comprehensively considering the correlation results in all the divided q sub-sequences, and the numerical correlation matrix of the sensor group is as shown in the following formula (3):

[0080] wherein

[0081] wherein, NCM is the numerical correlation matrix of the sensor group, R nn is the correlation coefficient of any two sensors.

[0082] In step S104, a numerical correlation graph is constructed according to the numerical correlation matrix, taking the sensors as nodes and the target correlation coefficients as edges, wherein the target correlation coefficients are the correlation coefficients whose absolute values are greater than or equal to the correlation coefficient threshold value in the correlation coefficients.

[0083] Specifically, in this embodiment, according to the numerical correlation matrix of the sensor group, the correlation coefficients whose absolute values are greater than or equal to the correlation coefficient threshold value are determined from all the correlation coefficients of the numerical correlation matrix, and the correlation coefficients whose absolute values are greater than or equal to the correlation coefficient threshold value are determined as the target correlation coefficients, and a numerical correlation graph is constructed by taking the sensors in the sensor group as nodes and the target correlation coefficients as edges.

[0084] In actual application, the correlation coefficient threshold value can be 0.7, so the correlation coefficients whose absolute values are greater than or equal to 0.7 are determined as the target correlation coefficients, and a numerical correlation graph is constructed by taking the target correlation coefficients as edges and the sensors as nodes, and the constructed numerical correlation graph is as shown in Figure 3 .

[0085] In step S105, according to the numerical correlation graph and the first data, the time offset between each pair of sensors with adjacent nodes in the numerical correlation graph is calculated, and a time-lag cross-correlation graph is constructed by taking the sensors with adjacent nodes in the numerical correlation graph as nodes and the time offset as edges.

[0086] Specifically, according to the numerical correlation graph and the first data in the preceding steps, we can further calculate the time offset between the sensors. In the numerical correlation graph, we focus on the relationship between the sensors with adjacent nodes because they have certain interaction between them. By analyzing the numerical correlation between them, we can infer the time offset between them.

[0087] The time offset represents the difference in time of the data observed between the sensors. By comparing the data between two sensors, we can determine the time difference between them. For example, if we observe in the numerical correlation graph that the data of the first sensor S i is ahead of the data of the second sensor Sj the data of the first sensor S i the data of the second sensor S j the data of the first sensor S

[0088] With the relationship between the sensors in the numerical correlation graph and the time information of the first data, we can calculate the time offsets between each pair of sensors. These time offsets can be used to construct the graph structure of the time-lag cross-correlation graph. In the time-lag cross-correlation graph, we take the sensors with adjacent relationships as nodes, and the time offsets as edges between nodes. In this way, we can form a graphical representation that shows the time-lag relationships between sensors.

[0089] In a preferred embodiment, according to the numerical correlation graph and the first data, the respective time offsets between each pair of sensors containing adjacent nodes in the numerical correlation graph are calculated, including:

[0090] determining the sensors containing adjacent nodes from the numerical correlation graph;

[0091] reading the respective third data of the sensors containing adjacent nodes from the first data, without time stamp alignment;

[0092] calculating the sampling frequencies of the respective sensors in the third data according to the third data;

[0093] determining the time-lag cross-correlation coefficients between any two sensors with the same sampling frequency according to the calculated sampling frequencies of the respective sensors in the third data;

[0094] determining the time offsets between any two sensors with the same sampling frequency according to the time-lag cross-correlation coefficients between any two sensors with the same sampling frequency.

[0095] Specifically, the determination of the time offsets between any two sensors with the same sampling frequency according to the time-lag cross-correlation coefficients between any two sensors with the same sampling frequency includes the following steps:

[0096] performing overall offset on a second sensor in the any two sensors with the same sampling frequency according to a plurality of offset values in a preset offset interval, to obtain time-lag cross-correlation coefficients of the first sensor and the second sensor under different offset values, taking the first sensor as a reference;

[0097] Based on the time delay cross-correlation coefficients of the first sensor and the second sensor at different offset values, determine the offset value corresponding to the largest time delay cross-correlation coefficient at different offset values;

[0098] Based on the offset value corresponding to the largest time delay cross-correlation coefficient under different offset values, the time offset of the first sensor and the second sensor with the same sampling frequency is determined.

[0099] Specifically, determining the time offset between the first sensor and the second sensor with the same sampling frequency based on the offset value corresponding to the largest time-delay cross-correlation coefficient under different offset values ​​includes:

[0100] Under the offset value with the largest value, obtain the timestamp of the first sensor at each recording point and the timestamp of the second sensor at each recording point respectively;

[0101] Calculate the average difference between the timestamp of the first sensor at each recording point and the timestamp of the second sensor at each recording point, and determine the average difference of the timestamps as the time offset between the first sensor and the second sensor.

[0102] Specifically, in this embodiment, sensors containing adjacent nodes are identified by reading the numerical correlation graph, such as... Figure 3 As shown, sensors containing adjacent nodes can be S1-S2, S2-S6, S4-S12, etc., which are not listed here, while S3 and S7 do not have adjacent nodes. Based on the numerical correlation graph, sensors containing adjacent nodes are identified. Third data corresponding to each sensor containing adjacent nodes, which has undergone data preprocessing but not timestamp alignment, is read from the first data. The sampling frequency of each sensor corresponding to each data point in the third data is calculated. Based on the calculated sampling frequencies of each data point, the time delay cross-correlation coefficient between any two sensors with the same sampling frequency is determined, thereby further determining the time offset between any two sensors with the same sampling frequency.

[0103] In practical applications, the third data is divided into several subsequences of length k, and the q-th subsequence is taken. timestamp Timestamp of each sensor subsequence It contains k data points, that is Each sensor subsequence is calculated using the following formula (4). sampling frequency

[0104]

[0105] wherein T avg is the average sampling time interval of the sensor in the qth sub-sequence D .

[0106] T is calculated by the following formula (5):

[0107]

[0108] wherein T itr is the sampling time interval of the k data points of the sensor in the qth sub-sequence D , and the calculation method is shown in the following formula (6):

[0109]

[0110] It should be noted that the sensors in the present application are all sensors with stable sampling frequency, and therefore the sampling frequency f q of the qth sub-sequence D of the sensor calculated is determined as the sampling frequency of the sensor, and the sampling frequencies of all sensors in the sensor group can be obtained by the above-mentioned formulas (4)-(6).

[0111] Further according to the calculated sampling frequencies of all sensors, first, all nodes corresponding to the sensors and all edges in the numerical correlation graph are marked as unvisited, an unvisited node sensor, i.e., a first sensor S i , is randomly selected and marked as visited, all adjacent nodes of the first sensor S i are traversed, and from the sensors corresponding to all adjacent nodes of the first sensor S i , a sensor with the same sampling frequency as the first sensor S i is screened out, if there is only one sensor with the same sampling frequency as the first sensor S i , the sensor is determined as a second sensor S j , and taking the first sensor S i as a reference, the second sensor S remove is offset by the length of p, i.e., T remove = [-P, -(P+1)…, P-1, P], p∈T j , the time-lag cross-correlation between the first sensor S i and the second sensor S j is calculated until p is completely valued in T remove , and the specific formula for calculating the time-lag cross-correlation coefficient is shown in the following formula (7):

[0112]

[0113] in, For the first sensor S i With the second sensor S j The time-delay cross-correlation coefficient at an offset of p.

[0114] Furthermore, the first sensor S is calculated using the formula (7) above. i With the second sensor S j The set of time-delay cross-correlation coefficients at different offsets The maximum value of the time-delay cross-correlation coefficient is determined from this set, and the corresponding offset p is determined based on the maximum value of the time-delay cross-correlation coefficient. rmax Calculate at offset p rmax Below, the first sensor S i With the second sensor S j The mean of the timestamp differences for each record point in the q-th subsequence is denoted as T. lag The calculation method is shown in the following formula (8):

[0115]

[0116] in, Indicates the first sensor S i The recording time at the g-th time point in the q-th subsequence. Indicates the second sensor S j On the q-th subsequence g+prmax The recording time at each time point. rmax ≥0 indicates the second sensor S j The record lags behind S i At this time T lag If negative, p rmax <0 indicates the second sensor S j The recording precedes that of the first sensor S. i At this time T lag It is positive.

[0117] By following the steps above, until all nodes and edges in the numerical correlation graph have been visited, draw the graph as shown below. Figure 4 The time-delay cross-correlation diagram shown is from... Figure 4 As can be seen from this, S6→S2, that is, S6 is the first sensor to be acquired and S2 is the second sensor to be acquired, and the time delay cross-correlation coefficient between the two is 0.21. The time delay cross-correlation coefficients of other sensors will not be elaborated here.

[0118] Step S106: Determine the timeliness evaluation value of each sensor data in the sensor group based on the time offset between each pair of sensors in the time delay cross-correlation diagram.

[0119] Specifically, the determining of the timeliness evaluation value of each sensor data in the sensor group according to the respective time offset between each two sensors in the time-lag cross-correlation diagram comprises:

[0120] obtaining an actual update time of a post-sampling sensor for which the timeliness evaluation value needs to be determined and an actual update time of a pre-sampling sensor adjacent to the post-sampling sensor, the post-sampling sensor representing a sensor with data updated later, and the pre-sampling sensor representing a sensor with data updated earlier;

[0121] determining a theoretical update time of the post-sampling sensor according to the actual update time of the pre-sampling sensor and the time offset between the post-sampling sensor and the pre-sampling sensor;

[0122] determining the timeliness evaluation value of the post-sampling sensor data according to the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor.

[0123] substituting the difference between the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor into a timeliness score equation to obtain the timeliness evaluation value of the post-sampling sensor, the timeliness score equation being shown in the following formula (9):

[0124]

[0125] wherein Q curr_tmp is the timeliness evaluation value of the sensor, diff(t, t') is the difference between the actual update time and the theoretical update time of the sensor, t is the actual update time of the sensor, t' is the theoretical update time of the sensor, e is a constant, and T avg is the average sampling time interval of the sensor, and it should be noted that the timeliness evaluation value of the sensor, the actual update time of the sensor, the theoretical update time of the sensor and the like mentioned in the present application can be understood as data parameters possessed by the sensor itself, rather than the sensor itself.

[0126] In the present embodiment, the theoretical update time of the post-sampling sensor is determined according to the actual update time of the pre-sampling sensor and the time offset between the post-sampling sensor and the pre-sampling sensor, and can be calculated by the following formula (10):

[0127] T(D after )=T′(D before )+T lag (10)

[0128] wherein T'(D before ) is the actual recording time of the current sampling of the pre-sampling sensor, and T(D after() represents the theoretical recording time of the sensor during this sampling.

[0129] In a preferred embodiment, the difference between the actual update time and the theoretical update time of the post-sampling sensor is limited. When the difference is less than 0, the timeliness evaluation value of the post-sampling sensor is 1; when the difference is greater than the average sampling time interval of the sensor, the timeliness evaluation value of the post-sampling sensor is 0. The timeliness scoring equation after the limitation is as follows: (11)

[0130]

[0131] Where d = diff(t, t').

[0132] Specifically, refer to Figure 5 The diagram illustrates a functional relationship of a timeliness scoring equation. It limits the difference between the actual update time and the theoretical update time of the post-sampling sensor, specifically between [0,1]. A condition is defined as follows: when the difference between the actual and theoretical update times of the post-sampling sensor is less than 0, or when the difference is greater than the average sampling time interval T of the post-sampling sensor... avg All of them will exceed the preset timeliness score range [0,1]. Therefore, formula (9) is transformed to obtain formula (11). The timeliness evaluation value of each sensor in the sensor group can be calculated through formula (11).

[0133] Step S107: Based on the timeliness assessment value of each sensor data in the sensor group, perform timeliness assessment on the sensor data in the sensor group.

[0134] Specifically, the timeliness assessment value of each sensor data in the sensor group is compared with the timeliness threshold.

[0135] The timeliness assessment value of each sensor data in the sensor group is compared with the timeliness threshold.

[0136] If the timeliness assessment value is greater than or equal to the timeliness threshold, it is determined that there is no abnormality in the timeliness of the sensor data;

[0137] If the timeliness assessment value is less than the timeliness threshold, it is determined that the timeliness of the sensor data is abnormal.

[0138] The embodiment of the present application provides a manufacturing big data quality identification method based on time-effect correlation rules, and the method comprises the following steps: acquiring original data recorded by each sensor in a sensor group, and performing data preprocessing on the original data to obtain first data; performing timestamp alignment on the first data to obtain second data, and retaining the first data; calculating respective corresponding correlation coefficients between data of all sensors in the sensor group according to the second data, to obtain a numerical correlation matrix of the sensor group; constructing a numerical correlation graph by taking sensors as nodes and taking target correlation coefficients as edges, wherein the target correlation coefficients are correlation coefficients with absolute values greater than or equal to a correlation coefficient threshold in the correlation coefficients; calculating respective corresponding time offsets between sensors with adjacent nodes in the numerical correlation graph according to the numerical correlation graph and the first data, and constructing a time-lag cross-correlation graph by taking sensors with adjacent nodes in the numerical correlation graph as nodes and taking time offsets as edges; determining time-effectiveness evaluation values of data of each sensor in the sensor group according to respective corresponding time offsets between two sensors in the time-lag cross-correlation graph; and performing time-effectiveness identification on sensor data in the sensor group according to the time-effectiveness evaluation values of data of each sensor in the sensor group. The original data in the sensor group is preprocessed and timestamp aligned, and the correlation coefficients between data of the sensors are calculated, so that the numerical correlation between the sensor data is determined. The time-lag cross-correlation between the data is determined according to the preprocessed original data and the numerical correlation between the data, and finally the time-effectiveness of the sensor data is identified according to the numerical correlation and the time-lag cross-correlation between the data.

[0139] Each embodiment in the specification focuses on the difference from other embodiments, and the same or similar parts between the embodiments can be referred to each other.

[0140] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, device or computer program product. Therefore, the embodiments of the present application can be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program codes.

[0141] The computer program instructions can also be loaded onto a computer or other programmable data processing terminal apparatus to cause a series of operational steps to be performed on the computer or other programmable data processing terminal apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable terminal apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.

[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.

[0143] The computer program instructions can also be loaded onto a computer or other programmable data processing terminal apparatus to cause a series of operational steps to be performed on the computer or other programmable data processing terminal apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable terminal apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.

[0144] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the embodiments by those of skill in the art once they have the benefit of the present disclosure. Therefore, the appended claims are intended to cover all such variations and modifications as falling within the scope of the application.

[0145] Finally, it is to be understood that the terms such as first and second, and the like, herein are used only to distinguish one from another entity or action, and do not necessarily require or imply such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0146] The above provides a kind of manufacturing big data quality identification method based on time limit correlation rule, has carried out detailed introduction, the principle and implementation mode of the present application are described in this paper with specific examples, the above example is only for helping to understand the method of the present application and its core idea;For the general technical personnel in the art, according to the idea of the present application, there will be changes in specific implementation mode and application range, as described above, the content of the specification should not be understood as the limitation of the present application.

Claims

1. A manufacturing big data quality identification method based on aging association rules, characterized in that, The method comprises: obtaining raw data recorded by each sensor in a sensor group, and performing data preprocessing on the raw data to obtain first data; aligning the first data with timestamps to obtain second data, and retaining the first data; calculating respective corresponding correlation coefficients between all sensors in the sensor group according to the second data, to obtain a numerical correlation matrix of the sensor group; constructing a numerical correlation graph with sensors as nodes and target correlation coefficients as edges according to the numerical correlation matrix, wherein the target correlation coefficients are correlation coefficients in the correlation coefficients whose absolute values are greater than or equal to a correlation coefficient threshold; calculating respective corresponding time offsets between sensors with adjacent nodes in the numerical correlation graph according to the numerical correlation graph and the first data, and constructing a time-lag cross-correlation graph with sensors with adjacent nodes in the numerical correlation graph as nodes and time offsets as edges; determining a time effectiveness evaluation value of data of each sensor in the sensor group according to respective corresponding time offsets between two sensors in the time-lag cross-correlation graph; performing time effectiveness identification on sensor data in the sensor group according to the time effectiveness evaluation value of data of each sensor in the sensor group.

2. The manufacturing big data quality authentication method based on aging association rules according to claim 1, characterized in that, The method comprises: obtaining actual update time of a post-sampling sensor and actual update time of a pre-sampling sensor that is adjacent to the post-sampling sensor, wherein the post-sampling sensor represents a sensor with data updated later, and the pre-sampling sensor represents a sensor with data updated earlier; determining theoretical update time of the post-sampling sensor according to the actual update time of the pre-sampling sensor and the time offset between the post-sampling sensor and the pre-sampling sensor; determining the time effectiveness evaluation value of data of the post-sampling sensor according to the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor.

3. The manufacturing big data quality authentication method based on aging association rules according to claim 2, characterized in that, The method comprises: in a case where a difference between the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor is greater than or equal to 0 and less than or equal to an average sampling time interval of the sensor, substituting the difference between the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor into a time effectiveness score equation to obtain the time effectiveness evaluation value of data of the post-sampling sensor, wherein the time effectiveness score equation is shown in the following formula: Wherein, Q curr_tmp is the sensor timeliness evaluation value, diff(t, t') is the difference between the actual update time and the theoretical update time of the sensor, t is the actual update time of the sensor, t' is the theoretical update time of the sensor, e is a constant, T avg is the average sampling time interval of the sensor.

4. The manufacturing big data quality authentication method based on aging association rules according to claim 3, characterized in that, in a case where the difference between the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor is less than 0, the time effectiveness evaluation value of data of the post-sampling sensor is 1, and in a case where the difference between the actual update time of the post-sampling sensor and the theoretical update time of the post-sampling sensor is greater than the average sampling time interval of the sensor, the time effectiveness evaluation value of data of the post-sampling sensor is 0. 5.The time-aging association rule based manufacturing big data quality identification method according to claim 1, wherein, According to the numerical correlation graph and the first data, a time offset between each pair of sensors containing adjacent nodes in the numerical correlation graph is calculated, including: determining the sensors containing adjacent nodes from the numerical correlation graph; reading third data corresponding to the sensors containing adjacent nodes from the first data according to the numerical correlation graph; calculating a sampling frequency of a sensor corresponding to each data in the third data according to the third data; determining a time lag cross-correlation coefficient between any two sensors with the same sampling frequency according to the sampling frequency of the sensor corresponding to each data; determining a time offset between any two sensors with the same sampling frequency according to the time lag cross-correlation coefficient between any two sensors with the same sampling frequency.

6. The manufacturing big data quality authentication method based on aging association rules according to claim 5, characterized in that, The determination of the time offset between any two sensors with the same sampling frequency according to the time lag cross-correlation coefficient between any two sensors with the same sampling frequency includes: performing an overall offset on a second sensor in the any two sensors with the same sampling frequency according to a plurality of offset values in a preset offset interval based on a first sensor in the any two sensors with the same sampling frequency, to obtain a time lag cross-correlation coefficient of the first sensor and the second sensor under different offset values; determining an offset value corresponding to a time lag cross-correlation coefficient with the largest value under different offset values according to the time lag cross-correlation coefficient of the first sensor and the second sensor under different offset values; determining a time offset between the first sensor and the second sensor with the same sampling frequency according to the offset value corresponding to the time lag cross-correlation coefficient with the largest value under different offset values. 7.The time-aging association rule-based manufacturing big data quality identification method according to claim 6, wherein, The determination of the time offset between the first sensor and the second sensor with the same sampling frequency according to the offset value corresponding to the time lag cross-correlation coefficient with the largest value under different offset values includes: obtaining a time stamp of the first sensor at each recording point and a time stamp of the second sensor at each recording point under the offset value with the largest value; calculating an average value of a difference between the time stamp of the first sensor and the time stamp of the second sensor at each recording point, and determining the average value of the difference between the time stamps as the time offset between the first sensor and the second sensor. 8.The time-aging association rule based manufacturing big data quality identification method of claim 5, wherein, The calculation of the sampling frequency of the sensor corresponding to each data in the third data according to the third data includes: obtaining a sampling time interval of a data point corresponding to each data in the third data; calculating an average sampling time interval according to the sampling time interval of the data point corresponding to each data in the third data; calculating the sampling frequency of the sensor corresponding to each data in the third data according to the average sampling time interval. 9.The time-aging association rule based manufacturing big data quality identification method of claim 1, wherein, The calculating, according to the second data, of respective correlation coefficients between data of all sensors in the sensor group in pairs to obtain a numerical correlation matrix of the sensor group comprises: dividing the second data into multiple sub-sequences, and calculating correlation coefficients between data in each sub-sequence; obtaining the numerical correlation matrix of the sensor group by using a covariance matrix according to the correlation coefficients between data in each sub-sequence. 10.The time-aging association rule based manufacturing big data quality identification method according to any one of claims 1-9, wherein, The timeliness identification of sensor data in the sensor group according to the timeliness evaluation value of each sensor data in the sensor group comprises: comparing the timeliness evaluation value of each sensor data in the sensor group with a timeliness threshold value respectively; in the case that the timeliness evaluation value is greater than or equal to the timeliness threshold value, determining that the timeliness of the sensor data is normal; in the case that the timeliness evaluation value is less than the timeliness threshold value, determining that the timeliness of the sensor data is abnormal.

Citation Information

Patent Citations

  • System and method for analyzing force sensor data

    CA3176028A1

  • Signal processor, signal processing method, program, and recording medium

    JP2005018536A