Industrial data acquisition and analysis system and method
Through sliding window algorithm and machine learning, the data association model is constructed, combined with the multi-stage exception judgment mechanism, the problem of insufficient accuracy and reliability of data repair in the existing technology is solved, and more efficient data repair and abnormal fluctuation point recognition is achieved, improving the accuracy and sensitivity of data processing.
Patent Information
- Application Number
- CN202510470352.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When facing a complex production environment, existing industrial data repair methods are susceptible to outliers, resulting in one-sided prediction results, lack of identification of typical data state, and fail to fully consider the multiple relationships of data, which reduces the accuracy and reliability of data repair.
The sliding window algorithm is used to collect data, combine machine learning and data mining technology to build a data correlation model, and generate the first and second inferred values by minimizing the weighted error function, comprehensively considering the overall distribution and association relationship of data points, and a multi-stage anomaly judgment mechanism is used to identify and repair data missing or errors, and combine the Kalman filtering algorithm to remove noise interference.
It improves the accuracy and reliability of data repair, can better reflect the overall trend of data, reduce storage costs, enhance the ability to capture data changes, and reduce the risk of misjudgment.
Smart Images

Figure CN120276398A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial data processing, and particularly relates to an industrial data acquisition and analysis system and method. Background Art
[0002] Industrial data refers to various types of data generated and collected in the industrial field, covering various information in the entire life cycle process from product design, production manufacturing, operation management to after-sales service. In the field of industrial data acquisition and analysis, with the higher requirements for the professionalism and accuracy of data application in the entire industrial manufacturing field. However, in the industrial environment, the production environment of enterprises is relatively complex, there are various interferences, and data loss or data errors are caused relatively frequently.
[0003] Existing data repair methods have certain limitations. For example, using the interpolation method for replacement only relies on individual data points for prediction. This method is easily affected by outliers and cannot explore data laws from the perspective of the overall data distribution, resulting in one-sided prediction results and unable to accurately reflect the overall trend of the data. Moreover, when determining the data points for prediction, there is a lack of effective identification of the typical state of the data, making the prediction results difficult to conform to the normal performance of the data and reducing the accuracy and reliability of the prediction. In addition, repairing data from a single perspective without comprehensively considering various correlation relationships of the data cannot comprehensively and accurately reflect the actual situation of industrial production. Therefore, the accuracy and reliability of data repair need to be improved. Summary of the Invention
[0004] To solve the technical problems in the background art, the present invention proposes an industrial data acquisition and analysis system and method.
[0005] An industrial data acquisition and analysis system proposed by the present invention includes: A data acquisition module: used to collect industrial equipment operation data and environmental data through sensors, and adopt a sliding window algorithm in the data acquisition module to set a sliding window in units of a fixed time interval; A data association model construction module: obtain the data collected by the data acquisition module and construct a data association model through a machine learning algorithm; A data detection module: identify whether data loss or error occurs through data mining techniques and statistical analysis methods; A data repair module: used to repair the missing data or error data detected by the data detection module; A data analysis module: analyze the data collected within the sliding window in the data acquisition module to obtain abnormal fluctuation points in the data.
[0006] Preferably, the data repair module includes: The first repair unit: used to generate the first speculation value for missing data or incorrect data; The second repair unit: used to generate the second speculation value for missing data or incorrect data; The data repair unit: used to take the average of the first speculation value and the second speculation value as the repair value of the missing data or incorrect data for filling or replacement.
[0007] Preferably, for the first repair unit, the generation of the first speculation value is as follows: Suppose a sliding window collects a series of industrial data points at time x1, denoted as (x1, y1), (x1, y2),..., (x1, yn), where yi is the corresponding industrial data, and i = 1, 2,..., n; Suppose that at time x1, data missing or data error occurs for y0; For each known data point (x1, yi), calculate the absolute value of the difference of each known data point : , where , that is, the average value of yi; Arrange the calculated absolute values of the differences of each known data point (x1, yi) in ascending order, and take the first k absolute values of the differences of the data points, denoted as key data points, denoted as (x1, yi1), (x1, yi2),..., (x1, yik); Calculate the weights of the k key data points , Determined according to the reciprocal of , that is: , where s = 1, 2,..., k; Suppose y0 = ; Determine the coefficients by minimizing the following weighted error function , ,..., , as follows: , where E represents the weighted error function, and t = 1, 2,..., k; Use the least squares method to solve and obtain the numerical values of the coefficients , ,..., ; Substitute the obtained numerical values of the coefficients , ,..., into the equation y0 = , calculate the predicted missing or incorrect value y0 at time x1, and take y0 as the first speculation value; Preferably, for the second repair unit, the generation of the second speculation value is as follows: Confirm the associated data through the associated data model; Assume There is data missing or data error, and obtain according to the associated data model The associated data of are respectively , ,..., ; Assume =ζ0+ζ1 +ζ2 +...+ζj + , where ζ0 is the intercept term, representing when , ,..., are all 0 The value of; ζ0, ζ1,..., ζj are respectively , ,..., The regression coefficients of are used to reflect The influence degree and direction of the associated data of on ; is the error term, used to represent the part in the formula that cannot be explained by the linear relationship. Assume Follows a normal distribution with a mean of 0; Obtain , , ,..., The historical data of, and use the least squares method to solve ζ0, ζ1, ζ2,..., ζj; Substitute the obtained values of ζ0, ζ1, ζ2,..., ζj into ζ0+ζ1 +ζ2 +...+ζj + to calculate and obtain The value of, which is the second speculation value.
[0008] Preferably, in the data analysis module, analyze whether there is abnormal fluctuation in the data within the sliding window, and the method is as follows: Calculate the mean, variance, and standard deviation of the data within the sliding window; Set the threshold range, and compare the statistics of the current window with the statistics of multiple historical windows; If one or more statistics of the current window exceed the threshold range of historical statistics, record these abnormal points as abnormal fluctuation points to be confirmed; For the abnormal fluctuation points to be confirmed that have not been judged, enter a first judgment: First smoothing line: Perform equal-weight smoothing on the data within the sliding window where the abnormal fluctuation point to be confirmed is located; Second smoothing line: The second smoothing line confirms the time decay weight by introducing a time decay coefficient, and performs weighted smoothing on the data within the sliding window where the abnormal fluctuation point to be confirmed is located according to the time decay weight; Calculate the first smoothing line value S1 and the second smoothing line value S2 of the window where the abnormal fluctuation point to be confirmed is located; If the deviation value between the abnormal fluctuation point to be confirmed and S1 and S2 is greater than or equal to the set threshold, determine the abnormal fluctuation point to be confirmed as a real abnormal fluctuation point; For the abnormal fluctuation points to be confirmed that have not been judged in the first judgment, enter a second judgment; Calculate the first smoothing line value and the second smoothing line value of the historical window data, and set the threshold T1 of the first smoothing line value of the historical window data and the threshold T2 of the second smoothing line value of the historical window data; Compare S1 and S2 of the window where the abnormal fluctuation point to be confirmed is located with T1 and T2 respectively; that is, compare S1 with T1 and S2 with T2; If S1 is greater than T1 or S2 is greater than T2, determine the abnormal fluctuation point to be confirmed as a real abnormal fluctuation point.
[0009] Preferably, in the data analysis module, when an abnormal fluctuation point appears, a safety warning is issued, and relevant personnel are notified by means of sound and light alarm, SMS notification, and system pop-up window.
[0010] Preferably, in the data acquisition module, the Kalman filter algorithm is combined to process the data within the sliding window to remove noise interference in the data.
[0011] An industrial data acquisition and analysis method includes the following steps: S1. Collect industrial equipment operation data and environmental data through sensors, and adopt a sliding window algorithm during the data collection process to set a sliding window in units of a fixed time interval; S2. Obtain the data collected in S1, and construct a data association model through a machine learning algorithm; S3. Use data mining techniques and statistical analysis methods to process the collected data to identify whether there is data loss or error; S4. Repair the missing data or error data detected in step 3; S5. Analyze the data collected within the sliding window in S1 to obtain the abnormal fluctuation points in the data.
[0012] In the present invention, the proposed industrial data acquisition and analysis system and method have the following beneficial technical effects: 1. The setting of the data repair module, the first repair unit, the generation of the first speculation value. During the calculation process, the average value of all known data points is considered, and based on this, the difference between each data point and the whole is measured. It can explore the data law from the perspective of the overall data distribution, avoiding the one-sidedness of only relying on individual data points for prediction, and making the prediction result better reflect the overall trend of the data. When determining the key data points, select the points with smaller differences from the average value. These points can better represent the normal or typical state of industrial data at the current time. It can make the prediction result more inclined to conform to the regular performance of the data, improving the accuracy and reliability of the prediction. By minimizing the weighted error function to determine the coefficients, a weighted linear combination relationship based on key data points is constructed. This method can automatically adjust the contribution degree of each key data point to the prediction result according to the characteristics and mutual relationships of the data itself, making the prediction model more in line with the actual data situation and enhancing the rationality and adaptability of the prediction; The second repair unit, the generation of the second speculation value. Confirm the associated data through the associated data model; obtain the speculation value of the missing data or incorrect data through the analysis of the associated data, improving the data accuracy of the second speculation value, and taking the average value of the first speculation value and the second speculation value as the repair value of the missing data or incorrect data for filling or replacement. Combining the advantages of the first repair unit and the second repair unit, balancing the deviation of the repair methods of the first repair unit and the second repair unit, and further improving the data accuracy and reliability.
[0013] 2. The setting of the data analysis module. Analyze whether there are abnormal fluctuations in the data within the sliding window. First, calculate the mean, variance, and standard deviation of the data within the sliding window, and compare the statistical quantities of the current window with those of multiple historical windows. Compared with single statistical quantity analysis, it can identify abnormal fluctuations in the data more comprehensively and accurately. Through a multi-stage abnormal judgment mechanism, perform a primary judgment and a secondary judgment on the abnormal fluctuation points to be confirmed in turn. In the primary judgment, calculate the first smoothed line value S1 and the second smoothed line value S2, and compare the deviation with the abnormal fluctuation points to be confirmed; in the secondary judgment, compare S1 and S2 of the window where the abnormal fluctuation points to be confirmed are located with the thresholds T1 and T2 of the historical window data respectively. This multi-stage judgment method can gradually screen and confirm the abnormal fluctuation points, avoiding misjudgment caused by the limitations of single judgment and improving the accuracy of judgment.
[0014] 3. In the data analysis module, two different smoothing line processing methods are introduced. The first smoothing line adopts equal-weight smoothing processing, which can reflect the medium- and long-term trend characteristics of the data, enabling analysts to understand the overall change trend of the data over a relatively long time range. The second smoothing line performs weighted smoothing by introducing a time decay coefficient, highlighting the influence of recent data and enhancing the response speed to trend changes. Compared with the first smoothing line, it can quickly capture the short-term change trend of the data. The second smoothing line assigns exponentially increasing weights to the latest data by introducing a time decay coefficient. This weight assignment method enables the smoothing line to more sensitively reflect the latest changes in the data. Especially when the data changes rapidly, it can quickly adjust the value of the smoothing line, thereby more accurately capturing the change trend of the data.
[0015] 4. In the data acquisition module, a sliding window algorithm is adopted and the sliding window is set in units of a fixed time interval. The sliding window algorithm can perform real-time processing on the data stream, and the sliding window only retains the data within the current window. As the window slides, the old data will be automatically removed. This greatly reduces the amount of data that needs to be stored and lowers the storage cost.
[0016] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Brief Description of the Drawings
[0017] Figure 1 is the principle block diagram of the system of the present invention; Figure 2 is the flowchart of the method of the present invention. Detailed Embodiments
[0018] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference signs denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0019] As Figure 1 shown, an industrial data acquisition and analysis system includes: A data acquisition module: used to collect industrial equipment operation data and environmental data through sensors, and adopt a sliding window algorithm in the data acquisition module, setting the sliding window in units of a fixed time interval. The data within the window is used for real-time analysis and decision-making; continuously sliding the window to continuously collect and update data, accurately capturing the real-time change trend of the data, and improving the accuracy of data acquisition; In the data acquisition module, a sliding window algorithm is adopted and a sliding window is set in units of a fixed time interval. The sliding window algorithm can process the data stream in real time, and the sliding window only retains the data within the current window. As the window slides, the old data will be automatically removed. This greatly reduces the amount of data that needs to be stored and lowers the storage cost.
[0020] In an optional embodiment, in the data acquisition module, the data within the sliding window is processed in combination with the Kalman filtering algorithm to remove the noise interference in the data and further improve the accuracy of the data; the Kalman filtering algorithm predicts the current state based on the previous measurement data and combines it with the current measurement value for optimization to obtain a more accurate data estimate value.
[0021] Data association model construction module: Obtain the data collected by the data acquisition module and construct a data association model through machine learning algorithms; Data detection module: Identify whether there is data loss or error through data mining techniques and statistical analysis methods; Data repair module: Used to repair the missing data or incorrect data detected by the data detection module; The data repair module includes: The first repair unit: Used to generate the first speculation value of the missing data or incorrect data; The second repair unit: Used to generate the second speculation value of the missing data or incorrect data; Data repair unit: Used to take the average of the first speculation value and the second speculation value as the repair value of the missing data or incorrect data for filling or replacement; In an optional embodiment, for the first repair unit, the generation of the first speculation value is as follows: Suppose a series of industrial data points are collected at time x1 in the sliding window, denoted as (x1, y1), (x1, y2),..., (x1, yn), where yi is the corresponding industrial data, i = 1, 2,..., n; Industrial data includes the output value and operating parameters of the device; Suppose that at time x1, y0 has data loss or data error; For each known data point (x1, yi), calculate the absolute value of the difference of each known data point : , where , that is, the average value of yi; Arrange the calculated absolute values of the differences of each known data point (x1, yi) in ascending order, and take the first k absolute values of the differences The data points, designated as key data points, are denoted as (x1, yi1), (x1, yi2),..., (x1, yik); Calculate the weights of the k key data points , Determined according to the reciprocal of , that is: , where s = 1, 2,..., k; thus, the points with smaller differences from the average value have larger weights; Let y0 = ; Determine the coefficients by minimizing the following weighted error function , ,..., , as follows: , where E represents the weighted error function and t = 1, 2,..., k; Use the least squares method to solve for the values of the coefficients , ,..., ; Substitute the obtained values of the coefficients , ,..., into the equation y0 = , calculate the predicted missing or incorrect value y0 at time x1, and take y0 as the first estimated value; In an alternative embodiment, for the second repair unit, the generation of the second estimated value is as follows: Confirm the associated data through the associated data model; Let have data missing or data error, and obtain the associated data of as , ,..., ; Let = ζ0 + ζ1 + ζ2 +... + ζj + , where ζ0 is the intercept term, representing the value of , ,..., when all are 0 ; ζ0, ζ1,..., ζj are respectively , ,..., regression coefficients, used to reflect associated data pairs degree and direction of influence; is the error term, used to represent the part in the formula that cannot be explained by the linear relationship. Let follow a normal distribution with a mean of 0; Obtain , , , ..., historical data of, and solve ζ0, ζ1, ζ2, ..., ζj using the least squares method; Substitute the obtained values of ζ0, ζ1, ζ2, ..., ζj into = ζ0 + ζ1 + ζ2 + ... + ζj + to calculate and obtain the value of, which is the second estimated value; Settings of the data repair module, generation of the first repair unit and the first estimated value. The calculation process takes into account the average value of all known data points and measures the difference between each data point and the whole based on this, enabling the discovery of data patterns from the perspective of the overall data distribution, avoiding the one-sidedness of relying only on individual data points for prediction, and making the prediction result better reflect the overall trend of the data.
[0022] When determining the key data points, select the points with smaller differences from the average value. These points can better represent the normal or typical state of industrial data at the current time, making the prediction result more inclined to conform to the regular performance of the data and improving the accuracy and reliability of the prediction.
[0023] Determine the coefficients by minimizing the weighted error function, constructing a weighted linear combination relationship based on key data points. This method can automatically adjust the contribution degree of each key data point to the prediction result according to the characteristics and mutual relationships of the data itself, making the prediction model more in line with the actual data situation and enhancing the rationality and adaptability of the prediction.
[0024] The second repair unit and generation of the second estimated value. Confirm the associated data through the associated data model; obtain the estimated values of the missing data or incorrect data through the analysis of the associated data, improve the data accuracy of the second estimated value, and take the average value of the first estimated value and the second estimated value as the repair value of the missing data or incorrect data for filling or replacement, integrating the advantages of the first repair unit and the second repair unit, balancing the deviation of the repair methods of the first repair unit and the second repair unit, and further improving the data accuracy and reliability.
[0025] Data analysis module: Analyze the data collected within the sliding window in the data collection module to obtain abnormal fluctuation points in the data; In an optional embodiment, in the data analysis module, to analyze whether there are abnormal fluctuations in the data within the sliding window, the method is as follows: Calculate the mean, variance, and standard deviation of the data within the sliding window; Set a threshold range, and compare the statistics of the current window with the statistics of multiple historical windows; If one or more statistics of the current window exceed the threshold range of the historical statistics, and record these abnormal points as abnormal fluctuation points to be confirmed; For the abnormal fluctuation points to be confirmed that have not been judged, enter a primary judgment: First smoothing line: Perform equal-weight smoothing on the data within the sliding window where the abnormal fluctuation point to be confirmed is located. This design reflects the medium- and long-term trend characteristics of the data. Equal-weight smoothing is a data processing method, which means that when smoothing the data, the same weight is assigned to each data point participating in the calculation, that is, the contribution degree of each data point in the smoothing result is the same.
[0026] Second smoothing line: The second smoothing line confirms the time decay weight by introducing a time decay coefficient, and performs weighted smoothing on the data within the sliding window where the abnormal fluctuation point to be confirmed is located according to the time decay weight. This design highlights the influence of recent data and enhances the response speed to trend changes.
[0027] The second smoothing line makes the sensitive smoothing line assign exponentially increasing weights to the latest data by introducing a time decay coefficient, so as to quickly capture data changes.
[0028] The time decay coefficient refers to the factor used to measure the influence of time on data in data analysis, and it is achieved by assigning weights to historical data; Calculate the first smoothing line value S1 and the second smoothing line value S2 of the window where the abnormal fluctuation point to be confirmed is located; If the deviation value between the abnormal fluctuation point to be confirmed and S1 and S2 is greater than or equal to the set threshold, determine the abnormal fluctuation point to be confirmed as a real abnormal fluctuation point; For the abnormal fluctuation points to be confirmed that have not been judged, enter a secondary judgment; The secondary judgment is: Calculate the first smoothing line value and the second smoothing line value of the historical window data, and set the threshold T1 of the first smoothing line value of the historical window data and the threshold T2 of the second smoothing line value of the historical window data; Compare S1 and S2 of the window where the abnormal fluctuation point to be confirmed is located with T1 and T2 respectively; that is, compare S1 with T1 and S2 with T2; If S1 is greater than T1 or S2 is greater than T2, then determine that the abnormal fluctuation point to be confirmed is a true abnormal fluctuation point; otherwise, do not process the abnormal fluctuation point to be confirmed; In an optional embodiment, in the data analysis module, when an abnormal fluctuation point appears, a safety warning is given, and relevant personnel are notified by means of sound and light alarm, SMS notification, and system pop-up window; In the setting of the data analysis module, it analyzes whether there are abnormal fluctuations in the data within the sliding window. First, it calculates statistical quantities such as the mean, variance, and standard deviation of the data within the sliding window, and compares the statistical quantities of the current window with those of multiple historical windows. By comprehensively considering multiple statistical indicators, it can describe the distribution characteristics and changes of the data from different dimensions. Compared with single-statistic analysis, it can identify abnormal fluctuations in the data more comprehensively and accurately. It adopts a multi-stage abnormal judgment mechanism and conducts a primary judgment and a secondary judgment on the abnormal fluctuation point to be confirmed in sequence. In the primary judgment, the first smoothed line value S1 and the second smoothed line value S2 are calculated and compared with the abnormal fluctuation point to be confirmed for deviation; in the secondary judgment, S1 and S2 of the window where the abnormal fluctuation point to be confirmed is located are respectively compared with the thresholds T1 and T2 of the historical window data. This multi-stage judgment method can gradually screen and confirm abnormal fluctuation points, avoiding misjudgment caused by the limitations of single judgment and improving the accuracy of judgment.
[0029] In the data analysis module, two different smoothed line processing methods are introduced. The first smoothed line adopts equal-weight smoothing processing, which can reflect the medium- and long-term trend characteristics of the data, enabling analysts to understand the overall change trend of the data in a relatively long time range. The second smoothed line performs weighted smoothing by introducing a time decay coefficient, highlighting the influence of recent data, enhancing the response speed to trend changes, and being able to quickly capture the short-term change trend of the data. The second smoothed line assigns exponentially increasing weights to the latest data by introducing a time decay coefficient. This weight allocation method enables the smoothed line to more sensitively reflect the latest changes in the data. Especially when the data changes rapidly, it can quickly adjust the value of the smoothed line, thereby more accurately capturing the change trend of the data.
[0030] Such as Figure 2 shown in an industrial data acquisition and analysis method, including the following steps: S1. Collect industrial equipment operation data and environmental data through sensors, and adopt a sliding window algorithm during the data collection process, setting a sliding window in units of a fixed time interval; S2. Obtain the data collected in S1 and construct a data association model through machine learning algorithms; S3. Use data mining techniques and statistical analysis methods to process the collected data and identify whether there are data missing or errors; S4. Repair the missing data or incorrect data detected in step S3; S5. Analyze the data collected within the sliding window in S1 to obtain the abnormal fluctuation points in the data.
[0031] Meanwhile, the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0032] In the embodiments provided by the present invention, it should be understood that the disclosed system or method can be implemented in other ways. For example, the above-described embodiments of the invention are merely illustrative. For example, the division of modules is only a logical function division, and there may be other division methods in actual implementation.
[0033] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0034] In addition, in each embodiment of the present invention, the functional modules can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.
[0035] For those skilled in the field of operation and maintenance, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and can be implemented in other specific forms without departing from the basic characteristics of the present invention.
[0036] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. An industrial data acquisition and analysis system, characterized in that, Including: Data acquisition module: used to collect the operating data and environmental data of industrial equipment through sensors, and adopt a sliding window algorithm in the data acquisition module to set a sliding window with a fixed time interval as the unit; Data association model construction module: obtains the data collected by the data acquisition module and constructs a data association model through machine learning algorithms; Data detection module: identifies whether data loss or errors occur through data mining techniques and statistical analysis methods; Data repair module: used to repair the missing data or incorrect data detected by the data detection module; Data analysis module: analyzes the data collected within the sliding window in the data acquisition module to obtain abnormal fluctuation points in the data.
2. The industrial data acquisition and analysis system according to claim 1, wherein The data repair module includes: The first repair unit: used to generate the first estimated value of the missing data or incorrect data; The second repair unit: used to generate the second estimated value of the missing data or incorrect data; Data repair unit: used to take the average of the first estimated value and the second estimated value as the repair value of the missing data or incorrect data for filling or replacement.
3. The industrial data acquisition and analysis system according to claim 2, wherein The first repair unit, the generation of the first estimated value is as follows: Suppose a series of industrial data points are collected at the same time x1 in the sliding window, denoted as (x1, y1), (x1, y2),...,(x1, yn), where yi is the corresponding industrial data, i = 1, 2,..., n; Suppose that at time x1, y0 has data loss or data error; For each known data point (x1, yi), calculate the absolute value of the difference for each known data point : , where , i.e., the average value of yi; The absolute value of the difference for each calculated known data point (x1, yi) is arranged in ascending order, and the first k absolute values of the differences are taken of the data points, which are set as key data points and denoted as (x1, yi1), (x1, yi2), ..., (x1, yik); Calculate the weights of k key data points , Determined according to the reciprocal of, that is: , where s = 1, 2, ..., k; Let y0 = ; The coefficients are determined by minimizing the following weighted error function , , ..., , as follows: , where E represents the weighted error function, t = 1, 2, ..., k; The coefficients are obtained by using the least squares method , ,..., , and their numerical values; The coefficients , ,..., are substituted into the equation y0 = to calculate the predicted missing or incorrect value y0 at time x1, and y0 is taken as the first speculation value.
4. The industrial data acquisition and analysis system according to claim 3, wherein The second repair unit, the generation of the second estimated value is as follows: Confirm associated data through the associated data model; Suppose data loss or data error has occurred, and the associated data of are respectively , ,..., ; Let = ζ0 + ζ1 + ζ2 +... + ζj + , where ζ0 is the intercept term, representing the value when , , ..., are all 0 ; ζ0, ζ1, ..., ζj are the regression coefficients of , , ..., respectively, which are used to reflect the influence degree and direction of the correlation data pair of on ; is an error term used to represent the part in the formula that cannot be explained by the linear relationship. Let follow a normal distribution with a mean of 0; Obtain , , ,... of historical data, and use the least squares method to solve for ζ0, ζ1, ζ2,..., ζj; Substitute the obtained values of ζ0, ζ1, ζ2, ..., ζj into = ζ0 + ζ1 + ζ2 +... + ζj + to calculate the value, which is the second estimated value.
5. The industrial data acquisition and analysis system according to claim 1, characterized in that, In the data analysis module, analyze whether there are abnormal fluctuations in the data within the sliding window, and the method is as follows: Calculate the mean, variance, and standard deviation of the data within the sliding window; Set a threshold range, and compare the statistical quantities of the current window with the statistical quantities of multiple historical windows; If one or more statistical quantities of the current window exceed the threshold range of the historical statistics, and record these abnormal points as abnormal fluctuation points to be confirmed; For the abnormal fluctuation points to be confirmed that have not been judged, enter a first judgment: The first smoothing line: perform equal-weight smoothing on the data within the sliding window where the abnormal fluctuation point to be confirmed is located; The second smoothing line: the second smoothing line confirms the time decay weight by introducing a time decay coefficient, and performs weighted smoothing on the data within the sliding window where the abnormal fluctuation point to be confirmed is located according to the time decay weight; Calculate the first smoothing line value S1 and the second smoothing line value S2 of the window where the abnormal fluctuation point to be confirmed is located; If the deviation value of the abnormal fluctuation point to be confirmed from S1 and S2 is greater than or equal to the set threshold, determine that the abnormal fluctuation point to be confirmed is a real abnormal fluctuation point; For the abnormal fluctuation points to be confirmed that have not been judged in the first judgment, enter a second judgment; Calculate the first smoothing line value and the second smoothing line value of the historical window data, and set the threshold T1 of the first smoothing line value of the historical window data and the threshold T2 of the second smoothing line value of the historical window data; Compare S1 and S2 of the window where the abnormal fluctuation point to be confirmed is located with T1 and T2 respectively; that is, compare S1 with T1 and S2 with T2; If S1 is greater than T1 or S2 is greater than T2, then it is determined that the abnormal fluctuation point to be confirmed is a true abnormal fluctuation point.
6. The industrial data acquisition and analysis system according to claim 1 or 5, characterized in that, In the data analysis module, when an abnormal fluctuation point appears, a safety warning is issued, and relevant personnel are notified by means of audible and visual alarms, text message notifications, and system pop-ups.
7. The industrial data acquisition and analysis system according to claim 1, wherein In the data acquisition module, the data within the sliding window is processed in combination with the Kalman filter algorithm to remove the noise interference in the data.
8. Using the industrial data acquisition and analysis method according to any one of claims 1-7, characterized in that, It includes the following steps: S1. Collect the operation data and environmental data of industrial equipment through sensors. During the data collection process, the sliding window algorithm is adopted, and the sliding window is set in units of a fixed time interval. S2. Obtain the data collected in S1 and construct a data association model through machine learning algorithms. S3. Use data mining techniques and statistical analysis methods to process the collected data to identify whether there is data loss or error. S4. Repair the missing data or incorrect data detected in step 3. S5. Analyze the data collected within the sliding window in S1 to obtain the abnormal fluctuation points in the data.
Citation Information
Patent Citations
Data restoration method and system for highway infrastructure multi-source heterogeneous data
CN109992579A
Urban traffic flow prediction method
CN115619052A
Anemograph fault early warning method and equipment based on adaptive sliding window division
CN116611244A
PMSM multi-mode switching model prediction control method and system and storage medium
CN117254734A
Vehicle comfort real-time evaluation method and system based on local vibration information
CN118196930A