Industrial high-quality management method and system taking compressed air quality data as core
By collecting and analyzing compressed air quality data, and constructing correlation and prediction models, the problem of insufficient compressed air quality monitoring in traditional manufacturing industries has been solved. This has enabled precise identification of abnormal causes and optimized decision-making, thereby improving production efficiency and quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-10
AI Technical Summary
In traditional manufacturing, monitoring data during production and operation often focuses on single energy consumption data or single production indicators, making it impossible to reduce costs and increase efficiency by controlling compressed air quality.
Data from the air compressor system, production environment, and production system are collected, a correlation model is built, the correlation between abnormal data and other data is calculated, abnormal data is located and optimization decisions are output, the correlation relationship is displayed through charts, a predictive model is constructed to predict future anomalies, and accurate cause location and optimization decision-making are achieved.
It enables precise identification and efficient rectification of industrial production anomalies, helping enterprises improve manufacturing processes, prevent production losses, and enhance energy consumption management and product quality.
Smart Images

Figure CN121836478A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of compressed air quality-related data set processing technology, and in particular to industrial high-quality management methods and systems based on compressed air quality data. Background Technology
[0002] Modern industrial manufacturing extensively utilizes automated equipment and advanced production processes, with electricity and gas serving as the fundamental power sources for their operation. Compressed air is a fundamental power source in modern industrial manufacturing, ranking second only to electricity. While compressed air contains harmful substances such as water, oil, and dust, oil and dust can be removed through filters. However, water is the most difficult substance to handle and poses the greatest threat to automated industrial production. Excessive water content leads to poor quality, equipment malfunctions, and damage, increasing defect rates, downtime, and overall energy consumption per unit of production. Currently, domestic enterprises are largely unaware of this aspect. Therefore, establishing a comprehensive real-time production management monitoring system, and through real-time monitoring and analysis of key data such as pressure and dew point, is crucial for improving manufacturing processes and preventing production losses.
[0003] Currently, monitoring data in the production and operation process of traditional manufacturing industries mostly focus on single energy consumption data or single production indicators, which cannot meet the needs of cost reduction and efficiency improvement by controlling compressed air quality. Summary of the Invention
[0004] The technical problem this application aims to solve is that, currently, monitoring data in the production and operation process of traditional manufacturing industries mostly focus on monitoring single energy consumption data or single production indicators, and cannot achieve the need for cost reduction and efficiency improvement by controlling compressed air quality.
[0005] To address, or at least partially address, the aforementioned technical problems, this application provides an industrial high-quality management method and system centered on compressed air quality data.
[0006] In a first aspect, this invention discloses an industrial high-quality management method based on compressed air quality data, comprising: Collect data related to the air compressor system, production environment data, and data generated by the production system, and preprocess the data to obtain a collection of data, which includes compressed air parameters, environmental parameters, and production process parameters. Retrieve abnormal data from the collected data set, construct a correlation model between the abnormal data and other data in the set one by one, and calculate the correlation between the abnormal data and other data in the set to obtain the correlation relationship between the abnormal data and the set data, and the correlation data between the abnormal data and the set data. Based on the correlation between abnormal data and aggregate data, we determine the data in the aggregate data that are related to the abnormal data to obtain fluctuation data. Based on the location of fluctuation data collection and the correlation between abnormal data and aggregate data, we output optimization decisions.
[0007] Preferably, the process involves collecting data related to the compressed air system, production environment data, and data generated by the production system, and preprocessing the data to obtain a collected data set. This collected data set includes compressed air parameters, environmental parameters, and production process parameters. Specifically, this includes the following steps: A parameter detection device is placed at the compressed air output port. The parameter detection device includes a dew point meter, an oil content analyzer, a laser particle analyzer, a flow meter, and a pressure sensor. It collects compressed air pressure dew point data, oil content data, particle size index data, flow rate data, and pressure data to obtain compressed air parameters. Environmental monitoring equipment, including temperature and humidity sensors and particle detectors, is installed in the product manufacturing workshop to collect temperature, humidity and particulate information in the production workshop environment and obtain environmental parameters. The backend server collects raw data, including abnormal material requisition cost data, finished product output value, defective product quantity, and downtime data. The raw data collection methods include obtaining data from ERP systems, MES systems, and equipment management systems, or directly inputting and generating data. After calculation, material loss values, equipment downtime, and unit production energy consumption values are obtained, thus yielding production process parameters.
[0008] Preferably, the step of retrieving abnormal data from the collected data set, constructing a correlation model between the abnormal data and other data in the set one by one, and calculating the correlation between the abnormal data and other data in the set to obtain the correlation relationship between the abnormal data and the set data, specifically includes the following steps: Extract abnormal data from the collected data set, and construct linear or nonlinear correlation models between the abnormal data and other data in the set to obtain the correlation relationship between the abnormal data and the set data. The correlation between the outlier data and the data constructed in the correlation model is calculated to obtain the correlation data between the outlier data and the set data.
[0009] Preferably, the step of retrieving abnormal data from the collected data set and constructing a linear or nonlinear correlation model between the abnormal data and other data in the set specifically includes the following steps: Determine whether there is a linear relationship between outlier data and other data in the set, and divide the data into data with linear relationships and data with non-linear relationships; Construct a linear regression analysis model to establish a linear relationship among data that have a linear relationship, and obtain the linear regression relationship of the data. A nonlinear model is constructed based on the mutual information model and the random forest model. Nonlinear relationships are then constructed from data that have nonlinear relationships, thus obtaining the nonlinear relationships of the data.
[0010] Preferably, the step of judging the data in the set data that are associated with the abnormal data based on the correlation between the abnormal data and the set data to obtain the fluctuation data, and outputting optimization decisions based on the location of the fluctuation data collection and the correlation between the abnormal data and the set data, specifically includes the following steps: Sort the data according to their correlation values, and then sort the aggregated data according to their correlation. Set a data threshold and capture the correlation data that is higher than the data threshold to obtain the fluctuation data. Based on the location of fluctuation data collection, the abnormal triggering device is located. By combining the abnormal triggering device with the correlation between abnormal data and aggregate data, the system adjustment actions are planned to obtain optimization decisions. The enterprise then adjusts the system and system parameters according to the optimization decisions.
[0011] Preferably, the following steps are then included: The relationship between abnormal data and the collected data set is displayed through charts.
[0012] Preferably, the following steps are then included: Construct a predictive model, input the real-time collected data set, predict the compressed air quality data for a future period, filter out abnormal data, filter out fluctuation data related to abnormal data, analyze abnormal equipment or environment based on fluctuation data, and output targeted future optimization decisions.
[0013] Secondly, the present invention discloses an industrial high-quality management system based on compressed air quality data, including the industrial high-quality management method based on compressed air quality data as described in any of the above claims, which includes: The data acquisition unit is used to acquire a set of collected data, including compressed air parameters, environmental parameters, and production process parameters. The data analysis and diagnosis unit extracts abnormal data from the data collection set acquired by the data acquisition unit, and performs correlation analysis on the abnormal data and other data in the set, including constructing a correlation model and calculating correlation data, to obtain the correlation between fluctuation data, abnormal data and set data, as well as the correlation data between abnormal data and set data, and locates the device that causes the abnormality based on the fluctuation data; The decision output unit combines the abnormal cause devices with the correlation between abnormal data and aggregate data to plan system adjustment actions and obtain optimized decisions.
[0014] Preferably, the data analysis and diagnostic unit includes: The association model construction sub-unit constructs linear and non-linear association models with abnormal data and other data in the set to obtain the association relationship between abnormal data and other data. The correlation calculation subunit calculates the correlation data between abnormal data and other data. It judges the strength of the correlation between abnormal data and other data through the correlation data, presets a data threshold, and captures the correlation data that exceeds the data threshold to obtain the fluctuation data. The trigger device positioning subunit is a device that collects positioning data through fluctuation data.
[0015] Preferably, it further includes: The data prediction unit constructs a data prediction model and uses real-time collected data sets to predict compressed air quality-related data for a future period of time.
[0016] The technical solution provided in this application has the following advantages compared with the prior art: This application provides an industrial high-quality management method and system centered on compressed air quality data. The method involves acquiring a dataset including compressed air parameters, environmental parameters, and production process parameters; retrieving abnormal data from the dataset; constructing correlation models between the abnormal data and other data in the dataset; calculating the correlation between the data; filtering data with strong correlation to the abnormal data to obtain fluctuation data; and establishing correlations between the fluctuation data, its acquisition location, and other data to output targeted equipment adjustment and optimization decisions. This enables precise identification and efficient rectification of industrial production anomalies. By analyzing abnormal data in the acquired dataset, calculating highly correlated fluctuation data, and locating the causes of anomalies through the relationships between fluctuation data and other data, and then outputting optimization decisions based on these causes, enterprises can improve manufacturing processes and prevent production losses.
[0017] The system provided by this invention offers modern industrial manufacturing operators real-time monitoring of key data such as energy consumption, output value, defect rate, and equipment failure rate in the production process. It also establishes a compressed air quality data monitoring software and hardware system that significantly impacts these data indicators. The main hardware includes a dew point sensor for detecting moisture content (pressure dew point), and smart meters for detecting energy consumption, such as electricity and water meters. This system helps enterprises establish a comprehensive real-time production management monitoring system. Furthermore, through real-time monitoring and analysis of key data such as pressure dew point, it helps enterprises improve manufacturing processes and prevent production losses. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating an intelligent predictive decision-making method for compressed air production provided in this application. Figure 1 ; Figure 2 A flowchart illustrating an intelligent predictive decision-making method for compressed air production provided in this application. Figure 2 ; Figure 3 A schematic diagram of the specific process of step S1 of the intelligent predictive decision-making method for compressed air production provided in this application; Figure 4 A detailed flowchart illustrating step S2 of the intelligent predictive decision-making method for compressed air production provided in this application; Figure 5 A schematic diagram of step S3 of the intelligent predictive decision-making method for compressed air production provided in this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] Firstly, see Figures 1-5 This invention discloses an industrial high-quality management method based on compressed air quality data, comprising: Step S1: Collect data related to the air compressor system, production environment data, and data generated by the production system, and preprocess the data to obtain a collection of data, which includes compressed air parameters, environmental parameters, and production process parameters. Step S2: Retrieve abnormal data from the collected data set, construct a correlation model between the abnormal data and other data in the set one by one, and calculate the correlation between the abnormal data and other data in the set to obtain the correlation relationship between the abnormal data and the set data, and the correlation data between the abnormal data and the set data. Step S3: Based on the correlation between abnormal data and aggregate data, determine the data in the aggregate data that are related to the abnormal data to obtain fluctuation data. Based on the location of fluctuation data collection and the correlation between abnormal data and aggregate data, output optimization decisions.
[0023] Specifically, in step S1, the collected dataset is acquired and preprocessed. First, compressed air parameters (pressure dew point, oil content, particle size, flow rate, and pressure) are collected using devices such as dew point meters, oil content analyzers, flow sensors, and pressure sensors. Environmental parameters (workshop temperature and humidity, environmental particle size) are collected using integrated sensors. Production process parameters (material loss values, equipment failure time, and unit production energy consumption) are obtained by connecting to ERP and MES systems. In addition, the system can collect values from smart water meters and smart electricity meters, which can calculate real-time energy consumption values. Then, the collected data is preprocessed, including removing outlier values and filling in missing values, to obtain a compressed air quality pipeline data set. The data covers the entire compressed air, environment, and production chain, ensuring the integrity of the collected data set. The preprocessing process is tailored to the characteristics of industrial data, effectively filtering out problematic data such as sensor failures and system delays, improving the analysis efficiency and accuracy of subsequent steps.
[0024] Specifically, in step S2, based on industry standards and enterprise stable period data, abnormal data is retrieved from the dataset. This abnormal data can be extracted in real-time by the system or collected in real-time by setting a threshold. Then, the abnormal data is paired with other data in the dataset one by one to construct a correlation model. The Pearson correlation coefficient is used to calculate linear correlation (|r|≥0.6 indicates strong linear correlation), and mutual information values are used to capture nonlinear correlation (mutual information >0.4 indicates strong nonlinear correlation). Finally, for all constructed models, the correlation strength value between abnormal data and correlated data is calculated, i.e., the correlation data. The final output is a table of correspondence between abnormal data and correlated data (e.g., abnormal material loss corresponds to pressure dew point fluctuations), and correlation data including correlation type (linear / nonlinear) and correlation strength value. The combination of linear and nonlinear relationship models takes into account both linear and nonlinear correlation scenarios, avoiding the analytical limitations of a single model. The use of correlation strength values quantifies the correlation data, resulting in highly interpretable results. It allows for the determination of the correlation strength and impact between abnormal data and fluctuating data. The quantitative results facilitate subsequent priority judgment, and the one-to-one pairing analysis method ensures that no potential correlated data is missed.
[0025] Specifically, in step S3, a correlation strength threshold is set (linear |r|≥0.6 or nonlinear mutual information value>0.4) to filter out aggregate data strongly correlated with abnormal data as fluctuation data; the cause is located based on the collection location of the fluctuation data, directly locating the data collection equipment; optimization decisions are output based on the correlation relationship, with single fluctuation data corresponding to equipment parameter adjustments, and multiple fluctuation data corresponding to multi-equipment collaborative adjustments, clearly defining equipment names, operating procedures, and verification standards. This precise location of abnormal causes transforms abstract correlation data into implementable equipment adjustment solutions, achieving a closed-loop connection between anomaly identification, cause location, and decision execution. Fluctuation data is filtered based on correlation strength, ensuring accurate cause location and avoiding blind investigation; decisions are deeply bound to the fluctuation data collection location and correlation relationship, resulting in strong targeting and high rectification efficiency; and it supports multi-cause collaborative adjustment scenarios, adapting to complex industrial production environments.
[0026] It is understandable that by collecting and preprocessing compressed air parameters, environmental parameters, and production process parameters, identifying abnormal data, and constructing linear and nonlinear models one by one with other data in the dataset, the correlation between the data is calculated. Strongly correlated data is filtered out to obtain fluctuation data. This fluctuation data, combined with its collection location and correlation relationships, outputs targeted equipment adjustment and optimization decisions, enabling precise location and efficient rectification of industrial production anomalies. By analyzing abnormal data in the collected dataset, calculating highly correlated fluctuation data, and combining abnormal and fluctuation data to locate the causes of anomalies, and outputting optimization decisions based on these causes, enterprises can improve manufacturing processes and prevent production losses. This allows enterprises to achieve energy conservation and consumption reduction in the production process, while also achieving the effect of digital management.
[0027] Step S2 is followed by the following steps: Step S4: Display the relationship between abnormal data and aggregated data, and collect the data set using charts.
[0028] Specifically, based on data characteristics and display needs, the appropriate chart type can be selected. Different types of graphs or tables can be used to display data and the relationships between data, depending on actual needs. For example, heatmaps can be used to show the correlation between abnormal data and aggregate data; scatter plots with regression lines can present the linear correlation details of strongly correlated data pairs; and time-series line graphs can show the dynamic changes of the collected data set. By transforming abstract correlated data and complex data sets into intuitive and visual graphical information, the understanding threshold for non-technical personnel is lowered, helping to quickly locate core correlations and data fluctuation patterns. At the same time, it provides a visual basis for anomaly review, decision verification, and production optimization. Using charts to display data ensures that workshop operators and managers can access information simultaneously, and the marked key information directly supports production decisions and problem tracing, enhancing the practical value of data visualization.
[0029] Step S3 is followed by the following steps: Step S5: Construct a prediction model, input the real-time collected data set, predict the compressed air quality data for a future period of time, filter out abnormal data, filter out fluctuation data related to abnormal data, analyze abnormal equipment or environment based on fluctuation data, and output targeted future optimization decisions.
[0030] Specifically, an LSTM (Long Short-Term Memory) network is selected to construct a time-series prediction model. The preprocessed historical data (including normal, abnormal, and fluctuating data) accumulated in step S1 over the past three months is divided into training, validation, and test sets in a 7:2:1 ratio. Mean squared error (MSE) is used as the loss function, and Adam is used as the optimizer (initial learning rate 0.001). The model is trained for 50 epochs and fine-tuned using an early stopping method (patience=5, stopping if the validation set loss does not decrease for 5 consecutive epochs) to ensure that the model's coefficient of determination R² ≥ 0.85 and mean absolute error (MAE) < 0.1%. The collected and preprocessed data set from step S1 (concatenated into an input vector based on the most recent 6 time steps) is input into the successfully trained prediction model, which outputs compressed air quality prediction data for the next 1-24 hours (supporting selection of time granularities such as 1 hour, 4 hours, and 8 hours). The prediction data includes predicted values and 95% confidence intervals for compressed air parameters, environmental parameters, and production process parameters for each future time period. Then, the predicted data is processed in steps S2-S3 to obtain the causes of future abnormal data. Based on these causes, the current system and its parameters are adjusted to avoid the generation of predicted abnormal data. By using a time-series prediction model to capture potential production anomalies in advance and accurately pinpoint the causes of potential fluctuations, optimization decisions are upgraded from passive response to proactive prevention. Potential anomalies are warned 1-24 hours in advance, addressing the pain point of traditional solutions that intervene only after anomalies occur, significantly reducing production losses. This avoids losses such as product defects and equipment failures caused by anomalies, while also allowing sufficient execution time for the enterprise (such as spare parts procurement and equipment downtime arrangements) to ensure production continuity. Furthermore, the combination of predicted data and correlation analysis provides clear logical support for proactive decision-making, rather than blind prevention.
[0031] Step S1 specifically includes the following steps: Step S11: Place parameter detection equipment at the compressed air output port. The parameter detection equipment includes a dew point meter, an oil content analyzer, a laser particle analyzer, a flow meter, and a pressure sensor. Collect compressed air pressure dew point data, oil content data, particle size index data, flow rate data, and pressure data to obtain compressed air parameters. Step S12: Install environmental monitoring equipment in the product manufacturing workshop. The environmental monitoring equipment includes temperature and humidity sensors and particle detectors to collect temperature, humidity and particulate matter data in the production workshop environment to obtain environmental parameters. Step S13: The backend server collects raw data, including abnormal material requisition cost data, finished product output value, defective product quantity, and downtime data. The raw data collection methods include obtaining data from the ERP system, MES system, and equipment management system, or directly inputting and generating data. After calculation, material loss value, equipment downtime, and unit production energy consumption value are obtained, thus obtaining production process parameters.
[0032] Specifically, specialized parameter monitoring equipment is precisely deployed at the compressed air output outlet. This includes a dew point meter to collect pressure dew point data, an oil content analyzer to collect oil content data, a laser particle detector to collect particle size data, a flow meter to collect flow rate data, and a pressure sensor to collect pressure data. All equipment collects and uploads data in real time at a frequency of once per minute to ensure data timeliness. Environmental monitoring equipment is also evenly installed in key areas of the compressed air production workshop, such as within a 3-meter radius of equipment, at workshop ventilation openings, and in areas with dense equipment. Temperature and humidity sensors collect real-time temperature and humidity data, and environmental particle detectors collect environmental particle indicators. This equipment collects data at a frequency of once every two minutes and uploads it synchronously with the compressed air parameters using timestamps. The backend server establishes real-time communication with the enterprise's ERP system, MES system, and equipment management system through standardized data interfaces (such as API interfaces) to synchronously obtain raw data such as abnormal material requisition cost data, finished product output value, defective product quantity, and downtime. In addition, raw data can also be obtained by directly inputting data into the backend server. The core production indicators are calculated according to a preset standardized algorithm: material loss value = (abnormal material requisition cost / finished product output value) × 100%, equipment downtime = cumulative downtime recorded by the system (unit: hours), and unit production energy consumption value = total production energy consumption / total quantity of finished products (unit: kWh / piece). The calculation results are updated once every 30 minutes and stored after being aligned with the time sequence of the first two types of parameters.
[0033] Step S2 specifically includes the following steps: Step S21: Retrieve abnormal data from the collected data set, and construct linear or nonlinear correlation models for the abnormal data and other data in the set one by one to obtain the correlation between the abnormal data and the set data. The correlation between the abnormal data and the set data includes linear regression relationship and nonlinear relationship of the data. Step S22: Calculate the correlation between the outlier data and the data used to construct the correlation model to obtain the correlation data between the outlier data and the set data.
[0034] Specifically, outlier data is retrieved from the collected dataset. This retrieval can be done automatically by the system or by automatically filtering out outlier data exceeding preset thresholds for various parameters. Each type of outlier data is paired with all other data in the dataset. The model type is selected based on the data distribution characteristics and potential correlation patterns. If the data exhibits a linear trend, a linear correlation model is constructed, using univariate / multivariate linear regression algorithms and least squares to fit the regression line. If the data does not show a clear linear pattern, a nonlinear correlation model is constructed, using a mutual information calculation model to quantify the dependencies between data based on information entropy theory, ultimately clarifying the correlation type for each pair of paired data. For the constructed linear correlation model, the Pearson correlation coefficient is calculated, and the regression equation and goodness of fit R² are output. For the nonlinear correlation model, the mutual information value is calculated, and the K-nearest neighbor algorithm is used to verify the reliability of the correlation. Finally, structured correlation data is output, including the names of the data in each pair, the correlation type, the correlation strength value, the reliability index, and a strong correlation judgment criterion is set. Transform abstract relationships into quantifiable and comparable numerical data to clarify the correlation strength between outlier data and other data.
[0035] Step S21 specifically includes the following steps: Step S211: Determine whether there is a linear relationship between the abnormal data and other data in the set, and divide the data into data with linear relationships and data with non-linear relationships; Step S212: Construct a linear regression analysis model, construct a linear relationship for the data that have a linear relationship, and obtain the linear regression relationship of the data; Step S213: Construct a nonlinear model based on the mutual information model and the random forest model, construct nonlinear relationships for data with nonlinear relationships, and obtain the nonlinear relationships of the data.
[0036] Specifically, abnormal data and other data to be correlated are extracted from the collected data set. For example, when the pressure dew point is abnormal, the data to be correlated includes oil content data, ambient humidity, material loss values, etc. First, a scatter plot of abnormal data and each data to be correlated is drawn to observe whether there is a linear clustering trend. Then, the F test (significance level α=0.05) is used to verify the significance of the linear relationship. If the F test p value is <0.05 and the scatter plot shows a clear linear trend, it is determined that there is a linear relationship. Otherwise, it is marked as having a non-linear relationship, forming two types of data lists and marking the judgment criteria. In step S212, for data pairs identified as having a linear relationship, a multiple linear regression analysis model is constructed using the least squares method, with outlier data as the dependent variable (e.g., material loss value) and correlated data as independent variables (e.g., pressure dew point, compressed air pressure). The model parameters are optimized using the gradient descent method, and a linear regression equation is output. For example, material loss value = 0.03 × pressure dew point + 0.01 × compressed air pressure + 0.5. The goodness-of-fit R² is calculated (R² ≥ 0.6 indicates a qualified model). Weakly fitted data pairs with R² < 0.6 are eliminated, and the linear regression relationship of the data is finally determined. The core formula for the goodness-of-fit R² is as follows: ;in, This represents the actual value of the dependent variable. Indicates the model's predicted value; This represents the mean of the dependent variable.
[0037] In step S213, for data pairs marked as having nonlinear relationships, a fusion nonlinear model of mutual information model and random forest model is constructed. First, the mutual information value between data is calculated through the mutual information model (based on information entropy theory, quantifying the dependency relationship between variables). Potentially strongly correlated data with mutual information values > 0.2 are initially screened. Then, with outlier data as target variables and correlated data as feature variables, a random forest model is constructed. The number of decision trees is set to 100, and the maximum depth is 8 to prevent overfitting. The model performance is verified through out-of-bag data, and feature importance scores are output. Data pairs with mutual information values > 0.2 and feature importance scores in the top 50% are identified as having nonlinear relationships, thus clarifying the nonlinear correlation pattern of the data. Specifically, firstly, variable roles are defined, with outlier data as the target variable and corresponding related data as feature variables. Missing values are filled using K-nearest neighbors, and the data is normalized using Min-Max to unify the units of measurement, with key fluctuation points marked (outliers are not blindly removed). The data is then split into training and testing sets in a 7:3 ratio. Next, using a mutual information model, the kernel density estimation method based on information entropy theory is employed to calculate the mutual information values between the target variable and each feature variable, screening out potentially strongly correlated features with MI(X,Y) > 0.2. Subsequently, a random forest model with 100 decision trees, a maximum depth of 8, and a minimum number of sample splits of 10 is constructed, and the model is tested using Bootstrap. p-sampling generates a training subset to train each decision tree, with minimizing the Gini coefficient as the node splitting objective. The optimized model is validated using out-of-bag (OOB) data (ensuring an OOB error rate ≤ 15%). Then, the prediction accuracy is verified by calculating MAE (≤ 10% of the target variable mean) and RMSE on the test set. Feature importance scores are output (normalized to 0-100), and partial dependency graphs and feature interaction graphs are drawn to visualize nonlinear trends. Finally, a structured output result is generated, including a list of core associated features (including mutual information values and feature importance scores), nonlinear association patterns interpreted in industrial terms, effective value ranges for features, and association strength levels.
[0038] Nonlinear models include multivariate nonlinear relationship models, nonlinear causal relationship models, nonlinear regression relationship models, logarithmic relationship models, exponential relationship models, quadratic relationship models, cubic relationship models, etc. After determining that the current model is nonlinear, the types of nonlinear relationships between abnormal data and fluctuating data can be determined again based on the relationships between the data, so as to output corresponding optimization decisions based on the data relationships in the future.
[0039] Step S22 specifically includes the following steps: Step S221: The data in the set is associated with the abnormal data one by one. The set is divided into abnormal data set and associated data set. The associated data set is associated with the abnormal data one by one to construct an association model. Step S222: Based on the correlation between the obtained abnormal data and the set data, calculate the correlation between all data in the associated data set and the abnormal data to obtain the correlation data.
[0040] Specifically, the collected data set is first split into abnormal data and associated data sets. Then, the data in the associated data set is paired with the abnormal data one by one, and a linear regression model or a nonlinear fusion model is automatically invoked. For linear data pairs, the Pearson correlation coefficient is calculated (the closer |r| is to 1, the stronger the linear correlation). For nonlinear data pairs, the feature importance score of the random forest model (normalized to 0-100 points) is used as the core indicator of correlation, and mutual information value is used to assist in verification. Finally, structured associated data is output, including associated data name, correlation type (linear / nonlinear), correlation strength value (correlation coefficient / feature score), verification index (R² / mutual information value), and strong correlation threshold is marked (linear |r|≥0.6, nonlinear feature score≥60 points).
[0041] Step S3 specifically includes the following steps: Step S31: Sort the correlation data according to the numerical value, thereby sorting the aggregated data by correlation, setting a data threshold, and capturing the correlation data that is higher than the data threshold to obtain the fluctuation data; Step S32: Locate the device causing the anomaly based on the location of the fluctuation data collection. Combine the device causing the anomaly with the correlation between the anomaly data and the aggregate data to plan system adjustment actions and obtain optimization decisions. The enterprise adjusts the system and system parameters according to the optimization decisions.
[0042] Specifically, the output structured correlation data is sorted in descending order of correlation strength, linear data is sorted by the absolute value of the Pearson correlation coefficient, and nonlinear data is sorted by the random forest feature importance score. Combining industry experience and the company's historical rectification effects, dynamic data thresholds are set: the linear correlation threshold is set to a Pearson correlation coefficient |r| ≥ 0.6, and the nonlinear correlation threshold is set to a feature importance score ≥ 60. The system automatically captures correlation data above the corresponding thresholds and extracts the underlying data of these correlations as fluctuation data. From massive amounts of correlation data, the system accurately filters out the core fluctuation data that has the most significant impact on abnormal data, eliminates weakly correlated interference data, and focuses on key clues for locating the causes of anomalies. The descending sorting method intuitively presents the priority of impact, the dynamic thresholds adapt to different parameter characteristics, and the capture logic is standardized and traceable, ensuring the core nature and accuracy of the fluctuation data and improving the efficiency of subsequent cause location. Based on the location, equipment, or environment from which fluctuation data is collected, the system identifies the equipment or environmental factors causing the anomaly. Combining this with the types of correlations, the system deepens the decision-making logic. If the data exhibits a linear relationship, corresponding parameter adjustments are implemented, such as raising the equipment regeneration temperature to 145℃ and lowering the pressure dew point. If the data exhibits a non-linear relationship, corresponding equipment status optimization and parameter coordination decisions are implemented, such as activating the workshop's standby dehumidifier and simultaneously lowering the air compressor's cooling water temperature by 2℃. When outputting decision content, the system must clearly specify the equipment name, operating steps, execution timeframe, and verification standards, and push this information to the enterprise's equipment management terminal in a graphical and textual format. The data correlation results are translated into actionable system adjustment solutions, directly guiding enterprises to resolve anomalies. The deep integration of decision logic and correlation types avoids generic suggestions, ensuring strong targeting and complete decision elements. Enterprises can directly execute the steps, highlighting practicality and operability, and effectively shortening the anomaly rectification cycle. Among them, the fluctuating data can be one data point or several data points. That is, the abnormal cause of the abnormal data can be one or more. The number of causes is determined according to the correlation model, and the corresponding optimization decision is output in combination with the causes. The optimization decision is not limited to processing one abnormal cause, but can process multiple abnormal causes.
[0043] Secondly, this invention discloses an industrial high-quality management system with compressed air quality data as its core, which includes: The data acquisition unit is used to acquire a set of collected data, including compressed air parameters, environmental parameters, and production process parameters. The data processing unit is used to perform secondary operations on the collected data set; The data analysis and diagnosis unit extracts abnormal data from the data collection set acquired by the data acquisition unit, and performs correlation analysis on the abnormal data and other data in the set, including constructing a correlation model and calculating correlation data, to obtain the correlation between fluctuation data, abnormal data and set data, as well as the correlation data between abnormal data and set data, and locates the device that causes the abnormality based on the fluctuation data; The decision output unit combines the abnormal cause devices with the correlation between abnormal data and aggregate data to plan system adjustment actions and obtain optimized decisions.
[0044] Specifically, the data acquisition unit acquires the collected data set and transmits it to the data analysis and diagnosis unit for analysis and diagnosis. It captures the fluctuation data associated with the abnormal data, locates the device or environment that caused the abnormality based on the abnormal data and fluctuation data, and outputs targeted optimization decisions for the cause of the abnormality.
[0045] The data acquisition unit includes a dew point meter, an oil content analyzer, a laser particle detector, a flow meter, a pressure sensor, a temperature and humidity sensor, a particle detector, and a system data acquisition unit. It can also directly connect to smart meters and water meters to calculate real-time energy consumption. The data acquisition unit collects a dataset including compressed air parameters, environmental parameters, and production process parameters. The data processing unit performs secondary calculations on the dataset, including calculating unit production energy consumption, gas-to-electricity ratio, and carbon emissions. The parameters obtained from these secondary calculations are categorized under the production process parameters in the dataset for subsequent analysis and diagnosis of abnormal data. Furthermore, the data processing unit preprocesses the dataset, including format conversion. The data analysis and diagnostic unit can extract fluctuation data related to abnormal data from the collected data. It employs a correlation model construction and data correlation calculation to acquire fluctuation data, locate abnormal devices and environments, and then reverse-match the data to the acquisition equipment and deployment location to pinpoint the source of the anomaly. This allows for deeper decision-making based on the type of correlation. For linear correlations, parameters are adjusted in reverse; for nonlinear correlations, device status and parameter coordination solutions are implemented, adjusting corresponding device parameters. In multi-cause scenarios, priority ranking decisions are output, clearly identifying and addressing core devices first.
[0046] The data analysis and diagnostic unit includes: The association model construction sub-unit constructs linear and non-linear association models with abnormal data and other data in the set to obtain the association relationship between abnormal data and other data. The correlation calculation subunit calculates the correlation data between abnormal data and other data. It judges the strength of the correlation between abnormal data and other data through the correlation data, presets a data threshold, and captures the correlation data that exceeds the data threshold to obtain the fluctuation data. The trigger device positioning subunit is a device that collects positioning data through fluctuation data.
[0047] The correlation model construction sub-unit first extracts abnormal data from the set output by the data acquisition unit based on preset thresholds (such as pressure dew point > -18℃, material loss value > 2.5%). Then, it uses scatter plot visualization and F test (α=0.05) to determine the relationship type between abnormal data and other data. If the F test p value < 0.05 and the scatter plot shows a linear trend, a linear correlation relationship is constructed by a linear regression model (least square fitting), and the regression equation and goodness of fit R² are output. If it is determined to be a non-linear relationship, a mutual information model and a random forest model are used for modeling. First, potential correlation data with MI > 0.2 are screened by mutual information, and then the non-linear law is learned by a random forest model with 100 decision trees, and the feature importance score is output. The correlation calculation subunit calculates the Pearson correlation coefficient (|r| ranges from 0 to 1) for linear correlation models and extracts the random forest feature importance score (normalized to 0-100) for nonlinear models, forming standardized correlation data. Dynamic thresholds are set based on the company's historical rectification effects (linear |r| ≥ 0.6, nonlinear score ≥ 60). The system automatically sorts the data in descending order of correlation strength and captures the basic data corresponding to data exceeding the threshold, which are the core fluctuation data causing the anomalies, such as pressure dew point (r=0.72) and ambient humidity (score 68). By back-matching the fluctuation data to the collection equipment and deployment location, the source of the anomaly can be located. If the fluctuation data is pressure dew point, the cause points to the dryer or workshop dehumidification system; if it is ambient humidity, the cause points to workshop ventilation or dehumidification equipment; if there are multiple fluctuations, the cause of multiple equipment synergies is identified.
[0048] The system also includes a data prediction unit, which constructs a data prediction model and inputs real-time collected data sets to predict compressed air quality-related data for a future period of time.
[0049] Specifically, an LSTM (Long Short-Term Memory) network was selected to construct a time-series prediction model. The input layer dimension consisted of 11 parameter variables × 6 time steps. Among them, the 11 core parameters were 5 types of compressed air, 3 types of environment, and 3 types of production process. A 64-neuron LSTM network and a 0.2 Dropout layer were used to prevent overfitting. The model was trained using three months of preprocessed historical data, with the training, validation, and test sets divided in a 7:2:1 ratio. The pass / fail criteria were MAE < 0.1% and R² ≥ 0.85. The model takes the latest 6-hour data collected in real-time as input and outputs the predicted compressed air quality data for the next 1-24 hours, including predicted values for each parameter and 95% confidence intervals. It automatically compares the data against preset anomaly thresholds, filters out potentially abnormal data and corresponding potential fluctuations, identifies potential triggering equipment based on the potential fluctuation data, and pushes preventative recommendations.
[0050] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0051] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0052] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0053] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0054] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.
[0055] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. The illustrative expressions of the above terms in this specification should not be construed as necessarily referring to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0056] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Since these modifications and variations fall within the scope of the claims and their equivalents, this invention also intends to include these modifications and variations.
[0057] The above description describes specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An industrial high-quality management method based on compressed air quality data, characterized in that, include, Collect data related to the air compressor system, production environment data, and data generated by the production system, and preprocess the data to obtain a collection of data, which includes compressed air parameters, environmental parameters, and production process parameters. Retrieve abnormal data from the collected data set, construct a correlation model between the abnormal data and other data in the set one by one, and calculate the correlation between the abnormal data and other data in the set to obtain the correlation relationship between the abnormal data and the set data, and the correlation data between the abnormal data and the set data. Based on the correlation between abnormal data and aggregate data, we determine the data in the aggregate data that are related to the abnormal data to obtain fluctuation data. Based on the location of fluctuation data collection and the correlation between abnormal data and aggregate data, we output optimization decisions.
2. The method according to claim 1, characterized in that, The process involves collecting data related to the compressed air system, production environment data, and data generated by the production system, and preprocessing the data to obtain a collected data set. This data set includes compressed air parameters, environmental parameters, and production process parameters. Specifically, the process includes the following steps: A parameter detection device is placed at the compressed air output port. The parameter detection device includes a dew point meter, an oil content analyzer, a laser particle analyzer, a flow meter, and a pressure sensor. It collects compressed air pressure dew point data, oil content data, particle size index data, flow rate data, and pressure data to obtain compressed air parameters. Environmental monitoring equipment, including temperature and humidity sensors and particle detectors, is installed in the product manufacturing workshop to collect temperature, humidity and particulate information in the production workshop environment and obtain environmental parameters. The back-end server collects raw data, including abnormal material requisition cost data, finished product output value, defective product quantity, and downtime data. The raw data collection methods include obtaining data from ERP systems, MES systems, and equipment management systems, or directly inputting and generating data. The raw data is then used to calculate material loss values, equipment downtime, and unit production energy consumption values, thus obtaining production process parameters.
3. The method according to claim 1, characterized in that, The process of retrieving abnormal data from the collected data set, constructing a correlation model between the abnormal data and other data in the set one by one, and calculating the correlation between the abnormal data and other data in the set to obtain the correlation relationship between the abnormal data and the set data, specifically includes the following steps: Extract abnormal data from the collected data set, and construct linear or nonlinear correlation models between the abnormal data and other data in the set to obtain the correlation relationship between the abnormal data and the set data. The correlation between the outlier data and the data constructed in the correlation model is calculated to obtain the correlation data between the outlier data and the set data.
4. The method according to claim 3, characterized in that, The process of retrieving abnormal data from the collected data set and constructing linear or nonlinear correlation models between the abnormal data and other data in the set to obtain the correlation between the abnormal data and the set data specifically includes the following steps: Determine whether there is a linear relationship between outlier data and other data in the set, and divide the data into data with linear relationships and data with non-linear relationships; Construct a linear regression analysis model to establish a linear relationship among data that have a linear relationship, and obtain the linear regression relationship of the data. A nonlinear model is constructed based on the mutual information model and the random forest model. Nonlinear relationships are then constructed from data that have nonlinear relationships, thus obtaining the nonlinear relationships of the data.
5. The method according to claim 1, characterized in that, The process involves determining the data in the set that are associated with the abnormal data based on the correlation between the abnormal data and the set data, obtaining fluctuation data, and then outputting optimization decisions based on the location of the fluctuation data collection and the correlation between the abnormal data and the set data. This process specifically includes the following steps: Sort the data according to their correlation values, and then sort the aggregated data according to their correlation. Set a data threshold and capture the correlation data that is higher than the data threshold to obtain the fluctuation data. Based on the location of fluctuation data collection, the abnormal triggering device is located. By combining the abnormal triggering device with the correlation between abnormal data and aggregate data, the system adjustment actions are planned to obtain optimization decisions. The enterprise then adjusts the system and system parameters according to the optimization decisions.
6. The method according to claim 1, characterized in that, The following steps are then included: The relationship between abnormal data and the collected data set is displayed through charts.
7. The method according to claim 1, characterized in that, The following steps are then included: Construct a predictive model, input the real-time collected data set, predict the compressed air quality data for a future period, filter out abnormal data, filter out fluctuation data related to abnormal data, analyze abnormal equipment or environment based on fluctuation data, and output targeted future optimization decisions.
8. An industrial high-quality management system based on compressed air quality data, comprising the industrial high-quality management method based on compressed air quality data as described in any one of claims 1-7, characterized in that, include; The data acquisition unit is used to acquire a set of collected data, including compressed air parameters, environmental parameters, and production process parameters. The data processing unit is used to perform secondary operations on the collected data set; The data analysis and diagnosis unit extracts abnormal data from the data collection set acquired by the data acquisition unit, and performs correlation analysis on the abnormal data and other data in the set, including constructing a correlation model and calculating correlation data, to obtain the correlation between fluctuation data, abnormal data and set data, as well as the correlation data between abnormal data and set data, and locates the device that causes the abnormality based on the fluctuation data; The decision output unit combines the abnormal cause devices with the correlation between abnormal data and aggregate data to plan system adjustment actions and obtain optimized decisions.
9. The system according to claim 8, characterized in that, The data analysis and diagnostic unit includes: The association model construction sub-unit constructs linear and non-linear association models with abnormal data and other data in the set to obtain the association relationship between abnormal data and other data. The correlation calculation subunit calculates the correlation data between abnormal data and other data. It judges the strength of the correlation between abnormal data and other data through the correlation data, presets a data threshold, and captures the correlation data that exceeds the data threshold to obtain the fluctuation data. The trigger device positioning subunit is a device that collects positioning data through fluctuation data.
10. The system according to claim 8, characterized in that, Also includes: The data prediction unit constructs a data prediction model and uses real-time collected data sets to predict compressed air quality-related data for a future period of time.