A Deep Concentration Production Scheduling Big Data Analysis System and Method

Through the deep concentration production scheduling big data analysis system, production data is collected in real time, time series analysis and index optimization are carried out, and key influencing factors are identified, which solves the problems of low data query efficiency and in-depth mining of production parameters quality relationships in the existing technology, and improves the timeliness and accuracy of production decisions.

CN120011373BActive Publication Date: 2025-07-08HUADIAN QINGDAO POWER GENERATION COMPANY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510502918.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-08
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The single data retrieval mechanism in the prior art leads to a decrease in query speed when the data volume is huge or the query conditions are complex, making it difficult to meet the requirements of fast and efficient decision-making, and the relationship between production parameters and product quality is not thorough enough, and there is a lack of effective identification of key influencing factors.

Method used

The deep-concentrated production scheduling big data analysis system is adopted to collect real-time production data through the data acquisition module, generate key performance snapshots, and perform time series analysis and quality fluctuation monitoring. Combining hash and B-tree indexing strategies to optimize query efficiency, and using decision trees and random forest models to analyze the impact of production parameters on product quality.

Benefits of technology

It realizes fine-grained capture and analysis of production information, improves trend prediction accuracy and early warning capabilities for quality abnormalities, improves production efficiency and product quality stability, and shortens the response time of multi-dimensional data query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011373B_ABST
    Figure CN120011373B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of big data analysis, and specifically to a deep concentration production scheduling big data analysis system and method. The system includes: a data collection module that collects real-time data of concentration production, including temperature, speed, and pressure information, and generates a real-time production data set. In the present invention, by collecting multi-dimensional data such as temperature, speed, and pressure in the production process in real time, key performance indicators are accurately extracted, fine-grained capture and analysis of production information are realized, a production performance snapshot is formed, and on this basis, in-depth time series analysis is carried out to identify and predict production trends and seasonal fluctuations, improving the accuracy and reliability of trend prediction; at the same time, combining the trend and seasonal analysis results, fine-grained quality fluctuation monitoring of the production process is carried out, the fluctuation range of quality anomalies is timely discovered and located, and early warning and efficient control of production quality risks are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data analysis, and particularly relates to a deep concentration production scheduling big data analysis system and method. Background Art

[0002] Big data analysis refers to the process of using advanced data storage, processing, and analysis technologies to extract valuable information, patterns, and knowledge from large and diverse data sets. Its technical core includes data collection, storage, preprocessing, mining, analysis, visualization, and real-time computing, etc., involving distributed computing frameworks, data mining algorithms, machine learning methods, and cloud computing platforms.

[0003] In the prior art, the data retrieval mechanism is relatively single, usually relying on linear or simple indexing methods, resulting in a significant reduction in query speed when the data volume is large or the query conditions are complex, making it difficult to meet the requirements of fast and efficient decision-making. Especially in emergency situations, the delay affects the timeliness and accuracy of decision-making. In addition, the relationship between production parameters and product quality is not deeply mined, often only staying at the surface correlation description, lacking effective identification of key influencing factors. Therefore, improvements are needed. Summary of the Invention

[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art and propose a deep concentration production scheduling big data analysis system and method.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A deep concentration production scheduling big data analysis system includes:

[0006] A data collection module, which collects real-time data of concentration production, including temperature, speed, and pressure information, generates a real-time production data set; extracts key performance indicators according to the real-time production data set, and generates a key performance snapshot;

[0007] A data processing module, based on the key performance snapshot, performs time series analysis to identify trends and seasonal fluctuations in production data, and generates a trend and seasonal analysis result; monitors the quality fluctuations of the key performance snapshot, and combines with the trend and seasonal analysis result to generate a quality monitoring result;

[0008] A data indexing module, based on the quality monitoring result, designs and implements hash and B-tree indexing strategies to optimize data query efficiency, and generates an optimized query path; performs composite index construction according to the optimized query path for multi-condition query optimization, and generates an index optimization result;

[0009] A decision support module, based on the index optimization result, uses decision tree and random forest models to analyze the impact of production parameters on product quality, and generates an influencing factor analysis result.

[0010] Preferably, the steps for obtaining the real-time production dataset are as follows:

[0011] Monitor the real-time data of temperature, speed, and pressure on the concentration production line, mark and format the time stamp for each collected data point to form the original real-time monitoring data;

[0012] Based on the original real-time monitoring data, remove error data and outliers, fill in the missing data points, and use statistical methods to evaluate the data consistency to generate the real-time production data after quality inspection;

[0013] Based on the real-time production data after quality inspection, perform data synchronization processing, align the data from different sensor sources according to the unified time stamp to form the real-time production dataset.

[0014] Preferably, the steps for obtaining the key performance snapshot are as follows:

[0015] Select the key performance indicators from the real-time production dataset, including the highest temperature, the lowest speed, and the average pressure, to generate the initially selected performance indicator data;

[0016] Based on the initially selected performance indicator data, construct the key performance snapshot, and the performance snapshot is presented in the form of views and charts by integrating the highest temperature, the lowest speed, and the average pressure indicators.

[0017] Preferably, the steps for obtaining the trend and seasonal analysis results are as follows:

[0018] Based on the key performance snapshot, arrange it in chronological order to form a time series dataset, analyze the change range and differences of each indicator in different time periods to obtain the preliminary trend and seasonal fluctuation data;

[0019] Calculate the trend-seasonal composite fluctuation factor according to the preliminary trend and seasonal fluctuation data;

[0020] Based on the trend-seasonal composite fluctuation factor, combined with the time series dataset, determine the trend and seasonal fluctuation patterns in the production data, and judge the change rate of the production data in different time cycles to obtain the trend and seasonal analysis results.

[0021] Preferably, the steps for obtaining the quality monitoring results are as follows:

[0022] Extract each performance indicator sequence in the key performance snapshot and calculate the control limits;

[0023] Based on the control limits, monitor the quality fluctuations in the real-time production data, mark the data points exceeding the control limits, and combine the trend and seasonal analysis results to generate the quality monitoring results.

[0024] Preferably, the steps for obtaining the optimized query path are as follows:

[0025] Based on the quality monitoring results, analyze the access patterns of production data queries, statistically analyze the distribution characteristics of query requests, screen the data fields applicable to hash indexes or B-tree indexes, determine the index construction plan, and generate an index strategy design plan;

[0026] According to the index strategy design plan, construct a hash index for processing equality queries, determine the hash function and calculate the key-value mapping relationship, and at the same time establish a B-tree index for range queries, perform index node splitting and adjustment operations, and generate an optimized query path.

[0027] Preferably, the steps for obtaining the index optimization result are as follows:

[0028] Based on the optimized query path, analyze the query execution plan, identify multiple fields included in the query pattern, extract the field combination with the highest query frequency, screen the field set applicable to the composite index, and generate a composite index construction plan;

[0029] Calculate the index adaptation degree score according to the composite index construction plan;

[0030] Based on the index optimization result, analyze the impacts of temperature, speed, and pressure on product quality, generate an impact factor analysis, and obtain an impact factor analysis result.

[0031] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0032] In the present invention, by collecting multi-dimensional data such as temperature, speed, and pressure in the production process in real time, accurately extracting key performance indicators, realizing fine-grained capture and analysis of production information, forming a production performance snapshot, and on this basis, conducting in-depth time series analysis to identify and predict production trends and seasonal fluctuations, improving the accuracy and reliability of trend prediction; at the same time, combining the trend and seasonal analysis results, conducting refined quality fluctuation monitoring on the production process, timely discovering and positioning the quality abnormal fluctuation range, and realizing early warning and efficient control of production quality risks. In addition, by using hash and B-tree index strategies to optimize the query path and implementing the construction of composite indexes for multiple conditions, the multi-dimensional data query response time is shortened, and the processing efficiency of complex queries is greatly improved; by analyzing the correlation degree between each production parameter and product quality through decision tree and random forest models, key influencing factors are effectively identified, helping production managers more accurately grasp the direction of process parameter adjustment, and further realizing the lean production goal, significantly improving production efficiency, reducing production costs, and enhancing the stability of product quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is the system flow chart of the present invention. Specific embodiments

[0034] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0035] Please refer to Figure 1 , the present invention provides a technical solution: a big data analysis system for deep concentration production scheduling includes:

[0036] A data collection module that collects real-time data of concentration production, including temperature, speed and pressure information, generates a real-time production data set; extracts key performance indicators based on the real-time production data set to generate a key performance snapshot;

[0037] A data processing module that, based on the key performance snapshot, performs time series analysis to identify trends and seasonal fluctuations in production data and generates trend and seasonal analysis results; monitors quality fluctuations of the key performance snapshot and combines the trend and seasonal analysis results to generate quality monitoring results;

[0038] A data indexing module that, based on the quality monitoring results, designs and implements hash and B-tree indexing strategies to optimize data query efficiency and generates an optimized query path; performs composite index construction according to the optimized query path for multi-condition query optimization and generates index optimization results;

[0039] A decision support module that, based on the index optimization results, uses decision tree and random forest models to analyze the impact of production parameters on product quality and generates factor analysis results.

[0040] The steps for obtaining the real-time production data set are as follows:

[0041] Monitor the real-time data of temperature, speed and pressure on the concentration production line, mark and format the time stamp for each collected data point to form the original real-time monitoring data;

[0042] Based on the original real-time monitoring data, remove error data and outliers, fill in missing data points, and use statistical methods to evaluate the consistency of the data to generate real-time production data after quality inspection;

[0043] Based on the real-time production data after quality inspection, perform data synchronization processing to align data from different sensor sources according to a unified time stamp to form a real-time production data set.

[0044] Specifically, monitor the real-time data of temperature, speed, and pressure on the concentration production line. Select multiple sensors installed at different positions on the production line and calibrate them in advance. The measurement range of each sensor is set based on historical test data. For example, set the temperature range to 0°C to 200°C and calculate the average value and standard deviation according to previous test samples to determine the upper limit threshold of temperature anomaly Ttemp = average temperature + 2 × standard deviation. If the measured temperature is higher than Ttemp, it is determined that it will not be included in the preliminary record for the time being. Similarly, set the speed range to 0 m / s to 5 m / s and the pressure range to 0 MPa to 2 MPa. After the preliminary calibration of each sensor at the acquisition end according to the above ranges, read the real-time data that has passed the calibration one by one and configure a unique timestamp. To demonstrate the threshold setting process, assume that the temperature sample has 50 observations, the average temperature is 70°C, and the standard deviation is 5°C. Then Ttemp = 70 + 2 × 5 = 80°C. If the observed value exceeds 80°C, it is temporarily regarded as an out-of-limit record at this stage. Each record consists of three measurement values of temperature, speed, and pressure and the corresponding time identifier, and the measurement units of temperature, speed, and pressure are written into the same data carrier in a unified format. Finally, after arranging all the records, the original real-time monitoring data is obtained.

[0045] Based on the above-obtained original real-time monitoring data, take the complete records of temperature, speed, and pressure respectively to detect error data and outliers. The judgment of outliers can adopt the method of average value plus or minus a fixed multiple of the standard deviation. For example, first perform statistical calculations on all temperature records, with an average value of 70°C and a standard deviation of 5°C, and set the threshold Ttemp = 70 + 2 × 5 = 80°C or below 60°C as the boundary. If the temperature is higher than 80°C or lower than 60°C, it is determined as an abnormal record. Obtain the speed anomaly boundary Vlimit and the pressure anomaly boundary Plimit for speed and pressure in the same way, and delete the detected abnormal records. Then, perform interpolation processing on the missing records. For example, use linear interpolation to estimate a temperature value for the vacancy between two adjacent normal temperature data. Fill the missing points for speed and pressure in the same way. Then, use the variance comparison method to evaluate data consistency, that is, perform variance analysis on temperature, speed, and pressure within the same time period and compare their discrete distributions in different time periods to view the fluctuation situation. If the dispersion degree of a certain time period exceeds the threshold obtained according to historical calculations, for example, when the variance is greater than 4, it is regarded as inconsistent data and can be deleted or reviewed again. Finally, the retained records can be summarized to obtain the real-time production data after quality inspection.

[0046] Based on the real-time production data after the above quality inspection, select the qualified records of temperature, speed, and pressure for data synchronization processing. Accurately round the timestamps corresponding to each record to the second level. Compare whether all the records of different sensors exist at the same time point. If there is data only from some sensors at a certain time point, the measurement results of the nearest time point can be used or interpolation can be performed based on the data at the current and adjacent moments to make up for it. For example, when the speed sensor lacks data at t = 10s while the temperature and pressure sensors have records at t = 10s, the speed value at t = 10s can be calculated by linear interpolation using the speed values at t = 9s and t = 11s. The threshold setting can be determined with reference to the statistical results to re-examine if the speed difference between adjacent moments is too large. For example, set the inspection standard of step size Δv = 0.5m / s. If the interpolated speed differs from the adjacent actual record by more than 0.5m / s, it is marked as suspected mismatched data and retained for subsequent review. All synchronized data is re-sorted in ascending order of timestamps and the three values of temperature, speed, and pressure are listed uniformly. Finally, the aligned data is merged into records under the same time reference to form a real-time production dataset.

[0047] The steps for obtaining the key performance snapshot are as follows:

[0048] From the real-time production dataset, screen the key performance indicators, including the highest temperature, the lowest speed, and the average pressure, to generate the performance indicator data of the preliminary screening;

[0049] Based on the performance indicator data of the preliminary screening, construct the key performance snapshot. The performance snapshot is presented in the form of views and charts by integrating the indicators of the highest temperature, the lowest speed, and the average pressure.

[0050] Specifically, in the real-time production dataset, the parameters such as temperature, speed, and pressure at several moments have been obtained. Based on the pre-established temperature range of 0°C to 200°C, speed range of 0m / s to 5m / s, and pressure range of 0MPa to 2MPa, all data points are compared item by item and the temperature, speed, and pressure at each moment are recorded. For the temperature data, the maximum value function in statistics can be used to compare all the temperature records within the specified range and store the function result as the highest temperature value. For the speed data, the minimum value function can be used to obtain the lowest speed value from the dataset within the range. For the pressure data, when calculating statistics, according to its total record length calculate the average value , where is the One qualified pressure record. Through the above calculations, comparisons, and statistics, representative values of three key indicators can be selected to form a set of performance indicators. If the temperature, speed, or pressure data at certain moments are not within the previously set ranges, they are regarded as invalid records and not included in the statistical scope. When calculating examples, if the maximum temperature is statistically 180°C, the minimum speed is 0.8 m / s, and the average pressure is 1.2 MPa, then 180°C, 0.8 m / s, and 1.2 MPa are written into the performance indicator set respectively. When wanting to set inspection thresholds again, historical records can be analyzed. For example, the temperature can be compared with the sum of the average value and twice the standard deviation when the sample size is 50 to obtain a temperature determination value. The speed and pressure are processed in the same way. Finally, all data passing the numerical verification are statistically analyzed for the highest temperature, lowest speed, and average pressure and registered in the same data table. The statistical results are finally uniformly recorded as the performance indicator data of the preliminary screening.

[0051] Based on the performance indicator data of the preliminary screening obtained in the previous step, the three values of the highest temperature, lowest speed, and average pressure are visually integrated as key contents. By respectively plotting various charts such as temperature curves, speed curves, and pressure curves under the same time axis for comparison and presentation, data points within a fixed time period can also be selected and highlighted using line charts or histograms. To expand the information volume, the historical highest temperature can be compared with the current highest temperature and the difference marked. The specific change trajectory of the lowest speed can also be compared on the same interface and interactively presented with the trend difference from the average pressure. To facilitate the identification of situations exceeding specific thresholds, a color marking method can be added to the charts. For example, if the temperature exceeds 200°C, it is marked with a red line segment; if the speed is lower than 0.5 m / s, it is marked with a yellow line segment; if the pressure exceeds 1.5 MPa, it is marked with an orange line segment. These thresholds can be determined according to the equipment manufacturer's suggestions or operating specifications and can be calculated through actual examples. For example, in the past week, 100 temperature records were collected and the average value was calculated as 80°C and the standard deviation was 6°C, then the threshold can be set as 80 + 2×6 = 92°C. If some points exceed 92°C, they are marked as records that need attention. The speed and pressure are also calculated for their average values and standard deviations in the same way to obtain thresholds and compare with the actual observed values. After all the charts and markings are summarized, a key performance snapshot containing the highest temperature, lowest speed, and average pressure is formed. The performance snapshot is presented in the form of views and charts by integrating the highest temperature, lowest speed, and average pressure indicators.

[0052] The steps to obtain the trend and seasonal analysis results are as follows:

[0053] Based on the key performance snapshot, a time series data set is formed by arranging in chronological order, and the change amplitudes and differences of each indicator in different time periods are analyzed to obtain preliminary trend and seasonal fluctuation data;

[0054] Based on the preliminary trend and seasonal fluctuation data, calculate the trend-seasonal composite fluctuation factor, and the calculation formula is:

[0055] ;

[0056] where represents the trend-seasonal composite fluctuation factor, represents the value of the trend component at the th time point, represents the value of the seasonal component at the th time point, is the total number of time points, represents the pressure change amplitude at the th time point, represents the speed change amplitude at the th time point, represents the temperature change amplitude at the th time point;

[0057] Based on the trend-seasonal composite fluctuation factor, combined with the time series data set, determine the trend and seasonal fluctuation patterns in the production data, judge the change rate of the production data in different time periods, and obtain the trend and seasonal analysis results.

[0058] Specifically, after the key performance snapshots have been obtained, first, the highest temperature, the lowest speed, and the average pressure data contained therein are arranged period by period and the corresponding time order is marked. Then, a specified time period is selected from the previously recorded temperature, speed, and pressure data as the comparison benchmark. For example, the temperature, speed, and pressure per hour within the first 72 hours are listed as an observation group. The difference magnitude between adjacent observation groups is compared and the difference range is recorded. If the temperature of a certain observation group increases by more than 10°C compared to the previous group, then this time period is marked as a high-fluctuation interval. This value of 10°C can be obtained based on the statistical conclusions of the previous equipment operation. For example, during the continuous operation of the equipment for 300 hours, the temperature fluctuation amplitude each time is concentrated between 0°C and 10°C. If the speed decreases by more than 0.3 m / s compared to the previous group, it can be regarded as a significant change in speed. The threshold value of 0.3 m / s can also be obtained by adding twice the standard deviation to the average speed difference. In terms of pressure, if the difference compared to the previous group exceeds 0.2 MPa, it indicates obvious pressure fluctuation. This upper limit of 0.2 MPa is set in the form of adding the standard deviation to the mean value of the difference distributions recorded by multiple pressure acquisitions. The time periods marked as high or low fluctuations are summarized and statistically analyzed, and the change amplitudes and differences of the above indicators within different time periods are calculated. Grouped by hour or day, the gap value between each group and the reference group is calculated to observe the changes in the morning and evening periods of the same day and the change amplitudes between different days. After comparing all the observation group time periods with the reference group time periods, a list of the amplitude changes of each indicator within different time periods can be obtained. By arranging the time series of the amplitude change list, the difference sequence can be further extracted and divided into ascending or descending types according to the fluctuation direction. Finally, summarizing such sequence data can obtain the preliminary trend and seasonal fluctuation data with phased characteristics.

[0059] The advantage of the formula lies in combining the change amounts in multiple aspects such as temperature, speed, and pressure to measure the comprehensive fluctuation characteristics, and by simultaneously considering the difference parts of the trend component and the seasonal component, optimizing the overall assessment of production fluctuations;

[0060] The steps for obtaining the parameters are as follows: In the time series dataset, the temperature values recorded at each time point are fitted with a long-term change curve in a predetermined manner to obtain the temperature trend component corresponding to this time point. Since the temperature changes within the range of 0°C to 200°C, it is necessary to read the temperature within 400 consecutive hours and calculate the average value per hour, and record each hourly average value as Then, use the polynomial fitting method to form the overall temperature trend function Taking and the minimum sum of squared differences as the convergence target, and finally, for each time point, is output as , for example, if 400 temperature means are collected within 400 hours, a cubic polynomial trend function is obtained after least-squares fitting. , according to the obtained coefficients , , , , then at the 100th hour, it can be calculated , and after calculation, can be obtained, and this value can represent the temperature trend component at the 100th hour;

[0061] The steps for obtaining the parameter are as follows: for the temperature records within the same time series, select the repeated fluctuation patterns within a specific period as the seasonal component, divide every 24 hours into a period segment, merge the corresponding temperatures within the same time period, and construct a daily periodic pattern. , perform a minimum difference fitting between the periodic pattern and the original temperature record. When the difference is close to the peak interval under the statistical distribution curve, record this curve as the final seasonal component, and thus the at the time point can be obtained. For example, for the temperature analysis of 28 days, it is statistically found that the temperature will show obvious stage increases and decreases at the 5th hour and the 18th hour of each day. Then, the corresponding peak and valley characteristics can be retained in the daily periodic function. Assuming that the 100th hour falls near the 4th hour of the day ((100 mod 24) = 4), then is the function value of the daily periodic function at the 4th hour, and may be obtained after determination by the aforementioned minimum difference principle;

[0062] The steps for obtaining the parameter are as follows: when considering the number of time points, directly count the number of time series participating in the evaluation. If one sampling point is taken per hour within 400 consecutive hours, can be obtained;

[0063] The steps for obtaining the parameter are as follows: calculate the absolute value of the difference in pressure relative to the previous time point for each time point and record it as the pressure change amplitude. By continuously monitoring the pressure value range from 0 MPa to 2 MPa, record the pressure readings of adjacent two points as and , and let . For example, the pressure at the 100th hour is 1.36 MPa, and the pressure at the 99th hour is 1.30 MPa, then , and after calculating for all time periods, the pressure change sequence can be obtained;

[0064] The steps for obtaining the parameter are as follows: calculate the amplitude for the change of speed in the time series, and record the speed readings of two adjacent points as and , and obtain the speed change amplitude at this moment in the way of . The speed is usually monitored in the range of 0 m / s to 5 m / s. For example, the speed is 2.6 m / s at the 100th hour and 2.4 m / s at the 99th hour, then . After obtaining the complete sequence, the total speed change amplitude can be statistically analyzed;

[0065] The steps for obtaining the parameter are as follows: the method for obtaining the temperature change amplitude is similar to that of pressure or speed. For the temperature readings at adjacent time points and , subtract them and take the absolute value to get . The temperature range is between 0 °C and 200 °C. By observing the differences in all adjacent time periods, a complete sequence of temperature change amplitudes can be obtained. For example, the temperature is 125.1 °C at the 100th hour and 123.4 °C at the 99th hour, then . Record this result in sequence and overwrite the entire time series;

[0066] Calculation process:

[0067] First step, calculate the numerator , and after obtaining from the previous steps and and , first sum all to and then take the absolute value. Second step, calculate the denominator . For the first part , sum the squares of all pressure change amplitudes and then take the square root. For the second part , combine the speed change amplitude and the temperature change amplitude and accumulate them item by item. Third step, divide the result of the numerator by the denominator to obtain . For example, after accumulation, the numerator gets 520, and for the first part of the denominator, taking the square root of the sum of squares of gets approximately 42.6. For the second part, after accumulating the speed and temperature changes, it gets approximately 147.4. After adding them, the denominator is approximately 190.0. Finally, can be calculated; ;

[0068] This result shows that after amplifying the difference between the trend and the seasonal component and combining it with the comprehensive changes of pressure, speed, and temperature, a numerical fluctuation factor can be obtained. If the value is higher than 2, it indicates that there are obvious multiple fluctuation patterns in the production data. If the value is lower than 1, it indicates that the overall fluctuation amplitude is small;

[0069] After the calculation of the trend-seasonal composite fluctuation factor is completed, it is necessary to combine the previously arranged time series data set to identify whether there are significant differences in trends and seasonal fluctuations within different time periods. First, re-compare the preliminary trend and seasonal fluctuation data obtained previously. Treat each period as an observation window and record the fluctuation patterns of temperature, speed, and pressure within that period. If the temperature gradually rises from 100°C to 120°C within a certain period, the speed changes multiple times within the range of 2.0 m / s to 2.5 m / s, and the pressure stabilizes around 1.2 MPa, then mark this period as a temperature fluctuation-dominated period, and distinguish each period by the actual observation window length, such as 24 hours or 48 hours. Conduct numerical statistics on the maximum and minimum differences and average differences of temperature, speed, and pressure in each period. If the calculated threshold is exceeded, note the anomaly during subsequent annotation. The threshold can be set in combination with the historical operation data of the equipment. For example, when the fluctuation amplitude of temperature within 24 hours is greater than 20°C, mark it as excessive fluctuation. After comparing the temperature, speed, and pressure data of all periods, count the number of periods whose fluctuation amplitude accounts for more than 30% of the total number of periods to determine whether there is concentrated fluctuation within a larger time range. Use the same method to make corresponding records for those periods dominated by speed or pressure fluctuations and include them in subsequent comparisons. If the temperature and pressure are both in the high-difference range in some periods, then classify them into the composite fluctuation period sequence. By comparing with other non-overlapping periods, the coverage distribution within a longer time period can be observed. After completing the above data collation, the change rate of production data within different time periods can be judged based on the differences between periods. By corroborating this judgment conclusion with the composite fluctuation factor results obtained previously, a set of trend and seasonal analysis results containing period segmentation, index fluctuation range, and rate differentiation can finally be formed.

[0070] The steps to obtain the quality control results are as follows:

[0071] Extract each performance index sequence in the key performance snapshot and calculate the control limit. The calculation formula is:

[0072] ;

[0073] Among them, is the control limit, is the selected performance index sequence, is the maximum value in the performance index sequence, is the minimum value in the performance index sequence, is the th data point in the performance index sequence, is the total number of performance indicators, is the average value of the performance indicators;

[0074] Based on the control limits, monitor the quality fluctuations in real-time production data, mark the data points exceeding the control limits, and combine the trend and seasonal analysis results to generate quality monitoring results.

[0075] Specifically, the benefit of the formula lies in comprehensively considering the range of the sequence and the cumulative degree of deviation from the average value, making the control limits not only depend on the maximum and minimum values in, but also reflect the deviation amount between each data point and the average value in the cube root term. Through this method, a more comprehensive measure of the overall distribution situation can be made.

[0076] is the selected sequence of performance indicators, which contains observation records of the same type. Any one of temperature, speed, or pressure can be selected. The core indicator sequences such as the highest temperature, the lowest speed, and the average pressure have been generated previously. Each sequence can be regarded as . In specific operations, the performance indicators to be concerned about will be selected from the key performance snapshots, and then the values of this indicator in a continuous time interval will be arranged into a sequence .

[0077] represents the maximum value in the sequence , which is the detection of the highest value of the monitored indicator in this interval. If the indicator is temperature, its value is usually in the range of 0°C to 200°C. By traversing the collected temperature sequence, the highest temperature observation value is extracted as . To obtain this maximum value, all will be compared first, and then the item with the highest value will be recorded. For example, in the 168-hour temperature sequence, traverse to find the maximum temperature, such as 175.3°C, then there is .

[0078] represents the minimum value in the sequence , which is the detection of the lowest value of the monitored indicator in this interval.

[0079] is the th data point in the sequence , and each corresponds to a specific observation time and measurement value, and it is necessary to ensure that the time stamps and the values are in one-to-one correspondence. To extract , the th sampling time can be located from the previous key performance snapshots and the corresponding observation value can be read. For example, in the pressure sequence, may correspond to the pressure value at the 10th hour segment of the continuous record. Obtain each After continuous data collection on-site during the process, all readings will be saved item by item and then numbered in chronological order on the computing end, starting from 1 to . For example, when the monitoring period is set to 48 hours and data is collected hourly, then , corresponds to the observed value at the 1st hour, corresponds to the 2nd hour, and so on. Eventually, 48 data points can be formed and arranged in sequence as .

[0080] is the total number of performance index sequences, which is a clear count of the number of data points included in the sequence and is usually determined by the preset sampling frequency and monitoring duration. Before starting the collection, the monitoring personnel will specify the observation period. For example, choosing 24 hours or 72 hours as a monitoring cycle and combining with the hourly frequency, the number of or data points can be obtained.

[0081] is the average value of the performance index sequence . To obtain , it is necessary to add up the values of all data points in the sequence and then divide by . Its value range is determined by the specific type of index. Temperature is measured in degrees Celsius and fluctuates within 0°C to 200°C, pressure can be between 0 MPa and 2 MPa, and speed can be in the range of 0 m / s to 5 m / s. During the calculation process, first perform for the sequence . Taking the temperature sequence as an example, if it contains 168 hourly temperature records and each value is between 65°C and 175.3°C, after calculating the total sum, the total temperature sum is 23400.0°C. Then .

[0082] Calculation process:

[0083] First step, calculate , for example , , and the difference is ;

[0084] Second step, calculate , taking and according to the above example take the absolute value of the difference between each and 139.29, then sum them up and divide by 168; for example, if the total sum of the absolute values is recorded as 5096.4, then ;

[0085] Third step, calculate ;

[0086] In the fourth step, add the results of the first two parts to obtain ;

[0087] This result indicates that when is 113.41, it can be used as the control limit for this temperature sequence. If a new temperature observation value is monitored to exceed 113.41, it can be determined that this value is higher than the control limit derived from the previous fluctuation range. If the monitored temperature observation value is lower than this value, it indicates that it is still relatively stable within the current range. Combining on-site experience, further correction can also be made to in subsequent steps.

[0088] Based on the control limit obtained previously , it is necessary to periodically compare the temperature, pressure, or speed data from the real-time production process, compare each record with . Records with a temperature greater than or a pressure greater than or a speed greater than are all regarded as out of range and need to be highlighted on the on-site monitoring interface. Here the value is taken from the calculation result of the previous paragraph and checked in combination with on-site experience. Each record is continuously collected on the production line to form a sequence with a fixed time interval. For example, it can be collected and written into a data table every hour or half an hour. Technicians first read all the observed values during the comparison period, and then check the difference between each one and one by one, distinguish the positive and negative signs of the difference and attach additional identification to the abnormal points. For the temperature monitoring scenario, records exceeding can be labeled with "too high temperature", in the speed monitoring scenario, a "too high speed" label is set, and in the pressure monitoring scenario, a "too high pressure" label is set. All the marked data is listed together and appended with a timestamp, equipment number, and observed value format for convenient subsequent review. In addition, the obtained trend and seasonal analysis results will also be synchronized for reference to check whether the situation of exceeding has occurred multiple times in a certain period of history. If it is found that the data exceeds this control limit in multiple adjacent observation cycles, it indicates that the fluctuation degree in this time period is relatively high. After the comparison and marking are completed, technicians will output the quality monitoring results and record the occurrence time and corresponding value of each abnormal point. The marked points and unmarked points together constitute the final quality monitoring results.

[0089] The steps to obtain the optimized query path are as follows:

[0090] Based on the quality monitoring results, analyze the access mode of production data queries, count the distribution characteristics of query requests, screen the data fields suitable for hash indexes or B-tree indexes, determine the index construction plan, and generate an index strategy design plan;

[0091] According to the index strategy design plan, construct a hash index to handle equality queries, determine the hash function and calculate the key-value mapping relationship. At the same time, establish a B-tree index for range queries, perform index node splitting and adjustment operations, and generate an optimized query path.

[0092] Specifically, based on the quality monitoring results, first collect all production data query requests within the past seven days and record their access times and usage frequencies. For each query, retrieve its read pattern for fields such as temperature, speed, or pressure. If the query request only uses equality conditions such as "temperature = 120°C" or "pressure = 1.2 MPa", it is classified into the equality query set. If the query request uses range conditions such as "speed > 1.0 m / s and speed < 3.0 m / s", it is classified into the range query set. After classifying all query requests, count the proportions of equality queries and range queries in chronological order. For example, if after comparing 10,000 query requests, it is found that approximately 5,000 belong to equality queries and most of them only lock the temperature or pressure fields, and another 3,000 belong to range queries and the more common case is that the speed field is filtered between 0 m / s and 5 m / s, and the remaining 2,000 are distributed in a small number of mixed conditions or other types of conditions. Then, merge these classification results with the query execution time-consuming information to evaluate where the access bottleneck lies. To refine the screening logic, a threshold can be set to distinguish whether the access volume is high enough. For example, if the frequency of a field participating in equality queries within seven days is greater than 2,000 times, it is determined that the access volume is high. This threshold is determined based on the statistical method obtained from the on-site detection log. More than 2,000 times is considered to have a greater impact on the retrieval efficiency. Similarly, for range queries, those with an access frequency of more than 1,000 times can be regarded as target fields for considering index construction. For those fields with low access volume or infrequent interval conditions, no indexing is done temporarily. Subsequently, further screen out the field names, query operation types, and occurrence times based on the statistical distribution characteristics of the access patterns, list the field combinations that meet the threshold in a list form and classify them into the candidate sets of hash index or B-tree index. When a field is mainly used for equality matching, it is included in the hash index candidate list. When a field is mainly used for interval matching, it is included in the B-tree index candidate list, and record its actual proportion such as "the number of range query times of the speed field accounts for 40% of the whole" and "the number of equality query times of the temperature field accounts for 50% of the whole", etc. Finally, refer to these candidate field combinations together with the quality monitoring results to confirm whether it is necessary to perform index acceleration on the production data analysis. After finishing the arrangement, the index construction plan can be obtained and placed in the plan description according to the classification of hash index or B-tree index, and then integrate the plan description into an index strategy design plan.

[0093] According to the index strategy design plan, it is necessary to establish a hash index for equality queries first and clarify the selection method of the hash function. First, select the most frequently used field from all field candidates. For example, for temperature or pressure, according to the statistical frequency obtained previously, the number of equality queries for it is greater than two thousand times. So, this field is used as the object for constructing the hash index. To generate the hash function, modulo or other hashing methods can be selected and combined with the value range of the specific field. For example, when the temperature is between 0°C and 200°C, the temperature value can be mapped to the hash bucket by taking the modulus of 201. Or, for pressure in the range of 0 MPa to 2 MPa, it can be discretized according to the decimal precision, and then the hashing function is executed for mapping to obtain the key values of each record. When hash conflicts occur, corresponding chaining or open-address strategies are needed to allocate bucket space. After completing the hash index, a B-tree index is established for range queries. Taking the speed field as an example, a large range search between 0 m / s and 5 m / s will occur for speed. So, the speed field is sorted in ascending order to construct the B-tree index nodes. Each node maintains a certain splitting threshold to control the number of child nodes. This splitting threshold can be set according to the established storage structure to at most accommodate fifty records per node. When a node is loaded with more than fifty data, it will automatically split into two child nodes or adjust the hierarchy. The pointers after splitting point to each speed range, and the keywords are reserved in the upper-level nodes for fast search. For the case of querying both speed and temperature simultaneously, the leaf nodes can also be refined according to the temperature distribution inside the B-tree, but it is necessary to ensure that the tree depth will not be increased excessively. For this purpose, the preset maximum level can be set to six layers, and when it exceeds six layers, node merging or re-partitioning is performed. Finally, according to the node structure after all hash indexes and B-tree indexes are established, an optimized query path for subsequent retrieval is generated according to the query fields.

[0094] The steps to obtain the index optimization result are as follows:

[0095] Based on the optimized query path, analyze the query execution plan, identify multiple fields included in the query pattern, extract the field combination with the highest query frequency, screen the set of fields applicable to the composite index, and generate a composite index construction plan;

[0096] According to the composite index construction plan, calculate the index adaptability score. The calculation formula is:

[0097] ;

[0098] Among them, is the index adaptability score, is the access frequency of the th query field, is the proportion of the th field in the data table, is the number of times the index of the th field has been hit, The total number of fields participating in the index construction;

[0099] Based on the index fitness score, the hierarchical relationship of the composite index is adjusted, the order of the index keys is optimized, and the index optimization results are generated.

[0100] Specifically, based on the optimized query path, it is necessary to sort out the existing access statistics data for different types of query requirements and identify the temperature fields, speed fields, and pressure fields involved in the query pattern. First, analyze the query execution plan one by one from the previously integrated hash index and B-tree index statistical records, and extract the fields in the data table whose specific usage frequency exceeds the predetermined threshold. The threshold can be set based on standards such as the access frequency is greater than two thousand times or the value accounts for more than 30% within seven days. These standards are established by doing statistics based on the query logs collected by the team and finding that more than two thousand times belong to a high access volume. Then, in such high-access fields, further check whether the query involves the joint conditions of multiple fields. If it is detected that the temperature and speed are combined for more than one thousand joint searches, the combination of these two fields is recorded as a candidate combination. If the number of joint query accesses of pressure and temperature within a week is also higher than the standard, it is also included in the candidate combination. All candidate combinations are sorted in descending order according to the total access frequency and those combinations with less than five hundred times are excluded. Confirm that the remaining combinations are all distributed as multiple fields appearing at the same time. High-frequency conditions, for these remaining combinations, their field names are compared with the data table storage structure to check the column position and field length of the field in the data table. The existing index structure in the hash index or B-tree index is also checked to see if it can cover multi-field filtering. If it cannot meet the requirements, it is marked as a potential composite index target. The sorting or grouping requirements that often appear in the query mode are further analyzed to see if they are consistent with the field combination. The fields that meet the sorting or grouping rules, such as temperature, time, and speed, are paired as another candidate combination. Finally, a series of high-access frequency field combinations that have been screened out above are summarized, such as "temperature-speed", "pressure-temperature" or "temperature-speed-time", etc., and the proportion of each field in the data table is checked. For example, the temperature column accounts for 20% in the table, which means that its data volume may be relatively medium, and the speed column accounts for 12%, which means that it occupies less space. For the fields that need to be jointly searched, they are combined and packaged into a set of indexes to be composited, and it is estimated that how many query steps can be reduced for each combination to enhance the retrieval efficiency. Combining this information can generate a composite index construction plan.

[0101] The formula is beneficial because it integrates the processing of multiple aspects of information such as field access frequency, data table proportion, and index hit count. This method magnifies the difference between high hits and low hits in access frequency, and uses the denominator to Avoid a single field having an overly large proportion, which may lead to an unbalanced score, and thus obtain a more balanced reference metric in the decision-making of multi-field composite indexes.

[0102] Indicates the access frequency of the th query field, which is the count value of the number of times this field is queried and used within a certain period. Obtaining this value depends on scanning each entry in the database query log. For example, all SQL queries are recorded within 240 hours of device operation. Each occurrence of "WHERE temperature = xxx" or "WHERE pressure < xxx" is counted as one access, and then the total access frequency is aggregated to each field, and this frequency is then distributed to . When a certain field is accessed 3500 times, then . In the case of multiple fields, statistics need to be done separately and presented in field order . A complete example is: When selecting three items, temperature, speed, and time, for a composite index, first count the access frequency of temperature as 2800 times, the access frequency of speed as 2200 times, and the access frequency of time as 1500 times. Then .

[0103] Is the proportion of the th field in the data table, indicating the storage proportion or data volume proportion that this field occupies among all records in the entire table. This value needs to be obtained from the data storage structure level. For example, first count the actual number of bytes or rows of the entire table, and then perform a division operation on the number of bytes or rows occupied by this field to obtain a fraction between 0 and 1. If the temperature column occupies 15% of the storage in the entire table, then , if the speed column occupies 10% then , and this kind of proportion can also be calculated by the number of rows or data chunks. Example: When the total size of the table is 100MB and the temperature column occupies 15MB, then .

[0104] Indicates the number of times the index of the th field has been hit, which is a number extracted from the database or system log, indicating the number of times the index for this field has been successfully used in a query within a certain period. For example, the hit frequency can be counted through the hash index or B-tree index generated previously. If the temperature field has been successfully retrieved by the index 2000 times within a week, then .

[0105] Indicates the total number of fields participating in index construction, which is determined according to the number of field combinations selected in the previous step. After data analysis, it is found that only three items, temperature, speed, and time, meet the combined requirements. Therefore .

[0106] Calculation process:

[0107] Step 1: Calculate the numerator , for example, there are three groups of fields, and its , , , then for each group, first calculate :

[0108] ;

[0109] Then calculate the numerator successively:

[0110] ;

[0111] Step 2: Calculate the denominator , first calculate each item :

[0112] ;

[0113] Then multiply the three and take the fourth root:

[0114] ;

[0115] Step 3: Divide the numerator by the denominator to get to be 0.00395;

[0116] This result indicates that when is approximately 0.00395, the adaptation score of the three-field composite index is relatively small, meaning that this combination is not optimal in terms of access frequency, data table occupancy, and hit count. If the calculated from other field combinations in the future is much greater than 0.00395, it is more worthy of being selected first. If exceeds 0.01, it can be regarded as a relatively high adaptation degree, indicating that this field combination may be more suitable for building a composite index in terms of index efficiency. Technicians can make a decision after comparing the values of multiple combinations.

[0117] Based on the index fitness score, we first conduct a hierarchical review of the currently established composite index structure, and check the number of fields and node distribution status contained in each layer of the index tree item by item. Fields with high access frequency and index hit times exceeding the specified threshold are placed in front of the upper nodes, and fields with relatively low access volume or a proportion that does not meet the screening criteria of the previous stage are placed in the back layer, so as to form a reconfiguration of the hierarchical nodes. Then we observe whether the excessive number of nodes causes the index level to be too deep. If the number of nodes in a certain layer is greater than the system experience value, such as fifty, the nodes are split and the field key value range covered by each node after the split is recorded. The experience value of fifty nodes is obtained by observing the past database retrieval performance. After each node is split, the number of pointers from the parent node to the child node must be checked. Whether the quantity matches the field order. If the original arrangement order of the index key is significantly different from the actual access mode, the order of the fields is adjusted. The number of query sorting or grouping times within seven days can be referred to to determine which field is more suitable to be ranked first. If the temperature field is used for sorting or grouping multiple times within seven days and the cumulative number exceeds two thousand times, the temperature field is prioritized in the arrangement of the index key. The speed or pressure fields that appear frequently are arranged in descending order according to the proportion. After the rearrangement is completed, the nodes are merged again to ensure that each node remains balanced in the index tree, not exceeding the maximum layer depth allowed by the system, such as six layers. Finally, all node pointers and field mapping relationships are integrated and the field order is marked to form an optimized multi-layer composite index description and record it as the index optimization result.

[0118] The steps to obtain the influencing factor analysis results are as follows:

[0119] Based on the index optimization results, extract production parameter data, including temperature, pressure and speed, remove abnormal and duplicate data, align production parameters with quality results, and establish a production parameter matching data set;

[0120] According to the production parameter matching data set, a decision tree model is constructed to analyze the quality impact of production parameters in different intervals, generate a classification structure based on the decision path, calculate the split contribution value of each production parameter in the classification structure, extract production variables, and obtain the classification parameter impact weight data;

[0121] Based on the classification parameter impact weight data, a random forest model is constructed, and multiple rounds of feature selection are performed to calculate the stability and classification contribution of production parameters under different decision tree combinations, analyze the interaction between different parameters, rank the importance of production parameters, and generate influencing factor analysis results.

[0122] Specifically, based on the index optimization results, it is necessary to select three parameters, namely temperature, pressure, and speed, from the historical production data as the objects of concern. First, arrange all the data previously collected by the equipment in ascending order of timestamp. For multiple duplicate records that appear at the same moment, retain the first one and eliminate the subsequent duplicate data. If abnormal observed values are found, conduct a check according to the preset threshold range. For example, compare the temperature with the range of 0°C to 200°C. If it exceeds 200°C, directly eliminate it. Compare the pressure with the range of 0 MPa to 2 MPa. If the value is less than 0 MPa or greater than 2 MPa, mark it as abnormal and exclude it. If a negative number or a value exceeding 5 m / s is detected for the speed within the range of 0 m / s to 5 m / s, also eliminate it. The above thresholds are obtained through statistics after multiple inspections of the safe operating range of the equipment. After all the cleaning is completed, read the product qualification rate or failure rate field in the quality result data within the same time period and match it with the timestamp, temperature, pressure, and speed. If the quality result record is missing within the same time period, retrieve the nearest quality result forward or backward and temporarily bind it to the temperature, pressure, and speed of this time period. After completing the record alignment for all time periods and confirming each item is correct one by one, aggregate the data of each moment into a row containing four items: temperature, pressure, speed, and quality result. Combine all the items into a production parameter matching data set.

[0123] According to the production parameter matching data set, before training the decision tree model, establish an interval mapping for the numerical values of the three parameters of temperature, pressure, and speed according to a fixed interval. For example, divide the temperature into twenty intervals of every ten degrees from 0°C to 200°C, divide the pressure into ten intervals of every 0.2 MPa from 0 MPa to 2 MPa, and divide the speed into five intervals of every 1 m / s from 0 m / s to 5 m / s. Correspondingly, divide the quality result field into two binary categories: qualified or unqualified. When the records in the data set enter the decision tree construction process, first read the interval where the temperature is located and split into two branches according to the binary quality result, and then perform the next-level split according to the interval where the pressure or speed is located. If the temperature interval in a split node statistically contains more qualified records, label it as the qualified main branch and calculate the corresponding split contribution value. Recursively establish multiple levels of decision paths downward in turn, and when calculating the split contribution value at each node, count in which interval segments the temperature, pressure, and speed appear in this node and the ratio of qualified or unqualified records in this interval segment. Finally, after the decision tree converges, one or more complete classification paths will be formed to judge the association between temperature and pressure or speed in a certain interval, and calculate the sum of the split contribution values of all nodes to measure the overall impact of the three parameters on the classification result. On this basis, extract the parameters of each high-contribution split node, form a production variable list for this paragraph of data, and summarize the comprehensive influence weight of each parameter, so as to obtain the classification parameter influence weight data.

[0124] Based on the influence weights of classification parameters, first select three parameters, namely temperature, pressure, and velocity, as the main features and configure several decision trees with different depths in the random forest model. Each decision tree performs node splitting according to the interval partitioning method obtained in the previous step. During training, randomly select 80% of the production parameter matching datasets to generate splitting rules and reserve 20% for verification. If the depth of a certain decision tree exceeds a predefined limit, such as ten levels, then prune the nodes of this tree to avoid overfitting. Run the same process multiple times for all the trained decision trees to accumulate the metrics of multiple trees. After summing up the temperature, pressure, and velocity splitting node counts used by each tree during classification, calculate their contribution degrees to the overall classification effect, and observe whether the fluctuation range of these parameters among different trees exceeds a predefined stability threshold, such as 5%. If the node contribution value of temperature remains around 30% in ten trees, it is regarded as a stable parameter; otherwise, it is regarded as a parameter with large fluctuations. Similarly, give the corresponding stability values for the analyzed pressure or velocity. Finally, calculate the importance ranking by comprehensively considering the contribution ratios and fluctuation situations of each parameter in all trees, so as to obtain a list of the most influential production parameters and use this ranking result as the analysis result of influencing factors.

Claims

1. A deep concentration production scheduling big data analysis system, characterized in that, The system includes: A data collection module that collects real-time data on concentration production, including temperature, speed, and pressure information, generates a real-time production data set; extracts key performance indicators based on the real-time production data set and generates a key performance snapshot; A data processing module that performs time series analysis based on the key performance snapshot, identifies trends and seasonal fluctuations in production data, and generates a trend and seasonal analysis result; monitors quality fluctuations of the key performance snapshot, combines the trend and seasonal analysis result, and generates a quality monitoring result; A data indexing module that designs and implements hash and B-tree indexing strategies based on the quality monitoring result, optimizes data query efficiency, and generates an optimized query path; performs composite index construction according to the optimized query path for multi-condition query optimization and generates an index optimization result; A decision support module that analyzes the impact of production parameters on product quality using decision tree and random forest models based on the index optimization result and generates an impact factor analysis result; The steps for obtaining the impact factor analysis result are: Based on the index optimization result, extract production parameter data, including temperature, pressure, and speed, remove abnormal data and duplicate data, align the production parameters with the quality results, and establish a production parameter matching data set; According to the production parameter matching data set, construct a decision tree model, analyze the quality impact of production parameters in different intervals, generate a classification structure based on the decision path, calculate the splitting contribution value of each production parameter in the classification structure, extract production variables, and obtain classification parameter impact weight data; Based on the classification parameter impact weight data, construct a random forest model, perform multiple rounds of feature selection, calculate the stability and classification contribution of production parameters under different decision tree combinations, analyze the interaction between different parameters, rank the importance of production parameters, and generate an impact factor analysis result.

2. The deep concentration production scheduling big data analysis system according to claim 1, wherein The steps for obtaining the real-time production data set are: Monitor the real-time data of temperature, speed, and pressure on the concentration production line, mark and format the time stamp for each collected data point to form the original real-time monitoring data; Based on the original real-time monitoring data, remove error data and outliers, fill in missing data points, and use statistical methods to evaluate the consistency of the data to generate real-time production data after quality inspection; Based on the real-time production data after quality inspection, perform data synchronization processing, align the data from different sensor sources according to a unified time stamp, and form a real-time production data set.

3. The deep concentration production scheduling big data analysis system according to claim 1, characterized in that The steps for obtaining the key performance snapshot are: From the real-time production data set, screen key performance indicators, including the highest temperature, the lowest speed, and the average pressure, and generate preliminarily screened performance indicator data; Based on the preliminarily screened performance indicator data, construct a key performance snapshot, and the performance snapshot is presented in the form of views and charts by integrating the highest temperature, the lowest speed, and the average pressure indicators.

4. The deep concentration production scheduling big data analysis system according to claim 1, characterized in that The steps for obtaining the trend and seasonal analysis result are: Based on the key performance snapshots, a time-series data set is formed in chronological order, and the change ranges and differences of each indicator in different time periods are analyzed to obtain preliminary trend and seasonal fluctuation data; According to the preliminary trend and seasonal fluctuation data, calculate the trend-seasonal composite fluctuation factor, and the calculation formula is: ; Among them, represents the trend seasonal composite fluctuation factor, represents the value of the trend component at the th time point, represents the value of the seasonal component at the th time point, m is the total number of time points, represents the pressure change amplitude at the th time point, represents the speed change amplitude at the th time point, represents the temperature change amplitude at the th time point; Based on the trend-seasonal composite fluctuation factor and combined with the time-series data set, determine the trend and seasonal fluctuation patterns in the production data, judge the change rate of the production data in different time periods, and obtain the trend and seasonal analysis results.

5. The deep concentration production scheduling big data analysis system according to claim 1, wherein The steps for obtaining the quality monitoring results are as follows: Extract each performance indicator sequence in the key performance snapshots and calculate the control limits. The calculation formula is: ; Among them, L is the control limit, B is the selected sequence of performance indicators, is the maximum value in the sequence of performance indicators, is the minimum value in the sequence of performance indicators, is the th data point in the sequence of performance indicators, n is the total number of performance indicators, is the average value of the performance indicators; Based on the control limits, monitor the quality fluctuations in the real-time production data, mark the data points exceeding the control limits, and combine with the trend and seasonal analysis results to generate the quality monitoring results.

6. The deep concentration production scheduling big data analysis system according to claim 1, characterized in that The steps for obtaining the optimized query path are as follows: Based on the quality monitoring results, analyze the access patterns of production data queries, statistically analyze the distribution characteristics of query requests, screen the data fields suitable for hash indexes or B-tree indexes, determine the index construction plan, and generate the index strategy design plan; According to the index strategy design plan, construct a hash index for processing equality queries, determine the hash function and calculate the key-value mapping relationship, and at the same time establish a B-tree index for range queries, and perform index node splitting and adjustment operations to generate an optimized query path.

7. The deep concentration production scheduling big data analysis system according to claim 1, wherein The steps for obtaining the index optimization results are as follows: Based on the optimized query path, analyze the query execution plan, identify the multiple fields included in the query pattern, extract the field combination with the highest query frequency, screen the set of fields suitable for composite indexes, and generate the composite index construction plan; According to the composite index construction plan, calculate the index adaptability score, and the calculation formula is: ; where R is the index adaptation degree score, is the access frequency of the k-th query field, is the proportion of the k-th field in the data table, is the number of times the index of the k-th field has been hit, is the total number of fields participating in index construction; Based on the index adaptability score, adjust the hierarchical relationship of the composite index, optimize the arrangement order of the index keys, and generate the index optimization results.

8. The deep concentration production scheduling big data analysis method of the deep concentration production scheduling big data analysis system according to any one of claims 1-7, characterized in that, Include the following steps: Collect concentrated real-time production data, including temperature, speed, and pressure, generate a real-time production data set, and based on the real-time production data set, extract temperature extremes, speed change rates, and pressure stability indicators to obtain key performance snapshots; Based on the key performance snapshots, identify the temperature extreme periodicity, speed change trend, and pressure fluctuation law to obtain the trend and seasonal analysis results; Conduct quality fluctuation monitoring on the trend and seasonal analysis results, and combine the quality monitoring values and fluctuation frequencies to obtain the quality monitoring results; Based on the quality monitoring results, design and implement hash and B-tree index strategies to optimize the data query efficiency and obtain the index optimization results; Based on the index optimization results, analyze the impacts of temperature, speed, and pressure on product quality, generate an impact factor analysis, and obtain the impact factor analysis results.

Citation Information

Patent Citations

  • Circuit board quality abnormity prediction system based on machine intelligent learning

    CN119129365A