Deep concentration production scheduling big data analysis system and method

Through the deep concentration production scheduling big data analysis system, production data is collected and analyzed in real time, combined with index optimization and machine learning models, the problems of slow query speed and in-depth relationship mining in the existing technology are solved, and efficient decision-making and lean production are achieved.

CN120011373AActive Publication Date: 2025-05-16HUADIAN QINGDAO POWER GENERATION COMPANY

Patent Information

Application Number
CN202510502918.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

When the existing big data analysis technology is large in volume or complex in query conditions, the query speed is significantly reduced, making it difficult to meet the requirements of fast and efficient decision-making, and the relationship between production parameters and product quality is not thorough enough.

Method used

Design a deep-concentrated production scheduling big data analysis system, collect production data in real time through the data acquisition module, extract key performance indicators, perform time series analysis and quality fluctuation monitoring, optimize data queries in combination with hash and B-tree index strategies, and use decision trees and random forest models to analyze the impact of production parameters on product quality.

Benefits of technology

It realizes fine-grained capture and analysis of production data, improves the accuracy and reliability of trend prediction, promptly detects and locates quality abnormalities, optimizes data query efficiency, identifies key influencing factors, improves production efficiency, reduces costs, and stabilizes product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011373A_ABST
    Figure CN120011373A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data analysis, in particular to a deep concentration production scheduling big data analysis system and method, and the system comprises a data collection module which collects concentration production real-time data, including temperature, speed and pressure information, and generates a real-time production data set; according to the method, multi-dimensional data such as temperature, speed and pressure in the production process are collected in real time, key performance indexes are accurately extracted, fine-grained capture and analysis of production information are achieved, production performance snapshots are formed, deep time sequence analysis is carried out on the basis, and the production trend and seasonal fluctuation are identified and predicted; the precision and reliability of trend prediction are improved; meanwhile, in combination with trend and seasonal analysis results, refined quality fluctuation monitoring is carried out on the production process, the fluctuation range of abnormal quality is found and positioned in time, and early warning and efficient management and control of production quality risks are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data analysis, and in particular to a deep concentrated production scheduling big data analysis system and method. Background Art

[0002] Big data analysis refers to the process of extracting valuable information, patterns and knowledge from large and diverse data sets using advanced data storage, processing and analysis technologies. Its core technology includes data collection, storage, preprocessing, mining, analysis, visualization and real-time computing, involving distributed computing frameworks, data mining algorithms, machine learning methods and cloud computing platforms.

[0003] In the existing technology, the data retrieval mechanism is relatively simple, usually relying on linear or simple indexing methods, which leads to a significant decrease in query speed when the amount of data is large or the query conditions are complex, making it difficult to meet the requirements of fast and efficient decision-making, especially in emergency conditions, where delays affect the timeliness and accuracy of decision-making. In addition, the relationship between production parameters and product quality is not explored in depth, and often only stays on the surface description of correlation, lacking effective identification of key influencing factors. Therefore, improvements are needed. Summary of the invention

[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a deep concentrated production scheduling big data analysis system and method.

[0005] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme: A deep concentrated production scheduling big data analysis system comprises: A data acquisition module collects and condenses real-time production data, including temperature, speed and pressure information, to generate a real-time production data set; extracts key performance indicators based on the real-time production data set to generate a key performance snapshot; The data processing module performs time series analysis based on the key performance snapshots, identifies trends and seasonal fluctuations in the production data, and generates trend and seasonal analysis results; monitors the quality fluctuations of the key performance snapshots, and generates quality monitoring results in combination with the trend and seasonal analysis results; The data index module designs and implements hash and B-tree index strategies based on the quality monitoring results, optimizes data query efficiency, and generates optimized query paths; executes composite index construction according to the optimized query paths, optimizes multi-condition queries, and generates index optimization results; The decision support module uses decision trees and random forest models to analyze the impact of production parameters on product quality based on the index optimization results and generates influencing factor analysis results.

[0006] Preferably, the steps of acquiring the real-time production data set are: Monitor real-time data of temperature, speed and pressure on the concentration production line, time-stamp and format each data point collected to form raw real-time monitoring data; Based on the original real-time monitoring data, error data and outliers are removed, missing data points are filled, and statistical methods are used to evaluate the consistency of the data to generate real-time production data after quality inspection; Based on the real-time production data after the quality inspection, data synchronization processing is performed to align data from different sensor sources according to a unified timestamp to form a real-time production data set.

[0007] Preferably, the steps of obtaining the key performance snapshot are: Filter key performance indicators from the real-time production data set, including maximum temperature, minimum speed and average pressure, to generate preliminary filtered performance indicator data; Based on the preliminary screened performance indicator data, a key performance snapshot is constructed, and the performance snapshot is presented in the form of views and charts by integrating the maximum temperature, minimum speed and average pressure indicators.

[0008] Preferably, the steps for obtaining the trend and seasonality analysis results are: Based on the key performance snapshots, the data are arranged in chronological order to form a time series data set, and the change range and difference of each indicator in different time periods are analyzed to obtain preliminary trend and seasonal fluctuation data; Calculate the trend-seasonal composite fluctuation factor based on the preliminary trend and seasonal fluctuation data; Based on the trend seasonal composite fluctuation factor, combined with the time series data set, the trend and seasonal fluctuation pattern in the production data are determined, the change rate of the production data in different time periods is judged, and the trend and seasonal analysis results are obtained.

[0009] Preferably, the steps for obtaining the quality monitoring results are: Extracting each performance indicator sequence in the key performance snapshot and calculating the control limit; Based on the control limits, quality fluctuations in real-time production data are monitored, data points exceeding the control limits are marked, and quality monitoring results are generated by combining the trend and seasonal analysis results.

[0010] Preferably, the steps for obtaining the optimized query path are: Based on the quality monitoring results, analyze the access mode of production data query, count the distribution characteristics of query requests, screen data fields suitable for hash index or B-tree index, determine the index construction plan, and generate the index strategy design plan; According to the index strategy design scheme, a hash index is constructed for processing equal value queries, a hash function is determined and a key-value mapping relationship is calculated, and a B-tree index is established for range queries, index node splitting and adjustment operations are performed, and an optimized query path is generated.

[0011] Preferably, the steps for obtaining the index optimization result are: Based on the optimized query path, the query execution plan is analyzed, multiple fields included in the query pattern are identified, the field combination with the highest query frequency is extracted, the field set suitable for the composite index is screened, and a composite index construction plan is generated; According to the composite index construction scheme, calculating the index fitness score; Based on the index optimization results, analyze the impact of temperature, speed, and pressure on product quality, generate an influencing factor analysis, and obtain the influencing factor analysis results.

[0012] Compared with the prior art, the advantages and positive effects of the present invention are: In the present invention, by collecting multidimensional data such as temperature, speed and pressure in the production process in real time, key performance indicators are accurately extracted, and fine-grained capture and analysis of production information is achieved to form a production performance snapshot, and on this basis, in-depth time series analysis is performed to identify and predict production trends and seasonal fluctuations, and improve the accuracy and reliability of trend prediction; at the same time, combined with the trend and seasonal analysis results, the production process is finely monitored for quality fluctuations, and the fluctuation range of quality anomalies is discovered and located in time, so as to achieve early warning and efficient control of production quality risks. In addition, hash and B-tree index strategies are used to optimize query paths, implement composite index construction for multiple conditions, shorten the query response time of multidimensional data, and greatly improve the processing efficiency of complex queries; the correlation between each production parameter and product quality is analyzed through decision trees and random forest models, and key influencing factors are effectively identified to help production managers more accurately grasp the direction of process parameter adjustment, thereby achieving lean production goals, and significantly improving production efficiency, reducing production costs and product quality stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a system flow chart of the present invention. DETAILED DESCRIPTION

[0014] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0015] See also Figure 1The present invention provides a technical solution: a deep concentrated production scheduling big data analysis system comprising: The data acquisition module collects and condenses real-time production data, including temperature, speed and pressure information, to generate real-time production data sets; extracts key performance indicators based on the real-time production data sets to generate key performance snapshots; The data processing module performs time series analysis based on key performance snapshots, identifies trends and seasonal fluctuations in production data, and generates trend and seasonal analysis results; monitors quality fluctuations of key performance snapshots, and generates quality monitoring results based on trend and seasonal analysis results; The data index module designs and implements hash and B-tree index strategies based on quality monitoring results, optimizes data query efficiency, and generates optimized query paths; executes composite index construction based on the optimized query path, optimizes multi-condition queries, and generates index optimization results; The decision support module uses decision trees and random forest models to analyze the impact of production parameters on product quality based on index optimization results and generates influencing factor analysis results.

[0016] The steps to obtain the real-time production data set are: Monitor real-time data of temperature, speed and pressure on the concentration production line, time-stamp and format each data point collected to form raw real-time monitoring data; Based on the original real-time monitoring data, remove error data and outliers, fill in missing data points, use statistical methods to evaluate data consistency, and generate real-time production data after quality inspection; Based on the real-time production data after quality inspection, data synchronization processing is performed to align data from different sensor sources according to a unified timestamp to form a real-time production data set.

[0017] Specifically, the real-time data of temperature, speed and pressure are monitored on the concentration production line. Multiple sensors installed at different positions of the production line are selected and calibrated in advance. The measurement range of each sensor is set according to the historical test data. For example, the temperature range is set to 0°C to 200°C and the average value and standard deviation are calculated based on the previous test samples to determine the upper limit threshold of temperature anomaly Ttemp = average temperature + 2×standard deviation. If the measured temperature is higher than Ttemp, it is determined that it will not be included in the preliminary record temporarily. Similarly, the speed range is set to 0m / s to 5m / s and the pressure range is set to 0MPa to 2MPa. Each sensor After preliminary verification according to the above range at the acquisition end, the verified real-time data will be read one by one and configured with a unique timestamp. In order to demonstrate the threshold setting process, it can be assumed that the temperature sample is 50 observations with an average temperature of 70°C and a standard deviation of 5°C, then Ttemp=70+2×5=80°C. If the observation value exceeds 80°C, it will be temporarily regarded as an out-of-limit record at this stage. Each record consists of three measurement values ​​of temperature, speed and pressure and the corresponding time stamp, and the measurement units of temperature, speed and pressure are written into the same data carrier in a unified format. Finally, all records are compiled to obtain the original real-time monitoring data.

[0018] Based on the original real-time monitoring data obtained above, the complete records of temperature, speed and pressure are taken to detect error data and abnormal values. The judgment of abnormal values ​​can be carried out by adding and subtracting a fixed multiple of the standard deviation of the average value. For example, the average value of all temperature records can be statistically calculated to be 70°C, the standard deviation is 5°C, and the threshold value Ttemp=70+2×5=80°C or below 60°C is set as the boundary. If the temperature is higher than 80°C or lower than 60°C, it is judged as an abnormal record. The speed and pressure are also obtained in the same way to get the speed abnormal boundary Vlimit and the pressure abnormal boundary Plimit, and the detected abnormal values ​​are analyzed. Normal records are eliminated, and then missing records are interpolated. For example, a temperature value is estimated by linear interpolation for the vacant position between two adjacent normal temperature data. The missing points for speed and pressure are filled in the same way. Then the variance comparison method is used to evaluate the data consistency, that is, variance analysis is performed on the temperature, speed and pressure in the same time period, and their discrete distribution in different time periods is compared to view the fluctuation. If the discreteness of a certain time period exceeds the threshold value obtained according to historical calculations, for example, when the variance is set to be greater than 4, it is considered as inconsistent data and can be eliminated or reviewed again. Finally, the retained records can be summarized to obtain real-time production data after quality inspection.

[0019] Based on the real-time production data after the above quality inspection, qualified records of temperature, speed and pressure are selected for data synchronization processing, and the timestamps corresponding to each record are accurate to the second level. Compare whether all the records of different sensors at the same time point exist. If only some sensors have data at a certain time point, the measurement results of the most recent time point or the data of the current and adjacent moments can be used for interpolation and supplement. For example, when the speed sensor is missing at t=10s and the temperature and pressure sensors have records at t=10s, the speed values ​​of t=9s and t=11s can be used to calculate the speed of t=10s through linear interpolation. The threshold setting can refer to the statistical results to determine if the speed difference between adjacent moments is too large. If the inspection standard of step size Δv=0.5m / s is set, if the difference between the interpolated speed and the adjacent actual record is greater than 0.5m / s, it is marked as suspected mismatched data and retained for subsequent review. All synchronized data are re-sorted in ascending order of timestamps and the three values ​​of temperature, speed and pressure are listed uniformly. Finally, the aligned data are merged into records under the same time base to form a real-time production data set.

[0020] The steps to obtain key performance snapshots are: Filter key performance indicators from real-time production data sets, including maximum temperature, minimum speed and average pressure, to generate preliminary filtered performance indicator data; Based on the preliminary screened performance indicator data, a key performance snapshot is constructed. The performance snapshot is presented in the form of views and charts by integrating the maximum temperature, minimum speed and average pressure indicators.

[0021] Specifically, the temperature, speed, pressure and other parameters recorded at several moments have been obtained in the real-time production data set. Based on the pre-established temperature range of 0°C to 200°C, speed range of 0m / s to 5m / s and pressure range of 0MPa to 2MPa, all data points are compared one by one and the temperature, speed and pressure at each moment are recorded. For the temperature data, the maximum value function in statistics can be used. Compare all temperature records that meet the specified range and store the function result as the highest temperature value. For speed data, the minimum value function can be used. Get the lowest speed value from the data set that meets the range, and calculate the pressure data according to its total record length when making statistics. Calculate the average ,in For the Qualified pressure records, through the above operation comparison and statistics, the representative values ​​of the three key indicators can be screened out and form a performance indicator set. If the temperature, speed or pressure data at certain moments are not within the above pre-established range, they are regarded as failure records and are not included in the statistical range. In the example calculation, if the maximum temperature is 180℃, the minimum speed is 0.8m / s, and the average pressure is 1.2MPa, 180℃, 0.8m / s and 1.2MPa are written into the performance indicator set respectively. If you want to set the inspection threshold again, you can analyze it based on the historical records. For example, the temperature can be compared with the average value of the sample number 50 and the sum of two times the standard deviation to obtain a temperature judgment value. The speed and pressure are also processed in the same way. Finally, all the data that pass the numerical verification are counted for the highest temperature, the lowest speed and the average pressure and registered in the same data table. Finally, the statistical results are uniformly recorded as the preliminary screened performance indicator data.

[0022] Based on the preliminary screening performance indicator data obtained in the previous step, the three values ​​of maximum temperature, minimum speed and average pressure are visualized and integrated as key contents. By drawing various charts such as temperature curve, speed curve and pressure curve on the same time axis for comparison and presentation, data points within a fixed time period can be selected and highlighted using line charts or histograms. In order to expand the amount of information, the historical maximum temperature can be compared with the current maximum temperature and the difference can be marked. The specific change trajectory of the minimum speed can also be compared on the same interface and the trend difference between the average pressure can be interactively displayed. In order to facilitate the identification of situations exceeding specific thresholds, color marking can be added to the chart. For example, if the temperature exceeds 200°C, it will be marked with a red line segment, and if the speed is less than 0.5m / s is marked with a yellow line segment, and pressure exceeding 1.5MPa is marked with an orange line segment. These thresholds can be determined according to the equipment manufacturer's recommendations or operating specifications and can be calculated through actual examples. For example, 100 records of temperature data were collected in the past week and the average value was 80℃ and the standard deviation was 6℃. The threshold can be set to 80+2×6=92℃. If some points exceed 92℃, they are marked as records that require attention. The speed and pressure are also calculated in the same way to obtain the threshold value after the average value and standard deviation are calculated and compared with the actual observation value. All charts and marks are summarized to form a key performance snapshot including the highest temperature, lowest speed and average pressure. The performance snapshot is presented in the form of views and charts by integrating the highest temperature, lowest speed and average pressure indicators.

[0023] The steps to obtain trend and seasonality analysis results are: Based on the key performance snapshots, the data are arranged in chronological order to form a time series data set, and the change range and difference of each indicator in different time periods are analyzed to obtain preliminary trend and seasonal fluctuation data; According to the preliminary trend and seasonal fluctuation data, the trend seasonal composite fluctuation factor is calculated. The calculation formula is: ; in, represents the trend seasonal composite volatility factor, Representative The trend component value at a time point, Representative The seasonal component value at a time point, is the total number of time points, Representative The pressure change at each time point, Representative The speed change at a time point, Representative The temperature change at each time point; Based on the trend-seasonal composite fluctuation factor and combined with the time series data set, the trend and seasonal fluctuation pattern in the production data are determined, the rate of change of the production data in different time periods is judged, and the trend and seasonal analysis results are obtained.

[0024] Specifically, after the key performance snapshot has been obtained, the maximum temperature, minimum speed and average pressure data contained therein are first arranged time period by time period and marked with the corresponding time sequence, and then the specified time period is selected from the previously completed records of the temperature, speed and pressure data as a comparison benchmark. For example, the temperature, speed and pressure per hour in the first 72 hours are listed as an observation group, and the difference between adjacent observation groups is compared and the difference range is recorded. If the temperature of an observation group increases by more than 10°C compared with the previous group, the period is marked as a high fluctuation interval. This value of 10°C can be obtained based on the previous statistical conclusions of the equipment operation. For example, during the continuous operation of the equipment for 300 hours, the temperature fluctuation range is concentrated between 0°C and 10°C each time, and if the speed is reduced by more than 0.3m / s compared with the previous group, it can be regarded as a significant change in speed. The 0.3m / s threshold can also be calculated by adding twice the standard to the average speed difference. The pressure fluctuation is obtained by the difference method. If the pressure difference exceeds 0.2MPa compared with the previous group, it indicates that the pressure fluctuation is obvious. The upper limit of 0.2MPa is set by multiple pressure collection tests and recording the mean plus standard deviation of the difference distribution. All time periods marked as high or low fluctuations are summarized and statistically analyzed, and the change range and difference of the above indicators in different time periods are calculated. The indicators are grouped in units of hours or days and the gap between each group and the reference group is calculated to observe the changes in the morning and evening periods of the same day and the change range between different days. After comparing all observation group time periods with the reference group time periods, the amplitude change list of each indicator in different time periods can be obtained. By arranging the time series of the amplitude change list, the difference sequence can be further extracted and divided into rising or falling types according to the fluctuation direction. Finally, summarizing these sequence data can obtain preliminary trends and seasonal fluctuation data containing stage characteristics.

[0025] The formula is beneficial in that it combines the changes in temperature, speed, and pressure to measure the comprehensive fluctuation characteristics, and optimizes the overall assessment of production fluctuations by considering the difference between the trend component and the seasonal component at the same time; The parameter acquisition step is to fit the temperature value recorded at each time point in the time series data set to a long-term change curve in a predetermined manner to obtain the temperature trend component corresponding to the time point. Since the temperature varies in the range of 0°C to 200°C, it is necessary to read the temperature for 400 consecutive hours and calculate the hourly average value, and record the hourly average value as , and then use the polynomial fitting method to form the overall temperature trend function ,by and The minimum sum of squares of the difference is the convergence target, and finally for each time point The output is For example, if 400 mean temperature values ​​are collected within 400 hours, a cubic polynomial trend function is obtained after least squares fitting. , according to the obtained coefficient , , , , then at the 100th hour we can calculate , after calculation, we can get , which can represent the temperature trend component at the 100th hour; The steps for obtaining the parameters are as follows: for the temperature records in the same time series, the repeated fluctuation patterns in a specific period are selected as seasonal components, each 24 hours is divided into a period segment, the corresponding temperatures in the same period are averaged and merged to construct a daily periodic pattern. , do the minimum difference fitting between the periodic pattern and the original temperature record, when the difference is close to the peak interval under the statistical distribution curve, record the curve as the final seasonal component, thus we can get the first time point For example, if we analyze the temperature periodically for 28 days, we can find that the temperature will rise and fall significantly at the 5th and 18th hours of each day. Then the corresponding peak and valley characteristics can be retained in the daily cycle function. Assuming that the 100th hour falls near the (100mod24)=4th hour, then That is, the function value of the daily cycle function at the 4th hour, which can be obtained by the aforementioned minimum difference principle ; The steps to obtain the parameters are as follows: when considering the number of time points, directly count the number of time series involved in the evaluation. If each hour is used as a sampling point for 400 consecutive hours, we can get ; The parameter acquisition step is to calculate the absolute value of the pressure difference relative to the previous time point at each time point and record it as the pressure change amplitude. By continuously monitoring the pressure value range from 0MPa to 2MPa, the pressure readings of two adjacent points are recorded as and ,make For example, the pressure at the 100th hour is 1.36 MPa, and the pressure at the 99th hour is 1.30 MPa, then After all the time periods are calculated, the pressure change sequence can be obtained. ; The steps to obtain the parameters are to calculate the amplitude of the speed change in the time series and record the speed readings of two adjacent points as and ,by The speed change amplitude at that moment is obtained by the method. The speed is usually monitored in the range of 0m / s to 5m / s. For example, if the speed is 2.6m / s at the 100th hour and 2.4m / s at the 99th hour, then , after obtaining the complete sequence, all speed changes can be counted; The steps for obtaining parameters are as follows: the temperature change amplitude is obtained in a similar way to pressure or speed, and the temperature readings at adjacent time points are and Subtract and take the absolute value, and we get , the temperature range is between 0℃ and 200℃. By observing the difference between all adjacent time periods, a complete temperature change amplitude sequence can be obtained. For example, the temperature is 125.1℃ at the 100th hour and 123.4℃ at the 99th hour. , record the results in sequence and cover the entire time series; Calculation process: The first step is to calculate the numerator ,by And obtained from the previous steps and After that, you can first arrive all Sum and take the absolute value. In the second step, calculate the denominator. , the first part of which Sum the squares of all pressure changes and then take the square root. Part 2 Then combine the speed change amplitude with the temperature change amplitude and add them up item by item. In the third step, divide the numerator result by the denominator to get For example, the numerator is added up to get 520, and the first part of the denominator is Taking the square root of the sum of the squares gives about 42.6. The second part adds the speed and temperature changes to get about 147.4. The denominator is about 190.0 after adding them together. Finally, we can calculate ; The results show that after amplifying the difference between the trend and seasonal components and combining them with the comprehensive changes in pressure, velocity, and temperature, a numerical fluctuation factor can be obtained. A value higher than 2 indicates that production data has a more obvious multiple fluctuation pattern. A value lower than 1 indicates that the overall fluctuation is small; After the trend seasonal composite fluctuation factor is calculated, it is necessary to combine the previously arranged time series data sets to identify whether there are significant differences in trends and seasonal fluctuations in different time periods. First, the preliminary trend and seasonal fluctuation data obtained previously are compared again, and each cycle is regarded as an observation window and the fluctuation patterns of temperature, speed and pressure in the cycle are recorded. If the temperature gradually increases from 100°C to 120°C in a certain cycle and the speed changes multiple times in the range of 2.0m / s to 2.5m / s while the pressure is stable at around 1.2MPa, then this cycle is marked as the temperature fluctuation dominant cycle, and the actual observation window length, such as 24 hours or 48 hours, is used to distinguish each cycle. The maximum and average differences of temperature, speed and pressure in each cycle are numerically counted. If the calculation threshold is exceeded, the abnormality is noted in the subsequent annotation. The threshold can be set in combination with the historical data of equipment operation, for example, When the temperature fluctuation is greater than 20℃ within 24 hours, it is marked as excessive fluctuation. After comparing the temperature, speed and pressure data of all cycles, the number of cycles with fluctuations exceeding 30% of the total number of cycles is counted to determine whether there are concentrated fluctuations within a larger time range. For those cycles dominated by speed or pressure fluctuations, the same method is used to make corresponding records and include them in subsequent comparisons. If the temperature and pressure in some cycles are in a high difference range at the same time, they are classified into a composite fluctuation cycle sequence. Compared with other non-overlapping cycles, their coverage distribution over a longer period of time can be observed. After completing the above data sorting, the rate of change of production data in different time periods can be judged based on the differences between cycles. This judgment conclusion can be verified with the composite fluctuation factor results obtained previously, and finally a set of trend and seasonal analysis results containing period segmentation, indicator fluctuation range and rate distinction can be formed.

[0026] The steps to obtain quality control results are: Extract each performance indicator sequence in the key performance snapshot and calculate the control limit. The calculation formula is: ; in, is the control limit, is the selected performance indicator sequence, is the maximum value in the performance indicator sequence, is the minimum value in the performance indicator sequence, is the first in the performance indicator sequence data points, is the total number of performance indicators, is the average value of the performance index; Based on control limits, monitor quality fluctuations in real-time production data, mark data points that exceed control limits, and combine trend and seasonal analysis results to generate quality monitoring results.

[0027] Specifically, the formula is beneficial in that it takes into account the range of the sequence and the cumulative degree of deviation from the mean, so that the control limit not only depends on The maximum and minimum values ​​in can also reflect the deviation of each data point from the mean in the cube root term, which can make a more comprehensive measure of the overall distribution.

[0028] is the selected performance indicator sequence, which contains the same type of observation records. You can select any one of temperature, speed or pressure. The previous article has generated core indicator sequences such as maximum temperature, minimum speed, and average pressure. Each sequence can be regarded as In the specific operation, the performance indicators that need to be paid attention to are selected from the key performance snapshots, and then the values ​​of these indicators in the continuous time interval are arranged into a sequence. .

[0029] Representative sequence The maximum value in is the highest value detection of the monitored indicator in this interval. If the indicator is temperature, its value is usually in the range of 0℃ to 200℃. By traversing the collected temperature series, the highest temperature observation value is extracted as In order to obtain the maximum value, all Compare and then record the one with the highest value. For example, for a 168-hour temperature sequence, traverse Find the maximum temperature, such as 175.3℃, then .

[0030] Representation sequence The minimum value in is the lowest value detection of the monitoring indicator in this interval.

[0031] Is a sequence Middle data points, each Each corresponds to a specific observation time and measurement value, and it is necessary to ensure that the timestamp and the value correspond one to one. , we can locate the first The sampling moment and read the corresponding observation value, for example, in the pressure series, It may correspond to the pressure value of the 10th hour period of continuous recording. The process will be to collect data continuously on site, save all readings one by one, and then number them in chronological order at the calculation end, from 1 to For example, if the monitoring period is set to 48 hours and data is collected once every hour, , Corresponding to the observation value at the first hour, Corresponding to the second hour, and so on, 48 data points can be formed and arranged in sequence as follows .

[0032] The total number of performance indicator sequences is a clear count of the number of data points in the sequence, which is usually determined by the pre-set sampling frequency and monitoring duration. The monitoring personnel will specify the observation period before starting the collection. For example, 24 hours or 72 hours is selected as a monitoring cycle. Combined with the frequency of once an hour, it can be obtained. or The number of data points.

[0033] is the performance indicator sequence In order to obtain the average value of , we need to add up the values ​​of all data points in the sequence and divide by The range of values ​​is determined by the specific indicator type. The temperature is measured in degrees Celsius and fluctuates between 0°C and 200°C. The pressure can be between 0MPa and 2MPa. The speed can be between 0m / s and 5m / s. implement , taking the temperature series as an example, if it contains 168 hourly temperature records and each value is between 65℃ and 175.3℃, the total temperature sum is 23400.0℃ after statistical summation, then .

[0034] Calculation process: The first step is to calculate ,For example , , the difference is ; The second step is to calculate ,by And according to the above example Each Take the absolute value of the difference with 139.29, then sum it up and divide it by 168; for example, if the sum of the absolute values ​​is 5096.4, then ; Step 3: Calculate ; The fourth step is to add the results of the first two parts to get ; The results show that when When the value is 113.41, it can be used as the control limit of the temperature series. If the new temperature observation value monitored exceeds 113.41, it can be determined that the value is higher than the control limit derived from the previous fluctuation range. If the monitored temperature observation value is lower than this value, it means that it is still relatively stable within the current range. Combined with field experience, it can also be used in subsequent steps. Make further corrections.

[0035] Based on the control limits obtained previously , it is necessary to periodically compare the temperature, pressure or speed data from the real-time production process, and compare each record with In comparison, the temperature is greater than or pressure greater than or speed greater than Any records that are out of range shall be highlighted in the on-site monitoring interface. The value is taken from the calculation result of the previous paragraph and verified with the field experience. Each record is collected continuously on the production line to form a sequence with a fixed time interval, for example, it can be collected and written into the data table every hour or half an hour. The technician reads all the observation values ​​in the comparison period, and then checks them one by one. The difference between the positive and negative signs of the difference is distinguished and the abnormal points are additionally marked. For temperature monitoring scenarios, The temperature is too high for the record, the speed is too high for the speed monitoring scene, and the pressure is too high for the pressure monitoring scene. All the annotated data are listed together with the timestamp, device number and observation value to facilitate subsequent review. In addition, the trend and seasonal analysis results that have been obtained are also included in the reference to check whether there have been multiple occurrences of more than 10,000 times in a period of time in history. If the data exceeds the control limit in multiple adjacent observation cycles, it means that the degree of fluctuation in this time period is high. The technician will output the quality monitoring results after completing the comparison and marking and record the occurrence time and corresponding value of each abnormal point. The marked points and unmarked points together constitute the final quality monitoring results.

[0036] The steps for obtaining the optimized query path are: Based on the quality monitoring results, analyze the access patterns of production data queries, count the distribution characteristics of query requests, select data fields suitable for hash indexes or B-tree indexes, determine the index construction plan, and generate the index strategy design plan; According to the index strategy design plan, a hash index is built to process equal value queries, the hash function is determined and the key-value mapping relationship is calculated. At the same time, a B-tree index is established for range queries, index node splitting and adjustment operations are performed, and an optimized query path is generated.

[0037] Specifically, based on the quality monitoring results, we first need to collect all production data query requests in the past seven days and record their access time and usage frequency. For each query, we retrieve its reading mode for fields such as temperature, speed or pressure. If the query request only uses equal conditions such as "temperature = 120℃" and "pressure = 1.2MPa", it will be classified into the equal value query set. If the query request uses range conditions such as "speed>1.0m / s and speed<3.0m / s", it will be classified into the range query set. After all query requests are classified, they are sorted in chronological order. Calculate the proportion of equivalent queries and range queries. For example, if after comparing 10,000 query requests, it is found that about 5,000 are equivalent queries and most of them only lock the temperature or pressure field, another 3,000 are range queries and the more common ones are speed fields filtered between 0m / s and 5m / s, and the remaining 2,000 are distributed in a small number of mixed conditions or other types of conditions. Combine these classification results with the query execution time information to evaluate the access bottleneck. In order to refine the screening logic, you can set a threshold to distinguish whether the access volume is high enough, such as setting a field to participate in the equivalent query frequency of more than 2,000 within seven days. The access volume is high when the number of accesses is more than 2,000. This is the threshold determined by the statistical method obtained from the on-site detection log. More than 2,000 times is considered to have a greater impact on the retrieval efficiency. Similarly, for range queries, those with an access frequency of more than 1,000 times can be regarded as target fields that can be considered for indexing. Fields with low access volume or infrequent interval conditions are temporarily not indexed. Then, according to the statistical distribution characteristics of the access mode, the field name, query operation type and number of occurrences are further screened. The field combinations that meet the threshold are listed in a list and summarized into the candidate set of hash index or B-tree index. When the field is mainly used for equal matching, it is included in the hash index candidate list. When the field is mainly used for interval matching, it is included in the B-tree index candidate list, and its actual share is recorded, such as "the speed field range query frequency accounts for 40% of the total" and "the temperature field equal value query frequency accounts for 50% of the total". Finally, these candidate field combinations are referred to together with the quality control results to confirm whether it is necessary to accelerate the indexing of production data analysis. After sorting, the index construction plan can be obtained and placed in the plan description according to the hash index or B-tree index classification, and then the plan description is integrated into an index strategy design plan.

[0038] According to the index strategy design plan, it is necessary to prioritize the establishment of hash indexes for equal value queries and clarify the selection of hash functions. First, select the most commonly used fields from all field candidates. For example, temperature or pressure is found to have more than two thousand equal value queries based on the statistical frequency obtained previously, so the field is used as the hash index construction object. In order to generate a hash function, you can choose modulus or other hashing methods and combine them with the value range of the specific field. For example, when the temperature is between 0℃ and 200℃, you can use the temperature value to 201 modulo to map it to the hash bucket, or discretize the pressure in the range of 0MPa to 2MPa according to decimal precision, and then execute the hash function mapping to obtain the key value of each record. When a hash result conflicts, a corresponding zipper or open address strategy is required to allocate bucket space. After the hash index is completed, a B-tree index is established for range queries. Taking the speed field as an example, the speed will appear In order to search a wide range of data between 0m / s and 5m / s, the speed field is sorted in ascending order to construct a B-tree index node. Each node maintains a certain split threshold to control the number of child nodes. This split threshold can be set according to the established storage structure to accommodate a maximum of 50 records per node. When the node is loaded with more than 50 data, it will automatically split into two child nodes or adjust the level. The pointer after the split will point to each speed interval and the keyword will be reserved in the upper node for fast search. In the case of simultaneous query of speed and temperature, the leaf node can be refined according to the temperature distribution within the B-tree, but it must be ensured that the tree depth will not be excessively increased. For this purpose, the preset maximum level can be set to six layers and nodes can be merged or re-divided after exceeding six layers. Finally, according to the node structure after all hash indexes and B-tree indexes are established, an optimized query path is generated according to the query field classification for reference in subsequent retrieval.

[0039] The steps to obtain index optimization results are: Based on the optimized query path, the query execution plan is analyzed, multiple fields included in the query pattern are identified, the field combination with the highest query frequency is extracted, the field set suitable for the composite index is screened, and the composite index construction plan is generated; According to the composite index construction plan, the index fitness score is calculated using the following formula: ; in, Score the index fitness, For the The access frequency of the query fields, For the The proportion of fields in the data table, For the The number of index hits for fields. The total number of fields involved in index building; Based on the index fitness score, the hierarchical relationship of the composite index is adjusted, the order of the index keys is optimized, and the index optimization results are generated.

[0040] Specifically, based on the optimized query path, it is necessary to sort out the existing access statistics data for different types of query requirements and identify the temperature fields, speed fields, and pressure fields involved in the query pattern. First, analyze the query execution plan one by one from the previously integrated hash index and B-tree index statistical records, and extract the fields in the data table whose specific usage frequency exceeds the predetermined threshold. The threshold can be set based on standards such as the access frequency is greater than two thousand times or the value accounts for more than 30% within seven days. These standards are established by doing statistics based on the query logs collected by the team and finding that more than two thousand times belong to a high access volume. Then, in such high-access fields, further check whether the query involves the joint conditions of multiple fields. If it is detected that the temperature and speed are combined for more than one thousand joint searches, the combination of these two fields is recorded as a candidate combination. If the number of joint query accesses of pressure and temperature within a week is also higher than the standard, it is also included in the candidate combination. All candidate combinations are sorted in descending order according to the total access frequency and those combinations with less than five hundred times are excluded. Confirm that the remaining combinations are all distributed as multiple fields appearing at the same time. High-frequency conditions, for these remaining combinations, their field names are compared with the data table storage structure to check the column position and field length of the field in the data table. The existing index structure in the hash index or B-tree index is also checked to see if it can cover multi-field filtering. If it cannot meet the requirements, it is marked as a potential composite index target. The sorting or grouping requirements that often appear in the query mode are further analyzed to see if they are consistent with the field combination. The fields that meet the sorting or grouping rules, such as temperature, time, and speed, are paired as another candidate combination. Finally, a series of high-access frequency field combinations that have been screened out above are summarized, such as "temperature-speed", "pressure-temperature" or "temperature-speed-time", etc., and the proportion of each field in the data table is checked. For example, the temperature column accounts for 20% in the table, which means that its data volume may be relatively medium, and the speed column accounts for 12%, which means that it occupies less space. For the fields that need to be jointly searched, they are combined and packaged into a set of indexes to be composited, and it is estimated that how many query steps can be reduced for each combination to enhance the retrieval efficiency. Combining this information can generate a composite index construction plan.

[0041] The formula is beneficial because it integrates the processing of multiple aspects of information such as field access frequency, data table proportion, and index hit count. This method magnifies the difference between high hits and low hits in access frequency, and uses the denominator to Avoid unbalanced scoring caused by a single field accounting for too large a proportion, so as to obtain a more balanced reference metric in the decision-making of multi-field composite indexes.

[0042] Indicates the access frequency of the th query field, which is the count value of the number of times the field is queried and used within a certain period. Obtaining this value depends on scanning each entry in the database query log. For example, all SQL queries are recorded within 240 hours of device operation, and each occurrence of "WHERE temperature = xxx" or "WHERE pressure < xxx" is counted as one access. Then, the total access frequency is aggregated for each field, and this frequency is distributed to . When a field is accessed 3500 times, is set. In the case of multiple fields, separate statistics need to be done and presented in field order . A complete example is: when selecting three items, temperature, speed, and time, for a composite index, first count the access frequency of temperature as 2800 times, the access frequency of speed as 2200 times, and the access frequency of time as 1500 times. Then .

[0043] Is the proportion of the th field in the data table, indicating what storage proportion or data volume proportion the field occupies among all records in the entire table. This value needs to be obtained from the data storage structure level. For example, first count the actual number of bytes or rows of the entire table, and then perform a division operation on the number of bytes or rows occupied by the field to obtain a fraction between 0 and 1. If the temperature column occupies 15% of the storage in the entire table, then , if the speed column occupies 10%, then , and this kind of proportion can also be calculated by the number of rows or data chunks. Example: When the total size of the table is 100MB and the temperature column occupies 15MB, then .

[0044] Indicates the number of times the index of the th field has been hit, which is a number extracted from the database or system log, representing the number of times the index for this field has been successfully used in a query within a period of time. For example, the hit frequency can be counted through the hash index or B-tree index generated previously. If the temperature field has been successfully retrieved by the index 2000 times within a week, then can be set.

[0045] Indicates the total number of fields participating in index construction, which is determined according to the number of field combinations selected in the previous step. After data analysis, it is found that only three items, temperature, speed, and time, meet the joint requirements. Therefore .

[0046] Calculation process: Step 1: Calculate the numerator , for example, there are three groups of fields, and their , , , then for each group first calculate : ; Then the numerator is calculated in turn: ; Step 2: Calculate the denominator , first calculate each : ; Then multiply the three and take the fourth root: ; Step 3: Divide the numerator and denominator to get is 0.00395; The results show that when When the score is about 0.00395, the fitness score of the three-field composite index is relatively small, which means that this combination is not optimal in terms of access frequency, data table proportion, and hit count. If it is much larger than 0.00395, it is more worthy of priority. If the index exceeds 0.01, it can be considered as a high degree of fit, indicating that the field combination may be more suitable for building a composite index in terms of index efficiency. Technical personnel can compare the indexes of multiple combinations. Make a decision after the value.

[0047] Based on the index fitness score, we first conduct a hierarchical review of the currently established composite index structure, and check the number of fields and node distribution status contained in each layer of the index tree item by item. Fields with high access frequency and index hit times exceeding the specified threshold are placed in front of the upper nodes, and fields with relatively low access volume or a proportion that does not meet the screening criteria of the previous stage are placed in the back layer, so as to form a reconfiguration of the hierarchical nodes. Then we observe whether the excessive number of nodes causes the index level to be too deep. If the number of nodes in a certain layer is greater than the system experience value, such as fifty, the nodes are split and the field key value range covered by each node after the split is recorded. The experience value of fifty nodes is obtained by observing the past database retrieval performance. After each node is split, the number of pointers from the parent node to the child node must be checked. Whether the quantity matches the field order. If the original arrangement order of the index key is significantly different from the actual access mode, the order of the fields is adjusted. The number of query sorting or grouping times within seven days can be referred to to determine which field is more suitable to be ranked first. If the temperature field is used for sorting or grouping multiple times within seven days and the cumulative number exceeds two thousand times, the temperature field is prioritized in the arrangement of the index key. The speed or pressure fields that appear frequently are arranged in descending order according to the proportion. After the rearrangement is completed, the nodes are merged again to ensure that each node remains balanced in the index tree, not exceeding the maximum layer depth allowed by the system, such as six layers. Finally, all node pointers and field mapping relationships are integrated and the field order is marked to form an optimized multi-layer composite index description and record it as the index optimization result.

[0048] The steps to obtain the influencing factor analysis results are as follows: Based on the index optimization results, extract production parameter data, including temperature, pressure and speed, remove abnormal and duplicate data, align production parameters with quality results, and establish a production parameter matching data set; According to the production parameter matching data set, a decision tree model is constructed to analyze the quality impact of production parameters in different intervals, generate a classification structure based on the decision path, calculate the split contribution value of each production parameter in the classification structure, extract production variables, and obtain the classification parameter impact weight data; Based on the classification parameter impact weight data, a random forest model is constructed, and multiple rounds of feature selection are performed to calculate the stability and classification contribution of production parameters under different decision tree combinations, analyze the interaction between different parameters, rank the importance of production parameters, and generate influencing factor analysis results.

[0049] Specifically, based on the index optimization results, it is necessary to select temperature, pressure and speed as the focus of historical production data. First, all data previously collected by the equipment are sorted in ascending order by timestamp. For multiple duplicate records appearing at the same time, the first one is retained and subsequent duplicate data are removed. If abnormal observations are found, they are checked according to the pre-set threshold range. For example, the temperature is compared with the range of 0℃ to 200℃. If it exceeds 200℃, it is directly removed. The pressure is compared with the range of 0MPa to 2MPa. If the value is less than 0MPa or greater than 2MPa, it is marked as abnormal and excluded. If a negative value or more than 5 is detected between 0m / s and 5m / s, the speed is detected. The value of m / s is also eliminated. The above thresholds are obtained based on statistics after multiple inspections of the equipment's safe working range. After all the cleaning is completed, the product qualification rate or failure rate field in the same time period is read from the quality result data and matched with the temperature, pressure, and speed according to the timestamp. If the quality result record in the same time period is missing, the most recent quality result is retrieved forward or backward and temporarily bound to the temperature, pressure, and speed of the time period. After the alignment of the records of all time periods is completed, each one is confirmed to be correct, and the data at each moment is aggregated into a row of entries containing four items of temperature, pressure, speed and quality results, and all entries are merged into a production parameter matching data set.

[0050] According to the production parameter matching data set, before training the decision tree model, the values ​​of the three parameters of temperature, pressure and speed are divided into fixed intervals to establish interval mapping. For example, the temperature is divided into twenty intervals from 0℃ to 200℃ with a ten-degree interval segment, the pressure is divided into ten intervals of 0.2MPa from 0MPa to 2MPa, and the speed is divided into five intervals of 1m / s from 0m / s to 5m / s. Correspondingly, the quality result field is divided into two values ​​according to qualified or unqualified. When the records in the data set enter the decision tree construction process, the temperature interval is read first and split into two branches according to the binary quality results, and then the next layer of splitting is performed according to the pressure or speed interval. If the temperature interval in a split node is statistically included The branches with more qualified records are marked as qualified main branches and the corresponding split contribution values ​​are calculated. The multi-layer decision paths are recursively established downwards and the temperature, pressure and speed in each node are calculated. The intervals and the ratio of qualified or unqualified records in this interval are counted. Finally, after the decision tree converges, one or more complete classification paths are formed to judge the correlation between temperature and pressure or speed in a certain interval, and the sum of the split contribution values ​​of all nodes is calculated to measure the overall influence of the three parameters on the classification results. On this basis, the parameters of each high-contribution split node are extracted to form a list of production variables for this segment of data and summarize the comprehensive influence weight of each parameter, thereby obtaining the classification parameter influence weight data.

[0051] Based on the classification parameter influence weight data, temperature, pressure and speed are first selected as the main features and several decision trees of different depths are configured in the random forest model. Each decision tree performs node splitting according to the interval partitioning method obtained in the previous step. During training, 80% of the production parameter matching data set is randomly selected to generate splitting rules and 20% is reserved for verification. If the depth of a decision tree exceeds the established limit, such as ten layers, the nodes of the tree are pruned to avoid overfitting. All trained decision trees are run multiple times according to the same process to accumulate the indicators of multiple trees, and the temperature used by each tree in classification is used as the threshold. After all the split node counts of temperature, pressure and speed are added up, their contribution to the overall classification effect is calculated, and it is observed whether the fluctuation range of these parameters between different trees exceeds the pre-given stability threshold, such as 5%. If the node contribution value of temperature remains around 30% in all ten trees, it is regarded as a stable parameter, otherwise it is regarded as a parameter with large fluctuations. The corresponding stability value is also given for the pressure or speed analyzed in the same way. Finally, the importance ranking of each parameter is calculated based on its contribution ratio and fluctuation in all trees, so as to obtain a list of the most influential production parameters and use this ranking result as the result of the influencing factor analysis.

Claims

1. A deep concentrated production scheduling big data analysis system, characterized in that: The system comprises: A data acquisition module collects and condenses real-time production data, including temperature, speed and pressure information, to generate a real-time production data set; extracts key performance indicators based on the real-time production data set to generate a key performance snapshot; The data processing module performs time series analysis based on the key performance snapshots, identifies trends and seasonal fluctuations in the production data, and generates trend and seasonal analysis results; monitors the quality fluctuations of the key performance snapshots, and generates quality monitoring results in combination with the trend and seasonal analysis results; The data index module designs and implements hash and B-tree index strategies based on the quality monitoring results, optimizes data query efficiency, and generates optimized query paths; executes composite index construction according to the optimized query paths, optimizes multi-condition queries, and generates index optimization results; The decision support module uses decision trees and random forest models to analyze the impact of production parameters on product quality based on the index optimization results and generates influencing factor analysis results.

2. The deep concentrated production scheduling big data analysis system according to claim 1 is characterized in that: The steps for obtaining the real-time production data set are: Monitor real-time data of temperature, speed and pressure on the concentration production line, time-stamp and format each data point collected to form raw real-time monitoring data; Based on the original real-time monitoring data, error data and outliers are removed, missing data points are filled, and statistical methods are used to evaluate the consistency of the data to generate real-time production data after quality inspection; Based on the real-time production data after the quality inspection, data synchronization processing is performed to align data from different sensor sources according to a unified timestamp to form a real-time production data set.

3. The deep concentrated production scheduling big data analysis system according to claim 1 is characterized in that: The steps for obtaining the key performance snapshot are: Filter key performance indicators from the real-time production data set, including maximum temperature, minimum speed and average pressure, to generate preliminary filtered performance indicator data; Based on the preliminary screened performance indicator data, a key performance snapshot is constructed, and the performance snapshot is presented in the form of views and charts by integrating the maximum temperature, minimum speed and average pressure indicators.

4. The deep concentrated production scheduling big data analysis system according to claim 1 is characterized in that: The steps for obtaining the trend and seasonality analysis results are as follows: Based on the key performance snapshots, the data are arranged in chronological order to form a time series data set, and the change range and difference of each indicator in different time periods are analyzed to obtain preliminary trend and seasonal fluctuation data; Calculate the trend-seasonal composite fluctuation factor based on the preliminary trend and seasonal fluctuation data; Based on the trend seasonal composite fluctuation factor, combined with the time series data set, the trend and seasonal fluctuation pattern in the production data are determined, the change rate of the production data in different time periods is judged, and the trend and seasonal analysis results are obtained.

5. The deep concentrated production scheduling big data analysis system according to claim 1 is characterized in that: The steps for obtaining the quality monitoring results are: Extracting each performance indicator sequence in the key performance snapshot and calculating the control limit; Based on the control limits, quality fluctuations in real-time production data are monitored, data points exceeding the control limits are marked, and quality monitoring results are generated by combining the trend and seasonal analysis results.

6. The deep concentrated production scheduling big data analysis system according to claim 1 is characterized in that: The steps for obtaining the optimized query path are: Based on the quality monitoring results, analyze the access mode of production data query, count the distribution characteristics of query requests, screen data fields suitable for hash index or B-tree index, determine the index construction plan, and generate the index strategy design plan; According to the index strategy design scheme, a hash index is constructed for processing equal value queries, a hash function is determined and a key-value mapping relationship is calculated, and a B-tree index is established for range queries, index node splitting and adjustment operations are performed, and an optimized query path is generated.

7. The deep concentrated production scheduling big data analysis system according to claim 1 is characterized in that: The steps for obtaining the index optimization result are: Based on the optimized query path, the query execution plan is analyzed, multiple fields included in the query pattern are identified, the field combination with the highest query frequency is extracted, the field set suitable for the composite index is screened, and a composite index construction plan is generated; According to the composite index construction scheme, calculating the index fitness score; Based on the index fitness score, the hierarchical relationship of the composite index is adjusted, the arrangement order of the index keys is optimized, and an index optimization result is generated.

8. The deep concentrated production scheduling big data analysis system according to claim 1 is characterized in that: The steps for obtaining the influencing factor analysis results are: Based on the index optimization results, extract production parameter data, including temperature, pressure and speed, remove abnormal data and duplicate data, align production parameters with quality results, and establish a production parameter matching data set; According to the production parameter matching data set, a decision tree model is constructed to analyze the quality impact of production parameters in different intervals, generate a classification structure based on the decision path, calculate the split contribution value of each production parameter in the classification structure, extract production variables, and obtain classification parameter impact weight data; Based on the classification parameter influence weight data, a random forest model is constructed, multiple rounds of feature selection are performed, the stability and classification contribution of production parameters under different decision tree combinations are calculated, the interaction between different parameters is analyzed, the importance of production parameters is ranked, and the influencing factor analysis results are generated.

9. The deep concentration production scheduling big data analysis method of the deep concentration production scheduling big data analysis system according to any one of claims 1 to 8 is characterized in that: The following steps are involved: Collect and condense real-time production data, including temperature, speed and pressure, to generate real-time production data sets. Based on the real-time production data sets, extract temperature extremes, speed change rates, and pressure stability indicators to obtain key performance snapshots. Based on key performance snapshots, identify the periodicity of temperature extremes, speed change trends, and pressure fluctuations, and obtain trend and seasonal analysis results; Conduct quality fluctuation monitoring on trend and seasonal analysis results, and obtain quality monitoring results by combining quality monitoring values ​​and fluctuation frequencies; Based on the quality monitoring results, hash and B-tree index strategies are designed and implemented to optimize data query efficiency and obtain index optimization results; Based on the index optimization results, analyze the impact of temperature, speed, and pressure on product quality, generate an influencing factor analysis, and obtain the influencing factor analysis results.

Citation Information

Patent Citations

  • Verifiable block chain indexing method supporting Boolean query and range query

    CN117131049A

  • Distributed market data acquisition management system

    CN117876016A

  • Intelligent production line testable digital twin modeling method

    CN117952009A

  • Circuit board quality abnormity prediction system based on machine intelligent learning

    CN119129365A

  • Food quality evaluation method and system based on big data

    CN119204850A

Cited By

  • Method, system and device for analyzing annotation data

    CN120597007A

  • Agricultural planting big data analysis task scheduling system

    CN121599431A