Industrial big data mining system and method based on artificial intelligence
Through an industrial big data mining system based on artificial intelligence, the problem of data real-time and fault prediction deviation in semiconductor production is solved, real-time monitoring and prediction of faults is realized, and production control efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510530495.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art cannot effectively deal with the problem of high real-time requirements of massive data in semiconductor production and prone to deviation in fault prediction, resulting in low production control efficiency.
The industrial big data mining system based on artificial intelligence is adopted, including data collection, storage, dimensionality reduction processing, artificial intelligence monitoring, early warning and deviation mining modules, and real-time fault monitoring, prediction and correction are carried out through principal component analysis and machine learning models to ensure the accuracy and timeliness of data analysis.
Real-time monitoring and prediction of faults in semiconductor production process is realized, the efficiency and accuracy of production control is improved, wrong predictions caused by deviation are avoided, and the demand for real-time and data integrity of semiconductor production is met.
Smart Images

Figure CN120448797A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data mining technology, and in particular to an industrial big data mining system and method based on artificial intelligence. Background Art
[0002] With the continuous evolution of semiconductor manufacturing technology and the continued expansion of industrial production, the manufacturing process generates massive amounts of data. This data comes from a wide range of sources, including production equipment operating parameters, process conditions, and product quality test results. Data mining of industrial big data is extremely necessary in semiconductor manufacturing. Due to the complexity and high-precision requirements of the manufacturing process, traditional data analysis methods struggle to cope. Data mining technology, however, can deeply uncover the patterns underlying this data. By analyzing large amounts of historical and real-time data, it can uncover subtle connections between different process steps, predict equipment failures, and provide early warning of potential product quality issues, thereby optimizing the production process and improving quality control.
[0003] Chinese Patent Publication No.: CN112181956B discloses a method for data mining based on hot rolling industry big data, which belongs to the technical field of hot rolling production line product quality determination. The technical solution of the present invention is: the secondary system data, tertiary MES data, and continuous value data parsed from DCA files of each hot rolling production line are collected from different databases into the postgreSQL database integrated by the big data platform mpp; the hot rolled steel coil data integrated in the big data platform are statistically analyzed to guide the production of hot rolled steel coils. However, this solution is not suitable for the production guidance process of semiconductors, and cannot adapt to the production characteristics of semiconductors to intelligently mine industrial big data in the semiconductor production process, and predict the occurrence of equipment failures based on the results of intelligent mining analysis, and give early warnings of possible product quality problems. Combined with the characteristics of the semiconductor production process that have high real-time requirements, fault monitoring is carried out in real time to achieve optimization of the semiconductor production process and improvement of quality control. Summary of the Invention
[0004] To this end, the present invention provides an industrial big data mining system and method based on artificial intelligence to overcome the problems in the prior art that the industrial big data of the semiconductor production process is large in quantity and diverse in types, and the semiconductor industrial production has high real-time requirements for fault monitoring and prediction, but fault prediction may have prediction deviations, resulting in low control efficiency of semiconductor industrial production.
[0005] To achieve the above objectives, the present invention provides an industrial big data mining system based on artificial intelligence, comprising:
[0006] A data acquisition module for collecting semiconductor production data;
[0007] A data storage module, used for storing the semiconductor production data;
[0008] A production database for storing the semiconductor production data;
[0009] A dimensionality reduction processing module, configured to perform dimensionality reduction processing on the semiconductor production data to obtain dimensionality-reduced production data;
[0010] An artificial intelligence monitoring module is used to perform real-time fault monitoring, fault prediction analysis, and correlation analysis based on the reduced-dimensional production data to obtain the fault condition, fault time point, and fault cause respectively;
[0011] A warning and alarm module is used to provide fault warning and fault alarm according to the fault condition, the fault time point and the fault cause;
[0012] The deviation mining module is used to analyze the deviation coefficient according to the deviation value and deviation error, and to judge the deviation situation according to the deviation coefficient. It is also used to perform deviation correction on the dimensionality reduction process and the fault prediction analysis process according to the deviation situation.
[0013] Furthermore, the dimensionality reduction processing module includes:
[0014] A data acquisition unit, configured to acquire semiconductor production data from a production database, and also configured to acquire learning data and extract learning data from a dimensionality reduction processing unit;
[0015] a dimensionality reduction method selection unit, configured to select a dimensionality reduction method according to the semiconductor production data;
[0016] A model building unit, configured to build a current dimensionality reduction model according to the dimensionality reduction method, and to train the current dimensionality reduction model according to the learning data;
[0017] a preprocessing unit, configured to preprocess the semiconductor production data to obtain preprocessed semiconductor production data;
[0018] A dimensionality reduction processing unit is used to perform dimensionality reduction processing on the pre-processed semiconductor production data according to the current dimensionality reduction model and output the dimensionality reduction data.
[0019] Furthermore, the dimensionality reduction method selection unit analyzes the semiconductor production data through the principal component analysis method to obtain the cumulative variance contribution rate G, and compares the cumulative variance contribution rate G with the preset cumulative variance contribution rate G0, judges the linearity of the semiconductor production data based on the comparison result, and selects the dimensionality reduction method of the semiconductor production data based on the judgment result.
[0020] Furthermore, the preprocessing unit preprocesses the semiconductor production data to obtain preprocessed semiconductor production data, the dimensionality reduction processing unit inputs the preprocessed semiconductor production data in the preprocessing unit into the dimensionality reduction model, and the dimensionality reduction model outputs the reduced dimensionality data, and the output reduced dimensionality data is passed to the artificial intelligence monitoring module.
[0021] Furthermore, the artificial intelligence monitoring module includes:
[0022] A real-time fault monitoring unit, configured to monitor the fault condition in real time based on the dimensionality reduction data;
[0023] A fault prediction analysis unit, configured to predict a fault time point based on the dimensionality reduction data;
[0024] A correlation analysis unit is used to analyze the cause of the fault based on the dimension reduction data.
[0025] Furthermore, the correlation analysis unit inputs the dimension-reduced data into a fault cause inference model, outputs the fault cause corresponding to the dimension-reduced data, and feeds the fault cause back to the early warning alarm module.
[0026] Furthermore, the early warning alarm module includes:
[0027] A fault warning unit, configured to provide a fault warning based on the fault time point and the fault cause;
[0028] The fault alarm unit is used to issue a fault alarm according to the fault condition and the fault cause.
[0029] Furthermore, the deviation mining module includes:
[0030] A mining analysis unit for analyzing the deviation coefficient based on the deviation value and the deviation error;
[0031] a mining judgment unit, configured to judge the deviation situation according to the deviation coefficient;
[0032] A deviation correction unit, configured to perform deviation correction on the dimensionality reduction process and the fault prediction analysis process according to the deviation;
[0033] a deviation optimization unit, configured to perform a real-time invalidity analysis on the semiconductor production data to obtain a real-time invalidity analysis result, optimize an analysis process of a deviation coefficient based on the real-time invalidity analysis result, and optimize storage of a production database based on the real-time invalidity analysis result;
[0034] The anti-mistake optimization unit is used to correct the optimization storage process of the production database according to the learning data, and is also used to adjust the deviation optimization process according to the update status of the learning data.
[0035] Furthermore, the deviation optimization unit calculates the effective coefficient EC based on the data missing rate DR, the data fluctuation coefficient FC, the data anomaly rate AR and the data consistency ratio CP, compares the effective coefficient EC with the preset effective coefficient EC0, judges the data validity based on the comparison result, and optimizes the analysis process of the deviation coefficient based on the judgment result.
[0036] On the other hand, the present invention also provides an industrial big data mining method based on artificial intelligence, comprising:
[0037] Step S1, collecting semiconductor production data through a data acquisition module;
[0038] Step S2, storing the semiconductor production data through a data storage module;
[0039] Step S3, performing dimensionality reduction processing on the semiconductor production data by a dimensionality reduction processing module to obtain dimensionality-reduced production data;
[0040] Step S4: Using an artificial intelligence monitoring module to perform real-time fault monitoring, fault prediction analysis, and correlation analysis based on the reduced-dimensional production data, the fault condition, fault time point, and fault cause are obtained.
[0041] Step S5: performing a fault warning and a fault alarm according to the fault condition, the fault time point and the fault cause through the early warning alarm module;
[0042] Step S6, analyzing the deviation coefficient according to the deviation value and the deviation error by the deviation mining module, and judging the deviation situation according to the deviation coefficient;
[0043] Step S7: performing deviation correction on the dimensionality reduction process and the fault prediction analysis process according to the deviation situation through the deviation mining module.
[0044] Compared with the prior art, the beneficial effects of the present invention are that the system realizes real-time collection of semiconductor production data through the data acquisition module, providing a reliable data basis for subsequent analysis, the system securely stores the collected semiconductor production data through the data storage module, ensuring the integrity and availability of the data, the system centrally manages and stores semiconductor production data through the production database, facilitating subsequent data retrieval and analysis, the system simplifies the data structure and reduces the data dimension through the dimensionality reduction processing module, making subsequent analysis more efficient and improving the speed and accuracy of data processing, the system uses the artificial intelligence monitoring module to use the reduced-dimensional production data to perform real-time monitoring, predictive analysis and correlation analysis of faults, and can promptly discover potential fault conditions, identify the time point of fault occurrence and analyze the cause of the fault, thereby improving the timeliness of fault handling, the system generates fault warning and alarm information according to the monitored fault conditions, time points and causes through the early warning alarm module, helping managers to take quick measures to reduce the impact of faults on production, and the system analyzes the deviation value and deviation error through the deviation mining module to determine the deviation in the production process and make corresponding corrections to ensure the accuracy of data analysis and avoid erroneous predictions caused by deviations. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a schematic diagram of the structure of the industrial big data mining system based on artificial intelligence in this embodiment;
[0046] Figure 2 Schematic diagram of the structure of the dimensionality reduction processing module of this embodiment;
[0047] Figure 3 This is a schematic diagram of the structure of the artificial intelligence monitoring module of this embodiment;
[0048] Figure 4 This is a structural diagram of the early warning alarm module of this embodiment;
[0049] Figure 5 This is a schematic diagram of the structure of the deviation mining module of this embodiment;
[0050] Figure 6 This is a flow chart of the industrial big data mining method based on artificial intelligence in this embodiment. DETAILED DESCRIPTION
[0051] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0052] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0053] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0054] See also Figure 1 As shown in FIG, which is a schematic diagram of the structure of the industrial big data mining system based on artificial intelligence in this embodiment, the system includes:
[0055] A data acquisition module for collecting semiconductor production data;
[0056] A data storage module, used for storing the semiconductor production data, the data storage module being connected to the data acquisition module;
[0057] A production database, used to store the semiconductor production data, the production database being connected to the data storage module;
[0058] A dimensionality reduction processing module is used to perform dimensionality reduction processing on the semiconductor production data to obtain reduced-dimensionality production data, and the dimensionality reduction processing module is connected to the production database;
[0059] An artificial intelligence monitoring module is used to perform real-time fault monitoring, fault prediction analysis, and correlation analysis based on the reduced-dimensionality production data, and obtain the fault situation, fault time point, and fault cause respectively. The artificial intelligence monitoring module is connected to the dimensionality reduction processing module;
[0060] A warning and alarm module is used to provide fault warning and fault alarm according to the fault condition, the fault time point and the fault cause, and the warning and alarm module is connected to the artificial intelligence monitoring module;
[0061] The deviation mining module is used to analyze the deviation coefficient based on the deviation value and the deviation error, and to judge the deviation situation based on the deviation coefficient. It is also used to perform deviation correction on the dimensionality reduction processing process and the fault prediction analysis process based on the deviation situation. The deviation mining module is connected to the data storage module, the dimensionality reduction processing module and the artificial intelligence monitoring module.
[0062] Specifically, the system is set in the semiconductor production equipment management terminal, and performs data mining on the industrial big data of the semiconductor production process to monitor and predict faults in the semiconductor industrial production process, meet the real-time requirements of semiconductor production monitoring, avoid fault prediction deviations and make corrections in time, and improve the efficiency of semiconductor industrial production control. The system realizes real-time collection of semiconductor production data through the data acquisition module, providing a reliable data basis for subsequent analysis. The system safely stores the collected semiconductor production data through the data storage module to ensure the integrity and availability of the data. The system centrally manages and stores semiconductor production data through the production database to facilitate subsequent data retrieval and analysis. The system uses the dimensionality reduction processing module to perform dimensionality reduction processing. Simplifying the data structure and reducing the data dimension make subsequent analysis more efficient and improve the speed and accuracy of data processing. The system uses the artificial intelligence monitoring module to use the reduced-dimensional production data to perform real-time monitoring, predictive analysis and correlation analysis of faults. It can promptly discover potential fault conditions, identify the time point when the fault occurs and analyze the cause of the fault, thereby improving the timeliness of fault handling. The system uses the early warning alarm module to generate fault warning and alarm information based on the monitored fault conditions, time points and causes, helping managers to take quick measures to reduce the impact of faults on production. The system uses the deviation mining module to analyze the deviation value and deviation error, judge the deviation in the production process, and make corresponding corrections to ensure the accuracy of data analysis and avoid erroneous predictions due to deviations.
[0063] Specifically, the data acquisition module collects semiconductor production data, which includes equipment operation status data, process data, product quality inspection data, environmental data and personnel operation data. The equipment operation status data includes the temperature, pressure, humidity, vibration amplitude, speed, current and voltage of the equipment. The process data includes the process parameters in each production process step, such as the etching gas flow, etching time, plasma density in the etching process, the deposition rate in the deposition process, the purity and ratio of the deposition material, the temperature curve and doping concentration in the diffusion process, etc. The product quality inspection data includes the electrical performance indicators of the chip, such as resistance, capacitance, inductance, etc. , the threshold voltage, on-resistance, breakdown voltage and physical property data of the transistor, such as the thickness, flatness, surface roughness, width and spacing of the metal wiring of the chip, the environmental data include the cleanliness, temperature and humidity, and static electricity level of the production workshop, the personnel operation data include recording the operating steps, operation time, operation sequence and adjustment information of the equipment parameters of the operator in the production process, this embodiment does not limit the collection method of the equipment operation status data, and those skilled in the art can freely set it according to actual conditions, as long as the collection requirements of the equipment operation status data are met, such as it can be set at the key heating parts of the equipment, such as motors, heating furnaces, and high-power electronic components. The temperature sensor measures the temperature change of the light source during operation in real time and converts the temperature signal into an electrical signal for output. This embodiment does not limit the method for collecting the process data. Those skilled in the art can freely configure it according to actual conditions. For example, the process data can be collected through the data recording function of the process equipment. Modern etching equipment usually has a built-in data recording system that directly reads parameters such as etching gas flow, etching time, and plasma density through an Ethernet interface. This embodiment does not limit the method for collecting the product quality inspection data. Those skilled in the art can freely configure it according to actual conditions. For example, automatic testing equipment can be used to test the electrical properties of the chip, such as parameters such as resistance, capacitance, inductance, and transistor threshold voltage. This embodiment does not limit the method for collecting the environmental data. Those skilled in the art can freely configure it according to actual conditions. For example, dust particle counters can be installed in key areas of the production workshop, such as the photolithography workshop and the packaging workshop, to monitor the number and size distribution of dust particles in the air. This embodiment does not limit the method for collecting the personnel operation data. Those skilled in the art can freely configure it according to actual conditions. For example, an operation terminal can be installed next to each production equipment. The operator logs into the system through the operation terminal to operate the equipment. The operation terminal can record the operator's login information (such as name, work number), operation time, operation steps, etc.
[0064] Specifically, the data storage module adopts a distributed storage architecture to store the semiconductor production data in a production database.
[0065] Specifically, this embodiment will disperse and store the semiconductor production data on each node in the production database based on the Hadoop distributed file system.
[0066] Specifically, the data storage module stores data in a dispersed manner on multiple nodes through a distributed storage architecture to easily cope with the growth of data volume. The storage capacity can be expanded by adding nodes to meet the increasing data storage needs in the semiconductor production process.
[0067] Specifically, the production database refers to a database with a distributed storage architecture for storing the semiconductor production data.
[0068] See also Figure 2 As shown in FIG, it is a schematic diagram of the structure of the dimensionality reduction processing module of this embodiment, and the dimensionality reduction processing module includes:
[0069] A data acquisition unit, configured to acquire semiconductor production data from a production database, and also configured to acquire learning data and extract learning data from a dimensionality reduction processing unit;
[0070] a dimensionality reduction method selection unit, configured to select a dimensionality reduction method according to the semiconductor production data;
[0071] A model construction unit, configured to construct a current dimensionality reduction model according to the dimensionality reduction method, and to perform learning and training on the current dimensionality reduction model according to the learning data, the model construction unit being connected to the data acquisition unit and the dimensionality reduction method selection unit;
[0072] A preprocessing unit, configured to preprocess the semiconductor production data to obtain preprocessed semiconductor production data, the preprocessing unit being connected to the data acquisition unit and the model building unit;
[0073] The dimensionality reduction processing unit is used to perform dimensionality reduction processing on the pre-processed semiconductor production data according to the current dimensionality reduction model and output the dimensionality reduction data. The dimensionality reduction processing unit is connected to the pre-processing unit.
[0074] Specifically, the data acquisition unit acquires semiconductor production data from a production database in real time.
[0075] It is understandable that this embodiment does not limit the method for real-time acquisition of the semiconductor production data. Those skilled in the art can freely set it according to actual conditions, and only need to meet the demand for real-time acquisition of the semiconductor production data. For example, data synchronization tools, including DataX and Sqoop, can be set to obtain the semiconductor production data in real time.
[0076] Specifically, the data acquisition unit obtains the learning data input by the user through the user interaction input window.
[0077] Specifically, the user interaction input window refers to the window interface through which the user inputs and interacts with the user.
[0078] Specifically, the data acquisition unit compares the extraction time interval tc with the preset extraction time interval tc0, and judges the extraction situation of the learning data according to the comparison result, where:
[0079] When tc < tc0, the data acquisition unit determines not to extract the learning data;
[0080] When tc ≥ tc0, the data acquisition unit determines to extract the learning data and extracts the learning data from the dimensionality reduction processing unit.
[0081] Specifically, the extraction time interval refers to the time interval elapsed from the last extraction of learning data to the current time, and the preset extraction time interval refers to the preset value of the time interval for which the learning data should be extracted, such as 20 hours. In this embodiment, the extraction method of the learning data is not limited, and those skilled in the art can freely set it according to the actual situation, as long as the interval extraction and update requirements of the learning data are met. For example, it can be set to randomly extract 3% of the dimensionality reduction data output by the dimensionality reduction processing unit within the extraction time interval as the learning data. The learning data includes the dimensionality reduction data and the semiconductor production data corresponding to the dimensionality reduction data.
[0082] Specifically, the dimensionality reduction method selection unit analyzes the semiconductor production data through the principal component analysis method to obtain the cumulative variance contribution rate G, and compares the cumulative variance contribution rate G with the preset cumulative variance contribution rate G0, G0 = 85%. According to the comparison result, the linear situation of the semiconductor production data is judged, and the dimensionality reduction method of the semiconductor production data is selected according to the judgment result, where:
[0083] When G ≥ G0, the dimensionality reduction method selection unit determines that the linear situation of the semiconductor production data is a linear structure, and selects the linear dimensionality reduction method as the dimensionality reduction method of the semiconductor production data;
[0084] When G < G0, the dimensionality reduction method selection unit determines that the linear situation of the semiconductor production data is a non-linear structure, and selects the non-linear dimensionality reduction method as the dimensionality reduction method of the semiconductor production data.
[0085] Specifically, the principal component analysis method refers to a commonly used multivariate statistical analysis method, which is mainly used for data dimensionality reduction, information extraction and visualization. This embodiment does not limit the specific implementation method of the principal component analysis method. Those skilled in the art can set it according to their own needs. It only needs to meet the calculation requirements of the cumulative contribution rate. The cumulative variance contribution rate refers to the cumulative variance contribution rate, which is the result of adding the variance contribution rates of the first K principal components, K = 1, 2, 3...n, n is the number of principal components, the variance contribution rate refers to the proportion of the variance of a principal component to the total variance, and the preset cumulative variance contribution rate refers to the preset value used to judge the linearity of semiconductor production data, and the linearity of semiconductor production data refers to The related data in the semiconductor production data show the characteristics of a linear relationship. The linear situation of the semiconductor production data includes the linear situation of the semiconductor production data as a linear structure and the linear situation of the semiconductor production data as a nonlinear structure. The linear dimensionality reduction method refers to a technology that maps high-dimensional data to a low-dimensional space through linear transformation, aiming to retain the key features and structure of the data while reducing the dimension of the data to solve the high computational complexity and visualization difficulties brought by high-dimensional data, such as linear discriminant analysis and multidimensional scaling analysis. The nonlinear dimensionality reduction method refers to a method used to process data with a nonlinear structure, mapping high-dimensional data to a low-dimensional space through nonlinear transformation to better reveal the intrinsic structure and characteristics of the data, such as isometric mapping and local linear embedding.
[0086] Specifically, the model building unit obtains the selected dimensionality reduction method from the dimensionality reduction method selection unit, initializes the structure of the dimensionality reduction model according to the selected dimensionality reduction method, and trains the dimensionality reduction model according to the learning data in the data acquisition unit, wherein:
[0087] When the dimension reduction method selection unit selects a linear dimension reduction method as the dimension reduction method for the semiconductor production data, the model construction unit selects the linear dimension reduction method as the structure of the dimension reduction model;
[0088] When the dimension reduction method selection unit selects a nonlinear dimension reduction method as the dimension reduction method for the semiconductor production data, the model construction unit selects the nonlinear dimension reduction method as the structure of the dimension reduction model;
[0089] The model construction unit constructs the dimensionality reduction model by using a dimensionality reduction model construction method, and the dimensionality reduction model construction method includes:
[0090] The learning data is divided into a 70% dimensionality reduction training set, a 20% dimensionality reduction validation set, and a 10% dimensionality reduction test set. The dimensionality reduction training set is input into the decision tree model to train the decision tree model, and the dimensionality reduction validation set is input into the trained decision tree model. The hyperparameters of the trained decision tree model are iteratively optimized, and the dimensionality reduction test set is input into the iteratively optimized decision tree model to perform dimensionality reduction test on the iteratively optimized decision tree model to obtain the dimensionality reduction test result. The total number of samples in the dimensionality reduction test set is set to f0, the number of correct dimensionality reduction test samples is set to f, and the dimensionality reduction test accuracy is F, F=f / f0. The dimensionality reduction test accuracy F is compared with the preset dimensionality reduction test accuracy F0. According to the comparison result, the training compliance of the iteratively optimized decision tree model is judged, and the judgment result is output, where:
[0091] When F≥F0, the model building unit determines that the iteratively optimized decision tree model training meets the standards, outputs the iteratively optimized decision tree model as a dimensionality reduction model, and provides the dimensionality reduction model to the dimensionality reduction processing unit;
[0092] When F<F0, the model construction unit determines that the iteratively optimized decision tree model training does not meet the standards, updates the learning data to obtain updated learning data, and trains the decision tree model, iteratively optimizes the hyperparameters, and analyzes and tests the decision tree model according to the updated learning data until the decision tree model training meets the standards.
[0093] Specifically, the dimensionality reduction model refers to a machine learning model that takes preprocessed semiconductor production data as input and outputs dimensionality reduction data, the dimensionality reduction training set refers to a data set divided from the learning data for training the decision tree model, the dimensionality reduction verification set refers to a data set divided from the learning data for verifying the training results of the decision tree model, the dimensionality reduction test set refers to a data set divided from the learning data for testing the decision tree model, the preset dimensionality reduction test accuracy refers to a preset value used to judge the training compliance status of the iteratively optimized decision tree model, the training compliance status of the iteratively optimized decision tree model refers to the accuracy compliance status of the trained decision tree model, and the training compliance status of the iteratively optimized decision tree model includes the training compliance status of the iteratively optimized decision tree model and the training failure of the iteratively optimized decision tree model.
[0094] Specifically, the preprocessing unit counts the data missing rate of each column and each row of production data in the semiconductor production data. When the data missing rate is greater than 30%, the column and the row of production data are deleted, and the outliers in the semiconductor production data are removed by statistical methods to obtain pure semiconductor production data. The pure semiconductor production data is standardized by the Z-score method to obtain second standardized semiconductor production data. The standardized semiconductor production data is also feature selected by a filtering method to remove redundant features to obtain preprocessed semiconductor production data.
[0095] Specifically, the data missing rate refers to the proportion of missing data in a single column and row in semiconductor production data to the total data in a single column and row. The statistical method refers to identifying and processing data points that deviate from the normal distribution through statistical principles and indicators to ensure data quality and the accuracy of subsequent analysis. This embodiment does not limit the specific method of the statistical method. For example, the interquartile range method can be used as a statistical method to remove outliers in semiconductor production data. The Z-score method refers to a statistical method for measuring the relative position of a data point in a data set. The calculation process is to take the difference between a number and the mean and divide it by the standard deviation, and convert the data set with different means and standard deviations into a standard normal distribution with a mean of 0 and a standard deviation of 1. The filtering method refers to a method used for feature selection in the data preprocessing stage, which measures the importance of the feature by calculating its statistic and removes redundant and unimportant features according to the set threshold.
[0096] Specifically, the dimensionality reduction processing unit inputs the semiconductor production data preprocessed in the preprocessing unit into the dimensionality reduction model, and the dimensionality reduction model outputs the dimensionality reduction data, and the output dimensionality reduction data is passed to the artificial intelligence monitoring module.
[0097] See also Figure 3 As shown in FIG, which is a schematic diagram of the structure of the artificial intelligence monitoring module of this embodiment, the artificial intelligence monitoring module includes:
[0098] A real-time fault monitoring unit, configured to monitor the fault condition in real time based on the dimensionality reduction data;
[0099] A fault prediction analysis unit, configured to predict a fault time point based on the dimensionality reduction data;
[0100] The correlation analysis unit is used to analyze the cause of the fault according to the dimension reduction data. The correlation analysis unit is connected to the real-time fault monitoring unit and the fault prediction analysis unit.
[0101] Specifically, the real-time fault monitoring unit inputs the dimension reduction data output from the dimension reduction processing unit into the fault monitoring model, outputs the semiconductor production fault situation, and feeds back the semiconductor production fault situation to the early warning alarm module.
[0102] Specifically, the semiconductor production failure condition refers to the failure condition corresponding to the reduced dimensionality data obtained by analyzing the reduced dimensionality data by the fault monitoring model. The semiconductor production failure condition includes no failure, general failure and serious failure. The no failure means that no failure occurs in semiconductor production. The general failure refers to a failure of semiconductor production within an acceptable range in semiconductor production, such as a decrease in the accuracy of the production machine. The serious failure refers to a failure that affects the production progress of semiconductor production, such as damage to the production machine. The fault monitoring model refers to a neural network learning model that takes the reduced dimensionality data as input and outputs the semiconductor production failure condition. The real-time fault monitoring unit constructs the fault monitoring model through a fault monitoring model construction method. The fault monitoring model construction method includes:
[0103] 70% of the historical fault data set is divided into a simulated training set and 30% of the historical fault data set is divided into a simulated validation set. A recurrent neural network model is selected as the neural network architecture of the fault monitoring model. The Adam optimizer and cross-entropy loss function are selected to train the recurrent neural network model. The simulated training set is loaded into the recurrent neural network model. Forward propagation is performed through the recurrent neural network model to calculate the output value of the fault monitoring model. The loss function value is calculated based on the output value of the recurrent neural network model and the true value. The gradient is calculated through the backpropagation algorithm, and the weights and bias of the recurrent neural network model are updated. The forward propagation, loss calculation and backpropagation process are repeated until the preset training rounds are reached. The accuracy of the recurrent neural network model is verified through the simulated validation set, and the recurrent neural network model with an accuracy rate of 90% is output as the fault monitoring model.
[0104] The historical fault data set refers to a data set used to train a fault monitoring model, and the historical fault data set includes historical dimensionality reduction data and historical semiconductor production failure situations corresponding to the historical dimensionality reduction data.
[0105] Specifically, the reduced-dimensional data is analyzed through the fault monitoring model to obtain the production failure situation of the semiconductor, so that the production of the semiconductor can be adjusted subsequently, avoiding the situation where the production of the semiconductor is stopped due to the failure to discover the fault in time, and timely maintenance of the semiconductor production can be carried out.
[0106] Specifically, the fault prediction analysis unit inputs the reduced-dimensionality data into a fault prediction model, outputs a future fault time point corresponding to the reduced-dimensionality data, and feeds the output future fault time point back to the early warning alarm module.
[0107] Specifically, the future failure time point refers to the time point at which the semiconductor production corresponding to the dimensionality reduction data fails in the future. The fault prediction model refers to a temporal convolutional network model that takes the dimensionality reduction data as input and the future failure time point as output. This embodiment does not limit the method for constructing the fault prediction model. Those skilled in the art can set it up by themselves according to actual needs, and only needs to meet the prediction requirements of the future failure time point. For example, the historical production data set can be divided into a front-end training set, a middle-end verification set, and a final test set in chronological order. The front-end training set accounts for 70% of the historical production data set, the middle-end verification set accounts for 15% of the historical production data set, and the final test set accounts for 15% of the historical production data set. The fault prediction model is trained using the front-end training set, verified using the middle-end verification set, and tested using the final test set until the test accuracy of the fault prediction model reaches 95%. The historical production data set refers to a data set in which historical semiconductor production data is converted into historical dimensionality reduction data and historical semiconductor production data fails at each time point.
[0108] Specifically, the correlation analysis unit is based on the i-th observation value of the independent variable X in the dimension reduction data, the mean value of the independent variable X The i-th observation value of the dependent variable Y in the historical failure cause and the mean value of the dependent variable Y Calculate the correlation coefficient r, i = 1, 2, 3...m, m is the observation order, set The correlation coefficient r is compared with the preset correlation coefficient r0, 0.8≤r0<1, and the correlation between the independent variable and the dependent variable is judged based on the comparison result. The fault cause inference model is trained using the independent variable and the dependent variable based on the judgment result, where:
[0109] When r<r0, the correlation analysis unit determines that the correlation between the independent variable and the dependent variable is irrelevant, and does not use the independent variable and the dependent variable to train the fault cause inference model;
[0110] When r≥r0, the correlation analysis unit determines that the correlation between the independent variable and the dependent variable is correlation, and trains the fault cause inference model using the independent variable and the dependent variable.
[0111] Specifically, the independent variables refer to input variables and influencing factors that may affect the occurrence of faults. The independent variables include equipment operation status data, process data, product quality inspection data, environmental data and personnel operation data. The dependent variables refer to output or response variables related to the fault. The dependent variables include fault types, and the fault types include motor faults, circuit faults and mechanical wear. The fault cause inference model refers to a structural equation model that combines factor analysis and path analysis methods, takes dimensionality reduction data as input, and takes fault causes as output. The fault cause refers to the cause of the fault when the dimensionality reduction data fails. This embodiment does not limit the training method of the fault cause inference model. Those skilled in the art can set it according to actual needs. It is sufficient to meet the output requirements for the cause of the fault. For example, the independent variables and dependent variables can be combined into a fault cause data set, 70% of the fault cause data set is divided into a fault cause training set to train the fault cause inference model, 20% of the fault cause data set is divided into a fault cause verification set to verify the results of the trained fault cause inference model, and 10% of the fault cause data set is divided into a fault cause test set to test the verified fault cause inference model. Repeat the above steps until the accuracy of the fault cause inference model test results reaches 97%. The correlation between the independent variable and the dependent variable refers to the association between the independent variable and the dependent variable. The correlation between the independent variable and the dependent variable includes the correlation between the independent variable and the dependent variable being uncorrelated and the correlation between the independent variable and the dependent variable being correlated.
[0112] Specifically, the correlation analysis unit inputs the dimension-reduced data into the fault cause inference model, outputs the fault cause corresponding to the dimension-reduced data, and feeds the fault cause back to the early warning alarm module.
[0113] See also Figure 4 As shown in FIG, which is a schematic diagram of the structure of the early warning alarm module of this embodiment, the early warning alarm module includes:
[0114] A fault warning unit, configured to provide a fault warning based on the future fault time point and the fault cause;
[0115] The fault alarm unit is used to issue a fault alarm according to the fault condition and the fault cause.
[0116] Specifically, the fault warning unit calculates the time difference Δty based on the future fault time point Ty output by the fault prediction and analysis unit and the actual time point Ts, sets Δty=Ty-Ts, and compares the time difference Δty with the preset time difference Δty0, 24h≤Δty0≤72h, where h is hours, and judges the necessity of warning based on the comparison result, and issues a fault warning based on the judgment result, wherein:
[0117] When Δty<Δty0, the fault warning unit determines that the warning necessity is unnecessary and does not perform a fault warning;
[0118] When Δty≥Δty0, the fault warning unit determines that the warning necessity is necessary, performs a fault warning, and transmits the future fault time point and the fault cause output by the correlation analysis unit to the control terminal for fault warning.
[0119] Specifically, the actual time point refers to the actual time when the fault unit judges the necessity of the warning, the preset time difference refers to the preset value used to judge the necessity of the warning, the warning necessity refers to whether the time difference is necessary for warning, the warning necessity includes the warning necessity being unnecessary and the warning necessity being necessary, and the control terminal refers to the terminal device that directly interacts with the user through the industrial big data mining system based on artificial intelligence. This embodiment does not limit the warning method of the fault warning. Those skilled in the art can set it according to actual needs. For example, the warning method of the fault warning can be set to send a warning message to the user's mobile phone.
[0120] Specifically, the fault alarm unit transmits the semiconductor production fault situation output by the real-time fault monitoring unit and the fault cause output by the correlation analysis unit to the control terminal for fault alarm.
[0121] See also Figure 5 , which is a schematic diagram of the structure of the deviation mining module of this embodiment, the deviation mining module includes:
[0122] A mining analysis unit for analyzing the deviation coefficient based on the deviation value and the deviation error;
[0123] a mining judgment unit, configured to judge the deviation according to the deviation coefficient, the mining judgment unit being connected to the mining analysis unit;
[0124] a deviation correction unit, configured to perform deviation correction on the dimensionality reduction process and the fault prediction analysis process according to the deviation, the deviation correction unit being connected to the mining judgment unit;
[0125] a deviation optimization unit, configured to perform real-time invalidity analysis on the semiconductor production data to obtain real-time invalidity analysis results, optimize the analysis process of the deviation coefficient based on the real-time invalidity analysis results, and optimize the storage of the production database based on the real-time invalidity analysis results, the deviation optimization unit being connected to the mining analysis unit;
[0126] The anti-error optimization unit is used to correct the optimization storage process of the production database according to the learning data, and is also used to adjust the deviation optimization process according to the update status of the learning data. The anti-error optimization unit is connected to the deviation optimization unit.
[0127] Specifically, the mining analysis unit calculates the deviation coefficient RMSE according to the deviation value αp of the semiconductor production data, the deviation error βp corresponding to the deviation value αp and the total number of deviation values f, where p = 1, 2, 3...z, where z is the order of the deviation values, and sets The deviation coefficient is obtained and fed back to the mining judgment unit.
[0128] Specifically, the deviation value of the semiconductor production data refers to the difference between the actual value of the semiconductor production data in the semiconductor production process and the preset reference standard. The deviation value of the semiconductor production data is calculated based on the actual value of the semiconductor production data in the semiconductor production process and the preset reference standard. Set αp=L1-L0, L1 is the actual value of the semiconductor production data in the semiconductor production process, L0 is the preset reference standard. This embodiment does not limit the specific value of the preset reference standard. Those skilled in the art can set it according to actual conditions, and only need to meet the calculation requirements of the deviation value of the semiconductor production data. For example, the specific value of the preset reference standard can be limited according to the equipment model of the semiconductor production equipment. The deviation error corresponding to the deviation value refers to the relative deviation value of the semiconductor production data and the preset reference standard in the semiconductor production process. The deviation error corresponding to the deviation value is calculated based on the semiconductor production data and the preset reference standard. Set The deviation coefficient refers to a quantitative indicator used to measure the degree to which semiconductor production data deviates from a reference standard.
[0129] Specifically, the mining judgment unit compares the deviation coefficient RMSE calculated in the mining analysis unit with the preset deviation coefficient RMSE0, 5%≤RMSE0≤15%, and judges the deviation of the semiconductor production data based on the comparison result, wherein:
[0130] When RMSE≤RMSE0, the mining judgment unit determines that the deviation of the semiconductor production data is normal, and transmits the judgment result to the deviation correction unit;
[0131] When RMSE>RMSE0, the mining judgment unit determines that the deviation of the semiconductor production data is abnormal, and transmits the judgment result to the deviation correction unit.
[0132] Specifically, the preset deviation coefficient refers to a preset value used to judge the deviation of semiconductor production data. The deviation of the semiconductor production data refers to the difference between the actual value of the semiconductor production data and the preset reference standard. The deviation of the semiconductor production data includes the deviation of the semiconductor production data being abnormal and the deviation of the semiconductor production data being normal.
[0133] Specifically, the deviation correction unit corrects the selection process of the dimensionality reduction method of the semiconductor production data according to the deviation judgment result of the semiconductor production data by the mining judgment unit, wherein:
[0134] If the mining judgment unit determines that the deviation of the semiconductor production data is normal, the deviation correction unit does not correct the selection process of the dimensionality reduction method of the semiconductor production data;
[0135] If the mining judgment unit determines that the deviation of the semiconductor production data is abnormal, the deviation correction unit corrects the selection process of the dimensionality reduction method of the semiconductor production data by the correction coefficient h=1.12-0.1×e -(RMSE-RMSE0) , e is the base of the natural logarithm, the preset cumulative variance contribution rate G0 is corrected to obtain the corrected preset cumulative variance contribution rate G0y, set G0y=G0×h, and replace the cumulative variance contribution rate G with the corrected preset cumulative variance contribution rate G0y to re-judge the linearity of the semiconductor production data.
[0136] Specifically, the deviation optimization unit calculates the effective coefficient EC according to the data missing rate DR, the data fluctuation coefficient FC, the data abnormality rate AR and the data consistency ratio CP, and sets EC=w1×(1-DR)+w2×(1-FC)+w3×(1-AR)+w4×(1-CP), w1=0.3, w2=0.25, w3=0.25, w4=0.2, w1 is the weight coefficient of the data missing rate, w2 is the weight coefficient of the data fluctuation coefficient, w3 is the weight coefficient of the data abnormality rate, and w4 is the weight coefficient of the data consistency ratio. The effective coefficient EC is compared with the preset effective coefficient EC0, 0.8≤EC0≤1, and the data validity is judged according to the comparison result, and the analysis process of the deviation coefficient is optimized according to the judgment result, wherein:
[0137] When EC≥EC0, the deviation optimization unit determines that the data validity is valid and does not optimize the analysis process of the deviation coefficient;
[0138] When EC<EC0, the deviation optimization unit determines that the data validity is invalid, optimizes the analysis process of the deviation coefficient, and optimizes the analysis coefficient OF=0.6-0.4×e-(EC0-EC) The deviation coefficient is optimized, e is the base of the natural logarithm, and the optimized deviation coefficient RMSEy is obtained. RMSEy = RMSE × OF is set, the deviation coefficient RMSE is replaced by the optimized deviation coefficient RMSEy, and the optimized deviation coefficient is fed back to the mining judgment unit.
[0139] Specifically, the data missing rate refers to the ratio of the number of missing valid data points to the total number of data points that should be collected within a certain time range. The data missing rate is calculated based on the number of missing data points and the total number of data points that should be collected. q is the number of missing data points, Q is the total number of data points that should be collected, and the data anomaly rate refers to the ratio of the number of abnormal data points that exceed the process allowable range and do not conform to the expected logic to the total number of data points. The data anomaly rate is calculated based on the number of abnormal data points z and the total number of data points Z. Set The data consistency ratio refers to the ratio of the number of records that meet the expected consistency between multi-source data and different dimensions of the same data to the total number of records. The data consistency ratio is calculated based on the number of consistent data records ay and the total number of data records Ay. The data fluctuation coefficient is an indicator of the degree of fluctuation of quantitative data in time series. The fluctuation coefficient is calculated based on the data standard deviation sx and the data mean Wx. This embodiment does not limit the method for obtaining the number of missing data points, the total number of data points to be collected, the number of abnormal data points, the total number of data points, the number of consistent data records, the total number of data records, the data standard deviation and the data mean. Those skilled in the art can set them according to actual needs. For example, the total number of collected data points can be obtained according to the sampling frequency and time window of the equipment. The sampling frequency refers to the frequency of collecting data points, such as setting the sampling frequency to 5 times per minute. The time window refers to the total time for collecting a batch of data points, such as setting the time window to 1 hour. The preset validity coefficient refers to a preset value used to judge the validity of the data. The data validity refers to whether the collected semiconductor production data is valid. The data validity includes data validity being valid and data validity being invalid.
[0140] Specifically, when the deviation optimization unit determines that the data validity is invalid, the production database is optimized and stored, and missing data points and abnormal data points are deleted to optimize the production database, and optimized data and analysis results are provided to support subsequent decision-making processes.
[0141] Specifically, the error-prevention optimization unit compares the learning data update time Tg with the preset learning data update time Tg0, 72h<Tg0<96h, and judges the validity of the learning data based on the comparison result, and corrects the optimization storage process of the production database based on the judgment result, wherein:
[0142] When Tg≥Tg0, the error-prevention optimization unit determines that the validity of the learning data is invalid and does not correct the optimization storage process of the production database;
[0143] When Tg<Tg0, the anti-error optimization unit determines that the validity of the learning data is valid, and corrects the optimization storage process of the production database. The correction scheme is to optimize the storage of the production database when the deviation optimization unit determines that the data validity is invalid, delete the missing data points and abnormal data points, and replace them with not deleting the missing data points and abnormal data points.
[0144] Specifically, the learning data update duration refers to the time value from the last time the learning data was updated to the present. This embodiment does not limit the method for obtaining the learning data update duration. Technical personnel in this field can set it according to actual needs. For example, the learning data update duration can be obtained based on the system time count. The preset learning data update duration refers to a preset value used to judge the validity of the learning data. The learning data validity refers to whether the learning data is valid learning data. The learning data validity includes the learning data validity being valid and the learning data validity being invalid.
[0145] Specifically, the anti-mistake optimization unit compares the learning data update times Kg with the preset learning data update times Kg0, 10 times < Kg0 < 15 times, and judges the learning data update status according to the comparison result, and adjusts the deviation optimization process according to the judgment result, wherein:
[0146] When Kg≥Kg0, the anti-error optimization unit determines that the learning data update is qualified and does not adjust the deviation optimization process;
[0147] When Kg<Kg0, the error-proof optimization unit determines that the learning data update is unqualified and adjusts the deviation optimization process by adjusting the deviation optimization coefficient NT=1.4-0.4×e -(Kg0-Kg) The preset effective coefficient EC0 is adjusted, e is the base of the natural logarithm, and the adjusted preset effective coefficient EC0N is obtained. EC0N=EC0×NT is set, the preset effective coefficient EC0 is replaced with the adjusted preset effective coefficient EC0N, and the effective coefficient EC is re-compared with the adjusted preset effective coefficient EC0N.
[0148] Specifically, the number of learning data updates refers to the number of learning data updates collected by the system. This embodiment does not limit the method for obtaining the number of learning data updates. Those skilled in the art can set it according to actual needs, such as obtaining the number of learning data updates based on the system's count of the number of learning updates. The preset number of learning data updates refers to a preset value for judging the learning data update status. The learning data update status refers to whether the learning data update is qualified. The learning data update status includes the learning data update status being qualified and the learning data update status being unqualified.
[0149] See also Figure 6 As shown in FIG, it is a flow chart of the industrial big data mining method based on artificial intelligence in this embodiment, including:
[0150] Step S1, collecting semiconductor production data through a data acquisition module;
[0151] Step S2, storing the semiconductor production data through a data storage module;
[0152] Step S3, performing dimensionality reduction processing on the semiconductor production data by a dimensionality reduction processing module to obtain dimensionality-reduced production data;
[0153] Step S4: Using an artificial intelligence monitoring module to perform real-time fault monitoring, fault prediction analysis, and correlation analysis based on the reduced-dimensional production data, the fault condition, fault time point, and fault cause are obtained.
[0154] Step S5: performing a fault warning and a fault alarm according to the fault condition, the fault time point and the fault cause through the early warning alarm module;
[0155] Step S6, analyzing the deviation coefficient according to the deviation value and the deviation error by the deviation mining module, and judging the deviation situation according to the deviation coefficient;
[0156] Step S7: performing deviation correction on the dimensionality reduction process and the fault prediction analysis process according to the deviation situation through the deviation mining module.
[0157] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
Claims
1. An industrial big data mining system based on artificial intelligence, characterized in that: include: A data acquisition module for collecting semiconductor production data; A data storage module, used for storing the semiconductor production data; A production database for storing the semiconductor production data; A dimensionality reduction processing module, configured to perform dimensionality reduction processing on the semiconductor production data to obtain dimensionality-reduced production data; An artificial intelligence monitoring module is used to perform real-time fault monitoring, fault prediction analysis, and correlation analysis based on the reduced-dimensional production data to obtain the fault condition, fault time point, and fault cause respectively; A warning and alarm module is used to provide fault warning and fault alarm according to the fault condition, the fault time point and the fault cause; The deviation mining module is used to analyze the deviation coefficient according to the deviation value and deviation error, and to judge the deviation situation according to the deviation coefficient. It is also used to perform deviation correction on the dimensionality reduction process and the fault prediction analysis process according to the deviation situation.
2. The industrial big data mining system based on artificial intelligence according to claim 1 is characterized in that: The dimensionality reduction processing module includes: A data acquisition unit, configured to acquire semiconductor production data from a production database, and also configured to acquire learning data and extract learning data from a dimensionality reduction processing unit; a dimensionality reduction method selection unit, configured to select a dimensionality reduction method according to the semiconductor production data; A model building unit, configured to build a current dimensionality reduction model according to the dimensionality reduction method, and to train the current dimensionality reduction model according to the learning data; a preprocessing unit, configured to preprocess the semiconductor production data to obtain preprocessed semiconductor production data; A dimensionality reduction processing unit is used to perform dimensionality reduction processing on the pre-processed semiconductor production data according to the current dimensionality reduction model and output the dimensionality reduction data.
3. The industrial big data mining system based on artificial intelligence according to claim 2 is characterized in that: The dimensionality reduction method selection unit analyzes the semiconductor production data through the principal component analysis method to obtain the cumulative variance contribution rate G, and compares the cumulative variance contribution rate G with the preset cumulative variance contribution rate G0, judges the linearity of the semiconductor production data based on the comparison result, and selects the dimensionality reduction method of the semiconductor production data based on the judgment result.
4. The industrial big data mining system based on artificial intelligence according to claim 3 is characterized in that: The preprocessing unit preprocesses the semiconductor production data to obtain preprocessed semiconductor production data, the dimensionality reduction processing unit inputs the preprocessed semiconductor production data in the preprocessing unit into the dimensionality reduction model, and the dimensionality reduction model outputs the reduced dimensionality data, and the output reduced dimensionality data is passed to the artificial intelligence monitoring module.
5. The industrial big data mining system based on artificial intelligence according to claim 1 is characterized in that: The artificial intelligence monitoring module includes: A real-time fault monitoring unit, configured to monitor the fault condition in real time based on the dimensionality reduction data; A fault prediction analysis unit, configured to predict a fault time point based on the dimensionality reduction data; A correlation analysis unit is used to analyze the cause of the fault based on the dimension reduction data.
6. The industrial big data mining system based on artificial intelligence according to claim 5 is characterized in that: The correlation analysis unit inputs the dimension-reduced data into the fault cause inference model, outputs the fault cause corresponding to the dimension-reduced data, and feeds back the fault cause to the early warning alarm module.
7. The industrial big data mining system based on artificial intelligence according to claim 1 is characterized in that: The early warning alarm module includes: A fault warning unit, configured to provide a fault warning based on the fault time point and the fault cause; The fault alarm unit is used to issue a fault alarm according to the fault condition and the fault cause.
8. The industrial big data mining system based on artificial intelligence according to claim 1 is characterized in that: The deviation mining module includes: A mining analysis unit for analyzing the deviation coefficient based on the deviation value and the deviation error; a mining judgment unit, configured to judge the deviation situation according to the deviation coefficient; A deviation correction unit, configured to perform deviation correction on the dimensionality reduction process and the fault prediction analysis process according to the deviation; The deviation optimization unit is used to perform real-time invalidity analysis on the semiconductor production data to obtain real-time invalidity analysis results, optimize the analysis process of the deviation coefficient based on the real-time invalidity analysis results, and optimize the storage of the production database based on the real-time invalidity analysis results.
9. The artificial intelligence-based industrial big data mining system according to claim 8, characterized in that: The deviation optimization unit calculates the effective coefficient EC based on the data missing rate DR, the data fluctuation coefficient FC, the data anomaly rate AR and the data consistency ratio CP, compares the effective coefficient EC with the preset effective coefficient EC0, judges the data validity based on the comparison result, and optimizes the analysis process of the deviation coefficient based on the judgment result.
10. A method for applying to the artificial intelligence-based industrial big data mining system according to any one of claims 1 to 9, characterized in that: include: Step S1, collecting semiconductor production data through a data acquisition module; Step S2, storing the semiconductor production data through a data storage module; Step S3, performing dimensionality reduction processing on the semiconductor production data by a dimensionality reduction processing module to obtain dimensionality-reduced production data; Step S4: Using an artificial intelligence monitoring module to perform real-time fault monitoring, fault prediction analysis, and correlation analysis based on the reduced-dimensional production data, the fault condition, fault time point, and fault cause are obtained. Step S5: performing a fault warning and a fault alarm according to the fault condition, the fault time point and the fault cause through the early warning alarm module; Step S6, analyzing the deviation coefficient according to the deviation value and the deviation error by the deviation mining module, and judging the deviation situation according to the deviation coefficient; Step S7: performing deviation correction on the dimensionality reduction process and the fault prediction analysis process according to the deviation situation through the deviation mining module.
Citation Information
Patent Citations
A Data Mining Method Based on Big Data from the Hot Rolling Industry
CN112181956B