Die casting defect prediction method based on integrated machine learning method
By integrating machine learning methods and feature engineering, and combining SMOTE and soft voting fusion technology, the problems of low efficiency and poor adaptability in die casting defect detection are solved, enabling real-time prediction and prevention, and improving the quality and production efficiency of die castings.
Patent Information
- Application Number
- CN202511236909.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-11-28
AI Technical Summary
Existing methods for detecting defects in die castings are inefficient, costly, and unreliable. They cannot achieve real-time prediction and intervention, and single models have limited ability to fit the nonlinear coupling relationship of multi-dimensional parameters, making it difficult to adapt to changes in the production environment.
An ensemble machine learning approach is adopted, combining feature engineering and hyperparameter optimization techniques. A synthetic balanced dataset is generated through SMOTE, and multiple machine learning models (such as random forest, support vector machine, and neural network) are selected and fused using soft voting to achieve real-time data acquisition and online learning.
It significantly improves the accuracy and robustness of defect prediction, adapts to complex and ever-changing production environments, enables real-time prediction and prevention, and enhances product quality and production efficiency.
Smart Images

Figure CN121032999A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of die casting defect prediction, and relates to a die casting defect prediction method based on an integrated machine learning method. BACKGROUND
[0002] Die casting manufacturing is a metal forming process widely used in the fields of automobiles, aerospace, electronics, etc., which can efficiently produce complex-shaped parts. However, due to the complex conditions such as high temperature, high pressure and rapid cooling in the die casting process, defects such as porosity, cracks, shrinkage, etc. often occur in die castings, affecting product performance and possibly causing safety hazards. Traditional defect detection methods mainly rely on manual visual inspection or non-destructive testing techniques (such as X-ray, ultrasonic testing), which have problems such as low efficiency, high cost, poor reliability, etc., and cannot realize real-time prediction and intervention.
[0003] With the development of industrial big data and machine learning technology, die casting defect prediction has gradually shifted to data-driven methods. In existing research, CN112347890A proposes a die casting defect detection method based on support vector machine (SVM), which collects die casting process parameters and inputs them into the SVM model to realize defect classification. Although this method has initially realized data-driven defect detection, it only relies on a single SVM model, and has limited fitting ability for the nonlinear coupling relationship of multi-dimensional parameters in the die casting process, and has not optimized the imbalance problem between defect samples and normal samples, resulting in a detection recall rate of less than 65% for small sample defects (such as micro-cracks). Another research "Neural Network-based Defect Prediction in Die Casting Processes" uses a multi-layer perceptron neural network to build a prediction model, which has certain advantages in feature extraction, but the model is easily affected by data distribution fluctuations, and does not introduce a dynamic learning mechanism. When the production environment (such as changes in raw material batches, equipment wear) changes, the model generalization error increases significantly, and the prediction accuracy in actual production decreases to less than 70%.
[0004] To solve the above problems, the present application proposes a die casting defect prediction method based on an integrated machine learning method. By integrating multiple machine learning models (such as random forests, support vector machines, neural networks, etc.), and combining feature engineering and hyperparameter optimization techniques, the accuracy and robustness of defect prediction are significantly improved. In addition, the present application introduces real-time data acquisition and online learning mechanisms, enabling the model to dynamically adapt to changes in the production environment, realizing real-time prediction and prevention of die casting defects, thereby improving product quality and production efficiency. This method is not only suitable for die casting manufacturing, but also can be extended to other metal forming processes, and has important industrial application value. SUMMARY
[0005] In view of the problems existing in the prior art, the application provides a die casting defect prediction method based on an integrated machine learning method, solves the problems of insufficient data quality, single model limitation and poor generalization ability in the prior art, and adapts to complex and changeable production environments.
[0006] In order to achieve the above-mentioned purpose, the technical scheme adopted is:
[0007] A die casting defect prediction method based on an integrated machine learning method, the die casting defect prediction method: first, collecting and storing die casting process parameter data through a target enterprise MES system to form an original database. Secondly, using the idea of SMOTE (synthetic minority over-sampling technique), generating synthetic samples to balance the sample quantity between different categories, thereby improving the performance of the model, and forming a historical database. Thirdly, based on the historical database, testing the performance of six different state-of-the-art machine learning classification algorithms in die casting defect classification: simple logistic regression model, SVM learning with random gradient descent optimization, multilayer perceptron network, random decision tree method, k-nearest neighbor classification algorithm, and naive Bayes classification. Fourthly, evaluating the above six classifiers, and selecting three classification algorithms as the best classifiers according to the best performance calculated by f-measure. Finally, the selected best classifiers are fused by soft voting, and the probability outputs of the three models are weighted and averaged to form a fused probability distribution, and the category with the maximum probability is selected as the final prediction result. The application integrates multiple machine learning models (such as random forest, support vector machine, neural network, etc.), and combines feature engineering and hyperparameter optimization technology, which significantly improves the accuracy and robustness of defect prediction. Specifically, the following steps are included:
[0008] Step 1, collecting die casting process data based on the MES system of an enterprise, in the application, collecting casting condition process variables, sensor data and fault data to form an original database. The collected casting condition process variables include injection speed V1, V2, V3, cycle time, spraying time, casting pressure, cake thickness, clamping force, part taking time, top opening time, pressure rising time and vacuum degree.
[0009] Step 2, since the sample quantity of the defect category of the original data set obtained in step 1 is much less than that of the normal part category, new minority class samples are generated by using the synthetic minority over-sampling method (SMOTE) to balance the sample quantity of different categories in the training data set, and a historical database is formed. The SMOTE method is defined as follows:
[0010] f new =f i +δ·(f zi -f i )
[0011] wherein f new denotes the generated new synthetic minority class sample point feature vector; f denotes the feature vector of the original minority class sample point; f zi denotes f i a randomly selected neighbor feature vector in the k-nearest neighbors of f i ; δ denotes a random number between 0 and 1, determining the interpolation proportion of the new sample;
[0012] Step 3, the historical database obtained in step 2 is subjected to feature preprocessing, and a minimum redundancy maximum correlation (mRMR) analysis method is used to perform feature screening among a plurality of process parameters affecting the forming quality of the casting. The judgment standard is that the input features need to maintain a high correlation with the forming quality of the casting while reducing the information redundancy with the selected feature subset as much as possible, so as to ensure that the selected feature subset can not only comprehensively represent the main influencing factors of the forming quality of the casting, but also avoid information repetition caused by high correlation between features. The selected feature subset is further divided into a training set and a test set for subsequent model construction and performance verification. Specifically as follows:
[0013] Step 3.1, calculate the maximum correlation D and the minimum redundancy R:
[0014]
[0015] wherein x i and x k denote the i-th and k-th features in the historical database; C denotes the target feature, i.e. the quality of the casting; I(x i ; C) denotes the correlation function of the i-th feature and the target feature; I(x i ; x k ) denotes the correlation function of the i-th feature and the k-th feature; T denotes the selected feature subset. The correlation function I is calculated by the following formula:
[0016]
[0017] wherein u and u' denote two random variables in the original database; P(·,·) denotes the probability.
[0018] Step 3.2, determine the minimum redundancy maximum correlation objective function by combining the above three formulas, defined as follows:
[0019]
[0020] Step 3.3, the ranking result of input features can be obtained by solving the minimum redundancy maximum relevance objective function using the forward selection algorithm. The above process is called the minimum redundancy maximum relevance method. Finally, the injection velocity V1, V2, V3, cycle time, spraying time, casting pressure, cake thickness, clamping force, pickup time, top opening time, pressure rise time, and vacuum degree are selected as 12 process parameters for defect prediction to obtain the final database, and the features are processed. Finally, the final database after feature processing is divided into training set and test set.
[0021] Further, the feature processing in step 3.3 is specifically: data cleaning, standardization and outlier processing are performed on the 12 input process parameters, and dimensionality reduction is performed when necessary to form a historical database that can be used for training set and test set.
[0022] Step 4, the performance of six different classification algorithms is evaluated on the historical database after feature processing in step 3, Figure 1 The performance of all 6 classifiers in F3 score is compared and the classification result is obtained.
[0023] The six different classification algorithms include simple logistic regression learning model, SVM learning with random gradient descent optimization, multilayer perceptron network, random decision tree method, k-nearest neighbor classification algorithm, and naive Bayes classification.
[0024] After analysis, it is found that the simple logistic regression (SLR) learning model, the SVM learning with random gradient descent (SGD) optimization and the multilayer perceptron network (MLP) show better performance in F3 score than the other five classification algorithms. After analyzing these results, F3 score is a variant of F-score used to measure the accuracy of the model, which is a weighted harmonic mean, mainly used to give more weight to recall (Recall) in evaluation. Based on the F3 score, the top three scores are selected as the best classifiers, i.e. simple logic learning, naive Bayes and random forest, and the classification results are passed to predict the final output.
[0025] Step 5, based on the three best classifiers selected in step 4, a new soft voting based ensemble classification method is proposed. The soft voting uses a probability average voting mechanism, that is, the prediction probabilities of each class of the three best classifiers are averaged to obtain the final classification probability. As shown in Figure 2 The structure and classification process of the ensemble framework are shown. The voting mechanism used is described in detail as follows:
[0026] Step 5.1, soft voting;
[0027] In soft voting, the average value of the selection probability is voted, and the final classification result is obtained by weighted average of the prediction probability of the selected three classifiers. Specifically, based on the prediction probability of each classifier for each category, the category with the highest score is selected as the final prediction result after weighted summation. The soft voting formula is as follows: let the prediction probability of three classifiers for category c be P1(c), P2(c) and P3(c), and the final prediction probability of the integrated classifier is:
[0028]
[0029] The integrated classifier determines the final classification result according to the maximum value of P final (c).
[0030] In order to evaluate the performance of the integrated classification method based on soft voting, the following classification evaluation indexes are commonly used:
[0031]
[0032] Among them: TP represents correctly predicting positive samples as positive; FN represents incorrectly predicting positive samples as negative; FP represents incorrectly predicting negative samples as positive; and TN represents correctly predicting negative samples as negative.
[0033] In the above evaluation indexes, the greater the estimated index value of precision, recall, F1 value, accuracy and, the better the prediction result of the soft voting integrated model.
[0034] Compared with the prior art, the present application has the following beneficial effects:
[0035] (1) The present application adopts the minimum redundancy maximum correlation method to determine the optimal input feature subset in the historical data set from the perspective of reality, removes redundant features on the premise of ensuring maximum correlation, can select the most relevant information, and can reduce redundant information.
[0036] (2) The present application improves data quality by introducing data enhancement and sample resampling technology, combines multi-model fusion methods such as logistic regression, random forest and support vector machine based on SGD optimization, uses soft voting mechanism for decision fusion, effectively overcomes the limitations of single model and overfitting problem, simultaneously selects the best sub-model based on F3 score, significantly improves the accuracy, robustness and generalization ability of the defect detection system, and is suitable for complex and variable actual die casting production environment.
[0037] (3) By integrating multiple classification models with excellent performance (such as random forest, SVM and neural network), the soft voting mechanism is used to fuse the prediction results of each model, and the advantages of different models in data feature extraction are fully utilized to improve the overall prediction performance. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 The evaluation index comparison chart of the algorithm of the application.
[0039] Figure 2 The overall flowchart of the application "a die casting defect prediction method based on an integrated machine learning method". DETAILED DESCRIPTION
[0040] The application provides a die casting defect prediction method based on an integrated machine learning method, good die casting forming rules are obtained by using a data mining algorithm, and the die casting defect classification prediction method is applied, and the following is specifically described in combination with the drawings and examples:
[0041] The overall process of the application is shown in the accompanying Figure 1 The specific implementation measures are as follows:
[0042] Step 1, collecting die casting process data based on an enterprise MES system, in the application, casting condition process variables, sensor data and fault data are used to develop a defect classification prediction model, and an original database is formed. The collected casting process variable data includes injection speed V1, V2 and V3, cycle time, spraying time, casting pressure, cake thickness, clamping force, part taking time, top opening time, pressure rising time and vacuum degree.
[0043] Specific data collection items are shown in Table 1:
[0044] Table 1, actual data collection items
[0045] NO Acquisition process parameter name NO Acquisition process parameter name 1 Injection speed V1 7 Briquetting thickness 2 Injection speed V2 8 Clamping force 3 Injection speed V3 9 Pick-up time 4 Cycle time 10 Ejection time 5 Spraying time 11 Boosting rise time 6 Casting pressure 12 Vacuum degree
[0046] Step 2, since the sample number of the data set defect category is far less than that of the normal part category, the original data obtained in step 1 is used by adopting the idea of SMOTE (synthetic minority over-sampling technique). SMOTE generates new samples between two minority class samples by linear interpolation, thereby effectively alleviating the overfitting problem caused by random over-sampling.
[0047] Finally, SMOTE over-sampling changes the data distribution of the unbalanced data set by adding generated minority class samples, and forms a historical database.
[0048] Step 3, the historical database obtained in step 2 is subjected to feature preprocessing, and a minimum redundancy maximum relevance method (mRMR) combined with a forward selection algorithm is used to screen a plurality of process parameters affecting the forming quality of the casting, the selected features need to maintain a high correlation with the forming quality of the casting while reducing the redundancy with the selected feature subset as much as possible, thereby forming a feature subset capable of effectively representing the main influencing factors of the forming quality of the casting; finally, injection speeds V1, V2 and V3, cycle time, spraying time, casting pressure, cake thickness, clamping force, part taking time, top opening time, pressure rise time and vacuum degree are selected, a total of 12 process parameters, and data cleaning, standardization and outlier processing are performed, and dimension reduction is performed if necessary, thereby forming a historical database that can be used for training set and test set.
[0049] Step 4, the performance of six different classification algorithms is respectively evaluated on the historical database after feature processing in step 3, and Table 2 compares the performance of all 6 classifiers in terms of F3 score. The simple logistic regression (SLR) learning model, the SVM learning with random gradient descent (SGD) optimization and the multilayer perceptron network (MLP) show better performance in terms of F3 score than the other five classification algorithms. After analyzing these results, three best classifiers are selected based on the F3 score, namely simple logistic learning, SVM (Support Vector Machine) learning with random gradient descent optimization and multilayer perceptron network, and the classification results are passed to different voting mechanisms to predict the final output.
[0050] Table 2, comparison of performance of 6 classifier algorithms
[0051]
[0052] Step 5, based on the selected classifier in step 4, a new integrated classification method based on soft voting is proposed. The soft voting consists of probability average voting, probability product voting, probability minimum voting or probability maximum voting;
[0053] The selected best classifiers are fused by soft voting, and the classification results of the three best classifiers are provided to the integrated classification method, which provides the final classification result. Table 3 shows the comparison of the fused probability and the prediction results of the algorithm before fusion.
[0054] Table 3, comparison of prediction results of different algorithms
[0055] Algorithm Accuracy Precision Recall rate F1 score F2 score F3 score After algorithm fusion 99.38% 0.9942 0.9942 0.9942 0.9942 0.9942 Simple logic learning 98.80% 0.983 0.982 0.988 0.988 0.988 Naive Bayes 98.08% 0.984 0.984 0.984 0.984 0.984 Random forest 98.38% 0.985 0.985 0.983 0.982 0.981
[0056] As can be clearly seen from Table 3, the integrated model using the soft voting fusion strategy is superior to the individual performance of each base classifier in multiple evaluation indicators (such as accuracy, recall rate, F3 score), indicating that the fusion mechanism effectively improves the overall discriminant ability of the model. All results are based on a 70:30 training-test data division, ensuring sufficient training samples and representative test results.
[0057] To further verify the stability and generalization ability of the model, the present application also evaluates it under different cross-validation settings. The results show that in 10-fold cross-validation, the model achieves an average accuracy of 98.63%; in 5-fold and 2-fold cross-validation, the accuracy is 98.45% and 99.26%, respectively. This indicates that the integrated algorithm has strong robustness and adaptability, and can maintain consistent and excellent prediction performance under different data division, suitable for data fluctuations and distribution changes that may occur in actual production.
[0058] Through the defect prediction algorithm proposed in the present application, enterprises can realize real-time optimization and dynamic adjustment of process parameters in the die casting production process, early warning of potential defect risks, and significant reduction of product defect rate, realizing the transformation of quality control mode from "after detection" to "before prevention". This method not only improves the consistency and reliability of products, but also improves the production efficiency and intelligent level, and has wide industrial application prospect.
[0059] The above-described embodiments only express the implementation of the present application, but should not be construed as limiting the scope of the present patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of protection of the present application.
Claims
1. A die casting defect prediction method based on an ensemble machine learning method, characterized by, The die casting defect prediction method comprises the following steps: Step 1, collecting die casting process data based on the MES system of the enterprise, in the present application, collecting casting condition process variables, sensor data and fault data to form an original database; Step 2, since the sample number of the defect category of the original data set obtained in step 1 is far less than that of the normal parts, a new minority class sample is generated by using a synthetic minority over-sampling technique (SMOTE) on the original database to form a historical database; Step 3, performing feature preprocessing on the historical database obtained in step 2, and performing feature screening on a plurality of process parameters affecting the forming quality of the casting by using a minimum redundancy maximum correlation analysis method; the feature subset obtained after screening is further divided into a training set and a test set for subsequent model construction and performance verification; Step 4, evaluating the performance of a plurality of different classification algorithms on the historical database after feature preprocessing in step 3, selecting the top three with the highest scores as the best classifiers, and passing the classification results to the prediction final output; Step 5, based on the three best classifiers selected in step 4, a new integrated classification method based on soft voting is proposed; the soft voting adopts a probability average voting mechanism, that is, the prediction probabilities of each category of the three best classifiers are averaged to obtain the final classification probability.
2. The die casting defect prediction method based on integrated machine learning method according to claim 1, characterized in that, In step 1, the collected casting condition process variables include injection speed V1, V2, V3, cycle time, spraying time, casting pressure, cake thickness, clamping force, part taking time, top opening time, and pressure rising time.
3. The die casting defect prediction method based on integrated machine learning method according to claim 1, characterized in that, In step 2, the SMOTE method is defined as follows: f new = f i + δ · (f zi - f i ) where f new represents the generated new synthetic minority class sample point feature vector; f i represents the feature vector of the original minority class sample point; f zi represents the k-nearest neighbor feature vector randomly selected from f i ; and δ represents a random number between 0 and 1, which determines the interpolation ratio of the new sample.
4. The die casting defect prediction method based on integrated machine learning method according to claim 3, characterized in that, Step 3 specifically comprises: Step 3.1, calculating the maximum correlation D and the minimum redundancy R: where x i and x k represent the i-th and k-th features in the historical database, respectively; C represents the target feature, i.e. the casting quality; I(x i ; C) represents the correlation function of the i-th feature and the target feature; I(x i ; x k ) represents the correlation function of the i-th feature and the k-th feature; T represents the selected feature subset; the correlation function I is calculated by the following formula: Wherein, u and u' represent two random variables in the original database; P(·,·) represents probability; Step 3.2, determine the minimum redundancy maximum correlation objective function by combining the above three formulas, which is defined as follows: Step 3.3, solve the minimum redundancy maximum correlation objective function by using a forward selection algorithm, and the ranking result of the input features can be obtained, the above process is called the minimum redundancy maximum correlation method; finally, 12 process parameters including injection speed V1, V2, V3, cycle time, spraying time, casting pressure, cake thickness, clamping force, part taking time, top opening time, and pressure rising time are selected for defect prediction to obtain the final database, and feature processing is performed on the final database, and finally the final database after feature processing is divided into a training set and a test set.
5. The die casting defect prediction method based on integrated machine learning method according to claim 4, characterized in that, In step 3.3, the feature processing specifically comprises: performing data cleaning, standardization and abnormal value processing on the 12 input process parameters, and performing dimension reduction when necessary to form a historical database that can be used for the training set and the test set.
6. The die casting defect prediction method based on integrated machine learning method according to claim 4, characterized in that, Step 4 specifically comprises: The plurality of different classification algorithms include a simple logistic regression learning model, a SVM learning with random gradient descent optimization, a multi-layer perceptron network, a random decision tree method, a k-nearest neighbor classification algorithm, and a naive Bayes classification. Based on the F3 score, the top three highest scores are selected as the best classifiers, including simple logic learning, naive Bayes and random forest, and the classification results are passed to predict the final output.
7. The die casting defect prediction method based on integrated machine learning method according to claim 4, characterized in that, The soft voting in step 5 is specifically: In the soft voting, the probability average voting method is selected, and the final classification result is obtained by weighted average of the prediction probabilities of the three selected classifiers; specifically, based on the prediction probability of each classifier for each category, the category with the highest score is selected as the final prediction result after weighted summation; the soft voting formula is as follows: let the prediction probabilities of three classifiers for category c be P1(c), P2(c) and P3(c), and the final prediction probability of the integrated classifier is: The ensemble classifier determines a final classification result from P final the maximum of (c).
8. The die casting defect prediction method based on the integrated machine learning method according to claim 7, characterized in that, In order to evaluate the performance of the integrated classification method based on soft voting, the following classification evaluation indicators are used: Wherein: TP represents correctly predicting positive samples as positive; FN represents incorrectly predicting positive samples as negative; FP represents incorrectly predicting negative samples as positive; and TN represents correctly predicting negative samples as negative; In the above evaluation indicators, the greater the estimated indicator values of precision, recall, F1 value, accuracy and AUC, the better the prediction result of the soft voting integrated model.