Aluminum alloy production and processing yield prediction method and system based on big data
Patent Information
- Application Number
- CN202611282409.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-24
- Publication Date
- 2026-09-18
AI Technical Summary
[0004]本发明解决的技术问题是:因为缺乏凝固热力学机理约束、预测框架无法跟随工况进行自适应调整、并且缺少物理过程一致性校验,导致跨牌号、跨工况预测失稳
[0016] The beneficial effects of this invention are as follows: By using the solidification temperature range as a core variable in metallurgy throughout the entire prediction process, unlike conventional algorithms that rely solely on statistical correlation coefficients, this invention introduces a physical rule table indexed by the solidification temperature range during feature selection. It compares the statistical direction with the physical theoretical direction for each feature, transforming pure statistical selection into a dual-criteria selection based on statistics and physics, fundamentally eliminating misjudgments of the direction of physical abrupt changes. During fusion, unlike traditional ensemble learning that indiscriminately participates in all basis functions, this invention introduces process distance and effective bandwidth, enabling adaptive activation of the best-performing basis functions in the current solidification temperature range for prediction. This transforms static weighting into dynamic weighting driven by operating conditions, avoiding hard extrapolation of fixed structures. During iterative verification, unlike relying solely on statistical labels to calculate residuals, this invention introduces physical residuals based on the CALPHAD solidification path to drive feature admission thresholds and negative feedback updates of basis function bandwidth. This transforms statistical residual feedback into dual-parameter feedback driven by physical residuals, ensuring stable predictions when the solidification range crosses critical values.
Smart Images

Figure CN122779720A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of production data prediction technology, and in particular to a method and system for predicting the output of aluminum alloy production and processing based on big data. Background Technology
[0002] In recent years, even for the same grade, the batch-to-batch fluctuations in melt chemical composition and the real-time coupling of dynamic parameters such as casting speed and cooling water flow rate during the direct-cooling semi-continuous casting process of aluminum alloys have led to significant drifts in the solidification temperature range. When the drift exceeds the process window, the risk of hot cracking and central shrinkage porosity increases significantly, directly determining the output of the batch. The current production line's commonly used output forecasting methods mostly treat each process parameter as an isolated feature input, without introducing the solidification temperature range as a core physical variable connecting composition, process, and defects into the forecasting loop. Furthermore, it fails to distinguish the heat transfer differences between different casting stages: start-up, steady state, and end-up. This results in forecasting deviations often exceeding the industrially permissible engineering threshold when switching alloy grades or adjusting casting speed, causing companies to rely on manual experience for repeated trial and error, severely restricting the achievement of continuous automatic scheduling and closed-loop quality control goals.
[0003] Existing statistical regression and conventional machine learning methods have the following inherent defects when dealing with aluminum alloy production forecasting. On the one hand, they lack prior constraints on the direction of solidification thermodynamics, making it very easy to misjudge statistical spurious correlations in production line noise as process regularities. When the solidification temperature range crosses a specific boundary, the correlation direction between casting speed and output may even reverse. Existing methods are unable to characterize such physical abrupt changes, resulting in prediction results that deviate significantly from reality. On the other hand, fixed structures perform well within the solidification range covered by the training samples, but once the solidification characteristics of a new batch of alloy deviate from the distribution of the training set, all basis functions still participate in the prediction with equal authority, lacking a structure adaptive mechanism based on the applicability of operating conditions. Furthermore, the residual verification of existing methods only stays at the level of statistical error of production label, without introducing thermodynamic process information such as solidification path and phase transformation sequence as internal consistency criteria. This results in high-confidence outputs in physically unreliable extrapolation regions, creating hidden dangers of quality control failure. Summary of the Invention
[0004] The technical problem solved by this invention is that the lack of solidification thermodynamic mechanism constraints, the inability of the prediction framework to adapt to the working conditions, and the lack of physical process consistency verification lead to instability in predictions across grades and working conditions.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a method for predicting the output of aluminum alloy production and processing based on big data, comprising the following steps: Step S1: Calculate the production data to obtain a standardized feature set and solidification temperature range; Step S2: Based on statistical and physical directions, perform consistency feature screening on the standardized feature set to obtain a high-confidence feature subset, and adaptively activate each candidate basis function in the candidate basis function library according to the solidification temperature range to obtain the activated candidate basis function group. Step S3: Input the high-confidence feature subset into the activation candidate basis function group for fusion prediction and generate physical residuals, and then perform negative feedback iteration to obtain the yield prediction result.
[0006] As a preferred embodiment of the big data-based aluminum alloy production and processing output prediction method of the present invention, step S1 specifically includes: Step S101: Obtain production data collected by the workshop IoT and the alloy grade of the current batch; Production data includes dynamic process parameters, melt chemical composition data, and ingot surface temperature field data; Using the batch ID as the association identifier, production data with different sampling frequencies are divided according to the casting stage based on the association identifier, thus obtaining the stage sensing data corresponding to each casting stage and forming a sensing dataset. Sliding window statistics were performed on the sensor data at each stage to obtain the sequence of window statistical values corresponding to each dynamic process parameter in each casting stage. The arithmetic mean and variance of the statistical value sequences of each window are calculated and used as the mean and variance of each dynamic process parameter in each casting stage, respectively, to obtain the statistical feature set. After identifying and removing outliers from the statistical feature set using the local outlier factor algorithm, Z-Score standardization is performed to obtain the standardized feature set.
[0007] As a preferred embodiment of the aluminum alloy production and processing output prediction method based on big data described in this invention, step S1 further includes: Step S102: Obtain the melt chemical composition data corresponding to the alloy grade and standardized feature set of the current batch; Using the alloy grade as an index, query the preset aluminum alloy phase diagram thermodynamic database to obtain the nominal composition, solidus temperature and liquidus temperature corresponding to the alloy grade, and use the solidus temperature and liquidus temperature as the initial solidus temperature and initial liquidus temperature. The deviation between the melt chemical composition data and the nominal composition in the preset aluminum alloy phase diagram thermodynamic database is determined, and the solidus temperature and liquidus temperature of the corresponding melt chemical composition data are obtained. Calculate the difference between the solidus temperature and the liquidus temperature to obtain the solidification temperature range corresponding to the melt chemical composition data; Based on the nominal composition corresponding to the alloy grade of the current batch, the CALPHAD method is used for forward modeling to obtain the CALPHAD standard thermodynamic equilibrium path under the nominal composition.
[0008] As a preferred embodiment of the big data-based aluminum alloy production and processing output prediction method of the present invention, step S2 specifically includes: Step S201 involves screening for consistency features based on statistical and physical directions, specifically including: Obtain standardized feature sets, solidification temperature ranges, and historical production labels; Calculate the statistical change direction of each feature in the standardized feature set relative to the historical output label; The correlation direction between each feature and historical output labels was calculated using the Pearson correlation coefficient, thus obtaining the statistical change direction of each feature; When the Pearson correlation coefficient is greater than 0, the corresponding feature is statistically positively correlated with the historical output label; When the Pearson correlation coefficient is less than 0, the corresponding feature is statistically negatively correlated with the historical output label; When the Pearson correlation coefficient is equal to 0, the corresponding feature and the historical output label have no statistically significant linear correlation and are marked as statistically neutral. By consulting the physical rule table based on the solidification temperature range, the direction of change of the physical theory of each characteristic can be obtained.
[0009] As a preferred embodiment of the aluminum alloy production and processing output prediction method based on big data described in this invention, step S201 further includes: If the statistical change direction of any feature in the standardized feature set is statistically neutral, or the physical theory change direction is physically uncorrelated, or the statistical change direction is different from the physical theory change direction, then the physical confidence of that feature in the k-th iteration is marked as low confidence, and that feature is removed from the current candidate basis function input queue. If the statistical change direction of any feature in the standardized feature set is the same as the change direction of the physical theory, then the physical confidence level of that feature in the k-th iteration is marked as high confidence, and that feature is retained. All features marked as high confidence are combined into a high confidence feature subset.
[0010] As a preferred embodiment of the aluminum alloy production and processing output prediction method based on big data described in this invention, step S2 further includes: Step S202 involves adaptively activating each candidate basis function in the candidate basis function library based on the solidification temperature range, specifically including: Pre-determine m heterogeneous candidate basis function libraries; The candidate basis function library includes basis functions for linear models and basis functions for nonlinear models; Each candidate basis function in the heterogeneous candidate basis function library has its corresponding best-performing solidification interval center value and effective bandwidth recorded in historical verification. Obtain the solidification temperature range for the current batch; Calculate the process distance between the solidification temperature range of the current batch and the center value of the solidification range with the best performance of each candidate basis function. When the process distance is greater than the effective bandwidth, the participation weight of the candidate basis function is reset to 0, and the candidate basis function does not participate in this round of prediction; When the process distance is less than or equal to the effective bandwidth, the basic participation weights of the candidate basis function are calculated. When the process distance is equal to 0, the basic participation weights of the candidate basis functions are: ; For all satisfied The basic participation weights of candidate basis functions that are less than or equal to the effective bandwidth are subjected to Softmax normalization to obtain the activated candidate basis function group and the normalized weights corresponding to each activated candidate basis function.
[0011] As a preferred embodiment of the aluminum alloy production and processing output prediction method based on big data described in this invention, step S3 specifically includes: Step S301, fusion prediction includes: The high-confidence feature subset is input into the activation candidate basis function set for solution, and the result is obtained. The fusion prediction output for round-iteration iterations is expressed as: ; in, Indicates the first The fusion prediction output of round iterations, Indicates the iteration round number. Indicates the candidate basis function number, and is the summation variable. Indicates the first The number of basis functions activated in the candidate basis function set during each round of iteration. Indicates the first One activation candidate basis function Indicates the first Normalized weights corresponding to each activation candidate basis function Indicates the first High-confidence feature subset in round iteration Indicates the first A set of activation candidate basis functions for high-confidence feature subsets The mapping output; The fusion predicted output is back-calculated using a physical inversion proxy model to obtain the theoretical solidification path corresponding to the fusion predicted output; Calculate the absolute value of the temperature difference between the theoretical solidification path and the CALPHAD standard thermodynamic equilibrium path at each corresponding time node, and take the maximum value among all absolute differences as the physical residual of the k-th iteration.
[0012] As a preferred embodiment of the big data-based aluminum alloy production and processing output prediction method of the present invention, step S3 further includes: Step S302, negative feedback iteration, specifically includes: Preset feature admission threshold; The preset feature admission threshold is updated based on the physical residual from the initial round, as expressed by: ; in, Indicates the first Feature admission threshold for each iteration Indicates the first The feature admission threshold updated in each iteration. Indicates the preset shrinkage coefficient. Indicates the first The physical residual of the round iteration, Represents the physical residual of the initial round. Indicates the first The ratio of the physical residual of the first iteration to the physical residual of the first iteration, with the upper limit of the ratio truncated to 10; And update the effective bandwidth of the m candidate basis functions, as expressed by: ; in, Indicates the first The candidate basis functions at the th... The effective bandwidth of the round iteration Indicates the first The candidate basis functions at the th... The effective bandwidth of each round This represents the preset minimum positive number, and m represents the total number of candidate basis functions;
[0013] Based on the The physical residuals of each iteration are determined using the iteration termination criterion.
[0014] As a preferred embodiment of the aluminum alloy production and processing output prediction method based on big data described in this invention, the iteration termination determination rule specifically includes: When the When the physical residual of the first iteration is less than or equal to the physical residual tolerance threshold, the iteration terminates and the second iteration is output. The fusion prediction output of round iterations is used as the final prediction result; When the When the number of iterations is greater than or equal to the maximum number of iterations, terminate the iteration and output the 1st iteration. The fusion prediction output of round iterations; When the th iteration in two consecutive rounds The physical residual of the round iteration is greater than or equal to the first iteration. When calculating the physical residual of the iteration, trigger exception protection, terminate the iteration, and output the first iteration. The wheel fusion predicts output and issues a physical mismatch warning signal; When the iteration terminates, the first... The fusion prediction output corresponding to each iteration is used as the average of the final prediction output, and the 1st iteration is used as the average of the fusion prediction output. The physical residuals corresponding to each iteration are used as the final physical residual values, which constitute the production prediction results.
[0015] A big data-based aluminum alloy production and processing output prediction system, which includes a calculation module, a screening module, and an iteration module; The calculation module performs calculations on the production data to obtain a standardized feature set and solidification temperature range; The screening module performs consistency feature screening on the standardized feature set based on statistical and physical directions to obtain a high-confidence feature subset, and adaptively activates each candidate basis function in the candidate basis function library according to the solidification temperature range to obtain an activated candidate basis function group. The iterative module inputs a high-confidence feature subset into the activation candidate basis function group for fusion prediction and generates physical residuals, and then performs negative feedback iteration to obtain the yield prediction results.
[0016] The beneficial effects of this invention are as follows: By using the solidification temperature range as a core variable in metallurgy throughout the entire prediction process, unlike conventional algorithms that rely solely on statistical correlation coefficients, this invention introduces a physical rule table indexed by the solidification temperature range during feature selection. It compares the statistical direction with the physical theoretical direction for each feature, transforming pure statistical selection into a dual-criteria selection based on statistics and physics, fundamentally eliminating misjudgments of the direction of physical abrupt changes. During fusion, unlike traditional ensemble learning that indiscriminately participates in all basis functions, this invention introduces process distance and effective bandwidth, enabling adaptive activation of the best-performing basis functions in the current solidification temperature range for prediction. This transforms static weighting into dynamic weighting driven by operating conditions, avoiding hard extrapolation of fixed structures. During iterative verification, unlike relying solely on statistical labels to calculate residuals, this invention introduces physical residuals based on the CALPHAD solidification path to drive feature admission thresholds and negative feedback updates of basis function bandwidth. This transforms statistical residual feedback into dual-parameter feedback driven by physical residuals, ensuring stable predictions when the solidification range crosses critical values. Attached Figure Description
[0017] Figure 1This is a flowchart illustrating the steps of a method for predicting aluminum alloy production and processing output based on big data, provided in one embodiment of the present invention.
[0018] Figure 2 This is a basic flowchart of a big data-based aluminum alloy production and processing output prediction system provided in one embodiment of the present invention.
[0019] Figure 3 This is a schematic diagram of a flowchart for screening statistical direction and physical direction consistency features, provided as an embodiment of the present invention. Detailed Implementation
[0020] The embodiments of this disclosure are described below with reference to the accompanying drawings. The terminology used in the Description of Embodiments section of this disclosure is for illustrative purposes only and is not intended to limit the scope of this disclosure.
[0021] The embodiments of this disclosure are described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure. Those skilled in the art will understand that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this disclosure are also applicable to similar technical problems.
[0022] The terms “first,” “second,” etc., used in this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of embodiments of this disclosure. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not expressly listed or inherent to those processes, methods, products, or apparatuses.
[0023] Example, refer to Figures 1-3 As an embodiment of the present invention, a method for predicting the output of aluminum alloy production and processing based on big data is provided, including the following steps: Step S1: Calculate the production data to obtain a standardized feature set and solidification temperature range; Step S2: Based on statistical and physical directions, perform consistency feature screening on the standardized feature set to obtain a high-confidence feature subset, and adaptively activate each candidate basis function in the candidate basis function library according to the solidification temperature range to obtain the activated candidate basis function group. Step S3: Input the high-confidence feature subset into the activation candidate basis function group for fusion prediction and generate physical residuals, and then perform negative feedback iteration to obtain the yield prediction result.
[0024] In one embodiment, the present invention provides a method for predicting the output of aluminum alloy production based on big data. First, a standardized feature set and solidification temperature range are calculated from production data collected by the Internet of Things (IoT) in the workshop. Then, the standardized feature set is screened for consistency based on statistical and physical directions to obtain a high-confidence feature subset. Next, each candidate basis function in the candidate basis function library is adaptively activated according to the solidification temperature range to obtain an activated candidate basis function group. Finally, the high-confidence feature subset is input into the activated candidate basis function group for fusion prediction, and after generating physical residuals, negative feedback iteration is performed to obtain the output prediction result. The method uses the solidification temperature range as the core physical variable throughout the entire prediction process. It uses statistical-physical dual criteria to screen and eliminate directional misjudgments of physical abrupt change points. It avoids hard extrapolation of fixed structures through dynamic weighted fusion driven by operating conditions. Furthermore, it uses the physical residual of the solidification path to drive the round-by-round negative feedback update of the feature admission threshold and basis function bandwidth. This effectively solves the problem of prediction instability across grades and operating conditions caused by the lack of solidification thermodynamic mechanism constraints, the inability of the prediction framework to adapt to operating conditions, and the lack of physical process consistency verification. It improves the accuracy, stability, and engineering applicability of aluminum alloy production and processing output prediction under different alloy grades and process conditions.
[0025] In the feature selection stage, unlike conventional algorithms that rely solely on statistical correlation coefficients, this invention introduces a physical rule table indexed by solidification temperature ranges. It compares the consistency between statistical and theoretical physical directions for each feature, transforming pure statistical selection into a dual-criteria selection based on both statistics and physics, thus eliminating directional misjudgments at physical abrupt change points. In the fusion prediction stage, unlike traditional ensemble learning that indiscriminately involves all basis functions, this invention introduces process distance and effective bandwidth. It adaptively activates the best-performing basis functions from the past based on the current solidification temperature range, transforming static weighting into condition-driven dynamic weighting, avoiding hard extrapolation when solidification characteristics deviate from the training set distribution. In the iterative verification stage… Unlike relying solely on statistical labels to calculate residuals, this invention introduces physical residuals based on the CALPHAD solidification path, which drive the dual-parameter negative feedback update of the feature admission threshold and basis function bandwidth. This transforms statistical residual feedback into a closed-loop correction driven by physical residuals, ensuring that the prediction remains stable when the solidification range crosses critical values. The above three stages form a progressive closed-loop prediction framework centered around the solidification temperature range, a core variable in metallurgy. The physical rule table provides physically reliable feature inputs for the fusion prediction, dynamic weighted fusion ensures the adaptability of the prediction structure to the given input under operating conditions, and physical residual feedback corrects the accumulated deviations of the first two stages in each round. The three stages work together to achieve a deep fusion of physical constraints and data-driven approaches.
[0026] Step S1 specifically includes: Step S101: Obtain production data collected by the workshop IoT and the alloy grade of the current batch; Production data includes dynamic process parameters, melt chemical composition data, and ingot surface temperature field data; Using the batch ID as the association identifier, production data with different sampling frequencies are divided according to the casting stage based on the association identifier, thus obtaining the stage sensing data corresponding to each casting stage and forming a sensing dataset. Sliding window statistics were performed on the sensor data at each stage to obtain the sequence of window statistical values corresponding to each dynamic process parameter in each casting stage. The arithmetic mean and variance of the statistical value sequences of each window are calculated and used as the mean and variance of each dynamic process parameter in each casting stage, respectively, to obtain the statistical feature set. After identifying and removing outliers from the statistical feature set using the local outlier factor algorithm, Z-Score standardization is performed to obtain the standardized feature set.
[0027] Step S1 also includes: Step S102: Obtain the melt chemical composition data corresponding to the alloy grade and standardized feature set of the current batch; Using the alloy grade as an index, query the preset aluminum alloy phase diagram thermodynamic database to obtain the nominal composition, solidus temperature and liquidus temperature corresponding to the alloy grade, and use the solidus temperature and liquidus temperature as the initial solidus temperature and initial liquidus temperature. The deviation between the melt chemical composition data and the nominal composition in the preset aluminum alloy phase diagram thermodynamic database is determined, and the solidus temperature and liquidus temperature of the corresponding melt chemical composition data are obtained. Calculate the difference between the solidus temperature and the liquidus temperature to obtain the solidification temperature range corresponding to the melt chemical composition data; Based on the nominal composition corresponding to the alloy grade of the current batch, the CALPHAD method is used for forward modeling to obtain the CALPHAD standard thermodynamic equilibrium path under the nominal composition.
[0028] In one embodiment, step S1 specifically includes: step S101, acquiring production data collected by the workshop IoT and the alloy grade of the current batch; the production data includes dynamic process parameters, melt chemical composition data and ingot surface temperature field data; Among them, the surface temperature field data of the ingot is collected and recorded by an infrared thermometer; the dynamic process parameters are obtained by the casting machine PLC, including casting speed, cooling water flow rate, pouring temperature and casting time; the chemical composition data of the melt is determined by a spectrometer, specifically the mass percentage of Si, Fe, Cu, Mn and Mg elements. Using the batch ID as the association identifier, production data with different sampling frequencies are divided according to the casting stage based on the association identifier, thus obtaining the stage sensing data corresponding to each casting stage and forming a sensing dataset. Each batch corresponds to a specific alloy grade. The casting process includes a start-up phase, a steady-state phase, and a finish-up phase. The start-up phase involves the casting speed gradually increasing from 0 after the casting machine starts, until the casting speed first exceeds or equals 95% of the set value. The steady-state phase involves calculating the difference between the maximum and minimum values of the casting speed within a 10-second time window when the casting speed exceeds or equals 95% of the set value. This difference is used as the speed fluctuation amplitude. The 10-second time window is determined by taking the upper limit of the main period of the casting speed fluctuation (5-8 seconds according to spectral analysis) and adding a 1.2-fold margin. If the speed fluctuation amplitude is less than or equal to ±5% of the set value within three consecutive 10-second time windows, the process is considered to have entered the steady-state phase. The three windows are defined as a 99.2% probability of steady-state operation when the fluctuation meets the conditions within three consecutive windows. The finish-up phase begins after the casting speed decreases by more than 5% of the set value and continues until the casting machine completely stops, reducing the casting speed to 0. The set value is the target casting speed given in the casting process specification. Sliding window statistics were performed on the sensor data at each stage to obtain the sequence of window statistical values corresponding to each dynamic process parameter in each casting stage. The sliding window statistical process specifically includes dividing the stage sensor data into multiple continuous time segments with a window length of 60 seconds. Each time segment slides along the time axis with a step size of 30 seconds. The statistical values of each dynamic process parameter within each time segment are calculated to obtain the window statistical value sequence corresponding to each dynamic process parameter in each casting stage. The window length and step size are set according to the duration of the casting stage and the sensor sampling frequency. The window length is set to 60 seconds. For example, when the sensor sampling frequency is 5Hz, 300 sampling points can be collected in this window, which is greater than the minimum sample size (30) required for statistical feature calculation, ensuring the statistical significance of the mean and variance. The higher the sampling frequency, the more sampling points are in the window, further enhancing the statistical reliability. The step size is set to 30 seconds, which is half of the window length, so that there is a 50% overlap between adjacent windows, maintaining the sample size while ensuring the smoothness of the feature sequence. The arithmetic mean and variance of the statistical value sequences of each window are calculated and used as the mean and variance of each dynamic process parameter in each casting stage, respectively, to obtain the statistical feature set; wherein, the statistical feature set corresponds one-to-one with the batch ID; After identifying and removing outliers from the statistical feature set using the local outlier factor algorithm, Z-Score standardization is performed to obtain the standardized feature set. The specific process includes: Before performing Z-Score standardization, features with zero variance in the statistical feature set after removing outliers are first identified. Specifically, this involves using the Local Outlier Factor algorithm to identify and remove outliers from the statistical feature set. The neighborhood parameter k of the Local Outlier Factor algorithm is set to 20, which is determined by rounding down the square root of the total sample size (usually greater than 500 batches), resulting in the statistical feature set after removing outliers. Features with zero variance in this statistical feature set are then removed, resulting in a non-zero variance statistical feature set. Z-Score standardization is then performed on the non-zero variance statistical feature set to obtain the standardized feature set. If a feature has zero variance in the full sample, it is removed before Z-Score standardization. Step S1 also includes: Step S102, obtaining the melt chemical composition data corresponding to the alloy grade and standardized feature set of the current batch; using the alloy grade as an index, querying the preset aluminum alloy phase diagram thermodynamic database to obtain the nominal composition, solidus temperature and liquidus temperature corresponding to the alloy grade, and using the solidus temperature and liquidus temperature as the initial solidus temperature and initial liquidus temperature. Among them, the preset aluminum alloy phase diagram thermodynamic database is a database that is pre-constructed and stored in the local storage medium based on the CALPHAD method; the database is indexed by aluminum alloy grade and stores the solidus temperature and liquidus temperature corresponding to the nominal composition of each alloy grade. The nominal composition refers to the nominal value of the content range of each element specified in the national standard GB / T 3190 "Chemical Composition of Wrought Aluminum and Aluminum Alloys" for the alloy grade. The deviation between the melt chemical composition data and the nominal composition in the preset aluminum alloy phase diagram thermodynamic database is determined, and the solidus temperature and liquidus temperature of the corresponding melt chemical composition data are obtained. Deviation determination includes: The absolute deviation between the mass percentage of each element in the melt chemical composition and the corresponding mass percentage in the nominal composition is calculated one by one. If the absolute deviation of any element exceeds the composition deviation tolerance threshold, it is determined that a deviation exists. A linear interpolation method is used to correct the initial solidus temperature and initial liquidus temperature using the melt chemical composition data as interpolation nodes, thus obtaining the solidus temperature and liquidus temperature of the corresponding melt chemical composition data. The composition deviation tolerance threshold is 0.1% by mass percentage. This value follows the industry standard value in GB / T 3190 and the field of aluminum alloy composition control, and is used to distinguish between normal composition fluctuations and significant deviations that require correction of solidus / liquidus temperatures. If the absolute deviation of all elements does not exceed the composition deviation tolerance threshold, it is determined that there is no deviation, and the initial solidus temperature and initial liquidus temperature are directly used as the solidus temperature and liquidus temperature of the corresponding melt chemical composition data. Calculate the difference between the solidus temperature and the liquidus temperature to obtain the solidification temperature range corresponding to the melt chemical composition data; based on the nominal composition corresponding to the alloy grade of the current batch, use the CALPHAD method to perform forward modeling to obtain the CALPHAD standard thermodynamic equilibrium path under the nominal composition. The forward modeling process of the CALPHAD method is as follows: taking the temperature range boundary of the nominal component of the current batch as 50°C above the liquidus temperature and 50°C below the solidus temperature, the phase transition data of the nominal component is calculated using the CALPHAD method, and the corresponding CALPHAD standard thermodynamic equilibrium path of temperature change over time is generated. The path is then discretized on the time axis at fixed intervals, with the time step set to 0.1 seconds. To address the issues in existing technologies, such as ambiguous production data stage divisions, lack of stage-specific feature extraction, and failure to incorporate the correlation between melt composition fluctuations and solidification thermodynamics into feature construction, this step clarifies the physical judgment boundaries and specific numerical judgment criteria for the three stages of startup, steady state, and termination. It introduces a standardized process of staged sliding window statistics, outlier removal, and zero-variance filtering. Combined with nominal composition deviation judgment and solidification temperature range calculation methods based on the CALPHAD phase diagram database, this process transforms raw sensor data into a standardized feature set carrying metallurgical physical information, thereby improving the physical reliability and cross-batch stability of production forecasts from the source.
[0029] Step S2 specifically includes: Step S201 involves screening for consistency features based on statistical and physical directions, specifically including: Obtain standardized feature sets, solidification temperature ranges, and historical production labels; Calculate the statistical change direction of each feature in the standardized feature set relative to the historical output label; The correlation direction between each feature and historical output labels was calculated using the Pearson correlation coefficient, thus obtaining the statistical change direction of each feature; When the Pearson correlation coefficient is greater than 0, the corresponding feature is statistically positively correlated with the historical output label; When the Pearson correlation coefficient is less than 0, the corresponding feature is statistically negatively correlated with the historical output label; When the Pearson correlation coefficient is equal to 0, the corresponding feature and the historical output label have no statistically significant linear correlation and are marked as statistically neutral. By consulting the physical rule table based on the solidification temperature range, the direction of change of the physical theory of each characteristic can be obtained.
[0030] Step S201 also includes: If the statistical change direction of any feature in the standardized feature set is statistically neutral, or the physical theory change direction is physically uncorrelated, or the statistical change direction is different from the physical theory change direction, then the physical confidence of that feature in the k-th iteration is marked as low confidence, and that feature is removed from the current candidate basis function input queue. If the statistical change direction of any feature in the standardized feature set is the same as the change direction of the physical theory, then the physical confidence level of that feature in the k-th iteration is marked as high confidence, and that feature is retained. All features marked as high confidence are combined into a high confidence feature subset.
[0031] Step S2 also includes: Step S202 involves adaptively activating each candidate basis function in the candidate basis function library based on the solidification temperature range, specifically including: Pre-determine m heterogeneous candidate basis function libraries; The candidate basis function library includes basis functions for linear models and basis functions for nonlinear models; Each candidate basis function in the heterogeneous candidate basis function library has its corresponding best-performing solidification interval center value and effective bandwidth recorded in historical verification. Obtain the solidification temperature range for the current batch; Calculate the process distance between the solidification temperature range of the current batch and the center value of the solidification range with the best performance of each candidate basis function. When the process distance is greater than the effective bandwidth, the participation weight of the candidate basis function is reset to 0, and the candidate basis function does not participate in this round of prediction; When the process distance is less than or equal to the effective bandwidth, the basic participation weights of the candidate basis function are calculated. When the process distance is equal to 0, the basic participation weights of the candidate basis functions are: ; For all satisfied The basic participation weights of candidate basis functions that are less than or equal to the effective bandwidth are subjected to Softmax normalization to obtain the activated candidate basis function group and the normalized weights corresponding to each activated candidate basis function.
[0032] In one embodiment, step S2 specifically includes: step S201, screening for consistency features based on statistical direction and physical direction, specifically including: Figure 3 This is a schematic diagram of the statistical direction and physical direction consistency feature screening process provided in one embodiment of the present invention. The process involves obtaining a standardized feature set, a solidification temperature range, and historical production labels; calculating the statistical change direction of each feature in the standardized feature set and the historical production labels one by one; and using the Pearson correlation coefficient to calculate the correlation direction between each feature and the historical production labels to obtain the statistical change direction of each feature. The Pearson correlation coefficient is calculated as follows: for any feature vector in the standardized feature set and the historical output label vector, calculate the covariance of the feature vector and the historical output label vector, the standard deviation of the feature vector and the standard deviation of the historical output label vector respectively, and divide the covariance by the product of the standard deviation of the feature vector and the standard deviation of the historical output label vector to obtain the Pearson correlation coefficient. When the Pearson correlation coefficient is greater than 0, the corresponding feature is statistically positively correlated with the historical production label; when the Pearson correlation coefficient is less than 0, the corresponding feature is statistically negatively correlated with the historical production label; when the Pearson correlation coefficient is equal to 0, the corresponding feature is statistically not significantly linearly correlated with the historical production label, and is marked as statistically neutral; the physical rule table is consulted according to the solidification temperature range to obtain the physical theoretical change direction of each feature; The physical rule table uses the range of solidification temperature as an index to store the physical theoretical change direction corresponding to each feature. The physical theoretical change direction includes physical positive correlation, physical negative correlation, and physical no obvious correlation. Among them, physical no obvious correlation means that based on the principles of solidification thermodynamics and metallurgy, this feature does not have a definite positive or negative physical effect on the yield within a given solidification temperature range. The physical rule table is pre-constructed based on the principles of solidification thermodynamics and metallurgy. The physical rule table is constructed by dividing the solidification temperature range into multiple continuous sub-ranges. The number of sub-ranges is determined according to the actual solidification temperature range to be covered. Each sub-range is continuous and non-overlapping. The boundaries of each sub-range are pre-set based on the solidification characteristics and thermodynamic principles of typical aluminum alloy grades. Based on the principles of solidification thermodynamics and metallurgy, the physical influence direction of each dynamic process parameter on the yield under different solidification temperature ranges is determined one by one. Using the solidification temperature range sub-ranges as indexes, the physical theoretical change directions corresponding to each feature are stored in the form of physical positive correlation, physical negative correlation, or no obvious physical correlation, thus forming the physical rule table. Specifically, taking a solidification temperature range of 50 degrees Celsius as an example, when the solidification temperature range is less than or equal to 50 degrees Celsius, the alloy solidification range is narrow, and increasing the casting speed helps to ensure complete mold filling, with a positive physical correlation. When the solidification temperature range is greater than 50 degrees Celsius, the alloy solidification range is wide, and increasing the casting speed will exacerbate the interdendritic tensile strain rate, leading to increased susceptibility to hot cracking, with a negative physical correlation. The physical direction of casting temperature is determined as follows: when the solidification temperature range is less than or equal to 50 degrees Celsius, there is a positive physical correlation, the principle being to improve fluidity; when the solidification temperature range is greater than 50 degrees Celsius, there is a negative physical correlation, the principle being to exacerbate shrinkage porosity. The physical direction of cooling water flow rate is positively correlated across the entire solidification temperature range, the principle being to refine grains and reduce segregation. However, there is a critical value for cooling water flow rate determined by both heat transfer conditions and ingot quality requirements. Beyond this critical value, the physical direction becomes insignificantly correlated. This critical value varies with the ingot size and alloy grade under different operating conditions. For example, the larger the ingot size, the higher the allowable critical value for cooling water flow rate, and vice versa. Exceeding this critical value results in excessive cooling capacity, which can lead to increased internal stress and segregation in the ingot, and no longer has a definite positive contribution to production increase. Step S201 further includes: if the statistical change direction of any feature in the standardized feature set is statistically neutral, or the physical theory change direction is physically uncorrelated, or the statistical change direction is different from the physical theory change direction, then the physical confidence of the feature in the k-th iteration is marked as low confidence, and the feature is removed from the current candidate basis function input queue. Specifically, the difference between the direction of statistical change and the direction of change in physical theory is manifested in the case that a positive statistical correlation corresponds to a negative physical correlation, or a negative statistical correlation corresponds to a positive physical correlation. If the statistical change direction of any feature in the standardized feature set is the same as the change direction of the physical theory, then the physical confidence level of that feature in the k-th iteration is marked as high confidence, and that feature is retained. The direction of statistical change is the same as the direction of change in physical theory. Specifically, a positive statistical correlation corresponds to a positive physical correlation, or a negative statistical correlation corresponds to a negative physical correlation, or a statistical neutrality corresponds to no significant physical correlation. All features marked as high confidence are combined into a high confidence feature subset; Step S202 involves adaptively activating each candidate basis function in the candidate basis function library according to the solidification temperature range. Specifically, this includes: pre-setting m heterogeneous candidate basis function libraries; where m is a positive integer greater than or equal to 2. The candidate basis function library includes basis functions for linear models and basis functions for nonlinear models; Among them, the basis functions of the linear model are linear regression basis functions, and the basis functions of the nonlinear model are random forest basis functions; The random forest uses CART regression trees as base learners, with mean squared error (MSE) as the splitting criterion. The number of decision trees is selected from {50, 100, 200} through ten-fold cross-validation. During training, Bootstrap sampling with replacement is used to sample the training set. The size of the input feature subset of each tree is set to 1 / 3 of the total number of features. Each decision tree grows independently and in parallel. The final prediction output is the arithmetic mean of the prediction results of all decision trees. Each candidate basis function in the heterogeneous candidate basis function library has its corresponding best-performing solidification interval center value and effective bandwidth recorded in historical verification. The effective bandwidth refers to the radius of the temperature range in which the candidate basis function maintains its effective predictive ability, centered on the center value of the solidification interval where the candidate basis function performs best. It is obtained by: evaluating each candidate basis function using 10-fold cross-validation during historical verification; traversing the prediction performance of each candidate basis function across different subsets of solidification temperature ranges; recording the center value of the solidification temperature range with the smallest prediction error as the center value of the solidification interval where the basis function performs best; and using this center value as a benchmark, recording the radius of the temperature range where the prediction error does not exceed 1.5 times the lowest prediction error as the effective bandwidth of the basis function. Historical verification refers to the process of evaluating each candidate basis function using 10-fold cross-validation on historical production batch data with labeled production volume. For each candidate basis function, traversing its prediction performance across different subsets of solidification temperature ranges; recording the center value of the solidification temperature range with the smallest prediction error as the center value of the solidification interval where the basis function performs best; and using this center value as a benchmark, recording the radius of the temperature range where the prediction error does not exceed 1.5 times the lowest prediction error as the effective bandwidth of the basis function. Obtain the solidification temperature range of the current batch; calculate the process distance between the solidification temperature range of the current batch and the center value of the solidification range with the best performance of each candidate basis function. The process distance is calculated as follows: the difference between the solidification temperature range and the center value of the solidification range is calculated, and the absolute value of this difference is the process distance; where the solidification temperature range is the range value, the center value of the solidification range is the point value, the arithmetic mean of the upper and lower limits of the solidification temperature range is taken as the midpoint value of the range, the difference between the midpoint value and the center value of the solidification range is calculated, and the absolute value of this difference is taken as the process distance. When the process distance is greater than the effective bandwidth, the participation weight of the candidate basis function is reset to 0, and the candidate basis function does not participate in this round of prediction; when the process distance is less than or equal to the effective bandwidth, the basic participation weight of the candidate basis function is calculated. The expression for calculating the fundamental participation weights of this candidate basis function is: ; in, This represents the basic participation weights of the candidate basis function. Indicates process distance, in degrees Celsius. This represents the first preset minimum positive number, in degrees Celsius. The value is 0.001℃, which is less than the usual order of magnitude of the process distance. Setting it to this value can prevent the denominator from being zero and causing calculation overflow when the process distance is 0. At the same time, when the process distance is greater than 0, the impact on the weight calculation result can be ignored. When the process distance is equal to 0, the basic participation weights of the candidate basis functions are: For all satisfying The basic participation weights of candidate basis functions that are less than or equal to the effective bandwidth are subjected to Softmax normalization to obtain the activation candidate basis function group and the normalized weights corresponding to each activation candidate basis function; the sum of the normalized weights of each activation candidate basis function is 1. To address the issues of lack of physical constraints in statistical screening and the inability of fixed prediction structures to adapt to operating conditions in existing technologies, this step introduces a physical rule table indexed by solidification temperature ranges for statistical-physical dual-criterion consistency screening, eliminating directional misjudgments at physical abrupt change points. Furthermore, it introduces process distance and effective bandwidth, enabling the adaptive activation of historically best-performing basis functions based on the current solidification temperature range to participate in prediction, transforming static weighting into operating condition-driven dynamic weighting. This provides physically reliable feature inputs and an operating condition-adaptive prediction structure for subsequent fusion prediction.
[0033] Step S3 specifically includes:
[0034] Step S301, fusion prediction includes:
[0035] The high-confidence feature subset is input into the activation candidate basis function set for solution, and the result is obtained. The fusion prediction output for round-iteration iterations is expressed as: ; in, Indicates the first The fusion prediction output of round iterations, Indicates the iteration round number. Indicates the candidate basis function number, and is the summation variable. Indicates the first The number of basis functions activated in the candidate basis function set during each round of iteration. Indicates the first One activation candidate basis function Indicates the first Normalized weights corresponding to each activation candidate basis function Indicates the first High-confidence feature subset in round iteration Indicates the first A set of activation candidate basis functions for high-confidence feature subsets The mapping output; The fusion predicted output is back-calculated using a physical inversion proxy model to obtain the theoretical solidification path corresponding to the fusion predicted output; Calculate the absolute value of the temperature difference between the theoretical solidification path and the CALPHAD standard thermodynamic equilibrium path at each corresponding time node, and take the maximum value among all absolute differences as the physical residual of the k-th iteration.
[0036] Step S3 also includes: Step S302, negative feedback iteration, specifically includes: Preset feature admission threshold; The preset feature admission threshold is updated based on the physical residual from the initial round, as expressed by: ; in, Indicates the first Feature admission threshold for each iteration Indicates the first The feature admission threshold updated in each iteration. Indicates the preset shrinkage coefficient. Indicates the first The physical residual of the round iteration, Represents the physical residual of the initial round. Indicates the first The ratio of the physical residual of the first iteration to the physical residual of the first iteration, with the upper limit of the ratio truncated to 10; And update the effective bandwidth of the m candidate basis functions, as expressed by: ; in, Indicates the first The candidate basis functions at the th... The effective bandwidth of the round iteration Indicates the first The candidate basis functions at the th... The effective bandwidth of each round This represents the preset minimum positive number, and m represents the total number of candidate basis functions; Based on the The physical residuals of each iteration are determined using the iteration termination criterion.
[0037] The specific rules for determining the termination of iterations include: When the When the physical residual of the first iteration is less than or equal to the physical residual tolerance threshold, the iteration terminates and the second iteration is output. The fusion prediction output of round iterations is used as the final prediction result; When the When the number of iterations is greater than or equal to the maximum number of iterations, terminate the iteration and output the 1st iteration. The fusion prediction output of round iterations; When the th iteration in two consecutive rounds The physical residual of the round iteration is greater than or equal to the first iteration. When calculating the physical residual of the iteration, trigger exception protection, terminate the iteration, and output the first iteration. The wheel fusion predicts output and issues a physical mismatch warning signal; When the iteration terminates, the first... The fusion prediction output corresponding to each iteration is used as the average of the final prediction output, and the 1st iteration is used as the average of the fusion prediction output. The physical residuals corresponding to each iteration are used as the final physical residual values, which constitute the production prediction results.
[0038] In one embodiment, step S3 specifically includes: step S301, the fusion prediction includes: inputting the high-confidence feature subset into the activation candidate basis function set for solving, to obtain the first... The fusion prediction output for round-iteration iterations is expressed as: ; in, Indicates the first The fusion prediction output of round iterations, This represents the iteration round number, k=0,1,2,…, Where k=0 is the initial round, Indicates the maximum number of iterations. =10. This value is set based on the statistical results of the amount of batch data and the iteration convergence speed in aluminum alloy production. Based on simulation experiments on historical batch data, the physical residual convergence under different iteration rounds is statistically analyzed. The experimental results show that more than 95% of batches can reduce the physical residual to below the physical residual tolerance threshold within 10 iterations. Therefore, setting the maximum number of iterations to 10 can control the computational cost while ensuring prediction accuracy. Indicates the candidate basis function number, and is the summation variable. Indicates the first The number of basis functions activated in the candidate basis function set during each round of iteration. Indicates the first One activation candidate basis function Indicates the first Normalized weights corresponding to each activation candidate basis function Indicates the first High-confidence feature subset in round iteration Indicates the first A set of activation candidate basis functions for high-confidence feature subsets The mapping output; The fusion predicted output is back-calculated using a physical inversion proxy model to obtain the theoretical solidification path corresponding to the fusion predicted output; The pre-construction method of the physical inversion proxy model specifically includes: coupling the solidification heat transfer model based on the CALPHAD method, performing forward modeling calculations of the solidification process for dynamic process parameter combinations, generating CALPHAD standard thermodynamic equilibrium path point sets and yield labels corresponding to temperature changes over time under each dynamic process parameter combination, forming a training dataset; discretizing the CALPHAD standard thermodynamic equilibrium path point set, extracting the temperature values at each time node as path parameter vectors; using the yield labels and path parameter vectors under the same dynamic process parameter combination as training samples, with the yield labels as input features and the path parameter vectors as prediction targets, and performing inverse mapping training using the Gaussian process regression method to obtain the mapping function from yield to solidification path parameters; wherein, the covariance function of the Gaussian process regression adopts the radial basis function (RBF), and the hyperparameters include signal variance and length scale, which are optimized by maximizing the log marginal likelihood; during the training process, for the input space after the yield labels are normalized... A Gaussian process prior is constructed, and the posterior predicted mean and variance are directly calculated through the closed-form solution of the Gaussian process regression without iterative training. During prediction, the predicted mean is used as the output of the path parameter vector, and the predicted variance is used to evaluate the uncertainty of path backpropagation. Within the production range covered by the training set, the solidification path is checked for thermodynamic phase transition order constraints, which include the liquidus temperature being greater than the solidus temperature and the cooling rate being always positive. If all training samples in a certain training round meet the phase transition order constraints, the verification is passed, and the current mapping function is output. If there are training samples that do not meet the phase transition order constraints, the samples are removed and retraining is performed. This process is repeated until no new samples do not meet the phase transition order constraints in two consecutive training rounds, or the preset maximum number of training rounds (200 rounds) is reached. The verified mapping function is used as a physical inversion surrogate model. Both the theoretical solidification path and the CALPHAD standard thermodynamic equilibrium path are sets of path points where the temperature changes over time, and both have been aligned to the same time node through linear interpolation.
[0039] Calculate the absolute value of the temperature difference between the theoretical solidification path and the CALPHAD standard thermodynamic equilibrium path at each corresponding time node, and take the maximum value among all absolute differences as the physical residual of the k-th iteration. The physical residual has the dimension of temperature. Step S302, negative feedback iteration, specifically includes: preset feature admission threshold; The initial value of the preset feature admission threshold is 1. This value is a dimensionless parameter and serves as the baseline value for the feature admission threshold. It should be noted that this initial value is not constant and will be adjusted in subsequent iterations according to the expression. The data is updated round by round. Those skilled in the art can also select other values as the initial values of the feature admission threshold based on the data characteristics and prediction accuracy requirements of the actual production line, such as 0.5 to 2.0. In the initial round, the consistency screening between the statistical direction and the physical direction is carried out according to the benchmark strictness. The preset feature admission threshold is updated based on the physical residual from the initial round, as expressed by: ; in, This represents the feature admission threshold for the k-th iteration. Indicates the first The updated feature admission threshold from each iteration is used for feature selection in the next iteration. The larger the value, the stricter the feature selection criteria. This represents the preset shrinkage coefficient, with a value range of [0.1, 0.5]. This represents the physical residual in the k-th iteration. Represents the physical residual of the initial round. This represents the ratio of the physical residual in the k-th iteration to the physical residual in the initial iteration. The upper limit of the ratio is truncated to 10. The truncation logic is as follows: If... ,but ;like If the physical residual is larger, the feature admission threshold increases cumulatively in each round according to the residual ratio, so that more features with questionable statistical directions are excluded in the next iteration; when the physical residual is smaller than the physical residual of the initial round, the increase of the feature admission threshold slows down, maintaining the stability of the screening. And update the effective bandwidth of the m candidate basis functions, as expressed by: ; in, Indicates the first The effective bandwidth of the candidate basis functions in the k-th iteration. Indicates the first The candidate basis functions at the th... The effective bandwidth of each round The preset minimum positive number is denoted as the second preset minimum positive number. The dimension of η is temperature, consistent with the dimension of the physical residual, and is set to 0.1℃. This value is set according to the conventional magnitude of the physical residual, which is usually 1 to 50℃. 0.1℃ is less than this conventional magnitude, which can prevent the denominator from being zero and causing bandwidth update overflow when the physical residual is zero. At the same time, when the physical residual is within the conventional magnitude range of 1 to 50℃, the impact on the bandwidth update result is negligible. m represents the total number of candidate basis functions. The dimension of the effective bandwidth is temperature, consistent with the dimension of the solidification temperature range. Based on the physical residual of the k-th iteration, the iteration termination determination rule is used for determination; The specific rules for determining the termination of an iteration include: when the... When the physical residual of the first iteration is less than or equal to the physical residual tolerance threshold, the iteration terminates and the second iteration is output. The fusion prediction output of round iterations is used as the final prediction result; The physical residual tolerance threshold is in the dimension of temperature and is set to 0.5℃. This value is set based on the statistical distribution of temperature measurement errors along the solidification path of aluminum alloys. The measurement error range of the infrared thermometer is ±0.3~±0.5℃, and the upper limit of 0.5℃ is taken as the tolerance threshold. When the physical residual is less than or equal to this value, the temperature deviation between the theoretical solidification path corresponding to the fusion prediction yield and the CALPHAD standard equilibrium path is already within the range of instrument measurement error, and further iteration cannot further improve the prediction accuracy. The physical residual tolerance threshold is in the dimension of temperature, which is consistent with the dimension of the physical residual. When the When the number of iterations is greater than or equal to the maximum number of iterations, terminate the iteration and output the 1st iteration. The fusion prediction output of the round of iterations; when the 1st round of iterations in two consecutive rounds... The physical residual of the round iteration is greater than or equal to the first iteration. When calculating the physical residual of the iteration, trigger exception protection, terminate the iteration, and output the first iteration. The fusion prediction yield of the round is used to issue a physical mismatch warning signal; when the iteration terminates, the first round will... The fusion prediction output corresponding to each iteration is used as the average of the final prediction output, and the 1st iteration is used as the average of the fusion prediction output. The physical residuals corresponding to each iteration are used as the final physical residual values to form the production prediction results; To address the issue that existing residual verification methods only address statistical errors in production labels and lack physical consistency criteria for solidification paths, this step introduces physical residuals based on CALPHAD solidification paths. These residuals drive the dual-parameter negative feedback update of feature admission thresholds and basis function bandwidth, achieving closed-loop correction driven by physical residuals. This ensures that predictions remain stable when crossing critical values in the solidification interval, avoids high-confidence erroneous outputs in physically unreliable regions, and effectively improves the physical consistency of prediction results.
[0040] Example 2, refer to Figure 2 In another embodiment of the present invention, which differs from the first embodiment, a big data-based aluminum alloy production and processing output prediction system is provided, including a calculation module, a screening module, and an iteration module. The calculation module performs calculations on the production data to obtain a standardized feature set and solidification temperature range; The screening module performs consistency feature screening on the standardized feature set based on statistical and physical directions to obtain a high-confidence feature subset, and adaptively activates each candidate basis function in the candidate basis function library according to the solidification temperature range to obtain an activated candidate basis function group. The iterative module inputs a high-confidence feature subset into the activation candidate basis function group for fusion prediction and generates physical residuals, and then performs negative feedback iteration to obtain the yield prediction results.
[0041] In one embodiment, the present invention provides a method for predicting the output of aluminum alloy production based on big data. First, a standardized feature set and solidification temperature range are calculated from production data collected by the Internet of Things (IoT) in the workshop. Then, the standardized feature set is screened for consistency based on statistical and physical directions to obtain a high-confidence feature subset. Next, each candidate basis function in the candidate basis function library is adaptively activated according to the solidification temperature range to obtain an activated candidate basis function group. Finally, the high-confidence feature subset is input into the activated candidate basis function group for fusion prediction, and after generating physical residuals, negative feedback iteration is performed to obtain the output prediction result. The method uses the solidification temperature range as the core physical variable throughout the entire prediction process. It uses statistical-physical dual criteria to screen and eliminate directional misjudgments of physical abrupt change points. It avoids hard extrapolation of fixed structures through dynamic weighted fusion driven by operating conditions. Furthermore, it uses the physical residual of the solidification path to drive the round-by-round negative feedback update of the feature admission threshold and basis function bandwidth. This effectively solves the problem of prediction instability across grades and operating conditions caused by the lack of solidification thermodynamic mechanism constraints, the inability of the prediction framework to adapt to operating conditions, and the lack of physical process consistency verification. It improves the accuracy, stability, and engineering applicability of aluminum alloy production and processing output prediction under different alloy grades and process conditions.
[0042] A comparative experiment on yield prediction was conducted according to the present invention. The experimental process and results are as follows.
[0043] Experimental Procedure: The experiment used historical data from 320 batches of continuous production from a direct-cooling semi-continuous casting line for aluminum alloys, covering six commonly used alloy grades (including 5083, 6061, and 7075). Each batch recorded complete dynamic process parameters (casting speed, cooling water flow rate, pouring temperature, and casting time) during the start-up, steady-state, and finish-up stages, as well as melt chemical composition (mass percentage of Si, Fe, Cu, Mn, and Mg) and ingot surface temperature field data. The dataset was randomly divided into a training set and a test set in a 7:3 ratio, with two additional batches of alloy grades not used in the training set reserved as a cross-grade generalization validation set.
[0044] The experiment is set up with four comparison schemes, all based on the same standardized feature set from step S1 as input: Scheme 1 (pure statistical screening): Feature screening is performed solely based on the Pearson correlation coefficient, without introducing physical rule table constraints, without performing statistical-physical dual-criterion consistency screening, and without performing negative feedback iteration. Scheme 2 (fixed-weighted fusion): Based on statistical-physical dual-criterion screening, all candidate basis functions are assigned fixed weights to participate in fusion prediction, without performing adaptive activation based on the effective bandwidth of process distance. Scheme 3 (no physical residual feedback): Based on statistical-physical dual-criterion screening and adaptive activation based on process distance, the mean square error of the production label is used as the residual feedback index, without constructing a physical inversion surrogate model to calculate solidification path deviation, i.e., without introducing physical residual-driven negative feedback iteration. This scheme: executes the entire process from S1 to S3, including statistical-physical dual-criterion consistency feature screening, adaptive activation based on process distance and effective bandwidth, physical residual calculation based on CALPHAD solidification path, and dual-parameter negative feedback iteration driven by physical residuals: feature admission threshold and effective bandwidth of basis functions.
[0045] In each scheme, the candidate basis function library is fixed to include two heterogeneous models: linear regression and random forest. The maximum number of iterations is set to 10, the physical residual tolerance threshold is set to 0.5℃, the initial value of the feature admission threshold is set to 1, the shrinkage coefficient is 0.3, the first preset minimum positive number is 0.001℃, and the second preset minimum positive number is 0.1℃. Each scheme is run repeatedly 10 times under the same hardware and software environment, and the average value of the evaluation index is taken as the final result. See Table 1 for a detailed comparison. Table 1. Comparison of yield forecasting performance under different forecasting schemes:
[0046] As shown in Table 1, Scheme 1, lacking constraints from the thermodynamic mechanism of solidification, relies solely on statistical correlation coefficients for feature selection. This leads to misjudgments where the statistical and physical directions are inconsistent when crossing critical values in the solidification temperature range. The physical direction misjudgment rate reaches 24.6%, with a determination coefficient R² of only 0.81. Scheme 2, while introducing physical rules to eliminate some directional misjudgments, suffers from a prediction deviation of 2.56t across grades because the candidate basis functions cannot adaptively adjust to the current solidification temperature range. Scheme 3 shows improved accuracy under normal operating conditions compared to the previous two, but lacks physical consistency verification based on the solidification path. The physical residuals cannot drive feature selection and basis function bandwidth correction round by round, resulting in a misjudgment rate across grades. The prediction error still reached 2.03t, and the accuracy rate of anomaly warning was only 72.0%. This solution eliminates directional misjudgment at physical abrupt change points through statistical-physical dual-criteria screening. It achieves adaptive activation of the basis function driven by the working condition through process distance and effective bandwidth. It corrects the deviation round by round through negative feedback iteration of the feature admission threshold and basis function bandwidth driven by physical residuals based on CALPHAD solidification path. It achieves the best results in terms of the accuracy indicators of the coefficient of determination and root mean square error. The prediction error across grades is controlled at 1.14t, the physical direction misjudgment rate is reduced to 4.2%, and the accuracy rate of working condition anomaly warning is improved to 91.0%, effectively ensuring the prediction reliability and stability across grades and working conditions.
[0047] This invention uses the solidification temperature range as a core metallurgical variable throughout the entire prediction process. In feature selection, unlike conventional algorithms that rely solely on statistical correlation coefficients, it introduces a physical rule table indexed by the solidification temperature range. It compares the statistical direction with the theoretical physical direction for each feature, transforming pure statistical selection into a dual-criteria selection based on statistics and physics, fundamentally eliminating misjudgments of the direction of physical abrupt changes. In fusion, unlike traditional ensemble learning that indiscriminately involves all basis functions, it introduces process distance and effective bandwidth, enabling adaptive activation of historically best-performing basis functions based on the current solidification temperature range. This transforms static weighting into dynamic weighting driven by operating conditions, avoiding hard extrapolation of fixed structures. In iterative verification, unlike relying solely on statistical labels to calculate residuals, it introduces physical residuals based on the CALPHAD solidification path to drive feature admission thresholds and negative feedback updates of basis function bandwidth. This transforms statistical residual feedback into dual-parameter feedback driven by physical residuals, ensuring prediction stability when the solidification range crosses critical values.
[0048] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0049] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the protection scope of the present invention.
Claims
1. A method for predicting aluminum alloy production and processing output based on big data, characterized in that, Includes the following steps: Step S1: Calculate the production data to obtain a standardized feature set and solidification temperature range; Specifically, step S1 includes: Step S101: Obtain production data collected by the workshop IoT and the alloy grade of the current batch; Production data includes dynamic process parameters, melt chemical composition data, and ingot surface temperature field data; Using the batch ID as the association identifier, production data with different sampling frequencies are divided according to the casting stage based on the association identifier, thus obtaining the stage sensing data corresponding to each casting stage and forming a sensing dataset. Sliding window statistics were performed on the sensor data at each stage to obtain the sequence of window statistical values corresponding to each dynamic process parameter in each casting stage. The arithmetic mean and variance of the statistical value sequences of each window are calculated and used as the mean and variance of each dynamic process parameter in each casting stage, respectively, to obtain the statistical feature set. After identifying and removing outliers from the statistical feature set using the local outlier factor algorithm, Z-Score standardization is performed to obtain the standardized feature set. Step S1 also includes: Step S102: Obtain the melt chemical composition data corresponding to the alloy grade and standardized feature set of the current batch; Using the alloy grade as an index, query the preset aluminum alloy phase diagram thermodynamic database to obtain the nominal composition, solidus temperature and liquidus temperature corresponding to the alloy grade, and use the solidus temperature and liquidus temperature as the initial solidus temperature and initial liquidus temperature. The deviation between the melt chemical composition data and the nominal composition in the preset aluminum alloy phase diagram thermodynamic database is determined, and the solidus temperature and liquidus temperature of the corresponding melt chemical composition data are obtained. Calculate the difference between the solidus temperature and the liquidus temperature to obtain the solidification temperature range corresponding to the melt chemical composition data; Step S2: Based on statistical and physical directions, perform consistency feature screening on the standardized feature set to obtain a high-confidence feature subset, and adaptively activate each candidate basis function in the candidate basis function library according to the solidification temperature range to obtain the activated candidate basis function group. Among them, the statistical direction is the statistical change direction obtained by calculating the standardized feature set and the historical production label using the Pearson correlation coefficient, and the physical direction is the physical theoretical change direction obtained by querying the physical rule table according to the solidification temperature range. Features with the same statistical change direction and physical theoretical change direction are marked as high confidence and retained, and the remaining features are removed from the current candidate basis function input queue. Step S3: Input the high-confidence feature subset into the activation candidate basis function group for fusion prediction and generate physical residuals, and then perform negative feedback iteration to obtain the yield prediction result.
2. The method for predicting aluminum alloy production and processing output based on big data as described in claim 1, characterized in that: Step S1 also includes: Based on the nominal composition corresponding to the alloy grade of the current batch, the CALPHAD method is used for forward modeling to obtain the CALPHAD standard thermodynamic equilibrium path under the nominal composition.
3. The method for predicting aluminum alloy production and processing output based on big data as described in claim 2, characterized in that: Step S2 specifically includes: Step S201 involves screening for consistency features based on statistical and physical directions, specifically including: Obtain standardized feature sets, solidification temperature ranges, and historical production labels; Calculate the statistical change direction of each feature in the standardized feature set relative to the historical output label; The correlation direction between each feature and historical output labels was calculated using the Pearson correlation coefficient, thus obtaining the statistical change direction of each feature; When the Pearson correlation coefficient is greater than 0, the corresponding feature is statistically positively correlated with the historical output label; When the Pearson correlation coefficient is less than 0, the corresponding feature is statistically negatively correlated with the historical output label; When the Pearson correlation coefficient is equal to 0, the corresponding feature and the historical output label have no statistically significant linear correlation and are marked as statistically neutral. By consulting the physical rule table based on the solidification temperature range, the direction of change of the physical theory of each characteristic can be obtained.
4. The method for predicting aluminum alloy production and processing output based on big data as described in claim 3, characterized in that: Step S201 also includes: If the statistical change direction of any feature in the standardized feature set is statistically neutral, or the physical theory change direction is physically uncorrelated, or the statistical change direction is different from the physical theory change direction, then the physical confidence of that feature in the k-th iteration is marked as low confidence, and that feature is removed from the current candidate basis function input queue. If the statistical change direction of any feature in the standardized feature set is the same as the change direction of the physical theory, then the physical confidence level of that feature in the k-th iteration is marked as high confidence, and that feature is retained. All features marked as high confidence are combined into a high confidence feature subset.
5. The method for predicting aluminum alloy production and processing output based on big data as described in claim 4, characterized in that: Step S2 also includes: Step S202 involves adaptively activating each candidate basis function in the candidate basis function library based on the solidification temperature range, specifically including: Pre-determine m heterogeneous candidate basis function libraries; The candidate basis function library includes basis functions for linear models and basis functions for nonlinear models; Each candidate basis function in the heterogeneous candidate basis function library has its corresponding best-performing solidification interval center value and effective bandwidth recorded in historical verification. Obtain the solidification temperature range for the current batch; Calculate the process distance between the solidification temperature range of the current batch and the center value of the solidification range with the best performance of each candidate basis function. When the process distance is greater than the effective bandwidth, the participation weight of the candidate basis function is reset to 0, and the candidate basis function does not participate in this round of prediction; When the process distance is less than or equal to the effective bandwidth, the basic participation weights of the candidate basis function are calculated. When the process distance is equal to 0, the basic participation weights of the candidate basis functions are: ; For all satisfied The basic participation weights of candidate basis functions that are less than or equal to the effective bandwidth are subjected to Softmax normalization to obtain the activated candidate basis function group and the normalized weights corresponding to each activated candidate basis function.
6. The method for predicting aluminum alloy production and processing output based on big data as described in claim 5, characterized in that: Step S3 specifically includes: Step S301, fusion prediction includes: The high-confidence feature subset is input into the activation candidate basis function set for solution, and the result is obtained. The fusion prediction output for round-iteration iterations is expressed as: ; in, Indicates the first The fusion prediction output of round iterations, Indicates the iteration round number. Represents the candidate basis function number, and is the summation variable. Indicates the first The number of basis functions activated in the candidate basis function set during each round of iteration. Indicates the first 100 candidate basis functions for activation Indicates the first Normalized weights corresponding to each activation candidate basis function Indicates the first High-confidence feature subset in round iteration Indicates the first A set of activation candidate basis functions for high-confidence feature subsets The mapping output; The fusion predicted output is back-calculated using a physical inversion proxy model to obtain the theoretical solidification path corresponding to the fusion predicted output; Calculate the absolute value of the temperature difference between the theoretical solidification path and the CALPHAD standard thermodynamic equilibrium path at each corresponding time node, and take the maximum value among all absolute differences as the physical residual of the k-th iteration.
7. The method for predicting aluminum alloy production and processing output based on big data as described in claim 6, characterized in that: Step S3 also includes: Step S302, negative feedback iteration, specifically includes: Preset feature admission threshold; The preset feature admission threshold is updated based on the physical residual from the initial round, as expressed by: ; in, Indicates the first Feature admission threshold for each iteration Indicates the first The feature admission threshold updated in each iteration. Indicates the preset shrinkage coefficient. Indicates the first The physical residual of the round iteration, Represents the physical residual of the initial round. Indicates the first The ratio of the physical residual of the first iteration to the physical residual of the first iteration, with the upper limit of the ratio truncated to 10; And update the effective bandwidth of the m candidate basis functions, as expressed by: ; in, Indicates the first The candidate basis functions at the th... The effective bandwidth of the round iteration Indicates the first The candidate basis functions at the th... The effective bandwidth of each round This represents the preset minimum positive number, and m represents the total number of candidate basis functions; Based on the The physical residuals of each iteration are determined using the iteration termination criterion.
8. The method for predicting aluminum alloy production and processing output based on big data as described in claim 7, characterized in that: The specific rules for determining the termination of iterations include: When the When the physical residual of the first iteration is less than or equal to the physical residual tolerance threshold, the iteration terminates and the second iteration is output. The fusion prediction output of round iterations is used as the final prediction result; When the When the number of iterations is greater than or equal to the maximum number of iterations, terminate the iteration and output the 1st iteration. The fusion prediction output of round iterations; When the th iteration in two consecutive rounds The physical residual of the round iteration is greater than or equal to the first iteration. When calculating the physical residual of the iteration, trigger exception protection, terminate the iteration, and output the first iteration. The wheel fusion predicts output and issues a physical mismatch warning signal; When the iteration terminates, the th The fusion prediction output corresponding to each iteration is used as the average of the final prediction output, and the 1st iteration is used as the average of the fusion prediction output. The physical residuals corresponding to each iteration are used as the final physical residual values, which constitute the output prediction results.
9. A big data-based aluminum alloy production and processing output prediction system, applied to the big data-based aluminum alloy production and processing output prediction method as described in any one of claims 1-8, characterized in that, It includes a calculation module, a filtering module, and an iteration module; The calculation module performs calculations on the production data to obtain a standardized feature set and solidification temperature range; The screening module performs consistency feature screening on the standardized feature set based on statistical and physical directions to obtain a high-confidence feature subset, and adaptively activates each candidate basis function in the candidate basis function library according to the solidification temperature range to obtain an activated candidate basis function group. The iterative module inputs a high-confidence feature subset into the activation candidate basis function group for fusion prediction and generates physical residuals, and then performs negative feedback iteration to obtain the yield prediction results.