Extrusion casting process data correctness detection method fusing domain knowledge and integrated model

By integrating the LOF, IForest, and Boxplot models and combining them with knowledge of the squeeze casting field, the problem of inconsistent data quality in the squeeze casting process is solved, efficient, multi-dimensional data correctness detection is achieved, and the quality and efficiency of data-driven applications are improved.

CN120744730APending Publication Date: 2025-10-03GUANGXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411723723.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The multi-source and heterogeneous nature of existing extrusion casting process data leads to inconsistent data quality. The single anomaly detection method has insufficient generalization ability and is difficult to detect abnormal data that is contrary to domain knowledge, which affects the correctness and efficiency of the data-driven model.

Method used

By integrating the local outlier factor (LOF), the isolation forest algorithm based on the isolation idea (IForest), and the box plot model based on statistical analysis, combined with the knowledge of squeeze casting, the data correctness is tested through the value selection rules of material composition, process parameters and casting performance data, and the voting method is used to synthesize the test results.

Benefits of technology

The recall rate and accuracy of abnormal data detection are improved, multi-dimensional data quality detection is realized, the quality of data-driven applications is ensured, and high-quality data support is provided for the design of extrusion casting process parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744730A_ABST
    Figure CN120744730A_ABST
Patent Text Reader

Abstract

The invention provides a squeeze casting process data correctness detection method fusing domain knowledge and an integrated model, which belongs to the technical field of squeeze casting process data processing, and comprises the following steps: verifying material component values and judging a sample material series, performing correctness detection based on domain knowledge, performing detection based on an LBI integrated model, and outputting a corresponding result label. And completing the detection. According to the method, knowledge and data feature two-dimensional detection is achieved, the dimension of data anomaly (correctness) detection is expanded, a new reference view angle is provided for application such as data cleaning, advantage complementation of multiple data correctness detection methods is achieved through the integrated model, and the generalization ability of the methods is improved. Wherein the integrated model can be independently used for anomaly detection of other data sets, experiments show that the method is high in efficiency, the detection recall rate and accuracy of the abnormal data are improved, and the average recall rate of the abnormal data can reach 95%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of squeeze casting process data processing, and in particular to a method for detecting correctness of squeeze casting process data by integrating domain knowledge and an integrated model. Background Art

[0002] Currently, data-driven science is becoming the fourth paradigm after experimental science, theoretical deduction, and simulation. Data-driven methods have become the new foundation and direction for development in various fields. For example, in the field of materials, scholars are using previous experimental data to achieve rapid design of new materials, effectively solving the problems of long cycles and high costs in traditional experimental research models. In the manufacturing field, data-driven manufacturing has also become a path to achieving intelligent manufacturing. However, the performance of data-driven methods is heavily dependent on the quality of their driving data samples (correctness, consistency, low missingness, etc.). Poor data quality not only leads to erroneous knowledge acquisition, but also affects the correctness, learning effect, and effectiveness of data-driven models and methods. It may even cause serious consequences and huge losses. For example, a machine learning prediction model trained on a data sample with erroneous data may have prediction results that are too different from the actual results and cannot be applied.

[0003] Squeeze casting is an advanced material preparation and near-net-shape manufacturing process that combines the advantages of casting and forging. To adapt to the development trend of intelligent manufacturing, a data-driven process parameter design method based on existing process data is needed to address the high cost and low efficiency of current process parameter design methods that rely on trial-and-error methods, thereby improving the production efficiency and part quality of the squeeze casting process. The process data relied on by data-driven methods can come from physical experiments or numerical simulation test records, theoretical calculations, industrial production, and existing literature. However, due to differences in production environments, experimental plans, measurement methods, and errors in data collection, recording, and publishing, squeeze casting process data collected from multiple sources inevitably conflict with each other or have certain errors compared to the actual data. In other words, some data may be abnormal or incorrect. To ensure the quality of the constructed data-driven process parameter design model for squeeze casting, it is crucial to effectively verify the correctness of the dependent data and then clean or improve it.

[0004] To implement data-driven applications, data is often collected from various channels or generated using various data augmentation algorithms. However, the quality of this data varies greatly. Currently, researchers in various fields are increasingly recognizing the importance of data quality and are employing anomaly detection techniques based on statistical analysis, clustering, density-based methods, and nearest neighbor methods to evaluate data accuracy, identify and remove (clean) anomalous data, and thereby acquire accurate understanding, knowledge, and models of the object. Typical data anomaly (i.e., incorrect data) detection methods include the density-based Local Outlier Factor (LOF), the isolation forest algorithm (IForest) based on isolation principles, and the statistical analysis-based box plot. However, existing data anomaly detection methods typically employ a single anomaly detection method to detect data samples. However, these single anomaly detection methods can only obtain single-angle measurements on a dataset, lack generalization capabilities, and are limited in applicable scenarios. For example, while the IForest method requires no prior knowledge of the data and has low computational complexity, its design principle dictates that it is only applicable to datasets with low levels of anomalies. It is ineffective for datasets with a high proportion of anomalies, and its results are not intuitive or easy to interpret. The density-based LOF method, while offering interpretable results, is heavily dependent on parameter selection and has high computational complexity, making it ineffective for high-dimensional datasets. The Boxplot method, however, is nonparametric and applicable to any data distribution, providing intuitive and easily comparable results. However, it is limited by the amount of data available and often lacks precision when the data size is small. Secondly, single anomaly detection methods use the same anomaly evaluation criteria for all data, failing to comprehensively consider both global and local information, potentially leading to problems such as insufficient detection accuracy and inefficiency. Thirdly, existing detection methods primarily rely on data distribution characteristics to verify data correctness, rarely incorporating data-related engineering knowledge to aid detection. This makes it difficult to detect anomalies that contradict domain knowledge. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for detecting the correctness of squeeze casting process data by integrating domain knowledge and integrated models, so as to solve the technical problems mentioned in the background technology.

[0006] Aiming at the data-driven R&D needs of the squeeze casting process, the present invention proposes a correctness detection method for process data (mainly consisting of material composition data, process parameter data and casting performance data) that integrates LOF, Boxplot and IForest algorithms and integrates squeeze casting domain knowledge (LBI-Based Data Correctness Detection Method Incorporating Squeeze Casting Domain Knowledge, LBISCDK) based on the characteristics of the squeeze casting process to ensure the quality of the multi-source and heterogeneous process data it relies on. This method comprehensively considers the advantages and disadvantages of different anomaly detection models and integrates squeeze casting domain knowledge. It detects the correctness of squeeze casting process data from both the process data itself and domain knowledge perspectives, thereby improving the inspection capability and detection efficiency of abnormal data, providing high-quality data samples for subsequent data-driven applications, and ensuring the quality of data-driven applications.

[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0008] A method for detecting correctness of squeeze casting process data by integrating domain knowledge and an integrated model, the method comprising the following steps:

[0009] Step 1: Verify the material composition value and determine the sample material series. First, use the material composition attribute value selection rules to verify the material composition data of the sample to ensure the accuracy of the material composition data. For samples with reasonable and correct material composition data values, compare the mass fractions of other elements in the alloy composition except the main element and take the maximum value to determine the material series to which the data sample belongs.

[0010] Step 2: Based on domain knowledge, the correctness test is performed on the sample data based on the process data attribute value type and empirical value range of the corresponding alloy series. If the sample data attribute value type matches and the value is within the empirical value range, it is marked as a normal sample and proceeds to the next step of testing. Otherwise, it is marked as abnormal and the test ends.

[0011] Step 3: Based on the LBI integrated model detection, after the detection in step 2, delete the data samples that violate the attribute value selection rules, and then use the LBI integrated model to perform a second detection on the remaining data samples, output the corresponding result labels, and complete the detection.

[0012] Furthermore, the specific process of step 2 is:

[0013] Each attribute in the squeeze casting process data (mainly including process parameter data, process influencing factor data, and casting performance data) has a specific meaning, data type, value range, and quantitative calculation method. For example, the pouring temperature must be higher than the solid-liquid phase temperature of the material. For a specific sample, it is necessary to determine whether the data type and value range of its attribute value are consistent with the domain process knowledge. The value selection rules of the squeeze casting process data attribute values ​​are defined to detect the correctness of the data. The details are as follows:

[0014] Represent the squeeze casting process data as a triple<As,T,V> , where As represents the attribute, T represents its data type, V represents the data range, As={A1,A2,…,A j ,…,A m}; T=(t1,t2,…,t j ,…,t m ), t j Indicates the data type of the jth attribute; V = {v1, v2, ..., v j ,…,v m}, is the empirical value range of the jth attribute, and are the minimum and maximum values ​​of the empirical values, respectively. Given a in DT ij , we can judge whether it is an abnormal data point according to formula (1):

[0015]

[0016] In the formula, type(a ij ) means to get a ij The data type of a ij The data type type(a ij ) and its attribute's empirical value type t j Match and If the result is 0, it is considered normal data; otherwise, it is considered abnormal data and is recorded as 1.

[0017] The data type and empirical value range of the squeeze casting process data attribute value are obtained by referring to the squeeze casting process related manuals and literature. Based on the obtained squeeze casting process data attribute value value range and numerical type information, the correctness of the data sample is tested. The squeeze casting pouring temperature data type is set to floating point type, and the empirical value range is 50-100℃ higher than the alloy liquidus line. That is, its value rule is determined as follows:

[0018] T p ∈[Tl,Tl+100](2)

[0019] Furthermore, the specific process of step 3 is:

[0020] At the same time, the local outlier factor, the isolation forest method based on the isolation idea, and the box plot model based on statistical analysis are used for separate detection. The results are y1, y2, and y3 respectively. The outputs of y1, y2, and y3 are all 1 or 0, where 1 represents abnormality and 0 represents correctness. The voting method is then used to synthesize the results according to formula (5) to obtain the final detection result:

[0021]

[0022] Furthermore, the specific process of using local outlier factor detection is as follows:

[0023] Step 3.1.1: Input the squeeze casting process dataset DT, the number of neighbors k, and the abnormal sample threshold e;

[0024] Step 3.1.2: Normalize the DT process data and convert the attribute values ​​to the [0, 1] interval. Then, abstract the attributes of the squeeze casting process data into coordinate axes to form process data sample points in several dimensional spaces in the local outlier factor.

[0025] Step 3.1.3: For each data sample S i ∈DT, calculate sample S i The Euclidean distance to other samples is sorted from small to large to obtain the kth distance k_dist(S i );

[0026] Step 3.1.4: According to k_dist(S i ), will be less than k_dist(S i ) is included in S i k-distance neighborhood N k (S i ), then take S i To any other sample S i' The distance and k_dist(S i ) is the largest value in the sample S i reachdist k (S i , S i' );

[0027] Step 3.1.5: Calculate sample S i The local reachable density lrd(S i ), lrd(S i ) is equal to N k (S i ) to sample S i The reciprocal of the average reachable distance is S i ;

[0028] Step 3.1.6: Calculate sample S i N k (S i ) of all samples in lrd(S i ) and sumlrd k (S i );

[0029] Step 3.1.7: Calculate sample S i The local outlier factor score of

[0030]

[0031] Step 3.1.8: Mark abnormal data samples according to the abnormal sample threshold e;

[0032] Step 3.1.9: If lof(S i )>eThe sample S i Marked as an abnormal data sample, the label is recorded as 1, that is, y1 = 1, otherwise the sample S i Marked as a normal data sample, the label is recorded as 0, that is, y1 = 0;

[0033] Step 3.1.10: Loop through steps 3.1.3 to 3.1.9 to calculate the local outlier factor score y1 for each data sample and complete the correctness check for all samples.

[0034] Furthermore, the specific process of the isolation forest method detection based on the isolation idea is as follows:

[0035] Step 3.2: Training phase: Generate a specified number of isolated trees as follows:

[0036] Step 3.2.1: Input the process data set DT, set the current tree height h, and the isolated tree limit height h lim , number of isolated trees n_estimators, subsampling sample size is an initial value;

[0037] Step 3.2.2: Normalize the process data and convert the attribute values ​​into the [0,1] interval;

[0038] Step 3.2.3: If h>h lim or|DT|≤1 ends, otherwise randomly select an attribute A from the data set DT j and a split point p, dividing the data set DT into two subsets DT_left, attribute A j The value of is less than p and DT_right, attribute A j The value of is greater than p;

[0039] Loop until the limit height of the isolated tree reaches h lim , the number of isolated trees is equal to or exceeds n_estimators, and the isolated forest iTrees is obtained;

[0040] Step 3.3: Evaluation phase: Give the sample anomaly score s as follows:

[0041] First, calculate the number of edges from the root node of iTree to the external node, that is, the path length, recorded as h(S i ), for a given sample size of The sample subspace and a sample data S i , and its anomaly score s is defined as shown in formula (3):

[0042]

[0043] Where E[h(S i )] represents the sample data S i The average path length in n iTrees, Defined as the average path length of failed searches in a binary search tree, normalized h(S i ), which is defined as shown in formula (4):

[0044]

[0045] in, is the harmonic series, ξ is Euler’s constant;

[0046] The closer the anomaly score s of a data sample is to 1, the more likely it is an outlier, and the sample label is recorded as 1. If the anomaly score s of most samples is close to 0.5, it means that there are no obvious outliers in the entire data set. When s is much less than 0.5, it means that the data sample is a normal value, the label is recorded as 0, and the output is y2.

[0047] Furthermore, the specific process of separate detection of the box plot model based on statistical analysis is as follows:

[0048] The box plot is used to intuitively understand the central tendency, dispersion and outliers of the data. The extreme values ​​that are abnormally greater than the set value or abnormally less than the set value are identified as outliers. The lower quartile Q1, the upper quartile Q3, the value of Q1 at the 25% position, and the value of Q3 at the 75% position are used. The interquartile range is defined as IQR=Q3-Q1, the upper edge of the attribute is Q3+1.5IQR, and the lower edge is Q1-1.5IQR. The data exceeding the upper and lower edge intervals are identified as outliers and marked as 1. The rest of the correct data are marked as 0. The specific steps are as follows:

[0049] Step 3.3.1: Input process data set DT;

[0050] Step 3.3.2: Sort the data of each attribute in ascending order;

[0051] Step 3.3.3: Calculate the minimum value Min, lower quantile Q1, median, upper quantile Q3 and maximum value Max of each attribute;

[0052] Step 3.3.4: Calculate the interquartile range of the attribute IQR = Q3-Q1;

[0053] Step 3.3.5: Calculate the upper limit of the attribute: Up = Q3 + 1.5IQR, and the lower limit: Low = Q1 - 1.5IQR;

[0054] Step 3.3.6: If a ij >Up or a ij <Low, the data point a ij The sample is marked as abnormal and the label is recorded as 1. Otherwise, the data point a ij The sample is marked as normal and the label is 0;

[0055] Step 3.3.7: After the loop of steps 3.3.1 to 3.3.6 is completed, the label y3 indicating whether each data sample is abnormal or not is output.

[0056] The present invention has the following beneficial effects due to the adoption of the above technical solution:

[0057] The present invention realizes dual-dimensional detection of knowledge and data features, expands the dimension of data anomaly (correctness) detection, provides a new reference perspective for applications such as data cleaning, and realizes the complementary advantages of multiple data correctness detection methods through an integrated model, thereby improving the generalization ability of the method. The integrated model can be used alone for anomaly detection of other data sets. Experiments show that the proposed method is highly efficient and improves the recall rate and accuracy of abnormal data detection. The average recall rate of abnormal data can reach 95%. It provides a support tool for multi-source channel collection and automatic and efficient cleaning of extrusion casting process data, laying the foundation for building high-quality extrusion casting process data sets and data-driven extrusion casting process parameter design methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a flow chart of the method of the present invention;

[0059] Figure 2 It is a schematic diagram of LOF of the present invention;

[0060] Figure 3 This is a schematic diagram of abnormal data detection using IForest of the present invention;

[0061] Figure 4This is a graph showing the average detection recall rate, precision rate, and accuracy rate on four different abnormal data sets of the present invention. DETAILED DESCRIPTION

[0062] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and by way of preferred embodiments. However, it should be noted that many of the details listed in this specification are merely provided to help the reader gain a thorough understanding of one or more aspects of the present invention, and these aspects of the present invention can be practiced even without these specific details.

[0063] like Figure 1 As shown, a method for detecting the correctness of squeeze casting process data that integrates domain knowledge and an integrated model outputs a label indicating whether it is abnormal, where label 1 represents abnormality and label 0 represents correctness. The method includes the following steps:

[0064] Step 1: Construct the squeeze casting process data structure

[0065] The squeeze casting process data set is set as DT, with n samples and m attributes; it mainly includes process parameter data, process influencing factor data and casting performance data. The process parameter data mainly includes pouring temperature, extrusion pressure, mold preheating temperature, etc. The process influencing factor data mainly consists of factors that affect the process parameters, including material composition and casting shape characteristic parameters, etc. The casting performance data consists of the mechanical properties of different castings, such as hardness and tensile strength. The corresponding categories in the data set can be called attributes. i (i=1,2,…,n) represents the i-th data sample, A j Indicates the jth attribute, a ij represents the jth attribute data corresponding to the i-th data sample, then DT={a ij |i=1,2,…,n,j=1,2,...m}, and its structure is shown in Table 1.

[0066] Table 1 Squeeze casting process data set

[0067]

[0068] Table 2 shows an example of a data set. The ground truth data is a near-liquidus squeeze casting process data set of four different series of aluminum alloys (Al-Si, Al-Mg, Al-Cu, and Al-Mg alloys) extracted and compiled from journals and master's and doctoral dissertations related to squeeze casting research. Each sample contains four process parameter attributes (pouring temperature, extrusion pressure, holding time, and mold preheating temperature), five casting geometric feature attributes (extrusion method, shape complexity, maximum and minimum wall thickness, and material composition), and three quality index attributes (tensile strength, yield strength, and shrinkage porosity).

[0069] Table 2 Example of squeeze casting process data

[0070]

[0071] Step 2: Verify material composition values ​​and determine the sample material series: First, apply the material composition attribute value selection rules to verify the sample material composition data to ensure the accuracy of the material composition data. For samples with reasonable and correct material composition data values, the mass fractions of elements other than the main element (such as Al) in the alloy composition are compared and the maximum value is taken to determine the material series to which the data sample belongs. For example, if the mass fraction of Si in a data sample is the maximum value other than Al, the sample is determined to be squeeze casting process data for an Al-Si alloy, and the attribute value selection rules and integrated model for the Al-Si alloy series are automatically selected.

[0072] Step 3: Correctness check based on domain knowledge: The sample data is checked based on the process data attribute value type and empirical value range of the corresponding alloy series. If the sample data attribute value type matches and the value is within the empirical value range, it is marked as a normal sample and proceeds to the next step of testing; otherwise, it is marked as abnormal and the test ends.

[0073] Each attribute in squeeze casting process data has a specific meaning, data type, value range, and quantitative calculation method. For example, the pouring temperature must be higher than the solid-liquid phase temperature of the material. For a specific sample, it is necessary to determine whether the data type and value range of its attribute value are consistent with the domain process knowledge. Therefore, we define the value selection rules for squeeze casting process data attributes to check the correctness of the data. The specific rules are as follows:

[0074] Represent the squeeze casting process data as a triple<As,T,V> , where As represents the attribute, T represents its data type, V represents the data range, As={A1,A2,…,A j ,…,A m}; T=(t1,t2,…,t j ,…,t m ), t j Indicates the data type of the jth attribute; V = {v1, v2, ..., v j ,…,v m}, is the empirical value range of the jth attribute, and are the minimum and maximum values ​​of the empirical values ​​respectively. ij , we can judge whether it is an abnormal data point according to formula (1):

[0075]

[0076] In the formula, type(a ij ) means to get a ij The data type of a ij The data type type(a ij ) and its attribute's empirical value type t j Match and If it is normal data, the result is recorded as 0; otherwise, it is abnormal data and recorded as 1.

[0077] The data type and empirical value range of the squeeze casting process attribute values ​​can be obtained by referring to the squeeze casting process related manuals, literature, etc. For example, usually combined with the casting shape, according to the near-liquidus method, a temperature 50-100°C higher than the alloy liquidus temperature Tl is selected as the pouring temperature T of the squeeze casting. p , and is generally controlled at a lower value to reduce energy waste; therefore, the data type of the squeeze casting pouring temperature is generally set to floating point type, and its empirical value range is 50-100℃ higher than the alloy liquidus line, that is, the value rule is determined as follows:

[0078] T p ∈[Tl,Tl+100] (2)

[0079] For example, a correctness check of the data samples based on the value range and value type of the squeeze casting process data attributes revealed four abnormal samples (see Table 3, where the label 0 represents normal and 1 represents abnormal). The pouring temperature data for samples 67 and 90 in the Al-Si alloy dataset deviated from the empirical range for Al-Si alloys. Verification of the data source literature revealed that these were semi-solid squeeze casting processes, not near-liquidus squeeze casting. The pouring temperature data for sample 46 in the Al-Cu alloy dataset deviated from the empirical range, and the Al material composition value was also significantly abnormal. Verification of the data source literature revealed that this was near-liquidus squeeze casting process data for AZ91D magnesium alloy. The pouring temperature value for sample 37 in the Al-Mg alloy dataset deviated from the empirical range.

[0080] Table 3 Abnormal samples that violate attribute value rules

[0081]

[0082] Step 4: LBI ensemble model detection: After the second step of detection, remove the data samples that violate the attribute value rules. Then, use the LBI ensemble model to perform a second detection on the remaining data samples and output the corresponding result labels. Complete the detection.

[0083] At the same time, the local outlier factor (LOF), the isolation forest algorithm (IForest) based on the isolation idea, and the box plot model based on statistical analysis are used for separate detection. The results are y1, y2, and y3 respectively. The outputs of y1, y2, and y3 are all 1 or 0, where 1 represents abnormality and 0 represents correctness. The voting method is then used to synthesize the results according to formula (5) to obtain the final detection result.

[0084]

[0085] (1) Detection based on LOF model

[0086] The LOF model calculates the sample point S i Neighborhood point N k (S i ) and the average value of the local reachability density of point S i The ratio of the local reachable density to determine S i The greater the value, the more likely the sample point is to be abnormal, as shown in the following example. Figure 2 .

[0087] The steps for detecting the correctness of squeeze casting process data based on the LOF model are as follows:

[0088] Step 4.1.1: Input the squeeze casting process dataset DT, the number of neighbors k, and the abnormal sample threshold e

[0089] Step 4.1.2: Normalize the DT process data and convert the values ​​of each attribute to the interval [0, 1]. Then, abstract the attributes of the squeeze casting process data into coordinate axes to form process data sample points in the multidimensional space of LOF.

[0090] Step 4.1.3: For each data sample S i ∈DT, calculate sample S i The Euclidean distance to other samples is sorted from small to large to obtain the kth distance k_dist(S i ).

[0091] Step 4.1.4: According to k_dist(S i ), will be less than k_dist(S i ) is included in S i k-distance neighborhood N k (S i ); then take S i To any other sample S i' The distance and k_dist(S i) is the larger value of sample S i reachdist k (S i , S i' ).

[0092] Step 4.1.5: Calculate sample S i The local reachable density lrd(S i ). lrd(S i ) is equal to N k (S i ) to sample S i The reciprocal of the average reachable distance is S i .

[0093] Step 4.1.6: Calculate sample S i N k (S i ) of all samples in lrd(S i ) and sumlrd k (S i ).

[0094] Step 4.1.7: Calculate sample S i LOF score

[0095] Step 4.1.8: Mark abnormal data samples according to the abnormal sample threshold e

[0096] Step 4.1.9: If lof(S i )>eThe sample S i Marked as an abnormal data sample, the label is recorded as 1, that is, y1 = 1, otherwise the sample S i Marked as a normal data sample, the label is recorded as 0, that is, y1=0.

[0097] Step 10: Loop through steps 3 to 9 to calculate the LOF score y1 for each data sample and complete the correctness check for all samples.

[0098] (2) Detection based on IForest model

[0099] The correctness detection process of squeeze casting process data based on IForest is mainly divided into two stages: training and evaluation, as shown in the figure below: Figure 3 .

[0100] Training phase: Generate a specified number of isolated trees. The steps are as follows:

[0101] Step 4.2.1: Input the process data set DT, set the current tree height h, and the isolated tree limit height h lim, number of isolated trees n_estimators, subsampling sample size is an initial value.

[0102] Step 4.2.2: Normalize the process data and convert the values ​​of each attribute to the range [0,1];

[0103] Step 4.2.3: If h>h lim or|DT|≤1 ends, otherwise randomly select an attribute A from the data set DT j and a split point p, the data set DT is divided into two subsets DT_left (attribute A j The value of p is less than that of DT_right (attribute A j The value is greater than p) loop until the limit height of the isolated tree reaches h lim , the number of isolated trees is equal to or greater than n_estimators, and the isolated forest iTrees is obtained

[0104] Evaluation phase: gives the anomaly score s of the sample.

[0105] First, calculate the number of edges from the root node of iTree to the external node, that is, the path length, recorded as h(S i ), for a given sample size of The sample subspace and a sample data S i , and its anomaly score s is defined as shown in formula (3):

[0106]

[0107] Where E[h(S i )] represents the sample data S i The average path length among n iTrees; It is defined as the average path length of failed searches in a binary search tree, and is mainly used to normalize h(S i ), which is defined as shown in formula (4):

[0108]

[0109] in, is the harmonic series, and ξ is Euler's constant.

[0110] The closer the abnormal score s of a data sample is to 1, the more likely it is an outlier, and the sample label is recorded as 1. If the abnormal score s of most samples is close to 0.5, it means that there are no obvious outliers in the entire data set. When s is much smaller than 0.5, it means that the data sample is very likely to be a normal value, and the label is recorded as 0. ,y2(3) Detection based on the Boxplot model

[0111] Boxplots can be used to intuitively understand the central tendency, dispersion, and outliers of the data, and identify abnormally large or small extreme values ​​as outliers. Through the lower quartile Q1 (values ​​at the 25% position) and the upper quartile Q3 (values ​​at the 75% position), the interquartile range is defined as IQR = Q3 - Q1. The upper edge of each attribute is Q3 + 1.5 IQR, and the lower edge is Q1 - 1.5 IQR. Data exceeding the upper and lower edge intervals are identified as outliers and marked as 1, and the remaining correct data are marked as 0. The steps are as follows:

[0112] Step 4.3.1: Input: process data set DT;

[0113] Step 4.3.2: Sort the data of each attribute in ascending order

[0114] Step 4.3.3: Calculate the minimum value (Min), lower quantile (Q1, the value at the 25% position), median (Median), upper quantile (Q3, the value at the 75% position), and maximum value (Max) of each attribute;

[0115] Step 4.3.4: Calculate the interquartile range (IQR) of each attribute: Q3 - Q1.

[0116] Step 4.3.5: Calculate the upper limit Up = Q3 + 1.5IQR and the lower limit Low = Q1 - 1.5IQR for each attribute;

[0117] Step 4.3.6: If a ij >Up or a ij <Low, the data point a ij The sample is marked as abnormal and the label is recorded as 1. Otherwise, the data point a ij The sample is marked as normal and the label is 0.

[0118] Step 4.3.7: After the above steps are completed, the label y3 indicating whether each data sample is abnormal or not is output.

[0119] Table 4 shows the abnormal samples detected and identified by the present invention for the example data set in Table 2.

[0120] Table 4 Abnormal samples detected by LBISCDK method

[0121]

[0122] Table 5 shows the detection index results of the LBISCDK method that integrates domain engineering knowledge and the purely data-driven LBI method.

[0123] Table 5 Anomaly detection results

[0124]

[0125] Figure 4 The average recall, precision, and accuracy of several methods on four different anomaly datasets are shown before and after integrating domain knowledge. +SCDK indicates integration with domain knowledge, while -SCDK indicates non-integration with domain knowledge.

[0126] This method achieves dual-dimensional detection of knowledge and data features, expands the dimension of data anomaly (correctness) detection, and provides a new reference perspective for applications such as data cleaning. Through the integrated model, it realizes the complementary advantages of multiple data correctness detection methods and improves the generalization ability of the method. The integrated model can be used alone for anomaly detection in other datasets. Experiments show that the proposed method is highly efficient and improves the detection recall rate and accuracy of abnormal data. The average recall rate of abnormal data can reach 95%. It provides a supporting tool for the multi-source channel collection and automatic and efficient cleaning of extrusion casting process data, laying the foundation for the construction of high-quality extrusion casting process data sets and data-driven extrusion casting process parameter design methods. It also improves the generalization ability of the three methods.

[0127] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for detecting the correctness of squeeze casting process data by integrating domain knowledge and an integrated model, characterized by: The method comprises the following steps: Step 1: Verify the material composition value and determine the sample material series. First, use the material composition attribute value selection rules to verify the material composition data of the sample to ensure the accuracy of the material composition data. For samples with reasonable and correct material composition data values, compare the mass fractions of other elements in the alloy composition except the main element and take the maximum value to determine the material series to which the data sample belongs. Step 2: Based on domain knowledge, the correctness test is performed on the sample data based on the process data attribute value type and empirical value range of the corresponding alloy series. If the sample data attribute value type matches and the value is within the empirical value range, it is marked as a normal sample and proceeds to the next step of testing. Otherwise, it is marked as abnormal and the test ends. Step 3: Based on the LBI integrated model detection, after the detection in step 2, delete the data samples that violate the attribute value selection rules, and then use the LBI integrated model to perform a second detection on the remaining data samples, output the corresponding result labels, and complete the detection.

2. The method for detecting correctness of squeeze casting process data by integrating domain knowledge and an integrated model according to claim 1, characterized in that: The specific process of step 2 is: Squeeze casting process data includes process parameter data, process influencing factor data, and casting performance data. Each attribute has a specific meaning, data type, value range, and quantitative calculation method. For a specific sample, it is necessary to determine whether the data type and value range of its attribute value are consistent with the domain process knowledge. The value selection rules of the squeeze casting process data attribute values ​​are defined to detect the correctness of the data. The details are as follows: Represent the squeeze casting process data as a triple<As,T,V> , where As represents the attribute, T represents its data type, V represents the data range, As={A1,A2,…,A j ,…,A m };T=(t1,t2,…,tj,…,tm),t j Indicates the data type of the jth attribute; V = {v1, v2, ..., v j ,…,v m }, is the empirical value range of the jth attribute, and are the minimum and maximum values ​​of the empirical values, respectively. Given a in DT ij , we can judge whether it is an abnormal data point according to formula (1): In the formula, type(a ij ) means to get a ij The data type of a ij The data type type(a ij ) and its attribute's empirical value type t j Match and If the result is 0, it is considered normal data; otherwise, it is considered abnormal data and is recorded as 1. The data type and empirical value range of the squeeze casting process data attribute value are obtained by referring to the squeeze casting process related manuals and literature. Based on the obtained squeeze casting process data attribute value value range and numerical type information, the correctness of the data sample is tested. The squeeze casting pouring temperature data type is set to floating point type, and the empirical value range is 50-100℃ higher than the alloy liquidus line. That is, its value rule is determined as follows: <h2 style=";text-align:left;direction:ltr">T<h2 style=";text-align:left;direction:ltr"> p <h2 style=";text-align:left;direction:ltr"> ∈[Tl,Tl+100](2) 3. The method for detecting correctness of squeeze casting process data by integrating domain knowledge and an integrated model according to claim 1, characterized in that: The specific process of step 3 is: At the same time, the local outlier factor, the isolation forest method based on the isolation idea, and the box plot model based on statistical analysis are used for separate detection. The results are y1, y2, and y3 respectively. The outputs of y1, y2, and y3 are all 1 or 0, where 1 represents abnormality and 0 represents correctness; Then, the voting method is used to synthesize the results according to formula (5) to obtain the final test result:

4. The method for detecting correctness of squeeze casting process data by integrating domain knowledge and an integrated model according to claim 3, characterized in that: The specific process of using local outlier factor detection is: Step 3.1.1: Input the squeeze casting process dataset DT, the number of neighbors k, and the abnormal sample threshold e; Step 3.1.2: Normalize the DT process data and convert the attribute values ​​to the [0, 1] interval. Then, abstract the attributes of the squeeze casting process data into coordinate axes to form process data sample points in several dimensional spaces in the local outlier factor. Step 3.1.3: For each data sample S i ∈DT, calculate sample S i The Euclidean distance to other samples is sorted from small to large to obtain the kth distance k_dist(S i ); Step 3.1.4: According to k_dist(S i ), will be less than k_dist(S i ) is included in S i The k-distance neighborhood Nk(Si), and then take Si to any other sample Si ' The larger value of the distance and k_dist(Si) is the reachable distance of sample Si reachdistk(Si,Si ' ); Step 3.1.5: Calculate sample S i The local reachable density lrd(S i ), lrd(S i ) is equal to N k (S i ) to sample S i The reciprocal of the average reachable distance is S i ; Step 3.1.6: Calculate sample S i N k (S i ) of all samples in lrd(S i ) and sumlrd k (S i ); Step 3.1.7: Calculate sample S i The local outlier factor score of Step 3.1.8: Mark abnormal data samples according to the abnormal sample threshold e; Step 3.1.9: If lof(S i )>e will sample S i Marked as an abnormal data sample, the label is recorded as 1, that is, y1 = 1, otherwise the sample S i Marked as a normal data sample, the label is recorded as 0, that is, y1 = 0; Step 3.1.10: Loop through steps 3.1.3 to 3.1.9 to calculate the local outlier factor score y1 for each data sample and complete the correctness check for all samples.

5. The method for detecting correctness of squeeze casting process data by integrating domain knowledge and an integrated model according to claim 3, characterized in that: The specific process of isolation forest method detection based on isolation idea is: Step 3.2: Training phase: Generate a specified number of isolated trees as follows: Step 3.2.1: Input the process data set DT, set the current tree height h, and the isolated tree limit height h lim , number of isolated trees n_estimators, subsampling sample size is an initial value; Step 3.2.2: Normalize the process data and convert the attribute values ​​into the [0,1] interval; Step 3.2.3: If h>h lim orDT|≤1 ends, otherwise randomly select an attribute A from the data set DT j and a split point p, dividing the data set DT into two subsets DT_left, attribute A j The value of is less than p and DT_right, attribute A j The value of is greater than p; Loop until the limit height of the isolated tree reaches h lim , the number of isolated trees is equal to or exceeds n_estimators, and the isolated forest iTrees is obtained; Step 3.3: Evaluation phase: Give the sample anomaly score s as follows: First, calculate the number of edges from the root node of iTree to the external node, that is, the path length, recorded as h(S i ), for a given sample size of The sample subspace and a sample data S i , and its anomaly score s is defined as shown in formula (3): Where E[h(S i )] represents the sample data S i The average path length in n iTrees, Defined as the average path length of failed searches in a binary search tree, normalized h(S i ), which is defined as shown in formula (4): in, is the harmonic series, ξ is Euler’s constant; The closer the anomaly score s of a data sample is to 1, the more likely it is an outlier, and the sample label is recorded as 1. If the anomaly score s of most samples is close to 0.5, it means that there are no obvious outliers in the entire data set. When s is much less than 0.5, it means that the data sample is a normal value, the label is recorded as 0, and the output is y2.

6. The method for detecting correctness of squeeze casting process data by integrating domain knowledge and an integrated model according to claim 3, characterized in that: The specific process of separate detection of the box plot model based on statistical analysis is as follows: The box plot is used to intuitively understand the central tendency, dispersion and outliers of the data. The extreme values ​​that are abnormally greater than the set value or abnormally less than the set value are identified as outliers. The lower quartile Q1, the upper quartile Q3, the value of Q1 at the 25% position, and the value of Q3 at the 75% position are used. The interquartile range is defined as IQR=Q3-Q1, the upper edge of the attribute is Q3+1.5IQR, and the lower edge is Q1-1.5IQR. The data exceeding the upper and lower edge intervals are identified as outliers and marked as 1. The rest of the correct data are marked as 0. The specific steps are as follows: Step 3.3.1: Input process data set DT; Step 3.3.2: Sort the data of each attribute in ascending order; Step 3.3.3: Calculate the minimum value Min, lower quantile Q1, median, upper quantile Q3 and maximum value Max of each attribute; Step 3.3.4: Calculate the interquartile range of the attribute IQR = Q3-Q1; Step 3.3.5: Calculate the upper limit of the attribute: Up = Q3 + 1.5IQR, and the lower limit: Low = Q1 - 1.5IQR; Step 3.3.6: If a ij >Up or a ij <Low, mark the sample to which the data point a ij belongs as abnormal, and label it as 1; otherwise, mark the sample to which the data point a ij belongs as normal, and label it as 0; Step 3.3.7: After the loop of steps 3.3.1 to 3.3.6 is completed, the label y3 indicating whether each data sample is abnormal or not is output.