A diagnostic method for abnormal fluctuations in quality characteristics based on heterogeneous data optimization
By constructing a diagnostic method for abnormal fluctuations in quality characteristics based on heterogeneous data optimization, eliminating the correlation between variables and sampling dimensions, extracting main features and performing dimensionality reduction, the problem of identifying abnormal fluctuations in multi-dimensional and multi-source data during the CNC machining of automobile crankshafts is solved, achieving more efficient anomaly localization and diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-05
- Publication Date
- 2026-04-03
AI Technical Summary
Existing quality diagnostic methods struggle to accurately identify the causes of abnormal fluctuations when faced with multidimensional, multi-source, and heterogeneous information data from CNC machine tools used in the machining of automotive crankshafts. Furthermore, traditional methods are ineffective in processing multidimensional and multi-source data, and are prone to false reporting and underreporting.
By constructing a diagnostic method for abnormal fluctuations in quality characteristics based on heterogeneous data optimization, the correlation between variables and sampling dimensions is eliminated, the main features are extracted and dimensionality is reduced, a multi-dimensional, multi-source diagnostic model for abnormal fluctuations in quality characteristics is constructed, and the abnormal location is located by using a variable correlation metric function model and time delay data blocks.
It improves the speed and accuracy of manufacturing process diagnosis, enables more precise location of the causes of abnormal fluctuations, and overcomes the shortcomings of traditional methods in processing multidimensional and multi-source data.
Smart Images

Figure CN116010766B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of manufacturing process quality diagnosis, and specifically relates to a diagnostic technology for abnormal fluctuations in quality characteristics. Background Technology
[0002] Control charts, as a widely used statistical process control tool, can analyze and infer the patterns or anomalies in data fluctuations based on statistical data, thereby achieving quality diagnosis and process improvement in the manufacturing process. Therefore, control charts are also widely used in quality diagnosis and analysis during the CNC machine tool machining of automotive crankshafts. However, traditional methods of quality control relying on control charts have many uncertainties and ambiguities in practical applications. In particular, while identifying abnormal fluctuations, they cannot pinpoint the causes of these fluctuations. Especially with the rapid development of CNC machine tool machining technology, the machining of automotive crankshafts involves a large amount of multi-dimensional, multi-source, and heterogeneous information data, making the diagnosis of quality characteristic fluctuations in the manufacturing process particularly difficult.
[0003] For analyzing quality anomaly fluctuations, there are methods combining SPC techniques (such as control charts) with control chart pattern recognition, methods based on GA for extracting quality defect association rules, and methods integrating CBR (Case Based Reasoning) and KNN (K-Nearest Neighbor Algorithm) algorithms to infer case similarity measures and using case databases for quality anomaly inference. However, many studies focus on the final output feature analysis, resulting in single models with poor interpretability, which is not conducive to quality decoupling and identification. While existing quality diagnostic methods can detect process anomalies, they still have shortcomings in analyzing and handling anomalies in multidimensional, multi-source, and heterogeneous data. For example, Shewhart's Statistical Process Control (SPC) is prone to false positives and false negatives; H. Hotelling's T2 diagram cannot identify the cause and location of anomalies; Bayesian networks, multilayer perceptrons, RNNs (Recurrent Neural Networks), and deep neural networks are widely applicable, but their performance is poor in handling the ambiguity of multi-layer, multidimensional, and multi-source data. Summary of the Invention
[0004] To more accurately locate manufacturing process anomalies and identify the uncontrolled factors causing abnormal fluctuations, this invention proposes a diagnostic method for abnormal fluctuations in quality characteristics based on heterogeneous data optimization. This method eliminates the correlation between variables and sampling dimensions, extracts a greater number of principal features, optimizes the dimensionality-reduced data, and then constructs a diagnostic model for abnormal fluctuations in quality characteristics based on the data's principal feature dimensionality reduction optimization method. The effectiveness of the proposed method is demonstrated through a case study of automotive crankshaft production.
[0005] The technical solution adopted in this invention is: a method for diagnosing abnormal fluctuations in quality characteristics based on heterogeneous data optimization, such as... Figure 1 As shown, it includes the following steps:
[0006] S1. Collect multi-dimensional, multi-source heterogeneous quality data of automotive crankshafts;
[0007] S2. Based on the collected multidimensional and multi-source heterogeneous mass data of automobile crankshafts, construct the main feature vector of the multidimensional and multi-source heterogeneous mass data of automobile crankshafts.
[0008] S3. The main feature vectors at different time points are calculated step by step using the variable correlation metric function model;
[0009] S4. Based on the main eigenvectors of the current observed multidimensional and multi-source mass data of the automobile crankshaft and the main eigenvectors of the previous d-step observed multidimensional and multi-source mass data of the automobile crankshaft, accumulate them together to construct the observation data matrix of the multidimensional and multi-source process X of the automobile crankshaft.
[0010] S5. The key principal feature analysis was performed on the observation data matrix of the multi-dimensional and multi-source process X of the automobile crankshaft to eliminate the correlation between the principal feature data.
[0011] S6. Based on the observation data matrix of the multidimensional and multi-source process X of the automobile crankshaft after step S5, establish the variable correlation measurement function model of the i-th dimension at time point k.
[0012] S7. Based on the variable correlation metric function model of the i-th dimension at time point k established in step S6, establish the corresponding time delay data block;
[0013] S8. Obtain the variable correlation matrix corresponding to each dimension based on the time delay data block;
[0014] S9. Perform dimensionality reduction on all main features based on the variable correlation matrix;
[0015] S10. Based on the variable correlation matrix of the main features processed in step S9, perform anomaly location.
[0016] Step S10 specifically includes the following sub-steps:
[0017] A1. Randomly initialize the aggregation positions of the main quality features to obtain the initial coordinates (X, Y) of each main feature:
[0018] Init X_axis; Init Y_axis
[0019] A2. Set the random direction and distance of the search target for each main feature to obtain the coordinates of each main feature in the current iteration;
[0020] A3. Based on the distance between the current iteration coordinates of the main feature and its initial coordinates, calculate the heterogeneous feature correlation analysis value corresponding to each main feature;
[0021] A4. Based on the correlation analysis values of the heterogeneous features corresponding to each main feature, obtain the correlation density of the corresponding heterogeneous location;
[0022] A5. Find the principal feature with the largest correlation density value and retain its initial coordinates.
[0023] The beneficial effects of this invention are as follows: This invention constructs a dimensionality reduction and optimization method for multidimensional and multi-source quality characteristic data, eliminating the correlation between variables and sampling dimensions, extracting more principal features, and optimizing the dimensionality-reduced data. Then, a diagnostic model for abnormal fluctuations in quality characteristics based on the data principal feature dimensionality reduction and optimization method is constructed, thereby obtaining the location results of abnormal quality in automotive crankshafts. The effectiveness of the proposed method is verified through examples. This invention overcomes the shortcomings of traditional principal component analysis's linear mapping in effectively handling the correlation between different batches of multidimensional and multi-source heterogeneous quality characteristic data, and can improve diagnostic speed and accuracy. Attached Figure Description
[0024] Figure 1 This is a flowchart of the solution of the present invention;
[0025] Figure 2 For the multivariate T of the training samples 2 Control charts;
[0026] Figure 3 Training graph for the OMFC model;
[0027] Figure 4 This is a training graph for a BP neural network.
[0028] Figure 5 Optimize the travel trajectory for multi-dimensional, multi-source master features. Detailed Implementation
[0029] To facilitate understanding of the technical content of this invention by those skilled in the art, the following description, in conjunction with the accompanying drawings, further illustrates the invention.
[0030] The solution of this invention is applied to the CNC machine tool machining process of automobile crankshafts. It acquires multi-dimensional, multi-source heterogeneous quality data of the automobile crankshaft through data acquisition devices such as sensors. This multi-dimensional, multi-source heterogeneous quality data includes the following three components: the diameter of the main journal of the automobile crankshaft (X1), the diameter of the connecting rod journal of the automobile crankshaft (X2), and the stroke of the crankshaft formed by the main journal and the connecting rod journal (X3). The diagnostic process of this invention includes the following steps:
[0031] 1. Reduced dimension main feature extraction (RDMFE) method for multidimensional, multi-source heterogeneous data in the manufacturing process.
[0032] The vector composed of the main features of the multidimensional, multi-source, heterogeneous quality data of automotive crankshafts is designed as follows:
[0033] X = [x1, x2, ..., x J ] T
[0034] In the formula, X represents that the vector has J principal features, and in this embodiment, J = 3.
[0035] Set {x t {ε} represents a multidimensional, multi-source, heterogeneous time series data, where t = 0, ±1, ±2, ... t If t = 0, ±1, ±2, ... is a valid data sequence, then the following equation holds:
[0036] x t =α1x t-1 +α2x t-2 +…+α p x t-p +ε t -βε t-1 -…-βε t-q
[0037] α(μ) and β(μ) have no common roots, therefore the above equation can be represented by the shift operators α() and β(), specifically expressed as follows:
[0038] α(B)x t =β(B)ε t
[0039] Among them, α(μ)=1-α1μ-α2μ 2 -…-α p μ p ,β(μ)=1-β1μ-β2μ 2 -…-β p μ p
[0040] Since α(μ) satisfies the minimum phase condition, we have ρ>1, such that α ∈ {μ:|μ|≤ρ}. -1 (μ)β(μ) is analytic, thus yielding the Taylor expansion.
[0041]
[0042] Therefore, we can obtain
[0043]
[0044] x t The variable correlation metric function is expressed as follows:
[0045]
[0046] Where, γ k It is x t The autocovariance of the function represents the correlation between the principal eigenvectors at time j and time j+k, where μ is a real coefficient and ψ is a variable. j For {x t The Wold coefficient, σ 2 The mean square error of the prediction in one step (predicting the next step).
[0047] Wherein, E(X) j -μ) represents X j The expected value of -μ, E(X) j+k -μ) represents X j+k The expected value of -μ;
[0048] The principal feature vectors at different time points are calculated step by step using a variable correlation metric function model.
[0049] Suppose we have the observed data of a multidimensional, multi-source process X, which is the sum of the principal eigenvectors of the current multidimensional, multi-source quality data and the principal eigenvectors of the quality data from the previous d-step observations. We can construct the following matrix:
[0050]
[0051] In the formula, the superscript T denotes transpose, and X(k) = [X 1,k ,X 2,k ,…X J,k ] T Let J represent the principal eigenvector of the observation at time point k.
[0052] For matrix X d Key principal feature analysis is performed to eliminate correlations among the collected quality principal feature data.
[0053] For all the obtained batches I of sample data, a variable correlation metric function model is established for the i-th batch of sample data at time point k, and its expression is as follows:
[0054] X i (k)=[(X i (k)) T ,(X i (k-1)) T ,…,(X i (kd)) T ]
[0055] In the formula, d represents the interval length of the dynamic observation process.
[0056] For example, suppose the i-th batch of data is as shown in Table 1:
[0057] Table 1 Data of the i-th batch
[0058]
[0059] This set of data indicates that the i-th batch of data collected over a time period of d. As shown in Table 1, the amount of data collected within the time period d is 10. Each index represents the data at each of the 10 times in the i-th batch of data, i.e., X in the above formula. i Each element in (k) is: (X i (kd)) T =[89.981,73.110,67.266]; (X i (k-(d-1))) T =[89.991,73.118,67.798];...;(X i (k-1)) T =[89.996,73.235,67.430]; (X i (k)) T =[89.997,73.735,67.452].
[0060] For the i-th batch of sample data, establish its corresponding... This is a time delay data block for all quality characteristic data measurements. (Process time delay data block) The expression is as follows:
[0061]
[0062] Based on the determined block interval length *d*, a multi-dimensional, multi-source master feature dataset based on the above formula can be obtained. The value of *d* is set according to actual needs. Furthermore, the real-time behavior of relevant quality characteristic parameters of the dynamic manufacturing process can be obtained from this dataset. The variable correlation matrix of the *i*th batch of sample data... It can be represented as:
[0063]
[0064] Variable correlation matrix The formula for each element in the formula is:
[0065]
[0066] Variable correlation matrix Each element in the variable is a variable J1 (i.e. ) and variable J2 (i.e. The correlation between the observations in the i-th batch of sample data is denoted as .
[0067] For all batches I of sample data, the data block can obtain relevant key feature information from the time-delayed data sequence of each batch of sample data: based on all the obtained batches I of sample data, the variable-related key feature information of each batch of sample data is aggregated to obtain the variable-related values of all operating dimensions, and then the average variable-related matrix can be obtained.
[0068]
[0069] Sort all eigenvalues λ of the correlation matrix of the average variable from largest to smallest, and then orthogonalize the eigenvectors corresponding to each eigenvalue using the Schmitt orthogonalization method to obtain the orthogonalized eigenvector matrix P.
[0070] From the formula Calculate the contribution rates B1,…,B for each eigenvalue λ. p The eigenvectors corresponding to the eigenvalues with a contribution rate greater than 85% in the eigenvector matrix P are the main features obtained after dimensionality reduction by the multidimensional multi-source heterogeneous data dimensionality reduction main feature extraction method, where p represents the number of eigenvalues in the average variable correlation matrix.
[0071] 2. Construction of a multidimensional, multi-source main feature correlation optimization quality characteristic diagnostic model (OMFC)
[0072] The anomaly location is identified within the principal feature vector obtained after dimensionality reduction, specifically including the following steps:
[0073] (1) Randomly initialize the aggregation positions of the main quality features to obtain the initial coordinates (X, Y):
[0074] Init X_axis; Init Y_axis
[0075] The specific values of X and Y are the offset data that occur when a CNC machine tool cuts and processes an automobile crankshaft.
[0076] (2) Set the random direction and distance of the search target for each quality main feature:
[0077] X i = X_axis + Random Value
[0078] Y i =Y_axis + Random Value
[0079] Random Value is a built-in random number generator in MATLAB. In this example, the random direction and distance are represented by positive and negative signs, and the distance is represented by a random numerical value. The range of Random Value is [-10, 10].
[0080] (3) Since the location of the main anomaly cannot be determined, the distance (Dist) between the main anomaly and the origin of the coordinate system is estimated first, and then the anomaly correlation discrimination value (S) is calculated. i ), and set it as the reciprocal of the distance:
[0081]
[0082]
[0083] (4) The heterogeneous characteristic correlation analysis value (S i Substituting the values into the differential characteristic correlation discriminant function, the correlation density Fit(S) at the location of the differential cause of this quality characteristic can be obtained. i ).
[0084]
[0085] Where: b = min(S) i ), α=S i -min(S i ), β takes the value of 2. (5) Find the single mass principal feature corresponding to the highest correlation density value in this set of mass principal features (find the maximum value):
[0086] max(Fit(S i ))
[0087] (6) Retain the current X and Y coordinates of the principal feature vector corresponding to the maximum correlation density value. At this time, the quality principal feature set starts to move to this position, that is, all the data in this set are transformed into the X and Y coordinates of the maximum correlation density value, and a new feature cluster central position is formed. After the new feature cluster central position is formed, step (7) is executed to perform iterative optimization. Step (2) is performed to add a random number Random Value to each coordinate in this set of data to form the next set of datasets. Continue to find the point with the highest correlation density in the next set of datasets.
[0088] X_axis = X(bestIndex)
[0089] Y_axis = Y(bestIndex)
[0090] (7) Enter iterative optimization, repeat steps (2) to (5), and determine whether the correlation density value is greater than the correlation density value of the previous iteration. If so, execute step (6), that is, the highest correlation density value calculated in the new round of data is lower than that in the previous round, and the iteration needs to be restarted. Add a random number (Random Value) to this set of data again. Through this operation process, the specific location of the key principal feature anomaly with the best correlation density value can be found.
[0091] The iteration can be stopped if the current iteration meets one of the following three conditions:
[0092] 1) The number of iterations reaches 1000;
[0093] 2) Substitute the parameters obtained through iteration into the model to estimate the results. When the root mean square error between the estimated result and the actual value is less than 10... -6 hour;
[0094] 3) The root mean square error change in the current iteration is less than 10 compared to the previous iteration. -8 hour;
[0095] The formula for calculating the root mean square error is:
[0096]
[0097] In the formula, n is the total number of samples used for estimation, and prediction i and target i These refer to the estimated value and the actual value of the i-th parameter, respectively.
[0098] The determination of whether this iteration is valid is as follows: if the correlation density corresponding to this iteration is greater than the correlation density value of the previous iteration, then this iteration is valid.
[0099] Taking the production data of a company's CNC machine tool machining of automobile crankshafts as the research object, the variables selected are as follows: X1 crankshaft main journal diameter, X2 crankshaft connecting rod journal diameter, X3 crankshaft stroke formed by the main journal and connecting rod journal. The dimensionality reduction steps of the main features of the training sample data are as follows:
[0100] (1) Generate random samples
[0101] Generate enough training sample data randomly.
[0102] (2) Through multivariate T 2 Control charts remove data that do not show anomalies
[0103] Table 2 Loss of Control Modes
[0104]
[0105] After training with training sample data, the model output data is categorized into different runaway modes and recorded in the database.
[0106] Fifty data sets were selected as training samples for each runaway mode exhibiting anomalies, as shown in Table 2 for the seven runaway modes, totaling 350 training samples. Table 3 shows the training data when X1, X2, and X3 shifted. Multivariate T-strain analysis of the anomaly data is also presented. 2 Control charts as follows Figure 2 As shown.
[0107] Table 3 Training data when X1, X2, and X3 are offset
[0108]
[0109] (3) Dimensionality reduction of main features of training sample data
[0110] The density values obtained after dimensionality reduction are shown in Table 4. The principal feature density values obtained after dimensionality reduction are shown in Table 5.
[0111] Table 4. Correlation density values of each principal feature
[0112]
[0113] Table 5. Correlation density values of principal features in dimensionality reduction data
[0114]
[0115] Application of multi-dimensional, multi-source principal feature optimization for quality characteristic diagnosis:
[0116] (1) Input and Output Design
[0117] Let the output variable be T = (T1, T2, T3). T There are a total of 7 distinct output patterns, as shown in Table 6.
[0118] Table 6. Seven Significant Output Patterns
[0119]
[0120] Wherein, T1, T2, and T3 represent the states of the three main features X1, X2, and X3 in this embodiment, respectively. A state value of 1 indicates out of control (a deviation has occurred, i.e., the main feature data is abnormal); a state value of 0 indicates under control (i.e., the main feature data is normal). (2) The preset parameter settings are shown in Table 7.
[0121] Table 7 Preset Parameters
[0122]
[0123] (3) Network training
[0124] Find the mapping relationship between the input data main features and the output runaway mode in the sample data main features.
[0125] The training results of the OMFC model and the BP neural network model are as follows: Figure 3 and Figure 4 As shown, it can be seen that if the training error is to meet the training requirements, OMFC needs to iterate 230 times and the BP neural network model needs to iterate 420 times. Therefore, the multidimensional multi-source main feature optimization model OMFC designed in this invention is better.
[0126] The training process of the multidimensional multi-source principal feature optimization model is as follows: Figure 5 As shown, by Figure 4 and Figure 5 It can be seen that the multidimensional multi-source principal feature optimization model requires fewer uncertain parameters and is easier to operate.
[0127] The diagnostic results of OMFC and the diagnostic results of the BP neural network model are shown in Table 8.
[0128] Table 8. Diagnostic results of OMFC and BP neural network model
[0129]
[0130] It can be seen that the multidimensional multi-source principal feature optimization model OMFC has more accurate diagnostic results than the BP neural network model. Therefore, in summary, the multidimensional multi-source principal feature optimization model OMFC has advantages over the BP neural network model, such as fewer iterations required to achieve optimal results, fewer uncertain parameters required for human input, easier operation, and more accurate diagnostic results.
[0131] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Various modifications and variations can be made to the invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the scope of the claims of the invention.
Claims
1. A method for diagnosing abnormal fluctuations in quality characteristics based on heterogeneous data optimization, characterized in that, Includes the following steps: S1. Collect multi-dimensional, multi-source heterogeneous quality data of automotive crankshafts; S2. Based on the collected multidimensional and multi-source heterogeneous mass data of automobile crankshafts, construct the main feature vector of the multidimensional and multi-source heterogeneous mass data of automobile crankshafts. S3. The main feature vectors at different time points are calculated step by step using the variable correlation metric function model; S4. Based on the currently observed multidimensional, multi-source mass data of the automobile crankshaft, the principal eigenvectors and previous... The main feature vectors of the multidimensional and multi-source quality data of the automobile crankshaft observed step by step are accumulated together to construct the observation data matrix of the multidimensional and multi-source process X of the automobile crankshaft. S5. The key principal feature analysis was performed on the observation data matrix of the multi-dimensional and multi-source process X of the automobile crankshaft to eliminate the correlation between the principal feature data. S6. Based on the observation data matrix of the multi-dimensional, multi-source process X of the automobile crankshaft processed in step S5, establish... The first point in time A variable correlation measurement function model for batch sample data; S7. Based on the steps established in S6 The first point in time A variable correlation metric function model for batch sample data is used to establish corresponding time-delay data blocks. S8. Obtain the variable correlation matrix corresponding to each batch of sample data based on the time delay data blocks; S9. Perform dimensionality reduction on all main features based on the variable correlation matrix; S10. Based on the variable correlation matrix of the main features processed in step S9, perform anomaly location; step S10 is detailed below. It includes the following steps: A1. Randomly initialize the aggregation positions of the main quality features to obtain the initial coordinates of each main feature. ): ; A2. Set the random direction and distance of the search target for each main feature to obtain the coordinates of each main feature in the current iteration; A3. Based on the distance between the current iteration coordinates of the main feature and its initial coordinates, calculate the heterogeneous feature correlation analysis value corresponding to each main feature; A4. Based on the correlation analysis values of the heterogeneous features corresponding to each main feature, obtain the correlation density of the corresponding heterogeneous location; the formula for calculating the correlation density of the heterogeneous location in step A4 is: ; in, This represents the correlation density at the heterogeneous location corresponding to the i-th principal feature. b represents the heterogeneous feature correlation analysis value corresponding to the i-th principal feature, where b = , , ; A5. Find the principal feature with the largest correlation density value and retain its coordinates in the current iteration.
2. The method for diagnosing abnormal fluctuations in quality characteristics based on heterogeneous data optimization according to claim 1, characterized in that, The dimensionality reduction process described in step S9 specifically includes: Let there be I batches of sample data. Based on the correlation matrix of the variables in each batch of sample data, we obtain the average correlation matrix, i.e.: ; in, Represents the correlation matrix of the average variables. Indicates the first The correlation matrix of variables corresponding to the batch sample data. Indicates the interval length of the dynamic observation process. This represents the observation data matrix of the multidimensional, multi-source process X of the automobile crankshaft. The dimension representing the principal feature vector of multidimensional, multi-source, heterogeneous quality data of automotive crankshafts. Indicates the first time; The contribution rate of each eigenvalue of the correlation matrix of the average variable is calculated according to the following formula, and the eigenvectors corresponding to the eigenvalues with a contribution rate greater than 85% are taken as the principal eigenvectors after dimensionality reduction: ; in, This represents the i-th eigenvalue of the correlation matrix of the average variables. , This represents the number of eigenvalues in the correlation matrix of the average variables. express The corresponding contribution rate.
3. The method for diagnosing abnormal fluctuations in quality characteristics based on heterogeneous data optimization according to claim 2, characterized in that, The expression for the time-delay data block established in step S7 is as follows: ; in, Indicates the first The delay data block corresponding to the batch sample data, with the superscript T indicating transpose. express time point 3D observation principal eigenvector, , They represent The first-dimensional principal feature observed at a given time point, The second-dimensional principal feature observed at a given time point. , The J-th dimension of the observed principal features at time point.
Citation Information
Patent Citations
Process monitoring method based on hybrid kernel PCA-CCA and kernel density estimation
CN111209973A
Method for diagnosing ADHD and related behavioral disorders
WO2011062890A1