Wind tunnel test data anomaly detection method fusing KPCA and fuzzy particle density
Through the combination of KPCA and fuzzy particle density, the problem of multimodal data feature extraction and fusion in wind tunnel tests is solved, and unsupervised wind tunnel test data abnormality detection is achieved, improving the accuracy and applicability of the detection.
Patent Information
- Application Number
- CN202510439976.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
The existing wind tunnel test abnormality detection methods are difficult to effectively process multimodal data characteristics, especially nonlinear data, and require a large amount of manual labeling of data, which limits the generalization ability and engineering applicability of the detection algorithm, and is difficult to meet the rapid migration needs of different test conditions.
KPCA is used for nonlinear dimensionality reduction, and the data is mapped to high-dimensional space through the radial basis function kernel, making it linearly divisible. The abnormal fraction is calculated by combining fuzzy particle density and fuzzy entropy to construct a fused abnormality detection method, without labeling data for model training, and unsupervised abnormality detection is achieved.
It realizes rapid and accurate abnormal detection of wind tunnel test data, effectively deals with fuzzy uncertain nonlinear data, improves the accuracy and applicability of detection, and is suitable for the extraction and fusion of multimodal data features.
Smart Images

Figure CN120372158A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method for detecting abnormal wind tunnel test data by integrating KPCA and fuzzy granular density. Background Technique
[0002] As an important means of aerodynamic research, the data quality of wind tunnel tests is directly related to the reliability of key links such as aircraft design and performance verification. In a complex aerodynamic load environment, test data is easily affected by multiple factors such as equipment vibration, turbulence interference, and sensor drift. If these abnormal data are not effectively identified, it will lead to calculation deviations of aerodynamic coefficients, seriously affecting the accuracy of engineering design. Therefore, developing an efficient method for detecting abnormal wind tunnel test data has important engineering significance.
[0003] Currently, most of the abnormal data in wind tunnel tests are detected manually, and then comprehensively judged based on expert experience, and finally accurate equipment fault causes and corresponding operation suggestions are given. However, the data in modern wind tunnel tests is large in quantity and diverse in type, and the test time is precious. Researchers hope to ensure that experimental measures can timely avoid abnormal phenomena or failures. When these data are detected and diagnosed manually, it is usually time-consuming and it is inevitable to miss or misdetect. To make up for the deficiencies of manual detection and diagnosis methods, in recent years, relevant research on introducing artificial intelligence technology into wind tunnel tests has also emerged. Specifically: The method based on statistical thresholds mainly uses the 3σ criterion or the interquartile range method to detect outliers in single-channel data. Although this kind of method is simple to implement, it is only applicable to Gaussian distribution data and ignores the spatial correlation of multi-sensor data, and the miss detection rate is relatively high. Traditional machine learning methods are divided into supervised learning and unsupervised learning methods. In supervised learning methods, algorithms such as support vector machine (SVM) and random forest need to rely on a historical sample library to construct a classification model, but in practice, the cost of reproducing abnormal working conditions in wind tunnel tests is extremely high, and the model is difficult to be migrated to different test projects. In unsupervised learning methods, principal component analysis (PCA) detects abnormalities through linear dimensionality reduction reconstruction error, but it is not sensitive enough to nonlinear phenomena such as shock oscillations in the transonic region. In modern deep learning methods, although the long short-term memory network (LSTM) can capture temporal features, it also faces challenges such as a large amount of training samples, lack of interpretability of hidden layer features, and long response delay to sudden changes in test parameters, and cannot meet the real-time monitoring requirements, and is not suitable for monitoring a large amount of wind tunnel test data.
[0004] Generally speaking, the existing anomaly detection methods for wind tunnel tests usually have the following problems: (1) Most of the existing methods are only applicable to data analysis at a single granularity and it is difficult to comprehensively utilize data features from different granularities; (2) For non-linear data with unlabeled, high-dimensional, and fuzzy uncertainties, it is difficult for existing detection methods to accurately capture the relationships between data; (3) Most of the anomaly detection methods for the wind tunnel test field require a large amount of manually labeled data, thus limiting the generalization ability and engineering applicability of the detection algorithm and making it difficult to meet the rapid migration requirements of different test conditions. Summary of the Invention
[0005] In view of the above deficiencies in the prior art, the present invention provides an anomaly detection method for wind tunnel test data that combines KPCA and fuzzy granular density, which solves the problem of extraction and fusion of multi-modal data features in existing wind tunnel test data anomaly detection, and can effectively process non-linear data with fuzzy uncertainties, thereby assisting wind tunnel test personnel to discover anomalies more quickly and accurately.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is: an anomaly detection method for wind tunnel test data that combines KPCA and fuzzy granular density, comprising the following steps:
[0007] S1. Obtain wind tunnel test data and perform normalization processing on the wind tunnel test data to obtain the normalized wind tunnel test data;
[0008] S2. According to the normalized wind tunnel test data, use KPCA for data dimensionality reduction and construct a KPCA anomaly score based on the reconstruction error;
[0009] S3. Perform normalization processing on the dimension-reduced wind tunnel test data and calculate the fuzzy relation matrix using the mixed fuzzy relation similarity;
[0010] S4. Calculate the fuzzy granular set of each attribute according to the fuzzy relation matrix;
[0011] S5. Calculate the fuzzy granular density of each object with respect to each attribute according to the fuzzy granular set, and calculate the fuzzy entropy and its corresponding weight;
[0012] S6. Calculate the fuzzy granular anomaly score of all wind tunnel test data according to the fuzzy granular density and its corresponding weight;
[0013] S7. Perform weighted fusion on the KPCA anomaly score and the fuzzy granular anomaly score to obtain the fused anomaly score;
[0014] S8. Judging one by one whether the abnormal score after the fusion of wind tunnel test data is greater than the threshold. If so, output the abnormal data and repeat S8 until all wind tunnel test data have been judged. Otherwise, regard it as normal data and judge the next wind tunnel test data until all wind tunnel test data have been judged.
[0015] The beneficial effects of the present invention are as follows: The present invention uses KPCA for non-linear dimensionality reduction, and maps the wind tunnel test data that is inseparable in the low-dimensional space to a high-dimensional space through the radial basis function kernel (RBF kernel) to make it linearly separable, solving the problem of insufficient feature extraction of the PCA method for non-linear data. At the same time, fuzzy information granules are constructed from the mixed fuzzy relation similarity to obtain a set of fuzzy information granules, solving the problem of feature extraction of multi-modal data in wind tunnel tests. The uncertainty and abnormal characteristics of the information granules are characterized by fuzzy granule density and fuzzy entropy respectively, and then the fuzzy entropy is weighted and summed to calculate the abnormal score of the sample, solving the problem of feature fusion of multi-modal data. The present invention does not require labeled data for model training and can effectively realize unsupervised anomaly detection of wind tunnel test data, thus assisting wind tunnel test personnel to discover anomalies more quickly and accurately.
[0016] Further, the specific content of S2 is as follows:
[0017] Set the radial basis function kernel parameter δ;
[0018] Based on the normalized wind tunnel test data, calculate the kernel matrix according to the radial basis function kernel;
[0019] Perform centering processing on the kernel matrix generated by the radial basis function kernel;
[0020] Perform eigenvalue decomposition on the obtained centered kernel matrix to obtain the eigenvectors belonging to the eigenvalues, and determine the target dimension d after dimensionality reduction according to the Bayesian information criterion;
[0021] According to the eigenvalues and eigenvectors, select the first d largest eigenvalues to form a subspace after dimensionality reduction, and obtain the wind tunnel test data after dimensionality reduction;
[0022] Map the wind tunnel test data after dimensionality reduction back to the original space, and construct the KPCA abnormal score according to the reconstruction error.
[0023] The beneficial effect of the above further solution is that the present invention solves the problem of insufficient feature extraction ability of traditional methods when dealing with non-linear data by using the radial basis function kernel for data dimensionality reduction and abnormal score construction.
[0024] Furthermore, the expression of the kernel matrix is as follows:
[0025] X = [x1, x2,..., x n
[0026]
[0027] Among them, X represents a matrix, each column of which represents a wind tunnel test data, and x n represents the nth normalized wind tunnel test data, φ(x) represents a non-linear mapping function, which maps the variable x in the matrix X i to a high-dimensional space, H represents the dimension after mapping, K represents the dimension before mapping, and K ij represents the value of the i-th row and j-th column of the kernel matrix, k() represents a similarity function, and k(x i , x j ) represents the similarity between the wind tunnel test data x i and x j in the high-dimensional space, γ represents the radial basis function kernel parameter, and ||x i - x j ||2 represents the Euclidean distance between the wind tunnel test data x i and x j ; represents the K-dimensional real vector space, represents the H-dimensional real vector space;
[0028] The expression of the centralized kernel matrix is as follows:
[0029]
[0030] Among them, K' represents the centralized kernel matrix, φ'(X) represents the matrix obtained by centralizing the wind tunnel test data matrix in the high-dimensional feature space, and φ'(X) T represents the matrix obtained by transposing the centralized matrix, 1 n×1 represents an n-dimensional column vector introduced for the simplified expression, each element of which is 1, and 1 n represents an n×n matrix, each element of which is K represents the kernel matrix, which is used to represent the similarity relationship of samples in the high-dimensional feature space;
[0031] The expression of the eigenvector is as follows:
[0032] K'v = λv
[0033] Among them, v represents the eigenvector belonging to λ, and λ represents the eigenvalue of K';
[0034] The expression of the wind tunnel test data after dimensionality reduction is as follows:
[0035] Y = V T K
[0036] Y = [y1, y2,..., yi ,..., y d
[0037]
[0038] V = [v1, v2,..., v d T
[0039] k(x i ) = [k(x i , x1), k(x i , x2),..., k(x i , x j ),..., k(x i , x n )] T
[0040] Among them, Y represents the matrix after dimensionality reduction of d×n dimensions, V represents the matrix of eigenvectors of n×d dimensions, y i and y d both represent the data representation after dimensionality reduction of the wind tunnel test, k(x i ) represents the column vector composed of the kernel function values between the wind tunnel test data x i and all data, k(x i , x n ) represents the similarity between the wind tunnel test data x i and x n in the high-dimensional feature space, v d represents the d-th eigenvector, representing the d-th direction of the principal component KPCA projection space;
[0041] The expression of the KPCA anomaly score is as follows:
[0042]
[0043] Among them, φ(x i ) represents the reconstructed representation of the wind tunnel test data y i in the high-dimensional space, y i'j' represents the projection value of the i'-th wind tunnel test data in the j'-th principal component direction in the matrix Y after dimensionality reduction, ||||2 represents the Euclidean distance, v j” represents the j''-th eigenvector of the kernel matrix K' in KPCA, d represents the target dimension after dimensionality reduction, and KPOS(x i ) represents the KPCA anomaly score of the wind tunnel test data x i .
[0044] The beneficial effects of the above further solution are as follows: By constructing a fuzzy relation matrix through the mixed fuzzy relation similarity, the present invention avoids the limitations of the single processing method of traditional similarity calculation methods and effectively retains the characteristic differences of multi-modal data in wind tunnel tests.
[0045] Furthermore, the expression of the fuzzy relation matrix is as follows:
[0046]
[0047] Wherein, represents the fuzzy similarity between wind tunnel test data x i and x j with respect to attribute c k , |x i (c k ) - x j (c k )| represents the difference between wind tunnel test data x i and x j on attribute c k , represents the threshold parameter, std(c k ) represents the standard deviation on attribute c k , δ represents the adjustable parameter, R B (x i , x j ) represents the fuzzy relation generated by attribute subset B, x j (c k ) represents the value of wind tunnel test data x j on attribute c k , x i (c k ) represents the data of wind tunnel test data x i on attribute c k and the value of x j on attribute c k , x i (c k ) represents the value of wind tunnel test data x i on attribute c k .
[0048] Furthermore, the expression of the fuzzy granule set is as follows:
[0049] G U / R B = {[x1] B , [x2] B ,..., [x i B ,..., [x n B}
[0050]
[0051] Among them, G U / R B represents the fuzzy granule set constructed under the fuzzy relation R B [x i B represents the fuzzy information granule centered on the wind tunnel test data x i generated by the mixed fuzzy relation similarity, [x n B represents the fuzzy information granule centered on the wind tunnel test data x n generated by the mixed fuzzy relation similarity. U represents the set of wind tunnel test data, B represents the attribute subset, and R B represents the fuzzy relation generated by the attribute subset B. G U represents the set of fuzzy information granules constructed on the wind tunnel test data set U, [x i B (x j ) is the membership function, indicating that the wind tunnel test data x i and x j have the fuzzy relation R B . R B (x i , x n ) represents the fuzzy similarity of the wind tunnel test data x i and x n under the attribute set B. R B (x i , x j ) represents the fuzzy similarity of the wind tunnel test data x i and x j under the attribute set B. represents the element in the i-th row and n-th column of the fuzzy relation matrix R B . represents the element in the i-th row and j-th column of the fuzzy relation matrix R B .
[0052] The beneficial effects of the above further solution are as follows: By constructing a multi-granularity fuzzy information granule set, the present invention realizes the fine-grained feature expression of wind tunnel test data in different attribute dimensions, and effectively solves the problem of insufficient representation ability of traditional methods for mixed attribute data.
[0053] Furthermore, the specific content of S5 is as follows:
[0054] Calculate the fuzzy granule density of each object with respect to each attribute according to the fuzzy granule set;
[0055] Calculate the fuzzy entropy and its corresponding weight according to the fuzzy granule density.
[0056] The beneficial effects of the above further solution are as follows: By introducing fuzzy granular density, the present invention realizes the precise quantification of the local distribution characteristics of data. The adaptive weight assignment mechanism based on fuzzy entropy overcomes the limitation of manually setting attribute weights in traditional methods and realizes the objective evaluation of the importance of different attributes.
[0057] Furthermore, the expression of the weight is as follows:
[0058]
[0059] Wherein, W(c k ) represents the weight function of the wind tunnel test data on attribute c k , C represents the set of all attributes with multiple attributes, k represents the wind tunnel test data attribute number, U represents the wind tunnel test data set, x i represents the wind tunnel test data, |[x i C | is the cardinality of the fuzzy information granule centered on the wind tunnel test data x i generated by the mixed fuzzy relation similarity, represents the fuzzy granular density of the wind tunnel test data x i on attribute c k , |U| represents the cardinality of the wind tunnel test data, represents the cardinality of the fuzzy granule , FE(c k ) represents the calculation function of the fuzzy entropy of the fuzzy relation matrix , FE(C) represents the calculation function of the fuzzy entropy of the fuzzy relation matrix R C .
[0060] Furthermore, the expression of the fuzzy granule anomaly score is as follows:
[0061]
[0062] Wherein, FGOS(x i ) represents the fuzzy granule anomaly score of the wind tunnel test data x i , C represents the set of all attributes with multiple attributes, W(c k ) represents the weight function of the wind tunnel test data on attribute c k , represents the fuzzy granular density of the wind tunnel test data x i on attribute c k .
[0063] Furthermore, the expression of the fusion anomaly score is as follows:
[0064] KFDOS(xi ) = KPOS(x i ) × FGOS(x i )
[0065] Wherein, KFDOS(x i ) represents the fusion anomaly score of the wind tunnel test data x i , FGOS(x i ) represents the fuzzy granule anomaly score of the wind tunnel test data x i , and KPOS(x i ) represents the KPCA anomaly score of the wind tunnel test data x i .
[0066] The beneficial effect of the above further solution is that by fusing the product of the KPCA anomaly score and the fuzzy granule anomaly score, a dual detection mechanism of global feature anomaly and local density anomaly is realized, effectively solving the misjudgment problem that may be caused by a single detection perspective. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0068] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0069] Embodiment
[0070] In this embodiment, KPCA is an important non - linear feature extraction model. It inherits the dimensionality reduction idea of traditional PCA and maps the data to a high - dimensional space through the introduction of a radial basis function kernel for non - linear feature extraction, effectively mining the potential information between samples. At the same time, the concept of fuzzy granules is used to perform a soft partition on the distribution of wind tunnel test data and quantify the local density, making it more flexible to handle noise, missing values, or boundary fuzziness in the data. In the anomaly detection of wind tunnel test data, the wind tunnel test data is imported into an information system, where each row represents a test data (i.e., a sample), and each column represents an attribute of the test data (i.e., a feature). An information system is represented by a quadruple FIS = (U, C, V, f), where U = {x1, x2, …, x n} represents a set composed of n test samples, called the universe of discourse, C = {c1, c2, …, c m} represents a set composed of m conditional attributes, V represents the attribute domain, that is, V = U c∈CV c where f represents a mapping function that associates each sample x i ∈U with the domain of attribute c k , i.e., f(x i (c k )) ∈ V c , where x i represents the i-th wind tunnel test data, and c k represents the k-th attribute of the k-th wind tunnel test data, and V c represents the value range of attribute c.
[0071] In this embodiment, first, the radial basis function kernel is selected to calculate the sample similarity in the high-dimensional feature space, solving the problem that it is difficult to characterize the similarity between complex non-linear data; then, the fuzzy information granules are constructed from the mixed fuzzy relation similarity to obtain a set of fuzzy information granules, solving the problem of extracting the features of multi-modal data in wind tunnel tests; the uncertainty and abnormal characteristics of the information granules are characterized by the fuzzy granule density and fuzzy entropy respectively, and then the fuzzy entropy is weighted and summed to calculate the abnormal score of the sample, solving the problem of feature fusion of multi-modal data; the present invention does not require labeled data for model training and can effectively achieve unsupervised anomaly detection of wind tunnel test data, thus assisting wind tunnel test personnel to discover anomalies more quickly and accurately.
[0072] As Figure 1 shown, the present invention provides a method for anomaly detection of wind tunnel test data by fusing KPCA and fuzzy granule density, and the implementation method is as follows:
[0073] S1. Obtain wind tunnel test data and perform normalization processing on the wind tunnel test data to obtain the normalized wind tunnel test data;
[0074] In this embodiment, the expression of the normalization processing is as follows:
[0075]
[0076] where F() represents the normalization processing function, and x i (c k ) represents the value of the wind tunnel test data x i on attribute c k , and respectively represent the maximum and minimum values of all wind tunnel test data on attribute c k .
[0077] In this embodiment, the value range of the numerical data is adjusted to the real number interval from 0 to 1 through the min-max normalization operation.
[0078] S2. According to the wind tunnel test data after normalization processing, use KPCA for data dimensionality reduction, and construct the KPCA anomaly score based on the reconstruction error. The implementation method is as follows:
[0079] Set the radial basis function kernel parameter δ according to expert experience or previous research;
[0080] Express the original data and calculate the kernel matrix according to the radial basis function kernel;
[0081] In this embodiment, calculate the kernel matrix according to the radial basis function kernel (RBF kernel):
[0082] X = [x1, x2,..., x n
[0083]
[0084] where X represents a matrix, each column of which represents a wind tunnel test data, x n represents the nth wind tunnel test data after normalization processing, φ(x) represents a non - linear mapping function that maps the variable x in matrix X i to a high - dimensional space, H represents the dimension after mapping, K represents the dimension before mapping, K ij represents the value of the i - th row and j - th column of the kernel matrix, k() represents a similarity function, k(x i , x j ) represents the similarity between the wind tunnel test data x i and x j in the high - dimensional space, γ represents the radial basis function kernel parameter, ||x i - x j ||2 represents the Euclidean distance between the wind tunnel test data x i and x j , represents the K - dimensional real vector space, represents the H - dimensional real vector space.
[0085] Perform centering processing on the kernel matrix generated according to the radial basis function kernel:
[0086]
[0087] where K' represents the centered kernel matrix, φ'(X) represents the matrix obtained by centering the wind tunnel test data matrix in the high - dimensional feature space, φ'(X) T represents the matrix obtained by transposing the centered matrix, 1 n×1 represents an n - dimensional column vector introduced for simplifying the expression, each element of which is 1, 1 n represents an n×n matrix, each element of which is Let \(K\) denote the kernel matrix, which is used to represent the similarity relationship of samples in the high-dimensional feature space;
[0088] Based on the obtained centralized kernel matrix, perform eigenvalue decomposition on it to obtain the eigenvectors belonging to the eigenvalues, and determine the target dimension \(d\) after dimensionality reduction according to the Bayesian information criterion;
[0089] In this embodiment, based on the obtained centralized kernel matrix, perform eigenvalue decomposition on it to obtain the eigenvector \(v\) belonging to the eigenvalue \(\lambda\). According to the Bayesian information criterion (BIC: Bayesian information criterion), the target dimension after dimensionality reduction can be determined:
[0090] \(K'v=\lambda v\)
[0091]
[0092] where \(v\) represents the eigenvector belonging to \(\lambda\), \(\lambda\) represents the eigenvalue of \(K'\), BIC represents the calculation formula of the Bayesian information criterion, \(|\Lambda|\) represents the number of non-zero eigenvalues of the model, CumulativeVar represents the cumulative variance of the model, \(k\) represents the number of parameters of the model, \(n\) represents the number of samples, and \(\ln\) represents the natural logarithm function.
[0093] According to the eigenvalues and eigenvectors, select the first \(d\) largest eigenvalues to form the subspace after dimensionality reduction, and obtain the wind tunnel test data after dimensionality reduction:
[0094] \(Y = V\) T \(K\)
[0095] \(Y = [y_1, y_2, \cdots, y\) i , \cdots, y\) d \)
[0096]
[0097] \(V = [v_1, v_2, \cdots, v\) d \) T
[0098] \(k(x\) i ) = [k(x\) i , x_1), k(x\) i , x_2), \cdots, k(x\) i , x\) j ), \cdots, k(x\) i , x\) n )]\) T
[0099] where \(Y\) represents the matrix after dimensionality reduction with dimensions \(d\times n\), \(V\) represents the matrix of eigenvectors with dimensions \(n\times d\), \(y\) i and \(y\)d Both represent the data representation after dimensionality reduction in the wind tunnel test. k(x i ) represents the column vector formed by the kernel function values between the wind tunnel test data x i and all data. k(x i , x n ) represents the similarity between the wind tunnel test data x i and x n in the high-dimensional feature space. v d represents the d-th eigenvector, representing the d-th direction of the principal component KPCA projection space.
[0100] Map the wind tunnel test data after dimensionality reduction back to the original space, and construct the KPCA anomaly score according to the reconstruction error:
[0101]
[0102] where φ(x i ) represents the reconstruction representation of the wind tunnel test data y i in the high-dimensional space. y i'j' represents the projection value of the i'-th wind tunnel test data in the reduced-dimensional matrix Y in the j'-th principal component direction. || ||2 represents the Euclidean distance. v j” represents the j''-th eigenvector of the kernel matrix K' in the principal component KPCA. d represents the target dimension after dimensionality reduction. KPOS(x i ) represents the KPCA anomaly score of the wind tunnel test data x i .
[0103] In this embodiment, KPCA is the core method for non-linear feature extraction and anomaly detection, which refers to the process of mapping the original data to the high-dimensional feature space through the kernel trick K ij = φ(x i ) T φ(x j ) and performing principal component analysis in this space. For any (x i , x j ) ∈ U×U, K ij represents the inner product similarity between the wind tunnel test data samples x i and x j in the high-dimensional feature space. The KPCA anomaly score is used to measure the degree to which a sample deviates from the principal component subspace in the high-dimensional feature space.
[0104] S3. Normalize the wind tunnel test data after dimensionality reduction, and calculate the fuzzy relation matrix using the hybrid fuzzy relation similarity;
[0105] In this embodiment, the fuzzy relation is the core concept of the fuzzy rough set theory, which refers to a fuzzy set \(R: U\times U\rightarrow[0,1]\) defined on \(U\times U\). For any \(R: U\times U\rightarrow[0,1]\), the membership function \(R(x,y)\) represents the degree to which the object \(x\) has a relationship \(R\) with the object \(y\). The fuzzy relation \(R\) on \(U\) can be represented by a fuzzy matrix \(M(R)\). Any attribute \(c\) k \(\in C\) can induce a fuzzy relation whose membership calculation formula is:
[0106]
[0107] Any subset of attributes can also induce a fuzzy relation \(R\) B , whose membership calculation formula is:
[0108]
[0109] wherein, represents the fuzzy similarity between the wind tunnel test data \(x\) i and \(x\) j with respect to the attribute \(c\) k , \(|x\) i (c k ) - x\) j (c k )| represents the difference between the wind tunnel test data \(x\) i and \(x\) j on the attribute \(c\) k , represents the threshold parameter, \(std(c k )\) represents the standard deviation of the attribute \(c\) k , \(\delta\) represents the adjustable parameter, \(R\) B (x i , x j ) represents the fuzzy relation generated by the attribute subset \(B\), \(x\) j (c k ) represents the value of the wind tunnel test data \(x\) j on the attribute \(c\) k , \(x\) i (c k ) represents the value of the wind tunnel test data \(x\) i on the attribute \(c\) k .
[0110] For the data after the above normalization processing, the fuzzy relation matrix \(M(R B )\) of any attribute subset \(B\) is calculated.
[0111] S4. Calculate the fuzzy granule sets of each attribute according to the fuzzy relation matrix;
[0112] In this embodiment, based on the fuzzy rough set theory, a multi-modal data feature extraction method is proposed. Fuzzy information granules (or information granules for short) are a set of objects induced by the fuzzy relationships between objects. The features of multi-modal data are extracted by constructing fuzzy relationships at multiple granularity levels, which specifically include the following steps:
[0113] Step 1: Selection of attribute subsets
[0114] Any set of attribute subsets can construct a group of fuzzy relationships, and through multiple fuzzy relationships, the fuzzy information granules of objects can be further constructed. To simplify the calculation, only one attribute is used to construct the information granules (i.e., basic fuzzy information granules). One attribute c is successively selected from all attributes C = {c1, c2,..., c m} to form an attribute subset B k = {c k}, thus obtaining m several attribute subsets k}
[0115] Step 2: Information granulation
[0116] The process of constructing information granules is also called information granulation. Information granulation is the process of dividing all objects into particles of different granularities (i.e., grain sizes). In the fuzzy rough set theory, a group of objects can be aggregated through the fuzzy relationships between objects to form information granules. Using any fuzzy relationship R B to partition all objects U, a set of generalized fuzzy equivalence classes is obtained, that is, a set of multi-granularity fuzzy information granules:
[0117] G U / R B = {[x1] B , [x2] B ,..., [x i B ,..., [x n B}
[0118] where G U / R B represents the set of fuzzy granules constructed under the fuzzy relationship R B , [x i B represents the fuzzy information granule centered on the wind tunnel test data x i generated by the similarity of the mixed fuzzy relationship, [x n B represents the fuzzy information granule centered on the wind tunnel test data x n generated by the similarity of the mixed fuzzy relationship, U represents the set of wind tunnel test data, B represents the attribute subset, and R B Denote the fuzzy relation generated by the attribute subset B, G U Denote the set of fuzzy information granules constructed on the wind tunnel test data set U. Obviously, the information granule [x i B is a fuzzy set on the set of all objects U, and its membership function is:
[0119]
[0120] The information granule [x i B The calculation formula for the base is:
[0121]
[0122] Obviously, it can be found that 1 ≤ [x i B ≤ |U|.
[0123] Step 3: Construction of multi-granularity fuzzy information granules
[0124]
[0125] The multi-granularity fuzzy information granule is a set of fuzzy information granules with multiple information granules. For any sample x i , respectively under m attribute subsets B c ∈ B, a set of multi-granularity fuzzy information granules of this sample can be obtained through information granulation.
[0126] S5. According to the set of fuzzy granules, calculate the fuzzy granule density of each object with respect to each attribute, and calculate the fuzzy entropy and its corresponding weight. The implementation method is as follows:
[0127] According to the set of fuzzy granules, calculate the fuzzy granule density of each object with respect to each attribute:
[0128]
[0129] According to the fuzzy granule density, calculate the fuzzy entropy and its corresponding weight:
[0130]
[0131] Among them, W(c k ) represents the weight function of the wind tunnel test data on the attribute c k , FE(c k ) represents the calculation function of the fuzzy entropy of the fuzzy relation matrix , FE(C) represents the calculation function of the fuzzy entropy of the fuzzy relation matrix R C , C represents the set of all attributes with multiple attributes, k represents the wind tunnel test data attribute number, U represents the wind tunnel test data set, x i Denote the wind tunnel test data, |x i C |The cardinality of the fuzzy information granule centered on the wind tunnel test data x i generated by the mixed fuzzy relation similarity degree, Denote the wind tunnel test data x i on the attribute c k The fuzzy granule density, |U| denotes the cardinality of the wind tunnel test data, Denote the cardinality of the fuzzy granule .
[0132] In the present invention, by introducing the fuzzy granule density, the accurate quantification of the local distribution characteristics of the data is realized. The adaptive weight assignment mechanism based on fuzzy entropy overcomes the limitation of manually setting the attribute weights in the traditional method and realizes the objective evaluation of the importance of different attributes.
[0133] In this embodiment, the fuzzy entropy is the core concept of the fuzzy set theory and refers to an important metric for quantifying the uncertainty of the fuzzy information granule. The larger the fuzzy entropy, the lower the discrimination degree of the attribute c k , that is, it has the greatest uncertainty.
[0134] In the above definition, the weight function W(c k ) uses the fuzzy entropy to realize the adaptive assignment of the attribute importance value. An attribute containing the maximum uncertainty will provide more abnormal features and will be given a greater weight in the abnormal score.
[0135] S6. Calculate the fuzzy granule abnormal scores of all wind tunnel test data according to the fuzzy granule density and its corresponding weight:
[0136]
[0137] where, FGOS(x i ) denotes the fuzzy granule abnormal score of the wind tunnel test data x i , C denotes the set of all attributes with multiple attributes, W(c k ) denotes the weight function of the wind tunnel test data on the attribute c k , Denote the wind tunnel test data x i on the attribute c k The fuzzy granule density, and k denotes the wind tunnel test data attribute number.
[0138] In this embodiment, the fuzzy granule abnormal score is used to measure the deviation degree of a sample relative to the overall data distribution, and the abnormal features from the fuzzy information granules of multiple granularities are fused by weighted summation of the multi-granularity information granule abnormality degrees.
[0139] S7. Perform weighted fusion on the KPCA anomaly score and the fuzzy granule anomaly score to obtain the fused anomaly score:
[0140] KFDOS(x i ) = KPOS(x i ) × FGOS(x i )
[0141] where KFDOS(x i ) represents the fused anomaly score of the wind tunnel test data x i , FGOS(x i ) represents the fuzzy granule anomaly score of the wind tunnel test data x i , and KPOS(x i ) represents the KPCA anomaly score of the wind tunnel test data x i .
[0142] In this embodiment, the anomaly score fusing KPCA and fuzzy granule density is used to measure the comprehensive anomaly degree of the sample in the multi-scale feature space. By simply multiplying and fusing the KPCA anomaly score and the fuzzy granule anomaly score, the collaborative detection of global non-linear anomaly and local density anomaly is achieved.
[0143] S8. Judging one by one whether the fused anomaly score of the wind tunnel test data is greater than the threshold. If so, output the abnormal data and repeat S8 until all the wind tunnel test data have been judged. Otherwise, regard it as normal data and judge the next wind tunnel test data until all the wind tunnel test data have been judged.
[0144] In this embodiment, let the outlier threshold of the abnormal point be ξ ∈ (0, 1). If the anomaly score KFDOS(x i ) of the sample x i ) > ξ, then x i is determined to be abnormal wind tunnel test data. By comparing the anomaly scores of all samples with the threshold ξ one by one, all the abnormal wind tunnel test data can be calculated.
[0145] In this embodiment, an information table containing 6 samples and 2 attributes is given, as shown on the left side of Table 1 (the right side is the normalized result). Table 1 is the table of original data and normalized results.
[0146] Table 1
[0147]
[0148] Step 1: Data normalization processing.
[0149] For subsequent calculations, first perform min-max normalization on the input data to convert the numerical data to between 0 and 1. The nominal attribute c3 remains unchanged. The processing results are shown on the right side of Table 1.
[0150] Step 2: Use KPCA for data dimensionality reduction and calculate the KPCA anomaly score.
[0151] Set the RBF kernel parameter γ = 0.1, calculate the kernel matrix K and centralize it to obtain the centralized kernel matrix K'.
[0152]
[0153] Calculate the eigenvalues and sort them from largest to smallest to obtain the set of non-zero eigenvalues Λ and the explained variance:
[0154] Λ = (λ1, λ2, λ3, λ4) ≈ (0.5022, 0.1569, 0.0152, 0.0083)
[0155]
[0156] CumulativeVar = (0.7356, 0.9654, 0.9877, 1.0000)
[0157] Among them, ExplanionVar represents the explained variance, and CumulativeVar represents the cumulative variance.
[0158] Determine the target dimension after dimensionality reduction according to the Bayesian Information Criterion (BIC):
[0159] n = min(|Λ|, |C|) = min(4, 3) = 3
[0160] Use:
[0161]
[0162] We can get:
[0163]
[0164] The selected dimension d = argmin(BIC(i)) = argmin(1.9214, 1.5270, 2.1290) = 2.
[0165] Set the RBF kernel parameter γ = 0.1, select the dimension d = 2 after dimensionality reduction, and the result of dimensionality reduction of the original data is:
[0166]
[0167] The result after reconstructing the data after dimensionality reduction is:
[0168]
[0169] By calculating the Euclidean distance between the original data and the reconstruction error, the KPCA anomaly score is obtained as follows:
[0170]
[0171] KPOF(x2) ≈ 0.7693
[0172] KPOF(x3) ≈ 1.4381
[0173] KPOF(x4) ≈ 0.8681
[0174] KPOF(x5) ≈ 0.4686
[0175] Step 3: Calculate the fuzzy relation matrix.
[0176] Let the fuzzy similarity radius δ = 0.5, and calculate the fuzzy relation matrices of attributes c1 and c2 respectively:
[0177]
[0178] Step 4: Construct fuzzy information granules.
[0179] Describe the uncertainty and anomaly characteristics of the data by constructing fuzzy information granules. For the fuzzy relation matrix Construct a fuzzy information granule for each row of the matrix:
[0180] Taking the sample x1 as an example below, from the fuzzy relation matrix It can be obtained that: From the fuzzy relation matrix It can be obtained that
[0181] Step 5: Calculate the fuzzy granule density, fuzzy entropy and their corresponding weights.
[0182] By calculating the density of the fuzzy information granule, the distribution compactness of the data in the local attribute space can be quantified, and the fuzzy entropy measures the uncertainty of the information it contains. The larger the fuzzy entropy, the lower the discrimination degree of attribute c k That is, it has the greatest uncertainty, so a higher weight is assigned. Calculate the density of each fuzzy granule as follows:
[0183]
[0184] Similarly, it can be obtained that:
[0185]
[0186] Calculate the fuzzy entropy of calculated attributes c1 and c2:
[0187]
[0188] FE(C2) ≈ 0.4991
[0189] Furthermore, obtain the weights of attributes c1 and c2:
[0190]
[0191] W(c1) ≈ 0.4131
[0192] Step 6: Calculate the fuzzy granule anomaly scores of all samples.
[0193] Calculate the fuzzy granule anomaly score of a sample by performing a weighted sum of the fuzzy granule density of a set of samples and the corresponding weights:
[0194]
[0195] Similarly, we can obtain:
[0196] FGOS(x2) ≈ 0.3725
[0197] FGOS(x3) ≈ 0.4415
[0198] FGOS(x4) ≈ 0.3711
[0199] FGOS(x5) ≈ 0.2932
[0200] Step 7: Calculate the fused KPCA and fuzzy granule anomaly scores of all samples.
[0201] Since there are differences in the numerical ranges and distribution characteristics between the KPCA anomaly scores (KPOS) and the fuzzy granule anomaly scores (FGOS), standardization is required to ensure fairness and comparability during fusion. Considering the subsequent form of product fusion and to avoid the excessive influence of extreme values on the results, we use a combination of the min-max normalization method and the Sigmoid function for standardization to improve the smoothness and robustness of the results. The normalized anomaly scores are:
[0202] KPOS(x1) ≈ 0.5157, FGOS(x1) ≈ 0.5000
[0203] KPOS(x2) ≈ 0.5769, FGOS(x2) ≈ 0.6663
[0204] KPOS(x3) ≈ 0.7311, FGOS(x3) ≈ 0.7311
[0205] KPOS(x4) ≈ 0.6016, FGOS(x4) ≈ 0.6649
[0206] KPOS(x5) ≈ 0.5000, FGOS(x5) ≈ 0.5834
[0207] Multiply them to obtain the fused anomaly score:
[0208] KFDOS(x1) = KPOS(x1) × FGOS(x1) ≈ 0.5157 × 0.5000 ≈ 0.2579
[0209] KFDOS(x2) ≈ 0.3844
[0210] KFDOS(x3) ≈ 0.5344
[0211] KFDOS(x4) ≈ 0.3999
[0212] KFDOS(x5) ≈ 0.2917
[0213] Step 8: Perform anomaly determination by comparing with a threshold.
[0214] Compare the anomaly scores of all samples. Obviously, the anomaly score of sample x3 of y is significantly higher than that of other samples. Let the anomaly score threshold ξ be 0.50. Compare the anomaly scores of all samples with the threshold ξ. Then x3 is determined as abnormal wind tunnel test data, output the information of sample x3, and remind the wind tunnel test personnel that there is an anomaly in this sample and further investigation measures need to be taken.
Claims
1. An abnormal detection method for wind tunnel test data that combines KPCA and fuzzy granular density, characterized in that, It includes the following steps: S1. Obtain the wind tunnel test data, and perform normalization processing on the wind tunnel test data to obtain the normalized wind tunnel test data; S2. According to the normalized wind tunnel test data, use KPCA for data dimensionality reduction, and construct a KPCA anomaly score based on the reconstruction error; S3. Perform normalization processing on the dimensionality-reduced wind tunnel test data, and calculate the fuzzy relation matrix using the mixed fuzzy relation similarity; S4. Calculate the fuzzy granule set of each attribute according to the fuzzy relation matrix; S5. Calculate the fuzzy granule density of each object with respect to each attribute according to the fuzzy granule set, and calculate the fuzzy entropy and its corresponding weight; S6. Calculate the fuzzy granule anomaly score of all wind tunnel test data according to the fuzzy granule density and its corresponding weight; S7. Perform weighted fusion on the KPCA anomaly score and the fuzzy granule anomaly score to obtain the fused anomaly score; S8. Judging one by one whether the fused anomaly score of the wind tunnel test data is greater than the threshold. If so, output the abnormal data, and repeat S8 until all wind tunnel test data have been judged. Otherwise, it is regarded as normal data, and the next wind tunnel test data is judged until all wind tunnel test data have been judged.
2. The abnormal detection method for wind tunnel test data integrating KPCA and fuzzy granular density according to claim 1, characterized in that The specific content of S2 is as follows: Set the radial basis function kernel parameter δ; Based on the normalized wind tunnel test data, calculate the kernel matrix according to the radial basis function kernel; Perform centering processing on the kernel matrix generated by the radial basis function kernel; According to the obtained centered kernel matrix, perform eigenvalue decomposition on it to obtain the eigenvectors belonging to the eigenvalues, and determine the target dimension d after dimensionality reduction according to the Bayesian information criterion; According to the eigenvalues and eigenvectors, select the first d largest eigenvalues to form a subspace after dimensionality reduction, and obtain the wind tunnel test data after dimensionality reduction; Map the wind tunnel test data after dimensionality reduction back to the original space, and construct a KPCA anomaly score based on the reconstruction error.
3. The abnormal detection method for wind tunnel test data integrating KPCA and fuzzy granular density according to claim 2, characterized in that, The expression of the kernel matrix is as follows: X = [x1, x2,..., x n Among them, X represents a matrix, and each column represents a wind tunnel test data, x n represents the nth normalized wind tunnel test data, φ(x) represents a non-linear mapping function, which maps the variable x in the matrix X i to a high-dimensional space, H represents the dimension after mapping, K represents the dimension before mapping, K ij represents the value of the i-th row and j-th column of the kernel matrix, k() represents a similarity function, k(x i , x j ) represents the similarity of the wind tunnel test data x i and x j in the high-dimensional space, γ represents the radial basis function kernel parameter, ||x i - x j ||2 represents the Euclidean distance between the wind tunnel test data x i and x j ; represents the K-dimensional real vector space, represents the H-dimensional real vector space; The expression of the centered kernel matrix is as follows: Among them, K' represents the centered kernel matrix, φ'(X) represents the matrix obtained by centering the wind tunnel test data matrix in the high-dimensional feature space, and φ'(X) T represents the matrix obtained by transposing the centered matrix, 1 n×1 represents the n-dimensional column vector introduced for the simplified expression, with each element being 1, 1 n represents an n×n matrix, with each element being K represents the kernel matrix, which is used to represent the similarity relationship of samples in the high-dimensional feature space; The expression of the eigenvector is as follows: K'v = λv where v represents the eigenvector belonging to λ, and λ represents the eigenvalue of K'; The expression of the wind tunnel test data after dimensionality reduction is as follows: Y = V T K Y = [y1, y2,..., y i ,..., y d V = [v1, v2,..., v d T k(x i ) = [k(x i , x1), k(x i , x2),..., k(x i , x j ),..., k(x i , x n )] T Among them, Y represents the matrix after dimensionality reduction of d×n dimensions, V represents the eigenvector matrix of n×d dimensions, y i and y d both represent the representation of the wind tunnel test data after dimensionality reduction. k(x i ) represents the column vector composed of the kernel function values between the wind tunnel test data x i and all data. k(x i , x n ) represents the similarity between the wind tunnel test data x i and x n in the high-dimensional feature space. v d represents the d-th eigenvector, representing the d-th direction of the principal component KPCA projection space; The expression of the KPCA anomaly score is as follows: Among them, φ(x i ) represents the reconstructed representation of the wind tunnel test data y i in the high-dimensional space. y i'j' represents the projection value of the i'-th wind tunnel test data in the j'-th principal component direction in the matrix Y after dimensionality reduction. ||||2 represents the Euclidean distance. v j” represents the j''-th eigenvector of the kernel matrix K' in KPCA. d represents the target dimension after dimensionality reduction. KPOS(x i ) represents the KPCA anomaly score of the wind tunnel test data x i .
4. The abnormal detection method for wind tunnel test data integrating KPCA and fuzzy granular density according to claim 1, characterized in that The expression of the fuzzy relation matrix is as follows: Among them, represents the wind tunnel test data x i and x j with respect to the property c k of the fuzzy similarity, |x i (c k ) - x j (c k )| represents the difference between the wind tunnel test data x i and x j on the property c k . represents the threshold parameter, std(c k ) represents the standard deviation of the property c k , δ represents the adjustable parameter, R B (x i , x j ) represents the fuzzy relationship generated by the attribute subset B, x j (c k ) represents the value of the wind tunnel test data x j on the attribute k , x i (c k ) represents the value of the wind tunnel test data x i on the attribute c k .
5. The abnormal detection method for wind tunnel test data integrating KPCA and fuzzy granular density according to claim 1, characterized in that The expression of the fuzzy granule set is as follows: G U / R B = {[x1] B , [x2] B ,..., [x i B ,..., [x n B} Among them, G U / R B denotes the fuzzy granule set constructed under the fuzzy relation R B [x i B denotes the fuzzy information granule centered on the wind tunnel test data x i generated by the similarity of the hybrid fuzzy relation, [x n B denotes the fuzzy information granule centered on the wind tunnel test data x n generated by the similarity of the hybrid fuzzy relation. U represents the set of wind tunnel test data, B represents the subset of attributes, R B denotes the fuzzy relation generated by the subset of attributes B, G U denotes the set of fuzzy information granules constructed on the set of wind tunnel test data U, [x i B (x j ) is the membership function, indicating that the wind tunnel test data x i and x j have the fuzzy relation R B to a certain degree. R B (x i ,x n ) represents the fuzzy similarity between the wind tunnel test data x i and x n under the attribute set B. R B (x i ,x j ) represents the fuzzy similarity between the wind tunnel test data x i and x j under the attribute set B. denotes the element in the i-th row and n-th column of the fuzzy relation matrix R B , denotes the element in the i-th row and j-th column of the fuzzy relation matrix R B . 6. The abnormal detection method for wind tunnel test data by integrating KPCA and fuzzy granular density according to claim 1, characterized in that The specific content of S5 is as follows: Calculate the fuzzy granule density of each object with respect to each attribute according to the fuzzy granule set; Calculate the fuzzy entropy and its corresponding weight according to the fuzzy granule density.
7. The abnormal detection method for wind tunnel test data integrating KPCA and fuzzy granular density according to claim 6, characterized in that The expression of the weight is as follows: Among them, W(c k ) represents the weight function of the wind tunnel test data on the attribute c k , C represents the set of all attributes with multiple attributes, k represents the wind tunnel test data attribute number, U represents the wind tunnel test data set, x i represents the wind tunnel test data, |[x i C | is the cardinality of the fuzzy information granule centered on the wind tunnel test data x i generated by the mixed fuzzy relation similarity, represents the fuzzy granule density of the wind tunnel test data x i on the attribute c k , |U| represents the cardinality of the wind tunnel test data, represents the cardinality of the fuzzy granule , FE(c k ) represents the calculation function of the fuzzy entropy of the fuzzy relation matrix , FE(C) represents the calculation function of the fuzzy entropy of the fuzzy relation matrix R C . 8. The abnormal detection method for wind tunnel test data integrating KPCA and fuzzy granular density according to claim 1, characterized in that The expression of the fuzzy granule anomaly score is as follows: Among them, FGOS(x i ) represents the fuzzy granular anomaly score of the wind tunnel test data x i , C represents the set of all attributes with multiple attributes, W(c k ) represents the weight function of the wind tunnel test data on the attribute c k , represents the fuzzy granular density of the wind tunnel test data x i on the attribute c k .
9. The abnormal detection method for wind tunnel test data integrating KPCA and fuzzy grain density according to claim 1, characterized in that The expression of the fused anomaly score is as follows: KFDOS(x i ) = KPOS(x i ) × FGOS(x i ) Among them, KFDOS(x i ) represents the fusion anomaly score of the wind tunnel test data x i , FGOS(x i ) represents the fuzzy granule anomaly score of the wind tunnel test data x i , and KPOS(x i ) represents the KPCA anomaly score of the wind tunnel test data x i .