Multi-level Track Defect Identification System Based on Vehicle Body Vibration Data
By adopting a multi-task cascading CNN architecture with spectral clustering and attention-guided in the orbital disease recognition system, the problem of orbital disease evaluation in complex environments is solved, and efficient and accurate multi-level disease recognition is achieved.
Patent Information
- Application Number
- CN202510190606.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The existing orbital state evaluation methods are difficult to achieve the real state evaluation of orbital diseases in complex environments, and the portable detection equipment is limited in obtaining full-line vibration data, resulting in a distribution difference between the vibration data and structural disease labels.
A multi-level track disease identification system based on vehicle body vibration data is adopted, including data acquisition module, environmental adaptive module, label enhancement module and disease identification module. Through a multi-task cascading CNN architecture with spectral clustering and attention-guided attention-guided, adaptive classification of orbital environments and multi-level disease recognition are realized.
It significantly improves the disease recognition effect in complex orbital environments, improves the accuracy and efficiency of detection, solves the problem of label sparseness, and realizes multi-level parallel recognition of orbital geometric and structural diseases.
Smart Images

Figure CN119669869B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of blasting auxiliary devices, and particularly to a multi-level track disease identification system based on vehicle body vibration data. Background Art
[0002] The timely detection and maintenance of track diseases are crucial for ensuring the safety of railway operations. With the rapid development of urban rail transit, higher requirements for timeliness and accuracy are imposed on track disease detection. Currently, track diseases mainly include two levels: geometric diseases and structural diseases. The geometric state of the track is mainly represented by the following geometric parameters: longitudinal level, alignment, cross-level, gauge, twist, etc. Each country has stipulated the deviation management values of track geometric parameters and classified the deviation degrees according to management requirements. The state exceeding these management values is a track geometric disease. Structural diseases mainly include defects on the surface of the rail (such as side wear, vertical wear, etc.). Similar to geometric diseases, each country has formulated corresponding detection standards and management systems. The inspection unit needs to classify and evaluate the diseases according to the specific damage type, location, and degree. These diseases will directly affect the safety of vehicle operation and the comfort of passengers.
[0003] Traditional track detection mainly relies on inspection vehicles and manual inspections. The inspection vehicle can combine various detection technologies such as ultrasonic and eddy current to accurately evaluate the track state. However, these detection methods can usually only be carried out during the maintenance window period, with limited detection frequency and high detection costs. Due to the significant correlation between the vehicle body vibration and the track state, high-frequency monitoring of the track state can be achieved by collecting vibrations through portable on-vehicle devices. To improve the detection efficiency, the track state evaluation method based on on-vehicle vibration data has gradually attracted attention.
[0004] Currently, the track state evaluation method based on on-vehicle vibration data still faces some technical problems in practical applications: The track infrastructure is widely distributed, and parameters such as the line profile, track composition, and operating speed vary significantly in space, forming a complex and diverse track environment. The heterogeneity of this environment makes it difficult for traditional disease identification methods to maintain stable detection effects under different line conditions. In addition, the popularization and application of portable detection devices have made it possible to obtain vibration data for the entire line. However, limited by human and time costs, on-site verification and annotation work are still mainly concentrated in key sections. This unbalanced data acquisition method leads to a significant distribution difference between vibration data and structural disease labels.
[0005] Current research mainly focuses on a single heterogeneous factor. The difference in the support stiffness of different types of track bed structures will lead to a frequency-dependent damping effect, thus affecting the mapping relationship between vehicle body vibration and track conditions. In the study of curve sections, it is found that a stronger self-excited tangential force will be generated at the wheel-rail contact interface, forming unique vibration characteristics. Although the research on these single factors reveals the basic laws of the vibration-disease relationship, it is difficult to comprehensively reflect the influence of complex environments. Therefore, it is difficult to evaluate the true state of track diseases by combining relevant parameters. Summary of the Invention
[0006] The present invention aims to provide a multi-level track disease identification system based on vehicle body vibration data, which solves the problem that the existing track condition evaluation methods cannot evaluate the true state of track diseases in complex environments.
[0007] To achieve the above object, the technical solution of the present invention is as follows: A multi-level track disease identification system based on vehicle body vibration data, comprising:
[0008] A data acquisition module for collecting heterogeneous data of the track environment;
[0009] An environment adaptation module for analyzing the correlation between vehicle body vibration data and heterogeneous data, constructing a dual-perspective similarity matrix with differentiated weights, and realizing the adaptive classification of the track environment through spectral clustering methods;
[0010] A label enhancement module, through the self-training method of the spectral normalized Gaussian process model, calculates the multi-dimensional uncertainty index of the samples, and gradually generates high-quality pseudo-labels through a dynamic threshold mechanism, and finally forms a complete labeled sample set;
[0011] A disease identification module, adopting an attention-guided multi-task cascaded CNN architecture, that is, detecting the disease location through a binary classification layer, filtering the detection results of the binary classification layer by means of a mask layer, evaluating the disease level through a multi-classification layer, and finally outputting a multi-level identification result including the disease location and severity.
[0012] Furthermore, the basic rules for dataset construction are as follows:
[0013] T = (a, b, l)|a ∈ A, b ∈ B, l ∈ L, m(a) = m(b) = m(l);
[0014] Wherein, A is the vehicle body vibration dataset, and each sampling point contains 6 vibration parameters A i = [a 1 ,..., a 6 , specifically: a 1 represents the longitudinal acceleration, a 2 represents the lateral acceleration, a 3Denote the vertical acceleration as \(a\). 4 Denote the longitudinal angular velocity as \(a\). 5 Denote the lateral angular velocity as \(a\). 6 Denote the vertical angular velocity; \(B\) is the dataset of heterogeneity factors, and each sampling point contains \(n\) environmental parameters \(B\). i \( = [b\) 1 ,..., \(b\) n ; To eliminate the dimensional differences between different parameters, [-1, 1] interval normalization is adopted; For the non-numerical ballast type parameter, numerical mapping is performed according to the following rules: the integral ballast is assigned -1, the rubber isolation pad ballast is assigned -0.33, the double-layer non-linear fastener ballast is assigned 0.33, and the steel spring floating slab ballast is assigned 1; \(L\) is the disease label set, and the label information of each sampling point is represented by \(L\). i \( = [l\) 1 ,..., \(l\) δ , where \(l\) 1-δ corresponds to \(\delta\) different types of track disease label values respectively;
[0015] To obtain statistically representative training samples, the method of constructing grid samples is as follows:
[0016] \(G = g\) 1 ,..., \(g\) i ,..., \(g\) n ;
[0017]
[0018] where \([·]\) represents the floor operation, \(g\) i represents the \(i\)-th grid sample, \(n\) is the total number of grids, which is determined by the total number of sampling points \(N\) and the number of sampling points \(m\) in a single grid: \(n = [N / m]\), and represent the vibration characteristics and heterogeneity factor information of the \(j\)-th sampling point in the \(i\)-th grid respectively.
[0019] Furthermore, the process of the spectral clustering algorithm is as follows: Analyze the correlation strength between the vehicle body vibration and the data heterogeneity data to determine the feature weights, and design an improved dual-perspective similarity matrix based on this to make full use of the differential information of the weights, and then perform spectral clustering to divide the line grid sample dataset \(G\) into different categories.
[0020] Furthermore, the method of constructing the dual-perspective similarity matrix is as follows: Divide the feature space into a high-weight perspective \(V\) h and a low-weight perspective \(V\) l based on the median of the feature weights. The high-weight perspective adopts larger connection parameter \(\alpha_V\) h and structure protection parameter \(\beta_V\) h, allowing more connections to be established between samples and protecting local structural characteristics; the low-weight perspective uses smaller connection parameter αV l and structural protection parameter βV l , restricting the connection between samples and reducing the dependence on local structures, and characterizing the local structural association through local density and reconstruction relationship within each perspective:
[0021] The local density is measured by the overlap degree of the k-nearest neighbor sets of samples:
[0022]
[0023] where N(i) represents the number d of the k-nearest neighbors of sample i, and d for each perspective is adjusted by α: d = d 0 ×(1 + αw s ), where d 0 takes 10;
[0024] The reconstruction relationship is calculated by normalizing the distance between samples:
[0025]
[0026] Design a Gaussian kernel function within the perspective based on local indicators:
[0027]
[0028] where w s is the feature weight vector, σ is the local adaptive bandwidth, taking the average distance of the k-nearest neighbors of the sample, and L(x, y) is the local structure protection term, composed of local density and reconstruction relationship.
[0029] Furthermore, fuse the similarity matrices to balance the contributions of the two perspectives to the final similarity:
[0030] S h +(1 - y)S l ;
[0031] where S h and S l represent the similarity matrices of the two perspectives respectively, γ is the fusion coefficient, used to control the dominant degree of the high-weight perspective; construct the degree matrix D based on the fused similarity matrix S:
[0032]
[0033] where the non-diagonal elements are all 0, calculate the Laplacian matrix L p using the symmetric normalization method:
[0034]
[0035] Among them, \(I\) is the identity matrix. The symmetric normalization method can maintain the geometric structure of high-dimensional data while having the advantages of high computational stability and effectively processing data of different scales. Calculate the first \(k\) eigenvectors \(\{v p _{1}\), \(v 1 _{2}\), \(\cdots\), \(v 2 _{k}\}\) corresponding to the smallest eigenvalues of \(L\), and construct the eigenvector matrix \(V\) with them as column vectors. Normalize each row of \(V\) to unit length to obtain the matrix \(Y\): k} and use them as column vectors to construct the eigenvector matrix \(V\). Normalize each row of \(V\) to unit length to obtain the matrix \(Y\):
[0036]
[0037] Cluster each row in \(Y\) as a \(k\)-dimensional point, and the category of sample \(i\) is the cluster it finally belongs to.
[0038] Furthermore, the pseudo-label generation method includes a spectral normalization Gaussian process model and an adaptive threshold training mechanism:
[0039] Spectral normalization Gaussian process model:
[0040] Use the Gaussian kernel function to measure the similarity between two points in the input space, which decays exponentially as the distance between the points increases:
[0041]
[0042] where \(l\) is the length scale parameter, and the initial value is set to \(0.1\); RFF approximates this kernel function through the following mapping function to approximate this kernel function:
[0043]
[0044] where \(W\) is a random projection matrix sampled from the normal distribution and \(b\) is a random bias vector sampled from the uniform distribution \(\cup[0, 2\pi]\), and \(D f is the feature dimension (512); through this mapping, the inner product of any two inputs \(x\) and \(x'\) satisfies:
[0045]
[0046] The predicted value is obtained through the linear combination of the mapped features, and the uncertainty can be calculated from the kernel matrix in the feature space;
[0047] Multi-dimensional uncertainty evaluation and adaptive threshold training mechanism:
[0048] Design a multi-dimensional uncertainty evaluation value μ based on the predicted value and uncertainty value output by the spectral normalization Gaussian process model, and comprehensively consider four metrics: E is entropy, representing distribution uncertainty; M is the marginal difference; C is the confidence level, taking the highest prediction probability; V is the prediction variance, representing the model stability;
[0049] E = -∑(p i × log(p i ));
[0050] M = p m - p s ;
[0051] C = max(p i );
[0052] V = diag(K)-∑p i 2 ;
[0053] μ = w 1 × E + w 2 × (1 - M)+ w 3 × (1 - C)+ w 4 × V;
[0054] Among them, diag(K) represents the diagonal elements of the kernel matrix, p i represents the predicted probability that the sample belongs to the i-th class, p m represents the highest predicted probability of the sample, p s represents the second-highest predicted probability of the sample; w 1 , w 2 , w 3 and w 4 are the weights of entropy, marginal difference, confidence level, and prediction variance respectively;
[0055] By determining whether the multi-dimensional uncertainty evaluation value μ exceeds the dynamic threshold to select the training set to be added to the next iteration, the specific method is as follows: Set the dynamic low threshold μ l , μ lower than μ l indicates that the sample is predicted to be a certain known class sample with a high certainty, that is, the certainty of the actual existence of the disease is relatively large, then set the sample predicted value as the pseudo-label; Set the dynamic high threshold μ h , μ higher than μ h indicates that the sample may have characteristics different from the known label samples and is very likely to belong to a new class, that is, the actual situation of no disease, then set the sample predicted value as the new class '0'.
[0056] Furthermore, the method for setting the initial threshold is as follows:
[0057]
[0058] Among them, μ and σ are the mean and standard deviation of the model uncertainty estimation respectively; the 3σ interval setting provides a reliable boundary in terms of statistics. Considering that the model prediction ability gradually improves with training iterations, the threshold needs to be adjusted accordingly to adapt to the change of model performance:
[0059]
[0060] Among them, is the threshold of the previous iteration, α·ΔC·σ(μ)<0.5σ(μ); is the average confidence of the newly labeled samples, and 0.95 is the baseline confidence; after reaching the predetermined number of iterations, the mature model is used to predict the remaining sample categories, and the samples of the new category in the remaining samples are screened according to the last threshold.
[0061] Furthermore, the binary classification layer adopts a weighted binary cross-entropy loss function, and by introducing a penalty factor w, the penalty for false negative samples is increased:
[0062]
[0063] Among them, y is the true label of each sample, that is, y = 1 indicates the presence of disease, y = 0 indicates the absence of disease, and y p is the predicted value, which is between 0 and 1; w is the penalty factor value. When y = 1, this binary cross-entropy loss function will be activated, and the closer the predicted value is to 0, the larger the loss function value.
[0064] Furthermore, the mask layer includes a screening layer and a multi-classification input layer. The screening layer applies a non-zero mask to the output of the binary classification layer, retains the positive predictions in the output, that is, the predicted value > 0.5, and sets all other predicted values to zero; the filtered output of the binary classification layer is multiplied element-wise with the original input to screen out the samples considered to have diseases by the binary classification layer; the screened samples form the multi-classification input layer to train the multi-classification layer to evaluate the disease level.
[0065] Furthermore, channel attention mechanism is adopted to weight the features of the output of the convolutional block. The channel attention weights in the first stage are non-linearly transformed through feature remapping, and sequentially pass through the dimension adjustment of the fully connected network and the ReLU activation function; the normalized guiding weights are generated through the sigmoid function, and the guiding weights are multiplied element-wise with the channel attention output in the multi-classification stage; the weighted features pass through the multi-layer fully connected network for feature transformation, and dropout is used in each layer to prevent overfitting; finally, according to the number of levels of different diseases, the multi-classification prediction results are output through the softmax function; the standard cross-entropy loss function is adopted in this stage, and is weighted and added to the loss function of the binary classification layer to jointly constitute the overall optimization objective of the model:
[0066] l t = α l × l b + β l × l m ;
[0067] where l b is the binary classification loss, and l m is the multi-classification loss, and α l and β l are the corresponding weight coefficients.
[0068] Compared with the prior art, the beneficial effects of this solution are as follows:
[0069] 1. The present invention proposes a correlation-driven dual-view spectral clustering algorithm, which realizes the adaptive classification of the track environment. Compared with the traditional method that only analyzes a single heterogeneity factor, this algorithm analyzes the correlation between vibration and multiple heterogeneity factors, and with the dual-view design of differential weights, it significantly improves the classification effect in a complex track environment, and the silhouette coefficient is significantly improved compared with the existing method.
[0070] 2. The present invention proposes a label enhancement method based on the adaptive threshold self-training method of spectral-normalized Gaussian process, which solves the problem of sparse labels for track structure diseases. By introducing the spectral-normalized Gaussian process and designing a dynamic threshold strategy, this method realizes high-quality annotation of unlabeled samples, and all evaluation indicators are significantly better than the existing methods.
[0071] 3. The present invention proposes an attention-guided multi-task cascaded CNN model, which realizes the multi-level parallel recognition of track geometry and structure diseases for the first time. Through the cascaded architecture and cross-level attention guidance mechanism, this model realizes the accurate evaluation of the severity while ensuring the accuracy of disease detection, and has achieved good performance in the recognition of geometric and structural diseases, providing more comprehensive technical support for track condition assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 is a flowchart of the multi-level track disease recognition system based on vehicle body vibration data of the present invention;
[0073] Figure 2 is a method for generating pseudo-labels of structural diseases based on the adaptive threshold self-training method of spectral-normalized Gaussian process in this embodiment;
[0074] Figure 3 is a multi-level track disease classification diagram of the attention-guided multi-task cascaded CNN in this embodiment;
[0075] Figure 4 is a feature-vibration correlation heat map in this embodiment;
[0076] Figure 5 It is the silhouette coefficient analysis chart for determining the optimal number of clusters in this embodiment;
[0077] Figure 6 It is the t-SNE visualization chart of the clustering results of grid samples and sampling points in this embodiment;
[0078] Figure 7 It is the performance comparison chart of K-means clustering, spectral clustering and density-based clustering methods in this embodiment;
[0079] Figure 8 It is the evolution chart of pseudo-label generation in cluster 0 and cluster 1 in this embodiment;
[0080] Figure 9 The training accuracy rate and loss function evolution curve of the same type of track diseases in this embodiment;
[0081] Figure 10 The evolution chart of attention weights in 100 rounds of training in this embodiment. Detailed implementation manners
[0082] The present invention will be further described in detail below through specific implementation manners:
[0083] Embodiment 1
[0084] As Figures 1 to 4 shown, a multi-level track disease recognition system based on vehicle body vibration data includes:
[0085] A data acquisition module, which is used to acquire the heterogeneous data of the track environment; in this embodiment, the data acquisition module uses a portable detector, and the size of the detector host is 270×170×91 mm. As Figure 1 (a-b) shown, the vehicle body vibration data is acquired by a portable detector on an ordinary passenger subway train. The vehicle body vibration data includes the longitudinal, lateral and vertical accelerations and angular velocities of the vehicle body, as well as the speed and mileage position information. The heterogeneous data of the track environment includes static and dynamic data. The static heterogeneous data is extracted from the line basic design drawing, including the ballast type, line curvature, and gradient; the dynamic heterogeneous data is the train running speed.
[0086] The basic rules for dataset construction are as follows:
[0087] T = (a, b, l)|a ∈ A, b ∈ B, l ∈ L, m(a) = m(b) = m(l) (1);
[0088] Among them, A is the vehicle body vibration dataset, and each sampling point contains 6 vibration parameters A i = [a 1,..., a 6 , specifically: a 1 represents the longitudinal acceleration, a 2 represents the lateral acceleration, a 3 represents the vertical acceleration, a 4 represents the longitudinal angular velocity, a 5 represents the lateral angular velocity, a 6 represents the vertical angular velocity. B is the dataset of heterogeneity factors, and each sampling point contains n environmental parameters B i = [b 1 ,..., b n . To eliminate the dimensional differences between different parameters, [-1, 1] interval normalization is adopted; for the non-numerical track bed type parameters, numerical mapping is performed according to the following rules: the integral track bed is assigned -1, the rubber isolation pad track bed is assigned -0.33, the double-layer non-linear fastener track bed is assigned 0.33, and the steel spring floating slab track bed is assigned 1. L is the disease label set, and the label information of each sampling point is represented by L i = [l 1 ,... l δ , where, l 1-δ correspond to δ different types of track disease label values respectively;
[0089] To obtain statistically representative training samples, the method for constructing grid samples is as follows:
[0090] G = g 1 ,..., g i ,..., g n (2);
[0091]
[0092] where, [·] represents the floor operation, g i represents the i-th grid sample, n is the total number of grids, which is determined by the total number of sampling points N and the number of sampling points m in a single grid: n = [N / m], and represent the vibration characteristics and heterogeneity factor information of the j-th sampling point in the i-th grid respectively. Through the above technical solutions, the present invention realizes the effective fusion of vibration data, heterogeneity factors and disease labels, laying a data foundation for subsequent disease identification.
[0093] The environment adaptive module is used to analyze the correlation between the vehicle body vibration data and the heterogeneity data, construct a dual-perspective similarity matrix with differentiated weights, and realize the adaptive classification of the track environment through spectral clustering method.
[0094] The process of the spectral clustering algorithm is as follows: First, analyze the correlation strength between vehicle body vibration and heterogeneity factors to determine the feature weights, and based on this, design an improved dual-perspective similarity matrix to make full use of the differential information of the weights, and then perform spectral clustering to divide the line grid sample data set G into different categories. Among them, the calculation process of the feature weights is as follows: Considering that there may be a non-linear relationship between heterogeneity factors and vehicle body vibration, the Spearman rank correlation coefficient and the Pearson correlation coefficient are used for calculation respectively.
[0095] Spearman rank correlation coefficient calculation formula:
[0096]
[0097] where, i represents the i-th heterogeneity factor, s represents the s-th vehicle body vibration; t′ i =Rank(t i ) and l′ i =Rank(l i ) respectively represent sorting t i and l i ; m represents that the data has m samples.
[0098] Pearson correlation coefficient calculation formula:
[0099]
[0100] where, i.e., the column mean of t i ; i.e., the column mean of l s ; m represents that the data has m samples. Take the larger value in formulas (4) and (5) as the correlation between the i-th heterogeneity factor and the s-th vehicle body vibration.
[0101]
[0102] The initial weight values are calculated as shown in the following formulas (7) and (8), and these two weights will be set as the basic weight values for calculating the weight values of each heterogeneity factor according to the designed feature weight algorithm. Since there are 6 vehicle body vibration characteristics, f takes the value of 6.
[0103] respectively.
[0104]
[0105] The inputs of the algorithm include heterogeneity data, vehicle body vibration data, the correlation matrix corr i-s between features and vehicle body vibration, and the predefined initial weight values w 1 and w 2 . The output is a list of heterogeneity factor weight values w s sorted in descending order of weights.。The core idea of the algorithm is to divide corr i-s into two groups according to the correlation strength, and assign w 1 and w 2 respectively. The specific steps are as follows: First, sort the correlation values corr i-s in descending order and group them. The first n are classified into the correlation group L 1 , and the rest are classified into the correlation group L 2 . Then, accumulate the weights of each feature. In the weight calculation stage, traverse all corr i-s . If it is in L 1 , then accumulate w 1 , and if it is in L 2 , then accumulate w 2 . Finally, sort the accumulated weights of each heterogeneity factor to obtain the final feature weight list w s .
[0106] The label enhancement module calculates the multi-dimensional uncertainty index of the samples through the self-training method of the spectral-normalized Gaussian process, and gradually generates high-quality pseudo-labels through the dynamic threshold mechanism, and finally forms a completely labeled sample set.
[0107] Among them, the method for constructing the dual-view similarity matrix is as follows: The feature space is divided into a high-weight view V h and a low-weight view V l based on the median of the feature weights. For these two views, different connection parameters α and structure protection parameters β are designed: The high-weight view adopts larger connection parameters αV h and structure protection parameters βV h to allow more connections to be established between samples and protect the local structural characteristics; the low-weight view adopts smaller connection parameters αV l and structure protection parameters βV l to limit the connections between samples and reduce the dependence on the local structure. This differential parameter design enables the two views to characterize the internal structure of the data from different levels. Inside each view, the local structural association is characterized by the local density and the reconstruction relationship:
[0108] The local density is measured by the overlap degree of the k-nearest neighbor sets of the samples:
[0109]
[0110] Among them, N(i) represents the number of k-nearest neighbors of sample i, d for each view is adjusted by α: d = d 0 ×(1 + αw s ), where d 0 takes 10;
[0111] The reconstruction relationship is calculated by normalizing the distance between samples:
[0112]
[0113] Based on the perspective of local index design, a Gaussian kernel function is as follows:
[0114]
[0115] where w s is the feature weight vector, σ is the local adaptive bandwidth, taking the average distance of the k-nearest neighbors of the sample, and L(x, y) is the local structure protection term, which consists of local density and reconstruction relationship.
[0116] To balance the contributions of the two perspectives to the final similarity, the similarity matrices are fused as follows:
[0117] S h +(1 - γ)S l (12);
[0118] where S h and S l represent the similarity matrices of the two perspectives respectively, and γ is the fusion coefficient, which is used to control the dominant degree of the high-weight perspective, so as to retain the auxiliary information provided by the low-weight features while maintaining the dominant role of the high-weight features. Based on the fused similarity matrix S, the degree matrix D is constructed:
[0119]
[0120] where the non-diagonal elements are all 0, and the Laplacian matrix L p is calculated from the similarity matrix and the degree matrix, using the symmetric normalization method:
[0121]
[0122] where I is the identity matrix. The symmetric normalization method can maintain the geometric structure of high-dimensional data while having the advantages of high computational stability and effective processing of data at different scales. Calculate the eigenvectors {v p corresponding to the first k (expected number of clusters) smallest eigenvalues of L 1 , v 2 , …, v k}, and construct the eigenvector matrix V with them as column vectors. Normalize each row of V to unit length to obtain the matrix Y:
[0123]
[0124] Each row in Y is clustered as a k-dimensional point, and the class of sample i is the cluster it finally belongs to. The above spectral clustering process classifies all line grid units, and the sampling points within the grid units are also divided into different classes, laying a foundation for subsequent disease recognition tasks for different classes.
[0125] The pseudo-label generation method includes a spectral normalization Gaussian process model and an adaptive threshold training mechanism:
[0126] As Figure 2 shown, the spectral normalization Gaussian process model (SNGP model) contains two residual blocks, and each residual block is configured with a spectral normalization layer. SN controls the Lipschitz continuity condition constant of the network by restricting the spectral norm of the weight matrix, and its calculation formula is as follows:
[0127]
[0128] where W is the weight matrix and σ(W) is its largest singular value.
[0129] The network parameters are shown in Table 1 below, and the output vector is connected to the random Fourier feature module. This module approximates the Gaussian process kernel function through random feature mapping to judge the similarity between unlabeled samples and labeled samples. Specifically, first use the Gaussian kernel function to measure the similarity between two points in the input space, which decays exponentially as the distance between the points increases:
[0130]
[0131] where is the length scale parameter, and the initial value is set to 0.1. RFF approximates this kernel function through the following mapping function to approximate this kernel function:
[0132]
[0133] where W is a random projection matrix sampled from the normal distribution and b is a random bias vector sampled from the uniform distribution U[0, 2π], D f is the feature dimension (512); through this mapping, the inner product of any two inputs x and x′ satisfies:
[0134]
[0135] The predicted value is obtained through the linear combination of the mapped features, and the uncertainty can be calculated from the kernel matrix in the feature space. For the evaluation of the SNGP output value in each iteration, this embodiment designs a multi-dimensional uncertainty evaluation strategy, such as Figure 2As shown, a multi-dimensional uncertainty evaluation value μ is designed based on the predicted value and uncertainty value output by SNGP, comprehensively considering four metrics: E is entropy, representing distribution uncertainty; M is the marginal difference; C is the confidence level, taking the highest prediction probability; V is the prediction variance, representing the model stability, as shown in the following formula:
[0136] E = -∑(p i × log(p i )) (20);
[0137] M = p m - p s (21);
[0138] C = max(p i ) (22);
[0139] V = diag(K) - ∑p i 2 (23);
[0140] μ = w 1 × E + w 2 × (1 - M) + w 3 × (1 - C) + w 4 × V (24);
[0141] Among them, diag(K) represents the diagonal elements of the kernel matrix, p i represents the predicted probability that the sample belongs to the i-th class, p m represents the highest predicted probability of the sample, p s represents the second-highest predicted probability of the sample; w 1 , w 2 , w 3 and w 4 are the weights of entropy, marginal difference, confidence level, and prediction variance respectively.
[0142] Furthermore, a threshold adaptive mechanism is designed. By determining whether the multi-dimensional uncertainty evaluation value μ exceeds the dynamic threshold, the training set to be added to the next iteration is selected. The specific method is as follows: Set the dynamic low threshold μ l , μ lower than μ l indicates that the sample is predicted to be a certain known class sample with a high degree of certainty, that is, the certainty of the actual existence of the disease is relatively large. Then, set the sample predicted value as the pseudo-label. Set the dynamic high threshold μ h , μ higher than μ h indicates that the sample may have characteristics different from the known label samples and is very likely to belong to a new class, that is, the actual situation of no disease. Then, set the sample predicted value as the new class '0'. Add the above samples to the training set for the next training. First, the initial threshold is set based on statistical principles as follows:
[0143]
[0144] Among them, μ and σ are the mean and standard deviation of the model uncertainty estimation respectively; the 3σ interval setting provides a reliable limit in a statistical sense and can effectively identify high-confidence samples and potential new-class samples. Considering that the model prediction ability gradually improves with training iterations, the threshold needs to be adjusted accordingly to adapt to the change of model performance:
[0145]
[0146] Among them, is the threshold of the previous iteration, α·ΔC·σ(μ)<0.5σ(μ); is the average confidence of the newly labeled samples, 0.95 is the baseline confidence; after reaching the predetermined number of iterations, the remaining sample categories are predicted by the mature model, and the samples of new categories among the remaining samples are screened according to the last threshold.
[0147] Table 1. Model architecture parameter configuration of the adaptive threshold self-training method based on spectral normalization Gaussian process
[0148]
[0149]
[0150] The disease identification module adopts an attention-guided multi-task cascaded CNN architecture, that is, the disease location is detected through a binary classification layer, the detection results are filtered by means of a mask layer, and the disease level is evaluated by a multi-classification layer. Finally, a multi-level identification result including the disease location and severity is output.
[0151] Among them, the binary classification layer is used to identify the disease location. The specific structure is as follows: a channel attention mechanism is introduced after the original input signal, and the importance weights of different vehicle body pose features are captured through parallel global average pooling and max pooling branches. After the features of the two branches are subjected to dimensionality reduction and non-linear transformation of the ReLU activation function, they are fused to generate attention weights, which are multiplied by the original input to obtain weighted features. The weighted features are subjected to feature extraction through a two-layer convolutional network, and each layer of convolution is followed by a max pooling operation to compress the feature dimension. After the extracted features are flattened, they are input into a three-layer fully connected network for feature transformation. The ReLU activation function is used in all cases, and dropout is adopted to prevent overfitting. Finally, a binary classification result is generated through an output node with a sigmoid activation function. Since misdiagnosing a diseased sample as disease-free (false negative) is more dangerous than misdiagnosing a disease-free sample as diseased (false positive). The binary classification layer adopts a weighted binary cross-entropy loss function, and the penalty factor w is introduced to increase the penalty for false negative samples:
[0152]
[0153] Among them, y is the true label of each sample, that is, y = 1 indicates the existence of a disease, and y = 0 indicates the absence of a disease. y p is the predicted value, which is between 0 and 1; w is the penalty factor value. When y = 1, the binary cross-entropy loss function will be activated. The closer the predicted value is to 0, the larger the loss function value.
[0154] The mask layer is used to screen the results of the binary classification layer. The mask layer includes a screening layer and a multi-classification input layer. The screening layer applies a non-zero mask to the output of the binary classification layer, retains the positive predictions in the output, that is, the predicted value > 0.5, and sets all other predicted values to zero; multiplies the filtered output of the binary classification layer with the original input element-wise to screen out the samples that the binary classification layer considers to have diseases; the screened samples form the multi-classification input layer for training to evaluate the disease level in the multi-classification layer.
[0155] The multi-classification layer is used to evaluate the disease level. The specific method is as follows: Feature weighting is performed on the output features of the convolutional block using the channel attention mechanism. The channel attention weights in the first stage are non-linearly transformed through feature remapping, and then sequentially pass through the dimension adjustment of the fully connected network and the ReLU activation function; a normalized guiding weight is generated through the sigmoid function, and the guiding weight is multiplied element-wise with the output of the channel attention in the multi-classification stage; the weighted features pass through a multi-layer fully connected network for feature transformation, and dropout is used in each layer to prevent overfitting; finally, according to the number of levels of different diseases, the multi-classification prediction results are output through the softmax function; the standard cross-entropy loss function is used in this stage, and it is weighted and added to the loss function of the binary classification layer to jointly form the overall optimization objective of the model:
[0156] l t = α l × l b + β l × l m (30);
[0157] Among them, l b is the binary classification loss, l m is the multi-classification loss, α l and β l are the corresponding weight coefficients.
[0158] Case analysis:
[0159] Select the subway Line 10 of a certain city as the specific test and verification object. Based on the above-mentioned disease identification system, conduct on-site data collection and disease identification experiments on this line. A total of 82,400 groups of complete test samples are obtained through a portable detector. Each group of samples includes vehicle body vibration data, heterogeneity data, and corresponding disease marking information. The data collection covers a range of 20.6095 kilometers of the entire line, including various typical working conditions such as different gradients and different track bed structures. The collected disease samples include two major categories: track geometry diseases and structural diseases. Among them, the geometry diseases cover five main types: undulation, alignment, gauge, cross level, and twist; the structural diseases select two typical disease forms: side wear and vertical wear of the rail.
[0160] Verification of the effectiveness of the dual-view spectral clustering algorithm:
[0161] In the data preprocessing stage, first, perform standardization processing on the dataset obtained by associating the vehicle body vibration data and the heterogeneity data. To eliminate the influence caused by different dimensions of various heterogeneity data and retain relatively complete information, normalize all data to the interval [-1, 1]. For the non-numerical track bed type parameters, special processing is required: the integral track bed, rubber vibration isolation pad track bed, double-layer non-linear fastener track bed, and steel spring floating slab track bed are respectively assigned values of 1, 2, 3, and 4, and after normalization, their values are -1, -0.33, 0.33, and 1 respectively. This mapping method maintains the relative relationship between different track bed types.
[0162] In terms of grid division and feature extraction, in this embodiment, the line is divided into grid units at a length of 200 meters, and a total of 103 grid samples are obtained. Each grid contains 800 sampling points (200 / 0.25 = 800). Calculate the mean values of the vehicle body vibration and heterogeneity data in each grid according to formulas (1)-(3) as the feature representation of the grid. Furthermore, calculate the correlation between each vibration feature and the heterogeneity data using formulas (4)-(8), as shown in Figure 4 the heat map shown, where the vertical axis represents four heterogeneity factors, the horizontal axis represents six vehicle body vibration parameters, blue represents negative correlation, red represents positive correlation, and the color depth reflects the intensity of the correlation. Based on the calculation results, obtain the weight values of each factor: speed 0.45, gradient 0.25, track curvature 0.25, track bed 0.05.
[0163] In terms of clustering parameter settings, set a larger connection parameter αV h = 0.9 and a structural protection parameter βV h = 0.8 for the high-weight perspective to allow more connections to be established between samples and protect the local structural characteristics; set smaller parameters αV l = 0.2 and βV l= 0.2 to limit the connection between samples and reduce the dependence on local structures; the fusion coefficient γ is set to 0.7 to balance the contributions of high- and low-weight perspectives. By comparing the silhouette coefficients of 2 - 10 clusters (as shown in Figure 5 ), the highest silhouette coefficient of 0.46 is obtained when the number of clusters is 2, and thus the optimal number of clusters is determined to be 2.
[0164] To visually display the clustering effect, the t-SNE technique is used to map the high-dimensional data to a two-dimensional space. As shown in Figure 6 , the blue cluster 0 contains 74 grids (59,200 sampling points), and the green cluster 1 contains 29 grids (23,200 sampling points). The two types of data show obvious separation in the feature space, indicating that the clustering model has successfully captured the data structure differences in the high-dimensional heterogeneous factor space.
[0165] To comprehensively evaluate the superiority of this scheme, a comparative experiment with traditional clustering algorithms was carried out. As shown in Figure 7 and Table 2 below, the weighted spectral clustering method proposed in this embodiment achieved a silhouette coefficient of 0.4567, which is a 35.7% improvement compared to the unweighted method (0.3365). The weighted spectral clustering, unweighted spectral clustering, and density-based spatial clustering all determined the optimal number of clusters to be 2, indicating that the spectral clustering method has good stability. Although K-means gave a more detailed four-class division, its lower silhouette coefficient (0.1774) implies the risk of overfitting. In terms of the balance of sample distribution, this method reasonably divides 103 samples into two groups of 29 and 74, while the distribution of density-based spatial clustering (86:17) shows obvious imbalance. Further verification through the visualization of clustering results shows that the clustering boundary obtained by this method is the clearest and the separation degree between clusters is the highest, while the results of other methods, especially K-means clustering, show relatively fuzzy cluster boundaries. The experimental results fully demonstrate the effectiveness and superiority of the spectral clustering algorithm in dealing with the heterogeneity problems in complex track environments.
[0166] Table 2 Comparison of clustering method performance and sample distribution
[0167]
[0168] Verification of the effectiveness of the adaptive threshold self-training method for structural disease samples based on spectral normalized Gaussian processes:
[0169] The random Fourier feature module sets the feature dimension to 512, and the initial value of the length scale parameter is 0.1. In the multi-dimensional uncertainty evaluation mechanism, both the entropy weight and the marginal difference weight are set to 0.3, and the confidence weight and the prediction variance weight are set to 0.2. As shown in Table 4 below, the label distributions of the two clustering samples are as follows (as shown in Table 3): Labels 0, 1, 2, and 3 represent that the samples have no diseases, disease level 1, 2, and 3 respectively. There are a large number of unlabeled samples in each sample set, and pseudo-labels need to be labeled through self-training. Since the number of samples with label 0 is 0, label 0 will be recognized as a new category during model training.
[0170] Table 3 Initial distribution of track structure disease labels in clustering samples
[0171]
[0172] Experimental data shows that: as shown in Table 4, the mean variance of the model in all scenarios is below 0.19, and the standard deviation is within 0.04. The average confidence value is 0.95. In terms of the length scale interval, the vertical wear interval of clustering cluster 0 is [0.20024903, 0.21437594], and the lateral wear interval of clustering cluster 1 is [0.1, 0.1457911].
[0173] Table 4 Performance indicators of the adaptive threshold self-training method based on spectral normalization Gaussian process
[0174]
[0175] The following data is obtained through the tracking analysis of the training process: as Figure 8 shown, in the lateral wear training of clustering cluster 0, the average confidence of the model increases from 0.8 to 0.95. In the first round of iteration, 13,486 samples are labeled, and 1,759 (13.04%) of them are identified as new categories. The new category identification ratio reaches 99.32% in the fourth round of iteration. For vertical wear training, the confidence increases from 0.5 to 0.9, and the identification ratio of new category samples in each round of iteration is in the range of 11 - 15%. In clustering cluster 1, 5,490 samples are labeled in the first round of lateral wear iteration, and the new category identification ratio is 16.67%. The new category identification ratio increases to 34.48% in the third round of iteration.
[0176] The final label distribution results are shown in Table 5 below. Among them, for cluster 0, the lateral wear and vertical wear contain 39,948 and 1,165 label 0 samples respectively, and for cluster 1, the lateral wear and vertical wear contain 5,198 and 472 label 0 samples respectively. To verify the effectiveness of the adaptive threshold self-training method based on spectral normalization Gaussian process in pseudo-labels, as shown in Table 6, a baseline model (convolutional neural network + self-training model), convolutional neural network + Monte Carlo sampling + self-training model, convolutional neural network + deterministic uncertainty estimation + self-training model, and convolutional neural network + spectral normalization Gaussian process + self-training model were selected for comparison. The adaptive threshold self-training method based on spectral normalization Gaussian process outperforms other methods in terms of four performance indicators: average confidence, computational efficiency, convergence performance, and confidence variance.
[0177] Table 5 Final Distribution of Track Structure Disease Labels Processed by the Adaptive Threshold Self-Training Method Based on Spectral Normalization Gaussian Process
[0178]
[0179]
[0180] Table 6 Performance Comparison of Different Models in Pseudo-Label Tasks
[0181]
[0182] Experimental data show that: The adaptive threshold self-training method based on spectral normalization Gaussian process can handle the problem of sparse track structure disease labels. Through the dynamic threshold adjustment mechanism and multi-dimensional uncertainty evaluation strategy, this method can perform data annotation for scenarios with different complexities. The dataset processed by the adaptive threshold self-training method based on spectral normalization Gaussian process can provide training samples for subsequent multi-level track disease recognition tasks.
[0183] Verification of the Effectiveness of Multi-Task Cascade CNN Guided by Multi-Level Disease Attention:
[0184] The multi-level track disease data that has been labeled is input into the attention-guided multi-task cascade CNN for training. As Figure 4 shown, the model contains seven branches, corresponding to the recognition of high and low, horizontal, gauge, alignment, twist, lateral wear, and vertical wear respectively. The penalty factor value of the binary classification loss function is set to 2. The weights of the binary classification loss and multi-classification loss in the total loss are set to 2 and 1 respectively to ensure that the model accurately detects the disease location first.
[0185] Parameter search is performed on two types of samples to determine the optimal hyperparameter combination. In cluster 0, the optimal hyperparameter combination is a learning rate of 0.001 and a regularization term of 0.01; in cluster 1, it is a learning rate of 0.01 and a regularization term of 0.1. The selection of hyperparameters for cluster 0 (59,200 samples) and cluster 1 (23,200 samples) is reasonable. The number of samples in cluster 0 is significantly larger, and the smaller learning rate and smaller regularization term enable the model to perform more detailed learning. In contrast, the number of samples in cluster 1 is smaller. The larger learning rate can accelerate the convergence speed of the model on less data, while the larger regularization term enables the model to better learn the true features of the data.
[0186] The number of training epochs is set to 100. As Figure 9 shown, the output loss function and accuracy varying with the training epochs demonstrate the recognition effects of different types of diseases. Figure 9 (a - d) show the training results of cluster 0, including unevenness ( Figure 9 -a), twist ( Figure 9 -b) in geometric diseases, and vertical wear ( Figure 9 -c) and side wear ( Figure 9 -d) in structural diseases. Figure 9 (e - h) show the training results of cluster 1, including gauge ( Figure 9 -e), alignment ( Figure 9 -f), unevenness ( Figure 9 -g) and vertical wear ( Figure 9 -h) in structural diseases. Each sub - figure is divided into upper and lower parts. The upper part shows the accuracy evolution curve, and the lower part shows the loss function evolution curve.
[0187] From the training effect, as shown in Table 7 below, the model has good performance in the classification of geometric and structural track diseases. It performs outstandingly in the binary classification task: the precision, recall rate, and F1 - score of all disease types are close to or reach 1.00. From the Kappa coefficient, the Kappa values of geometric diseases such as unevenness, level, gauge, and alignment are all above 0.9, while the Kappa values of structural diseases such as side wear and vertical wear are relatively low. In particular, the Kappa value of vertical wear in cluster 0 is 0.8010.
[0188] Table 7 Performance metrics of the attention - guided multi - task cascaded CNN model in the recognition of various track diseases
[0189]
[0190] The recognition performance in the multi-classification task remains at a relatively high level. The F1 scores of geometric disease types are generally above 0.95, while the F1 score of structural diseases such as side wear in cluster 1 is 0.93. Generally speaking, the model shows stable performance in both datasets, especially in the recognition of geometric diseases. Although the recognition of structural diseases is relatively challenging, its performance indicators still remain at an acceptably high level.
[0191] To further understand the learning mechanism of the model, this embodiment analyzes the evolution characteristics of attention weights in the recognition of unevenness and side wear. As Figure 10 shown, the model shows an accurate grasp of physical laws. Figure 10 -a shows that in the recognition of unevenness in cluster 1, the pitch angular velocity weight rises from 0.52 to 0.67 and remains stable. Figure 10 -b shows that in the recognition of side wear in cluster 1, the yaw angular velocity weight rapidly rises from 0.5 to 0.7, and the roll angular velocity weight steadily rises to 0.55. These changes in weight distribution reflect the model's effective recognition ability for different types of disease characteristics.
[0192] The above are only embodiments of the present invention, and common knowledge such as specific structures and / or characteristics well known in the art are not described in detail herein. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several modifications and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be subject to the content of its claims, and the specific implementation manners described in the specification can be used to interpret the content of the claims.
Claims
1. A multi-level track defect identification system based on vehicle body vibration data, characterized in that: include: The data acquisition module collects vehicle vibration data and track environment heterogeneity data; constructs a grid sample dataset including vehicle vibration dataset, heterogeneity factor dataset and disease label set; The environment adaptation module is used to analyze the correlation between the vehicle body vibration data and the heterogeneous data, construct a dual-view similarity matrix with differentiated weights, and realize the adaptive classification of the track environment through the spectral clustering method, that is, to analyze the correlation strength between the vehicle body vibration data and the heterogeneous data to determine the feature weight, divide the feature space into a high-weight perspective and a low-weight perspective based on the median of the feature weight, and fuse the similarity matrices representing the two perspectives respectively; then perform spectral clustering to divide the line grid sample data set G into different categories; The label enhancement module calculates the multi-dimensional uncertainty index of the sample through the self-training method of the spectral normalized Gaussian process model, and gradually generates high-quality pseudo-labels through a dynamic threshold mechanism, ultimately forming a complete annotated sample set; The defect recognition module adopts an attention-guided multi-task cascade CNN architecture, that is, the defect location is detected through a binary classification layer, the detection results of the binary classification layer are filtered with the help of a mask layer, the defect level is evaluated through a multi-classification layer, and finally a multi-level recognition result including the defect location and severity is output. The specific method is as follows: first, the fully annotated sample set is input into the binary classification layer, and the channel attention mechanism is introduced. The importance weights of different vehicle posture features are captured through parallel global average pooling and maximum pooling branches; the features of the two branches are fused to generate attention weights after dimensionality reduction and nonlinear transformation of the activation function, and multiplied with the original input to obtain weighted features; The weighted features are extracted through a two-layer convolutional network, and the disease location detection results are finally output. Secondly, the mask layer is used to filter the results of the binary classification layer. The mask layer includes a filtering layer and a multi-classification layer. The filtering layer applies a non-zero mask to the output of the binary classification layer, retains the positive predictions in the output, and sets all other prediction values to zero. The filtered binary classification layer output is element-wise multiplied with the original input to filter out samples that the binary classification layer believes to have diseases. The filtered samples are input into the multi-classification layer for multi-layer disease grade evaluation. Finally, a multi-level identification result including the location and severity of the disease is output.
2. The multi-level track defect identification system based on vehicle body vibration data according to claim 1 is characterized in that: The basic rules for dataset construction are as follows: T=(a,b,l)|a∈A, b∈B, l∈L, m(a)=m(b)=m(l); Among them, A is the vehicle body vibration data set, and each sampling point contains 6 vibration parameters A i =[a1, ..., a6], specifically: a1 represents longitudinal acceleration, a2 represents lateral acceleration, a3 represents vertical acceleration, a4 represents longitudinal angular velocity, a5 represents lateral angular velocity, and a6 represents vertical angular velocity; B is a heterogeneous factor data set, and each sampling point contains n environmental parameters B i =[b1, ..., b n ]; In order to eliminate the dimension differences between different parameters, the [-1, 1] interval normalization process is adopted; For non-numerical ballast type parameters, numerical mapping is performed according to the following rules: the integral ballast is assigned a value of -1, the rubber vibration isolation pad ballast is assigned a value of -0.33, the double-layer nonlinear fastener ballast is assigned a value of 0.33, and the steel spring floating plate ballast is assigned a value of 1; L is the disease label set, and the label information of each sampling point is represented by L i =[l1,...l δ ] indicates that, among which, l 1-δ They correspond to δ different types of track disease label values; In order to obtain statistically representative training samples, the method of constructing grid samples is as follows: G=g 1 ,...,g i ,...,g n ; Among them, [·] represents the rounding down operation, g i represents the i-th grid sample, n is the total number of grids, which is determined by the total number of sampling points N and the number of sampling points m of a single grid: n = [N / m], and They represent the vibration characteristics and heterogeneity factor information of the jth sampling point in the ith grid respectively.
3. The multi-level track defect identification system based on vehicle body vibration data according to claim 1 is characterized in that: The method of constructing the dual-view similarity matrix is as follows: the feature space is divided into high-weight view V based on the median of the feature weight h and low weight perspective V l , the high-weight view uses the connection parameter αV h and structural protection parameter βV h , to allow more samples to establish connections and protect local structural characteristics; the low-weight perspective uses the connection parameter αV l and structural protection parameter βV l , limiting the connection between samples and reducing the dependence on local structure, characterizing the local structure association through local density and reconstruction relationship within each perspective: The local density is measured by the overlap of the sample's k-nearest neighbor sets: Where N(i) represents the number d of k-nearest neighbors of sample i, and d for each view is adjusted by α: d = d0×(1+αw s ), where d0 is 10; The reconstruction relationship is calculated by normalizing the distance between samples: Gaussian kernel function in the design perspective based on local indicators: Among them, w s is the feature weight vector, σ is the local adaptive bandwidth, the average distance of the k nearest neighbors of the sample is taken, and L(x, y) is the local structure protection term, which consists of the local density and the reconstruction relationship.
4. The multi-level track defect identification system based on vehicle body vibration data according to claim 3 is characterized in that: The similarity matrices are fused to balance the contribution of the two perspectives to the final similarity: S h +(1-γ)S l ; Among them, S h and S l They represent the similarity matrices of the two perspectives respectively, and γ is the fusion coefficient, which is used to control the dominance of the high-weight perspective. Based on the fused similarity matrix S, the degree matrix D is: Among them, all non-diagonal elements are 0, and the Laplace matrix L is calculated by the similarity matrix and the degree matrix p , using the symmetric normalization method: Where I is the unit matrix. The symmetric normalization method can maintain the geometric structure of high-dimensional data while having the advantages of high computational stability and efficient processing of data of different scales. p The eigenvectors corresponding to the first k smallest eigenvalues of {v1, v2, ..., v k }, and use it as a column vector to construct the eigenvector matrix V, and normalize each row of V to unit length to obtain the matrix Y: Cluster each row in Y as a k-dimensional point, and the category of sample i is the cluster to which it finally belongs.
5. The multi-level track defect identification system based on vehicle body vibration data according to claim 1 is characterized in that: The pseudo-label generation method includes a spectral normalized Gaussian process model and an adaptive threshold training mechanism: Spectrally normalized Gaussian process model: The Gaussian kernel function is used to measure the similarity between two points in the input space, which decays exponentially as the distance between the points increases: Where l is the length scale parameter, and its initial value is set to 0.1; RFF is mapped by the following function To approximate the kernel function: Where W is the normal distribution The random projection matrix of the sample, b is the random bias vector U[0, 2π] sampled from a uniform distribution, D f is the feature dimension (512); through this mapping, the inner product of any two inputs x and x′ satisfies: The predicted value is obtained by linear combination of the mapped features, and the uncertainty can be calculated by the kernel matrix in the feature space; Multi-dimensional uncertainty assessment and adaptive threshold training mechanism: Based on the predicted value and uncertainty value output by the spectral normalized Gaussian process model, a multidimensional uncertainty assessment value μ is designed, taking into account four metrics: E is entropy, which represents distribution uncertainty; M is marginal difference; C is confidence, which takes the highest prediction probability; V is prediction variance, which represents model stability; E=-∑(p i ×log(p i )); M=p m -p s ; C=max(ρ i ); V=diag(K)-∑p i 2 ; μ=w1×E+w2×(1-M)+w3×(1-C)+w4×V; Among them, diag(K) represents the diagonal elements of the kernel matrix, p represents the predicted probability that the sample belongs to the i-th category, and p m Represents the highest predicted probability of the sample, p s represents the second highest predicted probability of the sample; w1, w2, w3 and w4 are the weights of entropy, marginal difference, confidence and prediction variance respectively; By determining whether the multidimensional uncertainty evaluation value μ exceeds the dynamic threshold, the training set to be added to the next iteration is selected. The specific method is: setting the dynamic low threshold μ l , μ is lower than μ l Indicates that the sample is predicted to be a known type of sample with high certainty, that is, the certainty of the real disease is high, then the sample prediction value is set as a pseudo label; set the dynamic high threshold μ h , μ is higher than μ h It means that the sample has characteristics different from the known labeled samples and belongs to a new category. In other words, if there is no disease, the sample prediction value is set to the new category '0'.
6. The multi-level track defect identification system based on vehicle body vibration data according to claim 1 is characterized in that: The initial threshold is set as follows: Among them, μ and σ are the mean and standard deviation of the model uncertainty estimate, respectively; the 3σ interval setting provides a statistically reliable limit. Considering that the model prediction ability gradually improves with training iterations, the threshold needs to be adjusted accordingly to adapt to changes in model performance: in, is the threshold of the previous iteration, α·ΔC·σ(μ)<0.5σ(μ); is the average confidence of the newly labeled samples, and 0.95 is the benchmark confidence; after reaching the predetermined number of iterations, the mature model is used to predict the categories of the remaining samples, and the samples of the new category in the remaining samples are screened according to the last threshold.
7. The multi-level track defect identification system based on vehicle body vibration data according to claim 1 is characterized in that: The binary classification layer uses a weighted binary cross entropy loss function, which increases the penalty for false negative samples by introducing a penalty factor w: Among them, y is the true label of each sample, that is, y = 1 means the disease exists, y = 0 means the disease does not exist, y p is the predicted value, between 0 and 1; w is the penalty factor value. When y=1, the binary cross entropy loss function will be activated. The closer the predicted value is to 0, the larger the loss function value.
8. The multi-level track defect identification system based on vehicle body vibration data according to claim 1 is characterized in that: The specific method of evaluating the disease level in the multi-classification layer is as follows: the channel attention mechanism is used to weight the output features of the convolution block, and the first-stage channel attention weights are nonlinearly transformed through feature remapping, and then pass through the dimension adjustment and ReLU activation function of the fully connected network in turn; the normalized guide weights are generated through the sigmoid function, and the guide weights are multiplied element-by-element with the channel attention output of the multi-classification stage; the weighted features are transformed through a multi-layer fully connected network, and dropout is used in each layer to prevent overfitting; finally, according to the number of different disease levels, the multi-classification prediction results are output through the softmax function; this stage uses the standard cross entropy loss function, which is weighted and added with the loss function of the binary classification layer to form the overall optimization goal of the model: l t =α l ×l b +β l ×l m ; Among them, l b is the binary classification loss, l m is the multi-classification loss, α l and β l is the corresponding weight coefficient.
Citation Information
Patent Citations
Weight self-updating multi-view spectral clustering method based on shared neighbors
CN111401468A
Pseudo-label remote sensing image scene classification method based on adaptive threshold
CN114549909A