Steel rail corrugation fault diagnosis method and training method of diagnosis model
By collecting acoustic signals on the train and performing multi-feature dimension extraction and optimization of XGBoost algorithm training, the low accuracy and environmental dependence of rail wave grinding fault detection in the existing technology is solved, and more efficient and accurate fault diagnosis is achieved.
Patent Information
- Application Number
- CN202510469954.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, the detection method of rail wave grinding failure has problems such as low accuracy, high cost and susceptible to environmental factors. In particular, the judgment basis based on the A-weighted sound pressure level is single, resulting in low recognition accuracy.
By collecting the acoustic signals when the train is traveling, using high-sensitivity acoustic sensor to obtain the acoustic signals to be processed, performing feature extraction in multiple feature dimensions, including time domain and frequency domain features, combined with the optimized XGBoost algorithm to train the diagnostic model, select a combination of feature dimensions with high importance for fault diagnosis.
It improves the accuracy of identification of rail wave grinding faults, reduces the dependence on the environment, ensures the safety of train operations, reduces the burden of data processing, and improves processing speed.
Smart Images

Figure CN120408144A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information processing, and in particular to a rail corrugation fault diagnosis method and a diagnostic model training method. Background Art
[0002] Rails, as a core component of rail transportation, bear all the loads generated by trains. However, as train mileage increases, rail surfaces may suffer from various damages such as uneven welds, track settlement, and rail corrugation.
[0003] Rail corrugation is a common surface damage that manifests as wavy wear along the length of the rail. When a train passes through a corrugated section, it causes significant vibration and noise, affecting passenger comfort and potentially accelerating wear of train and track components, jeopardizing operational safety. Therefore, timely and accurate detection of rail corrugation is crucial to preventing these problems. Summary of the Invention
[0004] A first aspect of the present application provides a rail corrugation fault diagnosis method, comprising:
[0005] Obtaining an acoustic signal to be processed, wherein the acoustic signal to be processed includes an acoustic signal generated when a train contacts a rail when the train is traveling;
[0006] Performing feature extraction on the acoustic signal to be processed based on a target feature dimension combination to obtain a target feature set; wherein the target feature dimension combination includes at least two feature dimensions, each feature dimension corresponds to an acoustic feature type, and the at least two feature dimensions in the target feature dimension combination are selected from a plurality of feature dimensions based on the importance of each feature dimension, and the importance of each feature dimension is determined based on a marginal contribution of each feature dimension to an output of a target diagnostic model;
[0007] The target feature set is processed using a pre-trained target diagnosis model to obtain a rail corrugation fault diagnosis result corresponding to the acoustic signal to be processed.
[0008] A second aspect of the present application provides a method for training a diagnostic model, wherein the diagnostic model is applied to a rail corrugation fault diagnosis method, and the training method comprises:
[0009] Obtaining a training set and a test set, wherein the training set includes a plurality of training sample pairs, each training sample pair includes a training sample and a training label, and the training label is used to indicate whether a rail corrugation fault occurs on the rail corresponding to the training sample;
[0010] Extracting features from the training samples in the training set based on a plurality of feature dimensions to obtain features corresponding to the training samples, where each feature corresponds to a feature dimension;
[0011] Select one from at least two candidate diagnostic models in sequence as the first candidate diagnostic model;
[0012] According to the above features, determine multiple candidate feature dimension combinations for the first candidate diagnostic model; wherein, each candidate feature dimension combination includes at least two candidate feature dimensions;
[0013] Use the training set, and train the first candidate diagnostic model respectively according to the above candidate feature dimension combinations to obtain the corresponding trained first candidate diagnostic model, and each trained first candidate diagnostic model corresponds to a candidate feature dimension combination;
[0014] Use the test set to determine the test performance of the trained first candidate diagnostic model;
[0015] Based on the test performance of the trained first candidate diagnostic model, select the target diagnostic model from the trained first candidate diagnostic models corresponding to each candidate diagnostic model respectively, and determine the candidate feature dimension combination corresponding to the target diagnostic model as the target feature dimension combination; the test performance of the target diagnostic model is better than that of the non-target diagnostic model; the non-target diagnostic model is other trained first candidate diagnostic models except the target diagnostic model among the trained first candidate diagnostic models corresponding to each candidate diagnostic model.
[0016] The third aspect of the present application provides a rail corrugation fault diagnosis device, including:
[0017] An acquisition module, configured to acquire an acoustic signal to be processed, and the acoustic signal to be processed includes the acoustic signal generated when the train travels in contact with the rail;
[0018] An extraction module, configured to extract features from the acoustic signal to be processed according to the target feature dimension combination to obtain a target feature set; wherein, the target feature dimension combination includes at least two feature dimensions, each feature dimension corresponds to an acoustic feature type, and at least two feature dimensions in the target feature dimension combination are selected from several feature dimensions according to the importance of each feature dimension, and the importance of each feature dimension is determined according to the marginal contribution of each feature dimension to the output of the target diagnostic model;
[0019] A processing module, configured to use the pre-trained target diagnostic model to process the target feature set to obtain a rail corrugation fault diagnosis result corresponding to the acoustic signal to be processed.
[0020] The fourth aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:
[0021] The memory is used to store a computer program;
[0022] The processor is used to execute the computer program, so that the electronic device can implement the rail corrugation fault diagnosis method according to the first aspect or any implementation manner of the first aspect.
[0023] The fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the rail corrugation fault diagnosis method according to the first aspect or any implementation manner of the first aspect.
[0024] A rail corrugation fault diagnosis method provided by the present application includes: after obtaining an acoustic signal to be processed, first performing feature extraction on it in multiple feature dimensions to obtain a target feature set. Each feature dimension corresponds to an acoustic feature type. At least two feature dimensions in the target feature dimension combination are selected from several feature dimensions according to the importance of each feature dimension. The importance of each feature dimension is determined according to the marginal contribution of each feature dimension to the output of the target diagnosis model. Then, the target diagnosis model is used to process the features in the obtained target feature set to obtain the rail corrugation fault diagnosis result of the acoustic signal to be processed. The above rail corrugation fault diagnosis method processes according to the features of multiple feature dimensions of the acoustic signal to be processed, and has a higher accuracy. Moreover, the feature dimensions corresponding to the multiple features obtained by extraction are the feature dimensions with high importance selected from several feature dimensions according to the marginal contribution to the output of the target diagnosis model, excluding the feature dimensions with low importance, reducing the data processing burden of the target diagnosis model, and improving the processing speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In combination with the drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.
[0026] Figure 1 is a schematic flow chart of a rail corrugation fault diagnosis method provided by an embodiment of the present application;
[0027] Figure 2 is a schematic flow chart of the process of pre-training a target diagnosis model provided by an embodiment of the present application;
[0028] Figure 3 is a schematic flow chart of determining a training set and a test set provided by an embodiment of the present application;
[0029] Figure 4It is a schematic flow chart for determining multiple candidate feature dimension combinations for the first candidate diagnosis model according to each of these features provided by an embodiment of the present application;
[0030] Figure 5 It is a schematic diagram of the ranking of feature importance scores for each feature dimension of the acoustic signal provided by an embodiment of the present application;
[0031] Figure 6 It is a schematic flow chart for determining the test performance of the first candidate diagnosis model after training by using a test set provided by an embodiment of the present application;
[0032] Figure 7 It is a schematic flow chart for the training method of the diagnosis model provided by an embodiment of the present application;
[0033] Figure 8 It is a simplified schematic flow chart for the training process of the diagnosis model provided by an embodiment of the present application;
[0034] Figure 9 It is a schematic structural diagram of a rail corrugation fault diagnosis device provided by an embodiment of the present application;
[0035] Figure 10 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0036] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, rather than to limit the present application.
[0037] The embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0038] The terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these process, method, product or device.
[0039] In the prior art, for the detection of rail corrugation faults, the detection methods include: acoustic detection, optical detection and vibration detection.
[0040] Among them, this optical detection uses triangulation and two-dimensional coordinate interpolation to evaluate the smoothness of the rail, which can provide intuitive visual feedback. However, it has a high cost and is vulnerable to factors such as weather changes and dust, restricting its application in the actual environment.
[0041] Among them, this vibration detection identifies the corrugation by analyzing the change in the characteristics of the acceleration signal of the axle box or bogie when the train passes through the corrugated section. This method has relatively low requirements for hardware and is suitable for real-time monitoring in a dynamic environment. However, improper installation positions of the sensors may lead to misjudgment, and in some cases, it may interfere with the normal operation of the train, increasing the maintenance difficulty.
[0042] To solve the above problems, acoustic detection is generally used to detect rail corrugation faults. The acoustic detection method mainly judges based on the difference in the sound pressure level of the in-vehicle noise between the corrugated section and the normal section of the vehicle. Specifically, it focuses on the A-weighted sound pressure level at a specific center frequency. However, judging only based on the A-weighted sound pressure level has a single judgment basis and a low recognition accuracy.
[0043] In the acoustic signal detection adopted in this application, the detection of the acoustic signal does not interfere with the normal operation of the train. Moreover, in this application, features of multiple characteristic dimensions of the acoustic signal to be processed are extracted, and based on the features of multiple characteristic dimensions, fault diagnosis of rail corrugation is carried out, enriching the judgment basis and improving the recognition accuracy.
[0044] Figure 1 It is a schematic flow chart of a method for diagnosing rail corrugation faults provided by an embodiment of this application. As Figure 1 shown, a method for diagnosing rail corrugation faults provided by an embodiment of this application may include steps 101 to 103, and the following will describe these steps in detail.
[0045] 101. Obtain the acoustic signal to be processed, where the acoustic signal to be processed includes the acoustic signal generated when the train travels and contacts the rail.
[0046] Among them, highly sensitive acoustic sensors are set on the train, and the acoustic sensors collect the sounds in the surrounding environment when the train travels. The sounds in the surrounding environment when the train travels include the sounds generated when the train travels and contacts the rail and the noise.
[0047] Among them, the acoustic sensors can be set inside the train carriage, and the installation position is not restricted. Different from the vibration sensors, installing them in specific positions may interfere with the normal operation of the train, and it can ensure the safety of the train's travel; the acoustic signals detected by the acoustic sensors are not affected by bad weather and are more widely applied in the actual environment.
[0048] In a possible implementation, when diagnosing rail corrugation faults in real time during train operation, an acoustic sensor is set on the running train to collect the acoustic signals in the surrounding environment during train operation in real time. For every set duration of the to-be-processed acoustic signals obtained, the subsequent steps 102-103 can be carried out.
[0049] In a possible implementation, when performing non-real-time processing on the acoustic signals generated during train operation, the to-be-processed acoustic signals collected over a period of time are intercepted into multiple segments according to the set duration, and the subsequent steps 102-103 are carried out for each segment of the to-be-processed acoustic signals with the set duration.
[0050] Among them, a time window can be determined according to the set duration, and interception is carried out based on this time window. During the interception process, two adjacent time windows overlap by half of the time window duration, ensuring that the two adjacent to-be-processed acoustic signals partially overlap, and preventing signal feature loss caused by non-periodic truncation.
[0051] In a possible implementation, the to-be-processed acoustic signals can be preprocessed first to improve their quality, and then feature extraction in multiple feature dimensions is carried out on the preprocessed acoustic signals.
[0052] As an example, the preprocessing of the to-be-processed acoustic signals can include: filtering, denoising, filling missing values, and normalization processing, etc.
[0053] In a possible implementation, cleaning and repair can be carried out through filtering techniques and statistical methods to achieve filtering, denoising, and filling missing values. The specific manner of this preprocessing is not limited in this application.
[0054] Among them, for the to-be-processed acoustic signals after preprocessing, their signal quality is better than that of the to-be-processed acoustic signals before preprocessing, which can be less noise, abnormal points are removed, there are no missing values, etc.
[0055] In a possible implementation, the normalization process in this preprocessing is specifically to convert acoustic signals from different sources into a unified form.
[0056] As an example, the unified form can be a unified sampling rate, the same quantization bit number, normalization to a specific amplitude range, etc.
[0057] In a possible implementation, the to-be-processed acoustic signals after this preprocessing can also be segmented, and subsequent processing is carried out on each segment of the acoustic signals, reducing the data processing volume of single feature extraction and also reducing the data processing volume of single prediction by the target diagnosis model.
[0058] As an example, the time period is set to 1 second, the time interval used for slicing is 0.5 seconds, and any adjacent slices have an overlapping area of 0.5 seconds, which can ensure the continuity of the acoustic signal segments and achieve higher accuracy in feature extraction and subsequent prediction.
[0059] 102. Based on the target feature dimension combination, feature extraction is performed on the acoustic signal to be processed to obtain a target feature set.
[0060] Among them, the target feature dimension combination includes at least two feature dimensions, each feature dimension corresponds to an acoustic feature type, and at least two feature dimensions in the target feature dimension combination are selected from several feature dimensions based on the importance of each feature dimension, and the importance of each feature dimension is determined based on the marginal contribution of each feature dimension to the output of the target diagnostic model.
[0061] Here, feature extraction of multiple feature dimensions is performed on the acoustic signal to be processed to obtain a target feature set. Correspondingly, the target feature set includes features of multiple feature dimensions.
[0062] Among them, each feature dimension corresponds to an acoustic feature type, and the features in the target feature set include multiple acoustic feature types, so that the subsequent target diagnosis model can combine the acoustic features of multiple feature dimensions for prediction.
[0063] Among them, feature extraction is performed on the acoustic signal to be processed in both time domain and frequency domain. One or more features can be extracted in the time domain, and one or more features can be extracted in the frequency domain, thereby realizing feature extraction in multiple feature dimensions in the time domain and frequency domain. The diagnostic model performs multi-faceted analysis on the features corresponding to the frequency domain and time domain, respectively, to deeply explore the performance of the acoustic signal to be processed in the time domain and frequency domain, thereby achieving more accurate corrugation recognition.
[0064] Among them, according to the determined at least two feature dimensions, feature extraction is performed in the preprocessed acoustic signal to be processed, and at least two features are extracted, and the at least two features include at least one time domain feature and at least one frequency domain feature, or at least two time domain features, or at least two frequency domain features.
[0065] Among them, the feature dimension corresponding to the extracted feature is predetermined. When determining the feature dimension, it is selected based on the importance of each feature dimension's marginal contribution to the output of the target diagnostic model. The importance of the selected feature dimension is higher than that of the unselected feature dimension.
[0066] Among them, the marginal contribution of a certain feature dimension to a certain diagnostic model is: the change in the output of the target diagnostic model caused by the features including this feature dimension and the features not including this feature dimension in the input of the diagnostic model. The greater the marginal contribution of a certain feature dimension to the output of the target diagnostic model, the higher the importance of this feature dimension.
[0067] Since the features of different feature dimensions of the acoustic signal have different marginal contributions to the output of the target diagnostic model, and the marginal contributions of the features of each feature dimension to the output of the target diagnostic model are affected by the diagnostic model adopted, the technical problem to be solved, the acoustic signal to be processed, etc. Therefore, the obtained target feature set is a combination of features of several feature dimensions with high importance, which is the one that best fits the current scenario of processing the acoustic signal to be processed and can provide the best input for realizing the accurate identification of rail corrugation faults.
[0068] In a possible implementation, the at least two feature dimensions include at least two of the following: A-weighted sound pressure level feature, root mean square energy, peak factor, impulse factor, margin factor, kurtosis factor, skewness factor, zero crossing rate, average spectral information entropy, average spectral centroid, average spectral kurtosis, Mel frequency cepstral coefficients.
[0069] Among them, the A-weighted sound pressure level feature is obtained by specifically weighting (A-weighting) and modifying the sound pressure levels of different frequency components in the sound to simulate the frequency perception characteristics of the human ear, and then summing the energies of all the modified frequency components (such as adding the squares and then taking the logarithm), resulting in the total sound pressure level; the root mean square (RMS) value is an important parameter used to describe the sound intensity; the crest factor is the ratio of the signal peak value to the effective value (RMS, Root Mean Square), representing the extreme degree of the peak in the waveform, and can be used to evaluate the statistical quantity of the relative intensity of the peak in the signal, and can be used to detect anomalies such as instantaneous or intermittent impacts in the signal; the impulse factor is the ratio of the signal peak value to the rectified average value (average of absolute values), and this ratio can reflect the proportion of the impact component or peak component in the signal relative to the overall signal intensity; the margin factor is the ratio of the signal peak value to the root mean square of the absolute value, and can be used to detect the condition of rail corrugation; the kurtosis factor can evaluate anomalies such as instantaneous impacts caused by rail corrugation faults in the signal; the skewness factor can evaluate anomalies such as misalignment and imbalance of the signal axis caused by rail corrugation; the zero-crossing rate (ZCR) is the number of times the signal passes through zero in each frame, and is a key feature for identifying the classification of knocking sounds; the average spectral information entropy is used to describe the uniformity of the energy distribution in the signal spectrum and can evaluate the distribution of different frequency components; the average spectral centroid is used to describe the average energy distribution of the signal spectrum and can evaluate long-term anomalies; the average spectral kurtosis can evaluate the overall kurtosis characteristics of the signal in the frequency domain; the Mel Frequency Cepstral Coefficients (MFCCs) reflect the energy distribution of the signal at different Mel frequencies.
[0070] Among them, the time-domain features can include the A-weighted sound pressure level feature, the root mean square value, the crest factor, the impulse factor, the margin factor, the kurtosis factor, the skewness factor, the zero-crossing rate, etc., and the frequency-domain features can include the average spectral information entropy, the average spectral centroid, the average spectral kurtosis, and the Mel Frequency Cepstral Coefficients. The above features can better capture the specific acoustic patterns of rail corrugation and improve the accuracy and reliability of identification.
[0071] Among them, the rail corrugation is a surface damage, presenting as wavy wear along the length of the rail, with a wavelength range between 25 millimeters and 1200 millimeters, and the induced vibration frequency is about 50 Hertz (Hz) to 1200 Hz. The A-weighted sound pressure level can be processed for specific frequencies, and the specific frequencies can be those in the range of 50 Hz to 1200 Hz, such as 315 Hz, 400 Hz, 500 Hz, 630 Hz, 800 Hz, 1000 Hz, etc.
[0072] Among them, the features of multiple feature dimensions extracted are used as the input information for the subsequent input diagnostic model. The diagnostic model processes based on the features of the multiple feature dimensions extracted to obtain the rail corrugation fault diagnosis result of the acoustic signal to be processed.
[0073] Among them, the process of determining the feature dimensions in the target feature set can be to first determine the marginal contribution of each feature dimension to each diagnostic model, and then determine the importance of the corresponding feature dimension based on the marginal contribution. Subsequently Figure 2 The corresponding relevant descriptions are elaborated in detail.
[0074] 103. Use the pre-trained target diagnostic model to process the target feature set to obtain the rail corrugation fault diagnosis result corresponding to the acoustic signal to be processed.
[0075] Among them, use the pre-trained target diagnostic model to process the features of multiple feature dimensions of the acoustic signal to be processed to obtain the rail corrugation fault diagnosis result of the acoustic signal to be processed.
[0076] Among them, the diagnosis result includes: rail corrugation fault occurs, rail corrugation fault does not occur. When the rail corrugation fault occurs, it can be divided into multiple levels according to the severity of the fault. Specifically, it can be set according to the actual situation and is not limited in this application.
[0077] In a possible implementation, the target diagnostic model can adopt the optimized XGBoost (eXtreme Gradient Boosting) algorithm.
[0078] The basic principle of the optimized XGBoost model is described as follows.
[0079] XGBoost is an efficient implementation method of the GBDT (Gradient Boosting Decision Tree) algorithm. The XGBoost model is a tree ensemble model, which consists of multiple decision trees. The final prediction result of the XGBoost model is the accumulation of the prediction scores of each tree, which is expressed by the following formula:
[0080] (1)
[0081] Among them, represents the number of decision trees in the XGBoost model, represents the th sample, is the th tree model, represents the The predicted class label of a sample is a collection of decision tree models
[0082] XGBoost uses the gradient boosting algorithm to train decision trees. The gradient boosting algorithm is an iterative algorithm that optimizes the model by continuously adding new decision trees. The construction of each new tree aims to correct the prediction errors of the previous tree. The objective function of XGBoost is defined as follows:
[0083] (2)
[0084] Where is the true label is the predicted label represents the total number of samples is the loss function is the regularization term that controls the model complexity. This regularization term is defined as:
[0085] (3)
[0086] Where is the prediction function is the number of leaf nodes in the tree is the regularization parameter is the regularization parameter is the weight of the
[0087] In formula (2), since the parameters contain functions, traditional optimization algorithms cannot be used for optimization. Assume is at the th iteration, then there is:
[0088] (4)
[0089] Where is the prediction function, that is, the tree model added in the th iteration. Expand formula (4) by Taylor's second-order expansion, and the expansion result is as follows:
[0090] (5)
[0091] Where and are the first-order derivative and second-order derivative of. Since is a constant term and can be removed, define As the subscript set of the samples under the th leaf node. Formula (5) can be transformed into:
[0092] (6)
[0093] Taking the derivative of can obtain
[0094] (7)
[0095] Substituting into formula (6), the minimum value of the objective function can be obtained as:
[0096] (8)
[0097] Among them, the optimized XGBoost performs in-depth optimization at the algorithm level, introduces a regularization term to control the complexity of the diagnostic model, uses the first-order and second-order derivatives of the objective function for quadratic Taylor expansion to capture the changes in the loss function, and improves the convergence speed of the diagnostic model. The optimized XGBoost algorithm is applicable to large-scale industrial data processing.
[0098] In a possible implementation, a target diagnostic model is pre-trained, and the Figure 3 related descriptions in the following
[0099] In this embodiment, after obtaining the acoustic signal to be processed, first perform feature extraction on it in multiple feature dimensions to obtain a target feature set. Each feature dimension corresponds to an acoustic feature type. At least two feature dimensions in the target feature dimension combination are selected from several feature dimensions according to the importance of each feature dimension. The importance of each feature dimension is determined according to the marginal contribution of each feature dimension to the output of the target diagnostic model. Then, use the target diagnostic model to process the features of multiple feature dimensions in the obtained target feature set to obtain the rail corrugation fault diagnosis result of the acoustic signal to be processed. Compared with only using the A-weighted sound pressure level to determine the rail corrugation fault, in this process, it is processed according to the features of multiple feature dimensions of the acoustic signal to be processed, and the accuracy is higher. Moreover, the multiple feature dimensions for extracting the features are the feature dimensions with high importance selected from several feature dimensions according to the marginal contribution to the output of the target diagnostic model, excluding the feature dimensions with low importance, reducing the data processing burden of the target diagnostic model and improving the processing speed.
[0100] Figure 2 is a schematic flowchart of the process for pre-training the target diagnostic model provided by the embodiment of this application, which may include steps 201 to 205. The following will describe these steps in detail.
[0101] 201. Obtain a training set and a test set. The training set contains a number of training sample pairs, and each training sample pair contains a training sample and a training label. The training label is used to characterize whether the rail corresponding to the training sample has a rail corrugation fault.
[0102] Among them, a number of features of multiple feature dimensions required during the application of the diagnostic model are determined in advance.
[0103] Specifically, when determining the feature dimensions in the target feature set, it is determined using the training set. First, obtain the training set, and then use the training set and the test set to train the candidate diagnostic model and perform a performance test to select the target diagnostic model.
[0104] Among them, the obtained training set contains a number of training sample pairs. The training sample pair contains a training sample and a training label. The training sample is an acoustic signal as a sample, and the training sample contains positive samples and negative samples. The label of the positive sample is that a rail corrugation fault has occurred, and the label of the negative sample is that no rail corrugation fault has occurred.
[0105] Among them, the obtained test set contains a number of test sample pairs. The test sample pair contains a test sample and a test label. The test sample is an acoustic signal as a sample and is used to be input into the first candidate diagnostic model trained by the training set to test the performance of the trained first candidate diagnostic model.
[0106] In a possible implementation, the obtained test set can be stored, and during the process of performing a performance test on the trained candidate diagnostic model, the test set is obtained from the storage location for performance testing.
[0107] In a possible implementation, the training set and the test set can be acoustic signals obtained from the real scenario of the train's movement. The acoustic signals are used as samples and are obtained without adding labels, as described below Figure 3 which illustrates the process of obtaining the training set and the test set.
[0108] Figure 3 is a schematic flowchart of the process for determining the training set and the test set provided by an embodiment of the present application, which may include steps 2011 to 2014. The following will describe these steps in detail.
[0109] 2011. Obtain the original sample acoustic signal.
[0110] Among them, the original sample acoustic signal can be a relatively long audio collected in the actual running scenario of the train, and the audio contains the audio with a rail corrugation fault and the audio without a rail corrugation fault.
[0111] Among them, the sound generated by the train contacting the rail can be collected in the traveling environment where the train is located, and this sound serves as the original sample acoustic signal.
[0112] Among them, highly sensitive acoustic sensors can be set in the running train carriages. These acoustic sensors collect the sounds in the surrounding environment when the train is traveling, and the sounds in the surrounding environment when the train is traveling include the sound generated by the train contacting the rail.
[0113] In specific implementation, the acoustic signals collected on-site may also include noise, outliers, or missing values, etc. The collected acoustic signals can be preprocessed in advance to improve the quality of the acoustic signals and provide a better training set and test set for subsequent training.
[0114] Among them, the process of preprocessing the collected original sample acoustic signals can be the same as the preprocessing process of the acoustic signals to be processed, which will not be elaborated here.
[0115] Among them, the quality of the original sample acoustic signals after preprocessing is better than that before preprocessing. For example, the signal-to-noise ratio of the original sample acoustic signals after preprocessing is higher than that before preprocessing.
[0116] 2012. Based on a set time period, intercept the original sample acoustic signals to obtain a sample acoustic signal set, and this sample acoustic signal set contains several sample acoustic signals.
[0117] Among them, this set time period can be used as a time window. The original sample acoustic signals are sliced using the time window to obtain multiple sample acoustic signals, and these multiple sample acoustic signals form a sample acoustic signal set.
[0118] Among them, since most of the rails are normal and only a small part will have corrugation faults, there may be a small part of fault samples with rail corrugation faults in the obtained sample acoustic signal set, and most of them are normal samples.
[0119] As an example, the set time period is 1 second, and the time interval for slicing is 0.5 second, which can ensure the continuity of the sample acoustic signals.
[0120] 2013. Add labels to each sample acoustic signal to obtain a data set, and this label is used to characterize whether the rail corresponding to the corresponding sample acoustic signal has a rail corrugation fault.
[0121] Among them, to provide sample pairs in the subsequent training process, labels need to be added to each sample acoustic signal, and this label is used to characterize whether the rail corresponding to the corresponding sample acoustic signal has a rail corrugation fault.
[0122] Among them, manual or automated methods can be used to add tags.
[0123] In one implementation, existing rail maintenance records or inspection reports can be utilized. Based on the timestamp and speed information of the train passing through a specific section, the mileage traveled by the train can be calculated. By comparing the calculated mileage with the mileage of known rail corrugation fault sections, it is possible to preliminarily determine which time period's acoustic signals correspond to the rail corrugation fault sections. For the acoustic signals of the preliminarily confirmed rail corrugation fault sections, an acoustic feature matching algorithm is then applied to further confirm whether there are typical acoustic features of rail corrugation faults, and the acoustic signal segments with typical acoustic features of rail corrugation faults are labeled to achieve automatic labeling.
[0124] In one implementation, a visualization tool can be developed to allow experts to view the original acoustic signals and their corresponding feature maps, so that experts can intuitively evaluate the results of automatic labeling. Moreover, this visualization tool can modify the automatic labeling by providing marking tools for experts to correct incorrect labels or add new labels.
[0125] 2014. Based on a set ratio, the data set is divided into a test set and a training set, and the ratio of the number of test sample pairs included in the test set to the number of training sample pairs included in the training set satisfies the set ratio.
[0126] Among them, the situations of the test sample pairs included in the test set and the training sample pairs included in the training set are set in advance, such as setting a ratio, setting the ratio of the number of test sample pairs included in the test set to the number of training sample pairs included in the training set.
[0127] Among them, after obtaining the data set, based on the set ratio, the sample pairs included in the data set are divided into a test set and a training set to ensure the accuracy and effectiveness of the subsequent training and performance evaluation of the diagnostic model.
[0128] Among them, the test set and the training set are mutually exclusive sets, and in the process of training the candidate diagnostic model, in a cross-validation manner, the sample pairs in the data set can be alternately used as the training set and the test set.
[0129] As an example, the set ratio is 2:8, 20% of the sample pairs in the data set are used as the test set, and 80% of the sample pairs are used as the training set. The specific value of the set ratio can be set according to the actual situation and is not limited in this application.
[0130] 202. Feature extraction is performed on the training samples in the training set based on several feature dimensions to obtain the respective features corresponding to the training samples, and each feature corresponds to a feature dimension.
[0131] Among them, the several feature dimensions can represent the acoustic characteristics of the acoustic signal from different feature dimensions, but different feature dimensions have different acoustic characterization situations for the acoustic signal, which can be specifically reflected in the determination of the marginal contribution to the target diagnosis model.
[0132] Among them, the several feature dimensions are all possible feature dimensions that can be used for feature extraction of the acoustic signal.
[0133] For example, the several feature dimensions include the following feature dimensions: A-weighted sound pressure level feature, root mean square value of energy, peak factor, impulse factor, margin factor, kurtosis factor, skewness factor, zero crossing rate, average spectral information entropy, average spectral centroid, average spectral kurtosis, Mel frequency cepstral coefficients. Of course, in specific implementation, other feature dimensions can be added according to the actual situation, which is not limited in this application.
[0134] Specifically, for each training sample in the training set, feature extraction is performed separately from the several feature dimensions to obtain the features of each feature dimension in the several feature dimensions of each training sample.
[0135] In this embodiment, some feature dimensions are determined from the several feature dimensions as the feature dimensions in the target feature set. The process of determining the feature dimensions is executed synchronously with the process of training and selecting the target diagnosis model. For the specific process, refer to steps 203 to 206.
[0136] 203. Select one from at least two candidate diagnosis models in sequence as the first candidate diagnosis model;
[0137] Among them, multiple candidate diagnosis models can be preset in advance. Each candidate diagnosis model can adopt different algorithms. Exemplarily, each diagnosis model can respectively adopt support vector machine (SVM, Support Vector Machine), decision tree (DT, Decision Tree), K-nearest neighbor (KNN, k-Nearest Neighbors), Gaussian naive Bayes (GNB, Gaussian Naive Bayes), random forest (RF, Random Forest), and extreme gradient boosting (XGBoost), etc.
[0138] Among them, in the multiple candidate diagnosis models, one is selected in sequence as the first candidate diagnosis model. The first candidate diagnosis model is used to explain the training process of each candidate diagnosis model, and the training process of each candidate diagnosis model can refer to the training process of the first candidate diagnosis model.
[0139] As an example, first use SVM as the first candidate diagnostic model for training to obtain multiple trained first candidate diagnostic models, and all of the multiple trained first candidate diagnostic models use the SVM algorithm; then use the decision tree as the first candidate diagnostic model for training to obtain multiple trained first candidate diagnostic models, and all of the multiple trained first candidate diagnostic models use the decision tree algorithm; and so on, to obtain multiple trained first candidate diagnostic models corresponding to each of the preset candidate diagnostic models respectively. Among the subsequent multiple trained first candidate diagnostic models, select the one with the best test performance as the target diagnostic model.
[0140] Among them, during the process of training the candidate diagnostic models, each of at least two candidate diagnostic models is sequentially used as the first candidate diagnostic model, and the subsequent steps 204-207 are respectively executed.
[0141] In a possible implementation, if there is only one candidate diagnostic model, then this candidate diagnostic model is used as the first candidate diagnostic model and trained using multiple candidate feature dimensions to determine the target diagnostic model from the multiple trained first candidate diagnostic models; among them, each candidate feature dimension is determined from several feature dimensions according to the importance of each feature dimension, and the specific details are described in the following detailed description.
[0142] 204. Determine multiple candidate feature dimension combinations for the first candidate diagnostic model according to these features; among them, each candidate feature dimension combination contains at least two candidate feature dimensions, and each candidate feature dimension is determined from several feature dimensions according to the importance of each feature dimension;
[0143] Among them, for the first candidate diagnostic model, multiple candidate feature dimension combinations are determined. The candidate feature dimensions included in each candidate feature dimension combination are determined from several feature dimensions according to their importance, and the importance of each candidate feature dimension is determined according to the marginal contribution of this feature dimension to the first candidate diagnostic model.
[0144] Among them, when the first candidate diagnostic model involves hyperparameters, multiple values of the hyperparameters can be preset for the first candidate diagnostic model, and each value of the hyperparameters is substituted into the first candidate diagnostic model to obtain the first candidate diagnostic model using the corresponding value combinations, so as to determine the corresponding multiple candidate feature dimension combinations for the first candidate diagnostic model of each value combination.
[0145] Among them, when the first candidate diagnostic model does not involve hyperparameters, there is no need to consider the values of the hyperparameters, and multiple candidate feature dimension combinations are determined for the first candidate diagnostic model.
[0146] In a possible implementation, multiple candidate feature dimension combinations corresponding to the first candidate diagnostic model can be determined by combining the hyperparameters in the first candidate diagnostic model, as follows Figure 4 The process of determining multiple candidate feature dimension combinations for the first candidate diagnostic model based on each of these features is described therein.
[0147] Figure 4 FIG. is a schematic flowchart of a process for determining multiple candidate feature dimension combinations for a first candidate diagnostic model according to each of these features in the case where the first candidate diagnostic model includes hyperparameters provided by an embodiment of the present application, and may include steps 2041 to 2046, which are described in detail below respectively.
[0148] 2041. Obtain at least two candidate value combinations of the hyperparameters in the first candidate diagnostic model;
[0149] Wherein, the hyperparameter is a parameter set by a human and cannot be automatically learned from data. Correspondingly, for at least two candidate diagnostic models, a trainer manually sets the value range of the hyperparameters in each candidate diagnostic model respectively.
[0150] Wherein, the hyperparameters of each diagnostic model in at least two candidate diagnostic models may be the same or different, and the present application does not limit the hyperparameters of the candidate diagnostic models and their value ranges.
[0151] In a possible implementation, the value range of the hyperparameters can be set for the at least two candidate diagnostic models first, then one of the at least two candidate diagnostic models is selected as the first candidate diagnostic model in sequence, and then multiple candidate value combinations are selected within the value range of the hyperparameters of the first candidate diagnostic model.
[0152] In a possible implementation, one of the at least two candidate diagnostic models can be selected as the first candidate diagnostic model in sequence first, then the value range of the hyperparameters of the first candidate diagnostic model is set, and then multiple candidate value combinations are selected from within the value range.
[0153] Wherein, each candidate diagnostic model in the at least two candidate diagnostic models needs to be used as the first candidate diagnostic model for relevant processing.
[0154] As an example, for the first candidate diagnostic model adopting the XGBoost algorithm, the set hyperparameters may include but are not limited to tree-related parameters, regularization parameters, learning rate, etc.
[0155] Wherein, for each first candidate diagnostic model, a hyperparameter grid is determined by using its respective hyperparameters and the corresponding value range, and each grid corresponds to a candidate value combination of the hyperparameters. During the process of training the first candidate diagnostic model, a grid search method is adopted to determine the candidate value combination.
[0156] In one possible implementation, the value range of the hyperparameters can be pre-set for the first candidate diagnostic model, and the selection step size of each hyperparameter can be set. The selection step size is used to determine the hyperparameter grid, and each grid corresponds to a candidate value combination of the hyperparameters. During the training of the first candidate diagnostic model, the grid search method is used to obtain multiple possible values of the hyperparameter. Similarly, the above method can be used to determine the possible values of each hyperparameter in the first candidate diagnostic model.
[0157] For example, grid search is used to optimize six key hyperparameters of the XGBoost model: number of iterations, learning rate, maximum tree depth, tree feature sampling ratio, sample sampling ratio, and minimum sample weight of leaf nodes.
[0158] Table 1 below is a list of main hyperparameter optimizations in an XGBoost algorithm provided in an embodiment of the present application, including the main hyperparameters of the XGBoost algorithm diagnostic model and the range of grid search.
[0159] Table 1
[0160]
[0161] The value range and step size of each hyperparameter in Table 1 can be used to determine the selectable value of each hyperparameter, and the selectable values of different hyperparameters can be combined to obtain a candidate value combination.
[0162] Among them, when determining the selectable values, the lower limit of the value range of the hyperparameter is taken as the first selectable value, and the step size is used as the increment. The first selectable value is increased by one increment to obtain the second selectable value; the first selectable value is increased by two increments to obtain the third selectable value, until the upper limit of the value range of the hyperparameter is reached. The final maximum selectable value can be the upper limit of the value range of the hyperparameter, or a value less than the upper limit of the value range of the hyperparameter.
[0163] As an example, the main hyperparameters of the XGBoost algorithm include: number of iterations, learning rate, and maximum tree depth. The range and step size of each hyperparameter can be found in Table 1. A grid search method is used to exhaustively enumerate possible combinations of these three hyperparameter values. The number of iterations can be selected from 10 possible values: 50, 100, 150, ..., 500; the learning rate can be selected from 5 possible values: 0.1, 0.2, 0.3, ..., 0.5; and the maximum tree depth can be selected from 8 possible values: 3, 4, 5, ..., 10. These three hyperparameters are combined according to their possible values to obtain a variety of candidate value combinations.
[0164] 2042. Determine one of the at least two candidate value combinations in sequence as a first candidate value combination;
[0165] Among them, for each candidate diagnostic model, during the process of selecting a candidate value combination of its hyperparameters, it is generally determined by traversing the hyperparameter combinations.
[0166] In this embodiment, during the process of training the first candidate diagnostic model, a method of fixing one combination and traversing the other combination can be adopted (that is, fixing the candidate value combination of the hyperparameters and traversing the candidate feature dimension combinations; or fixing the candidate feature dimension combinations and traversing the candidate value combinations of the hyperparameters), combining the candidate feature dimension combinations and the candidate value combinations, and training the first candidate diagnostic model with a candidate feature dimension combination and a candidate value combination obtained from the combination.
[0167] Among them, among the multiple candidate value combinations of the first candidate diagnostic model, one is determined as the first candidate value combination.
[0168] Among them, this first candidate value combination is only for facilitating the explanation of the process subsequently. In specific implementation, the following processing process is respectively performed on each candidate value combination in each candidate diagnostic model.
[0169] As an example, according to the description related to Table 1, when the main hyperparameters of the XGBoost algorithm include: the number of iterations, the learning rate, and the maximum depth of the tree, the three hyperparameters are combined according to the optional values respectively to obtain multiple candidate value combinations. For each candidate value combination, the subsequent processing steps are respectively executed to determine the importance degree of the feature dimension corresponding to each feature relative to the first candidate diagnostic model.
[0170] Among them, during the subsequent training process, the candidate feature dimension combinations corresponding to each candidate value combination of the hyperparameters are adopted to train the corresponding first candidate diagnostic model by using each candidate value combination and the corresponding multiple candidate feature dimension combinations.
[0171] In a possible implementation, when it is stipulated that the diagnostic model adopts a specific algorithm and there is only one candidate diagnostic model, then, candidate value combinations are respectively selected for the candidate diagnostic model, and multiple candidate feature dimensions are used for training each candidate value combination to determine the trained target diagnostic model.
[0172] 2043. Determine the marginal contribution of each feature dimension to the output of the first candidate diagnostic model that adopts the first candidate value combination according to each feature;
[0173] Among them, the marginal contribution of the feature of a certain feature dimension to the output of any candidate diagnostic model is the change in the output of the candidate diagnostic model caused by adding the feature of this feature dimension to the input of the candidate diagnostic model and not adding the feature of this feature dimension. The greater the marginal contribution of the feature of a certain feature dimension to the output of the candidate diagnostic model, the higher the importance of this feature dimension.
[0174] Specifically, when there are N feature dimensions, all the features of the N feature dimensions are used as the first feature combination. The first feature combination includes feature A, feature B, feature C, feature D....., feature N; based on the first feature combination, feature A in the first feature combination is removed to obtain the second feature combination; then based on the first feature combination, feature B (different from feature A) in the first feature combination is removed to obtain the third feature combination, and so on. One feature is removed from the first feature combination in turn to obtain N feature combinations; where N is a positive integer; the first feature combination, the second feature combination, the third feature combination....., the Nth feature combination are respectively input into the first candidate diagnostic model using the first candidate value combination to obtain the output results corresponding to each feature combination. The output result corresponding to the first feature combination is compared with the output results corresponding to each of the remaining feature combinations to determine the influence of each removed feature dimension on the output of the first candidate diagnostic model using the first candidate value combination. This output influence is the marginal contribution of the removed feature dimension to the output of the first candidate diagnostic model using the first candidate value combination.
[0175] As an example, the first candidate diagnostic model is the XGBoost model, and the selected first candidate value combination is {number of iterations = 50, learning rate = 0.2, maximum depth of the tree = 5, tree feature sampling ratio = 0.6, sample sampling ratio for each tree = 0.7, minimum sum of sample weights of leaf nodes = 3}. This first candidate value combination is used as the hyperparameter value of the XGBoost model to obtain an XGBoost model with known hyperparameters. According to the aforementioned first feature combination, second feature combination, ……, Nth feature combination, input them into the XGBoost model with known hyperparameters to obtain the output results. By comparing the output result corresponding to the second feature combination with the output result corresponding to the first feature combination, determine the influence of feature A on the output of the XGBoost model with known hyperparameters. This output influence is the marginal contribution of the removed feature dimension A to the output of the XGBoost model with known hyperparameters. Based on this, determine the marginal contribution of each feature dimension to the output of the XGBoost model with known hyperparameters.
[0176] In a possible implementation, the Shapley Additive exPlanation (SHAP value) can be used to represent the marginal contribution. For the features of each feature dimension, calculate the contribution degree of the features of this feature dimension to the prediction result of the corresponding first candidate diagnostic model, that is, calculate the SHAP value, so as to determine the marginal contribution corresponding to the features of each feature dimension.
[0177] 2044. Determine the importance of each feature to the first candidate diagnostic model using the first candidate value combination based on the marginal contribution of each feature dimension to the output of the first candidate diagnostic model using the first candidate value combination.
[0178] Among them, based on the marginal contribution of each feature dimension to the first candidate diagnostic model using the first candidate value combination, the importance of this feature dimension can be determined, and this importance is the marginal contribution of this feature dimension to the first candidate diagnostic model using the first candidate value combination.
[0179] Among them, for the same candidate diagnostic model, the hyperparameters adopt each candidate value combination, and the above process is executed respectively to determine the importance of the feature dimensions of the features corresponding to each candidate diagnostic model of each candidate value combination.
[0180] Similarly, for each candidate diagnostic model, the hyperparameters adopt each candidate value combination, and the above process is executed respectively to determine the importance of the feature dimensions of the features corresponding to each candidate diagnostic model of each candidate value combination.
[0181] Among them, after determining the marginal contribution of the features of each feature dimension to the output of the corresponding candidate diagnostic model, use this marginal contribution to determine the importance of the corresponding feature dimension.
[0182] Among them, the marginal contribution is positively correlated with the importance. The greater the marginal contribution, the greater the importance.
[0183] Among them, when it is stipulated that the diagnostic model adopts a specific algorithm, there is only one candidate diagnostic model. When determining different candidate value combinations for this candidate diagnostic model, calculate the marginal contribution of the features of each feature dimension of the training samples to its output. Correspondingly, when this candidate diagnostic model adopts different candidate value combinations, obtain the importance of each feature dimension, and then based on the importance of each feature dimension, determine at least two feature dimensions of the acoustic signal features included in the input information of this candidate diagnostic model; in the specific implementation manner, the feature dimension with high importance can be used as the feature dimension corresponding to the input information of the selected diagnostic model to improve the accuracy of rail corrugation fault diagnosis and recognition.
[0184] Among them, when there are multiple candidate diagnostic models, the marginal contribution of each feature dimension of the training samples to each candidate diagnostic model is determined respectively. Correspondingly, the importance of each feature dimension corresponding to each candidate diagnostic model is obtained.
[0185] Among them, when there are multiple candidate diagnostic models, for the same feature dimension, the marginal contribution output to different candidate diagnostic models may be the same or different, that is, the importance of this feature dimension to different candidate diagnostic models may be the same or different.
[0186] 2045. According to the importance of each feature to the first candidate diagnostic model using the first candidate value combination, select at least two feature dimensions from several feature dimensions corresponding to each feature to form at least one candidate feature dimension combination.
[0187] Among them, according to the importance of each feature dimension to the candidate diagnostic model, select two or more feature dimensions. Subsequently, in the actual application scenario, when processing the acoustic signal to be processed, use the at least two selected feature dimensions as the extraction feature dimensions for extracting the features of the acoustic signal to be processed, and obtain the target feature set.
[0188] In a possible implementation, when there are multiple candidate feature dimension combinations corresponding to the first candidate diagnostic model using the first candidate value combination, combination can be performed according to a threshold based on the importance of each feature dimension, or combination can be performed according to the importance ranking.
[0189] In a possible implementation, when there is only one candidate diagnostic model, use the candidate diagnostic model as the first candidate diagnostic model. According to the importance of each feature dimension, first sort the importance of each feature dimension, and select at least two feature dimensions that meet the preset importance rule as the target feature dimension combination.
[0190] Among them, the preset importance rule can be to select feature dimensions with importance greater than the set threshold, or to select a set number of feature dimensions ranked in the front from large to small in terms of importance, etc.
[0191] In a possible implementation, step 2045 can be implemented by the following process:
[0192] According to the importance of each feature dimension to the first candidate diagnostic model using the first candidate value combination, select at least two feature dimensions with importance greater than the preset importance threshold from several feature dimensions corresponding to each feature; based on the at least two feature dimensions, determine the candidate feature dimension combination corresponding to the first candidate diagnostic model using the first candidate value combination.
[0193] Among them, for the first candidate diagnosis model, an importance threshold can be preset. When selecting a candidate feature dimension combination, some feature dimensions among the several feature dimensions are excluded, and the importance of the excluded feature dimensions is not greater than the preset importance threshold. Correspondingly, each feature dimension with an importance greater than the preset importance threshold is retained, and the combination of each feature dimension with an importance greater than the preset importance threshold is used as the candidate feature dimension combination of the first candidate diagnosis model that adopts the first candidate value combination.
[0194] As an example, there are 100 feature dimensions, the preset importance threshold is 0.4, and the value range of the importance of each feature dimension is (0, 1). Among the 100 feature dimensions corresponding to the first candidate diagnosis model that adopts the first candidate value combination, the feature dimensions with an importance greater than 0.4 are selected, and 59 feature dimensions with an importance greater than 0.4 are obtained. The importance of the remaining 41 feature dimensions is not greater than 0.4. The 59 feature dimensions are combined into a candidate feature dimension combination, which is used as the candidate feature dimension combination of the first candidate diagnosis model that adopts the first candidate value combination.
[0195] In this application, in combination with the training process of the first candidate diagnosis model, a target feature dimension combination is selected from each candidate feature dimension combination corresponding to each candidate diagnosis model, and the target feature dimension combination is the candidate feature dimension combination corresponding to the selected target diagnosis model.
[0196] As an example, when there are 10 candidate feature dimension combinations, for a candidate value combination of the hyperparameters in the first candidate diagnosis model, using the training set, the features of the training set are respectively extracted with the 10 candidate feature dimension combinations to train the first candidate diagnosis model, and 10 trained candidate diagnosis models corresponding to the 10 candidate feature dimension combinations are obtained.
[0197] As an example, when there are 20 candidate value combinations of the hyperparameters of the first candidate diagnosis model, for the 20 candidate value combinations of the hyperparameters of the first candidate diagnosis model, using the training set, the features of the training set are respectively extracted with 10 candidate feature dimension combinations for each candidate value combination, and the first candidate diagnosis model is respectively trained to obtain 10 trained first candidate diagnosis models corresponding to the 10 candidate feature dimension combinations, and a total of 200 trained first candidate diagnosis models are obtained. Among them, the feature dimensions included in the candidate feature dimension combinations corresponding to the 20 candidate value combinations may be the same or different.
[0198] Specifically, the training process is sequentially executed for each value combination of the hyperparameters of each candidate diagnosis model, and finally, multiple trained first candidate diagnosis models corresponding to various value combinations of the hyperparameters and various candidate feature dimension combinations are obtained.
[0199] In a possible implementation, step 2045 can be implemented through the following process:
[0200] Rank the feature dimensions corresponding to each feature according to the importance of the first candidate diagnosis model using the first candidate value combination, and obtain the ranking result of each feature dimension; based on the ranking result, determine the first candidate feature dimension combination; wherein, the first candidate feature dimension combination includes two feature dimensions, and the importance of each feature dimension in the first candidate feature dimension combination is higher than that of other feature dimensions; the other feature dimensions include any feature dimension in the several feature dimensions except the two feature dimensions in the first candidate feature dimension combination; according to the order of importance from high to low in the ranking result, add one feature dimension to the first candidate feature dimension combination in turn to obtain at least two candidate feature dimension combinations corresponding to the first candidate diagnosis model using the first candidate value combination.
[0201] Among them, during the training process of the first candidate diagnosis model using the first candidate value combination, according to the importance corresponding to each feature dimension, the feature dimensions are sorted in descending order, first select the two feature dimensions with the highest ranking as the first candidate feature dimension combination, and then according to the importance of each feature dimension, select the one with the highest importance from the remaining unadded feature dimensions and add it to the current candidate feature dimension combination to obtain a new candidate feature dimension combination.
[0202] Among them, if there are N feature dimensions, N - 1 candidate feature dimension combinations are obtained.
[0203] Among them, for any two adjacent candidate feature combinations, the feature dimension added to the latter candidate feature combination compared with the former candidate feature combination has an importance lower than that of each feature dimension in the former candidate feature combination.
[0204] As an example, there are 20 feature dimensions. The 20 feature dimensions are sorted in descending order of importance as feature dimension 1, feature dimension 2, …, feature dimension 19, and feature dimension 20. First, it is determined that the candidate feature dimension combination includes {feature dimension 1, feature dimension 2}. Then, the feature dimension 3 with the highest importance among the remaining 18 feature dimensions is added, and the obtained candidate feature dimension combination includes {feature dimension 1, feature dimension 2, feature dimension 3}. Then, the feature dimension 4 with the highest importance among the remaining 17 feature dimensions is added, and the obtained candidate feature dimension combination includes {feature dimension 1, feature dimension 2, feature dimension 3, feature dimension 4}. And so on, the following candidate feature dimension combinations are obtained: {feature dimension 1, feature dimension 2}, {feature dimension 1, feature dimension 2, feature dimension 3}, …, {feature dimension 1, feature dimension 2, …, feature dimension 19}, {feature dimension 1, feature dimension 2, …, feature dimension 19, feature dimension 20}.
[0205] In a possible implementation, this step 2045 can be implemented by the following process:
[0206] According to the importance of each feature for the first candidate value combination of the first candidate diagnosis model, the feature dimensions corresponding to each feature are sorted to determine a second candidate feature dimension combination, and the second candidate feature dimension combination includes the several feature dimensions; according to the order of importance from low to high in the sorting result, one feature dimension is sequentially removed from the second candidate feature dimension combination from the several feature dimensions to obtain at least two candidate feature dimension combinations corresponding to the first candidate diagnosis model using the first candidate value combination.
[0207] Among them, during the training process of the first candidate diagnosis model using the first candidate value combination, first select the whole of the several feature dimensions to form a second candidate feature dimension combination, and then remove one feature dimension with the lowest importance in the second candidate feature dimension combination to obtain a new candidate feature dimension combination. According to the corresponding importance of each feature dimension, one is sequentially removed from the new candidate feature dimension combination in ascending order of importance to obtain multiple candidate feature dimension combinations.
[0208] Among them, for any two adjacent candidate feature combinations, the latter candidate feature combination has one less feature dimension than the former candidate feature combination, and the importance of the reduced feature dimension is higher than the importance of each feature dimension in the latter candidate feature combination.
[0209] Among them, if there are N feature dimensions, N - 1 candidate feature dimension combinations are obtained.
[0210] In a possible implementation, after sorting by importance, one feature dimension can be removed in ascending order one by one. Each time a feature dimension is removed, a candidate feature dimension combination is obtained, and multiple candidate feature dimension combinations are obtained.
[0211] As an example, there are 20 feature dimensions, which are sorted in ascending order of importance as feature dimension 20, feature dimension 19,..., feature dimension 2, feature dimension 1. First, determine the candidate feature dimension combination {feature dimension 20, feature dimension 19,..., feature dimension 2, feature dimension 1}. Remove the feature dimension 20 with the lowest importance from this candidate feature dimension combination to obtain the candidate feature dimension combination {feature dimension 19,..., feature dimension 2, feature dimension 1}. Then, remove the feature dimension 19 with the lowest importance from this candidate feature dimension combination to obtain the candidate feature dimension combination {feature dimension 18,..., feature dimension 2, feature dimension 1}, and so on, until the candidate feature dimension combination {feature dimension 2, feature dimension 1} with the highest importance is obtained, resulting in the following 19 candidate feature dimension combinations: {feature dimension 20, feature dimension 19,..., feature dimension 2, feature dimension 1}, {feature dimension 19, feature dimension 18,..., feature dimension 2, feature dimension 1}, {feature dimension 17, feature dimension 16,..., feature dimension 2, feature dimension 1},..., {feature dimension 3, feature dimension 2, feature dimension 1}, {feature dimension 2, feature dimension 1}.
[0212] In a possible implementation, the multiple candidate feature dimension combinations can be determined in advance. During the training of the first candidate diagnostic model, after each training is completed using a candidate feature dimension combination, the candidate feature dimension combination used last time can be changed in the above-mentioned manner to obtain a new candidate feature dimension combination, and the next training can be carried out. Among them, the way to change the feature dimensions included in the candidate feature dimension combination can be to remove a feature dimension with the lowest importance from the candidate feature dimension combination, or to select a feature dimension with the highest importance from the unadded feature dimensions and add it to the candidate feature dimension combination.
[0213] Figure 5 It is a schematic diagram of the sorting of the importance scores corresponding to the importance of the features of each feature dimension of the acoustic signal provided by the embodiments of the present application. In this schematic diagram, the importance of the features of each feature dimension is scored respectively and sorted in descending order of the scores. Among them, the higher the importance of the feature dimension, the greater its importance score.
[0214] According to this Figure 5As can be seen from the records, the effective value of the time-domain energy obtained the highest importance score. This is consistent with the actual physical phenomenon: when corrugation occurs on the rail, the vibration intensity during train travel increases, which in turn leads to a significant increase in the energy in the acoustic signal. In addition, the A-weighted sound pressure levels in the three frequency bands of 400 Hz, 500 Hz, and 315 Hz also showed relatively high importance scores. These frequency bands are usually related to the corrugation phenomenon and are its common frequency response ranges.
[0215] It should be noted that the above process of determining multiple candidate feature dimension combinations is performed for each candidate value combination of each first candidate diagnosis model, and the subsequent training and the process of determining the test performance are also performed.
[0216] 2046. Take the at least one candidate feature dimension combination as the at least one candidate feature dimension combination corresponding to the first candidate diagnosis model that adopts the first candidate value combination.
[0217] Among them, when training the first candidate diagnosis model that adopts the first candidate value combination, use the various candidate feature dimension combinations determined in the foregoing steps for training.
[0218] Among them, it can be to use any one of the candidate feature dimension combinations to extract features from the training samples in the training set, so as to use the features of each feature dimension obtained by extraction to train the corresponding first candidate diagnosis model that adopts the first candidate value combination. Moreover, also use the candidate feature dimension combination to extract features from the test samples in the test set, so as to use the features of each feature dimension obtained by extraction to test the performance of the corresponding first candidate diagnosis model that adopts the first candidate value combination. Correspondingly, for the first candidate diagnosis model that adopts the first candidate value combination, for each of its candidate feature dimension combinations, perform the above training and testing processes, and the subsequent steps will elaborate on this process.
[0219] 205. Use the training set to train the first candidate diagnosis model respectively according to each candidate feature dimension combination to obtain the corresponding trained first candidate diagnosis model, and each trained first candidate diagnosis model corresponds to a candidate feature dimension combination.
[0220] In a possible implementation, the first candidate diagnosis model does not include hyperparameters. Select a candidate feature dimension combination, use the candidate feature dimension combination to extract features from the training samples in the training set, input the obtained training sample feature set into the first candidate diagnosis model, the first candidate diagnosis model outputs a training diagnosis result, determine the training loss of this training according to the training diagnosis result and the training sample label, and use the training loss to adjust the parameter values in the first candidate diagnosis model. For the adjusted first candidate diagnosis model, use the training set for training again, and the training process is as described above until the training end condition is met.
[0221] In another possible implementation, the first candidate diagnostic model includes hyperparameters. Using the training set, according to the multiple candidate feature dimension combinations determined in the aforementioned step 204, the first candidate diagnostic model adopting the first candidate value combination is trained to obtain the trained first candidate diagnostic model. Each trained first candidate diagnostic model corresponds to a candidate feature dimension combination and adopts the first candidate value combination.
[0222] In a possible implementation, a candidate feature dimension combination is selected. Using this candidate feature dimension combination, feature extraction is performed on the training samples in the training set, and the obtained training sample feature set is input into the first candidate diagnostic model adopting the first candidate value combination. The first candidate diagnostic model adopting the first candidate value combination outputs a training diagnosis result. According to this training diagnosis result and the training sample labels, the training loss of this training is determined, and the parameter values in the first candidate diagnostic model adopting the first candidate value combination are adjusted using this training loss. For the adjusted first candidate diagnostic model adopting the first candidate value combination, the training set is used for training again, and the training process is as described above until the training end condition is met.
[0223] Among them, the training end condition can be that the training loss no longer converges, or the number of training times reaches the set number, etc. This application does not limit the training end condition.
[0224] Correspondingly, for the hyperparameters of the first candidate diagnostic model, each candidate value combination is used to respectively execute the above training process to obtain each trained first candidate diagnostic model. Each trained first candidate diagnostic model corresponds to a candidate value combination and a candidate feature dimension combination.
[0225] Similarly, for each candidate diagnostic model among multiple candidate diagnostic models, the above process is respectively executed to obtain the trained first candidate diagnostic model corresponding to each candidate diagnostic model. Each trained first candidate diagnostic model corresponds to a first candidate diagnostic model, a candidate value combination, and a candidate feature dimension combination.
[0226] 206. Using the test set, determine the test performance of the trained first candidate diagnostic model;
[0227] Among them, using the test set, the performance of each trained first candidate diagnostic model is tested to determine the prediction performance of the trained first candidate diagnostic model. The better the test performance, the more suitable the trained first candidate diagnostic model is for prediction in the field to which the test set belongs.
[0228] Among them, the test set includes several test samples to use these test samples to perform performance testing on the trained first candidate diagnostic model.
[0229] In a possible implementation, the performance of the first candidate diagnostic model after training can be tested in combination with test samples, as described in detail below Figure 6 in the following.
[0230] Figure 6 FIG. is a schematic flowchart of determining the test performance of the first candidate diagnostic model after training by using a test set provided by an embodiment of the present application, which may include steps 2061 to 2062, and the following describes these steps in detail respectively.
[0231] 2061. Input the test samples in the test set into each of the first candidate diagnostic models after training to obtain the prediction results of each of the first candidate diagnostic models after training. The test set includes a number of test sample pairs, and each test sample pair includes a test sample and a test label, and the test label is used to characterize whether the rail corresponding to the test sample has a rail corrugation fault;
[0232] Among them, the performance of each of the first candidate diagnostic models after training is tested by using the test set to obtain the test performance of each of the first candidate diagnostic models after training.
[0233] In a possible implementation, after determining the first candidate diagnostic model, a candidate value combination can be selected for the first candidate diagnostic model, and for the first candidate diagnostic model using the candidate value combination, training is respectively performed by using each candidate feature dimension combination. Correspondingly, when performing testing, the candidate feature dimension combination is used to extract the features of the test samples, and the extracted test sample features are input into the first candidate diagnostic model after training to obtain the prediction results of the first candidate diagnostic model after training.
[0234] As an example, for the XGBoost model, after determining a group of candidate value combinations, training is performed by using the candidate value combination and a certain determined candidate feature dimension combination to obtain a candidate XGBoost model after training. The test set is used for feature extraction by using the corresponding candidate feature dimension combination, and the features of multiple feature dimensions extracted are input into the candidate XGBoost model after training to obtain the prediction results.
[0235] Referring to the above example, after training is respectively performed for each candidate feature dimension combination and each candidate feature dimension combination of the XGBoost model, prediction is respectively performed by using the test set to obtain the prediction results.
[0236] 2062. Based on the prediction results and the corresponding test labels, determine the test performance of each of the first candidate diagnostic models after training.
[0237] Among them, quantitative calculations are performed on each prediction result and the corresponding test label to determine the test performance of each of the first candidate diagnostic models after training.
[0238] Among them, the test performance can be quantified as performance metrics, such as accuracy, F1 value, etc.
[0239] Among them, the F1 value is the harmonic mean of precision and recall. The precision represents the proportion of samples that are actually positive among the samples predicted as positive (with rail corrugation faults) by the model; the recall represents the proportion of samples that are actually positive among the samples correctly predicted as positive by the model.
[0240] In a possible implementation, the data set is divided into a training set and a test set. Therefore, in practical applications, a cross-validation combined with grid search can be used to train the diagnostic model until the training is completed, and the test performance of the candidate diagnostic models corresponding to each value combination and each candidate feature dimension combination is obtained.
[0241] In a possible implementation, training can be performed by using the training set and test set for verification. Specifically, the data set is divided into N mutually exclusive subsets. For a combination of candidate feature dimensions and a combination of hyperparameter values, one of the subsets is used as the test set in turn, and the other N - 1 subsets are used as the training set for N - iteration training and testing to determine the performance metrics of each test. The average value of the performance metrics of the multiple validations is taken to obtain the performance metrics of the combination of candidate feature dimensions and the combination of hyperparameter values used this time for training the first diagnostic model.
[0242] 207. Based on the test performance of the first candidate diagnostic model after training, among the first candidate diagnostic models after training corresponding to each candidate diagnostic model, select the target diagnostic model, and determine the candidate feature dimension combination corresponding to the target diagnostic model as the target feature dimension combination; the test performance of the target diagnostic model is better than that of the non-target diagnostic models; the non-target diagnostic models are the other first candidate diagnostic models after training among the first candidate diagnostic models after training corresponding to each candidate diagnostic model, excluding the target diagnostic model.
[0243] Among them, the target diagnostic model is one of the first candidate diagnostic models after training. The test performance of the target diagnostic model is better than that of the other first candidate diagnostic models after training. The other first candidate diagnostic models after training include the first candidate diagnostic models after training using the same algorithm as the target diagnostic model and also include the first candidate diagnostic models after training using different algorithms from the target diagnostic model.
[0244] In a possible implementation, for a first candidate diagnostic model, first, among the first candidate diagnostic models obtained by successively determining each candidate feature dimension combination corresponding to a value combination of hyperparameters and training them, the one with the optimal test performance is taken as the candidate diagnostic model corresponding to the optimal test performance of this value combination; then, the candidate diagnostic models corresponding to the optimal test performance of each value combination of hyperparameters are determined, and the first candidate diagnostic model with the optimal test performance among them is selected; finally, among the first candidate diagnostic models corresponding to the optimal test performance of each candidate diagnostic model, the one with the optimal test performance is selected as the target diagnostic model.
[0245] As an example, the candidate diagnostic models include an SVM diagnostic model and an XGBoost diagnostic model. Among them, the candidate value combinations of the SVM diagnostic model include 4 candidate value combinations A - D, and there are 5 candidate feature dimension combinations a - e in total; the candidate value combinations of the XGBoost diagnostic model include 5 candidate value combinations E - I in total, and there are 4 candidate feature dimension combinations f - i in total. Using each candidate value combination and candidate feature dimension combination, each candidate diagnostic model is trained to obtain 40 trained candidate diagnostic models. Among them, for the SVM diagnostic model using candidate value combination A and candidate feature dimension combination a - e for training, among the 5 trained candidate SVM diagnostic models, the candidate SVM diagnostic model with the optimal test performance is determined, which is the trained candidate SVM diagnostic model using candidate feature dimension combination b and candidate value combination A; similarly, the above selection is performed for candidate value combinations B - D respectively, and the optimal trained candidate SVM diagnostic models corresponding to each candidate value combination of the SVM diagnostic model (4 in total) are obtained respectively. Among the 5 optimal trained candidate SVM diagnostic models, the one with the optimal test performance is determined as the optimal trained candidate SVM diagnostic model of this SVM diagnostic model, which is the trained candidate SVM diagnostic model using candidate value combination C and candidate feature dimension combination e. Similarly, the optimal trained candidate XGBoost diagnostic model corresponding to the XGBoost diagnostic model is determined. Then, among this optimal trained candidate SVM diagnostic model and the optimal trained candidate XGBoost diagnostic model, the diagnostic model with the optimal test performance is selected as the finally selected target diagnostic model. For example, the finally selected target diagnostic model is the XGBoost diagnostic model using candidate value combination F and candidate feature dimension combination g.
[0246] Among them, the target diagnostic model is the first candidate diagnostic model with the optimal test performance after training. Correspondingly, the candidate feature dimension combination corresponding to this target diagnostic model is selected as the optimal candidate feature dimension combination.
[0247] In a possible implementation, each trained first candidate diagnosis model includes a candidate value combination. Correspondingly, based on the candidate value combination corresponding to each trained first candidate diagnosis model, the candidate value combination corresponding to the target diagnosis model is determined as the target value combination; based on the target value combination, the hyperparameters of the target diagnosis model are determined.
[0248] Among them, if the first candidate diagnosis model has hyperparameters, correspondingly, each trained first candidate diagnosis model includes a candidate value combination. After determining the target diagnosis model, the candidate value combination corresponding to the target diagnosis model is used as the target value combination, and the parameter values in the target value combination are used as the values of the hyperparameters in the target diagnosis model.
[0249] Among them, during the process of training the target diagnosis model, the corresponding candidate feature dimension combination and the value combination of the hyperparameters are used. Correspondingly, the test performance of the target diagnosis model is better than that of other first candidate diagnosis models, and it is also affected by its corresponding candidate feature dimension combination and the value combination of the hyperparameters. Therefore, the candidate feature dimension combination is selected as the target feature dimension combination, and the value combination is selected as the target value combination, and the situation of predicting the rail corrugation fault corresponding to the acoustic signal to be processed is carried out during the subsequent actual operation process.
[0250] As an example, there are 2 candidate diagnosis models, namely diagnosis model 1 and diagnosis model 2. The candidate value combination of the hyperparameters of diagnosis model 1 has 4, and the candidate feature dimension combination has 3. The candidate value combination of the hyperparameters of diagnosis model 2 has 2, and the candidate feature dimension combination has 3. The diagnosis models are trained respectively using the above-mentioned candidate value combination of the hyperparameters and the candidate feature dimension combination. It is determined that the trained diagnosis model 1 (obtained by training with the value combination 1 of the hyperparameters and the candidate feature dimension combination 2) has better test performance than the diagnosis model 1 obtained by training with other value combinations and candidate feature dimension combinations, and is also better than the diagnosis model 2 obtained by training with 2 candidate value combinations and 3 candidate feature dimension combinations. Correspondingly, the trained diagnosis model 1 is determined as the target diagnosis model, the candidate value combination 1 of the hyperparameters is used as the target value combination, and the candidate feature dimension combination 2 is used as the target feature dimension combination.
[0251] In a possible implementation, first select a candidate value combination to be fixed, and traverse the candidate feature dimension combinations. The test performance can be executed during the traversal process. When the test performance no longer improves significantly, determine the candidate feature dimension combination with the optimal test performance as the target feature dimension combination corresponding to this candidate value combination. In this process, the training process may end without training all the candidate feature dimension combinations, reducing the amount of data processing during the training process. Execute the above process for each candidate value combination of the current first candidate diagnostic model in turn to determine the best feature dimension combination and the best candidate value combination of the first candidate diagnostic model. According to the above process, train each candidate diagnostic model to determine the best feature dimension combination and the best candidate value combination of each candidate diagnostic model. Then, based on the test performance of each trained first candidate diagnostic model, select the trained candidate diagnostic model with the best test performance as the target diagnostic model, its corresponding best feature dimension combination as the target feature dimension combination, and its corresponding best value combination as the target value combination.
[0252] Among them, the determined target diagnostic model can be integrated into the actual application scenario. Specifically, it can be deployed through ways such as API (Application Program Interface), edge computing devices, or cloud platforms, etc. The deployment method is not limited in this application.
[0253] As shown in Table 2 below, it is a comparison list of various diagnostic models (diagnostic models using different algorithms). Table 2 shows the comparison of the accuracy rate and time of each diagnostic model using features of multiple feature dimensions and using traditional A-weighted sound pressure level prediction.
[0254] Table 2
[0255]
[0256] From the above Table 2, it can be concluded that if the diagnostic models use the same algorithm, using the features of multiple feature dimensions in the solution of this application, compared with the traditional A-weighted sound pressure level, the accuracy rate is higher; the prediction speed is slightly slower, but the lag time is at the millisecond level or even lower level, which is within the range that can be tolerated by the real-time requirement.
[0257] In this embodiment, a training set and a test set are obtained. The training set contains a number of training sample pairs, and each training sample pair includes a training sample and a training label. The training label is used to represent whether the rail corresponding to the training sample has a rail corrugation fault. Feature extraction is performed on the training samples in the training set based on a number of feature dimensions to obtain the respective features corresponding to the training samples, and each feature corresponds to a feature dimension. One is sequentially selected from at least two candidate diagnostic models as the first candidate diagnostic model. Based on the respective features, a plurality of candidate feature dimension combinations are determined for the first candidate diagnostic model. Among them, each candidate feature dimension combination includes at least two candidate feature dimensions, and each candidate feature dimension is determined from a number of feature dimensions according to the importance of each feature dimension. The training set is used to train the first candidate diagnostic model respectively according to the respective candidate feature dimension combinations to obtain the corresponding first candidate diagnostic model after training, and each first candidate diagnostic model after training corresponds to a candidate feature dimension combination. The test set is used to determine the test performance of the first candidate diagnostic model after training. Based on the test performance of the first candidate diagnostic model after training, among the first candidate diagnostic models after training corresponding to the respective candidate diagnostic models, a target diagnostic model is selected, and the candidate feature dimension combination corresponding to the target diagnostic model is determined as the target feature dimension combination. The test performance of the target diagnostic model is better than that of the non-target diagnostic models. The non-target diagnostic models are the other first candidate diagnostic models after training among the first candidate diagnostic models after training corresponding to the respective candidate diagnostic models, excluding the target diagnostic model. In this process, it is realized to traverse each value combination of hyperparameters and candidate feature dimension combinations, train the first candidate diagnostic model, obtain a number of first candidate diagnostic models after training, determine the one with the optimal test performance as the target diagnostic model, and determine the value combination and candidate feature dimension combination corresponding to the target diagnostic model as the finally selected hyperparameter values and feature dimensions, realizing the optimal configuration of the hyperparameters of the diagnostic model by grid search and selecting the optimal input feature dimension of the diagnostic model, and improving the accuracy of the diagnostic model prediction.
[0258] The above introduces a rail corrugation fault diagnosis method provided by an embodiment of the present application. The following will introduce the training process of the target diagnostic model used in the above rail corrugation fault diagnosis method.
[0259] Figure 7 It is a schematic flowchart of the training method of the diagnostic model provided by an embodiment of the present application, including steps 701 to 706, and the following will describe these steps in detail respectively.
[0260] 701. Obtain a training set and a test set. The training set contains a number of training sample pairs, and each training sample pair includes a training sample and a training label. The training label is used to represent whether the rail corresponding to the training sample has a rail corrugation fault.
[0261] Among them, in the training method of the diagnostic model provided in this embodiment, there are multiple candidate diagnostic models. During the training process, a target diagnostic model is determined from the multiple candidate diagnostic models, and moreover, during the training process, the hyperparameters and the input feature dimensions adopted by the target diagnostic model are also determined.
[0262] Among them, the sample pairs in the preset dataset can be divided into a test set and a dataset to obtain the test set and the training set required during the training process.
[0263] In a specific implementation, the original sample acoustic signals can be intercepted to obtain multiple sample acoustic signals, and labels are added to each sample acoustic signal to obtain this dataset.
[0264] Among them, the features of multiple feature dimensions required by each candidate diagnostic model during the application process are determined in advance.
[0265] Specifically, when determining the feature dimensions in the target feature set, it is determined using the training set. First, the training set is obtained, and then the candidate diagnostic models are trained and performance-tested using the training set and the test set to select the target diagnostic model.
[0266] Among them, the obtained training set contains several training sample pairs. The training sample pair contains a training sample and a training label. The training sample is an acoustic signal as a sample, and the training sample contains positive samples and negative samples. The label of the positive sample is that a rail corrugation fault occurs, and the label of the negative sample is that no rail corrugation fault occurs.
[0267] Among them, the obtained test set contains several test sample pairs. The test sample pair contains a test sample and a test label. The test sample is an acoustic signal as a sample, which is used to be input into the candidate diagnostic model trained by the training set to test the performance of the trained candidate diagnostic model.
[0268] In a possible implementation, the obtained test set can be stored, and during the process of testing the performance of the trained candidate diagnostic model when needed, the test set is obtained from the storage location for performance testing.
[0269] In a possible implementation, the training set and the test set can be acoustic signals obtained from the real scene of the train's travel. The acoustic signals are used as samples, and labels are added to them.
[0270] Among them, for the specific process of obtaining the training set and the test, reference can be made to the Figure 3 flow schematic diagram and the corresponding explanations above, which will not be elaborated here.
[0271] 702. Feature extraction is performed on the training samples in the training set based on several feature dimensions to obtain the respective features corresponding to the training samples, and each feature corresponds to a feature dimension.
[0272] Among them, the several feature dimensions can represent the acoustic characteristics of the acoustic signal from different feature dimensions, but different feature dimensions have different acoustic characterization situations for the acoustic signal, which can be specifically reflected in the determination of the marginal contribution to the corresponding candidate diagnostic model.
[0273] Among them, the several feature dimensions are all possible feature dimensions that can be used for feature extraction of the acoustic signal.
[0274] For example, the several feature dimensions include all the following feature dimensions: A-weighted sound pressure level feature, root mean square value of energy, peak factor, impulse factor, margin factor, kurtosis factor, skewness factor, zero crossing rate, average spectral information entropy, average spectral centroid, average spectral kurtosis, Mel frequency cepstral coefficients. Of course, in specific implementation, other feature dimensions can be added according to actual situations, which are not limited in this application.
[0275] Specifically, for each training sample in the training set, feature extraction is respectively performed from the several feature dimensions to obtain the features of each feature dimension in the several feature dimensions of each training sample, providing a basis for the input content required for subsequent training.
[0276] In this embodiment, some feature dimensions are determined from the several feature dimensions as the feature dimensions in the target feature set, and the process of determining the feature dimensions is executed synchronously with the process of training and selecting the target diagnostic model.
[0277] Among them, when there are hyperparameters in the candidate diagnostic model, the selection of hyperparameters is also executed synchronously with the process of training and selecting the target diagnostic model.
[0278] 703. Select one from at least two candidate diagnostic models in sequence as the first candidate diagnostic model.
[0279] Among them, multiple candidate diagnostic models are obtained. Different diagnostic models can adopt different algorithms. The diagnostic model can be a support vector machine, decision tree, K-nearest neighbor, Gaussian naive Bayes, random forest, extreme gradient boosting, etc. In this embodiment, the algorithm model most suitable for diagnosing rail corrugation faults is determined among these multiple algorithms.
[0280] Among them, for each candidate diagnostic model in advance, hyperparameters and the value ranges of hyperparameters are set. The hyperparameters and their value ranges of any two candidate diagnostic models may be the same or different, and the hyperparameters and their value ranges of each candidate diagnostic model are not limited in this application.
[0281] Among them, one of the at least two candidate diagnostic models can be selected as the first candidate diagnostic model to perform the training and prediction processes for the first candidate diagnostic model. Each candidate diagnostic model among the at least two candidate diagnostic models is respectively selected as the first candidate diagnostic model.
[0282] Among them, for the first candidate diagnostic model, a value combination of its hyperparameters is selected as the candidate value combination, and during the selection process, it is determined by traversing the hyperparameter combinations.
[0283] In a possible implementation, a value step size can be set for each hyperparameter, and multiple values are selected within the value range at this step size to obtain the selectable values of the hyperparameter. The selectable values of multiple hyperparameters are combined to obtain multiple value combinations of the hyperparameters.
[0284] In a possible implementation, for the first candidate diagnostic model, using its respective hyperparameters and the corresponding value ranges, a hyperparameter grid is determined. Each grid corresponds to a candidate value combination of the hyperparameters. During the process of training the first candidate diagnostic model, the grid search method is adopted, and by traversing the value combination method of each hyperparameter in the hyperparameter grid, the optimal combination of the hyperparameters of each first candidate diagnostic model can be determined.
[0285] As an example, the first candidate diagnostic model adopts the XGBoost model, and the hyperparameters of the first candidate diagnostic model include: the number of iterations, the learning rate, the maximum tree depth, the tree feature sampling ratio, the sample sampling ratio, and the minimum sample weight of the leaf nodes.
[0286] 704. According to each of these features, multiple candidate feature dimension combinations are determined for the first candidate diagnostic model; among them, each candidate feature dimension combination contains at least two candidate feature dimensions.
[0287] Among them, for the first candidate diagnostic model, multiple candidate feature dimension combinations are determined. The candidate feature dimensions included in each candidate feature dimension combination are determined according to their importance degrees, and the importance degree of each candidate feature dimension is determined according to the marginal contribution of this feature dimension to the first candidate diagnostic model.
[0288] Among them, when the first candidate diagnostic model involves hyperparameters, the value of the hyperparameters can be set for the first candidate diagnostic model in advance. Each value of the hyperparameters is substituted into the first candidate diagnostic model to obtain the first candidate diagnostic model using this value combination, so as to determine the corresponding multiple candidate feature dimension combinations for the first candidate diagnostic model of each value combination.
[0289] Among them, when the first candidate diagnostic model does not involve hyperparameters, the value of the hyperparameters does not need to be considered, and multiple candidate feature dimension combinations are determined for the first candidate diagnostic model.
[0290] In a possible implementation, multiple candidate feature dimension combinations of the first candidate diagnostic model can be determined in combination with the hyperparameters in the first candidate diagnostic model.
[0291] In a possible implementation, the first candidate diagnostic model includes hyperparameters. Embodiments of the present application provide a process for determining multiple candidate feature dimension combinations for the first candidate diagnostic model based on each feature, which may include steps 7041 to 7045. These steps will be described in detail below.
[0292] 7041. Obtain at least two candidate value combinations of the hyperparameters in the first candidate diagnostic model;
[0293] Among them, the hyperparameter is a parameter set by a person and cannot be automatically learned from data.
[0294] During the training process, the trainer pre-sets the value range of the hyperparameters for each candidate diagnostic model.
[0295] Among them, the hyperparameters of each diagnostic model in at least two candidate diagnostic models can be the same or different. In the present application, there is no limitation on the hyperparameters of each candidate diagnostic model and their value ranges.
[0296] In a possible implementation, the value range of the hyperparameters can be set for the at least two candidate diagnostic models first, then one of the at least two candidate diagnostic models is selected as the first candidate diagnostic model, and then multiple candidate value combinations are selected within the value range of the hyperparameters of the first candidate diagnostic model.
[0297] In a possible implementation, one of the at least two candidate diagnostic models can be selected as the first candidate diagnostic model first, then the value range of the hyperparameters of the first candidate diagnostic model is set, and then multiple candidate value combinations are selected from within the value range.
[0298] As an example, for the first candidate diagnostic model using the XGBoost algorithm, the set hyperparameters may include, but are not limited to, tree-related parameters, regularization parameters, learning rate, etc.
[0299] Among them, for the first candidate diagnostic model, using its respective hyperparameters and the corresponding value ranges, a hyperparameter grid is determined. Each grid corresponds to a candidate value combination of the hyperparameters. During the process of training the first diagnostic model, the grid search method is used to traverse the values of each hyperparameter to determine multiple candidate value combinations.
[0300] In a possible implementation, for each candidate diagnostic model serving as a first candidate diagnostic model, a value range of hyperparameters can be preset, a selection step size for each hyperparameter can be set, and using this selection step size, a hyperparameter grid is determined. Each grid corresponds to a combination of candidate values of the hyperparameters. During the process of training the first candidate diagnostic model, a grid search method is adopted to obtain multiple available values of the hyperparameters. Similarly, for each hyperparameter in the first candidate diagnostic model, its respective available values are determined.
[0301] For example, grid search is used to optimize six key hyperparameters of the XGBoost model: the number of iterations, the learning rate, the maximum tree depth, the tree feature sampling ratio, the sample sampling ratio, and the minimum sample weight of leaf nodes.
[0302] 7042. Sequentially determine one of the at least two combinations of candidate values as the first combination of candidate values;
[0303] Among them, during the process of training the first candidate diagnostic model, a way of fixing one combination and traversing the other combination can be adopted to combine the candidate feature dimension combination and the combination of candidate values. Using one candidate feature dimension combination and one combination of candidate values obtained from the combination, the first candidate diagnostic model is trained.
[0304] Among them, for each first candidate diagnostic model, a combination of candidate values of its hyperparameters is selected. During the selection process, it is determined by traversing the hyperparameter combinations.
[0305] Among them, one of the multiple combinations of candidate values of the first candidate diagnostic model is determined as the first combination of candidate values.
[0306] Among them, the naming of the first combination of candidate values is only for the convenience of explaining the process subsequently. In specific implementation, the following processing process is performed for each combination of candidate values in each first candidate diagnostic model.
[0307] As an example, for the scenario where the main hyperparameters of the first candidate diagnostic model using the XGBoost algorithm include: the number of iterations, the learning rate, and the maximum tree depth, the three hyperparameters are combined according to the optional values to obtain multiple combinations of candidate values. For each combination of candidate values, the subsequent processing steps are respectively executed to determine the importance of each feature dimension of the features relative to the candidate diagnostic model for each combination of candidate values.
[0308] Among them, during the subsequent training process, for each combination of candidate values of the hyperparameters, the corresponding candidate feature dimension combination is used to train the corresponding first candidate diagnostic model using each combination of candidate values and the corresponding multiple candidate feature dimension combinations.
[0309] In a possible implementation, when it is specified that a candidate diagnostic model adopts a specific algorithm and there is only one candidate diagnostic model, then, candidate value combinations are respectively selected for the candidate diagnostic model, and multiple candidate feature dimensions are used for training for each candidate value combination to determine the trained target diagnostic model.
[0310] 7043. Determine the marginal contribution of each feature dimension to the output of the first candidate diagnostic model adopting the first candidate value combination according to each feature;
[0311] Among them, the marginal contribution of the feature of a certain feature dimension to the output of any candidate diagnostic model is the change in the output of the candidate diagnostic model caused by adding the feature of this feature dimension and not adding the feature of this feature dimension to the input of the candidate diagnostic model. The greater the marginal contribution of the feature of a certain feature dimension to the output of the candidate diagnostic model, the higher the importance of this feature dimension.
[0312] Specifically, when there are N feature dimensions, all the features of the N feature dimensions are used as the first feature combination. The first feature combination includes feature A, feature B, feature C, feature D....., feature N; based on the first feature combination, feature A in the first feature combination is removed to obtain the second feature combination; then based on the first feature combination, feature B (this feature B is different from feature A) in the first feature combination is removed to obtain the third feature combination, and so on. One feature is removed from the first feature combination in turn to obtain N feature combinations; where N is a positive integer; the first feature combination, the second feature combination, the third feature combination....., the Nth feature combination are respectively input into the first candidate diagnostic model adopting the first candidate value combination to obtain the output results corresponding to each feature combination, and the output result corresponding to the first feature combination is compared with the output results corresponding to each remaining feature combination to determine the output influence of each removed feature dimension on the output of the first candidate diagnostic model adopting the first candidate value combination, and this output influence is the marginal contribution of the removed feature dimension to the output of the first candidate diagnostic model adopting the first candidate value combination.
[0313] As an example, the first candidate diagnostic model is an XGBoost model, and the first candidate value combination selected is {50 (number of iterations), 0.2 (learning rate), 5 (maximum depth of the tree), 0.6 (sampling ratio of tree features), 0.7 (sampling ratio of samples per tree), 3 (minimum sum of sample weights of leaf nodes)}. Taking this first candidate value combination as the hyperparameter values of the XGBoost model, an XGBoost model with known hyperparameters is obtained. Based on the aforementioned first feature combination, second feature combination, ……, Nth feature combination, inputting them into the XGBoost model with known hyperparameters to obtain the output results. By comparing the output results corresponding to the second feature combination with the output results corresponding to the first feature combination, the influence of feature A on the output of the XGBoost model with known hyperparameters is determined. This output influence is the marginal contribution of the removed feature dimension A to the output of the XGBoost model with known hyperparameters. Based on this, the marginal contributions of each feature dimension to the output of the XGBoost model with known hyperparameters are determined.
[0314] In a possible implementation, the Shapley Additive Explanation value (SHAP value) can be used to represent this marginal contribution. For the features of each feature dimension, calculate the contribution degree of the features of this feature dimension to the prediction results of the corresponding candidate diagnostic model, that is, calculate the SHAP value, to determine the marginal contributions corresponding to the features of each feature dimension.
[0315] 7044. Determine the importance of each feature to the first candidate diagnostic model that adopts the first candidate value combination according to the marginal contributions of each feature dimension to the output of the first candidate diagnostic model that adopts the first candidate value combination;
[0316] Among them, according to the marginal contribution of each feature dimension to the first candidate diagnostic model that adopts the first candidate value combination, the importance of this feature dimension can be determined, and this importance is the marginal contribution of this feature dimension to the first candidate diagnostic model that adopts the first candidate value combination.
[0317] Among them, for the same first candidate diagnostic model, adopt each candidate value combination and execute the above process respectively to determine the importance of the feature dimensions of the features corresponding to the first candidate diagnostic model of each candidate value combination.
[0318] Similarly, for each candidate diagnostic model among at least two candidate diagnostic models, adopt each candidate value combination and execute the above process respectively to determine the importance of the feature dimensions of the features corresponding to the candidate diagnostic model of each candidate value combination.
[0319] Among them, after determining the marginal contribution of the features of each feature dimension to the output of the corresponding candidate diagnostic model, use this marginal contribution to determine the importance of the corresponding feature dimension.
[0320] Among them, the marginal contribution is positively correlated with the importance. The greater the marginal contribution, the greater the importance.
[0321] Among them, when it is stipulated that the candidate diagnostic model adopts a specific algorithm, there is only one such candidate diagnostic model. When determining different candidate value combinations for this candidate diagnostic model, the marginal contribution of the features of each feature dimension of the training sample to its output is considered. Accordingly, when obtaining different candidate value combinations for this candidate diagnostic model, the importance of each feature dimension is determined, and then at least two feature dimensions of the acoustic signal features included in the input information of this candidate diagnostic model are determined.
[0322] Among them, when there are multiple candidate diagnostic models, the marginal contribution of the features of each feature dimension of the training sample to each first candidate diagnostic model is determined respectively. Accordingly, what is obtained is the importance of each feature dimension corresponding to each first candidate diagnostic model.
[0323] Among them, when there are multiple candidate diagnostic models, for the same feature dimension, the marginal contribution to the outputs of different candidate diagnostic models may be different, and its importance to different candidate diagnostic models is also different.
[0324] 7045. According to the importance of each feature to the first candidate diagnostic model adopting the first candidate value combination, among the several feature dimensions corresponding to each feature, at least two feature dimensions are selected to form at least one candidate feature dimension combination.
[0325] Among them, according to the importance of each of these feature dimensions to the first candidate diagnostic model, two or more feature dimensions are selected. Subsequently, in an actual application scenario, when processing the acoustic signal to be processed, the at least two selected feature dimensions are used as the extraction feature dimensions for extracting the features of the acoustic signal to be processed, and a target feature set is obtained.
[0326] In a possible implementation, when dealing with multiple candidate feature dimension combinations corresponding to the first candidate diagnostic model adopting the first candidate value combination, the combinations can be made according to the importance of each feature by using a threshold, or the combinations can be made according to the importance ranking.
[0327] In a possible implementation, when there is only one candidate diagnostic model, this diagnostic model is used as the first candidate diagnostic model. According to the importance of each feature dimension, the importance of each feature dimension is sorted first, and at least two feature dimensions that meet the preset importance rule are selected as the target feature dimension combination.
[0328] Among them, the preset importance rule can be to select the feature dimensions with an importance greater than the set threshold, or to select the set number of feature dimensions ranked in the front from large to small in terms of importance, etc.
[0329] In this application, in combination with the training process of the first candidate diagnostic model, among the candidate feature dimension combinations corresponding to each candidate diagnostic model, a target feature dimension combination is selected, and this target feature dimension combination is the candidate feature dimension combination corresponding to the selected target diagnostic model.
[0330] In a possible implementation, this step 7045 can be implemented by the following process:
[0331] According to the importance of each feature for the first candidate diagnostic model using the first candidate value combination, among the several feature dimensions corresponding to each feature, at least two feature dimensions with importance greater than a preset importance threshold are selected; based on these at least two feature dimensions, the candidate feature dimension combination corresponding to the first candidate diagnostic model using the first candidate value combination is determined.
[0332] Among them, for the first candidate diagnostic model, an importance threshold can be preset in advance. When selecting the candidate feature dimension combination, some feature dimensions among these several feature dimensions are excluded, and the importance of the excluded feature dimensions is not greater than the preset importance threshold. Correspondingly, each feature dimension with importance greater than the preset importance threshold is retained, and the combination of these feature dimensions with importance greater than the preset importance threshold is used as the candidate feature dimension combination of the first candidate diagnostic model using the first candidate value combination.
[0333] As an example, there are 50 feature dimensions, the preset importance threshold is 0.5, and the value range of the importance of each feature dimension is (0, 1). Among the 2,000 feature dimensions corresponding to the first candidate diagnostic model using the first candidate value combination, the feature dimensions with importance greater than 0.5 are selected, and 30 feature dimensions with importance greater than 0.5 are obtained, and the importance of the remaining 20 feature dimensions is not greater than 0.5. These 30 feature dimensions are combined into a candidate feature dimension combination, which is used as the candidate feature dimension combination of the first candidate diagnostic model using the first candidate value combination.
[0334] As an example, when there are 10 feature dimension combinations and 20 candidate value combinations of the hyperparameters of the first candidate diagnostic model, for the 20 candidate value combinations of the hyperparameters in the first candidate diagnostic model, using the training set, each candidate value combination extracts the features of the training set with 10 candidate feature dimension combinations respectively, and the first candidate diagnostic model is trained respectively, and 10 trained first candidate diagnostic models corresponding to the 10 candidate feature dimension combinations are obtained. Based on the 20 candidate value combinations of the hyperparameters, 200 trained first candidate diagnostic models can be obtained. Among them, the feature dimensions included in the candidate feature dimension combinations corresponding to the 20 candidate value combinations may be different.
[0335] Specifically, for each combination of values of each hyperparameter of each first candidate diagnostic model, the training process is performed respectively, and finally, the trained first candidate diagnostic models corresponding to various combinations of values of hyperparameters and various combinations of candidate feature dimensions are obtained.
[0336] Among them, according to a set threshold and based on the importance of each candidate feature dimension, for the first candidate diagnostic model using the first candidate value combination, multiple feature dimensions with an importance greater than the set threshold are selected as the candidate feature dimension combination of the first candidate diagnostic model using the first candidate value combination.
[0337] Among them, it is also possible to determine multiple candidate feature dimension combinations of the first candidate diagnostic model using the first candidate value combination in the way of cyclically increasing or decreasing the feature dimensions.
[0338] In a possible implementation, in a possible implementation, step 7045 can be implemented by the following process:
[0339] Based on the importance of each feature to the first candidate diagnostic model using the first candidate value combination, the feature dimensions corresponding to each feature are sorted respectively to obtain the sorting result of each feature dimension; based on this sorting result, a first candidate feature dimension combination is determined; among them, the first candidate feature dimension combination includes two feature dimensions, and the importance of each feature dimension in the first candidate feature dimension combination is higher than that of other feature dimensions; the other feature dimensions include any feature dimension in the several feature dimensions except the two feature dimensions in the first candidate feature dimension combination; according to the order of importance from high to low in the sorting result, one feature dimension is sequentially added to the first candidate feature dimension combination to obtain at least two candidate feature dimension combinations corresponding to the first candidate diagnostic model using the first candidate value combination.
[0340] Among them, during the training process of the first candidate diagnostic model using the first candidate value combination, each feature dimension is sorted in descending order according to the corresponding importance. First, the two feature dimensions ranked in the front are selected as the first candidate feature dimension combination, and then, according to the importance of each feature dimension, the one with the highest importance is sequentially selected from the remaining unadded feature dimensions and added to the current candidate feature dimension combination to obtain a new candidate feature dimension combination.
[0341] Among them, if there are N feature dimensions, N - 1 candidate feature dimension combinations are obtained.
[0342] In a possible implementation, in a possible implementation, step 7045 can be implemented by the following process:
[0343] According to the importance of each feature for the first candidate diagnosis model using the first candidate value combination, sort the feature dimensions corresponding to each feature to determine the second candidate feature dimension combination, and the second candidate feature dimension combination includes the several feature dimensions; according to the order of importance from low to high in the sorting result, sequentially remove one feature dimension from the second candidate feature dimension combination from the several feature dimensions to obtain at least two candidate feature dimension combinations corresponding to the first candidate diagnosis model using the first candidate value combination.
[0344] Among them, during the training process of the first candidate diagnosis model using the first candidate value combination, first select the several feature dimensions as a whole to form the second candidate feature dimension combination, then remove the feature dimension with the lowest importance from the second candidate feature dimension combination to obtain a new candidate feature dimension combination, and then, in accordance with the corresponding importance of each feature dimension, sequentially remove one from the new candidate feature dimension combination in the order from low to high to obtain multiple candidate feature dimension combinations.
[0345] In a possible implementation, after sorting by importance, one feature dimension can be sequentially removed in the order from low to high, and each time a feature dimension is removed, a candidate feature dimension combination is obtained, resulting in multiple candidate feature dimension combinations.
[0346] 7046. Take the at least one candidate feature dimension combination as the at least one candidate feature dimension combination corresponding to the first candidate diagnosis model using the first candidate value combination.
[0347] Among them, when training the first candidate diagnosis model using the first candidate value combination, use the candidate feature dimension combinations determined in the foregoing steps for training.
[0348] Among them, any one of the candidate feature dimension combinations can be used to extract features from the training samples in the training set, so as to train the first candidate diagnosis model corresponding to the first candidate value combination using the features of each extracted feature dimension. Moreover, the candidate feature dimension combination is also used to extract features from the test samples in the test set, so as to test the performance of the first candidate diagnosis model corresponding to the first candidate value combination using the features of each extracted feature dimension. Correspondingly, for the first candidate diagnosis model using the first candidate value combination, for each of its candidate feature dimension combinations, perform the above training and testing processes, and the subsequent steps will elaborate on this process in detail.
[0349] 705. Use the training set to train the first candidate diagnosis model respectively according to the candidate feature dimension combinations to obtain the corresponding trained first candidate diagnosis model, and each trained first candidate diagnosis model corresponds to a candidate feature dimension combination.
[0350] Among them, using the training set, according to the multiple candidate feature dimension combinations determined in the foregoing step 704, train the first candidate diagnosis model using the first candidate value combination to obtain the trained first candidate diagnosis model. Each trained first candidate diagnosis model corresponds to a candidate feature dimension combination and a first candidate value combination.
[0351] Among them, for the explanation of training the first candidate diagnosis model using a candidate feature dimension combination and the first candidate value combination, please refer to the foregoing Figure 2 explanation in step 205, which will not be elaborated here.
[0352] Correspondingly, for the hyperparameters of the first candidate diagnosis model, use each candidate value combination to separately execute the above training process to obtain each trained first candidate diagnosis model. Each trained first candidate diagnosis model corresponds to a candidate value combination and a candidate feature dimension combination.
[0353] Similarly, for each of the two candidate diagnosis models, separately execute the above process to obtain the trained first candidate diagnosis model corresponding to each candidate diagnosis model. Each trained first candidate diagnosis model corresponds to a first candidate diagnosis model, a candidate value combination, and a candidate feature dimension combination.
[0354] 706. Use the test set to determine the test performance of the trained first candidate diagnosis model;
[0355] Among them, the test set contains test samples and test labels, and the test labels represent whether the rails corresponding to the training samples have rail corrugation faults
[0356] Among them, use the test set to test the performance of each trained first candidate diagnosis model to determine the prediction performance of the trained first candidate diagnosis model. The better the test performance, the more suitable the trained first candidate diagnosis model is for prediction in the field to which the test set belongs.
[0357] Among them, the test set contains a number of test samples to use the test samples to test the performance of the trained first candidate diagnosis model.
[0358] In a possible implementation, it is possible to combine using test samples to test the performance of the trained first candidate diagnosis model. For the specific explanation of this performance test process, please refer to the foregoing Figure 6 explanation, which will not be elaborated here.
[0359] 707. Based on the test performance of the first candidate diagnostic model after training, among the first candidate diagnostic models after training corresponding to each candidate diagnostic model, select the target diagnostic model, and determine the candidate feature dimension combination corresponding to the target diagnostic model as the target feature dimension combination; the test performance of the target diagnostic model is better than that of the non-target diagnostic models; the non-target diagnostic models are the other first candidate diagnostic models after training among the first candidate diagnostic models after training corresponding to each candidate diagnostic model, excluding the target diagnostic model.
[0360] Among them, select the one with the best test performance from multiple trained first candidate diagnostic models as the target diagnostic model.
[0361] Among them, the target diagnostic model corresponds to a target value combination and a target feature dimension combination. By training the first candidate diagnostic model, the target diagnostic model is obtained. The hyperparameters of the target diagnostic model adopt the target value combination, and the input of the target diagnostic model is the features of multiple feature dimensions extracted from the acoustic signal using the target feature dimension combination.
[0362] Among them, the test performance of the target diagnostic model is better than that of other trained first candidate diagnostic models, and the target diagnostic model adopts the target value combination and the target feature dimension combination, and its test performance is better than any one of other value combinations and other candidate feature dimension combinations applied to the target diagnostic model.
[0363] Among them, the target diagnostic model is one of the first candidate diagnostic models after training, and the test performance of the target diagnostic model is better than that of other trained candidate diagnostic models.
[0364] In a possible implementation, each model to be diagnosed is sequentially used as the first model to be diagnosed. For each first candidate diagnostic model, first, among the first candidate diagnostic models obtained by training with each candidate feature dimension combination corresponding to a value combination of the hyperparameters in sequence, select the one with the best test performance as the candidate diagnostic model with the best test performance corresponding to this value combination; then, among the candidate diagnostic models with the best test performance corresponding to each value combination, select the first candidate diagnostic model with the best test performance; finally, among the first candidate diagnostic models with the best test performance corresponding to each candidate diagnostic model, select the one with the best test performance as the target diagnostic model.
[0365] Among them, the target diagnostic model is the first candidate diagnostic model after training with the best test performance. Correspondingly, select the candidate feature dimension combination corresponding to the target diagnostic model as the target feature dimension combination, select the candidate value combination corresponding to the target diagnostic model as the target value combination, the hyperparameters of the target diagnostic model adopt the target value combination, and the input information of the target diagnostic model adopts the features of the target feature dimension combination.
[0366] Among them, the combination of the values of the hyperparameters corresponding to the target diagnostic model is selected as the optimal hyperparameter combination.
[0367] Among them, during the process of training the target diagnostic model, the corresponding combination of candidate feature dimensions and the combination of the values of the hyperparameters are used. Correspondingly, the test performance of the target diagnostic model is better than that of other first candidate diagnostic models after training, and it is affected by the corresponding combination of candidate feature dimensions and the combination of the values of the hyperparameters. Therefore, this combination of candidate feature dimensions is selected as the target feature dimension combination, and this combination of values is selected as the target value combination. During the subsequent actual operation process, the situation of the rail corrugation fault corresponding to the acoustic signal to be processed is predicted.
[0368] As an example, there are 2 candidate diagnostic models, namely diagnostic model 1 and diagnostic model 2. There are 4 combinations of candidate values of the hyperparameters of diagnostic model 1 and 3 combinations of candidate feature dimensions. There are 2 combinations of candidate values of the hyperparameters of diagnostic model 2 and 3 combinations of candidate feature dimensions. The diagnostic model 1 and the diagnostic model 2 are respectively trained by using the above combinations of candidate values of the hyperparameters and the combinations of candidate feature dimensions. It is determined that the test performance of the trained diagnostic model 1 (obtained by training with the combination of the value of the hyperparameter 1 and the combination of candidate feature dimensions 2) is better than that of the diagnostic model 1 obtained by using other combinations of values and combinations of candidate feature dimensions, and is also better than that of the diagnostic model 2 obtained by using 2 combinations of candidate values and 3 combinations of candidate feature dimensions. Correspondingly, the trained diagnostic model 1 is determined as the target diagnostic model, the combination of candidate values of the hyperparameter 1 as the target value combination, and the combination of candidate feature dimensions 2 as the target feature dimension combination.
[0369] In a possible implementation, first select a combination of candidate values to be fixed, traverse the combinations of candidate feature dimensions, and the test performance can be executed during the traversal. When the test performance no longer improves significantly, the combination of candidate feature dimensions with the optimal test performance is determined as the target feature dimension combination corresponding to this combination of candidate values. During this process, the training process may end without training all the combinations of candidate feature dimensions, reducing the amount of data processing during the training process. The above process is sequentially executed for each combination of candidate values of the current first candidate diagnostic model to determine the best feature dimension combination and the best combination of candidate values of this first candidate diagnostic model. According to the above process, each candidate diagnostic model is trained to determine the best feature dimension combination and the best combination of candidate values of each candidate diagnostic model. Then, according to the test performance of each first candidate diagnostic model after training, the first candidate diagnostic model with the best test performance after training is selected as the target diagnostic model, its corresponding best feature dimension combination as the target feature dimension combination, and its corresponding best combination of values as the target value combination.
[0370] Among them, the determined target diagnostic model can be integrated into the actual application scenario, and can be specifically deployed through methods such as APIs, edge computing devices, or cloud platforms. The deployment method is not limited in this application.
[0371] In this application, since the determined target diagnostic model diagnoses rail corrugation faults and is specifically applied to classification problems, the finally selected target diagnostic model is an algorithm model that is better in classification problems.
[0372] In this embodiment, a training set and a test set are obtained. The training set contains a number of training sample pairs, and each training sample pair contains a training sample and a training label. The training label is used to represent whether the rail corresponding to the training sample has a rail corrugation fault; feature extraction is performed on the training samples in the training set based on a number of feature dimensions to obtain the corresponding features of the training samples, and each feature corresponds to a feature dimension; one is sequentially selected from at least two candidate diagnostic models as the first candidate diagnostic model; according to the respective features, a plurality of candidate feature dimension combinations are determined for the first candidate diagnostic model; wherein, each candidate feature dimension combination contains at least two candidate feature dimensions; using the training set, according to the respective candidate feature dimension combinations, the first candidate diagnostic model is respectively trained to obtain the corresponding trained first candidate diagnostic model, and each trained first candidate diagnostic model corresponds to a candidate feature dimension combination; using the test set, the test performance of the trained first candidate diagnostic model is determined; based on the test performance of the trained first candidate diagnostic model, among the trained first candidate diagnostic models respectively corresponding to the candidate diagnostic models, the target diagnostic model is selected, and the candidate feature dimension combination corresponding to the target diagnostic model is determined as the target feature dimension combination; the test performance of the target diagnostic model is better than that of the non-target diagnostic model; the non-target diagnostic model is the other trained first candidate diagnostic models except the target diagnostic model among the trained first candidate diagnostic models respectively corresponding to the candidate diagnostic models. In this process, it is realized to traverse each value combination of hyperparameters and candidate feature dimension combinations, train the candidate diagnostic models, obtain a plurality of trained candidate diagnostic models, determine the one with the best test performance as the target diagnostic model, and determine the value combination and candidate feature dimension combination corresponding to the target diagnostic model as the finally selected hyperparameter value and feature dimension, realizing the optimal configuration of the hyperparameters of the diagnostic model by grid search and selecting the optimal input feature dimension of the diagnostic model, and improving the accuracy of the diagnostic model prediction.
[0373] Figure 8 It is a simplified flowchart of the training process of the diagnostic model provided by the embodiment of the present application, including steps 801 to 804, which will be described in detail below.
[0374] 801. Data preprocessing;
[0375] Among them, through filtering technology and statistical methods, the acquired acoustic signals collected during the train's movement are cleaned and repaired, and the acoustic signals with longer durations are split into several acoustic samples with shorter durations, and labels are added to them to mark whether the acoustic sample has a rail corrugation fault.
[0376] Moreover, the acoustic sample is split into multiple subsets, and any subset can be used as a training set or a test set.
[0377] 802. Determination of feature dimension combinations;
[0378] Among them, for multiple candidate diagnostic models, multiple candidate feature dimension combinations that can be used as their input features are determined in sequence. The features include time-domain features and frequency-domain features.
[0379] Among them, a candidate feature dimension combination composed of several feature dimensions with relatively high importance can be determined from several feature dimensions. The specific process can refer to the process recorded in the foregoing embodiments.
[0380] Among them, multiple candidate feature dimension combinations can be determined in sequence from several feature dimensions according to the importance level. The feature dimensions included in each candidate feature dimension combination are different. The specific process can refer to the process recorded above.
[0381] 803. Training of diagnostic models;
[0382] Among them, using the subsets obtained by splitting in the foregoing data preprocessing, the training sets and test sets required for training and predicting each candidate diagnostic model.
[0383] Among them, using the training set, a corresponding candidate diagnostic model is trained in combination with a candidate feature dimension combination.
[0384] Among them, when training the candidate diagnostic model, the hyperparameters of the candidate diagnostic model are also determined. For the specific process, please refer to the process explanation of training the first diagnostic model above.
[0385] 804. Performance testing of candidate diagnostic models.
[0386] Among them, after training the candidate diagnostic model, the performance of the trained candidate diagnostic model is tested using the test set.
[0387] In specific implementation, a cross-validation strategy can be adopted. Using multiple subsets obtained by preprocessing, the test set and the training set are respectively determined to train the candidate diagnostic model and perform performance testing. The average value of the results of multiple performance tests is used as the test performance of the trained candidate diagnostic model.
[0388] Subsequently, during the training process of the candidate diagnostic model, the best combination is determined among several time-domain features and frequency-domain features to enhance the expressiveness and generalization ability of the candidate diagnostic model. The expressiveness includes high accuracy and strong real-time performance, so that it can accurately identify the rail corrugation fault.
[0389] Among them, for multiple candidate diagnostic models, one is selected for training each time, and steps 802-804 are executed cyclically until all candidate diagnostic models are completely trained. Using the test performance of each trained candidate diagnostic model, the one with the best test performance is selected as the finally selected target diagnostic model, and the combination of candidate feature dimensions used in the training process of the target diagnostic model is used as the combination of feature dimensions required for acoustic signal feature extraction.
[0390] In this embodiment, multiple combinations of candidate feature dimensions corresponding to each candidate diagnostic model are first determined, and then based on these multiple combinations of candidate feature dimensions, the diagnostic model is trained using the cross-validation strategy. During this training process, the cross-validation strategy is combined to select the optimal combination of hyperparameters, ensuring that the trained diagnostic model can not only achieve a high prediction accuracy but also meet the real-time requirements.
[0391] The above introduces a rail corrugation fault diagnosis method provided by an embodiment of the present application. Next, a device for executing the above rail corrugation fault diagnosis method will be introduced.
[0392] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a rail corrugation fault diagnosis device provided by an embodiment of the present application. As Figure 9 shown, the rail corrugation fault diagnosis device 900 includes:
[0393] An acquisition module 901, configured to acquire an acoustic signal to be processed, where the acoustic signal to be processed includes the acoustic signal generated when the train travels in contact with the rail;
[0394] An extraction module 902, configured to perform feature extraction on the acoustic signal to be processed according to the target feature dimension combination to obtain a target feature set; wherein, the target feature dimension combination includes at least two feature dimensions, each feature dimension corresponds to an acoustic feature type, and at least two feature dimensions in the target feature dimension combination are selected from several feature dimensions according to the importance of each feature dimension, and the importance of each feature dimension is determined according to the marginal contribution of each feature dimension to the output of the target diagnostic model;
[0395] A processing module 903, configured to use the pre-trained target diagnostic model to process the target feature set to obtain the rail corrugation fault diagnosis result corresponding to the acoustic signal to be processed.
[0396] In a possible implementation, it further includes:
[0397] A training module for the process of pre-training to obtain a target diagnostic model;
[0398] The training module includes:
[0399] An obtaining unit for obtaining a training set and a test set. The training set contains a number of training sample pairs, and each training sample pair contains a training sample and a training label, where the training label is used to characterize whether the rail corresponding to the training sample has a rail corrugation fault;
[0400] An extraction unit for extracting features of the training samples in the training set based on a number of feature dimensions to obtain the corresponding features of the training samples, and each feature corresponds to a feature dimension;
[0401] A first selection unit for sequentially selecting one from at least two candidate diagnostic models as the first candidate diagnostic model;
[0402] A first determination unit for determining multiple candidate feature dimension combinations for the first candidate diagnostic model according to the respective features; among them, each candidate feature dimension combination contains at least two candidate feature dimensions, and each candidate feature dimension is determined from a number of feature dimensions according to the importance of each feature dimension;
[0403] A training unit for using the training set to train the first candidate diagnostic model respectively according to the respective candidate feature dimension combinations to obtain the corresponding trained first candidate diagnostic model, and each trained first candidate diagnostic model corresponds to a candidate feature dimension combination;
[0404] A second determination unit for using the test set to determine the test performance of the trained first candidate diagnostic model;
[0405] A second selection unit for selecting a target diagnostic model from the trained first candidate diagnostic models respectively corresponding to the candidate diagnostic models based on the test performance of the trained first candidate diagnostic model, and determining the candidate feature dimension combination corresponding to the target diagnostic model as the target feature dimension combination; the test performance of the target diagnostic model is better than that of the non-target diagnostic model; the non-target diagnostic model is other trained first candidate diagnostic models except the target diagnostic model among the trained first candidate diagnostic models respectively corresponding to the candidate diagnostic models.
[0406] In a possible implementation, the second determination unit includes:
[0407] An input subunit, configured to respectively input the test samples in the test set into each of the first candidate diagnosis models after training, and obtain the prediction results of each of the first candidate diagnosis models after training. The test set includes a plurality of test sample pairs, and each test sample pair includes a test sample and a test label, where the test label is used to characterize whether the rail corresponding to the test sample has a rail corrugation fault.
[0408] A first determination subunit, configured to determine the test performance of each of the first candidate diagnosis models after training based on the prediction results and the corresponding test labels.
[0409] In a possible implementation, the first candidate diagnosis model includes hyperparameters. The first determination unit includes:
[0410] An obtaining subunit, configured to obtain at least two candidate value combinations of the hyperparameters in the first candidate diagnosis model;
[0411] A second determination subunit, configured to sequentially determine one of the at least two candidate value combinations as a first candidate value combination;
[0412] A third determination subunit, configured to determine the first marginal contribution of each feature dimension pair to the output of the first candidate diagnosis model using the first candidate value combination according to each feature;
[0413] A fourth determination subunit, configured to determine the importance of each feature to the first candidate diagnosis model using the first candidate value combination according to the first marginal contribution;
[0414] A selection subunit, configured to select at least two feature dimensions from the several feature dimensions corresponding to each feature to form at least one candidate feature dimension combination according to the importance of each feature to the first candidate diagnosis model using the first candidate value combination;
[0415] A fifth determination subunit, configured to use the at least one candidate feature dimension combination as the at least one candidate feature dimension combination corresponding to the first candidate diagnosis model using the first candidate value combination.
[0416] In a possible implementation, the selection subunit is specifically configured to:
[0417] Select at least two feature dimensions with importance greater than a preset importance threshold from the several feature dimensions corresponding to each feature according to the importance of each feature to the first candidate diagnosis model using the first candidate value combination; and determine the candidate feature dimension combination corresponding to the first candidate diagnosis model using the first candidate value combination based on the at least two feature dimensions; or
[0418] According to the importance of each feature for the first candidate diagnosis model using the first candidate value combination, sort the feature dimensions corresponding to each feature to obtain the sorting result of each feature dimension; based on the sorting result, determine the first candidate feature dimension combination; wherein, the first candidate feature dimension combination includes two feature dimensions, and the importance of each feature dimension in the first candidate feature dimension combination is higher than that of other feature dimensions; the other feature dimensions include any feature dimension in the several feature dimensions except the two feature dimensions in the first candidate feature dimension combination; according to the order of importance from high to low in the sorting result, sequentially add one feature dimension to the first candidate feature dimension combination to obtain at least two candidate feature dimension combinations corresponding to the first candidate diagnosis model using the first candidate value combination; or
[0419] According to the importance of each feature for the first candidate diagnosis model using the first candidate value combination, sort the feature dimensions corresponding to each feature to determine the second candidate feature dimension combination, and the second candidate feature dimension combination includes the several feature dimensions; according to the order of importance from low to high in the sorting result, sequentially remove one feature dimension from the second candidate feature dimension combination from the several feature dimensions to obtain at least two candidate feature dimension combinations corresponding to the first candidate diagnosis model using the first candidate value combination.
[0420] In a possible implementation, each trained first candidate diagnosis model includes a candidate value combination, and further includes:
[0421] A sixth determination subunit, configured to determine the candidate value combination corresponding to the target diagnosis model as the target value combination based on the candidate value combination corresponding to each trained first candidate diagnosis model;
[0422] A seventh determination subunit, configured to determine the hyperparameters of the target diagnosis model based on the target value combination.
[0423] In a possible implementation, the obtaining unit is specifically configured to:
[0424] Obtain the original sample acoustic signal;
[0425] Based on a set time period, intercept the original sample acoustic signal to obtain a sample acoustic signal set, and the sample acoustic signal set includes several sample acoustic signals;
[0426] Add labels to each sample acoustic signal to obtain a data set, and the label is used to characterize whether the rail corresponding to the corresponding sample acoustic signal has a rail corrugation fault;
[0427] Based on a set ratio, divide the data set into a test set and a training set, and the ratio of the number of test sample pairs included in the test set to the number of training sample pairs included in the training set satisfies the set ratio.
[0428] In a possible implementation, the at least two feature dimensions include at least two of the following: A-weighted sound pressure level feature, root mean square value of energy, peak factor, impulse factor, margin factor, kurtosis factor, skewness factor, zero crossing rate, average spectral information entropy, average spectral centroid, average spectral kurtosis, Mel frequency cepstral coefficients.
[0429] It should be noted that for the functional explanations of the components of a rail corrugation fault diagnosis device provided in this embodiment, please refer to the explanations in the foregoing method embodiment, and will not be elaborated in this embodiment.
[0430] In this embodiment, a training set and a test set are obtained. The training set contains a number of training sample pairs, and each training sample pair contains a training sample and a training label, where the training label is used to represent whether the rail corresponding to the training sample has a rail corrugation fault; feature extraction is performed on the training samples in the training set based on a number of feature dimensions to obtain the corresponding features of the training samples, and each feature corresponds to a feature dimension; one is sequentially selected from at least two candidate diagnostic models as the first candidate diagnostic model; based on the respective features, a number of candidate feature dimension combinations are determined for the first candidate diagnostic model; wherein, each candidate feature dimension combination contains at least two candidate feature dimensions; the training set is used to train the first candidate diagnostic model respectively according to the respective candidate feature dimension combinations to obtain the corresponding trained first candidate diagnostic model, and each trained first candidate diagnostic model corresponds to a candidate feature dimension combination; the test set is used to determine the test performance of the trained first candidate diagnostic model; based on the test performance of the trained first candidate diagnostic model, among the trained first candidate diagnostic models respectively corresponding to the candidate diagnostic models, a target diagnostic model is selected, and the candidate feature dimension combination corresponding to the target diagnostic model is determined as the target feature dimension combination; the test performance of the target diagnostic model is better than that of the non-target diagnostic models; the non-target diagnostic models are the other trained first candidate diagnostic models except the target diagnostic model among the trained first candidate diagnostic models respectively corresponding to the candidate diagnostic models. This device realizes traversing each value combination of hyperparameters and candidate feature dimension combinations, training the first candidate diagnostic model to obtain a number of trained first candidate diagnostic models, determining the one with the best test performance as the target diagnostic model among them, and determining the value combination and candidate feature dimension combination corresponding to the target diagnostic model as the finally selected hyperparameter value and feature dimension, realizing the optimal configuration of the hyperparameters of the diagnostic model by grid search and selecting the optimal input feature dimension of the diagnostic model, and improving the accuracy of the prediction of the diagnostic model.
[0431] An electronic device is also provided in an embodiment of the present application. Refer to Figure 10As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the rail corrugation fault diagnosis method in the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 10 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0432] As Figure 10 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1008 into the random access memory (RAM) 1003. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 1003. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.
[0433] Generally, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a memory card, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 10 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0434] In the embodiments of the present application, there is also provided a computer program product including computer-readable instructions. When the computer-readable instructions run on the electronic device, the electronic device is enabled to implement any one of the rail corrugation fault diagnosis methods provided in the embodiments of the present application.
[0435] In the embodiments of the present application, there is also provided a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device can be enabled to implement any one of the rail corrugation fault diagnosis methods provided in the embodiments of the present application.
[0436] It should be further noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.
[0437] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for this application, software program implementation is a better implementation method in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.
[0438] In the above embodiments, it can be implemented in whole or in part through software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0439] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
Claims
1. A method for diagnosing rail corrugation faults, characterized in that, Including: Obtain an acoustic signal to be processed, where the acoustic signal to be processed includes an acoustic signal generated when a train contacts a rail during travel; According to a target feature dimension combination, perform feature extraction on the acoustic signal to be processed to obtain a target feature set; wherein, the target feature dimension combination includes at least two feature dimensions, each feature dimension corresponds to an acoustic feature type, and at least two feature dimensions in the target feature dimension combination are selected from several feature dimensions according to the importance of each feature dimension, and the importance of each feature dimension is determined according to the marginal contribution of each feature dimension to the output of the target diagnosis model; Use the pre-trained target diagnosis model to process the target feature set to obtain the rail corrugation fault diagnosis result corresponding to the acoustic signal to be processed.
2. The rail corrugation fault diagnosis method according to claim 1, wherein The process of pre-training the target diagnosis model includes: Obtain a training set and a test set, where the training set includes several training sample pairs, each training sample pair includes a training sample and a training label, and the training label is used to represent whether the rail corresponding to the training sample has a rail corrugation fault; Based on several feature dimensions, perform feature extraction on the training samples in the training set to obtain the features corresponding to the training samples, and each feature corresponds to a feature dimension; Select one from at least two candidate diagnosis models in turn as the first candidate diagnosis model; According to the features, determine multiple candidate feature dimension combinations for the first candidate diagnosis model; wherein, each candidate feature dimension combination includes at least two candidate feature dimensions, and each candidate feature dimension is determined from several feature dimensions according to the importance of each feature dimension; Use the training set to train the first candidate diagnosis model respectively according to the candidate feature dimension combinations to obtain the corresponding trained first candidate diagnosis model, and each trained first candidate diagnosis model corresponds to a candidate feature dimension combination; Use the test set to determine the test performance of the trained first candidate diagnosis model; Based on the test performance of the trained first candidate diagnosis model, select the target diagnosis model from the trained first candidate diagnosis models corresponding to the candidate diagnosis models respectively, and determine the candidate feature dimension combination corresponding to the target diagnosis model as the target feature dimension combination; the test performance of the target diagnosis model is better than that of the non-target diagnosis model; the non-target diagnosis model is other trained first candidate diagnosis models except the target diagnosis model among the trained first candidate diagnosis models corresponding to the candidate diagnosis models.
3. The method for diagnosing rail corrugation faults according to claim 2, characterized in that, The step of using the test set to determine the test performance of the trained first candidate diagnosis model includes: Input the test samples in the test set into each trained first candidate diagnosis model respectively to obtain the prediction results of each trained first candidate diagnosis model. The test set includes several test sample pairs, each test sample pair includes a test sample and a test label, and the test label is used to represent whether the rail corresponding to the test sample has a rail corrugation fault; Based on the prediction results and the corresponding test labels, determine the test performance of each trained first candidate diagnosis model.
4. The rail corrugation fault diagnosis method according to claim 2, characterized in that The first candidate diagnostic model includes hyperparameters. Based on the respective features, determining multiple candidate feature dimension combinations for the first candidate diagnostic model includes: Obtaining at least two candidate value combinations of the hyperparameters in the first candidate diagnostic model; Sequentially determining one of the at least two candidate value combinations as the first candidate value combination; Based on the respective features, determining the first marginal contribution of each feature dimension pair to the output of the first candidate diagnostic model using the first candidate value combination; Based on the first marginal contribution, determining the importance of each feature to the first candidate diagnostic model using the first candidate value combination; Based on the importance of each feature to the first candidate diagnostic model using the first candidate value combination, selecting at least two feature dimensions from among the several feature dimensions corresponding to each feature to form at least one candidate feature dimension combination; Regarding the at least one candidate feature dimension combination as the at least one candidate feature dimension combination corresponding to the first candidate diagnostic model using the first candidate value combination.
5. The rail corrugation fault diagnosis method according to claim 4, wherein Based on the importance of each feature to the first candidate diagnostic model using the first candidate value combination, selecting at least two feature dimensions from among the several feature dimensions corresponding to each feature to form at least one candidate feature dimension combination, including: Based on the importance of each feature to the first candidate diagnostic model using the first candidate value combination, selecting at least two feature dimensions with importance greater than a preset importance threshold from among the several feature dimensions corresponding to each feature; based on the at least two feature dimensions, determining the candidate feature dimension combination corresponding to the first candidate diagnostic model using the first candidate value combination; or Based on the importance of each feature to the first candidate diagnostic model using the first candidate value combination, sorting the feature dimensions corresponding to each feature to obtain the sorting result of the respective feature dimensions; based on the sorting result, determining a first candidate feature dimension combination; wherein, the first candidate feature dimension combination includes two feature dimensions, and the importance of each feature dimension in the first candidate feature dimension combination is higher than that of any other feature dimension; the other feature dimensions include any feature dimension among the several feature dimensions other than the two feature dimensions in the first candidate feature dimension combination; according to the order of importance from high to low in the sorting result, sequentially adding one feature dimension to the first candidate feature dimension combination to obtain at least two candidate feature dimension combinations corresponding to the first candidate diagnostic model using the first candidate value combination; or Based on the importance of each feature to the first candidate diagnostic model using the first candidate value combination, sorting the feature dimensions corresponding to each feature to determine a second candidate feature dimension combination, the second candidate feature dimension combination including the several feature dimensions; according to the order of importance from low to high in the sorting result, sequentially removing one feature dimension from the second candidate feature dimension combination among the several feature dimensions to obtain at least two candidate feature dimension combinations corresponding to the first candidate diagnostic model using the first candidate value combination.
6. The rail corrugation fault diagnosis method according to any one of claims 4-5, characterized in that Each trained first candidate diagnostic model includes a candidate value combination. The method further includes: Based on the candidate value combinations corresponding to each trained first candidate diagnostic model, determine the candidate value combination corresponding to the target diagnostic model as the target value combination; Based on the target value combination, determine the hyperparameters of the target diagnostic model.
7. The rail corrugation fault diagnosis method according to claim 2, characterized in that, The training set and the test set are determined in the following manner: Obtain the original sample acoustic signals; Based on a set time period, intercept the original sample acoustic signals to obtain a sample acoustic signal set, where the sample acoustic signal set contains a number of sample acoustic signals; Add labels to each sample acoustic signal to obtain a data set, where the labels are used to characterize whether the rail corresponding to the corresponding sample acoustic signal has a rail corrugation fault; Based on a set ratio, divide the data set into a test set and a training set, where the ratio of the number of test sample pairs contained in the test set to the number of training sample pairs contained in the training set satisfies the set ratio.
8. The rail corrugation fault diagnosis method according to claim 1, wherein The at least two feature dimensions include at least two of the following: A-weighted sound pressure level feature, root mean square value of energy, peak factor, impulse factor, margin factor, kurtosis factor, skewness factor, zero crossing rate, average spectral information entropy, average spectral centroid, average spectral kurtosis, Mel frequency cepstral coefficients.
9. A training method for a diagnostic model, the diagnostic model being used for diagnosing rail corrugation faults, characterized in that, The training method includes: Obtain a training set and a test set, where the training set contains a number of training sample pairs, and each training sample pair contains a training sample and a training label, and the training label is used to characterize whether the rail corresponding to the training sample has a rail corrugation fault; Based on a number of feature dimensions, perform feature extraction on the training samples in the training set to obtain the respective features corresponding to the training samples, and each feature corresponds to a feature dimension; Sequentially select one from at least two candidate diagnostic models as the first candidate diagnostic model; According to the respective features, determine multiple candidate feature dimension combinations for the first candidate diagnostic model; among them, each candidate feature dimension combination contains at least two candidate feature dimensions; Use the training set, and respectively train the first candidate diagnostic model according to the respective candidate feature dimension combinations to obtain the corresponding trained first candidate diagnostic model, and each trained first candidate diagnostic model corresponds to a candidate feature dimension combination; Use the test set to determine the test performance of the trained first candidate diagnostic model; Based on the test performance of the trained first candidate diagnostic model, select a target diagnostic model from the trained first candidate diagnostic models corresponding to the respective candidate diagnostic models, and determine the candidate feature dimension combination corresponding to the target diagnostic model as the target feature dimension combination; the test performance of the target diagnostic model is better than that of the non-target diagnostic models; the non-target diagnostic models are the other trained first candidate diagnostic models except the target diagnostic model among the trained first candidate diagnostic models corresponding to the respective candidate diagnostic models.
10. The training method of the diagnostic model according to claim 9, wherein, The first candidate diagnostic model contains hyperparameters. According to the respective features, determining multiple candidate feature dimension combinations for the first candidate diagnostic model includes: Obtain at least two candidate value combinations of the hyperparameters in the first candidate diagnostic model; Sequentially determine one from the at least two candidate value combinations as the first candidate value combination; Determine the first marginal contribution of each feature dimension to the output of the first candidate diagnostic model using the first candidate value combination according to each feature; Determine the importance of each feature to the first candidate diagnostic model using the first candidate value combination according to the first marginal contribution; Select at least two feature dimensions from the several feature dimensions corresponding to each feature to form at least one candidate feature dimension combination according to the importance of each feature to the first candidate diagnostic model using the first candidate value combination; Use the at least one candidate feature dimension combination as the at least one candidate feature dimension combination corresponding to the first candidate diagnostic model using the first candidate value combination.
Citation Information
Patent Citations
GIS high-voltage isolation switch unknown fault self-learning method
CN117312965A