Method for predicting service life of high-speed train bearing and related equipment

A high-speed train bearing life prediction model is constructed through macro and micro strategy networks, and multi-objective reward function optimization is used to solve the problem of low accuracy in existing technologies, achieving efficient and accurate life prediction, which is suitable for edge computing environments.

CN120805696APending Publication Date: 2025-10-17CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510938937.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies have the problem of low accuracy in predicting the life of high-speed train bearings, especially in the case of high-dimensional, long-time series industrial equipment data. Traditional methods have high computational costs, redundant model structures and difficulty in lightweight deployment, and the black box search mechanism lacks explanatory power, resulting in large prediction errors.

Method used

The macro-strategy network is used to obtain the model architecture and the micro-strategy network is used to obtain the model parameters. A bearing life prediction model is constructed, and the model is optimized through a multi-objective reward function to achieve hierarchical construction and parameter optimization to improve prediction accuracy.

Benefits of technology

The accuracy and computational efficiency of high-speed train bearing life prediction are improved, the inference time and parameter quantity of the model are reduced, the lightweight requirements of edge computing are met, and the interpretability of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805696A_ABST
    Figure CN120805696A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of high-speed train bearing life prediction, and provides a high-speed train bearing life prediction method and related equipment, and the method comprises the steps: obtaining a plurality of fault vibration signals of a high-speed train bearing; acquiring a model architecture by using the macroscopic strategy network, and acquiring model parameters by using the microcosmic strategy network; constructing a bearing life prediction model based on the model architecture and the model parameters; carrying out life prediction according to each fault vibration signal by utilizing a bearing life prediction model, and constructing a multi-target reward function based on reasoning time corresponding to all life prediction results; and optimizing the bearing life prediction model based on the multi-target reward function to obtain an optimized bearing life prediction model, and performing life prediction on the to-be-predicted high-speed train bearing by using the optimized bearing life prediction model. According to the method, the accuracy of predicting the service life of the high-speed train bearing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bearing life prediction of high-speed trains, and particularly relates to a bearing life prediction method of high-speed trains and related equipment. BACKGROUND

[0002] In the field of industrial equipment health management, rotating machinery plays an irreplaceable key role. The bearing of a high-speed train, as one of the core rotating components in a bogie, usually operates under harsh conditions such as high speed and heavy load, and its failure can cause system-level failure or even catastrophic accidents. Therefore, reliable condition monitoring is of great significance to the safe operation of high-speed trains. Remaining useful life (RUL) prediction is a core technology to ensure the safe operation of equipment.

[0003] Because the axle box bearing of a high-speed train bears all vertical and lateral loads transmitted from the car body and the bogie during operation, it also bears certain longitudinal loads along the track direction under frequent acceleration and braking conditions. The life prediction methods for ordinary bearings in the past are not completely applicable. With the rapid development of deep learning technology, bearing life prediction technology based on deep learning has been widely applied. For example, convolutional neural networks have strong ability to explore hidden information in original signals without knowing the damage mechanism of components.

[0004] Convolutional neural networks perform well in bearing life prediction, but the form and structure of neural networks will vary according to different tasks and requirements. For example, classic convolutional neural network models such as AlexNet, VGGNet and GoogLeNet, etc. These models need extensive hyperparameter tuning and architecture adjustment when facing different task requirements, otherwise the model performance will be greatly discounted. However, designing specific architectures for different tasks and requirements will consume a lot of expert time, and the whole process is time-consuming and laborious.

[0005] With the application of deep learning technology, life prediction models based on neural architecture search (NAS) have gradually become a research hotspot, but the existing technology still has the following problems: 1. The current mainstream NAS method faces the dual challenges of efficiency and practicality in industrial scenarios. Typical industrial equipment data (such as high-frequency vibration signals) have the characteristics of high dimension and long time sequence. The traditional search strategy needs to traverse a large number of architecture combinations, resulting in a sharp increase in computing cost. The generated model generally has the problems of large parameter quantity and redundant structure, which is difficult to meet the lightweight deployment requirements of the edge computing end. In addition, the black box search mechanism lacks the ability to explain the architecture decision, for example, it cannot quantify the influence of the depth of the convolutional layer or the type of recurrent neural network unit on the prediction result, making it difficult for the model to meet the requirements of the industrial certification standard. 2. The strong time sequence correlation of the equipment degradation process requires the model to have long-term dependence capturing ability. In the existing method, the traditional jump connection cannot dynamically adapt to the change of time sequence feature scale in the stacked architecture. Research shows that the transmission efficiency of time sequence features in deep network architecture will decrease significantly with the increase of network depth, directly leading to a large increase in prediction error. 3. The strategy of increasing network depth to improve accuracy in the existing technology leads to gradient path blockage. In the back propagation process of deep network, the gradient amplitude of low-level parameters decays exponentially, causing the degradation of front-end feature extraction capability. As can be seen, the performance of the life prediction model obtained by the existing technology is insufficient, resulting in the problem of low accuracy of life prediction of high-speed train bearings. SUMMARY

[0006] The application provides a life prediction method for high-speed train bearings and related equipment, which can solve the problem of low accuracy of life prediction of high-speed train bearings.

[0007] In a first aspect, the embodiments of the application provide a life prediction method for high-speed train bearings, which comprises:

[0008] Obtaining a plurality of fault vibration signals of high-speed train bearings;

[0009] Obtaining a model architecture using a macro-strategy network and obtaining model parameters using a micro-strategy network;

[0010] Constructing a bearing life prediction model based on the model architecture and the model parameters;

[0011] Using the bearing life prediction model to perform life prediction according to each fault vibration signal, and constructing a multi-objective reward function based on the reasoning time corresponding to all life prediction results; the multi-objective reward function is used to describe the calculation efficiency of the bearing life prediction model;

[0012] The bearing life prediction model is optimized based on a multi-objective reward function, and an optimized bearing life prediction model is obtained, and the optimized bearing life prediction model is used to predict the life of the bearing of the high-speed train.

[0013] Optionally, the model architecture includes a number of CNN layers, a number of GRU layers, a connection type, and an attention type.

[0014] The model architecture is obtained by using a macro policy network, including:

[0015] The environment state is input into the macro policy network for calculation, and the number of CNN layers, the number of GRU layers, the connection type, and the attention type are output; the environment state includes a four-dimensional position encoding vector corresponding to the number of CNN layers, a four-dimensional position encoding vector corresponding to the number of GRU layers, a four-dimensional position encoding vector corresponding to the connection type, and a four-dimensional position encoding vector corresponding to the attention type.

[0016] Optionally, the model parameters include a CNN channel number, a convolution kernel size, and a GRU hidden unit number.

[0017] The model parameters are obtained by using a micro policy network, including:

[0018] The model architecture is input into the micro policy network for calculation, and the CNN channel number, the convolution kernel size, and the GRU hidden unit number are output.

[0019] Optionally, the bearing life prediction model is constructed based on the model architecture and the model parameters, including:

[0020] The N CNN layers and the M GRU layers are sequentially connected according to the connection type, and the attention mechanism in each CNN layer and each GRU layer is configured according to the attention type; each CNN layer is configured according to the CNN channel number and the convolution kernel size, and each GRU layer is configured according to the GRU hidden unit number, to obtain the bearing life prediction model; N represents the number of CNN layers, M represents the number of GRU layers, the input end of the first CNN layer is the input end of the bearing life prediction model, and the output end of the last GRU layer is the output end of the bearing life prediction model.

[0021] Optionally, the multi-objective reward function is:

[0022]

[0023] wherein R represents a multi-objective reward function value, a, β, and γ represent weight coefficients, MSE represents a mean square error, θ represents a model parameter quantity, θ target represents a target parameter quantity, t inf represents an average value of inference times corresponding to all life prediction results, and t ref represents a reference inference time.

[0024] Optionally, the bearing life prediction model is optimized based on a multi-objective reward function to obtain an optimized bearing life prediction model, including:

[0025] Determine whether the multi-objective reward function value is greater than the preset reward value;

[0026] If so, the bearing life prediction model is used as the optimized bearing life prediction model;

[0027] Otherwise, the macro policy network and the micro policy network are updated based on the attention mechanism, and the steps of using the macro policy network to obtain the model architecture and using the micro policy network to obtain the model parameters are returned.

[0028] Optionally, the macro-policy network and micro-policy network are updated based on the attention mechanism, including:

[0029] Calculate the policy update gradient based on the attention mechanism;

[0030] The parameters in the macro policy network and the micro policy network are updated according to the policy update gradient.

[0031] Optionally, calculate the policy update gradient based on the attention mechanism, including:

[0032] By formula:

[0033]

[0034] w i =softmax(f θ (|o i ;a i |))

[0035] Calculate policy update gradients

[0036] Among them, A(s,a i ) represents the advantage function, s represents the environmental state, π θ represents the action probability distribution, w i represents the attention weight, represents the gradient operator, f θ Represents the parameterized function of the policy network, a i represents the i-th action, o i represents the i-th observation value.

[0037] In a second aspect, an embodiment of the present application provides a life prediction device for a high-speed train bearing, comprising:

[0038] A first acquisition module is used to acquire multiple fault vibration signals of high-speed train bearings;

[0039] The second obtaining module is configured to obtain a model architecture by using the macro strategy network and obtain model parameters by using the micro strategy network.

[0040] The construction module is configured to construct the bearing life prediction model based on the model architecture and the model parameters.

[0041] The life prediction module is configured to perform life prediction according to each fault vibration signal by using the bearing life prediction model, and construct a multi-objective reward function based on reasoning time corresponding to all life prediction results. The multi-objective reward function is used to describe the calculation efficiency of the bearing life prediction model.

[0042] The optimization module is configured to optimize the bearing life prediction model based on the multi-objective reward function, obtain an optimized bearing life prediction model, and perform life prediction on the bearing to be predicted of the high-speed train by using the optimized bearing life prediction model.

[0043] In a third aspect, an embodiment of the present application provides a terminal device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the life prediction method of the bearing of the high-speed train is implemented.

[0044] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, the life prediction method of the bearing of the high-speed train is implemented.

[0045] The above-mentioned scheme of the present application has the following advantages:

[0046] In the embodiment of the present application, by acquiring a plurality of fault vibration signals of the high-speed train bearing, then acquiring the model architecture by using the macro strategy network, and acquiring the model parameters by using the micro strategy network, the bearing life prediction model is constructed based on the model architecture and the model parameters, then the bearing life prediction model is used to predict the life according to each fault vibration signal, and the multi-objective reward function is constructed based on the reasoning time corresponding to all life prediction results, finally the bearing life prediction model is optimized based on the multi-objective reward function, and the optimized bearing life prediction model is obtained, and the optimized bearing life prediction model is used to predict the life of the high-speed train bearing to be predicted. Wherein, the bearing life prediction model is constructed by using the macro strategy network and the micro strategy network, which can respectively calculate the architecture and the parameters of the model, realize the hierarchical construction of the bearing life prediction model, construct the multi-objective reward function based on the reasoning time, which can quantify the calculation efficiency of the constructed bearing life prediction model, and optimize the bearing life prediction model by using the multi-objective reward function, which can effectively optimize the prediction accuracy and reasoning time of the bearing life prediction model, and further improve the accuracy of the life prediction of the high-speed train bearing.

[0047] Other benefits of the present application will be described in detail in the subsequent specific embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0049] Figure 1 The flowchart of the bearing life prediction method for an embodiment of the present application is provided.

[0050] Figure 2 The structural schematic diagram of the bearing life prediction model provided by an embodiment of the present application is provided.

[0051] Figure 3 The importance score schematic diagram provided by an embodiment of the present application is provided.

[0052] Figure 4 The quantitative index comparison schematic diagram provided by an embodiment of the present application is provided.

[0053] Figure 5 The self-attention heat map provided by an embodiment of the present application is provided.

[0054] Figure 6 The structural schematic diagram of the bearing life prediction device provided by an embodiment of the present application is provided.

[0055] Figure 7 The structural schematic diagram of a terminal device provided for an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0056] In the following description, for the purposes of explanation and not limitation, specific details are set forth such as particular architectures, techniques, etc. in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, circuits, and

[0057] It should be understood that the term "comprises" when used in this specification and the appended claims, specifies the presence of stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0058] It should also be understood that the term "and / or" when used in this specification and the appended claims, means any one or more of the associated listed items can be present, and includes multiples of any one or more of the associated listed items.

[0059] As used in this specification and the appended claims, the term "if" can be interpreted as meaning "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [the described condition or event]" or "in response to detecting [the described condition or event]," depending on the context.

[0060] In addition, in the description of the specification and the appended claims, the terms "first," "second," "third," etc. are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0061] Reference within the specification of this application to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in additional embodiments," and so on, in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily referring to some, but not all, embodiments. The terms "including," "comprising," "having," and variations thereof are meant to encompass the items listed thereafter, but do not exclude other items from also being present.

[0062] To solve the problem of low accuracy of existing life prediction of high-speed train bearings, the embodiments of the present application provide a life prediction method of high-speed train bearings. The life prediction method obtains a plurality of fault vibration signals of high-speed train bearings, then obtains a model architecture by using a macro-strategy network, obtains model parameters by using a micro-strategy network, constructs a bearing life prediction model based on the model architecture and the model parameters, then uses the bearing life prediction model to predict the life according to each fault vibration signal, constructs a multi-objective reward function based on the reasoning time corresponding to all life prediction results, finally optimizes the bearing life prediction model based on the multi-objective reward function, obtains an optimized bearing life prediction model, and uses the optimized bearing life prediction model to predict the life of the high-speed train bearing to be predicted. Wherein, the macro-strategy network and the micro-strategy network are used to construct the bearing life prediction model, which can perform strategy calculation on the architecture and parameters of the model respectively, realize hierarchical construction of the bearing life prediction model, construct the multi-objective reward function based on the reasoning time, which can quantify the calculation efficiency of the constructed bearing life prediction model, and use the multi-objective reward function to optimize the bearing life prediction model, which can effectively optimize the prediction accuracy and reasoning time of the bearing life prediction model, thereby improving the accuracy of life prediction of high-speed train bearings.

[0063] Next, the life prediction method of high-speed train bearings provided by the present application is exemplarily described.

[0064] As shown in Figure 1 The life prediction method of high-speed train bearings provided by the present application includes the following steps:

[0065] Step 11, obtaining a plurality of fault vibration signals of high-speed train bearings.

[0066] The above fault vibration signal includes vibration signals of high-speed train bearings in historical time periods. The plurality of fault vibration signals correspond one-to-one to a plurality of historical time periods.

[0067] In some embodiments of the present application, the fault vibration signal of the high-speed train bearing can be obtained by a sensor or the like.

[0068] Step 12, obtain the model architecture by using the macro strategy network, and obtain the model parameters by using the micro strategy network.

[0069] The above model architecture includes the number of CNN layers, the number of GRU layers, the connection type (such as residual connection, dense connection, etc.), and the attention type (such as channel attention, self-attention, etc.), and the model parameters include the number of CNN channels, the size of the convolution kernel, and the number of GRU hidden units.

[0070] In some embodiments of the present application, the macro strategy network is used to output the model architecture according to the environment state, which can be a multilayer perceptron (MLP, Multilayer Perceptron), and the environment state includes a four-dimensional position encoding vector corresponding to the number of CNN layers, a four-dimensional position encoding vector corresponding to the number of GRU layers, a four-dimensional position encoding vector corresponding to the connection type, and a four-dimensional position encoding vector corresponding to the attention type. The micro strategy network is used to output the model parameters according to the model architecture, which can be a long short-term memory network (LSTM, Long Short-Term Memory). The four-dimensional position encoding vector is generated by a sine / cosine function and one-hot encoding, the first two dimensions encode the ordinal relationship of the decision options (such as the number of layers), the third dimension identifies the decision type (CNN / GRU / connection / attention), and the fourth dimension normalizes the numerical range. The macro strategy network (MLP) outputs discrete architecture based on the four-dimensional vector, and the micro strategy network (long short-term memory network (LSTM, Long Short-Term Memory)) further generates continuous parameters to realize hierarchical optimization. The decision dimensions of the macro strategy network (processed by MLP) are shown in Table 1.

[0071] Table 1

[0072]

[0073] The decision dimensions of the micro strategy network (processed by LSTM) are shown in Table 2.

[0074] Table 2

[0075]

[0076] The above step of obtaining the model architecture by using the macro strategy network and obtaining the model parameters by using the micro strategy network includes:

[0077] First, obtain the model architecture by using the macro strategy network.

[0078] Specifically, the environment state is input into the macro strategy network for calculation, and the number of CNN layers, the number of GRU layers, the connection type and the attention type are output.

[0079] For example, the macro strategy network is a multi-layer perception, the environment state is input into the multi-layer perception, the multi-layer perception calculates the environment state, and a Gumbel-Softmax function is used to sample the calculation result, and the model architecture is output.

[0080] The macro strategy network is used to determine the basic architecture configuration of the model, and the implementation process first receives an environment state composed of four independent four-dimensional position encoding vectors, each vector corresponds to an architecture decision dimension (the number of CNN layers, the number of GRU layers, the connection type, and the attention type), and each dimension input is processed through a shared feature extraction layer (linear layer + ReLU activation); then in the decision generation stage, an independent decision head network (fully connected layer) is configured for each decision dimension, a Gumbel-Softmax sampling method with added Gumbel noise is used to ensure exploration during training, a temperature parameter is used to control the sampling sharpness and keep the gradient propagable to realize end-to-end training, and during inference, the architecture option with the maximum probability is directly selected; finally, the output format is a four-tuple [number of CNN layers, number of GRU layers, connection type, attention type], and each dimension option comes from a predefined search space.

[0081] Second, the micro strategy network is used to obtain the model parameters.

[0082] Specifically, the model architecture is input into the micro strategy network for calculation, and the number of CNN channels, the size of the convolution kernel and the number of GRU hidden units are output.

[0083] For example, the micro strategy network is a long short-term memory network, the model architecture is input into the long short-term memory network, the probability distribution of the model parameters is generated through linear transformation and Softmax activation function in the network, and the final output model parameters are obtained according to the probability distribution. The probability distribution describes the generation probability of multiple model parameters, and the final output model parameters are selected according to the generation probability.

[0084] It should be noted that the output of the macro strategy network is a discrete action obtained by Gumbel-Softmax sampling, the output of the micro strategy network is a continuous parameter obtained by linear transformation and Softmax activation, and the decision process of the macro strategy network to generate the model architecture and the decision process of the micro strategy network to generate the model parameters satisfy Markov property, and the expression is:

[0085] π θ (a t |s)=π macro (o|s)·π micro (a|o,s)

[0086] wherein, π θ denotes the action probability distribution, a t denotes the complete action (i.e. model architecture and model parameters), π macro denotes the macro action probability distribution, π micro denotes the micro action probability distribution, s denotes the environment state, o denotes the macro action (i.e. model architecture), and a denotes the micro action (i.e. model parameters).

[0087] For example, the connection type and attention type in the model architecture correspond to the CNN layer number and GRU layer number in the model architecture, and the model parameters. For example, if the CNN layer number in the model architecture is 2 and the GRU layer number is 2, the CNN channel number in the model parameters includes the channel number of the first CNN layer and the channel number of the second CNN layer, the convolution kernel size includes the convolution kernel size of the first CNN layer and the convolution kernel size of the second CNN layer, and the GRU hidden unit number includes the GRU hidden unit number of the first GRU layer and the GRU hidden unit number of the second GRU layer. For example, if the CNN layer number is 2, the GRU layer number is 2, and there are three connections, the connection type in the connection type includes the connection type of the three connections, such as residual connection, residual connection, and dense connection, which correspond to the connection between the first CNN layer and the second CNN layer, the connection between the second CNN layer and the first GRU layer, and the connection between the first GRU layer and the second GRU layer, respectively. The attention type includes four attention types, such as self-attention, channel attention, channel attention, and self-attention, which correspond to the attention type in the first CNN layer, the attention type in the second CNN layer, the attention type in the first GRU layer, and the attention type in the second GRU layer, respectively.

[0088] Step 13, constructing the bearing life prediction model based on the model architecture and the model parameters.

[0089] Specifically, the N CNN layers and the M GRU layers are connected in sequence according to the connection type, and the attention mechanism in each CNN layer and each GRU layer is configured according to the attention type. Each CNN layer is configured according to the CNN channel number and the convolution kernel size, and each GRU layer is configured according to the GRU hidden unit number, to obtain the bearing life prediction model.

[0090] N denotes the CNN layer number, M denotes the GRU layer number, the input end of the first CNN layer is the input end of the bearing life prediction model, and the output end of the last GRU layer is the output end of the bearing life prediction model.

[0091] It should be noted that the CNN layer includes a convolution block, a pooling block and an attention block connected in sequence, the GRU layer includes a recurrent block and an attention block connected in sequence, the operation of the attention block in the CNN layer and the GRU layer is an attention mechanism, the attention mechanism is set to the corresponding attention type; the convolution block in the CNN layer is a convolution operation, the number of channels of the convolution operation in each CNN layer is set to the corresponding CNN channel number, and the convolution kernel of the convolution operation is set to the corresponding convolution kernel size; the recurrent block in the GRU layer is a gated recurrent unit (GRU), and the number of hidden units therein is set to the GRU hidden unit number.

[0092] For example, the number of CNN layers is 2, the number of GRU layers is 2, and the connection type is residual connection, residual connection and dense connection, then the 2 CNN layers and the 2 GRU layers are connected in sequence according to the order of residual connection, residual connection and dense connection. The attention mechanism in the CNN layer is a channel attention mechanism, and the attention mechanism in the GRU layer is a self-attention mechanism, and the expression is:

[0093] ConvBlock i =ReLU(Conv1D(x))⊙σ(ChannelAttention(x))

[0094]

[0095] wherein, Convblock i represents the calculation result of the CNN layer, x represents the input data of the CNN layer, ReLU represents the ReLU activation function, Conv1D represents the convolution operation, σ represents the Sig activation function, ChannelAttention represents the channel attention mechanism, and h t represents the calculation result of the GRU layer, GRU represents the gated recurrent unit operation, x t represents the input data of the GRU layer, h t-1 represents the calculation result of the previous GRU layer, represents a dimension matching function, II() represents an indicator function, C type represents the connection type, and residual represents the residual connection.

[0096] Step 14, using the bearing life prediction model, performing life prediction according to each fault vibration signal, and constructing a multi-objective reward function based on the inference time corresponding to all life prediction results.

[0097] The multi-objective reward function is used to describe the calculation efficiency of the bearing life prediction model. The inference time is the time from the start of the calculation of the bearing life prediction model to the output of the life prediction result.

[0098] Specifically, the signal features of each fault vibration signal are extracted, and the bearing life prediction model is used to calculate each signal feature to obtain the high-speed train bearing life prediction result corresponding to each fault vibration signal (used to describe the difference between the end time of the time period corresponding to the sampled fault vibration signal and the time when the high-speed train bearing cannot be used, i.e. the remaining life, which can be a time length, such as 190 days), and record the inference time of the bearing life prediction model for calculating each signal feature. The multi-objective reward function is constructed based on the inference time corresponding to all life prediction results.

[0099] The multi-objective reward function is:

[0100]

[0101] Wherein, R represents the value of the multi-objective reward function, α, β, γ represent the weight coefficients, MSE represents the mean square error (calculated according to all high-speed train bearing life prediction results), θ represents the model parameter quantity, θ target represents the target parameter quantity, t inf represents the average value of the inference time corresponding to all life prediction results, t ref represents the reference inference time.

[0102] It should be noted that the signal features include the kurtosis, entropy, fractal dimension, peak factor, pulse factor, crest factor, energy ratio, spectral flatness, mean, variance, skewness, peak vibration, and effective value vibration of the signal, which can be obtained by analyzing the fault vibration signal using functions in matlab and other software.

[0103] Step 15, based on the multi-objective reward function, the bearing life prediction model is optimized to obtain the optimized bearing life prediction model, and the optimized bearing life prediction model is used to predict the life of the high-speed train bearing to be predicted.

[0104] The above-mentioned high-speed train bearing to be predicted is a high-speed train bearing that needs to be life predicted.

[0105] In some embodiments of the present application, the step of optimizing the bearing life prediction model based on the multi-objective reward function to obtain the optimized bearing life prediction model, and using the optimized bearing life prediction model to predict the life of the high-speed train bearing to be predicted comprises:

[0106] The first step is to determine whether the multi-objective reward function value is greater than the preset reward value.

[0107] If so, the bearing life prediction model is used as the optimized bearing life prediction model.

[0108] Otherwise, the macro policy network and the micro policy network are updated based on the attention mechanism, and the steps of using the macro policy network to obtain the model architecture and using the micro policy network to obtain the model parameters are returned.

[0109] It should be noted that the above-mentioned steps of updating the macro-policy network and the micro-policy network based on the attention mechanism include: calculating the policy update gradient based on the attention mechanism; and updating the parameters in the macro-policy network and the micro-policy network according to the policy update gradient.

[0110] Specifically, through the formula:

[0111]

[0112] w i =softmax(f θ (|o i ;a i |))

[0113] Calculate policy update gradients

[0114] Among them, A(s,a i ) represents the advantage function, s represents the environmental state, π θ represents the action probability distribution, w i represents the attention weight, represents the gradient operator, i.e. the partial derivative of the policy network parameter θ, f θ Represents the parameterized function of the policy network, defined by the parameters θ, a i represents the i-th action, which represents the discrete or continuous action selected by the policy network in state s (such as the number of CNN layers, the number of GRU hidden units, etc.), o i Represents the i-th observation value, usually refers to the environment's response to action a i Feedback (such as reward value or new state).

[0115] Exemplarily, the parameters in the macro-policy network and the micro-policy network are updated using a gradient ascent method based on the policy update gradient. The advantage function is the difference between the multi-objective reward function value and the baseline value. The baseline value is the value of a preset function (usually a state-value function). When the parameters in the macro-policy network and the micro-policy network are updated, the baseline value can be updated synchronously.

[0116] Secondly, the optimized bearing life prediction model is used to predict the life of the high-speed train bearing to be predicted.

[0117] Specifically, the fault vibration signal of the high-speed train bearing to be predicted is obtained, and the fault vibration signal is input into the optimized bearing life prediction model for calculation to obtain the high-speed train bearing life prediction result of the high-speed train bearing to be predicted.

[0118] It is worth mentioning that the bearing life prediction model is constructed by using the macro strategy network and the micro strategy network, the architecture and the parameters of the model can be calculated respectively, the hierarchical construction of the bearing life prediction model is realized, the calculation efficiency of the constructed bearing life prediction model is quantified based on the inference time, the multi-objective reward function is constructed, the prediction accuracy and the inference time of the bearing life prediction model are effectively optimized by using the multi-objective reward function to optimize the bearing life prediction model, and the accuracy of the life prediction of the high-speed train bearing is improved.

[0119] The method of the present application will be exemplarily described below in conjunction with a specific example.

[0120] Taking the working condition one (the rotating speed is 2100r / min, and the radial force is 12kN) of the data set XJTU-SY bearing data set provided by Xi'an Jiaotong University and Shenyang University of Technology as an example. First, the original data in the 5 bearing monitoring CSV files is loaded, the remaining useful life (RUL, Remaining Useful Life) is extracted as a label, and the sliding window method (the window size is 5) is used to convert the time series data into a supervised learning format; then, the data of each bearing is standardized, and for the working condition one, the bearings 1, 2, 4 and 5 are combined as a training set, and the bearing 3 is reserved as a test set.

[0121] The macro strategy network receives a four-dimensional position encoding vector [0.0, 1.0, 2.0, 3.0] as an environment state, first maps the position encoding to a feature space through a Linear(1→16) linear layer, and then applies a ReLU activation function to enhance the non-linear expression capability. The decision head after that contains four parallel output layers: a CNN layer number decision layer (Linear(16→3), corresponding to [1, 2, 3] layers), a GRU layer number decision layer (Linear(16→2), corresponding to [1, 2] layers), a connection type decision layer (Linear(16→3), corresponding to ['none','residual', 'dense']), and an attention type decision layer (Linear(16→3), corresponding to ['none','self', 'channel']). The sampling mechanism uses Gumbel-Softmax with a temperature parameter of 0.1 for differentiable sampling, and uses a hard sampling strategy during training and directly selects the maximum probability decision during inference. Among them, none means nothing, residual is residual connection, dense is dense connection, self is self-attention, and channel is channel attention.

[0122] The micro strategy network receives a 32-dimensional embedding vector (4 decisions x 8 dimensions) output by the macro strategy network, captures the temporal dependence between decisions through an LSTM memory unit with a single layer of 32-dimensional hidden units, and then outputs specific parameters through five parameter decision heads: CNN channel decision (Linear(32→3), options [16, 32, 64]), convolution kernel decision (Linear(32→3), options [3, 5, 7]), pooling type decision (Linear(32→3), options ['none','max', 'avg']), GRU hidden unit decision (Linear(32→3), options [32, 64, 128]), and GRU bidirectionality decision (Linear(32→2), options [False, True]). Among them, max is max pooling, and avg is average pooling.

[0123] Multi-objective reward calculation integrates indicators in three dimensions of model accuracy, complexity, and efficiency to generate the final reward value using a weighted combination method. The accuracy weight α is set to 10.0, the complexity weight β is set to 1.0, and the efficiency weight γ is set to 0.5. The target parameter quantity is set to 20000, and the reference inference time is set to 10ms.

[0124] The attention weighting mechanism evaluates the importance of macro and micro decisions and weights the gradients to improve the search efficiency of the model. The 20-dimensional input vector containing 4 macro decisions and 16 micro parameters is first processed by a Linear(20→32) linear layer and a ReLU activation function to build a hidden layer, which enhances the feature expression capability. Then, a Linear(32→20) linear layer and a Softmax activation function are used to output the attention weights of the 20 decision dimensions. In the gradient update phase, a learnable baseline value (initialized to 0) is introduced to estimate the state value, and the advantage of the current decision relative to the baseline is calculated by the advantage function.

[0125] Based on the decision results of the macro and micro strategy networks, the actual network architecture (i.e., the bearing life prediction model) is generated according to the following configuration rules. The CNN part supports a maximum of 3 layers deep, each layer can choose a channel number from [16, 32, 64] and a convolution kernel size from [3, 5, 7], and uses the same padding mode to ensure the feature map size unchanged. The GRU part supports a maximum of 2 layers structure, the hidden unit number can be selected from [32, 64, 128], and whether to enable the bidirectional mode can be selected. In terms of connection mode, when the macro strategy chooses residual connection, 1×1 convolution is automatically used to adjust the dimension to match the input and output; when dense connection is selected, feature map truncation or zero padding operation is used to realize dimension alignment. In terms of attention module, if channel attention is selected, the default compression ratio (reduction_ratio=16) is used to balance the calculation overhead and feature screening capability; if self-attention is selected, the scale=sqrt(dim) scaling mechanism is applied to stabilize the gradient propagation. The whole network construction process dynamically combines each component according to the strategy network output to form a neural network architecture that meets the needs of multi-objective optimization.

[0126] The strategy network (i.e., macro strategy network and micro strategy network) is gradually optimized through 60 rounds of neural network architecture search (NAS) iteration strategies. In each iteration, the model architecture is first sampled, and then the corresponding bearing life prediction model is instantiated, trained for 5 epochs with a batch size of 16 (initial learning rate 0.0005). After training, the model mean square error (MSE) is calculated on the validation set, and the model inference time is accurately measured through the GPU Event mechanism for managing and synchronizing GPU tasks. Based on these indicators, a multi-objective reward function is calculated, and the parameters are updated using gradient clipping (max_norm = 1.0) and a learning rate of 0.0005. After each iteration, the current optimal architecture is saved based on the Pareto optimality principle. In the final training phase, the best-performing model is selected from all saved architectures, trained for a complete 20 epochs with a learning rate of 0.0003, and an early stopping mechanism is enabled (training is stopped when the validation loss does not decrease for 3 consecutive rounds) to avoid overfitting, thus obtaining a final model that balances performance and efficiency.

[0127] The structure of the optimal bearing life prediction model under working condition 1 is shown in Figure 2 , which includes two CNN layers and two GRU layers connected in sequence. The two CNN layers have convolutions of Conv 64x3 (64 channels, kernel size 3) and Conv 64x5 (64 channels, kernel size 5), respectively. An average pooling operation is performed between the two convolutions, and the subsequent features enter the two GRU layers, which are GRU 32 (hidden units 32) unidirectional and GRU 32 bidirectional, respectively. The final output is the prediction result.

[0128] The importance evaluation results of different decision types are shown in Figure 3 , where the horizontal axis represents the importance score, and the vertical axis represents different decision types, including pooling type, attention type, connection type, CNN layer number, GRU layer number, GRU bidirectional, CNN channel number, convolution kernel size, and GRU hidden unit. Among them, the importance score of the pooling type is the highest, close to 0.06; the importance score of the GRU hidden unit is the lowest, about 0.03.

[0129] The quantitative indicators of the original bearing life prediction model and the NAS-optimized bearing life prediction model are shown in Figure 4 , Figure 4 a is the comparison of mean square error, the horizontal axis is the original model and the NAS model (i.e., the NAS-optimized bearing life prediction model), and the vertical axis represents the value of the mean square error, Figure 4 b is the comparison of root mean square error, the horizontal axis is the original model and the NAS model, and the vertical axis represents the value of the root mean square error, Figure 4c is the average absolute error comparison, the horizontal axis is the original model and the NAS model, and the vertical axis represents the value of the average absolute error, Figure 4 d is the parameter quantity comparison, the horizontal axis is the original model and the NAS model, and the vertical axis represents the parameter quantity, Figure 4 e is the inference time comparison, the horizontal axis is the original model and the NAS model, and the vertical axis represents the inference time in milliseconds ms, Figure 4 f is the prediction result comparison, the horizontal axis represents the sample, and the vertical axis represents the value of the life, the solid line is the true value, which represents the true remaining life of the sample, and the two dotted lines respectively represent the remaining life predicted by the original model and the remaining life predicted by the NAS model. In the mean square error comparison, the original model is 0.0023, the NAS model is 0.0003, and the improvement is 89.2%; in the root mean square error comparison, the original model is 0.0484, the NAS model is 0.0159, and the improvement is 67.1%; in the average absolute error comparison, the original model is 0.0402, the NAS model is 0.0128, and the improvement is 68.1%. In terms of inference time, the original model is 1.2800ms, the NAS model is 0.8571ms, and the improvement is 33.0%; in terms of parameter quantity, the original model is 107169, the NAS model is 45441, and the improvement is 57.6%. It can be seen from the prediction result comparison that the NAS model prediction is closer to the true value.

[0130] The self-attention heat map output by the GRU layer is as shown in Figure 5 The horizontal and vertical axes are both sequence positions, i.e. time step sequence positions of the input signal, corresponding to the time sequence dimension of the hidden state output by the GRU layer, and different pixel values represent different self-attention weights. As can be seen from the figure, the self-attention weights of most positions are concentrated between 0.135-0.155. By visualizing the heat map, the attention degree of the model to the features at different time points can be analyzed. The uniformly distributed weights indicate that there is no obvious local focusing of the fault features in the time sequence.

[0131] The method provided in the present application combines the explainable mechanism, the dynamic connection structure and the hybrid attention mechanism, and is used for optimizing the bearing life prediction model in the bearing life prediction task. A hierarchical control mechanism with explainability is designed; a dynamic connection and hybrid attention collaborative optimization strategy is proposed; and the importance evaluation of architecture decision is realized. Efficient architecture search is realized through a hierarchical strategy network, the attention weighting mechanism enhances the update strength of key decisions, improves the model performance while providing explainability; the multi-objective reward function balances the prediction accuracy and resource consumption, and meets the actual needs of the high-speed train scene.

[0132] The life prediction device for the high-speed train bearing provided in the present application is exemplarily described below.

[0133] As Figure 6As shown, an embodiment of the present application provides a life prediction device for a high-speed train bearing. The life prediction device 600 for a high-speed train bearing includes:

[0134] A first acquisition module 601 is used to acquire multiple fault vibration signals of high-speed train bearings;

[0135] A second acquisition module 602 is used to acquire a model architecture using a macro-strategy network and to acquire model parameters using a micro-strategy network;

[0136] A construction module 603 is used to construct a bearing life prediction model based on the model architecture and model parameters;

[0137] Life prediction module 604 is used to use the bearing life prediction model to predict the life of each fault vibration signal and construct a multi-objective reward function based on the inference time corresponding to all life prediction results; the multi-objective reward function is used to describe the computational efficiency of the bearing life prediction model;

[0138] The optimization module 605 is used to optimize the bearing life prediction model based on the multi-objective reward function to obtain the optimized bearing life prediction model, and use the optimized bearing life prediction model to predict the life of the high-speed train bearing to be predicted.

[0139] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0140] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0141] like Figure 7 As shown, an embodiment of the present application provides a terminal device, and the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 7The apparatus D1 00 is configured to execute any of the above method embodiments, and comprises at least one processor D100, a memory D101, and a computer program D102 stored in the memory D101 and run on the at least one processor D100. The processor D100 implements the steps in any of the above method embodiments when executing the computer program D102.

[0142] Specifically, the processor D100, when executing the computer program D102, obtains a plurality of fault vibration signals of a high-speed train bearing, then obtains a model architecture by using a macro strategy network, obtains model parameters by using a micro strategy network, constructs a bearing life prediction model based on the model architecture and the model parameters, then performs life prediction according to each fault vibration signal by using the bearing life prediction model, constructs a multi-objective reward function based on reasoning time corresponding to all life prediction results, finally optimizes the bearing life prediction model based on the multi-objective reward function, obtains an optimized bearing life prediction model, and performs life prediction on a high-speed train bearing to be predicted by using the optimized bearing life prediction model. Wherein, the bearing life prediction model is constructed by using the macro strategy network and the micro strategy network, which can respectively perform strategy calculation on the architecture and the parameters of the model, realize hierarchical construction of the bearing life prediction model, the multi-objective reward function is constructed based on the reasoning time, which can quantify the calculation efficiency of the constructed bearing life prediction model, the bearing life prediction model is optimized by using the multi-objective reward function, which can effectively optimize the prediction accuracy and reasoning time of the bearing life prediction model, and further improve the accuracy of life prediction of the high-speed train bearing.

[0143] The processor D100 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or can also be any conventional processor.

[0144] The storage D101 can be an internal storage unit of the terminal device D10 in some embodiments, such as a hard disk or a memory of the terminal device D10. The storage D101 can also be an external storage device of the terminal device D10 in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like. Further, the storage D101 can include both an internal storage unit and an external storage device of the terminal device D10. The storage D101 is used to store an operating system, an application program, a boot loader, data, and other programs, such as program codes of the computer program, and the like. The storage D101 can also be used to temporarily store data that has been output or will be output.

[0145] The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in each of the above method embodiments.

[0146] The computer program product, when executed on a terminal device, causes the terminal device to implement the steps in each of the above method embodiments.

[0147] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the present application implements all or part of the processes in the above embodiments, which can be completed by instructing related hardware through a computer program. The computer program can be stored in a computer readable storage medium, and the computer program, when executed by a processor, can implement the steps in each of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer readable medium at least includes any entity or device capable of carrying the computer program code of the life prediction method and device of a high-speed train bearing, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. For example, a U disk, a mobile hard disk, a magnetic disk or an optical disk, and the like. In some jurisdictions, according to legislation and patent practice, the computer readable medium cannot be an electrical carrier signal and a telecommunications signal.

[0148] In the above embodiments, the description of each embodiment is focused on, and the part not described or recorded in a certain embodiment can be referred to the relevant description of other embodiments.

[0149] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0150] The above is the preferred embodiment of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the principles described in the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.

Claims

1. A method for predicting the life of a high-speed train bearing, characterized in that: include: Obtain multiple fault vibration signals of high-speed train bearings; Use the macro-policy network to obtain the model architecture, and use the micro-policy network to obtain the model parameters; Constructing a bearing life prediction model based on the model architecture and the model parameters; Using the bearing life prediction model, life prediction is performed based on each of the fault vibration signals, and a multi-objective reward function is constructed based on the inference time corresponding to all life prediction results; the multi-objective reward function is used to describe the computational efficiency of the bearing life prediction model; The bearing life prediction model is optimized based on the multi-objective reward function to obtain an optimized bearing life prediction model, and the optimized bearing life prediction model is used to predict the life of the high-speed train bearing to be predicted.

2. The life prediction method according to claim 1, characterized in that: The model architecture includes the number of CNN layers, GRU layers, connection type, and attention type; The method of obtaining the model architecture by using the macro strategy network includes: The environment state is input into the macro-strategy network for calculation, and the number of CNN layers, the number of GRU layers, the connection type and the attention type are output; the environment state includes the four-dimensional position coding vector corresponding to the number of CNN layers, the four-dimensional position coding vector corresponding to the number of GRU layers, the four-dimensional position coding vector corresponding to the connection type and the four-dimensional position coding vector corresponding to the attention type.

3. The life prediction method according to claim 2, characterized in that: The model parameters include the number of CNN channels, convolution kernel size, and the number of GRU hidden units; The method of obtaining model parameters by using a micro-strategy network includes: The model architecture is input into the micro-strategy network for calculation, and the number of CNN channels, convolution kernel size, and number of GRU hidden units are output.

4. The life prediction method according to claim 3, characterized in that: The constructing of a bearing life prediction model based on the model architecture and the model parameters includes: According to the connection type, N CNN layers and M GRU layers are connected in sequence, and the attention mechanism in each CNN layer and each GRU layer is configured according to the attention type. Each CNN layer is configured according to the number of CNN channels and the convolution kernel size, and each GRU layer is configured according to the number of GRU hidden units to obtain the bearing life prediction model; N represents the number of CNN layers, M represents the number of GRU layers, the input end of the first CNN layer is the input end of the bearing life prediction model, and the output end of the last GRU layer is the output end of the bearing life prediction model.

5. The life prediction method according to claim 1, characterized in that: The multi-objective reward function is: Among them, R represents the value of the multi-objective reward function, α, β, γ represent weight coefficients, MSE represents mean square error, θ represents the model parameter, θ target represents the target parameter, t inf represents the average inference time corresponding to all life prediction results, t ref Represents the reference inference time.

6. The life prediction method according to claim 1, characterized in that: The optimizing the bearing life prediction model based on the multi-objective reward function to obtain an optimized bearing life prediction model includes: Determining whether the multi-objective reward function value is greater than a preset reward value; If yes, the bearing life prediction model is used as the optimized bearing life prediction model; Otherwise, the macro policy network and the micro policy network are updated based on the attention mechanism, and the process returns to the step of obtaining the model architecture using the macro policy network and obtaining the model parameters using the micro policy network.

7. The life prediction method according to claim 6, characterized in that: The updating of the macro-strategy network and the micro-strategy network based on the attention mechanism includes: Calculate the policy update gradient based on the attention mechanism; Parameters in the macro policy network and the micro policy network are updated according to the policy update gradient.

8. The life prediction method according to claim 7, characterized in that: The attention-based strategy update gradient calculation includes: By formula: w i =softmax(f θ (|o i ;a i |)) Calculate policy update gradients Among them, A(s,a i ) represents the advantage function, s represents the environmental state, π θ represents the action probability distribution, w i represents the attention weight, represents the gradient operator, f θ Represents the parameterized function of the policy network, a i represents the i-th action, o i represents the i-th observation value.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the life prediction method for a high-speed train bearing according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the life prediction method for a high-speed train bearing according to any one of claims 1 to 8 is implemented.