A Predictive Maintenance Method for a Mechanical Device

The multi-sensory features of mechanical equipment state data are extracted through the state transition network model, and combined with real-time conversion-near-end strategy to optimize the network model, the problem of over-complexity of predictive maintenance models in the existing technology is solved, and simplified real-time maintenance decisions and efficient predictive maintenance are achieved.

CN119515353BActive Publication Date: 2025-06-27HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411627614.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-06-27
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

The prior art when integrating data-driven mechanical equipment residual service life (RUL) prediction with maintenance planning, the model is too cumbersome, resulting in complex predictive maintenance.

Method used

The state transition network model is used to extract the multi-sensory features of mechanical equipment state data, and combined with real-time conversion-near-end strategy to optimize the network model to realize real-time ordering and maintenance decisions.

Benefits of technology

Reduces the complexity of predictive maintenance, simplifies algorithms through effective feature extraction and decision-making frameworks, and improves the efficiency and accuracy of predictive maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119515353B_ABST
    Figure CN119515353B_ABST
Patent Text Reader

Abstract

The present invention discloses a predictive maintenance method for mechanical equipment, which relates to the technical field of aero-engine maintenance. The method includes: obtaining the state data of the mechanical equipment at multiple moments, and inputting the state data into a state conversion network model, extracting multi-sensory features of different scales of the state data through the state data to obtain a belief state; the belief state is a feature representation of the state data; then inputting the belief state into a pre-trained real-time conversion - proximal policy optimization network model to obtain the predictive maintenance result of the mechanical equipment. This method can reduce the complexity of the predictive maintenance of mechanical equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aero-engine maintenance, and particularly to a predictive maintenance method for mechanical equipment driven by real-time sensors. Background Art

[0002] Currently, the application scope of Prognostics Health Management (PHM) is very wide, and it can be applied to fields such as aerospace, military, industrial manufacturing, energy, and transportation. For example, for a Boeing 787, approximately 1,000 sensors record the real-time operating parameters of the engine, and these parameters indicate the potential state of the flight equipment. These data are the basis for predicting the Remaining Useful Life (RUL) and the predictive aircraft maintenance plan. In the aerospace field, PHM can help aircraft and spacecraft monitor their health status in real time during flight, predict potential failures, and take corresponding measures in a timely manner to ensure flight safety and reliability. However, the purpose of predicting the RUL of mechanical equipment is to provide a scientific basis for realizing the dynamic real-time maintenance of mechanical equipment. A simple RUL algorithm is not complete because it lacks a maintenance plan. After obtaining the RUL of each state of the mechanical equipment, it is necessary for the staff to calculate the optimal ordering time and maintenance time, and such manual operations will undoubtedly cause meaningless cost losses. PHM should integrate the two, so that it can detect the health status of the equipment in real time or near real time, predict potential failures and RUL, and combine the maintenance cost and the failure cost for system real-time maintenance and real-time decision-making.

[0003] Currently, only a few studies have integrated data-driven RUL prediction into the maintenance plan. However, for the method that combines real-time ordering and real-time maintenance, there are still deficiencies, and the existing model applications are too cumbersome. Because the existing models often retain all the RULs, but the data of the mechanical equipment monitored during long-term operation is a relatively long time series. The early operation performance of the equipment is relatively stable, and the performance degradation is not obvious. The RUL of the mechanical equipment decreases with the increase of the service time. Especially in the early stage of service, extremely complex algorithms are required to obtain the accurate remaining useful life, resulting in the complexity of the predictive maintenance of mechanical equipment. Summary of the Invention

[0004] Based on this, it is necessary to provide a predictive maintenance method for mechanical equipment aiming at the above technical problems, and this method can reduce the complexity of the predictive maintenance of mechanical equipment.

[0005] The present invention provides a predictive maintenance method for mechanical equipment, including:

[0006] Obtaining the state data of the mechanical equipment at multiple moments;

[0007] Input the status data into the state conversion network model, extract multi-sensory features of different scales of the status data through the state conversion network model, and obtain the belief state; the belief state is the feature representation of the status data.

[0008] Input the belief state into the pre-trained real-time conversion-proximal policy optimization network model to obtain the predictive maintenance result of the mechanical equipment.

[0009] Preferably, obtain the status data of the mechanical equipment at multiple moments, including:

[0010] Obtain the original status data collected in real time by multiple sensors on the mechanical equipment at multiple moments;

[0011] Perform normalization processing on the original status data to obtain the status data of the mechanical equipment at multiple moments.

[0012] Preferably, the state conversion network model includes a convolutional neural network, a residual neural network, and a self-attention mechanism module. Extract multi-sensory features of different scales of the status data through the state conversion network model to obtain the belief state, including:

[0013] Input the status data into the convolutional neural network to obtain the high-order features of the status data;

[0014] Input the high-order features of the status data into the residual neural network to obtain the high-order optimized features of the status data;

[0015] Input the high-order optimized features of the status data into the self-attention mechanism module to obtain the high-order belief state;

[0016] Input the high-order belief state of the status data into the linear neural network to obtain the belief state.

[0017] Preferably, the state conversion network model further includes a bidirectional gated recurrent unit, and the method further includes:

[0018] Input the belief state into the bidirectional gated recurrent unit to obtain the remaining service life of the mechanical equipment.

[0019] Preferably, the real-time conversion-proximal policy optimization network model includes a real-time ordering framework and a real-time maintenance framework; input the belief state into the pre-trained real-time conversion-proximal policy optimization network model to obtain the predictive maintenance result of the mechanical equipment, including:

[0020] Input the belief state into the real-time ordering framework to obtain the ordering strategy of the mechanical equipment, and determine the ordering strategy as the predictive ordering result;

[0021] Input the belief state into the real-time maintenance framework to obtain the maintenance strategy of the mechanical equipment, and determine the predictive maintenance result with the maintenance strategy.

[0022] Preferably, the construction process of the real-time conversion-proximal policy optimization network model includes:

[0023] Obtain the sample belief state of the sample mechanical equipment;

[0024] Input the sample belief state into the policy network to obtain the sample predictive maintenance result of the sample mechanical equipment;

[0025] Input the current state and the sample predictive maintenance result into the value network, and determine the evaluation value through the reward function in the value network;

[0026] Determine the policy loss according to the sample predictive maintenance result and the sample true maintenance result, and determine the value loss according to the evaluation value;

[0027] Update the parameters of the policy network according to the policy loss, and update the parameters of the value network according to the value loss;

[0028] Determine the updated policy network and value network as the real-time conversion-proximal policy optimization network model.

[0029] Preferably, inputting the sample belief state into the policy network to obtain the sample predictive maintenance result of the sample mechanical equipment includes:

[0030] Input the sample belief state into the policy network to obtain the action data of the sample mechanical equipment; the action data is the probability value of each behavior of the sample mechanical equipment; the behaviors include continued use, maintenance, and ordering;

[0031] Determine the sample predictive maintenance result through the stochastic policy according to the probability value of each behavior.

[0032] Preferably, the method further includes:

[0033] Obtain the test belief state of the test mechanical equipment;

[0034] Input the test belief state into the real-time conversion-proximal policy optimization network model to obtain the action data of the test mechanical equipment; the action data is the probability value of each behavior of the test mechanical equipment; the behaviors include warning ordering, shutdown maintenance, and continued operation;

[0035] Determine the test sample predictive maintenance result through the stochastic policy;

[0036] Determine the test sample true maintenance result through the deterministic policy;

[0037] Determine the performance evaluation result of the real-time conversion-proximal policy optimization network model according to the predictive maintenance result of the test sample and the true maintenance result of the test sample.

[0038] Preferably, the real-time conversion-proximal policy optimization network model includes a real-time ordering framework, a reward function, including:

[0039]

[0040] where r t is the evaluation value, t is the service time of the sample mechanical equipment, a t is the sample predictive maintenance result of the sample mechanical equipment, C is the total ordering cost of the sample mechanical equipment, T is the mechanical equipment life, L is the advance ordering period, SST is the state windowing, Order represents warning ordering, and Continue represents continued operation.

[0041] Preferably, the real-time conversion-proximal policy optimization network model includes a real-time maintenance framework, a reward function, including:

[0042]

[0043] where r t is the evaluation value, t is the service time of the sample mechanical equipment, a t is the sample predictive maintenance result of the sample mechanical equipment, cr is the replacement cost of the sample mechanical equipment, cf is the reset cost, SST is the state windowing, Replace represents warning maintenance, and Continue represents continued operation.

[0044] The present invention provides a predictive maintenance device for mechanical equipment, including:

[0045] An acquisition module, configured to acquire state data of the mechanical equipment at multiple moments;

[0046] An extraction module, configured to input the state data into a state conversion network model, extract multi-sensory features of different scales of the state data through the state data, and obtain a belief state; the belief state is a feature representation of the state data;

[0047] A prediction module, configured to input the belief state into a pre-trained real-time conversion-proximal policy optimization network model to obtain a predictive maintenance result of the mechanical equipment.

[0048] The present invention provides a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned predictive maintenance method for mechanical equipment is implemented.

[0049] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the predictive maintenance method of the above-mentioned mechanical equipment is implemented.

[0050] The above at least one technical solution adopted by the present invention can achieve the following beneficial effects:

[0051] By extracting multi-sensory features of different scales through the state transition network model, it is possible to obtain a rich feature representation of the mechanical equipment state data at a relatively small computational time cost. Such feature extraction can reduce the amount of raw data required for subsequent processing and lower the computational load. Based on effective feature extraction, the input dimension of the real-time conversion-proximal policy optimization network model is reduced, and the predictive maintenance result is directly determined through the pre-trained real-time conversion-proximal policy optimization network model, which overall reduces the complexity of the algorithm, thereby making the predictive maintenance process simpler. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where:

[0053] Figure 1 is a schematic flow chart of a predictive maintenance method for a mechanical equipment provided by the present invention;

[0054] Figure 2 is a state transition network framework model diagram provided by the present invention;

[0055] Figure 3 is a schematic flow chart of another predictive maintenance method for a mechanical equipment provided by the present invention;

[0056] Figure 4 is a schematic diagram of a simulation aeroengine;

[0057] Figure 5 is a schematic flow chart of another predictive maintenance method for a mechanical equipment provided by the present invention;

[0058] Figure 6 is a structural block diagram of a predictive maintenance device for a mechanical equipment provided by the present invention;

[0059] Figure 7 is a schematic diagram of a computer device for a predictive maintenance method for a mechanical equipment provided by the present invention.

[0060] Description of the reference numerals:

[0061] 1. Fan; 2. Combustion chamber; 3. Low-pressure rotor speed; 4. Low-pressure turbine; 5. Low-pressure compressor; 6. High-pressure compressor; 7. High-pressure rotor speed; 8. High-pressure turbine; 9. Nozzle. Detailed implementation manner

[0062] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0063] The predictive maintenance method of the mechanical equipment provided by the present invention can be executed by a server. The server can be a server set up on a business platform or a device such as a desktop computer or a laptop computer that can execute the solution of the present invention. For the convenience of description, the following only uses the server as the execution subject for description.

[0064] Figure 1 It is a flow diagram of a predictive maintenance method for a mechanical equipment in this specification, specifically including the following steps:

[0065] S101. Obtain the state data of the mechanical equipment at multiple moments.

[0066] The mechanical equipment can be an aeroengine, military equipment, industrial manufacturing equipment, transportation equipment, energy equipment, etc.

[0067] In an exemplary embodiment, obtaining the state data of the mechanical equipment at multiple moments includes: obtaining the original state data collected in real time by multiple sensors on the mechanical equipment at multiple moments, and performing normalization processing on the original state data to obtain the state data of the mechanical equipment at multiple moments.

[0068] There are multiple sensors on the mechanical equipment, and the multiple sensors are used to collect the state data of the mechanical equipment in real time. When predictive maintenance needs to be performed on the mechanical equipment, the server can obtain the original state data collected in real time by the sensors at multiple moments from the sensors. The original state data can be expressed as, X = [x1, x2,..., x LD T , where NC represents the number of sensors, and LD represents the length of the sensor data sequence. Among them represents the i-th sensor data in the k-th time period.

[0069] Input the original state data at each moment into the normalization algorithm, and perform normalization processing on the original state data at each moment through the normalization algorithm to obtain the state data at multiple moments. Among them, the normalization algorithm can adopt the Min-max scaling method for processing. ​

[0070] S102. Input the status data into the status conversion network model, extract multi-sensory features at different scales of the status conversion network model, and obtain the belief state; the belief state is the feature representation of the status data.

[0071] The multi-sensory features are data obtained by multiple sensors in multiple dimensions, used to represent the features of the system in multiple dimensions. For an aero-engine, multiple sensors are distributed at different positions of the engine to collect data at different positions of the engine, and the data of multiple sensors are integrated to represent the features of the engine in multiple dimensions. The belief state is a probabilistic representation of the state of the environment where the agent is located. The belief state synthesizes all past observation information and prior knowledge about the environmental dynamics, and is used to estimate the probability distribution of various true states that the environment may be in, and is the data representation after removing the noise of the original data. For an aero-engine, the original data of the sensor contains a lot of noise, and the potential degradation state of the engine cannot be directly observed. The probability distribution of the potential state of the engine can be estimated, and this probability distribution is the belief state of the engine.

[0072] As Figure 2 shown, Figure 2 is the structural schematic diagram of the status conversion network model. The status conversion network model includes a convolutional neural network, a residual neural network, a self-attention mechanism module, and a linear neural network module. Therefore, in an exemplary embodiment, extracting multi-sensory features at different scales of the status data through the status conversion network model to obtain the belief state includes: inputting the status data into the convolutional neural network to obtain the high-order features of the status data; and inputting the high-order features of the status data into the residual neural network to obtain the high-order optimized features of the status data; then inputting the high-order optimized features of the status data into the self-attention mechanism module to obtain the high-order belief state, and finally, inputting the high-order belief state of the status data into the linear neural network to obtain the belief state. Among them, the status conversion network model also includes a bidirectional gated recurrent unit, and inputting the belief state into the bidirectional gated recurrent unit to obtain the remaining useful life of the mechanical equipment.

[0073] In an exemplary embodiment, the self-attention mechanism module is a deep autoencoder with an attention strategy, used to selectively extract multi-sensory features at different scales, and the bidirectional gated recurrent unit module is used to comprehensively extract degradation features at different time scales. The status conversion network model fuses the degradation features and predicts the RUL.

[0074] During the training process of the state transition model, a Mean Squared Error (MSE) loss function is constructed to calculate the loss and backpropagate the error to update the neural network parameters of the TransStateNet (State Transition Network). The belief state of the TransStateNet is used as the input to the real-time conversion-proximal policy optimization network space. The belief state will maintain the same size as the original mechanical equipment data, i.e., ζ = x1, x2, …, x LD T ,

[0075] The TransStateNet is constructed to perform state transitions on the initial state of the original mechanical equipment. Compared with the original data, the TransStateNet extracts the degradation features of multi-sensor data at different time scales, eliminates the influence of redundant features, and reduces the difficulty of fitting the state distribution in reinforcement learning. The belief state of the TransStateNet is input into a real-time maintenance or real-time decision-making framework, which decides to continue operation, issue a warning, or stop for maintenance in real time based on the processed state. The pseudocode of the TransStateNet is as follows:

[0076]

[0077]

[0078] S103: Input the belief state into the pre-trained real-time conversion-proximal policy optimization network model to obtain the predictive maintenance result of the mechanical equipment.

[0079] The predictive maintenance results of the mechanical equipment include warning order, shutdown maintenance, and continued operation.

[0080] The real-time conversion-proximal policy optimization network model includes a real-time order framework and a real-time maintenance framework. Input the belief state into the real-time order framework to obtain the order policy of the mechanical equipment, and determine the order policy as the predictive order result; input the belief state into the real-time maintenance framework to obtain the maintenance policy of the mechanical equipment, and determine the maintenance policy as the predictive maintenance result.

[0081] The predictive order results include warning order and continued operation, and the predictive maintenance results include shutdown maintenance and continued operation.

[0082] ​In an exemplary embodiment, the construction process of the real-time conversion-proximal policy optimization network model includes: obtaining the sample belief state of the sample mechanical equipment, and inputting the sample belief state into the initial real-time conversion-proximal policy optimization network to obtain the sample predictive maintenance result of the sample mechanical equipment; inputting the current state and the sample predictive maintenance result into the value network, and determining the evaluation value through the reward function in the value network; determining the policy loss according to the sample predictive maintenance result and the sample true maintenance result, and determining the value loss according to the evaluation value; updating the parameters of the initial proximal policy network according to the policy loss, and updating the parameters of the value network according to the value loss; determining the updated initial proximal policy network model as the real-time conversion-proximal policy optimization network model.

[0083] Wherein, before constructing the real-time conversion-proximal policy optimization network model, the sample original state data respectively collected by multiple sensors on the sample mechanical equipment at multiple moments can be obtained first, and then the sample original state data is normalized to obtain the sample belief state of the sample mechanical equipment; the method of normalizing the sample original state data is the same as the method of normalizing the original state data in the above embodiment, and this embodiment will not be elaborated here.

[0084] Optionally, inputting the sample belief state into the initial proximal policy optimization network to obtain the sample predictive maintenance result of the sample mechanical equipment includes: inputting the sample belief state into the initial proximal policy network to obtain the action data of the sample mechanical equipment; the action data is the probability value of each behavior of the sample mechanical equipment; the behaviors include early warning ordering, shutdown maintenance, and continued operation; according to the probability value of each behavior, the sample predictive maintenance result is determined through a stochastic policy.

[0085] Inputting the sample belief state into the initial proximal policy network to obtain the action data of the sample mechanical equipment output by the initial proximal policy network.

[0086] Determining the sample predictive maintenance result through a stochastic policy according to the probability value of each behavior includes: randomly obtaining one behavior from all behaviors as the sample predictive maintenance result; for example, randomly taking continued operation as the sample predictive maintenance result of the sample mechanical equipment.

[0087] The proximal policy optimization network includes a real-time ordering framework and a real-time maintenance framework. Therefore, inputting the sample belief state into the initial proximal policy network to obtain the sample predictive maintenance result of the sample mechanical equipment, that is, inputting the sample belief state into the real-time ordering framework to obtain the sample predictive maintenance result of the sample mechanical equipment, and the sample predictive maintenance result includes continuing to run and warning for ordering; inputting the sample belief state into the real-time maintenance framework to obtain the sample predictive maintenance result of the sample mechanical equipment, and the sample predictive maintenance result includes continuing to run and warning for maintenance.

[0088] Both the real-time ordering framework and the real-time maintenance framework include a value network. Inputting the sample current state and the predictive maintenance result into the value network, and determining the evaluation value through the reward function in the value network, including: inputting the sample current state and the predictive maintenance result into the value network of the real-time ordering framework, and determining the evaluation value of the real-time ordering framework through the reward function in the value network of the real-time ordering framework; inputting the current state and the sample predictive maintenance result into the value network of the real-time maintenance framework, and determining the evaluation value of the real-time maintenance framework through the reward function in the value network of the real-time maintenance framework; the evaluation value represents the value of adopting the sample predictive maintenance result under the current state.

[0089] In an exemplary embodiment, the construction process of the real-time ordering framework is as follows:

[0090] If the framework issues an order demand at t = T - L, the order arrives when the remaining useful life (RUL) of the engine is 0. At this time, there is no holding cost and shortage cost, and the total cost C is the spare part ordering price c.

[0091] If the framework issues an order demand at t < T - L, the order arrives when the RUL of the engine > 0. At this time, there is a holding cost, and the total cost can be expressed as:

[0092] C = c + c·I·(T - L - t) (1)

[0093] Where, I represents the rate of return on funds, t is the service time, T is the life of the mechanical equipment, and L is the order lead time.

[0094] If the framework issues an order demand at t > T - L and t <= T, the order arrives when the RUL of the engine < 0. At this time, there is a shortage cost, and the total cost can be expressed as:

[0095] C = c + R·(t + L - T) (2)

[0096] Where, R is the unit profit of the mechanical equipment.

[0097] In the concept of reinforcement learning, the current state is a summary of all previous states. That is, the next state depends only on the current state. Therefore, unlike ordinary time series prediction, it is not possible to set a time sliding window to improve prediction stability. However, to reduce the impact of special cases on system maintenance and improve the robustness of the system, the present invention designs a state sliding window (SST) to replace the time sliding window. Generally speaking, the state of the entire interval [T - L - SST, T - L + SST] represents the state at time T - L. When the service time t of the real-time maintenance framework reaches this interval segment, it is considered that the framework predicts RUL = L.

[0098] This real-time ordering framework sets a negative reward function. For example, when RUL = L, the framework can choose two actions at this time: continue to run or give a warning for ordering. When the action is to continue running, the mechanical equipment continues to work without incurring costs. When the action is to give a warning for ordering, the mechanical equipment sends an ordering request, and the total cost is c. The negative reward function r = -c is set. To improve the training effect and correctly guide the actions of the intelligent agent, the present invention slightly modifies the cost function and converts it into a reward function. In the above situation, when the framework predicts RUL = L and chooses to give a warning for ordering, it is the most ideal situation. Choosing to stop for maintenance in other states is a sub-optimal strategy. Therefore, the reward function of the real-time ordering framework can be defined as follows:

[0099]

[0100] where r t is the evaluation value, t is the service time of the sample mechanical equipment, a t is the sample predictive maintenance result of the sample mechanical equipment, C is the total ordering cost of the sample mechanical equipment, T is the life of the mechanical equipment, L is the advance ordering period, SST is the state sliding window, Order represents giving a warning for ordering, and Continue represents continuing to run.

[0101] Among them, the process of the real-time ordering framework is as follows:

[0102]

[0103]

[0104] Input of the real-time ordering framework: belief state ζ = [x1, x2,..., x LD T , a t-1 (the action selected in the previous state) c, I, L, R.

[0105] ​Output of the real-time ordering framework: the trained DT-PPO, updated policy network parameters and value network parameters, ordering cost C, and ordering point T0* = t. It should be noted that the ordering point is the moment when the predicted order is issued, and the policy network is the real-time ordering framework.

[0106] Specifically, the real-time ordering framework conducts a total of max_epoch training sessions. Each training session processes LD data sequences, from t = 1 to t = LD, and terminates when the action a t = Order, at which time T* = t represents the ordering point.

[0107] According to the stochastic policy function π θ (a t |s t ) Formula 4 selects the action a, where π θ is the policy model determined by the parameter θ, usually a neural network. The stochastic policy adopted by the Proximal Policy Optimization (PPO) algorithm means that the agent selects the action a according to the probability distribution of the current state s.

[0108] π θ (a t |s t ) = P(a t = a|s t = s) (4)

[0109] The present invention designs an s-gpolicy, which combines a deterministic policy and a stochastic policy. The stochastic policy is used in the training stage for exploration purposes, and the deterministic policy is adopted in the testing stage for exploitation purposes. The s-gpolicy can be expressed as follows:

[0110]

[0111] Among them, softmax is the stochastic policy, and greedy is the deterministic policy. Select the action a according to the s-gpolicy t , and execute the action a t .

[0112] The advantage function A can be calculated according to formula (6) t GAE(λ) .

[0113]

[0114]

[0115] Among them, r is the discount factor, and λ is the GAE parameter. is the V - value part of the error at time step t + l, and the calculation of the V - value part uses the value function V(s).

[0116] Calculate the ratio according to the clip method in formula (8), and the clipped ratio is used for the calculation of the policy function loss and the value function loss.

[0117]

[0118] where, π θ (a t |s t ) represents the probability of choosing action a t in state s t according to the new policy parameter θ. π θold (a t |s t ) is the probability of choosing action a old in state s t according to the old policy function θ t . r t (θ) represents the unclipped version of the ratio calculated according to the new policy parameter θ multiplied by the advantage function. The role of the Clip function is to limit this ratio within the interval [1 - ε, 1 + ε] to ensure that the policy update is not too drastic.

[0119] In an exemplary embodiment, the construction process of the real - time maintenance framework is as follows:

[0120] If the predictive maintenance result of the real - time maintenance framework is to select shutdown maintenance when t < T, there is a replacement cost at this time, and the total cost can be expressed as:

[0121] C = cr (9)

[0122] Regardless of whether the predictive maintenance result of the real - time maintenance framework is to select continuous operation or shutdown maintenance when t = T, there are replacement costs and reset costs at this time, and the total cost can be expressed as:

[0123] C = cr + cf (10)

[0124] The real-time maintenance framework sets a negative reward function. For example, when t = T, the real-time maintenance framework can choose two actions at this time: continue to run or stop for maintenance. When the action is to continue running, the mechanical equipment continues to work without incurring costs. When the action is to stop for maintenance, the mechanical equipment stops for repair, and the total cost is cr + cf. The negative reward function is set as r = -(cr + cf). To improve the training effect and correctly guide the actions of the intelligent agent, the present invention slightly modifies the cost function and converts it into a reward function. In the above situation, it is the most ideal situation for the framework to choose to stop for maintenance at t = T - 1, at which time the ordered spare parts just arrive. Choosing to stop for repair in other states is a sub-optimal strategy. The real-time maintenance framework also uses a state sliding window instead of a time sliding window. The reward function of the real-time maintenance framework can be defined as follows:

[0125]

[0126] where r t is the evaluation value, t is the service time of the sample mechanical equipment, a t is the sample predictive maintenance result of the sample mechanical equipment, cr is the replacement cost of the sample mechanical equipment, cf is the reset cost, SST is the state sliding window, Replace represents preventive maintenance, and Continue represents continue to run.

[0127] The process of the real-time maintenance framework is as follows:

[0128]

[0129]

[0130] The input of the real-time maintenance framework: belief state ζ = [x1, x2,..., x LD T , α t-1 (the action selected in the previous state), cr, cf.

[0131] The output of the real-time maintenance framework: the trained DT-PPO, update the policy network parameters and value network parameters, and the maintenance point T0* = t; it should be noted that the maintenance point is the moment when preventive maintenance is issued. The policy network in this embodiment is the real-time maintenance framework.

[0132] Specifically, the real-time maintenance framework conducts a total of max_epoch times of training. Each time of training processes LD data sequences. From t = 1 to t = LD, it terminates when the action a t = "Replace", and at this time T* = t represents the maintenance point.

[0133] According to formula (4), the policy function π θ (a t |s​t ) Select action a, where π θ is a policy model determined by parameter θ, usually a neural network, a stochastic policy adopted by the PPO algorithm, that is, the agent selects action a according to the probability distribution of the current state s.

[0134] Calculate the advantage function according to formula (6) Calculate the ratio and the clipped ratio according to the clip method in formula (8) for the calculation of the policy function loss; then calculate the policy function loss and the value function loss according to the min method in formula 8.

[0135] Input the current state of the sample and the predictive maintenance result into the value network, and determine the evaluation value through the reward function in the value network; determine the policy loss according to the sample predictive maintenance result and the sample true maintenance result, and determine the value loss according to the evaluation value; update the parameters of the initial proximal policy network according to the policy loss, and update the parameters of the value network according to the value loss; determine the updated initial proximal policy network model as the real-time conversion-proximal policy optimization network model.

[0136] In an exemplary embodiment, the sample predictive maintenance result is the action output by the policy function, and the reward function in the value network is r used in the ordering policy and the maintenance policy above t , determine the policy loss and the value loss according to the comparison between the action output by the policy function and the action output by s-gpolicy above; it should be noted that PPO is the real-time ordering framework and the real-time maintenance framework.

[0137] In an exemplary embodiment, this embodiment includes: obtaining the test belief state of the test mechanical equipment; inputting the test belief state into the real-time conversion-proximal policy optimization network model to obtain the action data of the test mechanical equipment; the action data is the determined value of each behavior of the test mechanical equipment; the behaviors include early warning ordering, shutdown maintenance, and continued operation; determine the test sample predictive maintenance result through a stochastic policy; determine the test sample true maintenance result through a deterministic policy; determine the performance evaluation result of the real-time conversion-proximal policy optimization network model according to the test sample predictive maintenance result and the test sample true maintenance result.

[0138] Specifically, the prediction accuracy of the real-time conversion-proximal policy optimization network model can be determined based on the predictive maintenance results of the test samples and the true maintenance results of the test samples. If the accuracy is greater than the preset accuracy threshold, the performance evaluation result of the real-time conversion-proximal policy optimization network model is determined to be excellent. If the prediction accuracy of the real-time conversion-proximal policy optimization network model is less than or equal to the preset accuracy threshold, the performance evaluation result of the real-time conversion-proximal policy optimization network model is determined to be poor, that is, it is necessary to retrain the real-time conversion-proximal policy optimization network model.

[0139] In an exemplary embodiment, the present invention also provides a flowchart of a predictive maintenance method for real-time sensor-driven, as Figure 3 shown. Specifically, this embodiment includes: collecting the initial states of multiple sensors of a mechanical device in different operating states, using a state conversion network to convert the initial states into corresponding belief states, inputting training samples into DT-PPO for quality evaluation of the generated samples, and using test samples to output predictive maintenance decisions for the test samples.

[0140] To verify the effectiveness of the predictive maintenance method for mechanical devices provided by the present invention, in the present invention, a dataset (Commercial Modular Aero-Propulsion System Simulation, C-MAPSS) is used to carry out method verification. C-MAPSS is derived from a turbofan engine. A turbofan engine is a modern gasoline turbine engine, which is used by the National Aeronautics and Space Administration (NASA) of the United States. The structural diagram of the simulated aeroengine is as Figure 4 shown. The simulated aeroengine includes a fan 1, a combustion chamber 2, a low-pressure rotor speed 3, a low-pressure turbine 4, a low-pressure compressor 5, a high-pressure compressor 6, a high-pressure rotor speed 7, a high-pressure turbine 8, and a nozzle 9. The dataset includes the time series of each engine. All engines are of the same type, but during the manufacturing process, the initial wear degree and differences of each engine are different, which are unknown to the user. There are three optional configurations that can be used to change the performance of each engine. Each engine has 21 sensors. When the engine is running, the sensors collect measurement data related to the engine state. There is some sensor noise in the collected data. Gradually, each engine will have some deficiencies, which can be detected from the sensor readings. The time series ends at a certain time before a failure occurs. Table 1 shows the simulated aeroengine dataset, including an introduction to the C-MAPSS dataset.

[0141] In actual monitoring, since the outputs of some sensors are fixed and cannot provide effective information for the framework. To reduce the computational complexity of the framework, the present invention selects 14 sensor parameters to construct samples, and extracts degradation feature information through the proposed framework MCIBIGRUN. The corresponding numbers of sensors are 2, 3, 4, 7, 8, 9, 11, 12, 13, 14, 15, 17, 20, and 21 respectively.

[0142] Table 1

[0143]

[0144] Then, a state transition network model is established, and the training sample set is used for network model training. The state transition network, as a deep autoencoder, selectively extracts multi-sensory features at different scales and outputs the belief state. The belief state of TransStateNet remains the same size as the initial state and serves as the input to the state space of DT-PPO. Among them, the Adam optimizer is used to perform stochastic gradient optimization on the network. The initial learning rate is 0.01, the weight decay rate is 0.01, and it is iteratively trained 1000 times, and gradient penalty regularization is used to improve the training stability.

[0145] To prove the effectiveness of the proposed real-time ordering framework, the present invention first defines a set of performance indicators for comparison. The first is the root mean square error (RMSE), the second is the scoring function (Score), and the third is the mean bias error (MBE). The experimental results of the real-time ordering framework are related to the ordering price of mechanical equipment, the capital utilization rate of the market, and the profit generated per unit service cycle. The prices of turbofan engines vary greatly due to factors such as type, performance, manufacturer, degree of newness, and whether after-sales support is included. The capital utilization rate is the interest rate of the central bank and is relatively stable. The unit profit is affected by multiple factors, including customer income, operating costs, and market strategies. To make a better comparison, the present invention considers different cost combinations c = {100, 500}, I = 0.05, R = {0.5, 1}.

[0146] The experimental results are shown in Table 2. Table 2 presents the experimental results of the real-time ordering framework, showing the RMSE, Score, and MBE of the real-time ordering framework under different cost combinations. When c = 100, an increase in R may help reduce the prediction error. When c = 500, an increase in R may lead to an increase in the prediction error, and all three metrics show such a trend. When R = 0.5, an increase in c may help reduce the prediction error. When R = 1, an increase in c may lead to an increase in the prediction error, and the three metrics also show such a trend. This may mean that under the condition of a constant parameter, the impact of the change of another parameter on the performance metric is not monotonic, and further analysis may be needed to determine its impact pattern. In addition, under all cost combinations, MBE is greater than 0, indicating that the framework always tends to issue an order warning in advance.

[0147] Table 2

[0148] c I R L PM MSE 100 0.05 0.5 20 191.40 100 0.05 1 20 147.20 500 0.05 0.5 20 136.40 500 0.05 1 20 240.40 MAPE 100 0.05 0.5 20 20% 100 0.05 1 20 20% 500 0.05 0.5 20 20% 500 0.05 1 20 20% MBE 100 0.05 0.5 20 5.40 100 0.05 1 20 2.80 500 0.05 0.5 20 2.80 500 0.05 1 20 7.20

[0149] To prove the effectiveness of the proposed control decision framework, the present invention first defines a set of performance metrics for comparison. The first is the average maintenance cost rate (AMCR), the second is the average wasted operating life (AWOL), the third is the average scheduled replacement times (ASRT), and the fourth is the average unscheduled replacement times (AUSRT). To prove the effectiveness of the proposed control decision, the present invention first defines a set of benchmark strategies for comparison. The first is the ideal maintenance strategy (IMC), the second is the corrective maintenance strategy (CMC), and the third is the scheduled maintenance strategy (SM).

[0150] The experimental results are shown in Table 3. Table 3 presents the experimental results of the real-time ordering framework, showing the AC, AWOL, ASRT, and AUSRT of the framework under different cost combinations. The first result is the average maintenance cost rate for five different maintenance strategies and four different cost combinations. As expected, the ideal alternative policy IMC provides the lowest cost. The present invention provides the second-best result and performs better than the scheduled maintenance strategy and the corrective maintenance strategy. In addition, the average wasted operating life of the four benchmark strategies is also reported. As shown in the table, there is no wasted operating life in the corrective maintenance strategy. The ideal maintenance cost strategy IMC provides the second-best result. The predictive maintenance strategy of the present invention ranks after the two and is better than the two cases of the scheduled maintenance strategy. Finally, the framework of the present invention performs well in terms of the average scheduled replacement times and the average unscheduled replacement times.

[0151] Table 3

[0152]

[0153] The content of the present invention is mainly divided into two parts. The first part is the state transition network, which is used to extract multi-dimensional features of sensors, eliminate redundant degradation information and convert the initial state into a belief state; the second part is to establish an improved state transition-proximal policy optimization network model, including a real-time ordering framework and a real-time maintenance framework, integrating the belief state of mechanical equipment into the real-time ordering and real-time maintenance frameworks. The real-time ordering framework is used to issue order requests; the real-time maintenance framework is used to decide when to stop the equipment for maintenance.

[0154] Among them, the convolutional neural network in the state transition network model is the CNN module, the residual neural network in the state transition network model is the ResNet module, the self-attention mechanism module in the state transition network model is the Self-Attention module, and the bidirectional gated recurrent unit in the state transition network model is the BiGRU module.

[0155] Refer to Figure 5 As shown, PPO adopts the Actor-Critic architecture. The present invention uses a two single-hidden-layer neural network structure to test the method we proposed to find the best network representing the correct mapping between states and actions. In addition to the state processed by the state transition network, the input layer also includes the action selected by PPO according to the previous state. The two neural networks have the same number of input and hidden layer neurons. The policy network outputs the action probability distribution, that is, the probability of performing each action in a given state. This network is usually the output of the softmax function. The value function network outputs the estimated state value function, that is, the expected value of the total future reward obtained in the current state. This is a scalar output, reflecting how much reward the agent is expected to accumulate in the current state. Considering the limitations of each architecture, the number of hidden neurons and hidden layers of the Actor-Critic architecture is determined through cross-validation to better represent the interaction between the input data and its target value.

[0156] The beneficial effects of the predictive maintenance method driven by real-time sensors in the present invention are as follows: relying on the real-time data sharing of sensors, the powerful feature extraction ability of deep learning, and the dynamic decision-making ability of reinforcement learning, this method converts the real-time sensor data collected from monitoring sensors into the basis for real-time maintenance and real-time decision-making. Through the state transition network module, the framework can effectively process data from multiple sensors, perform feature extraction and prediction of degradation features, and solve the high-dimensional and multi-scale problems that are difficult to handle in traditional methods. It can be seen from the experimental results that the framework shows good robustness under different cost combinations and failure modes and can adapt to different operating conditions and external environments. This method can make full use of existing fault samples, flexibly mine sample features, and accurately match the local feature positions in images by embedding a multi-head attention mechanism, significantly improving the feature extraction performance of the generative adversarial network and the local generation quality of samples. It effectively improves the accuracy and stability of the deep learning fault diagnosis model in the case of extremely small samples. Although the implementation embodiments of the present invention are disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the illustrations shown and described here.

[0157] In the predictive maintenance method of mechanical equipment provided by the present invention, a state transition network is built as a deep autoencoder to selectively extract multi-sensory features of different scales of sensors and output a belief state; a real-time conversion-proximal policy optimization network model based on the belief state is designed to establish an order spare parts or shutdown maintenance strategy; an s-g policy is designed for the real-time conversion-proximal policy optimization network model to guide action selection and a state sliding window is used to process multi-sensor state sequence data, reducing the uncertainty of the action space of the real-time conversion-proximal policy optimization network model, improving the robustness of the maintenance strategy, and avoiding accidents. This method solves the problem that existing models for integrating data-driven remaining useful life into maintenance plans are too complex.

[0158] When applying the predictive maintenance method of a mechanical equipment provided by the present invention, it is not necessary to execute according to Figure 1 the order of the steps shown. The specific execution order of each step can be determined as needed, and the present invention does not limit this.

[0159] The above is a predictive maintenance method of a mechanical equipment provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding predictive maintenance device for a mechanical equipment, as Figure 6 shown.

[0160] Figure 6Schematic diagram of a predictive maintenance device for a mechanical device provided by the present invention. The device 600 includes:

[0161] An acquisition module 601, configured to acquire status data of the mechanical device at multiple moments;

[0162] An extraction module 602, configured to input the status data into a status conversion network model, extract multi-sensory features of different scales of the status data through the status data, and obtain a belief state; the belief state is a feature representation of the status data;

[0163] A prediction module 603, configured to input the belief state into a pre-trained real-time conversion-proximal policy optimization network model to obtain a predictive maintenance result of the mechanical device.

[0164] For the specific limitations of a predictive maintenance device for a mechanical device, reference may be made to the limitations of a predictive maintenance method for a mechanical device in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned predictive maintenance device for a mechanical device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0165] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program can be used to execute the above-mentioned Figure 7 A predictive maintenance method for a mechanical device provided.

[0166] The present invention also provides Figure 7 The structural schematic diagram of the computer device shown, as Figure 7 shown, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, there may also be other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned Figure 1 A predictive maintenance method for a mechanical device provided.

[0167] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0168] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded by the present invention.

Claims

1. A predictive maintenance method for mechanical equipment, characterized in that: include: Obtain status data of mechanical equipment at multiple times; The state data is input into the state transition network model, and the multi-sensory features of different scales of the state data are extracted through the state transition network model to obtain the belief state; the belief state is a feature representation of the state data; the state transition network model includes a convolutional neural network, a residual neural network, a self-attention mechanism module and a linear neural network module, and the multi-sensory features of different scales of the state data are extracted through the state transition network model to obtain the belief state, including: inputting the state data into the convolutional neural network to obtain the high-order features of the state data; inputting the high-order features of the state data into the residual neural network to obtain the high-order optimization features of the state data; inputting the high-order optimization features of the state data into the self-attention mechanism module to obtain the high-order belief state; inputting the high-order belief state of the state data into the linear neural network to obtain the belief state; the state transition network model also includes a bidirectional gated recurrent unit, and the method also includes: inputting the belief state into the bidirectional gated recurrent unit to obtain the remaining service life of the mechanical equipment; The belief state is input into a pre-trained real-time conversion-proximal policy optimization network model to obtain the predictive maintenance result of the mechanical equipment; the real-time conversion-proximal policy optimization network model includes a real-time ordering framework and a real-time maintenance framework; the belief state is input into a pre-trained real-time conversion-proximal policy optimization network model to obtain the predictive maintenance result of the mechanical equipment, including: inputting the belief state into the real-time ordering framework to obtain the ordering strategy of the mechanical equipment, and determining the ordering strategy as the predictive ordering result; inputting the belief state into the real-time maintenance framework to obtain the maintenance strategy of the mechanical equipment, and determining the maintenance strategy as the predictive maintenance result; The construction process of the real-time conversion-proximal policy optimization network model includes: obtaining a sample belief state of a sample mechanical equipment; inputting the sample belief state into the policy network to obtain a sample predictive maintenance result of the sample mechanical equipment; inputting the current state and the sample predictive maintenance result into the value network, and determining an evaluation value through a reward function in the value network; determining a policy loss based on the sample predictive maintenance result and the sample actual maintenance result, and determining a value loss based on the evaluation value; updating the parameters of the policy network based on the policy loss, and updating the parameters of the value network based on the value loss; determining the updated policy network and value network as the real-time conversion-proximal policy optimization network model.

2. The method according to claim 1, characterized in that The obtaining of state data of the mechanical equipment at multiple times includes: Acquire original state data collected in real time by multiple sensors on the mechanical equipment at the multiple moments; The original state data is normalized to obtain state data of the mechanical equipment at multiple moments.

3. The method according to claim 1, characterized in that The step of inputting the sample belief state into the policy network to obtain the sample predictive maintenance result of the sample mechanical equipment includes: Input the sample belief state into the policy network to obtain the action data of the sample mechanical equipment; the action data is the probability value of each behavior of the sample mechanical equipment; the behavior includes early warning ordering, shutdown maintenance and continued operation; According to the probability value of each behavior, the sample predictive maintenance result is determined through a random strategy.

4. The method according to claim 1, characterized in that: The method further comprises: Get the test belief status of the test mechanical equipment; Inputting the test belief state into the real-time conversion-proximal strategy optimization network model to obtain the action data of the test mechanical equipment; the action data is a determined value of each behavior of the test mechanical equipment; the behavior includes continued use, maintenance and ordering; Determine the predictive maintenance results of the test samples through randomization strategy; Determine the true maintenance results of the test samples through a deterministic strategy; According to the test sample predictive maintenance result and the test sample actual maintenance result, the performance evaluation result of the real-time conversion-proximal strategy optimization network model is determined.

5. The method according to claim 1, characterized in that The real-time conversion-proximal strategy optimization network model includes a real-time ordering framework, and the reward function includes: Among them, r t is the evaluation value, t is the service life of the sample mechanical equipment, a t is the sample predictive maintenance result of the sample mechanical equipment, C is the total ordering cost of the sample mechanical equipment, T is the life of the mechanical equipment, L is the advance ordering period, SST is the status window, Order means warning ordering, Continue means continued operation, I represents the rate of return on capital, and R is the unit profit of the mechanical equipment.

6. The method according to claim 1, characterized in that The real-time conversion-proximal strategy optimization network model includes a real-time maintenance framework, and the reward function includes: Among them, r t is the evaluation value, t is the service life of the sample mechanical equipment, a t is the sample predictive maintenance result of the sample mechanical equipment, cr is the replacement cost of the sample mechanical equipment, cf is the replacement cost, SST is the status window, Replace indicates early warning maintenance, and Continue indicates continued operation.

Citation Information

Patent Citations

  • Mechanical equipment dynamic maintenance decision-making method based on evaluation and prediction information

    CN117454771A

  • Engine remaining service life prediction method based on attention multi-dimensional dynamic convolutional network

    CN118484645A