Operation and maintenance method and system based on deep reinforcement learning, and storage medium
Through the operation and maintenance method based on deep reinforcement learning, the equipment is multi-dimensional time series data to build a comprehensive health index and multi-time scale matrix, and the strategy network and value network are trained, which solves the problems of fault diagnosis limitations and insufficient decision-making accuracy in equipment operation and maintenance, and achieves more efficient operation and maintenance decision-making and fault analysis.
Patent Information
- Application Number
- CN202510390918.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has problems such as limitations in fault diagnosis and insufficient decision-making accuracy in equipment operation and maintenance, especially the inability to effectively link the equipment health index with specific root causes of failures, and ignores the dynamicity and multi-scale characteristics of the state data.
Using an operation and maintenance method based on deep reinforcement learning, we use the multi-dimensional time series data of the equipment to build a comprehensive health index sequence and a multi-time scale matrix sequence, train the strategy network and value network, output risk levels and exception types, and generate operation and maintenance reports.
The problem of sparse feedback in the deep reinforcement learning process is solved, the accuracy and robustness of operation and maintenance decisions are improved, and the equipment failure factors can be analyzed in more detail and corresponding operation and maintenance measures are given.
Smart Images

Figure CN120218909A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of operation and maintenance, and in particular, to an operation and maintenance method, system, and storage medium based on deep reinforcement learning. Background Art
[0002] Traditional device operation and maintenance methods mainly rely on manual operation and maintenance. The disadvantage is that operation and maintenance measures are only taken after obvious abnormal indicators are observed or the device has already shown abnormalities. Moreover, operation and maintenance personnel can often only focus on the software level of the device and cannot perceive problems such as device hardware aging and environmental deterioration. As a result, it is impossible to effectively and comprehensively evaluate the device operation status and prevent problems before they occur. To achieve good operation and maintenance effects, manual operation and maintenance require a large amount of manpower and time, such as frequently querying device alarm performance and other indicators, and even going to the site.
[0003] In the prior art, the deep reinforcement learning method is introduced in the intelligent operation and maintenance process. The model is iteratively trained by collecting device status data. When the model converges, it can make effective feedback on the input, so as to provide effective operation and maintenance measures when the device is at risk. In general operation and maintenance scenarios, the device health index is usually regarded as an indicator of whether the device status is healthy. Although the device health index can reflect the overall operation status of the device, it cannot effectively link the health status with the specific root cause of the failure. That is to say, when the device health index drops, operation and maintenance personnel can know that there is a problem with the device, but it is difficult to quickly and accurately locate the specific location of the failure, resulting in obvious deficiencies in fault diagnosis and possible delays in operation and maintenance response. In addition, existing operation and maintenance models usually only focus on the status data of the device at a certain moment, ignoring the dynamic and multi-scale characteristics of the status data, resulting in the operation and maintenance model being unable to comprehensively understand the change law of the device status, thus affecting the accuracy of decision-making. Therefore, the prior art lacks an operation and maintenance means that can solve the limitations of fault diagnosis and improve the accuracy of decision-making at the same time. Summary of the Invention
[0004] The purpose of the present invention is to provide an operation and maintenance method, system, and storage medium based on deep reinforcement learning, aiming to solve the limitations of fault diagnosis by solving the problem of sparse training feedback of the operation and maintenance model, and improving the accuracy of operation and maintenance decision-making by using the dynamic data of the device status.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] In the first aspect, an operation and maintenance method based on deep reinforcement learning is provided, including:
[0007] Collect multi-dimensional time series data of the device and perform preprocessing to construct a comprehensive health index sequence for representing the device health status and a multi-time scale matrix sequence for representing the device operation status;
[0008] Taking the comprehensive health index sequence and the multi-time scale matrix sequence as inputs, and the risk level and abnormal type as outputs, train a policy network. The policy network outputs an action a based on the state S(t) at time t. The agent obtains the next state S(t + 1) and the comprehensive reward r associated with the comprehensive health index through the interaction of the action a with the environment. t , store the experience data [S(t), a, r t , S(t + 1)] in the experience replay pool;
[0009] Extract experience data from the experience replay pool, and train the value network and the policy network. The value network updates the Q-value based on an improved Q-value function, updates the parameters of the value network by minimizing the loss function of the value network, and updates the parameters of the policy network by the policy gradient method to obtain an optimal decision-making model;
[0010] Preprocess the real-time status data of the device and input it into the optimal decision-making model, and generate an operation and maintenance report according to the probability distribution of the risk level and abnormal type output by the policy network.
[0011] Furthermore, the multi-dimensional time series data of the device is the device status data collected from historical monitoring data or real-time monitoring data, including: operating environment data, hardware performance data, system performance data, software performance data, current operating alarm data, and historical operating alarm data.
[0012] Furthermore, construct a comprehensive health index sequence for representing the health status of the device, including:
[0013] Obtain the time series data of multiple performance indicators related to the health status from the preprocessed multi-dimensional time series data;
[0014] Calculate the health score of each performance indicator at time t:
[0015]
[0016] Calculate the comprehensive health index EHI(t) of the device at time t for representing the health status of the device:
[0017]
[0018] where H i (t) is the health score of the i-th performance indicator at time t, and the value range is [0, 1];
[0019] x i (t) is the value of the i-th performance indicator at time t;
[0020] x iref is the reference value of the i-th performance metric;
[0021] σ i is the standard deviation of the i-th performance metric;
[0022] w i is the weight of the i-th performance metric;
[0023] Construct a multi-time scale matrix sequence for representing the operating state of the device, including:
[0024] Obtain the time series data of multiple variables related to the operating state from the preprocessed multi-dimensional time series data, and perform scale division according to a preset time window to obtain several multi-dimensional sequence segments;
[0025] Calculate the eigenvalues for each multi-dimensional sequence segment of each time window, including the mean, variance, and autocorrelation coefficient, and construct a feature matrix:
[0026]
[0027] where, F k is the feature matrix corresponding to the k-th time window;
[0028] μ i,k σ 2 i,k ρ i,k respectively represent the mean, variance, and autocorrelation coefficient of the i-th variable in the k-th time window;
[0029] Obtain a multi-time scale matrix sequence Z based on the feature matrices corresponding to each time window, and the multi-time scale matrix sequence Z is used to represent the operating state of the device at different time scales:
[0030] Z = [F1, F2,...... F m .
[0031] Furthermore, the risk level f output by the policy network includes:
[0032] Low risk, indicating that the device is operating normally and there are no obvious fault signs in the short term;
[0033] Medium risk, indicating that the device may have potential problems, but will not cause serious faults in the short term;
[0034] High risk, indicating that the device has obvious fault signs or abnormalities and may malfunction in the short term;
[0035] Emergency, indicating that the device state is very bad, a fault is about to occur, and immediate action is required;
[0036] The abnormal types m output by the policy network include but are not limited to:
[0037] m = 0 indicates that the device is operating normally;
[0038] m = 1001 indicates that the fan speed is abnormal, resulting in too high a temperature;
[0039] m = 1002 indicates that the fan speed is abnormal, resulting in too low a temperature;
[0040] m = 1003 indicates that the device is running at high speed continuously, resulting in too high a temperature;
[0041] m = 1004 indicates that the fan failure leads to too high humidity;
[0042] m = 1005 indicates that the fan failure leads to too low humidity;
[0043] m = 2001 indicates that the memory card has too many read / write times;
[0044] m = 2002 indicates hardware aging attenuation;
[0045] m = 3001 indicates that too high a temperature leads to a high CPU occupancy rate;
[0046] m = 3002 indicates an abnormal i / o interface;
[0047] m = 4001 indicates that the abnormal operation of the device leads to a high CPU occupancy rate;
[0048] m = 4002 indicates that the process restarts repeatedly;
[0049] m = 5001 indicates a link interruption;
[0050] m = 5002 indicates data packet loss;
[0051] m = 5003 indicates data error code;
[0052] m = 6001 indicates that the device is offline;
[0053] m = 6002 indicates that the device data writing fails;
[0054] m = 6003 indicates that the device protection switching fails.
[0055] Furthermore, the comprehensive reward r t = r + λ·EHI(t);
[0056] where r represents the immediate reward for executing the current action a at state S(t), and λ represents the weight of the comprehensive health index EHI(t);
[0057] The improved Q-value function is:
[0058] Q(S(t), a) ← Q(S(t), a) + α[(r + λ·EHI(t)) + γ max a ′Q(S(t + 1), a′) - Q(S(t), a)]
[0059] Among them, Q(S(t), a) represents the Q - value when the agent executes the current action a in state S(t);
[0060] Q(S(t + 1), a′) represents the Q - value when the agent executes the next action a′ in state S(t + 1);
[0061] α represents the learning rate;
[0062] γ represents the reward discount factor;
[0063] max a′ Q(S(t + 1), a′) represents the maximum Q - value among all possible actions taken in state S(t + 1).
[0064] Furthermore, the loss function of the value network is the mean square error between the current Q - value and the target Q - value;
[0065] Calculate the TD error according to the current Q - value and the target Q - value:
[0066] TD = α[(r + λ·EHI(t)) + γ max a′ Q(S(t + 1), a′) - Q(S(t), a)].
[0067] In a second aspect, a maintenance and operation system based on deep reinforcement learning is provided, including:
[0068] A data processing module, configured to collect multi - dimensional time - series data of the device and perform pre - processing, and construct a comprehensive health index sequence for representing the health state of the device and a multi - time - scale matrix sequence for representing the operation state of the device;
[0069] A pre - training module, configured to use the comprehensive health index sequence and the multi - time - scale matrix sequence as inputs, and use the risk level and the abnormal type as outputs to train a policy network. The policy network outputs an action a based on the state S(t) at time t, and feeds back a comprehensive reward r to the agent t , and the agent obtains the state S(t + 1) at the next moment through the interaction between the action a and the environment, and stores the experience data [S(t), a, r t , S(t + 1)] into the experience replay pool;
[0070] A repeated training module, which is used to extract experience data from the experience replay pool, train a value network and the policy network. The value network updates the Q value based on an improved Q value function, updates the parameters of the value network by minimizing the loss function of the value network, and updates the parameters of the policy network by a policy gradient method to obtain an optimal decision-making model;
[0071] An operation and maintenance module, which is used to preprocess the real-time status data of the device and input it into the optimal decision-making model, and generate an operation and maintenance report according to the probability distribution of the risk level and abnormal type output by the policy network.
[0072] Further, the comprehensive reward r t = r + λ·EHI(t);
[0073] where r represents the immediate reward for executing the current action a in the state S(t), and λ represents the weight of the comprehensive health index EHI(t);
[0074] The improved Q value function is:
[0075] Q(S(t),a) ← Q(S(t),a) + α[(r + λ·EHI(t)) + γmax a′ Q(S(t + 1),a′) - Q(S(t),a)]
[0076] where Q(S(t),a) represents the Q value of the agent executing the current action a in the state S(t);
[0077] Q(S(t + 1),a′) represents the Q value of the agent executing the next action a′ in the state S(t + 1);
[0078] α represents the learning rate;
[0079] γ represents the reward discount factor;
[0080] maxa′Q(S(t + 1),a′) represents the maximum Q value among all possible actions in the state S(t + 1).
[0081] Further, the loss function of the value network is the mean square error between the current Q value and the target Q value;
[0082] Calculate the TD error according to the current Q value and the target Q value:
[0083] TD = α[(r + λ·EHI(t)) + λmax a′ Q(S(t + 1),a′) - Q(S(t),a)].
[0084] Based on the same inventive concept, the present invention also provides a computer storage medium storing computer-executable instructions, and when the computer-executable instructions are executed, the foregoing operation and maintenance method is implemented.
[0085] Technical effects and advantages of the present invention:
[0086] (1) The comprehensive health index sequence and the multi-time scale matrix sequence are introduced into the policy network for training. During training, the comprehensive reward r associated with the comprehensive health index is fed back to the agent, t solving the problem of sparse feedback in the process of deep reinforcement learning;
[0087] (2) The multi-time scale matrix sequence extracts more effective state representations from the multi-dimensional time series data of the device, enabling the model to better understand the dynamic behavior of the device, thereby more effectively capturing the trend changes and abnormal fluctuations in the device operation, and improving the accuracy and robustness of operation and maintenance decisions;
[0088] (3) The process of deep reinforcement learning does not depend on the data at a certain moment, but takes into account the long-term influence of various external factors, preventing misjudgment of the health state of the device, and when the health index of the device is reduced due to various factors, it can more detailedly analyze the device failure factors and give corresponding operation and maintenance measures.
[0089] Other features and advantages of the present invention will be described in the subsequent description, and part of them will become obvious from the description, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures pointed out in the description, the claims and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0091] Figure 1 It is a schematic flowchart of the operation and maintenance method in the embodiment of the present invention;
[0092] Figure 2 It is a schematic diagram of the association of operation and maintenance devices in a specific embodiment of the present invention;
[0093] Figure 3 It is a schematic diagram of data processing in a specific embodiment of the present invention;
[0094] Figure 4 It is a schematic flowchart of the operation and maintenance process in a specific embodiment of the present invention;
[0095] Figure 5 Schematic diagram of the principle of the operation and maintenance method in a specific embodiment of the present invention;
[0096] Figure 6 Schematic diagram of the change of the comprehensive health index EHI over time t;
[0097] Figure 7 Schematic diagram of the change of the original Q value over the training time;
[0098] Figure 8 Schematic diagram of the change of the updated Q value over the training time;
[0099] Figure 9 Comparison diagram of the change of the original Q value and the updated Q value over the training time. Detailed implementation manners
[0100] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0101] As Figure 1 shown, an operation and maintenance method based on deep reinforcement learning disclosed in an embodiment of the present invention includes:
[0102] Step S1: Collect multi-dimensional time series data of the device and perform preprocessing to construct a comprehensive health index sequence for representing the health state of the device and a multi-time scale matrix sequence for representing the operating state of the device;
[0103] Step S2: Use the comprehensive health index sequence and the multi-time scale matrix sequence as inputs, and the risk level and the abnormal type as outputs to train a policy network and construct an experience replay pool;
[0104] The policy network outputs an action a based on the state S(t) at time t, and feeds back a comprehensive reward r associated with the comprehensive health index to the intelligent agent t , and the intelligent agent obtains the state S(t + 1) at the next moment through the interaction between the action a and the environment, and stores the experience data [S(t), a, r t , S(t + 1)] into the experience replay pool;
[0105] Step S3: Extract experience data from the experience replay pool to train a value network and the policy network. The value network updates the Q value based on an improved Q value function, updates the parameters of the value network and the policy network, and obtains an optimal decision-making model;
[0106] Update the parameters of the value network and the policy network, including: updating the parameters of the value network by minimizing the loss function of the value network, and updating the parameters of the policy network by the policy gradient method.
[0107] Step S4: Preprocess the real-time status data of the device and input it into the optimal decision-making model, and generate an operation and maintenance report according to the probability distribution of the risk level and abnormal type output by the policy network.
[0108] In the embodiment of the present invention, dynamic information on the health and status of the device is extracted from the multi-dimensional time-series data of the device. By calculating the comprehensive health index sequence and the multi-time scale matrix sequence, it is introduced into the policy network for training. During training, a comprehensive reward r associated with the comprehensive health index is fed back to the intelligent agent. t , which solves the problem of sparse feedback in the process of deep reinforcement learning; since the multi-time scale matrix sequence extracts more effective state representations from the multi-dimensional time-series data of the device, the model can better understand the dynamic behavior of the device, thereby more effectively capturing the trend changes and abnormal fluctuations in the device operation, and improving the accuracy and robustness of the operation and maintenance decision-making.
[0109] The following details this method through a specific embodiment:
[0110] As Figure 2 shown, the central control center is used to monitor the automatic operation and maintenance units deployed on each device. Each automatic operation and maintenance unit includes: a data acquisition module, an algorithm logic module, and a feedback module.
[0111] The main functions of the data acquisition module include: 1) collecting the historical monitoring data or real-time monitoring data of the device; 2) preprocessing the data; 3) using the preprocessed data as the training data of the model.
[0112] The main functions of the algorithm logic module include: 1) training the model; 2) predicting and reporting the device risk situation based on the data collected in real time by the device.
[0113] The main function of the feedback module includes: making corresponding policy operations according to the device risk situation predicted by the algorithm logic module to achieve automatic operation and maintenance.
[0114] As Figures 3 to 5 shown, the process of device automatic operation and maintenance is as follows:
[0115] 1. Data collection;
[0116] Collect device status data from the historical monitoring data or real-time monitoring data of the device to obtain multi-dimensional time series data of the device, including operating environment data (temperature, humidity, etc.), hardware performance data (number of reads and writes, etc.), system performance data (memory occupancy, CPU occupancy, etc.), software performance data (memory occupancy, etc.), current operation alarm data, historical operation alarm data, and so on.
[0117] 2. Data preprocessing;
[0118] 2-1. Data cleaning, used to remove invalid, incorrect or duplicate data;
[0119] 2-2. Data standardization, used to scale different types of data to the same scale. For example, use the z-score formula for data standardization processing;
[0120] 2-3. Data normalization, used to convert data into values between 0 and 1;
[0121] 2-4. Feature extraction, used to select the feature attributes with the largest weights from the current features;
[0122] 2-5. Feature dimensionality reduction, used to perform data dimensionality reduction according to the principal component analysis method.
[0123] The above data processing methods are all conventional data processing methods, and the specific steps will not be elaborated. Use the above methods to process the multi-dimensional time series data of the device to obtain a multi-dimensional vector matrix.
[0124] 3. Construct a comprehensive health index sequence;
[0125] 3-1. Obtain the time series data of multiple performance indicators related to the health status from the preprocessed multi-dimensional time series data;
[0126] 3-2. Calculate the health score of each performance indicator at time t:
[0127]
[0128] 3-3. Calculate the comprehensive health index EHI(t) of the device at time t:
[0129]
[0130] Among them, H i (t) is the health score of the i-th performance indicator at time t, and the value range is [0, 1];
[0131] x i (t) is the value of the i-th performance indicator at time t;
[0132] x iref is the reference value of the i-th performance metric;
[0133] σ i is the standard deviation of the i-th performance metric;
[0134] w i is the weight of the i-th performance metric.
[0135] For example, obtain the time-series data of performance metrics related to the device health status, such as the flow rate x1(t), temperature x2(t), disk usage rate x3(t), and CPU usage rate x4(t).
[0136] Assume that the weights of the above four performance metrics are respectively:
[0137] w1 = 0.3, w2 = 0.2, w3 = 0.3, w4 = 0.2
[0138] Calculate their health scores at time t according to the time-series data of the above performance metrics respectively:
[0139] H1(t) = 0.9, H2(t) = 0.8, H3(t) = 0.7, H4(t) = 0.85
[0140] Calculate the comprehensive health index EHI(t) of the device at time t:
[0141]
[0142] According to the above method, by calculating the comprehensive health index EHI(t) of the device at time t, construct a comprehensive health index sequence, which is used to represent the dynamic health status of the device, Figure 6 is the curve graph of the comprehensive health index EHI changing with time t.
[0143] 4. Construct a multi-time scale matrix sequence;
[0144] 4-1. Obtain the time-series data of multiple variables related to the running state from the preprocessed multi-dimensional time-series data, and perform scale division according to a preset time window to obtain several multi-dimensional sequence segments;
[0145] 4-2. Calculate the eigenvalues of each multi-dimensional sequence segment in each time window, including the mean value, variance, and autocorrelation coefficient, and construct a feature matrix:
[0146]
[0147] where, F k is the feature matrix corresponding to the k-th time window;
[0148] μ i,k 、σ 2i,k , ρ i,k respectively represent the mean, variance, and autocorrelation coefficient of the i-th variable in the k-th time window;
[0149] 4-3. Obtain the multi-time scale matrix sequence Z based on the feature matrices corresponding to each time window:
[0150] Z = [F1, F2,......F m
[0151] For example, obtain the time series data of flow rate, temperature, and CPU usage generated during device operation:
[0152] x1(t) Flow rate: 100, 105, 102, 98, 110, 108, 104 Gbps
[0153] x2(t) Temperature: 40, 42, 41, 39, 43, 44, 40 °C
[0154] x3(t) CPU usage: 70, 75, 72, 68, 77, 74, 71%
[0155] Divide into a time window with three time nodes as a unit, and calculate the mean μ1 of the flow rate x1(t) in the first time window, calculate the mean μ2 of the temperature x2(t) in the first time window, and calculate the mean variance σ3 of the CPU usage x3(t) in the first time window:
[0156]
[0157] Calculate the variance σ1 of the flow rate x1(t) in the first time window 2 , calculate the variance σ2 of the temperature x2(t) in the first time window 2 , calculate the mean variance σ3 of the CPU usage x3(t) in the first time window 2 :
[0158]
[0159] Assume that the autocorrelation coefficients of the flow rate x1(t), temperature x2(t), and CPU usage x3(t) in the first time window are all 0.5:
[0160] ρ1 = 0.5, ρ2 = 0.5, ρ3 = 0.5
[0161] Obtain the feature matrix F1 of the first time window as:
[0162]
[0163] Using the method provided in this step, a feature matrix F for each time window is constructed based on the mean, variance, and autocorrelation coefficient of each variable in each time window. k A multi-time scale matrix sequence Z is constructed based on the feature matrices corresponding to each time window, which is used to represent the operating state of the device at different time scales.
[0164] Reinforcement learning is an artificial intelligence technology aimed at enabling an agent to learn and make decisions in an environment to maximize cumulative rewards. The core idea of reinforcement learning is that through the interaction between the agent and the environment, the agent can learn the optimal decision-making strategy. Deep reinforcement learning combines the advantages of deep learning and reinforcement learning, using neural networks to estimate state values and policies, thus achieving efficient learning and decision-making.
[0165] Deep reinforcement learning learns how to achieve specific goals through the interaction between the agent and the environment. The agent performs actions (Action) in the environment to change the state (State) and obtains rewards (Reward) based on state transitions. The goal of the agent is to maximize its long-term cumulative rewards, that is, the rule for choosing the best action in a given state.
[0166] The key components of reinforcement learning include:
[0167] State, that is, the environment where the agent is located;
[0168] Action, that is, the actions that the agent can perform in a specific state;
[0169] Reward, that is, the feedback obtained by the agent from the environment after performing an action, used to evaluate the quality of the action;
[0170] Policy, that is, the rule or probability for the agent to choose actions;
[0171] Value Function, that is, predicting the cumulative rewards that the agent can obtain starting from a certain state and following a specific policy.
[0172] The policy network and value network play a key role in the learning process of the agent.
[0173] Policy Network: A neural network used to estimate the policy. The policy refers to the probability distribution of the actions that the agent takes in a given state. The policy network can help the agent achieve a dynamic decision-making strategy, thus learning and making decisions more effectively in the environment. It usually adopts a deep neural network structure and can handle continuous or discrete action spaces.
[0174] Value Network: A neural network used to estimate state values. The state value refers to the expected value of the cumulative reward (Q-value) after an agent takes a certain action in a given state. The value network can help the agent understand the value of each state, thus enabling better decision-making. It usually adopts a deep neural network structure and can handle continuous or discrete state spaces.
[0175] 5. Train the policy network and the value network;
[0176] 5-1. The policy network is used to generate actions based on the input state data. It is configured with an input layer, an output layer, and hidden layers. The input layer is used to input the state set composed of the comprehensive health index sequence and the multi-time scale matrix sequence. The output layer is used to output the risk level and the abnormal type of the device. The two hidden layers use the activation function ReLU to perform non-linear transformation on the input data.
[0177] For example, the data dimension of the input layer of the policy network is set to 6, and the input data is a feature vector group (6 dimensions of operating environment data, hardware performance data, system performance data, software performance data, current operating alarm data, and historical operating alarm data). Both hidden layers contain 64 neurons and use the activation function ReLU for non-linear transformation. The data dimension of the output layer is set to 2, which is used to output two-dimensional data: the risk level f and the abnormal type m. The policy network realizes the mapping from state input to action output.
[0178] In the embodiment of the present invention, the risk level f output by the policy network includes:
[0179] Low risk, indicating that the device is operating normally and there are no obvious fault signs in the short term;
[0180] Medium risk, indicating that the device may have potential problems, but will not cause serious faults in the short term;
[0181] High risk, indicating that the device has obvious fault signs or abnormalities and may malfunction in the short term;
[0182] Emergency, indicating that the device state is very bad, a fault is about to occur, and immediate action is required.
[0183] The abnormal type m output by the policy network includes but is not limited to: the fan fails, resulting in too high / low environmental temperature and too high / low environmental humidity, excessive hardware erasure times, hardware aging, system anomalies, cpu anomalies, too high software memory, software process restart, and so on.
[0184] Exemplarily:
[0185] m = 0, indicating that the device is operating normally;
[0186] m = 1001 indicates that the abnormal fan speed leads to overheating;
[0187] m = 1002 indicates that the abnormal fan speed leads to low temperature;
[0188] m = 1003 indicates that the continuous high - speed operation of the device leads to overheating;
[0189] m = 1004 indicates that the fan failure leads to high humidity;
[0190] m = 1005 indicates that the fan failure leads to low humidity; ...
[0192] m = 2001 indicates that the memory card has been read and written too many times;
[0193] m = 2002 indicates hardware aging and attenuation; ...
[0195] m = 3001 indicates that overheating leads to a high CPU occupancy rate;
[0196] m = 3002 indicates an abnormal i / o interface; ...
[0198] m = 4001 indicates that the abnormal operation of the device leads to a high CPU occupancy rate;
[0199] m = 4002 indicates that the process restarts repeatedly; ...
[0201] m = 5001 indicates a link interruption;
[0202] m = 5002 indicates data packet loss;
[0203] m = 5002 indicates data error; ...
[0205] m = 6001 indicates that the device is offline;
[0206] m = 6002 indicates that the device data writing fails;
[0207] m = 6002 indicates that the device protection switching fails; ...
[0209] In this embodiment, some values of the exception type m and their meaning expressions are given. However, in actual applications, the exception type m not only includes the above examples. The operation and maintenance personnel can also expand it according to the actual situation, and they are not listed one by one here.
[0210] Taking the comprehensive health index sequence and the multi-time scale matrix sequence as inputs, and the risk level f and the anomaly type m as outputs, train the policy network, and construct an experience replay pool according to the empirical data obtained during the training process, including:
[0211] 1) Initialize the value network and the policy network;
[0212] 2) Select the current state S(t) for the agent, where the state S(t) is the set of states composed of the comprehensive health index sequence and the multi-time scale matrix sequence at time t;
[0213] 3) The policy network outputs an action a according to the input state S(t), specifically the risk level f and the anomaly type m;
[0214] 4) The agent obtains the next state S(t + 1) through the interaction between the action a and the environment, and calculates the comprehensive reward r according to the reward function t ;
[0215] The comprehensive reward r t = r + λ·EHI(t);
[0216] where r represents the immediate reward for the agent to execute the current action a in the state S(t), and λ represents the weight of the comprehensive health index EHI(t);
[0217] 5) Store the empirical data [S(t), a, r t , S(t + 1)] into the experience replay pool.
[0218] 5-2. Extract empirical data;
[0219] Adopt the prioritized experience replay method to extract empirical data from the experience replay pool. By marking the priority of each piece of empirical data, non-uniform sampling is achieved. Prioritized experience replay can effectively reduce the number of experiences required for learning and reduce the number of interactions with the environment.
[0220] The traditional Q-learning formula:
[0221] Q(S, a) ← Q(S, a) + α[r + γmax a′ Q(S, a′) - Q(S, a)]
[0222] It mainly relies on the immediate reward r (usually the running performance of the device at a certain moment, such as reduced power consumption, improved performance, etc.) to update the Q value. This kind of reward is usually based on short-term benefits and performance. For example, the device saves power and improves efficiency at a certain moment, but this kind of immediate reward ignores the long-term health status of the device. If the model only focuses on short-term benefits, it may lead to increased wear and tear of the device over a long period, a gradual deterioration of the health status, and ultimately serious failures. Since the reward mechanism of traditional Q-learning often relies on the immediate feedback after the device fails or its performance significantly deteriorates, this means that the model can only learn and adjust after the failure occurs and cannot prevent device failures in advance. Moreover, before the device fails or issues an alarm, the traditional reinforcement learning model lacks effective feedback signals. This problem of feedback sparsity leads to low training efficiency of the model, poor learning effects, and difficulty in quickly converging to the optimal strategy.
[0223] As Figure 7 shown, the original Q value calculated by using the traditional Q-learning method gradually increases with the training time, representing that the accuracy of the model gradually improves with training, but the overall fluctuation is large and the convergence is slow.
[0224] In the implementation of the present invention, the device health index is introduced into Q-learning to solve the problem of sparse feedback in model training, and when the device health is low caused by various factors, it can more detailedly analyze the device failure factors and give corresponding operation and maintenance measures or suggestions.
[0225] 5-3. Update network parameters;
[0226] 1) Extract experience data [S(t), a, r t , S(t + 1)] from the experience replay pool to train the value network, and the value network updates the Q value based on the improved Q value function;
[0227] The improved Q value function is:
[0228] Q(S(t), a) ← Q(S(t), a) + α[(r + λ·EHI(t)) + γmax a′ Q(S(t + 1), a′) - Q(S(t), a)]
[0229] where Q(S(t), a) represents the Q value when the agent executes the current action a in state S(t);
[0230] Q(S(t + 1), a′) represents the Q value when the agent executes the next action a′ in state S(t + 1);
[0231] α represents the learning rate;
[0232] γ represents the reward discount factor;
[0233] max a′ Q(S(t + 1), a′) represents the maximum Q-value among all possible actions taken at state S(t + 1).
[0234] As Figure 8 and Figure 9 shown, the updated Q-value calculated using the improved Q-value function gradually increases with the training time, indicating that the model accuracy gradually improves with the training time. Compared with the original Q-value, the updated Q-value converges faster.
[0235] 2) The value network calculates the value of the current action at state S(t), i.e., the current Q-value. The target network is used to calculate the Q-value of the next state, i.e., the target Q-value. The parameters of the value network are updated by minimizing the loss function of the value network;
[0236] The loss function of the value network is the mean squared error between the current Q-value and the target Q-value;
[0237] Calculate the TD error based on the current Q-value and the target Q-value:
[0238] TD = α[(r + θ·EHI(t)) + γmax a′ Q(S(t + 1), a′) - Q(S(t), a)]
[0239] The parameters of the value network are updated by minimizing the mean squared error loss value of the TD error.
[0240] 3) Update the parameters of the policy network through the policy gradient method;
[0241] A policy is a rule by which an agent selects actions based on the current state. The policy gradient method is to define a parameter, such as using a neural network to represent the policy, and then directly adjust these parameters to maximize the cumulative reward obtained by the policy in the environment.
[0242] The policy gradient method uses gradient ascent to optimize the policy. The policy gradient is calculated based on the rewards obtained by executing the policy in the environment. This gradient represents the direction of parameter change, making the policy adjust in the direction that can obtain higher rewards. Simply put, if the current policy obtains a higher reward when executed in the environment, the adoption rate of this policy is increased by adjusting the parameters, making it more likely to be adopted in the future. Conversely, if the reward is low, the probability of this policy being adopted is reduced by adjusting the parameters. The policy gradient method gradually learns the optimal policy by continuously calculating the policy gradient and updating the parameters.
[0243] 4) Through iterative training, until the policy network and the value network converge or reach the preset number of training times, the trained policy network and value network are obtained, and the optimal decision-making model is generated.
[0244] 6. Equipment Risk Prediction and Automatic Operation and Maintenance
[0245] Regularly update the parameters of the optimal decision-making model to adapt to the latest changes in the equipment. Preprocess the real-time status data of the equipment and input it into the optimal decision-making model, and generate an operation and maintenance report according to the probability distribution of the risk level and abnormal type output by the policy network.
[0246] Calculate the comprehensive health index and multi-time scale matrix based on the real-time status data or stage data of the equipment, and input them into the policy network. If the policy network predicts that the equipment is currently in "medium risk" or "high risk", generate an alarm notification in the operation and maintenance report. If it is predicted that the equipment status will deteriorate further and the risk level may be upgraded, update the evaluation results in real time, issue a higher-level alarm and give real-time feedback.
[0247] For example, according to the policy network prediction, if the current risk level of the equipment is high risk, the abnormal type is that the temperature is too high due to the failure of the fan speed regulation, and the CPU is overloaded due to the high temperature, the operation and maintenance report generates an alarm notification, and it is judged that the performance trend of the equipment is that the temperature and CPU usage have continued to rise in the past week, exceeding the normal range. The automatic operation and maintenance model outputs the recommended operation and maintenance operations: first check the operation of the fan, find the reason for the speed regulation failure and perform maintenance if necessary.
[0248] In the embodiment of the present invention, by calculating the comprehensive health index sequence and multi-time scale matrix sequence and combining with the training of the deep reinforcement learning model, the operation and maintenance strategy of the equipment can be effectively optimized, the real-time monitoring and efficient management of the equipment can be realized, and a closed-loop system is formed from the process of data acquisition to model application, continuously improving the stability and performance of the system. It can also dynamically adapt to the changes in the equipment status, does not rely on the data at a certain moment, but takes into account the long-term influence of various external factors, prevents misjudging the health status of the equipment, and can more detailedly analyze the equipment failure factors and give corresponding operation and maintenance measures when the equipment health index is reduced due to various factors.
[0249] The second embodiment of the present invention also provides an operation and maintenance system based on deep reinforcement learning, including:
[0250] A data processing module, configured to collect multi-dimensional time-series data of the equipment and perform preprocessing, and construct a comprehensive health index sequence for representing the health status of the equipment and a multi-time scale matrix sequence for representing the operation status of the equipment.
[0251] A pre-training module, configured to use the comprehensive health index sequence and the multi-time scale matrix sequence as inputs, and use the risk level and abnormal type as outputs to train a policy network. The policy network outputs an action a based on the state S(t) at time t and feeds back a comprehensive reward r to the intelligent agent t, the agent interacts with the environment according to the action a to obtain the state S(t + 1) at the next moment, and stores the experience data [S(t), a, r t , S(t + 1)] in the experience replay pool;
[0252] The repeated training module is used to extract experience data from the experience replay pool, train the value network and the policy network. The value network updates the Q value based on the improved Q value function, updates the parameters of the value network by minimizing the loss function of the value network, and updates the parameters of the policy network by the policy gradient method to obtain the optimal decision-making model;
[0253] The operation and maintenance module is used to preprocess the real-time state data of the device and input it into the optimal decision-making model, and generate an operation and maintenance report according to the probability distribution of the risk level and abnormal type output by the policy network.
[0254] In this embodiment, the multi-dimensional time-series data of the device is the device state data collected from historical monitoring data or real-time monitoring data, including: operating environment data, hardware performance data, system performance data, software performance data, current operation alarm data, and historical operation alarm data.
[0255] The data processing module constructs a comprehensive health index sequence, specifically including:
[0256] Obtain the time-series data of multiple performance indicators related to the health state from the preprocessed multi-dimensional time-series data;
[0257] Calculate the health score of each performance indicator at time t:
[0258]
[0259] Calculate the comprehensive health index EHI(t) of the device at time t, which is used to represent the health state of the device:
[0260]
[0261] Among them, H i (t) is the health score of the i-th performance indicator at time t, and the value range is [0, 1];
[0262] x i (t) is the value of the i-th performance indicator at time t;
[0263] x i ref is the reference value of the i-th performance indicator;
[0264] σ i is the standard deviation of the i-th performance indicator;
[0265] wi is the weight of the i-th performance metric.
[0266] Furthermore, construct a multi-time scale matrix sequence, specifically including:
[0267] Obtain the time series data of multiple variables related to the running state from the preprocessed multi-dimensional time series data, and perform scale division according to a preset time window to obtain several multi-dimensional sequence segments;
[0268] Calculate the eigenvalues for each multi-dimensional sequence segment of each time window, including the mean, variance, and autocorrelation coefficient, and construct a feature matrix:
[0269]
[0270] where, F k is the feature matrix corresponding to the k-th time window;
[0271] μ i,k 、σ 2 i,k 、ρ i,k respectively represent the mean, variance, and autocorrelation coefficient of the i-th variable in the k-th time window;
[0272] Obtain the multi-time scale matrix sequence Z according to the feature matrices corresponding to each time window, and the multi-time scale matrix sequence is used to represent the running state of the device at different time scales:
[0273] Z = [F1, F2,......F m .
[0274] The risk levels f output by the policy network include:
[0275] Low risk, indicating that the device is running normally and there are no obvious signs of failure in the short term;
[0276] Medium risk, indicating that the device may have potential problems, but will not cause serious failures in the short term;
[0277] High risk, indicating that the device has obvious signs of failure or abnormalities and may fail in the short term;
[0278] Emergency, indicating that the device state is very bad, a failure is about to occur, and immediate action is required.
[0279] The abnormal types m output by the policy network include but are not limited to:
[0280] m = 0, indicating that the device is running normally;
[0281] m = 1001, indicating that the fan speed is abnormal, resulting in too high temperature;
[0282] m = 1002 indicates that the abnormal fan speed leads to too low temperature;
[0283] m = 1003 indicates that the continuous high-speed operation of the device leads to too high temperature;
[0284] m = 1004 indicates that the fan failure leads to too high humidity;
[0285] m = 1005 indicates that the fan failure leads to too low humidity;
[0286] m = 2001 indicates that the memory card has been read and written too many times;
[0287] m = 2002 indicates hardware aging attenuation;
[0288] m = 3001 indicates that too high temperature leads to a high CPU occupancy rate;
[0289] m = 3002 indicates an abnormal i / o interface;
[0290] m = 4001 indicates that the abnormal operation of the device leads to a high CPU occupancy rate;
[0291] m = 4002 indicates that the process restarts repeatedly;
[0292] m = 5001 indicates a link interruption;
[0293] m = 5002 indicates data packet loss;
[0294] m = 5003 indicates data error code;
[0295] m = 6001 indicates that the device is offline;
[0296] m = 6002 indicates that the device data writing fails;
[0297] m = 6003 indicates that the device protection switching fails.
[0298] In the pre-training module, using the comprehensive health index sequence and the multi-time scale matrix sequence as inputs, and the risk level f and the abnormal type m as outputs, train the policy network, and construct an experience replay pool according to the empirical data obtained during the training process, including:
[0299] 1. Initialize the value network and the policy network;
[0300] 2. Select the current state S(t) for the agent, where the state S(t) is the state set composed of the comprehensive health index sequence and the multi-time scale matrix sequence at time t;
[0301] 3. The policy network outputs an action a according to the input state S(t), specifically the risk level f and the abnormal type m;
[0302] 4. The agent interacts with the environment according to action a to obtain the next state S(t+1), and calculates the comprehensive reward r according to the reward function t ;
[0303] The comprehensive reward r t = r + λ·EHI(t);
[0304] where r represents the immediate reward of the agent when executing the current action a in state S(t), and λ represents the weight of the comprehensive health index EHI(t);
[0305] 5. Store the experience data [S(t), a, r t , S(t+1)] into the experience replay pool.
[0306] In the repeated training module:
[0307] 1) Extract the experience data [S(t), a, r t , S(t+1)] from the experience replay pool to train the value network, and the value network updates the Q value based on the improved Q value function;
[0308] The improved Q value function is:
[0309] Q(S(t), a) ← Q(S(t), a) + α[(r + λ·EHI(t)) + γmax a′ Q(S(t+1), a′) - Q(S(t), a)]
[0310] where Q(S(t), a) represents the Q value of the agent when executing the current action a in state S(t);
[0311] Q(S(t+1), a′) represents the Q value of the agent when executing the next action a′ in state S(t+1);
[0312] α represents the learning rate;
[0313] γ represents the reward discount factor;
[0314] max a′ Q(S(t+1), a′) represents the maximum Q value among all possible actions in state S(t+1).
[0315] 2) The value network calculates the value of the current action in state S(t), that is, the current Q value, and the target network is used to calculate the Q value of the state at the next moment, that is, the target Q value, and updates the parameters of the value network by minimizing the loss function of the value network;
[0316] The loss function of the value network is the mean square error between the current Q value and the target Q value;
[0317] Calculate the TD error based on the current Q value and the target Q value:
[0318] TD = α[(r + λ·EHI(t)) + γmax a′ Q(S(t + 1), a′) - Q(S(t), a)]
[0319] Update the parameters of the value network by minimizing the mean squared error loss value of the TD error.
[0320] 3) Update the parameters of the policy network by the policy gradient method.
[0321] A policy is a rule for an agent to select actions based on the current state. The policy gradient method defines a parameter, such as using a neural network to represent the policy, and then directly adjusts these parameters to maximize the cumulative reward obtained by the policy in the environment.
[0322] Regarding the system in the above embodiments, the specific ways in which each unit module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0323] The operation and maintenance decision-making model provided by the present invention realizes real-time monitoring and efficient management of devices. From the process of data acquisition to model application, a closed-loop system is formed, continuously improving the stability and performance of the system, and being able to dynamically adapt to changes in device states.
[0324] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, including: a memory, a processor, and the processor is configured to read and execute a computer program stored in the memory to implement the foregoing device operation and maintenance method.
[0325] Based on the same inventive concept, an embodiment of the present invention further provides a computer storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed, the foregoing device operation and maintenance method is implemented.
[0326] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical or other forms.
[0327] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module, that is, it may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, each functional module can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0328] If the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0329] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily all essential to the present invention.
[0330] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An operation and maintenance method based on deep reinforcement learning, characterized in that: The method comprises: Collect and preprocess the multi-dimensional time series data of the equipment, and construct a comprehensive health index sequence for representing the health status of the equipment and a multi-time scale matrix sequence for representing the operating status of the equipment; The comprehensive health index sequence and the multi-time scale matrix sequence are used as input, and the risk level and the abnormal type are used as output to train the policy network. The policy network outputs action a based on the state S(t) at time t, and the agent obtains the state S(t+1) at the next moment and the comprehensive reward r associated with the comprehensive health index according to the interaction between action a and the environment. t , the empirical data [S(t),a,r t ,S(t+1)] is stored in the experience replay pool; Extracting experience data from the experience replay pool, training the value network and the policy network, updating the Q value of the value network based on an improved Q value function, updating the parameters of the value network by minimizing the loss function of the value network, and updating the parameters of the policy network by a policy gradient method to obtain an optimal decision model; The real-time status data of the equipment is pre-processed and input into the optimal decision model, and an operation and maintenance report is generated according to the probability distribution of the risk level and abnormal type output by the strategy network.
2. The method according to claim 1, characterized in that The multi-dimensional time series data of the device is the device status data collected from historical monitoring data or real-time monitoring data, including: operating environment data, hardware performance data, system performance data, software performance data, current operation alarm data and historical operation alarm data.
3. The method according to claim 1, characterized in that Construct a comprehensive health index sequence to represent the health status of the equipment, including: Obtaining time series data of multiple performance indicators related to health status from the preprocessed multi-dimensional time series data; Calculate the health score of each performance indicator at time t: Calculate the comprehensive health index EHI(t) of the device at time t to indicate the health status of the device: Among them, H i (t) is the health score of the i-th performance indicator at time t, ranging from [0,1]; x i (t) is the value of the i-th performance indicator at time t; is the reference value of the i-th performance indicator; σ i is the standard deviation of the ith performance indicator; w i is the weight of the i-th performance indicator; Construct a multi-time scale matrix sequence to represent the equipment operation status, including: Obtaining time series data of multiple variables related to the operating status from the preprocessed multi-dimensional time series data, dividing the data into scales according to a preset time window, and obtaining a number of multi-dimensional sequence segments; The characteristic values of the multidimensional sequence segments in each time window are calculated separately, including the mean, variance, and autocorrelation coefficient, and the characteristic matrix is constructed: Among them, F k is the feature matrix corresponding to the kth time window; μ i,k , σ 2 i,k , i,k Respectively represent the mean, variance, and autocorrelation coefficient of the i-th variable in the k-th time window; A multi-time-scale matrix sequence Z is obtained according to the characteristic matrix corresponding to each time window. The multi-time-scale matrix sequence Z is used to represent the operating status of the device at different time scales: Z=[F1,F2,......F m ]。 4. The method according to claim 1, characterized in that: The risk level f output by the policy network includes: Low risk means the equipment is operating normally and there are no obvious signs of failure in the short term; Medium risk means that the equipment may have potential problems, but will not cause serious failures in the short term; High risk, indicating that the equipment has obvious signs of failure or abnormality and may fail in the short term; Emergency, indicating that the equipment is in very bad condition, failure is imminent, and immediate action is required; The abnormal types m output by the policy network include but are not limited to: m=0, indicating that the equipment is operating normally; m=1001, indicating that the temperature is too high due to abnormal fan speed; m=1002, indicating that the temperature is too low due to abnormal fan speed; m=1003, indicating that the device continues to run at high speed, resulting in excessive temperature; m=1004, indicating that the humidity is too high due to fan failure; m=1005, indicating that the humidity is too low due to fan failure; m=2001, indicating that the memory card has been read and written too many times; m = 2002, indicating hardware aging attenuation; m=3001, indicating that the CPU usage is high due to high temperature; m=3002, indicating an I / O interface abnormality; m=4001: indicates that the CPU usage is high due to abnormal operation of the device. m=4002, indicating that the process is restarted repeatedly; m=5001, indicating link interruption; m=5002, indicating data packet loss; m=5003, indicating data error; m=6001, indicating that the device is offline; m=6002, indicating that device data writing failed; m=6003, indicating that the equipment protection switching fails.
5. The method according to claim 1, characterized in that Comprehensive Rewards t =r+λ·EHI(t); Among them, r represents the immediate reward for executing the current action a in state S(t), and λ represents the weight of the comprehensive health index EHI(t); The improved Q value function is: Q(S(t),a)←Q(S(t),a)+α[(r+λ·EHI(t))+γmax a′ Q(S(t+1),a′)-Q(S(t),a)] Among them, Q(S(t),a) represents the Q value of the agent performing the current action a in state S(t); Q(S(t+1),a′) represents the Q value of the agent performing the next action a′ in state S(t+1); α represents the learning rate; γ represents the reward discount factor; max a′ Q(S(t+1), a′) represents the maximum Q value among all possible actions taken in state S(t+1).
6. The method according to claim 5, characterized in that The loss function of the value network is the mean square error between the current Q value and the target Q value; Calculate the TD error based on the current Q value and the target Q value: TD=α[(r+λ·EHI(t))+γmax a′ Q(S(t+1),a′)-Q(S(t),a)]。 7. An operation and maintenance system based on deep reinforcement learning, characterized in that: The system comprises: A data processing module is used to collect and preprocess the multi-dimensional time series data of the equipment, and to construct a comprehensive health index sequence for representing the health status of the equipment and a multi-time scale matrix sequence for representing the operating status of the equipment; A pre-training module is used to take the comprehensive health index sequence and the multi-time scale matrix sequence as input, and take the risk level and abnormality type as output to train a policy network, wherein the policy network outputs an action a based on the state S(t) at time t, and feeds back a comprehensive reward r to the agent. t , the agent obtains the next state S(t+1) according to the interaction between action a and the environment, and converts the experience data [S(t), a, r t ,S(t+1)] is stored in the experience replay pool; A repeated training module, used to extract experience data from the experience replay pool, train the value network and the policy network, update the Q value of the value network based on the improved Q value function, update the parameters of the value network by minimizing the loss function of the value network, and update the parameters of the policy network by the policy gradient method to obtain the optimal decision model; The operation and maintenance module is used to pre-process the real-time status data of the equipment and input it into the optimal decision model, and generate an operation and maintenance report according to the probability distribution of the risk level and abnormal type output by the policy network.
8. The system according to claim 7, characterized in that Comprehensive Rewards t =r+λ·EHI(t); Among them, r represents the immediate reward for executing the current action a in state S(t), and λ represents the weight of the comprehensive health index EHI(t); The improved Q value function is: Q(S(t),a)←Q(S(t),a)+α[(r+λ·EHI(t))+γmax a′ Q(S(t+1),a′)-Q(S(t),a)] Among them, Q(S(t),a) represents the Q value of the agent performing the current action a in state S(t); Q(S(t+1),a′) represents the Q value of the agent performing the next action a′ in state S(t+1); α represents the learning rate; γ represents the reward discount factor; max a′ Q(S(t+1), a′) represents the maximum Q value among all possible actions taken in state S(t+1).
9. The system according to claim 8, characterized in that The loss function of the value network is the mean square error between the current Q value and the target Q value; Calculate the TD error based on the current Q value and the target Q value: TD=α[(r+λ·EHI(t))+γmax a′ Q(S(t+1),a′)-Q(S(t),a)]。 10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions, when executed, implement the method according to any one of claims 1 to 6.