A model optimization method and device for equipment fault prediction
By integrating action confidence across multiple deep reinforcement learning networks, the method addresses the challenge of imbalanced data in fault prediction, improving robustness and accuracy in complex communication systems.
Patent Information
- Application Number
- CN202410377178.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-03-29
AI Technical Summary
Existing deep reinforcement learning models have limited capabilities when processing unbalanced data, resulting in insufficient robustness in fault prediction in industrial systems and are prone to overfitting or underfitting.
By integrating the action probability distribution and state action value function of multiple deep learning networks, the action confidence strategy integration method is used to calculate the action confidence and state action confidence, and optimize the model parameters to improve the robustness and adaptability of the model.
It enhances the adaptability of the model in the face of a diverse environment, avoids the limitations of a single strategy, improves the accuracy and reliability of failure prediction, and reduces the risk of poor performance of a single strategy in a specific environment.
Smart Images

Figure CN118211060B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and in particular, to a method and device for optimizing a model for equipment fault prediction. Background Art
[0002] With the continuous evolution of communication technologies, the scale and complexity of equipment have increased significantly, which has also posed a series of challenges. Problems such as equipment aging, network congestion, and software failures constantly threaten the normal operation of communication systems, potentially affecting the user experience and service quality. Traditional fault handling takes mitigation measures only after observing severe symptoms. Although this post-fault mitigation strategy will surely reduce the costs associated with unnecessary mitigation operations, users often have to endure a poor service experience before the mitigation operations take effect. Fault prediction has shifted from traditional passive operation and maintenance to proactive operation and maintenance, making operation and maintenance not just a response after problems occur, but an early prediction and proactive intervention of system problems. Through fault prediction, the system operation and maintenance team can identify potential problems earlier and take measures before the problems escalate into severe faults, thereby reducing service interruptions and performance degradation that users may face. This proactive operation and maintenance method enables the system to more flexibly cope with the increasingly complex communication environment. Fault prediction not only focuses on the current system state, but also on the challenges that the system may face in the future, enabling the operation and maintenance team to formulate more effective coping strategies. In addition, fault prediction also provides a powerful tool for system optimization. As shown in the fault prediction framework Figure 1 , by analyzing historical data, training a deep reinforcement learning model, and optimizing algorithms, the prediction results of equipment faults can be output to better understand system behavior and propose improvement suggestions. This iterative process enables communication systems to continuously evolve and better adapt to new technologies and user needs. Therefore, fault prediction is not only a means of problem response, but also an indispensable part of the development of communication technologies.
[0003] In industrial systems, equipment fault prediction is a key task, and the challenge lies in the general lack of handling uncertainty in traditional methods for fault prediction. This uncertainty covers the uncertainty of the deep reinforcement learning model itself and the dynamic uncertainty of the environment, thus significantly affecting the performance of the deep reinforcement learning model in practical applications. Currently, deep reinforcement learning models for fault prediction generally exhibit problems of insufficient robustness, especially with relatively limited ability to handle imbalanced data and being prone to falling into the dilemmas of overfitting or underfitting. Therefore, current technologies often fail to achieve satisfactory prediction effects when facing complex environments and uncertainties in industrial systems.
[0004] In view of this, how to overcome the defects of the existing technologies and solve the phenomenon that the existing deep reinforcement learning models have limited ability to handle imbalanced data is a problem to be solved in this technical field. Summary of the Invention
[0005] In view of the above deficiencies or improvement requirements of the prior art, the present invention solves the problem that the existing deep reinforcement learning model has limited ability to process unbalanced data.
[0006] The embodiments of the present invention adopt the following technical solutions:
[0007] In a first aspect, the present invention provides a model optimization method for equipment fault prediction, specifically: obtaining input data of the model according to equipment status information, inputting the input data into at least one group of deep learning networks in the model, and obtaining the action probability distribution and state-action value output by each group of deep learning networks; calculating the first action confidence of the corresponding deep learning network in the form of a ratio according to the action probability distribution; calculating the second action confidence of the corresponding deep learning network in the form of a ratio according to the state-action value; performing policy integration on the action probability distributions of all deep learning networks according to the first action confidence to obtain the overall action probability distribution of the model; performing policy integration on the state-action values of all deep learning networks according to the second action confidence to obtain the overall state-action value of the model; calculating the optimization parameters of the model according to the overall action probability distribution and the overall state-action value, and using the optimization parameters to optimize the model.
[0008] Preferably, the obtaining of the input data of the model according to the equipment status information specifically includes: processing the equipment status information into data in a unified format, preprocessing the unified data, mapping the preprocessed data to a standard overall distribution, and normalizing the mapped data; using the method of principal component analysis to extract features from the normalized data, and using the extracted vector group as the input data.
[0009] Preferably, each group of the deep learning networks includes a policy network and a value network. The obtaining of the action probability distribution and state-action value output by each group of deep learning networks specifically includes: the policy network maps the input data to the first output data through a probability distribution function, and uses the first output data as the action probability distribution; the value network maps the input data to the second output data through a state-action value function, and uses the second output data as the state-action value.
[0010] Preferably, according to the action probability distribution, the first action confidence of the corresponding deep learning network is calculated in the form of a ratio; according to the state-action value, the second action confidence of the corresponding deep learning network is calculated in the form of a ratio, which specifically includes: obtaining the first maximum value and the second maximum value of the action probability distribution, calculating the difference value between the first maximum value and the second maximum value, and using the ratio of the difference value to the first maximum value as the first action confidence corresponding to the action probability distribution; obtaining the second maximum value and the second second maximum value of the state-action value, calculating the difference value between the second maximum value and the second second maximum value, and using the ratio of the difference value to the second maximum value as the second action confidence corresponding to the state-action value.
[0011] Preferably, the action probability distributions of all deep learning networks are integrated according to the first action confidence to obtain the action probability distribution of the overall model; the state-action values of all deep learning networks are integrated according to the second action confidence to obtain the state-action value of the overall model, which specifically includes: using the first action confidence of each group of deep learning networks as the first weight, calculating the first weighted sum of the probability distribution functions of all deep learning networks according to the first weight, and taking the first weighted sum as the overall action probability distribution; using the second action confidence of each group of deep learning networks as the second weight, calculating the second weighted sum of the action value functions of all deep learning networks according to the second weight, and taking the second weighted sum as the overall state-action value.
[0012] Preferably, according to the overall action probability distribution and the overall state-action value, the optimization parameters of the model are calculated, and the model is optimized using the optimization parameters, which specifically includes: for each moment of the action-state transition, obtaining the second action confidence of the current moment, calculating the temporal error of the current moment based on the second action confidence of the current moment; calculating the priority sampling probability of the current moment based on the temporal error and the second action confidence of the current moment; updating the network parameters based on the temporal error and the priority sampling probability, and optimizing the model using the updated network parameters.
[0013] Preferably, calculating the temporal error of the current moment based on the second action confidence of the current moment specifically includes: calculating the product of the state value function of the next moment and a specified decay factor, obtaining the difference between the sum of the reward value of the next moment and the product and the state value function of the current moment; using the specified learning rate as the weight of the difference value, using the specified influence factor as the weight of the second action confidence of the current moment, calculating the weighted sum of the state-action value function of the current moment, the difference value and the second action confidence of the current moment, and taking the weighted sum as the temporal error.
[0014] Preferably, calculating the priority sampling probability at the current moment based on the timing error and the second action confidence at the current moment specifically includes: calculating a first value according to the timing error and the hyperparameters of priority sampling; calculating a second value according to the timing error, the hyperparameters of priority sampling, a specified influence factor, and the second action confidence at the current moment; and taking the ratio of the first value and the second value as the priority sampling probability.
[0015] Preferably, updating the network parameters based on the timing error and the priority sampling probability specifically includes: calculating the difference between the current estimated value and the target value according to the timing error, and taking the difference as the gradient for updating the neural network parameters; sampling the samples according to the priority sampling to improve the attention to the samples with larger timing errors during training.
[0016] On the other hand, the present invention provides a model optimization device for equipment fault prediction, specifically including: at least one processor and a memory, which are connected by a data bus. The memory stores instructions executable by the at least one processor. After the instructions are executed by the processor, they are used to complete the model optimization method for equipment fault prediction in the first aspect.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: Based on the action confidences of different deep reinforcement learning networks, integrating the action probability distributions and state-action value functions of multiple deep reinforcement learning networks enhances the adaptability of the model in dealing with diverse environments and avoids various limitations and deficiencies that a single strategy may face. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a schematic diagram of the fault prediction process in the prior art;
[0020] Figure 2 It is a schematic diagram of the model training process in the method provided by the embodiment of the present invention;
[0021] Figure 3 It is a flowchart of a model optimization method for equipment fault prediction provided by the embodiment of the present invention;
[0022] Figure 4 It is a flowchart of another model optimization method for equipment fault prediction provided by the embodiment of the present invention;
[0023] Figure 5 Schematic diagram of the deep learning network structure used in the embodiments of the present invention;
[0024] Figure 6 Flowchart of another model optimization method for device fault prediction provided by the embodiments of the present invention;
[0025] Figure 7 Schematic diagram of the data processing process of the policy network used in the embodiments of the present invention;
[0026] Figure 8 Schematic diagram of the structure of a model optimization device for device fault prediction provided by the embodiments of the present invention;
[0027] Wherein, the reference numerals are as follows:
[0028] 11: Processor; 12: Memory. Detailed implementation manners
[0029] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0030] The present invention is an architecture of a specific functional system. Therefore, in specific embodiments, the functional logical relationships of each structural module are mainly described, and specific software and hardware implementation manners are not limited.
[0031] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0032] Embodiment 1:
[0033] In order to solve the problems existing in using the existing deep reinforcement learning model for device fault prediction, such as Figure 2 shown, this embodiment provides a more adaptable deep reinforcement learning model and a model training optimization method. The model trained by this optimization method can more effectively handle uncertainties, improve the robustness of the model, and achieve more accurate and reliable prediction results in an unbalanced data scenario.
[0034] As Figure 3 shown, the specific steps of the model optimization method for device fault prediction provided by the embodiments of the present invention are as follows.
[0035] Step 101: Obtain the input data of the model according to the device status information, input the input data into at least one group of deep learning networks in the model, and obtain the action probability distribution and state-action value output by each group of deep learning networks.
[0036] In order to train and optimize the deep reinforcement learning model and achieve better prediction results, it is first necessary to collect device status information and use the device status information as sample data for model training. During the sample data collection stage, it is necessary to obtain a sample data set from each device, continuously train the model using the sample data set, and continuously optimize the model during the training process to improve the accuracy and robustness of fault prediction, so as to ensure that the model adapts to the changing working environment and effectively capture potential fault patterns.
[0037] In order to complete the optimization of the model, it is also necessary to obtain the output of the model as the basis for optimization after each sample data is input into the model. In this embodiment, the output of the model is the action probability distribution and state-action value. In the method provided in this embodiment, in order to avoid the influence of unbalanced data on the prediction result, multiple groups of deep learning networks are used in each model to obtain the outputs of multiple groups of deep learning networks, and all the outputs are integrated to obtain the final optimization parameters and prediction data.
[0038] Step 102: Calculate the first action confidence of the corresponding deep learning network in the form of a ratio according to the action probability distribution; calculate the second action confidence of the corresponding deep learning network in the form of a ratio according to the state-action value.
[0039] In the method provided in this embodiment, in order to overcome the potential risks that may be brought by the poor performance of a single model in a specific fault prediction environment or task, a technical solution for integrating action confidence strategies is proposed. After obtaining the action probability distribution and state-action value of each deep learning network in the model, it is also necessary to calculate the action confidence of each deep learning network according to the output, and then integrate the action confidence of all deep learning networks to make up for the limitations of a single strategy and reduce the potential risks brought by its poor performance in a specific environment or task. Further, in order to avoid the excessive influence of extreme values on the action confidence, the first action confidence of the action probability distribution and the second action confidence of the state-action value are calculated in the form of a ratio respectively, so as to eliminate the influence of extreme values on the overall distribution.
[0040] Step 103: Integrate the action probability distributions of all deep learning networks according to the first action confidence to obtain the overall action probability distribution of the model; integrate the state-action values of all deep learning networks according to the second action confidence to obtain the overall state-action value of the model.
[0041] After obtaining the action confidence of each group of deep learning networks, it is also necessary to calculate the overall action confidence of the model based on the integration strategy provided in this embodiment, and adjust the output strategy according to the overall action confidence to obtain the overall action probability distribution and the overall state-action value, so as to synthesize the action probability distributions and state-action value functions of multiple networks and generate more credible decision results. At the same time, the corresponding sample data and output strategy can also be stored in the experience pool for subsequent model learning and iterative optimization.
[0042] In the method provided in this embodiment, the calculation method based on the action confidence strategy fusion realizes a more comprehensive decision-making process by integrating the action probability distributions and state-action value functions of multiple networks. When the strategy of a certain group of deep learning networks fails in a specific scenario, other strategies can complement each other, improving the robustness of the overall system and thus significantly enhancing the accuracy of fault prediction. This integration method not only enhances the adaptability of the model in dealing with diverse environments but also effectively avoids the limitations and deficiencies that a single strategy may face, bringing a more reliable and efficient solution approach to the field of fault prediction.
[0043] Step 104: Calculate the optimization parameters of the model according to the overall action probability distribution and the overall state-action value, and use the optimization parameters to optimize the model.
[0044] After obtaining the overall action probability distribution and the overall state-action value output by the model, the fault can be predicted based on the output data, and the actual fault alarm is compared with the fault prediction output by the strategy. Then, the corresponding temporal difference error is calculated, and the strategy is updated accordingly. Through this process, the system continuously optimizes the strategy to achieve higher fault prediction accuracy in similar states.
[0045] After going through steps 101 - 104 provided in this embodiment, the optimization of the deep reinforcement learning model for fault prediction can be completed.
[0046] As Figure 4 shown, the input data to be input into the model can be obtained in the following way.
[0047] Step 201: Process the device status information into data in a unified format, preprocess the unified data, map the preprocessed data onto a standard overall distribution, and normalize the mapped data.
[0048] In actual implementation, the device status information obtained from the system may include: network topology, operation logs, alarms / events, system performance status (traffic, temperature, optical power, CPU, etc.), and service configuration / status, etc. To facilitate the model to process the data, first, the collected device status information needs to be processed into a unified data format, and the device status information in the unified data format is used as the model input data. Then, preprocessing is performed on the device status information in the same data format. Available preprocessing methods include: removing invalid, incorrect, or duplicate data. Then, by calculating the difference between the standard deviation of the data points and the average value, the data is mapped onto the standard normal distribution. And the data of different types is scaled to the same scale for data normalization, and the data is converted into values between 0 and 1.
[0049] Step 202: Use the method of principal component analysis to perform feature extraction on the normalized data, and use the vector group obtained by the feature extraction as the input data.
[0050] Among the obtained device status information, the information in different dimensions has different degrees of influence on the prediction result when performing fault prediction. To reduce the data processing volume, the dimension that has the greatest influence on the prediction result can be selected as the input data through feature extraction. For example, use principal component analysis (PCA for short) for dimensionality reduction, reduce the dimension of the data, select the feature attribute that has the greatest influence on the target value from the existing features, and finally generate a vector group (s1, s2, s3, s4, s5) corresponding to the input of five types of information. In the general scenario of network communication device fault prediction, each dimension in the vector group can respectively correspond to the following device status information: (network topology, operation logs, alarms / events, system performance status, service configuration / status).
[0051] After going through steps 201 - 202 provided in this embodiment, the input data to be input into the model can be obtained.
[0052] In the model of this embodiment, multiple groups of deep learning networks are used for prediction. As Figure 5 shown, each group of deep learning networks contains a policy network and a value network. As Figure 6 shown, each group of deep learning networks can output the action probability distribution and the state-action value network in the following manner.
[0053] Step 301: The policy network maps the input data into the first output data through the probability distribution function, and uses the first output data as the action probability distribution.
[0054] The data processing process of the policy network is as Figure 7As shown, the dimension of the output layer of the policy network is set according to actual needs. In a general scenario of device fault prediction, the dimension of the output layer of the policy network is set to 2, corresponding to the fault prediction category and the recommended action respectively. The policy network realizes a mapping from input to output, mapping from a vector group of device state information to an action probability distribution of fault prediction information. When optimizing the model by the method provided in this embodiment, the network parameters of the policy network are continuously adjusted to optimize the correctness of the policy network so that it can achieve a better policy output.
[0055] Step 302: The value network maps the input data into second output data through the state-action value function and uses the second output data as the state-action value.
[0056] In order to evaluate each action in the policy network, it is also necessary to obtain the value of each state-action through the value network so as to adjust the policy network according to the state-action value and make the actions output by it reach a higher state-action value.
[0057] After steps 301 - 302 provided in this embodiment, the action probability distribution and the state-action value output by each group of deep learning networks can be obtained.
[0058] Furthermore, there is a hidden layer between each output layer of the policy network and the value network, and the hidden layer uses a specified activation function to perform a non-linear transformation on the input data. The number of neurons and the type of activation function of the hidden layer can be selected according to actual needs. In a certain actual scenario, the hidden layer contains 64 neurons, and between each layer, the ReLU activation function f(x) = max(0, x) is used for non-linear transformation.
[0059] In the existing fault prediction model, the action confidence is calculated by the difference between the maximum state-action value and the second-largest state-action value under the input data s, the maximum action probability It is obtained by the difference. When the maximum value and the second - largest value are equal, the confidence level will be 0, and the output strategy will be greatly affected. The method of simply calculating the difference between the maximum value and the second - largest value may ignore the distribution information of some data. In some cases, the maximum value and the second - largest value may be too extreme, resulting in their excessive contribution to the weight, while other data points are ignored. For example, consider a fault scenario in an optical transmission network. Improper configuration settings may cause the branch adaptation module to fail to meet the actual requirements, thus leading to module failure. At the same time, long - term high - load operation or too high environmental temperature may cause the branch adaptation module to overheat, thus triggering a failure. When these two factors exist simultaneously, they lead to the failure state of the branch adaptation module, making the difference between the maximum value and the second - largest value zero in the confidence level calculation. In this case, these two extreme values contribute too much to the overall strategy weight, and the calculated value of the action confidence level is instead 0, and the low weight leads to incorrect strategy output.
[0060] In the method provided in this embodiment, using the form of the denominator ratio can avoid this situation because it considers the distribution of all data points, thus calculating the weight more stably. The specific calculation method of the action confidence level is as follows: Obtain the first maximum value and the first second - largest value of the action probability distribution, calculate the difference value between the first maximum value and the first second - largest value, and use the ratio of the difference value to the first maximum value as the first action confidence level corresponding to the action probability distribution; Obtain the second maximum value and the second second - largest value of the state - action value, calculate the difference value between the second maximum value and the second second - largest value, and use the ratio of the difference value to the second maximum value as the second action confidence level corresponding to the state - action value. By calculating the action confidence level in this way, it is possible to avoid key information for fault prediction from obtaining a low weight.
[0061] In specific implementation, the action confidence level calculation formulas for the state - action value and the probability distribution P are defined as follows:
[0062]
[0063]
[0064] Among them, Q max(s|·) is the first maximum value of the state - action value under the input data s, P max(s|·) is the second maximum value of the action probability under the input data s, Q second(s|·) is the first second - largest value of the state - action value under the input data s, P second(s|·) is the second second - largest value of the action probability under the input data s. ε is a small positive number. Usually, a very small positive number is taken, such as 0.0001 or 0.001 to prevent the denominator from being zero.
[0065] In the calculation of action confidence, when operating on the difference value between the maximum value and the second-largest value, if the difference value is directly used, it may be affected by extreme values. In the method provided in this embodiment, by using a ratio form, the difference value is combined with the overall distribution information, and the resulting result is more representative. It can reduce the influence of extreme values caused by various factors in the case of the same fault, improve the stability of the calculation, and thus better adapt to various data distribution situations. Therefore, the action confidence calculation method using the ratio form is more robust and helps to prevent problems that may occur in numerical calculations.
[0066] In this embodiment, multiple deep learning networks are used to output policies. The input data needs to be input into n policy networks and n value networks in the model respectively, and n action probability distributions P(a|s) and state-action values are output. Then, the n policy action probability distributions and state-action values are integrated to obtain the overall policy action probability and state-action value of the model. Specifically: taking the first action confidence of each group of deep learning networks as the first weight, calculating the first weighted sum of the probability distribution functions of all deep learning networks according to the first weight, and taking the first weighted sum as the overall action probability distribution; taking the second action confidence of each group of deep learning networks as the second weight, calculating the second weighted sum of the action value functions of all deep learning networks according to the second weight, and taking the second weighted sum as the overall state-action value.
[0067] In actual implementation, the overall policy action probability distribution and state-action value output by the model are as follows:
[0068]
[0069]
[0070] Among them, is the i-th probability distribution function under the input data s, is the first action confidence of the i-th probability distribution network, which is used as the first weight in the above formula, is the i-th state-action value function under the input data s, is the second action confidence of the i-th state-action value function, which is used as the second weight in the above formula.
[0071] Through the above action probability distribution, the policy function output fault prediction results can be obtained: the fault type f and the recommended action a. In a certain actual scenario, f being 0 represents normal, 1 represents sub-channel out-of-frame, 2 represents management control module failure, 3 represents branch adaptation module failure, and so on. The action a is a subset of the set of operable operations.
[0072] After obtaining the output result, action sampling can be performed according to the policy action probability distribution P, transferring the model to the next state, storing the sample data in the database, and then updating the policy of the model based on the state-action value to optimize the model. Specifically: for each moment of action-state transition, obtain the second action confidence at the current moment, calculate the temporal error at the current moment based on the second action confidence at the current moment; calculate the priority sampling probability at the current moment based on the temporal error and the second action confidence at the current moment; update the network parameters based on the temporal error and the priority sampling probability, and optimize the model using the updated network parameters.
[0073] Among them, the specific calculation method of the temporal error is as follows: calculate the product of the state value function at the next moment and the specified decay factor, and obtain the difference between the sum of the reward value at the next moment and the product and the state value function at the current moment; use the specified learning rate as the weight of the difference, and use the specified influence factor as the weight of the second action confidence at the current moment, calculate the weighted sum of the state-action value function at the current moment, the difference, and the second action confidence at the current moment, and take the weighted sum as the temporal error.
[0074] In actual implementation, the temporal error is calculated as follows:
[0075] w = Q(S t ) + α[r t+1 + γQ(S t+1 ) - Q(S t )] + βψ t ;
[0076] Among them, Q(S t ) is the state-action value function in state S t , Q(S t+1 ) is the state-action value function in state S t+1 , α is the learning rate used to control the update step size, which is used as the weight of the difference in the above formula, r t+1 is the reward obtained at time t + 1, γ is the decay factor used to consider future rewards, β is the influence factor used to adjust the confidence, which is used as the weight of the action confidence at the current moment in the above formula, ψ t is the second action confidence at time t. By introducing action confidence, the importance of each action for prediction accuracy is considered. In this way, the model can more flexibly adjust the focus of learning, thereby improving the learning effect of key actions.
[0077] The specific calculation method of the priority sampling probability is as follows: calculate the first value according to the timing error and the hyperparameters of priority sampling; calculate the second value according to the timing error, the hyperparameters of priority sampling, the specified influence factor, and the second action confidence at the current moment; use the ratio of the first value to the second value as the priority sampling probability.
[0078] In actual implementation, the priority sampling probability is as follows:
[0079]
[0080] Among them, w is the timing error, ε is a small positive number used to prevent division by zero, η is the hyperparameter of priority sampling for adjusting the priority sampling probability, j is the index of the sample, β is used to adjust the influence of confidence, that is, β is the specified influence factor, and ψ t is the action confidence at time t. In this embodiment, by introducing the action confidence, the system is allowed to more flexibly adjust the learning priority when dealing with abnormal situations such as large changes in action confidence, and avoid over-relying on abnormal actions to affect the learning process.
[0081] The way to update the network parameters is specifically as follows: calculate the difference between the current estimated value and the target value according to the timing error, and use the difference as the gradient for updating the neural network parameters to improve the prediction accuracy; sample the samples according to the priority sampling to increase the attention to the samples with larger timing errors during training, thereby improving the learning efficiency.
[0082] In actual implementation, the network parameter update is as follows:
[0083]
[0084] Among them, θ is the network parameter, α is the learning rate, is the gradient of the loss function with respect to the network parameter.
[0085] In specific implementation, the gradient of the loss function with respect to the network parameter can be obtained by calculating the partial derivative of the loss function with respect to the network parameter. In specific deep learning frameworks such as TensorFlow or PyTorch, automatic differentiation tools (Autograd) are usually used to automatically calculate these gradients.
[0086] Through this process, the model continuously optimizes the policy in the policy network to achieve higher fault prediction accuracy in similar states.
[0087] The model optimization method for equipment fault prediction provided in this embodiment provides a more comprehensive perspective for the decision-making process by integrating the action probability distributions and state-action value functions of multiple deep learning networks. It overcomes the limitations of a single strategy and aims to reduce the potential risks that may arise due to the poor performance of a certain strategy in a specific environment or task. When a certain strategy fails in a specific situation, other strategies can be relied on to complement each other, thus significantly improving the robustness of the overall system. This method not only enhances the adaptability of the model in dealing with diverse environments but also effectively avoids various limitations and deficiencies that a single strategy may face, and can handle critical fault prediction tasks in industrial systems.
[0088] Embodiment 2:
[0089] Based on the model optimization method for equipment fault prediction provided in the above Embodiment 1, the present invention further provides a model optimization device for equipment fault prediction that can be used to implement the above method, as Figure 8 shown, which is a schematic diagram of the device architecture of an embodiment of the present invention. The model optimization device for equipment fault prediction in this embodiment includes one or more processors 11 and a memory 12. Among them, Figure 8 One processor 11 is taken as an example here.
[0090] The processor 11 and the memory 12 can be connected through a bus or other means, Figure 8 Taking connection through a bus as an example here.
[0091] The memory 12, as a non-volatile computer-readable storage medium for the model optimization method for equipment fault prediction, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the model optimization method for equipment fault prediction in Embodiment 1. The processor 11 executes various functional applications and data processing of the model optimization device for equipment fault prediction by running the non-volatile software programs, instructions, and modules stored in the memory 12, that is, implements the model optimization method for equipment fault prediction in Embodiment 1.
[0092] The memory 12 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 12 may optionally include a memory remotely located relative to the processor 11, and these remote memories can be connected to the processor 11 through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0093] The program instructions / modules are stored in the memory 12 and, when executed by one or more processors 11, perform the model optimization method for device fault prediction in the above-mentioned Embodiment 1. For example, perform each of the steps shown in Figure 3 , Figure 4 and Figure 6 above.
[0094] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, which can include: Read Only Memory (ROM), Random Access Memory (RAM), magnetic disk or optical disk, etc.
[0095] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A model optimization method for equipment fault prediction, characterized in that Including: Obtain the input data of the model according to the device status information, input the input data into at least one set of deep learning networks in the model, and obtain the action probability distribution and state-action value output by each set of deep learning networks; wherein, in the scenario of network communication device fault prediction, the device status information obtained from the system includes: network topology, operation logs, alarms, events, system performance status, service configuration, and service status; among the obtained device status information, the information in different dimensions has different degrees of influence on the prediction result during fault prediction. To reduce the amount of data processing, the dimension that has the greatest influence on the prediction result is selected as the input data through feature extraction. Calculate the first action confidence of the corresponding deep learning network in the form of a ratio according to the action probability distribution; calculate the second action confidence of the corresponding deep learning network in the form of a ratio according to the state-action value; wherein, the second action confidence First action confidence Among them, Q max(s|·) is the first maximum value of the state-action value under the input data s, and P max(s|·) is the second maximum value of the action probability under the input data s, Q second(s|·) is the first large value of the state-action value under the input data s, and P second(s|·) is the second large value of the action probability under the input data s, and ε is a small positive number; Perform policy integration on the action probability distributions of all deep learning networks according to the first action confidence to obtain the overall action probability distribution of the model; perform policy integration on the state-action values of all deep learning networks according to the second action confidence to obtain the overall state-action value of the model. Calculate the optimization parameters of the model according to the overall action probability distribution and the overall state-action value, and use the optimization parameters to optimize the model. Among them, through the above action probability distribution, the fault type f and the recommended action a are obtained, where f being 0 represents normal, 1 represents sub-channel out-of-frame, 2 represents management control module failure, and 3 represents tributary adaptation module failure; the action a is a subset of the set of operations that can be performed; perform action sampling according to the action probability distribution, transfer the model to the next state, store the sample data in the database, and then update the policy of the model according to the state-action value to achieve the optimization of the model; continuously optimize the model during the training process to improve the accuracy and robustness of fault prediction, so as to ensure that the model adapts to the changing working environment and effectively capture potential fault patterns.
2. The model optimization method for equipment fault prediction according to claim 1, characterized in that The obtaining the input data of the model according to the device status information specifically includes: Process the device status information into data in a unified format, preprocess the unified data, map the preprocessed data to a standard overall distribution, and normalize the mapped data. Use the method of principal component analysis to extract features from the normalized data, and use the extracted vector group as the input data.
3. The model optimization method for equipment fault prediction according to claim 1, characterized in that, Each set of the deep learning networks includes a policy network and a value network. The obtaining the action probability distribution and state-action value output by each set of deep learning networks specifically includes: The policy network maps the input data to the first output data through a probability distribution function, and uses the first output data as the action probability distribution. The value network maps the input data to the second output data through a state-action value function, and uses the second output data as the state-action value.
4. The model optimization method for equipment fault prediction according to claim 1, wherein The performing policy integration on the action probability distributions of all deep learning networks according to the first action confidence to obtain the overall action probability distribution of the model; performing policy integration on the state-action values of all deep learning networks according to the second action confidence to obtain the overall state-action value of the model specifically includes: Take the first action confidence of each group of deep learning networks as the first weight, calculate the first weighted sum of the probability distribution functions of all deep learning networks according to the first weight, and take the first weighted sum as the overall action probability distribution; Take the second action confidence of each group of deep learning networks as the second weight, calculate the second weighted sum of the action value functions of all deep learning networks according to the second weight, and take the second weighted sum as the overall state-action value.
5. The model optimization method for equipment fault prediction according to claim 1, wherein Calculate the optimization parameters of the model according to the overall action probability distribution and the overall state-action value, and use the optimization parameters to optimize the model. Specifically, it includes: For each moment of action-state transition, obtain the second action confidence at the current moment, and calculate the temporal error at the current moment based on the second action confidence at the current moment; Calculate the priority sampling probability at the current moment based on the temporal error and the second action confidence at the current moment; Update the network parameters based on the temporal error and the priority sampling probability, and use the updated network parameters to optimize the model.
6. The model optimization method for equipment fault prediction according to claim 5, wherein The calculating the temporal error at the current moment based on the second action confidence at the current moment specifically includes: Calculate the product of the state value function at the next moment and the specified discount factor, and obtain the difference between the sum of the reward value and the product at the next moment and the state value function at the current moment; Take the specified learning rate as the weight of the difference, take the specified influence factor as the weight of the second action confidence at the current moment, calculate the weighted sum of the state-action value function at the current moment, the difference, and the second action confidence at the current moment, and take the weighted sum as the temporal error.
7. The model optimization method for equipment fault prediction according to claim 5, wherein, The calculating the priority sampling probability at the current moment based on the temporal error and the second action confidence at the current moment specifically includes: Calculate the first value according to the temporal error and the hyperparameters of priority sampling; Calculate the second value according to the temporal error, the hyperparameters of priority sampling, the specified influence factor, and the second action confidence at the current moment; Take the ratio of the first value and the second value as the priority sampling probability.
8. The model optimization method for equipment fault prediction according to claim 5, wherein The updating the network parameters based on the temporal error and the priority sampling probability specifically includes: Calculate the difference between the current estimated value and the target value according to the temporal error, and take the difference as the gradient for updating the neural network parameters; Sample the samples according to the priority sampling to improve the attention to the samples with larger temporal errors during training.
9. A model optimization device for equipment fault prediction, characterized in that: It includes at least one processor and a memory, the at least one processor and the memory are connected through a data bus, the memory stores instructions that can be executed by the at least one processor, and after the instructions are executed by the processor, they are used to complete the model optimization method for equipment fault prediction according to any one of claims 1-8.
Citation Information
Patent Citations
Integration-based cooperative multi-agent deep reinforcement learning method
CN116468107A
Information system anomaly detection method, device, equipment and medium
CN117033048A