Aluminum electrolysis fault identification method based on deep reinforcement learning and automatic parameter adjustment

By using deep reinforcement learning to automatically adjust the parameters of convolutional neural networks in aluminum electrolytic fault recognition, the problem of low accuracy in aluminum electrolytic fault recognition is solved, and high-precision fault recognition and model structure optimization are achieved.

CN114595750BActive Publication Date: 2025-05-13XIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210186157.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-05-13
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

In the prior art, the accuracy of aluminum electrolytic fault data identification is not high, and the fault recognition model structure is difficult to determine.

Method used

The aluminum electrolytic fault identification method based on deep reinforcement learning parameters is adopted. Through a one-dimensional convolutional neural network and a softmax classifier, the aluminum electrolytic data is extracted and classified, and the network structure is automatically optimized to improve the recognition accuracy.

Benefits of technology

It improves the accuracy of aluminum electrolytic fault identification, optimizes the recognition accuracy, simplifies the network parameter adjustment process, and saves computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114595750B_ABST
    Figure CN114595750B_ABST
Patent Text Reader

Abstract

The aluminum electrolysis fault identification method based on deep reinforcement learning parameter automatic adjustment disclosed by the present invention comprises the following steps: step 1, processing the collected aluminum electrolysis data; step 2, selecting a one-dimensional convolutional neural network as the network model, and determining the classifier as softmax; step 3, optimizing and determining the network structure of the convolutional neural network model; step 4, training the classification model; step 5, judging the termination condition; step 6, inputting the sample to be identified into the classification; step 7, calculating the recognition rate. The aluminum electrolysis fault identification method based on deep reinforcement learning parameter automatic adjustment of the present invention solves the problem that the accuracy of aluminum electrolysis fault data identification in the prior art is not high and the fault identification model structure is difficult to determine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of aluminum electrolysis, and in particular relates to an aluminum electrolysis fault identification method based on automatic adjustment of deep reinforcement learning parameters. Background Art

[0002] The aluminum electrolysis industry is one of my country's strategically important basic industries. With the continuous development of modern automation industry, the production process and technology of aluminum are becoming more and more complicated, and the possibility of failure is also increasing. Small failures often trigger chain reactions, which have a serious impact on the national economy and national security.

[0003] In the aluminum electrolysis production process, the aluminum electrolysis cell is accompanied by complex physical and chemical reactions. The industry often uses monitoring systems to monitor the status of the electrolysis cell, collect data from the electrolysis cell and other related equipment, and conduct modeling and analysis to determine whether an electrolysis failure has occurred. Faced with massive data, it is very difficult to identify electrolysis faults using traditional manual feature extraction methods, and the recognition effect is not ideal. Therefore, researching and using advanced theories and methods to accurately identify data samples has become a new problem facing aluminum electrolysis fault identification. To address this problem, a method for aluminum electrolysis fault identification based on automatic adjustment of reinforcement learning parameters is used to extract features and classify the collected data to effectively solve the problem of aluminum electrolysis fault identification.

[0004] The convolutional neural network was proposed by Yann Lecun in 1998. Its main function is to perform convolution operations on input data by setting convolution kernels, and then extract more abstract and comprehensive data features by increasing the number of convolution kernel layers. It has a wide range of applications in data feature extraction problems.

[0005] Deep reinforcement learning was proposed by Mnih V. et al. in 2015. Its principle is to introduce deep neural networks into the reinforcement learning process to replace the traditional reinforcement learning policy function for action selection. The intelligent agent gradually optimizes the strategy by continuously interacting with the environment, and finally obtains an optimal strategy to maximize the benefits of completing the decision-making process. It has a good ability to solve dynamic decision-making problems.

[0006] Due to the large number of electrolytic equipment and complex processes, the dimension of data features is increasing. The identification of fault samples is mainly to extract target features. Only better feature extraction can achieve better recognition results. However, the existing recognition methods fail to fully utilize this feature of aluminum electrolysis data. In addition, most of the current recognition methods set related parameters based on empirical values ​​only, and fail to exert the optimal performance of the algorithm. Therefore, the accuracy of aluminum electrolysis data recognition is not high. Summary of the invention

[0007] The purpose of the present invention is to provide an aluminum electrolysis fault identification method based on automatic adjustment of deep reinforcement learning parameters, which solves the problems in the prior art that the accuracy of aluminum electrolysis fault data identification is low and the fault identification model structure is difficult to determine.

[0008] The technical solution adopted by the present invention is an aluminum electrolysis fault identification method based on automatic adjustment of deep reinforcement learning parameters, comprising the following steps:

[0009] Step 1, processing the collected aluminum electrolysis data;

[0010] Step 2: Select a one-dimensional convolutional neural network as the network model and set the classifier to softmax;

[0011] Step 3: Optimize and determine the network structure of the convolutional neural network model;

[0012] Step 4: Train the classification model

[0013] Training the classification model is divided into network initialization, data feature extraction, establishing the objective function, and optimizing parameters using the gradient descent method;

[0014] Step 5: Determine the termination condition

[0015] Set the number of consecutive occurrences of optimal performance to N best , determine whether the number of consecutive occurrences of the optimal performance of the recognition network reaches the maximum number of times I, if not reached, return to step 3 to re-optimize the structural parameters; if satisfied, proceed to step 6;

[0016] Step 6: Input the sample to be identified into the classification

[0017] Input the aluminum electrolysis data test samples into the trained network model, and use the trained network to classify the samples to be identified;

[0018] Step 7: Calculate the recognition rate

[0019] Formula (4) is used to calculate the optimal recognition accuracy of the network for aluminum electrolysis data:

[0020]

[0021] Among them, n is the total number of test samples, and b is the correctly classified samples to be identified.

[0022] The present invention is also characterized in that

[0023] The specific implementation of step 1 is:

[0024] The collected aluminum electrolysis data is processed into input signals that can be recognized by the network, the number of categories and sample dimensions of the samples to be recognized are determined, and the aluminum electrolysis data is divided into training samples and test samples according to a certain ratio. The processed training sample set is {(x (1) ,y (1) ),...,(x (m) ,y (m) )}, where m is the number of training samples, x (i) is the i-th training sample, y (i) ∈{1,2,...,k} is the label of the i-th training sample, k is the number of classes of samples to be identified in the aluminum electrolysis data; the processed test sample set is {x (1) ,x (2) ,...,x (n)}, where n is the number of training samples, x (i) is the ith test sample.

[0025] The specific implementation of step 2 is:

[0026] Step 2.1) Determine the network model as a one-dimensional convolutional neural network

[0027] Step 2.2) Determine the classifier as softmax

[0028] The classifier in this step uses the softmax classifier. When the training sample set is {(x (1) ,y (1) ),...,(x (m) ,y (m) )}, the softmax classifier is used to classify the features of the extracted data to be identified as follows:

[0029]

[0030] In which, it is assumed that the vector h θ (x (i) ) every element p(y (i) =j|x (i) ; θ) represents the feature x of the sample object to be identified (i) The probability of belonging to the jth class, j∈{1,2,…,k}, the larger the probability, the more likely the sample x to be identified is (i) The greater the probability of belonging to the jth class, where θ1,θ2,...,θ k is the parameter vector of the convolutional neural network.

[0031] The specific implementation of step 3 is:

[0032] 3.1) Determine the number of input layer nodes,

[0033] The number of nodes in the input layer of the convolutional neural network is related to the dimension d of the object data to be identified, and is the input of the feature number of the object to be identified;

[0034] 3.2) Determine the number of output layer nodes,

[0035] For the entire convolutional neural network model, the number of output layer nodes is the number k of classes of samples to be identified in the aluminum electrolysis data;

[0036] 3.3) Determine and optimize the hidden layer network structure.

[0037] The specific implementation of step 3.3 is:

[0038] 3.3.1) Network initialization,

[0039] Initialize the convolutional neural network using a set of hyperparameters, define a reasonable state space based on the set of hyperparameters, and discretize the state space;

[0040] 3.3.2) Building a deep reinforcement learning model,

[0041] Define the agent as a deep neural network with a three-layer fully connected structure, which is used to select actions according to the strategy under different states. The strategy is defined as the weight parameter of the neural network, and the state s is defined as a tuple of the hyperparameter combination of the convolutional neural network, as shown in (2).

[0042] s={e,f,o,h} (2)

[0043] Where e represents the size of the convolution kernel, f represents the number of convolution kernels, o represents the number of fully connected layers, and h represents the number of nodes in the fully connected layer. The optimal hyperparameter combination is defined, that is, the classification performance of the convolutional neural network reaches the optimal hyperparameter combination s*. Action a is defined as the adjustment of the current convolutional neural network hyperparameters, that is, within the state space defined in step 3.3.1, according to the guidance of the policy function, one of the hyperparameters is selected, and an adjustment direction is selected to adjust the parameter by one unit, and P is used. best represents the best performance of the classifier, N best represents the number of times the best performance appears continuously, and the maximum number of times the best performance appears continuously is set to 1;

[0044] 3.3.3) Dynamic decision making of intelligent agents,

[0045] When the agent is in state s, the ε-greedy algorithm is used to select action a to adjust a parameter in the hyperparameter combination by one unit in a certain direction. After completing action a, a new set of hyperparameter combinations s' is obtained, that is, the agent is transferred from state s to state s', and the fault recognition network is retrained in state s';

[0046] 3.3.4) Performance evaluation,

[0047] The environmental reward r is defined as the weighted sum of the recognition accuracy of the convolutional neural network for samples of different categories under the current parameter configuration. r is used to evaluate the performance of the fault recognition network in state s.

[0048] 3.3.5) Update N p ,

[0049] During the iteration process, the best performance of the classification model is recorded as P best , if r>P best , then N best =0, if r <P best , then let N best =N best +1, used to count the number of times the best performance occurs, and then determine whether the model converges to the best performance, that is, N best Whether the threshold I is reached;

[0050] 3.3.6) Update the agent strategy,

[0051] The classification result r of the convolutional neural network is used as the feedback value of the environment, and the agent continuously optimizes the strategy based on the feedback value;

[0052] 3.3.7) Determine whether the termination condition is met,

[0053] Determine the number of times the optimal performance appears consecutively N best Whether the preset maximum number of times I is reached, if not, return to 3.3.3), let N best =N best +1, and use the agent to re-determine the network structure parameters. If it has been reached, go to step 3.3.8);

[0054] 3.3.8) Output the optimal hyperparameter combination,

[0055] Output the optimal hyperparameter combination s for the fault identification network model * , and the optimal hyperparameter combination s * As the optimal structure of the convolutional neural network, it is substituted into subsequent training.

[0056] The specific implementation of step 4 is as follows:

[0057] 4.1) Network initialization,

[0058] Use Status * Define a one-dimensional convolutional neural network;

[0059] 4.2) Data feature extraction,

[0060] Feature extraction is completed by the convolution operation in the network. For all convolutional layers, the filling method is set to all zero filling, the moving step size of the convolution kernel is set to 1, the activation function is set to the relu function, and a batch normalization layer is set before each convolutional layer;

[0061] 4.3) Establish the objective function,

[0062] The objective function is the loss function Loss, which is defined as the mean square error of the classifier:

[0063]

[0064] Among them, y i represents the true category label of the i-th sample, represents the predicted category label of the fault recognition network for the i-th sample;

[0065] 4.4) Gradient descent method to optimize parameters,

[0066] Back propagation is used to calculate the gradient variables and optimize the weight parameters in the convolutional neural network to minimize the objective function.

[0067] The beneficial effects of the present invention are as follows: by processing and analyzing the aluminum electrolysis data, due to the high-dimensional characteristics of the processed data, a convolutional neural network is used for feature extraction and classification, and the recognition result is better than the traditional neural network (BP) and traditional classification methods (SVM, etc.); at the same time, considering that the structure and parameter configuration of the convolutional neural network have a great influence on the recognition accuracy, a deep reinforcement learning algorithm is introduced to realize the recognition of aluminum electrolysis data samples by the automatically adjusted convolutional neural network, so that the recognition accuracy is optimized. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a schematic diagram of the overall implementation process of the method of the present invention;

[0069] Figure 2 It is a simplified diagram of the convolutional neural network structure in the method of the present invention;

[0070] Figure 3 It is a simplified diagram of the application process of the deep reinforcement learning algorithm in the method of the present invention. DETAILED DESCRIPTION

[0071] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0072] The present invention provides an aluminum electrolysis fault identification method based on deep reinforcement learning parameter automatic adjustment, such as Figure 1-3 , including the following steps:

[0073] Step 1, processing the collected aluminum electrolysis data;

[0074] The specific implementation of step 1 is:

[0075] The collected aluminum electrolysis data is processed into input signals that can be recognized by the network, the number of categories and sample dimensions of the samples to be recognized are determined, and the aluminum electrolysis data is divided into training samples and test samples according to a certain ratio (7:3 or 8:2). The processed training sample set is {(x (1) ,y (1) ),...,(x (m) ,y (m) )}, where m is the number of training samples, x (i) is the i-th training sample, y (i) ∈{1,2,...,k} is the label of the i-th training sample, k is the number of classes of samples to be identified in the aluminum electrolysis data; the processed test sample set is {x (1) ,x (2) ,...,x (n)}, where n is the number of training samples, x (i) is the ith test sample.

[0076] Step 2: Select a one-dimensional convolutional neural network as the network model and set the classifier to softmax;

[0077] The specific implementation of step 2 is:

[0078] Step 2.1) Determine the network model as a one-dimensional convolutional neural network (1D_CNN)

[0079] Since neural networks have strong self-learning and fault-tolerant capabilities, they have obvious advantages over traditional machine learning algorithms in solving pattern recognition problems. In classification problems, accurate extraction of representative features of data is crucial. Since one-dimensional convolutional neural networks have unique advantages in feature extraction of serialized data and have the characteristics of weight sharing, they can accelerate the convergence of the model. Therefore, one-dimensional convolutional neural networks are used as network models.

[0080] Step 2.2) Determine the classifier as softmax

[0081] The classifier in this step uses the softmax classifier. When the training sample set is {(x (1) ,y (1) ),...,(x (m) ,y (m) )}, the softmax classifier is used to classify the features of the extracted data to be identified as follows:

[0082]

[0083] In which, it is assumed that the vector h θ (x (i) ) every element p(y (i) =j|x (i) ; θ) represents the feature x of the sample object to be identified (i) The probability of belonging to the jth class (vector h θ (x (i) )The sum of each element is 1), j∈{1,2,…,k}, the larger the probability, the more likely the sample x to be identified is. (i) The greater the probability of belonging to the jth class, where θ1,θ2,...,θ k is the parameter vector of the convolutional neural network.

[0084] Step 3: Optimize and determine the network structure of the convolutional neural network model;

[0085] The convolutional neural network model is mainly composed of two parts: convolutional layer and fully connected layer. The number of input layers is the number of features of the data to be classified, d, and the number of output layers is the number of sample classes to be identified in the aluminum electrolysis data, k.

[0086] The optimization of network structure mainly involves the optimization of convolutional layers and fully connected layers. Generally, the network model can be divided into input layer, hidden layer and output layer. Therefore, the determination of network structure is divided into three aspects: the number of input layer nodes, the number of output layer nodes and the optimization of hidden layer structure.

[0087] The specific implementation of step 3 is:

[0088] like Figure 2 As shown in the figure, the convolutional neural network model includes convolutional layers and fully connected layers, where the number of input layers is the number of features of the data to be classified d, and the number of output layers is the number of sample classes to be identified in the aluminum electrolysis data k.

[0089] The optimization of the network structure mainly involves the optimization of the hyperparameter configuration of the convolutional layer and the fully connected layer. Generally, the network model can be divided into three modules: the input layer, the hidden layer (i.e., the convolutional layer and the fully connected layer), and the output layer. Correspondingly, the optimization of the network structure can also be divided into three aspects: the number of input layer nodes, the number of output layer nodes, and the hidden layer structure optimization, as follows:

[0090] 3.1) Determine the number of input layer nodes,

[0091] The number of nodes in the input layer of the convolutional neural network is related to the dimension d of the object data to be identified, and is the input of the feature number of the object to be identified;

[0092] 3.2) Determine the number of output layer nodes,

[0093] For the entire convolutional neural network model, the number of output layer nodes is the number k of classes of samples to be identified in the aluminum electrolysis data;

[0094] 3.3) Determine and optimize the hidden layer network structure.

[0095] For traditional convolutional neural networks, the various hyperparameters of the hidden layer structure need to be manually set based on multiple experiments and experience, which is time-consuming, labor-intensive, and wastes computing resources. However, the quality of the hidden layer structure setting directly affects the performance of the network model. Therefore, it is necessary to determine and optimize the hidden layer structure of the network.

[0096] This step is based on the basic principles of deep reinforcement learning, and innovates a convolutional neural network algorithm based on automatic adjustment of deep reinforcement learning parameters, which realizes the determination of the hidden layer structure of the convolutional network, eliminates complicated manual adjustment steps, saves time, and saves precious computing resources; at the same time, it can effectively improve the recognition accuracy of aluminum electrolysis fault samples, and the implementation process is simple, with significant effects on fault identification problems.

[0097] Reference Figure 3 , the specific process of using deep reinforcement learning algorithm to optimize the hidden layer structure of convolutional neural network is:

[0098] The specific implementation of step 3.3 is:

[0099] 3.3.1) Network initialization,

[0100] A set of hyperparameters is used to initialize the convolutional neural network. In order to reduce the amount of algorithm calculation, a reasonable state space is defined based on this set of hyperparameters. At the same time, in order to ensure the convergence of the algorithm, the state space needs to be limited to a reasonable range and discretized. For example, the reasonable value range of the convolutional neural network learning rate is [0,1], so it cannot be adjusted to a negative value or a value greater than 1, and the value type of this hyperparameter is a floating point type. Therefore, during the execution of the algorithm, the hyperparameter needs to be discretized to avoid the situation where the state space is too large and the algorithm is difficult to converge;

[0101] 3.3.2) Building a deep reinforcement learning model,

[0102] The agent is defined as a deep neural network with a three-layer fully connected structure, which is used to select actions according to strategies in different states. The strategy is defined as the weight parameters of the neural network.

[0103] The state s is defined as a tuple of the hyperparameters of the convolutional neural network, as shown in (2).

[0104] s={e,f,o,h} (2)

[0105] Among them, e represents the size of the convolution kernel, f represents the number of convolution kernels, o represents the number of fully connected layers, and h represents the number of nodes in the fully connected layer. The optimal hyperparameter combination is defined as s * , that is, the classification performance of the convolutional neural network reaches the optimal hyperparameter combination s*, and action a is defined as the adjustment of the current convolutional neural network hyperparameters, that is, within the state space defined in step 3.3.1, according to the guidance of the policy function, one of the hyperparameters is selected, and an adjustment direction is selected to adjust the parameter by one unit, and P is used. best represents the best performance of the classifier, N best represents the number of times the best performance appears continuously, and the maximum number of times the best performance appears continuously is set to 1;

[0106] 3.3.3) Dynamic decision making of intelligent agents,

[0107] When the agent is in state s, the ε-greedy algorithm is used to select action a to adjust a parameter in the hyperparameter combination by one unit in a certain direction. After completing action a, a new set of hyperparameter combinations s' is obtained, that is, the agent is transferred from state s to state s', and the fault recognition network is retrained in state s';

[0108] 3.3.4) Performance evaluation,

[0109] The environmental reward r is defined as the weighted sum of the recognition accuracy of the convolutional neural network for samples of different categories under the current parameter configuration. r is used to evaluate the performance of the fault recognition network in state s.

[0110] 3.3.5) Update N p ,

[0111] During the iteration process, the best performance of the classification model is recorded as P best , if r>P best , then N best =0, if r <P best , then let N best =N best +1, used to count the number of times the best performance occurs, and then determine whether the model converges to the best performance, that is, N best Whether the threshold I is reached;

[0112] 3.3.6) Update the agent strategy,

[0113] The ultimate goal of agent training is to obtain the decision-making method that enables the classifier to achieve the best performance. The classification result r of the convolutional neural network is used as the feedback value of the environment, and the agent continuously optimizes the strategy based on the feedback value;

[0114] After a period of exploration and trial and error, the intelligent agent gradually forms an optimal strategy. When the intelligent agent is in a new state, it can autonomously choose an action that can bring the greatest benefit according to the learned strategy to adjust the hyperparameter combination, thereby obtaining a network structure with the best recognition performance.

[0115] 3.3.7) Determine whether the termination condition is met,

[0116] Determine the number of times N that the optimal performance appears consecutively best Whether the preset maximum number of times I is reached, if not, return to 3.3.3), let N best =N best +1, and use the agent to re-determine the network structure parameters. If it has been reached, go to step 3.3.8);

[0117] 3.3.8) Output the optimal hyperparameter combination,

[0118] Output the optimal hyperparameter combination s for the fault identification network model * , and the optimal hyperparameter combination s * As the optimal structure of the convolutional neural network, it is substituted into subsequent training.

[0119] Step 4: Train the classification model

[0120] Training the classification model is divided into network initialization, data feature extraction, establishing the objective function, and optimizing parameters using the gradient descent method;

[0121] The specific implementation of step 4 is as follows:

[0122] 4.1) Network initialization,

[0123] Use Status * Define a one-dimensional convolutional neural network;

[0124] 4.2) Data feature extraction,

[0125] Feature extraction is completed by the convolution operation in the network. In order to reduce model parameters and speed up model convergence, for all convolutional layers, the padding method (padding) is set to all zero padding, the movement stride (stride) of the convolution kernel is set to 1, and the activation function (activation) is set to the relu function. In order to reduce the amount of calculation, a batch normalization (BN) layer is set in front of each convolutional layer.

[0126] 4.3) Establish the objective function,

[0127] The objective function is the loss function Loss, which is defined as the mean square error of the classifier:

[0128]

[0129] Among them, y i represents the true category label of the i-th sample, Represents the predicted category label of the fault identification network for the i-th sample; the smaller the Loss, the better the performance of the fault identification network.

[0130] 4.4) Gradient descent method to optimize parameters,

[0131] Back propagation is used to calculate the gradient variables and optimize the weight parameters in the convolutional neural network to minimize the objective function.

[0132] Step 5: Determine the termination condition

[0133] Set the number of consecutive occurrences of optimal performance to N best , determine whether the number of consecutive occurrences of the optimal performance of the recognition network reaches the maximum number of times I, if not reached, return to step 3 to re-optimize the structural parameters; if satisfied, proceed to step 6;

[0134] Step 6: Input the sample to be identified into the classification

[0135] Input the aluminum electrolysis data test samples into the trained network model, and use the trained network to classify the samples to be identified;

[0136] Step 7: Calculate the recognition rate

[0137] In order to more intuitively represent the classification effect of the network on the samples to be identified, the optimal recognition accuracy of the network for aluminum electrolysis data is calculated using formula (4):

[0138]

[0139] Among them, n is the total number of test samples, and b is the correctly classified samples to be identified.

[0140] Advantages of the invention method The specific advantages include the following aspects:

[0141] 1) It has strong intelligence and can automatically adjust the network structure, eliminating the time-consuming process of manually adjusting network parameters, while saving computing resources and achieving the best classification accuracy.

[0142] 2) It is highly practical and has good recognition accuracy for fault identification problems in the aluminum electrolysis process, and the implementation process is simple.

[0143] 3) It has strong universal adaptability, not only in the identification of aluminum electrolysis faults, but also in other pattern recognition problems, and can still achieve relatively ideal recognition performance.

Claims

1. Aluminum electrolysis fault identification method based on deep reinforcement learning parameter automatic adjustment, characterized in that: The following steps are involved: Step 1, processing the collected aluminum electrolysis data; Step 2: Select a one-dimensional convolutional neural network as the network model and set the classifier to softmax; Step 3: Optimize and determine the network structure of the convolutional neural network model; The specific implementation of step 3 is: 3.1) Determine the number of input layer nodes, The number of nodes in the input layer of the convolutional neural network is related to the dimension d of the object data to be identified, and is the input of the feature number of the object to be identified; 3.2) Determine the number of output layer nodes, For the entire convolutional neural network model, the number of output layer nodes is the number k of classes of samples to be identified in the aluminum electrolysis data; 3.3) Determine and optimize the hidden layer network structure; The specific implementation of step 3.3 is: 3.3.1) Network initialization, Initialize the convolutional neural network using a set of hyperparameters, define a reasonable state space based on the set of hyperparameters, and discretize the state space; 3.3.2) Building a deep reinforcement learning model, The agent is defined as a deep neural network with a three-layer fully connected structure, which is used to select actions according to strategies in different states. The strategy is defined as the weight parameters of the neural network. The state s is defined as a tuple of the hyperparameters of the convolutional neural network, as shown in (2). s={e,f,o,h} (2) Where e represents the size of the convolution kernel, f represents the number of convolution kernels, o represents the number of fully connected layers, and h represents the number of nodes in the fully connected layer. The optimal hyperparameter combination is defined, that is, the classification performance of the convolutional neural network reaches the optimal hyperparameter combination s*. Action a is defined as the adjustment of the current convolutional neural network hyperparameters, that is, within the state space defined in step 3.3.1, according to the guidance of the policy function, one of the hyperparameters is selected, and an adjustment direction is selected to make a unit adjustment to the parameter, and P is used. best represents the best performance of the classifier, N best represents the number of times the best performance appears continuously, and the maximum number of times the best performance appears continuously is set to 1; 3.3.3) Dynamic decision making of intelligent agents, When the agent is in state s, the ε-greedy algorithm is used to select action a to adjust a parameter in the hyperparameter combination by one unit in a certain direction. After completing action a, a new set of hyperparameter combinations s' is obtained, that is, the agent is transferred from state s to state s', and the fault recognition network is retrained in state s'; 3.3.4) Performance evaluation, The environmental reward r is defined as the weighted sum of the recognition accuracy of the convolutional neural network for samples of different categories under the current parameter configuration. r is used to evaluate the performance of the fault recognition network in state s. 3.3.5) Update N p , During the iteration process, the best performance of the classification model is recorded as P best , if r>P best , then N best =0, if r <P best , then let N best =N best +1, used to count the number of times the best performance occurs, and then determine whether the model converges to the best performance, that is, N best Whether the threshold I is reached; 3.3.6) Update the agent strategy, The classification result r of the convolutional neural network is used as the feedback value of the environment, and the agent continuously optimizes the strategy based on the feedback value; 3.3.7) Determine whether the termination condition is met, Determine the number of times the optimal performance appears consecutively N best Whether the preset maximum number of times I is reached, if not, return to 3.3.3), let N best =N best +1, and use the agent to re-determine the network structure parameters. If it has been reached, go to step 3.3.8); 3.3.8) Output the optimal hyperparameter combination, Output the optimal hyperparameter combination s for the fault identification network model * , and the optimal hyperparameter combination s * As the optimal structure of the convolutional neural network, it is substituted into subsequent training; Step 4: Train the classification model Training the classification model is divided into network initialization, data feature extraction, establishing the objective function, and optimizing parameters using the gradient descent method; Step 5: Determine the termination condition Set the number of consecutive occurrences of optimal performance to N best , determine whether the number of consecutive occurrences of the optimal performance of the recognition network reaches the maximum number of times I, if not reached, return to step 3 to re-optimize the structural parameters; if satisfied, proceed to step 6; Step 6: Input classification of samples to be identified Input the aluminum electrolysis data test samples into the trained network model, and use the trained network to classify the samples to be identified; Step 7: Calculate the recognition rate Formula (4) is used to calculate the optimal recognition accuracy of the network for aluminum electrolysis data: Among them, n is the total number of test samples, and b is the correctly classified samples to be identified.

2. The aluminum electrolysis fault identification method based on deep reinforcement learning parameter automatic adjustment according to claim 1 is characterized in that: The specific implementation method of step 1 is: The collected aluminum electrolysis data is processed into input signals that can be recognized by the network, the number of categories and sample dimensions of the samples to be recognized are determined, and the aluminum electrolysis data is divided into training samples and test samples according to a certain ratio. The processed training sample set is {(x (1) ,y (1) ),...,(x (m) ,y (m) )}, where m is the number of training samples, x (i) is the i-th training sample, y (i) ∈{1,2,...,k} is the label of the i-th training sample, k is the number of classes of samples to be identified in the aluminum electrolysis data; the processed test sample set is {x (1) ,x (2) ,...,x (n) }, where n is the number of training samples, x (i) is the ith test sample.

3. The aluminum electrolysis fault identification method based on deep reinforcement learning parameter automatic adjustment according to claim 2 is characterized in that: The specific implementation of step 2 is: Step 2.1) Determine the network model as a one-dimensional convolutional neural network Step 2.2) Determine the classifier as softmax The classifier in this step uses the softmax classifier. When the training sample set is {(x (1) ,y (1) ),...,(x (m) ,y (m) )}, the softmax classifier is used to classify the features of the extracted data to be identified as follows: In which, it is assumed that the vector h θ (x (i) ) of each element p(y (i) =jx (i) ; θ) represents the feature x of the sample object to be identified (i) The probability of belonging to the jth class, j∈{1,2,...,k}, the larger the probability, the more likely the sample x to be identified is (i) The greater the probability of belonging to the jth class, where θ1,θ2,...,θ k is the parameter vector of the convolutional neural network.

4. The aluminum electrolysis fault identification method based on deep reinforcement learning parameter automatic adjustment according to claim 1 is characterized in that: The specific implementation of step 4 is as follows: 4.1) Network initialization, Use Status * Define a one-dimensional convolutional neural network; 4.2) Data feature extraction, Feature extraction is completed by the convolution operation in the network. For all convolutional layers, the filling method is set to all zero filling, the moving step size of the convolution kernel is set to 1, the activation function is set to the relu function, and a batch normalization layer is set before each convolutional layer; 4.3) Establish the objective function, The objective function is the loss function Loss, which is defined as the mean square error of the classifier: Among them, y i represents the true category label of the i-th sample, represents the predicted category label of the fault recognition network for the i-th sample; 4.4) Gradient descent method to optimize parameters, Back propagation is used to calculate the gradient variables and optimize the weight parameters in the convolutional neural network to minimize the objective function.

Citation Information

Patent Citations

  • Electroencephalogram signal emotion recognition method based on network structure search

    CN113516101A

  • Energy internet hybrid energy system and scheduling method thereof

    CN113991654A