Accelerated training neural network-based discharge pulse signal identification method and system
Through the improved multi-layer convolutional neural network and optimizer, combined with softmax and second-order derivative method, the problem of computing resource consumption and training time for discharge pulse signal recognition in the prior art is solved, and efficient and accurate automatic discharge pulse signal recognition is achieved.
Patent Information
- Application Number
- CN202510356148.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
When the prior art recognizes the discharge pulse signal of the passivation layer interface of the IGBT chip, there are problems such as large computer memory calculation, long training time, high cost and unstable optimizer updates, making it difficult to effectively identify the discharge pulse signal.
The multi-layer convolutional neural network based on the output layer activation function softmax and the second-order derivative method are used to improve the optimizer's multi-layer convolutional neural network, combined with the MultiFocalLossLayer loss function, the optimizer corrects the weight update formula through the control factor α to improve training speed and recognition accuracy.
It realizes automatic identification of discharge pulse signals in current signals, simplifies the recognition process, improves recognition accuracy, and reduces computing resource consumption and training time.
Smart Images

Figure CN120296557A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of discharge pulse signal recognition, and relates to a method and system for recognizing discharge pulse signals based on an accelerated training neural network. Background Art
[0002] Under the action of a positive repetitive square wave pulse voltage, surface discharges are likely to occur at the interface between the outermost polyimide (PI) material of the passivation layer of the IGBT chip and the encapsulating silicone gel. Due to the rapid polarization and depolarization caused by the rapid change of the electric field at the rising and falling edges of the square wave, the surface discharge current signal between the interfaces will generate a positive-polarity polarization current at the rising edge of the square wave and a negative-polarity depolarization current at the falling edge of the square wave. The polarization current and the depolarization current are collectively referred to as displacement current. The forward surface discharge pulses at the silicone gel-PI interface are concentrated at the rising edge of the square wave, and a few are distributed in the high-level region. The reverse surface discharge pulses are concentrated at the falling edge, and a few are distributed in the low-level region. This results in the overlap of the surface discharge pulses and the displacement current, bringing interference to the analysis of the discharge pulses. The voltage, current, and PMT (photomultiplier tube) output signals of the surface discharge are as Figure 1 shown.
[0003] Since the discharge pulses overlap with the polarization current, it is necessary to correctly identify the discharge pulse signals in the current signal. In previous studies, the general idea was to first locate the peak positions of the discharge pulses in the current signal sequence by means of extreme points, and then define the current signal data within a period of time before and after the peak of the discharge pulse as the discharge pulse signal. Specifically, there are two methods: one is as Figure 2 shown, which requires using the PMT output optical pulse signal to correspond one-to-one with the discharge pulse to identify the discharge pulse. This method requires a computer to process the PMT output optical pulse signal, increasing the computer memory operation amount and being limited by the threshold of the optical pulse signal peak. The other is as Figure 3 shown, which requires directly using the maximum value recognition method in the current signal to set reasonable threshold parameters to identify the discharge pulse. This method requires repeated experiments and comparisons to set reasonable maximum value thresholds, and it is necessary to distinguish noise in order to identify the correct position of the discharge pulse peak. This process requires a large amount of manpower and time. Using deep learning methods often requires large-scale data samples for training, resulting in a large amount of time and economic costs consumed during the training process, restricting its application in discharge pulse signal recognition. The current random gradient algorithm (SGD) used as an optimizer for training deep learning models has deficiencies such as unstable updates, slow convergence speed, the need to adjust the learning rate, and being prone to falling into local optima. And using ADAM as an optimizer requires more control factors to be tuned, and the convergence speed is slow, etc. Summary of the Invention
[0004] To address the deficiencies in the existing technology, the present invention provides a method and system for identifying discharge pulse signals based on an accelerated training neural network.
[0005] The present invention adopts the following technical solutions.
[0006] In the first aspect of the present invention, a method for identifying discharge pulse signals based on an accelerated training neural network is proposed. The optimizer is improved and optimized based on the output layer activation function softmax and the second derivative method. The improved optimizer is used to accelerate the training of a multi-layer convolutional neural network, and the trained multi-layer convolutional neural network is used to identify discharge pulse signals.
[0007] Among them, improving the optimizer based on the output layer activation function softmax and the second derivative method includes:
[0008] First, based on the first and second derivatives of the activation function softmax, a preliminary weight update formula is obtained:
[0009]
[0010] In the formula, δ k The weight value corresponding to the kth category, x k Represents the value input to the activation function softmax for the kth category;
[0011] Then, the control factor α is used to correct the preliminary weight update formula to obtain the update formula for the weights of the multi-layer convolutional neural network by the optimizer:
[0012]
[0013] In the formula, the value range of α is (0, 1].
[0014] Preferably, the method includes:
[0015] Step 1: Obtain current signals of multiple cycles and perform polarization current and discharge pulse classification markings on the current signals of each cycle to form a data set;
[0016] Step 2: Build a multi-layer convolutional neural network for identifying discharge pulse signals;
[0017] Step 3: Configure network training options, and use the data set and the improved optimizer to train the multi-layer convolutional neural network;
[0018] Step 4: Use the trained multi-layer convolutional neural network to identify discharge pulse signals.
[0019] Preferably, the multi-layer convolutional neural network includes an input layer, a convolutional feature extraction module, multiple fully connected layers, an activation layer, and an output layer;
[0020] Among them, the convolutional feature extraction module includes multiple convolutional modules. Each convolutional module performs feature extraction by continuously applying a convolutional layer, a normalization layer, and an activation function with the same parameters three times. The output of the convolutional feature extraction module is connected to multiple fully connected layers, and each fully connected layer is processed by a normalization layer, a fully connected layer, and an activation function.
[0021] Preferably, the multi-layer convolutional neural network adopts the MultiFocalLossLayer loss function, which is specifically as follows:
[0022] FL(p t ) = -α t (1 - p t ) γ log(p t )
[0023] Among them, FL(p t ) is the sample loss value; p t is the prediction probability of the network for the true label; α t is the loss weight; γ is the weight exponent.
[0024] Preferably, the configuration of the network training options in step 3 includes:
[0025] Configure the control factors of the improved optimizer, set the maximum number of training epochs to 300, the number of samples for each parameter update to 32, the initial learning rate to 0.0003, the learning rate decay factor to 0.8, the learning rate decay period to 1, and the L2 regularization strength to 0.1.
[0026] The second aspect of the present invention proposes a discharge pulse signal recognition system based on an accelerated training neural network. The system includes:
[0027] A data acquisition module, which is used to acquire current signals of multiple cycles and classify and label the polarization current and discharge pulses of the current signals of each cycle to form a data set;
[0028] A network construction module, which is used to build a multi-layer convolutional neural network for discharge pulse signal recognition;
[0029] A network training module, which is used to configure network training options and train the multi-layer convolutional neural network using the data set and the improved optimizer;
[0030] A signal recognition module, which is used to recognize discharge pulse signals using the trained multi-layer convolutional neural network.
[0031] Preferably, the multi-layer convolutional neural network built in the network construction module includes an input layer, a convolutional feature extraction module, multiple fully connected layers, an activation layer, and an output layer;
[0032] The convolutional feature extraction module includes multiple convolutional modules. Each convolutional module performs feature extraction by continuously applying a convolutional layer, a normalization layer, and an activation function with the same parameters three times. The output of the convolutional feature extraction module is connected to multiple fully connected layers, and each fully connected layer is processed by a normalization layer, a fully connected layer, and an activation function.
[0033] Preferably, the multi-layer convolutional neural network adopts the MultiFocalLossLayer loss function, which is specifically as follows:
[0034] FL(p t ) = -α t (1 - p t ) γ log(p t )
[0035] Where FL(p t ) is the sample loss value; p t is the prediction probability of the network for the true label; α t is the loss weight; γ is the weight exponent.
[0036] Preferably, the optimizer adopted in the network training module is improved based on the output layer activation function softmax and the second derivative method, including:
[0037] First, obtain the preliminary weight update formula according to the first-order and second-order derivatives of the activation function softmax:
[0038]
[0039] In the formula, δ k is the weight value corresponding to the k-th category, and x k represents the value input to the activation function softmax for the k-th category;
[0040] Then, use the control factor α to correct the preliminary weight update formula to obtain the update formula of the optimizer for the weights of the multi-layer convolutional neural network:
[0041]
[0042] In the formula, the value range of α is (0, 1].
[0043] The third aspect of the present invention proposes a terminal, including a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method.
[0044] The fourth aspect of the present invention proposes a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method are implemented.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] The multi-layer convolutional neural network based on training acceleration in the present invention can automatically identify the discharge pulse signal sequence by inputting only the current signal, which solves the defect that previous studies can only locate the extreme point of the discharge pulse first, and then need to identify all the discharge pulses before and after the extreme point, making it easier to identify the discharge pulse.
[0047] The present invention utilizes the multi-order differentiability of the softmax function, adopts the softmax function as the activation function of the output layer of the discharge pulse signal recognition and classification model, and combines the second-order derivative method to improve the optimizer of the deep learning classification model training process, thereby improving the convergence speed of the model training process. By retaining the advantages of the second-order derivative of sofmax as the optimizer of the deep learning classification model training process as much as possible and overcoming its shortcomings of large computational complexity and high memory consumption, the deep learning model is trained, thereby effectively accelerating the training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 The voltage, current and PMT output signal of the surface discharge;
[0049] Figure 2 This is a schematic diagram of using the PMT extreme point to locate the extreme point of the discharge pulse;
[0050] Figure 3 A schematic diagram of locating a discharge pulse using the extreme point of the discharge pulse itself;
[0051] Figure 4 A schematic diagram of the functional principle of the method of the present invention;
[0052] Figure 5 Classify single cycle current signals;
[0053] Figure 6 It is the CNN network structure diagram;
[0054] Figure 7 is the classification result of the network for the untrained current signal;
[0055] Figure 8 is the confusion matrix of the test set.
[0056] Figure 9 It is a comparison curve of the training loss function of the ADAM optimizer and the optimizer of the present invention. DETAILED DESCRIPTION
[0057] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. The embodiments described in this application are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0058] Embodiment 1 of the present invention provides a method for identifying discharge pulse signals based on an accelerated training neural network. A convolutional neural network that automatically identifies discharge pulses and polarization current signals in current signals is generated through deep learning. The optimizer is improved and optimized based on the output layer activation function softmax and the second derivative method. The improved optimizer is used to accelerate the training of the multi-layer convolutional neural network, and the trained multi-layer convolutional neural network is used to identify discharge pulse signals;
[0059] Among them, improving the optimizer based on the output layer activation function softmax and the second derivative method includes:
[0060] First, based on the first and second derivatives of the activation function softmax, a preliminary weight update formula is obtained:
[0061] softmax function where k represents the kth category and N is the total number of classification categories;
[0062] The first derivative of the softmax function is:
[0063]
[0064] The corresponding second derivative is:
[0065]
[0066] On the basis of using softmax as the output layer activation function and combining it with the second derivative method for improvement, replace the weight update formula in the mini-batch rule with where e k represents the error between the target value and the output value after softmax activation, that is, the weight update formula based on the second derivative method is as shown in the formula:
[0067]
[0068] In formula (1), since the value range of softmax(x k ) is (0, 1), the value range of 1 - 2 * softmax(x k ) is (-1, 1). At this time, the value method of δ k includes negative values and may have a division by zero phenomenon;
[0069] Then, based on the consideration of the iterative descent in the optimization algorithm (i.e., the value of δ k is less than 0) and that the denominator cannot be 0, a correction is made in the denominator of Equation (1) to ensure that its value range is greater than 0. One correction idea is The control factor here is the same as α in the Stochastic Gradient Descent (SGD) method, and its value range is a positive number in (0, 1]. To facilitate the performance comparison with SGD, one control factor of the relatively common ADAM method is streamlined here, and there is less need for manual intervention. The improved one in the present invention is named the MSOFTMAXOPT optimizer.
[0070] Based on the above algorithm principle of the improved second-order derivative method optimizer with softmax as the activation function of the output layer, as Figure 4 shown, the method specifically includes:
[0071] Step 1: Obtain current signals of multiple cycles and classify and label the polarization current and discharge pulses of the current signal of each cycle to form a data set;
[0072] Further preferably, the data preparation process is as follows: Manually classify the current signals of 100 cycles into polarization current and discharge pulses per cycle on matlab. The current signal of each cycle is a one-dimensional signal, and at the same time, label the discharge pulse signal and the polarization current signal respectively. The result is as Figure 5 shown.
[0073] Step 2: Build a multi-layer convolutional neural network for identifying discharge pulse signals;
[0074] Further preferably, build a Convolutional Neural Network (CNN) on matlab;
[0075] As Figure 6 shown, the network structure diagram shows a multi-level Convolutional Neural Network (CNN), and the main structure of each layer includes a one-dimensional convolutional layer (Conv1D), a normalization layer (IN Layer), and an activation function (LeakyReLU). The following is a detailed analysis of this network structure:
[0076] I. Input layer Input:
[0077] The input is: The input dimension of the network is [B, N, 1], where: B represents the batch size, that is, the number of samples input at one time; N represents the length of the input data (usually the time step or sequence length); 1 represents the feature dimension, indicating that each time step of the input data has only one feature.
[0078] II. Structure of the Convolutional Layer (Conv1D):
[0079] Each convolutional layer module extracts features through a combination of three convolutions, normalization, and activation functions.
[0080] Each convolutional module in each layer has different filter sizes and numbers. The specific parameter explanations are as follows:
[0081] Conv1D parameters [filter_size, filter_num, stride, padding]: filter_size: the size of the convolutional kernel (i.e., the receptive field size of each convolution operation); filter_num: the number of convolutional kernels, which determines the dimension of the output features; stride: the stride of the convolution, i.e., the step length by which the convolutional kernel slides each time; padding: the number of elements padded around the input to ensure the dimension of the output; IN Layer (Instance Normalization): the output after each layer of convolution is normalized to help accelerate training and improve the stability of the model; LeakyReLU activation function: the LeakyReLU activation function is applied after the convolutional layer, with parameter α = 0.2, meaning the slope of the negative part is 0.2.
[0082] The convolutional structures of each layer are as follows:
[0083] First layer: Input dimension [B, N, 1], convolutional kernel parameters [3, 4, 1, 1], output dimension [B, N, 4].[[]]
[0084] Second layer: Input dimension [B, N, 4], convolutional kernel parameters [7, 8, 1, 3], output dimension [B, N, 8].[[]]
[0085] Third layer: Input dimension [B, N, 8], convolutional kernel parameters [15, 16, 1, 7], output dimension [B, N, 16].[[]]
[0086] Fourth layer: Input dimension [B, N, 16], convolutional kernel parameters [31, 32, 1, 15], output dimension [B, N, 32].[[]]
[0087] Fifth layer: Input dimension [B, N, 32], convolutional kernel parameters [63, 64, 1, 31], output dimension [B, N, 64].[[]]
[0088] Sixth layer: Input dimension [B, N, 64], convolutional kernel parameters [127, 128, 1, 63], output dimension [B, N, 128].[[]]
[0089] The triple convolution operation refers to the continuous application of the same convolution, normalization, and activation functions three times. That is, this module will repeatedly apply the convolution operation with the same parameters three times, such as the convolution kernel size, the number of convolution kernels, the stride, and the padding, etc., remaining unchanged.
[0090] The further explanation is as follows:
[0091] 1. The first convolutional layer module: Take the input [B, N, 1] as an example.
[0092] The first convolution operation will use the parameters [3, 4, 1, 1], that is, the convolution kernel size is 3, the number of convolution kernels is 4, the stride is 1, and the padding is 1.
[0093] After this convolution operation, the output dimension becomes [B, N, 4].
[0094] Then, the output of [B, N, 4] will pass through the normalization layer (IN Layer) and the activation layer (LeakyReLU), and then be input into the second convolution operation with the same convolution parameters.
[0095] 2. Inside this module, the [B, N, 4] obtained after the first convolution is continuously fed into the second same convolution structure. The parameters of this second convolution are still [3, 4, 1, 1], and the output is still [B, N, 4].
[0096] 3. Finally, the same [B, N, 4] output enters the third convolution, the parameters are still [3, 4, 1, 1], and the final output dimension is still [B, N, 4].
[0097] There are normalization and activation functions after each convolution operation. The triple convolution only further extracts the depth and performs non - linear transformation on the feature map, without changing the number of output channels or dimensions. Therefore, after the input [B, N, 1] enters the first convolutional layer module and undergoes three convolutions with the same parameters, the output dimension is [B, N, 4] (that is, the number of convolution kernels determines the dimension of the output features).
[0098] III. Fully - connected layer (FC) and activation layer:
[0099] The output [B, N, 128] of the convolution feature extraction module is connected to the multi - layer fully - connected layer module.
[0100] The fully - connected layer parameter
[128] indicates that the dimension of the output hidden units is 128.
[0101] Each fully - connected layer is processed by an Instance Normalization layer and a LeakyReLU activation function (α = 0.2).
[0102] IV. Output layer:
[0103] Final output layer: The multi-layer fully connected layers are connected to the last output layer, and finally an output of size [B, N, 2] is obtained.
[0104] The FC parameter [2] indicates that the output feature dimension is 2, representing the final classification of the network.
[0105] Softmax layer: After the output layer, the Softmax activation function is used to normalize the output and convert it into a probability distribution (suitable for classification tasks).
[0106] The above describes the structure of the network and the functions of each layer. However, in actual problems, the amount of discharge pulse data is much smaller than that of polarization current data, which is a typical imbalanced data set. The discharge pulse data are minority class samples, and the polarization current data are majority class samples. To effectively handle the classification problem on the imbalanced data set, further optimize the performance of the model, and avoid the traditional CNN network loss function being too focused on the contribution of majority class samples to the loss, resulting in the inability to identify minority class samples. The present invention adopts a custom loss function, namely the MultiFocalLossLayer. This loss function can enhance the model's recognition ability for minority class data, that is, discharge pulses, thus improving the overall classification effect.
[0107] During the training process of the network, the MultiFocalLossLayer loss function is used to calculate the error between the model prediction result and the true label. The design of this loss function takes into account the characteristics of class imbalance, and dynamically adjusts the loss weight by introducing loss weights and weight exponents, making the model pay more attention to minority class samples. Specifically, the calculation formula of MultiFocalLossLayer is as follows:
[0108] FL(p t )=-α t (1-p t ) γ log(p t )
[0109] Where FL(p t ) is the sample loss value, which is an index to measure the gap between the model prediction output and the actual label; p t is the predicted probability of the model for the true class; α t is the loss weight, used to balance the importance of different classes; γ is the weight exponent, which controls the penalty for difficult samples.
[0110] Step 3: Configure the network training options, and train the multi-layer convolutional neural network using the said data set and the improved optimizer;
[0111] Further preferably, configure the network training options: use the MSOFTMAXOPT optimizer, and set parameters such as the maximum number of training epochs to 300, the number of samples for each parameter update to 32, the initial learning rate to 0.0003, the learning rate decay factor to 0.8, the learning rate decay period to 1, and the L2 regularization strength to 0.1, to ensure that the model can converge effectively.
[0112] Train the network: Input the current signals of 100 epochs in Step 1 into the convolutional neural network, divide them into 85 training set data and 15 test set data, and perform network training;
[0113] Step 4: Use the trained multi-layer convolutional neural network to identify the discharge pulse signals.
[0114] Through the training of the current signal data of 100 epochs, a convolutional neural network with the ability to automatically identify and classify discharge pulses and polarization currents is generated. Load this network in matlab, then load the current signals that have not been trained and tested, and use the trained network to classify the current signals, and the accurate identification of discharge pulses can be achieved. The effect is as Figure 7 shown. It can be seen that the automatic classification effect of the network is very accurate. The network has the following advantages:
[0115] (1) First, after the network is trained, only need to input the current signal into the convolutional neural network to identify the discharge pulse data; while the previous methods using PMT to assist in locating the discharge pulse and directly using the maximum value identification method in the current signal to identify the discharge pulse both need to first locate the peak of the discharge pulse, and then set parameters to find all the discharge pulse signals.
[0116] (2) The PMT-assisted location of the discharge pulse requires a relatively large peak value of the discharge pulse itself, that is, the amplitude of the optical signal generated by the discharge is large enough to correspond to the discharge pulse, which results in that some discharge pulses with too small amplitudes cannot be identified by this method. The present invention improves the accuracy rate.
[0117] (3) If the threshold value is not set reasonably when directly using the maximum value identification method to identify the discharge pulse in the current signal, it may cause some discharge pulses with too small amplitudes to be confused with the oscillations in the trailing part of the discharge pulses with larger amplitudes, resulting in these discharge pulses not being able to be identified. The present invention improves the accuracy rate.
[0118] (4) Use the trained convolutional neural network to classify on the test set, represent the classification results with a confusion matrix, and the larger the percentage of the main diagonal of the confusion matrix, the higher the accuracy rate. Figure 8 It can be seen that the recognition accuracy rate of this network is extremely high.
[0119] (5) Use a convolutional neural network with a contrastive ADAM optimizer and the MSOFTMAXOPT optimizer of the present invention to train on the training set. Figure 9 It can be seen that the convolutional neural network converges faster when using the MSOFTMAXOPT optimizer.
[0120] Embodiment 2 of the present invention provides a discharge pulse signal recognition system based on accelerating the training of a neural network, including:
[0121] A data acquisition module, configured to acquire current signals of multiple cycles and classify and label the polarization current and discharge pulses of the current signals of each cycle to form a data set.
[0122] A network construction module, configured to build a multi-layer convolutional neural network for discharge pulse signal recognition, including an input layer, a convolutional feature extraction module, a multi-layer fully connected layer, an activation layer, and an output layer.
[0123] The convolutional feature extraction module includes multiple convolutional modules, and each convolutional module performs feature extraction by continuously applying a convolutional layer, a normalization layer, and an activation function with the same parameters three times; the output of the convolutional feature extraction module is connected to the multi-layer fully connected layer, and each fully connected layer is processed by a normalization layer, a fully connected layer, and an activation function.
[0124] Further preferably, the multi-layer convolutional neural network adopts a MultiFocalLossLayer loss function, specifically as follows:
[0125] FL(p t )=-α t (1 - p t ) γ log(p t )
[0126] Among them, FL(p t ) is the sample loss value; p t is the prediction probability of the network for the true label; α t is the loss weight; γ is the weight exponent.
[0127] A network training module, configured to configure network training options and train the multi-layer convolutional neural network using the data set and the improved optimizer; the optimizer is improved based on the softmax activation function of the output layer and the second derivative method, including:
[0128] First, obtain a preliminary weight update formula according to the first and second derivatives of the softmax activation function:
[0129]
[0130] In the formula, δ kThe weight value corresponding to the k-th category, x k represents the value input to the activation function softmax for the k-th category;
[0131] Then, the control factor α is used to correct the preliminary weight update formula to obtain the weight update formula of the optimizer for the multi-layer convolutional neural network:
[0132]
[0133] In the formula, the value range of α is (0, 1].
[0134] The signal recognition module is used to recognize the discharge pulse signal by using the trained multi-layer convolutional neural network.
[0135] Embodiment 3 of the present invention provides a terminal, including a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method.
[0136] Embodiment 4 of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method are implemented.
[0137] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0138] Based on the multi-layer convolutional neural network with accelerated training, the present invention can automatically recognize the discharge pulse signal sequence only by inputting the current signal, solving the defect in the previous research that only the extreme points of the discharge pulse can be located first, and then all discharge pulses need to be recognized before and after the extreme points, making the recognition of discharge pulses more convenient.
[0139] The present invention utilizes the multi-order differentiability of the softmax function, uses the softmax function as the activation function of the output layer of the discharge pulse signal recognition and classification model, and at the same time combines the second derivative method to improve the optimizer in the training process of the deep learning classification model, improving the convergence speed of the model training process. By trying to retain the second-order derivative of sofmax as the advantage of the optimizer in the training process of the deep learning classification model and overcoming its deficiencies in large computational complexity and large memory consumption, the deep learning model is trained, which can effectively accelerate the training process.
[0140] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer-readable storage medium, on which computer-readable program instructions are loaded for causing a processor to implement various aspects of the present disclosure.
[0141] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as an instantaneous signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0142] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0143] Computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for identifying discharge pulse signals based on accelerating the training of a neural network, characterized in that Improve and optimize the optimizer based on the softmax activation function of the output layer and the second derivative method. Use the improved optimizer to accelerate the training of a multi-layer convolutional neural network, and use the trained multi-layer convolutional neural network to identify discharge pulse signals; Among them, improving and optimizing the optimizer based on the softmax activation function of the output layer and the second derivative method includes: Obtain a preliminary weight update formula according to the first and second derivatives of the softmax activation function: where δ k is the weight value corresponding to the k-th category, and x k represents the value input from the k-th category to the activation function softmax; Use a control factor α to correct the preliminary weight update formula to obtain the update formula of the optimizer for the weights of the multi-layer convolutional neural network: In the formula, the value range of α is (0, 1].
2. The method for identifying discharge pulse signals based on an accelerated training neural network according to claim 1, characterized in that It includes: Step 1: Obtain current signals of multiple cycles and classify and label the polarization current and discharge pulses of the current signals of each cycle to form a data set; Step 2: Build a multi-layer convolutional neural network for identifying discharge pulse signals; Step 3: Configure network training options, and use the data set and the improved optimizer to train the multi-layer convolutional neural network; Step 4: Use the trained multi-layer convolutional neural network to identify discharge pulse signals.
3. A method for identifying discharge pulse signals based on accelerating the training of a neural network according to claim 1 or 2, characterized in that: The multi-layer convolutional neural network includes an input layer, a convolutional feature extraction module, a multi-layer fully connected layer, an activation layer, and an output layer; among them, the convolutional feature extraction module includes multiple convolutional modules, and each layer of convolutional module performs feature extraction by continuously applying a convolutional layer, a normalization layer, and an activation function with the same parameters three times; the output of the convolutional feature extraction module is connected to the multi-layer fully connected layer, and each layer of fully connected layer is processed by a normalization layer, a fully connected layer, and an activation function.
4. A method for identifying discharge pulse signals based on accelerating the training of a neural network according to claim 1 or 2, characterized in that: The loss function adopted by the multi-layer convolutional neural network is as follows: FL(p t ) = -α t (1 - p t ) γ log(p t ) Among them, FL(p t ) is the sample loss value; p t is the predicted probability of the network for the true label; α t is the loss weight; γ is the weight exponent.
5. A method for identifying discharge pulse signals based on accelerating the training of a neural network according to claim 2, characterized in that: The configuration of the network training options in step 3 includes: Configure the control factor of the improved optimizer, set the maximum number of training cycles to 300, the number of samples for each parameter update to 32, the initial learning rate to 0.0003, the learning rate decay factor to 0.8, the learning rate decay period to 1, and the L2 regularization strength to 0.
1.
6. A discharge pulse signal recognition system based on an accelerated training neural network, characterized in that, The system includes: A data acquisition module for obtaining current signals of multiple cycles and classifying and labeling the polarization current and discharge pulses of the current signals of each cycle to form a data set; A network construction module for building a multi-layer convolutional neural network for identifying discharge pulse signals; A network training module for configuring network training options and using the data set and the improved optimizer to train the multi-layer convolutional neural network; A signal recognition module for using the trained multi-layer convolutional neural network to identify discharge pulse signals.
7. A system for identifying discharge pulse signals based on accelerating the training of a neural network according to claim 6, characterized in that: The multi-layer convolutional neural network built in the network construction module includes an input layer, a convolutional feature extraction module, multiple fully connected layers, an activation layer, and an output layer; Among them, the convolutional feature extraction module includes multiple convolutional modules. Each convolutional module performs feature extraction by continuously applying a convolutional layer, a normalization layer, and an activation function with the same parameters three times. The output of the convolutional feature extraction module is connected to the multiple fully connected layers, and each fully connected layer is processed by a normalization layer, a fully connected layer, and an activation function.
8. The discharge pulse signal recognition system based on an accelerated training neural network according to claim 7, characterized in that: The loss function adopted by the multi-layer convolutional neural network is as follows: FL(p t ) = -α t (1 - p t ) γ log(p t ) Among them, FL(p t ) is the sample loss value; p t is the predicted probability of the network for the true label; α t is the loss weight; γ is the weight exponent.
9. The discharge pulse signal recognition system based on an accelerated training neural network according to claim 6, characterized in that: The optimizer adopted in the network training module is improved based on the output layer activation function softmax and the second derivative method, and includes: Obtain a preliminary weight update formula according to the first and second derivatives of the activation function softmax: where δ k is the weight value corresponding to the k-th category, and x k represents the value input from the k-th category to the activation function softmax; Use a control factor α to correct the preliminary weight update formula to obtain the update formula of the optimizer for the weights of the multi-layer convolutional neural network: In the formula, the value range of α is (0, 1].
10. A terminal, including a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1-5.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1-5 are implemented.