A transformer fault soundprint detection method based on long short-term memory network optimization

Through the LSTM model and the Black Hawk algorithm optimized by the Sobol sequence, the feature extraction difficulties and gradient problems in transformer voiceprint signal detection are solved, achieving more efficient fault detection.

CN120564761BActive Publication Date: 2025-10-10JIAXING HENGCHUANG ELECTRIC POWER DESIGN & RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511061569.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-10
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

In transformer fault detection, traditional methods have difficulty in effectively extracting the temporal characteristics of transformer voiceprint signals, resulting in low detection accuracy, and the RNN model is prone to falling into local optimality and gradient vanishing or exploding problems.

Method used

A fault voiceprint detection method based on LSTM is adopted. By extracting the entropy features of sample data and using the Sobol sequence to generate hyperparameters, it is trained in combination with the Black Hawk optimization algorithm to avoid gradient vanishing and explosion, and improve the efficiency and accuracy of model training.

Benefits of technology

It effectively captures the timing characteristics of the transformer voiceprint signal, improves the accuracy of fault detection, avoids gradient vanishing and explosion phenomena, and improves the effect of transformer fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564761B_ABST
    Figure CN120564761B_ABST
Patent Text Reader

Abstract

The application discloses a transformer fault sound print detection method based on long short-term memory network optimization, and relates to the technical field of transformer fault detection.The application detects transformer fault sound print signals by training an LSTM model.In the LSTM model training process, an initial population of hyperparameters in the LSTM model is generated by a Zobor sequence pseudo-random number in each round of training, and the hyperparameters are optimized by a black eagle optimization algorithm; and the random function of each optimization stage in the black eagle optimization algorithm is replaced by the Zobor sequence, so that the training process of the LSTM model is prevented from falling into a local optimal solution, the training efficiency is improved, the gradient disappearance and gradient explosion phenomena that are prone to occur in the optimization method that depends on gradient information are avoided, and the fault detection model obtained after training can well extract the time sequence features in the sound print signals of the transformer to be detected, accurate fault detection results are obtained, and the accuracy of transformer fault detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of transformer fault detection, and in particular to a transformer fault soundprint detection method based on long short-term memory network optimization. BACKGROUND

[0002] At present, when detecting the fault of a transformer, the soundprint signal of the transformer can be detected. The soundprint signal of the transformer is a nonlinear time series. Traditional time domain feature extraction methods, such as mean, root mean square, kurtosis and skewness, and frequency domain feature extraction methods, such as power spectrum, envelope spectrum analysis and fast Fourier transform, are used. The traditional transformer fault soundprint detection method is essentially a multi-classification problem.

[0003] In the prior art, a multi-classification machine learning algorithm is generally needed to learn the internal correlation rules between the soundprint features and the fault labels from the training data, and then test the effect of learning the internal rules by the multi-classification machine learning algorithm through test data, so that the multi-classifier can have the ability to intelligently diagnose the transformer soundprint signal. From the existing technical path, a static data-driven supervised classic machine learning algorithm, such as decision tree, ensemble learning and support vector machine, is usually used for training and prediction.

[0004] However, because the soundprint signal of the transformer is a nonlinear time series information, the above method destroys the potential rules contained in the time series, so that only the correlation rules of the static data are learned. Even if RNN is used to learn the time series features of the soundprint signal, RNN is restricted by its own algorithm structure, and the learning effect of the soundprint signal of the transformer is poor, the memory is short, and RNN is prone to gradient disappearance and gradient explosion in the training process. The method of extracting the time series features of the soundprint signal by the traditional LSTM is easy to fall into local optimum in the training, and the optimization method depending on the gradient information is also easy to produce the phenomena of gradient disappearance and gradient explosion, which leads to low accuracy of transformer fault detection in actual application. SUMMARY

[0005] Therefore, it is necessary to provide a transformer fault soundprint detection method based on long short-term memory network optimization in view of the above technical problems.

[0006] The present application adopts the following technical solutions:

[0007] The present invention provides a transformer fault voiceprint detection method based on long short-term memory network optimization. The method comprises the following steps: first, obtaining historical voiceprint signals of the transformer as sample data, and marking the real fault type corresponding to each sample data; then, extracting entropy value features of the sample data, and inputting the entropy value features of the sample data into an LSTM model to obtain a predicted fault type; then, performing multiple rounds of training on the LSTM model with the optimization goal of minimizing the deviation between the real fault type and the predicted fault type to obtain a fault detection model; and during the training, generating an initial population of hyperparameters in the LSTM model by using a Sobol sequence to generate quasi-random numbers, and optimizing the hyperparameters by using a Black Hawk optimization algorithm; replacing the random function of each optimization stage in the Black Hawk optimization algorithm with a Sobol sequence; finally, extracting the entropy value features of the voiceprint signal of the transformer to be detected in an application and inputting the entropy value features into the fault detection model, and determining the fault type of the transformer to be detected by using the fault detection model.

[0008] The present invention provides a transformer fault detection device, comprising:

[0009] The acquisition module is used to obtain the transformer's historical voiceprint signals as sample data and mark the actual fault type corresponding to each sample data;

[0010] The prediction module is used to extract the entropy features of the sample data and input the entropy features of the sample data into the LSTM model to obtain the predicted fault type;

[0011] A training module is configured to perform multiple rounds of training on the LSTM model with the optimization objective of minimizing the deviation between the actual fault type and the predicted fault type to obtain a fault detection model. During the training, the module generates an initial population of hyperparameters in the LSTM model using a Sobol sequence to generate quasi-random numbers, and optimizes the hyperparameters using a Black Hawk optimization algorithm. The random function in each optimization stage of the Black Hawk optimization algorithm is replaced with a Sobol sequence.

[0012] The detection module is used to extract the entropy value characteristics of the voiceprint signal of the transformer to be detected and input it into the fault detection model, and determine the fault type of the transformer to be detected through the fault detection model.

[0013] The present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method for detecting transformer fault voiceprints based on long short-term memory network optimization is implemented.

[0014] The present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned transformer fault voiceprint detection method based on long short-term memory network optimization is implemented.

[0015] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects:

[0016] The present invention detects transformer fault voiceprint signals by training an LSTM model. During the LSTM model training process, a quasi-random number is generated through a Sobol sequence to generate an initial population of hyperparameters in the LSTM model, and the hyperparameters are optimized through a Black Hawk optimization algorithm. Moreover, the random function of each optimization stage in the Black Hawk optimization algorithm is replaced with a Sobol sequence to generate high-quality hyperparameters that approximate a uniform distribution, thereby better covering the parameter space and improving the algorithm diversity, thereby preventing the LSTM model training process from falling into a local optimal solution, improving the training efficiency while avoiding the gradient vanishing and gradient explosion phenomena that are easily generated by the optimization method that relies on gradient information, so that the fault detection model obtained after training can well learn the mapping relationship between the transformer voiceprint signal and the transformer fault type, thereby improving the accuracy of transformer fault detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0018] Figure 1 A flowchart of a transformer fault voiceprint detection method based on long short-term memory network optimization provided by the present invention;

[0019] Figure 2 A schematic diagram of an LSTM model provided by the present invention;

[0020] Figure 3 A schematic diagram of a transformer fault classification data processing flow provided by the present invention;

[0021] Figure 4 A schematic diagram of the structure of a transformer fault voiceprint detection algorithm based on a novel Black Hawk optimized long short-term memory network provided by the present invention;

[0022] Figure 5 This is a schematic diagram of a transformer fault detection device provided by the present invention. DETAILED DESCRIPTION

[0023] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0024] Currently, traditional time- and frequency-domain feature extraction methods for transformer fault detection often fail to effectively extract hidden fault signatures, resulting in reduced fault identification rates. Different system responses to fault shocks can lead to varying degrees of confusion in the monitoring data. Analysis methods based on entropy theory can directly measure data complexity without decomposing or transforming the data, enabling the identification of different faults. Methods such as sample entropy, permutation entropy, and spread entropy have been proposed and applied to fault diagnosis.

[0025] Existing technical approaches use static data-driven supervised classical machine learning algorithms, such as decision trees, ensemble learning, and support vector machines, for training and prediction, ignoring the semantic information of the sequence context. Even when using RNNs to learn the temporal characteristics of voiceprint signals, RNNs are constrained by their inherent algorithmic structure and have poor learning effects on transformer voiceprint signals. They also have short memory and are prone to vanishing and exploding gradients during training.

[0026] Specifically, RNN-based transformer fault detection usually has the following defects:

[0027] (1) Short-term memory. When RNN processes long sequences, information is transmitted through hidden states. As time goes by, information from earlier time steps may gradually disappear or be overwritten when it is transmitted to later time steps. This makes it difficult for RNN to capture and utilize long-term dependencies in the sequence, thus limiting its performance in processing complex tasks.

[0028] (2) Gradient vanishing and gradient exploding phenomena. During the backpropagation process of an RNN, the gradient will gradually vanish (become very small) or explode (become very large) as time steps pass. Gradient vanishing makes it difficult for the RNN to learn long-term dependencies during training because the gradient information of earlier time steps is almost zero when backpropagated to the initial layer. Gradient exploding may lead to unstable training process, excessive weight updates, and even numerical overflow.

[0029] The technical solutions provided by various embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0030] Figure 1 The figure is a flow chart of a transformer fault voiceprint detection method based on long short-term memory network optimization in the present invention, which specifically includes the following steps:

[0031] S101: Obtain historical voiceprint signals of the transformer as sample data, and mark the actual fault type corresponding to each sample data.

[0032] Generally, transformer mechanical defect faults are generally caused by overcurrent and fault current caused by short circuit or surge, which can cause stress between wires, leads and coils, causing winding deformation, core loosening, accessory loosening and other faults. Design defects, insulation aging, overload and other factors can also cause winding deformation. The above problems are generally detected by using voiceprint technology, and the fault types are generally divided into harmonic current, DC bias, fan aging, core loosening, heavy load, fan abnormal sound, partial discharge and the like.

[0033] Therefore, in one or more embodiments of the application, the server of the business platform can first obtain the historical voiceprint signal of the transformer as sample data, wherein the historical voiceprint signal of the normal transformer and the historical voiceprint signal of the fault transformer corresponding to multiple fault types can be included. Each sample data corresponding to the real fault type is also labeled to train and apply the selected model through the sample data.

[0034] The server mentioned in the application can be a server arranged in a business platform, or a device such as a desktop computer, a notebook computer and the like capable of executing the scheme of the application. For the convenience of description, only the server is taken as the execution subject for description below.

[0035] S102: Extract the entropy value feature of the sample data, and input the entropy value feature of the sample data into the LSTM model to obtain the predicted fault type.

[0036] After obtaining the sample data, the features of the sample data can be further extracted to facilitate the subsequent model fault feature extraction. At present, some signal analysis methods based on Shannon information entropy value theory are proposed, but there are usually the following shortcomings:

[0037] (1) Shannon information entropy converts the continuous body through integration, and defines the thermodynamic entropy as the expected value of the logarithm of the probability density. However, directly applying this definition is unreasonable in many practical situations, especially in the absence of any prior information about the underlying distribution.

[0038] (2) Entropy estimation techniques are also known as geometric estimators. These methods still have a fundamental dependence on the frequency-based view, and require a large amount of data to obtain useful results, which leads to long computation time.

[0039] (3) A common challenge for frequency-based estimators is handling outliers. It ignores the true impact of outliers by creating ambiguity between sparse sampling, low-probability regions of the measurement space, and truly disallowed measurements; on the other hand, geometric estimators limit the impact of outliers by using statistical averaging over the entire sample.

[0040] (4) There is still a single feature problem as with ordinary entropy, which cannot characterize time series from multiple fractional orders.

[0041] Based on this, in one or more embodiments of this specification, to address the shortcomings of analysis methods based on Shannon's information entropy theory, the present invention also proposes using Boltzmann–Shannon interaction entropy (BSIE) to extract entropy features from transformer voiceprint signals. The introduction of geometric partitioning entropy provides a better solution for quantifying the uncertainty of sample estimates extracted from a one-dimensional bounded continuous probability distribution.

[0042] Compared to traditional entropy estimation methods, this paper offers several improvements, including a better integration of effectiveness for sparse samples and the impact of extreme outliers, to define an entropy measure that is not biased by sample size. This paper also introduces a fractional-order calculation method into the Boltzmann-Shannon interaction entropy, further accounting for fractional-order information and capable of measuring the dynamic changes of fractional-order time series.

[0043] Specifically, the frequency and measurement-based entropy estimation method provides a new tool for analyzing the distribution of samples in a continuous state space: the Boltzmann-Shannon interaction entropy. In one or more embodiments of the present invention, the Boltzmann-Shannon interaction entropy of the sample data can be extracted as the entropy value feature of the sample data using the following formula:

[0044] ,

[0045] Where, is the sample data, is the JS divergence, is the equal length histogram segmentation distribution of the sample data, is the quantile split distribution of the sample data, is the entropy characteristic of the sample data. and Both are empirical distributions constructed based on sample data and are used to describe the distribution characteristics of data, but their construction methods and emphases are different.

[0046] in, The range of the sample data (minimum and maximum values) is divided into several equal-width intervals (bins), and then the frequency (or probability) of the data points in each interval is calculated. Each bin has the same width, but the frequency may be different. The sample data is used to determine the data range (minimum and maximum values) and the bin width. The frequency of each bin directly depends on the number of sample data points that fall into that bin. And is a discrete distribution that reflects the density approximation of sample data on an equal width interval. Therefore, Similar to the histogram estimator, it approximates the probability density function of the sample data population.

[0047] The sample data is divided into intervals based on their quantiles, so that each interval contains a roughly equal proportion of data points (i.e., equal frequency partitioning). Each interval has the same number of data points, but the interval widths may vary. The sample data is used to calculate the quantiles, which define the boundaries of the intervals. The probability mass of each interval is roughly equal. The interval widths are automatically adjusted based on the data distribution: smaller in areas with dense data and larger in areas with sparse data. is also a discrete distribution, but it is more directly related to the cumulative distribution function of the sample data. Based on the empirical distribution function, its quantile is an unbiased estimate of the population quantile of the sample data (when the sample size is large enough).

[0048] yes and The average distribution of , , the JS divergence is the average of these two KL divergences, that is .

[0049] The subsequent server can input the time domain and frequency domain features of the sample data into the LSTM model to obtain the predicted fault type corresponding to the sample data.

[0050] Figure 2 This is a schematic diagram of an LSTM model in this manual. Figure 2 middle, represents the forget gate, represents the input gate, represents the output gate, Indicates long-term memory, " represents the sigmoid activation function, represents the input features, Represents the hidden layer features. Specifically, the LSTM algorithm structure includes:

[0051] (1) Input gate: Determines which new information should be added to the memory cell. The input gate consists of a sigmoid activation function and a tanh activation function. The sigmoid function determines which information is important, while the tanh function generates new candidate information. The output of the input gate is multiplied by the candidate information, and the result will be considered when the memory cell is updated.

[0052] (2) Forget gate: Determines which old information should be forgotten or removed from the memory cell. The forget gate consists of only a sigmoid activation function. The output of the sigmoid function is directly multiplied by the current state of the memory cell to determine which information should be retained and which should be forgotten. Information with output values ​​closer to 1 will be retained, while information with output values ​​closer to 0 will be forgotten.

[0053] (3) Output gate: This gate determines which information in the memory cell should be output to the hidden state of the current time step. The output gate also consists of a sigmoid activation function and a tanh activation function. The sigmoid function determines which information should be output, while the tanh function processes the state of the memory cell to prepare for output. The output of the sigmoid function is multiplied by the memory cell state processed by the tanh function, and the result is the hidden state of the current time step.

[0054] LSTMs selectively retain important information while ignoring irrelevant details, and then perform subsequent processing based on this information. This mechanism enables them to efficiently process and output key information, solving the problems faced by RNNs when processing long sequences.

[0055] S103: The LSTM model is trained for multiple rounds with minimizing the deviation between the actual fault type and the predicted fault type as the optimization goal to obtain a fault detection model. During the training, the initial population of hyperparameters in the LSTM model is generated by generating quasi-random numbers using a Sobol sequence, and the hyperparameters are optimized using a Black Hawk optimization algorithm. The random function in each optimization stage of the Black Hawk optimization algorithm is replaced with a Sobol sequence.

[0056] S104: extracting the entropy value feature of the voiceprint signal of the transformer to be detected and inputting it into a fault detection model, and determining the fault type of the transformer to be detected through the fault detection model.

[0057] However, due to the nested layers of neurons, the LSTM loss function is a strongly nonlinear, non-convex loss function. For this non-convex optimization problem, existing optimization methods that rely on gradient information struggle to find a global optimal solution and are prone to falling into local minima. Furthermore, due to the LSTM's recurrent structure, existing optimization methods that rely on gradient information are prone to vanishing and exploding gradients during model training, resulting in prohibitively high training costs and difficulty.

[0058] Therefore, in one or more embodiments of the present invention, after obtaining the predicted fault type of the sample data as described above, the server can minimize the deviation between the actual fault type and the predicted fault type as the optimization goal, and optimize the LSTM model through the Black Hawk optimization algorithm. Figure 3A transformer fault classification data processing flow diagram in the application.

[0059] The loss function of the LSTM is a strong nonlinear non-convex loss function, and the essence of network learning is to solve a non-convex optimization problem, and due to the bottleneck of the existing optimization method relying on gradient information, the application will adopt a new Black Eagle optimization algorithm for global optimization of parameters, which is a meta-heuristic algorithm and also a derivative-free optimization method.

[0060] Black Eagle Optimization (BEO) algorithm. The BEO algorithm combines the biological laws of Black Eagle and mathematical transformations to guide the search behavior of particles. The traditional Black Eagle optimization increases the randomness of samples by using a large number of traditional Rand functions to generate pseudo-random numbers to approximate uniform distribution, but the pseudo-random numbers generated by the Rand function cannot approximate the uniform distribution with high quality, resulting in slow search speed and insufficient algorithm diversity. In view of the above problems, the application introduces Sobol sequence to generate pseudo-random numbers to form the initial population of LSTM hyperparameters and replace the Rand function in the traditional Black Eagle optimization algorithm, and the essence is to use Sobol sequence to enhance the uniform distribution of random parameters to improve the optimization efficiency.

[0061] Specifically, Sobol sequence:

[0062] The Sobol sequence is a sequence with the smallest prime number 2 as the base. To generate a random sequence , first, a irreducible polynomial with base 2 (the highest order is ) is needed to generate directional numbers.

[0063] The primitive polynomial is: .

[0064] The directional number is: .

[0065] Wherein, , there is the following relationship between the sequence .

[0066] .

[0067] In the formula, is any positive integer, satisfying is an odd number, , and are initially given, that is, are calculated initially. If the number of required directions is greater than :

[0068] .

[0069] The random sequence generated The sequence value is ,in, for The first binary digit in the binary representation of is in the order from low to high. The first Bit.

[0070] The mathematical model of the SSBEO algorithm is constructed based on the Black Hawk's tracking, hovering, capturing, grabbing, warning, migration, courtship and hatching behaviors.

[0071] SSBEO is a novel intelligent optimization algorithm based on the biological behavior of black hawks, designed to solve complex engineering optimization problems. The algorithm simulates the hunting, migration, and breeding behaviors of black hawks and combines mathematical transformations to guide the behavior of search particles. BEO possesses powerful global optimization and local search capabilities, making it suitable for solving large-scale, nonlinear, and nonconvex complex optimization problems.

[0072] Black Hawk Behavior: The Black Hawk's behavior is the core inspiration for the BEO algorithm, mainly including the following:

[0073] Predatory behaviors: include stalking, circling, capturing, snatching, and warning.

[0074] Migratory behavior: adapting to environmental changes.

[0075] Reproductive behavior: includes courtship and incubation.

[0076] The mathematical model of SSBEO is based on the Black Hawk's behavior described above and consists of the following components:

[0077] (1) Initialization:

[0078] To increase sample randomness during population initialization and subsequent strategies, traditional BEOs rely heavily on the traditional Rand function to generate quasi-random numbers that approximate a uniform distribution. However, these quasi-random numbers cannot approximate a uniform distribution with high quality. This results in a highly random population, but not necessarily uniformly distributed across the solution space. This leads to slow population search speeds and insufficient algorithmic diversity. To address these issues, the present invention introduces a Sobol sequence to generate quasi-random numbers for population initialization and replaces other Rand functions in traditional BEOs.

[0079] Assume that the number of black hawk individuals is n and the dimension of the hyperparameters in the LSTM model is d. Using the matrix X Represents the initial population, and each column in the population corresponds to the position of each black hawk:

[0080] ,

[0081] .

[0082] in, For the population j Individual i The value of the dimension, is the lower bound of the search space for hyperparameters in the LSTM model, is the upper bound of the search space for hyperparameters in the LSTM model, is the first j The Sobol sequence corresponding to the individuals generates the quasi-random number vector i A random number between 0 and 1.

[0083] (2) Tracking:

[0084] ,

[0085] ,

[0086] .

[0087] in, For the t A random position in the search space for the first iteration, is the step size parameter, is a quasi-random number between 0 and 1 generated by the Sobol sequence, is a quasi-random number between 0 and 1 generated by the Sobol sequence, For the t The position of a random black hawk in the iteration, The next position of a random black hawk standing on high ground, searching in the direction of potential prey, is a reverse learning strategy, that is, searching for the next position in the opposite direction, For the t +1 iteration position, To select the position with the minimum fitness, use it to update the t +1 iteration position, is the current optimal position.

[0088] (3) Circling:

[0089] ,

[0090] .

[0091] in, 、 yes The maximum distance to the search frontier, is a rotation matrix used to simulate rotation searches in high-dimensional spaces.

[0092] (4) Capture:

[0093] ,

[0094] .

[0095] in, is the distance between all individuals and the current optimal position. The position after the scale is enlarged, and is the position adjustment factor, A random vector with elements between 0.5 and 1 generated for the Sobol sequence.

[0096] (5) Snatch:

[0097] .

[0098] in, is a random vector that follows a normal distribution.

[0099] (6) Warning:

[0100] , ,

[0101] .

[0102] in, To guide the position after the movement by Poisson distribution, In accordance with the close order, rearrange The position matrix after is the step size parameter, is based on the parameters of the Poisson distribution, is the distribution function of the Poisson distribution.

[0103] (7) Migration:

[0104] , .

[0105] in, is the fitness function, is the worst fitness value, For the j The fitness value of an individual. is a random vector generated by the Sobol sequence, is a pseudorandom number between 0.4 and 1 generated by the Sobol sequence.

[0106] (8) Courtship:

[0107] .

[0108] in, , is the maximum number of iterations, For the t Iteration No. j individual positions, is a random vector generated by the Sobol sequence.

[0109] (9) Incubation:

[0110] .

[0111] in, is an array of random numbers that follows a normal distribution.

[0112] Figure 4 This is a schematic diagram of the structure of a transformer fault voiceprint detection algorithm based on a novel Black Hawk-optimized long-short-term memory network. The above process optimizes the hyperparameters of the long-short-term memory network, resulting in a fault detection model. Finally, the entropy characteristics of the transformer's voiceprint signal are extracted. The fault detection model then extracts fault features from these entropy characteristics to determine the transformer's fault type.

[0113] based on Figure 1 The transformer fault voiceprint detection method based on long short-term memory network optimization shown in the figure detects transformer fault voiceprint signals by training an LSTM model. During the LSTM model training process, each round of training generates an initial population of hyperparameters in the LSTM model through a quasi-random number generated by a Sobol sequence, and the hyperparameters are optimized through a Black Hawk optimization algorithm; and the random functions in each optimization stage of the Black Hawk optimization algorithm are replaced by a Sobol sequence, thereby avoiding the LSTM model training process from falling into a local optimal solution, improving training efficiency while avoiding the gradient vanishing and gradient explosion phenomena that are prone to occur in optimization methods that rely on gradient information. Therefore, the fault detection model obtained after training can well extract the time series features in the voiceprint signal of the transformer to be detected, obtain accurate fault detection results, and improve the accuracy of transformer fault detection.

[0114] The present invention uses the LSTM model to extract features from the nonlinear time series information of transformer voiceprint signals, which can effectively capture the time series characteristics and improve the feature learning ability of transformer voiceprint signals. At the same time, the Sobol sequence is used to generate quasi-random numbers to replace the Rand function in the Black Hawk optimization algorithm and apply it to the LSTM model training, which improves the training efficiency while avoiding the gradient vanishing and gradient explosion phenomena that are prone to occur in optimization methods that rely on gradient information.

[0115] When applying the transformer fault voiceprint detection method based on long short-term memory network optimization provided by the present invention, it is not necessary to Figure 1 The steps are executed in the order shown. The specific execution order of the steps can be determined according to needs, and the present invention does not limit this.

[0116] The above is a transformer fault voiceprint detection method based on long short-term memory network optimization provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding transformer fault detection device, such as Figure 5 shown.

[0117] Figure 5 A schematic diagram of a transformer fault detection device provided by the present invention includes:

[0118] The acquisition module 201 is used to obtain the historical voiceprint signal of the transformer as sample data and mark the actual fault type corresponding to each sample data;

[0119] Prediction module 202, used to extract entropy features of sample data and input the entropy features of the sample data into the LSTM model to obtain a predicted fault type;

[0120] The training module 203 is configured to perform multiple rounds of training on the LSTM model with minimizing the deviation between the actual fault type and the predicted fault type as the optimization goal to obtain a fault detection model. During the training, the LSTM model is initialized by generating a quasi-random number using a Sobol sequence, and the hyperparameters are optimized using a Black Hawk optimization algorithm. The random function in each optimization stage of the Black Hawk optimization algorithm is replaced with a Sobol sequence.

[0121] The detection module 204 is used to extract the entropy value characteristics of the voiceprint signal of the transformer to be detected and input the entropy value characteristics into the fault detection model, and determine the fault type of the transformer to be detected through the fault detection model.

[0122] The specific definition of the transformer fault detection device can be referred to the definition of the transformer fault sound print detection method based on the long short-term memory network optimization in the above, and will not be described here. Each module in the above transformer fault detection device can be realized by software, hardware and combination thereof in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so that the processor calls and executes the operation corresponding to each module.

[0123] The application further provides a computer readable storage medium, which stores a computer program, and the computer program can be used to execute the above Figure 1 The application provides a transformer fault sound print detection method based on a long short-term memory network optimization.

[0124] The application further provides a computer device, which comprises a processor, an internal bus, a network interface, a memory and a nonvolatile memory at the hardware level, and can further comprise other hardware required by business. The processor reads the corresponding computer program from the nonvolatile memory into the memory and then runs to implement the above Figure 1 The application provides a transformer fault sound print detection method based on a long short-term memory network optimization.

[0125] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a nonvolatile computer readable storage medium, and when the computer program is executed, the computer program can include the processes of the above-mentioned embodiment methods. In the embodiments of the application, any reference to the memory, storage, database or other medium can include at least one of the nonvolatile and volatile memories. The nonvolatile memory can include a read-only memory (ROM), a tape, a floppy disk, a flash memory or an optical memory. The volatile memory can include a random access memory (RAM) or an external cache memory. As an illustration but not limitation, the RAM can be in various forms, such as a static random access memory (SRAM) or a dynamic random access memory (DRAM).

[0126] Each technical feature of the above embodiments can be combined arbitrarily. In order to make the description simple, all possible combinations of each technical feature in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered that the combination is within the scope of the present application.

Claims

1. A transformer fault voiceprint detection method based on long short-term memory network optimization, characterized in that: include: Obtain the transformer's historical voiceprint signals as sample data and mark the actual fault type corresponding to each sample data; Extract the entropy features of the sample data and input the entropy features of the sample data into the LSTM model to obtain the predicted fault type; The LSTM model is trained multiple times with the optimization goal of minimizing the deviation between the actual fault type and the predicted fault type to obtain a fault detection model. During training, a quasi-random number population of hyperparameters in the LSTM model is generated using a Sobol sequence, and the hyperparameters are optimized using a Black Hawk optimization algorithm. The random function in each optimization stage of the Black Hawk optimization algorithm is replaced with a Sobol sequence. The entropy value characteristics of the voiceprint signal of the transformer to be detected are extracted and input into the fault detection model, and the fault type of the transformer to be detected is determined by the fault detection model.

2. The transformer fault voiceprint detection method based on long short-term memory network optimization according to claim 1 is characterized in that: The step of extracting entropy features of the sample data specifically includes: The Boltzmann-Shannon interaction entropy of the sample data is extracted as the entropy value feature of the sample data through the following formula: ; in, is the sample data, is the JS divergence, is the equal length histogram segmentation distribution of the sample data, is the quantile split distribution of the sample data, is the entropy characteristic of the sample data.

3. The transformer fault voiceprint detection method based on long short-term memory network optimization according to claim 1 is characterized in that: The method of generating the initial population of hyperparameters in the LSTM model by using the quasi-random number generator of the Sobol sequence specifically includes: The Sobol sequence is used to generate quasi-random numbers to generate the initial population of hyperparameters in the LSTM model using the following formula: , ; in, is the initial population, For the population j Individual i The value of the dimension, is the lower bound of the search space for hyperparameters in the LSTM model, is the upper bound of the search space for hyperparameters in the LSTM model, is the first j The Sobol sequence corresponding to the individuals generates the quasi-random number vector i dimensional random numbers, is the dimension of the hyperparameters in the LSTM model, is the number of individuals in the initial population.

4. The transformer fault voiceprint detection method based on long short-term memory network optimization according to claim 1 is characterized in that: The actual fault types include: one or more of: winding deformation, loose core, loose accessories, harmonic current, DC bias, fan aging, overload, abnormal fan noise and partial discharge.

5. A transformer fault detection device, characterized in that: include: The acquisition module is used to obtain the transformer's historical voiceprint signals as sample data and mark the actual fault type corresponding to each sample data; The prediction module is used to extract the entropy features of the sample data and input the entropy features of the sample data into the LSTM model to obtain the predicted fault type; A training module is configured to perform multiple rounds of training on the LSTM model with the optimization objective of minimizing the deviation between the actual fault type and the predicted fault type to obtain a fault detection model. During the training, the module generates an initial population of hyperparameters in the LSTM model using a Sobol sequence to generate quasi-random numbers, and optimizes the hyperparameters using a Black Hawk optimization algorithm. The random functions in each optimization stage of the Black Hawk optimization algorithm are replaced with Sobol sequences. The detection module is used to extract the entropy value characteristics of the voiceprint signal of the transformer to be detected and input it into the fault detection model, and determine the fault type of the transformer to be detected through the fault detection model.

6. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

7. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 4 when executing the computer program.

Citation Information

Patent Citations

  • Medium and long term hydrological probability forecasting method and system considering interpretable deep learning model nested combination

    CN119204350A

  • Transformer fault detection method and device, medium and equipment

    CN120220700A