Transformer state identification method and device, electronic equipment and storage medium

By training a voiceprint extraction model and utilizing registered voiceprints to identify transformer states, the problem of recognition accuracy caused by insufficient training data was solved, achieving high-accuracy recognition with a small amount of labeled data.

CN119763604BActive Publication Date: 2025-11-18IFLYTEK CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411610719.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-11-18
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of transformer condition recognition models decreases when training data is insufficient.

Method used

A voiceprint extraction model is trained based on a large amount of first-sample audio data without state labels. A registered voiceprint is determined using a small amount of second-sample audio data with state labels. The registered voiceprint is then combined with the transformer state recognition.

Benefits of technology

The accuracy of transformer state identification was improved with a limited amount of state-labeled sample audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119763604B_ABST
    Figure CN119763604B_ABST
Patent Text Reader

Abstract

The application provides a transformer state recognition method and device, electronic equipment and storage medium, and relates to the technical field of electric power. The method comprises the following steps: obtaining audio data to be tested of a transformer; inputting the audio data to be tested into a voiceprint extraction model to obtain a target voiceprint output by the voiceprint extraction model; the voiceprint extraction model is trained based on first sample audio data of first sample transformers with stateless labels in a first quantity; and the state of the transformer is recognized based on the target voiceprint and at least one registered voiceprint. The application is characterized in that the voiceprint extraction model is first trained based on a large amount of first sample audio data with stateless labels, and then at least one registered voiceprint is determined based on a small amount of second sample audio data with state labels by using the voiceprint extraction model, so that the state of the transformer is recognized in combination with the registered voiceprint. In the case that only a small amount of sample audio data with state labels is needed, the accuracy of the transformer state recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power technology, and in particular to a transformer condition identification method, device, electronic device, and storage medium. Background Technology

[0002] In the power and other energy sectors, with the increasing demand for the construction of high-voltage and ultra-high-voltage power grids, higher and higher requirements are being placed on the reliability of power systems, especially transformers (if a failure occurs, it will lead to economic losses and even cause large-scale social shutdowns). Therefore, early warning of transformer failures is becoming increasingly important.

[0003] In related technologies, such as Figure 1 As shown, a deep neural network model is typically trained in a supervised manner based on transformer audio data labeled with normal state and transformer audio data labeled with fault state to obtain a transformer state recognition model. The audio data of the transformer under test is input into the trained transformer state recognition model to obtain the posterior probability of each state category. The state category corresponding to the maximum posterior probability is determined as the state of the transformer under test.

[0004] However, in the aforementioned related technologies, when training the transformer state recognition model in a supervised manner, a large amount of labeled transformer audio data is required as training data. If the training data is insufficient, the accuracy of transformer state recognition will be reduced. Summary of the Invention

[0005] This invention provides a transformer state identification method, device, electronic device, and storage medium to address the shortcomings of existing technologies where insufficient training data reduces the accuracy of transformer state identification.

[0006] This invention provides a transformer status identification method, comprising the following steps.

[0007] Acquire the audio data to be tested from the transformer;

[0008] The audio data to be tested is input into the voiceprint extraction model to obtain the target voiceprint output by the voiceprint extraction model; the voiceprint extraction model is trained based on the first sample audio data of the first sample transformer without state labels.

[0009] The state of the transformer is identified based on the target voiceprint and at least one registered voiceprint; the at least one registered voiceprint is a voiceprint obtained by inputting a second number of second sample audio data, including state labels, of a second number of second sample transformers into the voiceprint extraction model, where the second number is less than the first number.

[0010] According to a transformer state identification method provided by the present invention, the step of identifying the state of the transformer based on the target voiceprint and at least one registered voiceprint includes:

[0011] For each of the registered voiceprints, determine the similarity between the target voiceprint and the registered voiceprint;

[0012] The state of the transformer is determined based on the target state label of the second sample audio data corresponding to the registered voiceprint with the highest similarity.

[0013] According to the transformer state identification method provided by the present invention, the voiceprint extraction model is trained based on the following method:

[0014] Obtain the first sample audio data without state tags from the first sample transformer;

[0015] Based on the first sample audio data, a training dataset is determined, wherein the training dataset includes sample audio data of at least two preset sampling rates corresponding to the first sample audio data;

[0016] For each preset sampling rate sample audio data in the training dataset, feature extraction is performed on the sample audio data at the preset sampling rate to obtain the sample audio features of the sample audio data at the preset sampling rate;

[0017] The sample audio features at each of the preset sampling rates are input into the initial encoder to obtain the predicted coding features corresponding to each of the preset sampling rates output by the initial encoder;

[0018] Based on the predictive coding features of the first sample audio data at different preset sampling rates, the first loss information is determined.

[0019] Based on the first loss information, the voiceprint extraction model is determined.

[0020] According to a transformer state identification method provided by the present invention, determining the voiceprint extraction model based on the first loss information includes:

[0021] The predicted coding features corresponding to each preset sampling rate are input into the initial decoder to obtain the predicted decoding features corresponding to each preset sampling rate output by the initial decoder.

[0022] Based on the predicted decoding features and sample audio features corresponding to the same preset sampling rate of the first sample audio data, the second loss information is determined.

[0023] Based on the second loss information and the first loss information, the network parameters of the initial encoder are adjusted to obtain the voiceprint extraction model.

[0024] According to a transformer state identification method provided by the present invention, the step of extracting features from sample audio data at a preset sampling rate to obtain sample audio features of the sample audio data at the preset sampling rate includes:

[0025] Feature extraction is performed on the sample audio data at the preset sampling rate based on at least two types of filters. For each type of filter, the sample audio features of the sample audio data at the preset sampling rate corresponding to the filter are obtained. The filters include at least two types of equal-height Mel filters, equal-height frequency filters, and equal-height inverse Mel filters.

[0026] The step of inputting the sample audio features of each preset sampling rate into the initial encoder to obtain the predicted coding features corresponding to each preset sampling rate output by the initial encoder includes:

[0027] The sample audio features corresponding to all the filters at each preset sampling rate are concatenated to obtain combined audio features;

[0028] The combined audio features are input into the initial encoder to obtain the predictive coding features corresponding to each preset sampling rate of each filter output by the initial encoder.

[0029] According to a transformer state identification method provided by the present invention, determining the training dataset based on the first sample audio data includes:

[0030] Acquire the third sample audio data of the devices other than the first sample transformer without state tags;

[0031] Based on the first sample audio data and the third sample audio data, the training dataset is determined, and the training dataset also includes sample audio data with at least two preset sampling rates corresponding to the third sample audio data.

[0032] According to a transformer state identification method provided by the present invention, the sampling rates of the first sample audio data and the third sample audio data are both greater than the maximum preset sampling rate among all preset sampling rates;

[0033] The step of determining the training dataset based on the first sample audio data and the third sample audio data includes:

[0034] The first sample audio data is downsampled to obtain sample audio data of at least two preset sampling rates corresponding to the first sample audio data.

[0035] The third sample audio data is downsampled to obtain sample audio data of at least two preset sampling rates corresponding to the third sample audio data;

[0036] The training dataset is determined based on the sample audio data of at least two preset sampling rates corresponding to the first sample audio data and the sample audio data of at least two preset sampling rates corresponding to the third sample audio data.

[0037] The present invention also provides a transformer status identification device, comprising:

[0038] The acquisition unit is used to acquire the audio data to be tested from the transformer.

[0039] An extraction unit is used to input the audio data to be tested into a voiceprint extraction model to obtain the target voiceprint output by the voiceprint extraction model; the voiceprint extraction model is trained based on a first number of stateless label-free first sample audio data of a first sample transformer.

[0040] The identification unit is used to identify the state of the transformer based on the target voiceprint and at least one registered voiceprint; the at least one registered voiceprint is a voiceprint obtained by inputting a second number of second sample audio data, including state labels, of a second number of second sample transformers into the voiceprint extraction model, wherein the second number is less than the first number.

[0041] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the transformer state identification method as described above.

[0042] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the transformer state identification method as described above.

[0043] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the transformer state identification method as described above.

[0044] The transformer state identification method, apparatus, electronic device, and storage medium provided by this invention input the acquired audio data of the transformer to be tested into a voiceprint extraction model to obtain a target voiceprint output by the voiceprint extraction model. Based on the target voiceprint and at least one registered voiceprint, the state of the transformer is identified. The voiceprint extraction model is trained based on a first number of unlabeled first sample audio data, and the registered voiceprint is obtained by inputting a second number of second sample audio data, including state labels, into the voiceprint extraction model. It can be seen that this invention first trains a voiceprint extraction model based on a large amount of unlabeled first sample audio data, then uses the voiceprint extraction model to determine at least one registered voiceprint based on a small amount of labeled second sample audio data, and finally combines the registered voiceprint to achieve transformer state identification. This improves the accuracy of transformer state identification with only a small amount of labeled sample audio data. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0046] Figure 1 This is a flowchart illustrating the transformer status identification process in related technologies.

[0047] Figure 2 This is a flowchart illustrating the transformer state identification method provided in an embodiment of the present invention.

[0048] Figure 3 This is a spectrum diagram of the equal-height Mel filter provided in an embodiment of the present invention.

[0049] Figure 4 This is one of the flowcharts for the training method of the voiceprint extraction model provided in the embodiments of the present invention.

[0050] Figure 5 This is the second flowchart of the training method for the voiceprint extraction model provided in this embodiment of the invention.

[0051] Figure 6 This is a schematic diagram of the network model provided in an embodiment of the present invention.

[0052] Figure 7 This is a schematic diagram of the transformer status identification device provided in an embodiment of the present invention.

[0053] Figure 8 This is a schematic diagram of the physical structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0055] Currently, intelligent manufacturing in industry has entered an era of rapid development. With the continuous advancement of science and technology and automation, industries such as manufacturing and energy are gradually moving towards automation and intelligence. Therefore, the normal operation of industrial equipment is becoming increasingly important. Sudden failures in industrial equipment can not only cause huge losses to a region's industrial manufacturing but also create significant secondary safety hazards, resulting in substantial economic losses and potentially even catastrophic casualties or severe social impacts. In energy scenarios such as electricity, the increasing demand for high-voltage and ultra-high-voltage power grid construction places increasingly higher demands on the reliability of power systems, especially transformers (failures can lead to economic losses and even widespread work stoppages). Therefore, early warning systems for transformer failures are becoming increasingly crucial.

[0056] Currently, transformer fault conditions can be identified through machine learning based on the sounds emitted during a fault, i.e., through state recognition technology. Common transformer fault conditions include DC bias, short-circuit impact, partial discharge, and heavy overload. The specific detection principle for transformer fault condition identification is as follows: The transformer's internal structure vibrates under electrical, magnetic, and mechanical stresses. The resulting mechanical waves are transmitted through a solid-liquid-gas medium and converted into audio signals. These audio signals are captured by acoustic sensors (microphones), yielding audio data containing a wealth of time-frequency domain feature information. When a transformer malfunctions, the sound changes accordingly. By combining this with artificial intelligence techniques such as deep learning, the transformer's state can be effectively identified.

[0057] However, in related technologies, training a transformer state recognition model in a supervised manner requires a large amount of labeled transformer audio data as training data. If the training data is insufficient, the accuracy of transformer state recognition will be reduced.

[0058] Based on this, the present invention proposes a transformer state recognition method. First, a voiceprint extraction model is trained based on a large amount of first sample audio data without state labels. Then, the voiceprint extraction model determines at least one registered voiceprint based on a small amount of second sample audio data with state labels. Finally, the registered voiceprint is combined to realize transformer state recognition. This method improves the accuracy of transformer state recognition when only a small amount of sample audio data with state labels is required.

[0059] The following is combined with Figures 2-6 The present invention describes a transformer status identification method. The executing entity of this transformer status identification method can be an electronic device such as a terminal, computer, or server, or a transformer status identification device installed in the electronic device. This transformer status identification device can be implemented through software, hardware, or a combination of both.

[0060] Figure 2 This is a flowchart illustrating the transformer state identification method provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the transformer condition identification method includes the following steps:

[0061] Step 201: Obtain the audio data to be tested from the transformer.

[0062] For example, when a transformer is operating normally or abnormally, the transformer is subjected to electrical, magnetic, and mechanical stresses, which cause vibrations. The resulting mechanical waves are transmitted through a solid-liquid-gas medium and converted into audio signals. These audio signals are captured by an acoustic sensor (microphone), thereby obtaining the transformer's audio data to be measured.

[0063] Step 202: Input the audio data to be tested into the voiceprint extraction model to obtain the target voiceprint output by the voiceprint extraction model; the voiceprint extraction model is trained based on the first sample audio data of the first sample transformer without state labels.

[0064] For example, when acquiring the audio data to be tested from the transformer, the audio data to be tested is downsampled to at least two preset sampling rates according to the actual sampling rate. At least two types of filters are used to extract features from the downsampled audio data to be tested from the transformer, and the audio features to be tested from the downsampled audio data to be tested corresponding to each filter are obtained. The audio features to be tested from the downsampled audio data to be tested corresponding to each filter are concatenated to obtain the concatenated features to be tested. The concatenated features to be tested are input into a pre-trained voiceprint extraction model. The voiceprint extraction model analyzes the input concatenated features to be tested, and finally obtains the target voiceprint corresponding to the concatenated features to be tested output by the voiceprint extraction model.

[0065] It should be noted that if the actual sampling rate of the audio data to be tested on the transformer is the same as the preset sampling rate, then downsampling is not required.

[0066] It should be noted that when at least two types of filters are included, the two types of filters include at least two of the following: equal-height Mel filters, equal-frequency filters, and equal-height inverse Mel filters. Preferably, at least two types of filters include three types of filters, namely equal-height Mel filters, equal-frequency filters, and equal-height inverse Mel filters. Among them, the equal-frequency filters are triangular filters arranged at equal intervals, and the equal-height inverse Mel filters exhibit the characteristic of dense filters at high frequencies and sparse filters at low frequencies. The calculation formulas for the Mel frequency of the equal-height Mel filters and the inverse Mel frequency of the equal-height inverse Mel filters can be expressed by the following formulas (1) and (2), respectively:

[0067] f_mel = 2595 log10(1 + f / 700) (1)

[0068] f_imel = 700 (10 ^ (f / 2595) - 1) (2)

[0069] Where f_mel represents the Mel frequency of the equal-height Mel filter, f represents the normal frequency corresponding to the equal-height frequency filter, and f_imel represents the inverse Mel frequency of the equal-height inverse Mel filter. Figure 3 This is a spectrum diagram of the equal-height Mel filter provided in an embodiment of the present invention, such as... Figure 3 As shown, features can be extracted from the audio data of the transformer based on the spectrogram of the equal-height Mel filter.

[0070] Step 203: Based on the target voiceprint and at least one registered voiceprint, identify the state of the transformer; the at least one registered voiceprint is a voiceprint obtained by inputting a second number of second sample audio data, including state labels, of the second sample transformer into the voiceprint extraction model, where the second number is less than the first number.

[0071] For example, after the voiceprint extraction model is trained, a second batch of second sample audio data, including status labels, is downsampled to at least two preset sampling rates based on the actual sampling rate, such as 32kbps and / or 16kbps. This enriches the content information of the sample audio data; for example, 48kbps second sample audio data is downsampled to 32kbps and / or 16kbps, and 20kbps second sample audio data is downsampled only to 16kbps. Then, all the downsampled second sample audio data is input into the voiceprint extraction model, and the voiceprint is extracted through inference by the voiceprint extraction model. The extracted voiceprint is called the registered voiceprint, and the registered voiceprint is stored in correspondence with the second sample audio data including status labels.

[0072] When obtaining the target acoustic signature corresponding to the splicing features of the transformer under test, the target acoustic signature is compared with each pre-registered acoustic signature. The target state label of the second sample audio data corresponding to the registered acoustic signature that is closest to the target acoustic signature is determined, and the state represented by the target state label is determined as the state of the transformer. For example, if the target state label represents a DC bias, the state of the transformer is determined to be DC bias; if the target state label represents a short circuit impact, the state of the transformer is determined to be short circuit impact; if the target state label represents a partial discharge, the state of the transformer is determined to be partial discharge; if the target state label represents a heavy overload, the state of the transformer is determined to be heavy overload. Of course, if the target state label represents a normal state, the state of the transformer is also normal, indicating that the transformer has not failed.

[0073] It should be noted that the second number of sample audio data, including status labels, is determined based on the audio signals collected when the transformer experiences all types of faults. When a new fault is subsequently discovered, new sample audio data needs to be determined based on the new audio signals collected when the transformer experiences a new fault. Then, the new sample audio features corresponding to the new sample audio data are input into the voiceprint extraction model to obtain a new registered voiceprint. The new registered voiceprint is then stored in correspondence with the new sample audio data including status labels, continuously improving the registered voiceprint and further enhancing the accuracy of transformer status identification.

[0074] The transformer state identification method provided by this invention inputs the acquired audio data of the transformer to be tested into a voiceprint extraction model to obtain the target voiceprint output by the voiceprint extraction model. Based on the target voiceprint and at least one registered voiceprint, the state of the transformer is identified. The voiceprint extraction model is trained based on a first number of unlabeled first sample audio data, and the registered voiceprint is obtained by inputting a second number of state-labeled second sample audio data into the voiceprint extraction model. It can be seen that this invention first trains the voiceprint extraction model based on a large amount of unlabeled first sample audio data, then uses the voiceprint extraction model to determine at least one registered voiceprint based on a small amount of state-labeled second sample audio data, and finally combines the registered voiceprint to achieve transformer state identification. This improves the accuracy of transformer state identification with only a small amount of state-labeled sample audio data.

[0075] In one embodiment, step 203 above identifies the state of the transformer based on the target voiceprint and at least one registered voiceprint, which can be implemented in the following way:

[0076] For each registered voiceprint, the similarity between the target voiceprint and the registered voiceprint is determined; based on the target state label of the second sample audio data corresponding to the registered voiceprint with the highest similarity, the state of the transformer is determined.

[0077] For example, for each registered voiceprint, the cosine distance between the target voiceprint and the registered voiceprint is calculated. The cosine distance is used to characterize the similarity between the target voiceprint and the registered voiceprint. The registered voiceprint corresponding to the maximum cosine distance is determined as the voiceprint that is closest to the target voiceprint. The target state label of the second sample audio data corresponding to the registered voiceprint with the maximum cosine distance is obtained. The state represented by the target state label is determined as the state of the transformer.

[0078] In this embodiment, the state of the transformer can be determined based on the similarity between the target voiceprint and each registered voiceprint, thus improving the accuracy of transformer state identification.

[0079] In one embodiment, Figure 4 This is one of the flowcharts of the training method for the voiceprint extraction model provided in this embodiment of the invention, such as... Figure 4 As shown, the voiceprint extraction model is trained in the following manner:

[0080] Step 401: Obtain the first sample audio data of the first sample transformer without state tags.

[0081] For example, a large amount of stateless audio data of the first sample transformer during operation and some open source collections of audio data are collected. All the collected stateless audio data are used as the first sample audio data, and the sampling rate of the first sample audio data is greater than the preset sampling rate.

[0082] Step 402: Based on the first sample audio data, determine the training dataset, wherein the training dataset includes sample audio data of at least two preset sampling rates corresponding to the first sample audio data.

[0083] For example, when all the first sample audio data are obtained, each first sample audio data can be downsampled to obtain sample audio data at least two preset sampling rates corresponding to the first sample audio data. Following the same method, sample audio data at least two preset sampling rates corresponding to each first sample audio data can be obtained. The set of sample audio data at least two preset sampling rates corresponding to all the first sample audio data is determined as the training dataset. Taking two preset sampling rates as an example, 32k and 16k respectively, the training dataset includes 32k sample audio data and 16k sample audio data corresponding to all the first sample audio data.

[0084] Step 403: For the sample audio data at each preset sampling rate in the training dataset, perform feature extraction on the sample audio data at the preset sampling rate to obtain the sample audio features of the sample audio data at the preset sampling rate.

[0085] For example, for sample audio data at each preset sampling rate in the training dataset, feature extraction is performed on the sample audio data at each preset sampling rate to obtain sample audio features of sample audio data at each preset sampling rate in the training dataset. Then, the sample audio features of 32k sample audio data and 16k sample audio data of the same first sample audio data are bound together to ensure that the 32k sample audio features and 16k sample audio features exist in each batch and correspond one-to-one.

[0086] It should be noted that when there are sample audio data in the training dataset whose length is less than the preset length, the sample audio data that is less than the preset length needs to be padded with zeros to reach the preset length.

[0087] Step 404: Input the sample audio features of each preset sampling rate into the initial encoder to obtain the predictive coding features corresponding to each preset sampling rate output by the initial encoder.

[0088] For example, the network model uses a masked autoencoder (MAE) to mask 70% of the features. The initial encoder can adopt a Transformer structure. The sample audio features at each preset sampling rate are input into the initial encoder. The initial encoder performs encoding analysis on the sample audio features at each preset sampling rate and finally outputs the predicted encoded features corresponding to each preset sampling rate.

[0089] Step 405: Determine the first loss information based on the predictive coding features of different preset sampling rates corresponding to the first sample audio data.

[0090] For example, after obtaining the predicted coding features corresponding to all preset sampling rates in the training dataset, for the predicted coding features of different preset sampling rates corresponding to each first sample audio data, the mean square error (MSE) is used to calculate the first loss information based on the predicted coding features of different preset sampling rates corresponding to the first sample audio data. The first loss information is the distance error of the predicted coding features of different preset sampling rates corresponding to the first sample audio data. The first loss information can also be called the contrast loss, which aims to bring the representation vectors of samples with different sampling rates closer together.

[0091] Step 406: Determine the voiceprint extraction model based on the first loss information.

[0092] For example, when calculating the first loss information, the distance error representing the first loss information is compared with a preset error. When the distance error representing the first loss information is greater than the preset error, the network parameters of the initial encoder are adjusted and iterated continuously until the distance error representing the first loss information is less than or equal to the preset error. Then, the final initial encoder is determined as the voiceprint extraction model.

[0093] In this embodiment, based on the predictive coding features of different preset sampling rates corresponding to the first sample audio data, the first loss information is calculated using MSE loss. Based on the first loss information, the representation vectors of samples with different sampling rates are narrowed down, and the final voiceprint extraction model has the ability to adapt to the input of audio data with different sampling rates, and can compare voiceprint similarity between audio data with different sampling rates.

[0094] In one embodiment, Figure 5 This is the second flowchart of the training method for the voiceprint extraction model provided in this embodiment of the invention, as follows: Figure 5 As shown, step 406 above determines the voiceprint extraction model based on the first loss information, which can be implemented through the following steps:

[0095] Step 4061: Input the predictive coding features corresponding to each preset sampling rate into the initial decoder to obtain the predictive decoding features corresponding to each preset sampling rate output by the initial decoder.

[0096] The initial decoder can be constructed using a Convolutional Neural Network (CNN) structure.

[0097] For example, when the predictive coding features corresponding to each preset sampling rate are obtained from the output of the initial encoder, the predictive coding features corresponding to each preset sampling rate are then input into the initial decoder. The initial decoder decodes the predictive coding features corresponding to each preset sampling rate and finally outputs the predictive decoding features corresponding to each preset sampling rate.

[0098] Step 4062: Determine the second loss information based on the predicted decoding features and sample audio features corresponding to the same preset sampling rate of the first sample audio data.

[0099] For example, when the predicted decoding features corresponding to each preset sampling rate of the initial decoder output are obtained, for each preset sampling rate, based on the predicted decoding features of the first sample audio data at that preset sampling rate and the sample audio features at that preset sampling rate, the second loss information is calculated using MSE. The second loss information can also be called reconstruction loss. For example, the second loss information is calculated based on the difference between the 32k predicted decoding features corresponding to the first sample audio data and the input 32k sample audio features, and the difference between the 16k predicted decoding features corresponding to the first sample audio data and the input 16k sample audio features.

[0100] Step 4063: Based on the second loss information and the first loss information, adjust the network parameters of the initial encoder to obtain the voiceprint extraction model.

[0101] For example, when the first loss information and the second loss information are obtained, the network parameters of the initial encoder are adjusted based on the first loss information and the second loss information, and the network parameters of the initial decoder are adjusted based on the second loss information. This process is iterated until the convergence condition is met. Finally, the initial encoder obtained by training is used as the voiceprint extractor, and the feature input is no longer masked. That is, the initial encoder obtained by training is determined as the voiceprint extraction model. Figure 6 This is a schematic diagram of the network model provided in an embodiment of the present invention, such as... Figure 6 As shown, the network model includes an initial encoder and an initial decoder. Sample audio features at each preset sampling rate are input into the initial encoder to obtain the predicted coding features corresponding to each preset sampling rate output by the initial encoder. The predicted coding features corresponding to each preset sampling rate output by the initial encoder are input into the initial decoder to obtain the predicted decoding features corresponding to each preset sampling rate output by the initial decoder. Based on the predicted coding features of the first sample audio data at different preset sampling rates, a first loss information is determined. Based on the predicted decoding features and sample audio features at the same preset sampling rate corresponding to the first sample audio data, a second loss information is determined. Based on the second loss information and the first loss information, the network parameters of the initial encoder are adjusted to obtain the voiceprint extraction model.

[0102] In this embodiment, based on the first loss information, the second loss information is determined based on the predicted decoding features and sample audio features corresponding to the same preset sampling rate of the first sample audio data. The self-supervised training of the voiceprint extraction model is realized based on the first loss information and the second loss information, thereby improving the robustness and accuracy of the voiceprint extraction model.

[0103] In one embodiment, the feature extraction of the sample audio data at the preset sampling rate in step 403 above, to obtain the sample audio features of the sample audio data at the preset sampling rate, can be implemented in the following way:

[0104] Feature extraction is performed on the sample audio data at the preset sampling rate based on at least two types of filters. For each type of filter, the sample audio features of the sample audio data at the preset sampling rate corresponding to the filter are obtained. The filters include at least two types of equal-height Mel filters, equal-height frequency filters, and equal-height inverse Mel filters.

[0105] Accordingly, step 404 above inputs the sample audio features of each preset sampling rate into the initial encoder to obtain the predictive coding features corresponding to each preset sampling rate output by the initial encoder. This can be achieved in the following ways:

[0106] The sample audio features corresponding to each preset sampling rate of all the filters are concatenated to obtain combined audio features; the combined audio features are input into the initial encoder to obtain the predictive coding features corresponding to each preset sampling rate of each filter output by the initial encoder.

[0107] For example, taking three types of filters—equal-height Mel filters, equal-height frequency filters, and equal-height inverse Mel filters—as an example, feature extraction is performed on sample audio data at a preset sampling rate based on each filter. This yields sample audio features of the sample audio data at the preset sampling rate corresponding to each filter. These sample audio features can be filter bank features. The sample audio features output by the three types of filters are then concatenated to form a 3D model. T The combined audio features of F, where 3 represents the channel dimension, T represents the time dimension, and F represents the frequency dimension, are input into the initial encoder to obtain the predictive coding features corresponding to each preset sampling rate of each filter output by the initial encoder.

[0108] In this embodiment, traditional filterbank features employ equal-height Mel filters, with dense filters at low frequencies and sparse filters at high frequencies, corresponding to the objective law that the human ear becomes less sensitive to higher frequencies. However, transformers, in addition to low-frequency signals, also exhibit high-frequency signals such as partial discharge. Using only filterbank features that fit the characteristics of the human ear cannot capture information about these high-frequency signals. Therefore, this invention employs at least two types of filters from equal-height Mel filters, equal-frequency filters, and equal-height inverse Mel filters to extract sample audio features, thereby obtaining a combination of features at multiple frequency scales. This results in more complete information about the combined audio features input to the initial encoder, reduces information loss, and further improves the accuracy of the final trained voiceprint extraction model.

[0109] In one embodiment, step 402 above determines the training dataset based on the first sample audio data, which can be implemented in the following way:

[0110] Acquire stateless audio data of other devices besides the first sample transformer; determine the training dataset based on the first sample audio data and the third sample audio data, wherein the training dataset also includes sample audio data of at least two preset sampling rates corresponding to the third sample audio data.

[0111] Optionally, the sampling rates of the first sample audio data and the third sample audio data are both greater than the maximum preset sampling rate among all preset sampling rates.

[0112] The first sample audio data is downsampled to obtain sample audio data at least two preset sampling rates corresponding to the first sample audio data; the third sample audio data is downsampled to obtain sample audio data at least two preset sampling rates corresponding to the third sample audio data; the training dataset is determined based on the sample audio data at least two preset sampling rates corresponding to the first sample audio data and the sample audio data at least two preset sampling rates corresponding to the third sample audio data.

[0113] Other equipment may be electrical equipment and / or industrial equipment, such as switch cabinets and reactors, and industrial equipment such as motors and air compressors.

[0114] For example, when all the first sample audio data are obtained, the first sample audio data can be downsampled for each first sample audio data to obtain sample audio data with at least two preset sampling rates corresponding to the first sample audio data. In the same way, sample audio data with at least two preset sampling rates corresponding to each first sample audio data can be obtained.

[0115] It can also acquire the third sample audio data of the stateless tags of other devices besides the first sample transformer, that is, acquire the third sample audio data of the stateless tags of electrical equipment and / or industrial equipment. The sampling rate of the third sample audio data is greater than the preset sampling rate. For each third sample audio data, the third sample audio data can be downsampled to obtain sample audio data of at least two preset sampling rates corresponding to the third sample audio data. Following the same method, at least two preset sampling rates corresponding to each third sample audio data can be obtained.

[0116] The training dataset is defined as the set of sample audio data corresponding to at least two preset sampling rates for all first sample audio data and the set of sample audio data corresponding to at least two preset sampling rates for all third sample audio data. For example, with two preset sampling rates of 32k and 16k, the training dataset includes 32k and 16k sample audio data corresponding to all first sample audio data, and 32k and 16k sample audio data corresponding to all third sample audio data.

[0117] In this embodiment, a training dataset is determined based on sample audio data at at least two preset sampling rates corresponding to the first sample audio data and sample audio data at at least two preset sampling rates corresponding to the third sample audio data. This training dataset includes sample audio data from devices other than transformers, enriching the data sources of the training dataset. This allows the voiceprint extraction model trained on the training dataset to adapt to unknown audio data to be tested, thereby improving the generalization of the voiceprint extraction model. In addition, the sample audio data is downsampled to at least two preset sampling rates according to the actual sampling rate, making the retrieval of the closest voiceprint more accurate.

[0118] The transformer status identification device provided by the present invention is described below. The transformer status identification device described below and the transformer status identification method described above can be referred to in correspondence.

[0119] Figure 7 This is a schematic diagram of the transformer status identification device provided in an embodiment of the present invention, as shown below. Figure 7 As shown, the transformer status identification device 700 includes an acquisition unit 701, an extraction unit 702, and an identification unit 703, wherein:

[0120] Acquisition unit 701 is used to acquire the audio data to be tested from the transformer;

[0121] Extraction unit 702 is used to input the audio data to be tested into the voiceprint extraction model to obtain the target voiceprint output by the voiceprint extraction model; the voiceprint extraction model is trained based on the first sample audio data of the first sample transformer without state labels.

[0122] The identification unit 703 is used to identify the state of the transformer based on the target voiceprint and at least one registered voiceprint; the at least one registered voiceprint is a voiceprint obtained by inputting a second number of second sample audio data including state labels of a second number of second sample transformers into the voiceprint extraction model, and the second number is less than the first number.

[0123] The transformer state identification device provided by this invention inputs the acquired audio data of the transformer to be tested into a voiceprint extraction model to obtain a target voiceprint output by the voiceprint extraction model. Based on the target voiceprint and at least one registered voiceprint, the state of the transformer is identified. The voiceprint extraction model is trained based on a first number of unlabeled first sample audio data, and the registered voiceprint is obtained by inputting a second number of labeled second sample audio data into the voiceprint extraction model. It can be seen that this invention first trains the voiceprint extraction model based on a large amount of unlabeled first sample audio data, then uses the voiceprint extraction model to determine at least one registered voiceprint based on a small amount of labeled second sample audio data, and finally combines the registered voiceprint to achieve transformer state identification. This improves the accuracy of transformer state identification with only a small amount of labeled sample audio data.

[0124] Based on any of the above embodiments, the identification unit 703 is specifically used for:

[0125] For each of the registered voiceprints, determine the similarity between the target voiceprint and the registered voiceprint;

[0126] The state of the transformer is determined based on the target state label of the second sample audio data corresponding to the registered voiceprint with the highest similarity.

[0127] Based on any of the above embodiments, the voiceprint extraction model is trained in the following manner:

[0128] Obtain the first sample audio data without state tags from the first sample transformer;

[0129] Based on the first sample audio data, a training dataset is determined, wherein the training dataset includes sample audio data of at least two preset sampling rates corresponding to the first sample audio data;

[0130] For each preset sampling rate sample audio data in the training dataset, feature extraction is performed on the sample audio data at the preset sampling rate to obtain the sample audio features of the sample audio data at the preset sampling rate;

[0131] The sample audio features at each of the preset sampling rates are input into the initial encoder to obtain the predicted coding features corresponding to each of the preset sampling rates output by the initial encoder;

[0132] Based on the predictive coding features of the first sample audio data at different preset sampling rates, the first loss information is determined.

[0133] Based on the first loss information, the voiceprint extraction model is determined.

[0134] Based on any of the above embodiments, determining the voiceprint extraction model based on the first loss information includes:

[0135] The predicted coding features corresponding to each preset sampling rate are input into the initial decoder to obtain the predicted decoding features corresponding to each preset sampling rate output by the initial decoder.

[0136] Based on the predicted decoding features and sample audio features corresponding to the same preset sampling rate of the first sample audio data, the second loss information is determined.

[0137] Based on the second loss information and the first loss information, the network parameters of the initial encoder are adjusted to obtain the voiceprint extraction model.

[0138] Based on any of the above embodiments, the step of extracting features from the sample audio data at the preset sampling rate to obtain sample audio features of the sample audio data at the preset sampling rate includes:

[0139] Feature extraction is performed on the sample audio data at the preset sampling rate based on at least two types of filters. For each type of filter, the sample audio features of the sample audio data at the preset sampling rate corresponding to the filter are obtained. The filters include at least two types of equal-height Mel filters, equal-height frequency filters, and equal-height inverse Mel filters.

[0140] The step of inputting the sample audio features of each preset sampling rate into the initial encoder to obtain the predicted coding features corresponding to each preset sampling rate output by the initial encoder includes:

[0141] The sample audio features corresponding to all the filters at each preset sampling rate are concatenated to obtain combined audio features;

[0142] The combined audio features are input into the initial encoder to obtain the predictive coding features corresponding to each preset sampling rate of each filter output by the initial encoder.

[0143] Based on any of the above embodiments, determining the training dataset based on the first sample audio data includes:

[0144] Acquire the third sample audio data of the devices other than the first sample transformer without state tags;

[0145] Based on the first sample audio data and the third sample audio data, the training dataset is determined, and the training dataset also includes sample audio data with at least two preset sampling rates corresponding to the third sample audio data.

[0146] Based on any of the above embodiments, the sampling rates of the first sample audio data and the third sample audio data are both greater than the maximum preset sampling rate among all preset sampling rates;

[0147] The step of determining the training dataset based on the first sample audio data and the third sample audio data includes:

[0148] The first sample audio data is downsampled to obtain sample audio data of at least two preset sampling rates corresponding to the first sample audio data.

[0149] The third sample audio data is downsampled to obtain sample audio data of at least two preset sampling rates corresponding to the third sample audio data;

[0150] The training dataset is determined based on the sample audio data of at least two preset sampling rates corresponding to the first sample audio data and the sample audio data of at least two preset sampling rates corresponding to the third sample audio data.

[0151] Figure 8 This is a schematic diagram of the physical structure of the electronic device provided in the embodiments of the present invention, such as... Figure 8 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a transformer state identification method, which includes: acquiring the transformer's audio data to be measured;

[0152] The audio data to be tested is input into the voiceprint extraction model to obtain the target voiceprint output by the voiceprint extraction model; the voiceprint extraction model is trained based on the first sample audio data of the first sample transformer without state labels.

[0153] The state of the transformer is identified based on the target voiceprint and at least one registered voiceprint; the at least one registered voiceprint is a voiceprint obtained by inputting a second number of second sample audio data, including state labels, of a second number of second sample transformers into the voiceprint extraction model, where the second number is less than the first number.

[0154] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0155] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, the computer program being executed by a processor, the computer being able to execute the transformer state identification method provided by the above methods, the method including: acquiring the transformer's audio data to be tested;

[0156] The audio data to be tested is input into the voiceprint extraction model to obtain the target voiceprint output by the voiceprint extraction model; the voiceprint extraction model is trained based on the first sample audio data of the first sample transformer without state labels.

[0157] The state of the transformer is identified based on the target voiceprint and at least one registered voiceprint; the at least one registered voiceprint is a voiceprint obtained by inputting a second number of second sample audio data, including state labels, of a second number of second sample transformers into the voiceprint extraction model, where the second number is less than the first number.

[0158] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the transformer state identification method provided by the above methods, the method comprising: acquiring the transformer's audio data to be tested;

[0159] The audio data to be tested is input into the voiceprint extraction model to obtain the target voiceprint output by the voiceprint extraction model; the voiceprint extraction model is trained based on the first sample audio data of the first sample transformer without state labels.

[0160] The state of the transformer is identified based on the target voiceprint and at least one registered voiceprint; the at least one registered voiceprint is a voiceprint obtained by inputting a second number of second sample audio data, including state labels, of a second number of second sample transformers into the voiceprint extraction model, where the second number is less than the first number.

[0161] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying the state of a transformer, characterized in that, include: Acquire the audio data to be tested from the transformer; The audio data to be tested is input into the voiceprint extraction model to obtain the target voiceprint output by the voiceprint extraction model; The voiceprint extraction model is trained based on the first sample audio data of the first sample transformer without state labels. The state of the transformer is identified based on the target voiceprint and at least one registered voiceprint; The at least one registered voiceprint is a voiceprint obtained by inputting a second number of second sample audio data, including status labels, of a second number of second sample transformers into the voiceprint extraction model, wherein the second number is less than the first number. The voiceprint extraction model was trained in the following manner: Obtain the first sample audio data without state tags from the first sample transformer; Based on the first sample audio data, a training dataset is determined, wherein the training dataset includes sample audio data of at least two preset sampling rates corresponding to the first sample audio data; For each preset sampling rate sample audio data in the training dataset, feature extraction is performed on the sample audio data at the preset sampling rate to obtain the sample audio features of the sample audio data at the preset sampling rate; The sample audio features at each of the preset sampling rates are input into the initial encoder to obtain the predicted coding features corresponding to each of the preset sampling rates output by the initial encoder; Based on the predictive coding features of the first sample audio data at different preset sampling rates, the first loss information is determined. Based on the first loss information, the voiceprint extraction model is determined.

2. The transformer status identification method according to claim 1, characterized in that, The process of identifying the state of the transformer based on the target voiceprint and at least one registered voiceprint includes: For each of the registered voiceprints, determine the similarity between the target voiceprint and the registered voiceprint; The state of the transformer is determined based on the target state label of the second sample audio data corresponding to the registered voiceprint with the highest similarity.

3. The transformer status identification method according to claim 1, characterized in that, The step of determining the voiceprint extraction model based on the first loss information includes: The predicted coding features corresponding to each preset sampling rate are input into the initial decoder to obtain the predicted decoding features corresponding to each preset sampling rate output by the initial decoder. Based on the predicted decoding features and sample audio features corresponding to the same preset sampling rate of the first sample audio data, the second loss information is determined. Based on the second loss information and the first loss information, the network parameters of the initial encoder are adjusted to obtain the voiceprint extraction model.

4. The transformer status identification method according to claim 1, characterized in that, The step of extracting features from the sample audio data at the preset sampling rate to obtain sample audio features of the sample audio data at the preset sampling rate includes: Feature extraction is performed on the sample audio data at the preset sampling rate based on at least two types of filters. For each type of filter, the sample audio features of the sample audio data at the preset sampling rate corresponding to the filter are obtained. The filters include at least two types of equal-height Mel filters, equal-height frequency filters, and equal-height inverse Mel filters. The step of inputting the sample audio features of each preset sampling rate into the initial encoder to obtain the predicted coding features corresponding to each preset sampling rate output by the initial encoder includes: The sample audio features corresponding to all the filters at each preset sampling rate are concatenated to obtain combined audio features; The combined audio features are input into the initial encoder to obtain the predictive coding features corresponding to each preset sampling rate of each filter output by the initial encoder.

5. The transformer state identification method according to any one of claims 1-4, characterized in that, The step of determining the training dataset based on the first sample audio data includes: Acquire the third sample audio data of the devices other than the first sample transformer without state tags; Based on the first sample audio data and the third sample audio data, the training dataset is determined, and the training dataset also includes sample audio data with at least two preset sampling rates corresponding to the third sample audio data.

6. The transformer state identification method according to claim 5, characterized in that, The sampling rates of the first sample audio data and the third sample audio data are both greater than the maximum preset sampling rate among all preset sampling rates; The step of determining the training dataset based on the first sample audio data and the third sample audio data includes: The first sample audio data is downsampled to obtain sample audio data of at least two preset sampling rates corresponding to the first sample audio data. The third sample audio data is downsampled to obtain sample audio data of at least two preset sampling rates corresponding to the third sample audio data; The training dataset is determined based on the sample audio data of at least two preset sampling rates corresponding to the first sample audio data and the sample audio data of at least two preset sampling rates corresponding to the third sample audio data.

7. A transformer status identification device, characterized in that, include: The acquisition unit is used to acquire the audio data to be tested from the transformer. The extraction unit is used to input the audio data to be tested into the voiceprint extraction model to obtain the target voiceprint output by the voiceprint extraction model; The voiceprint extraction model is trained based on the first sample audio data of the first sample transformer without state labels. An identification unit is used to identify the state of the transformer based on the target voiceprint and at least one registered voiceprint; The at least one registered voiceprint is a voiceprint obtained by inputting a second number of second sample audio data, including status labels, of a second number of second sample transformers into the voiceprint extraction model, wherein the second number is less than the first number. The voiceprint extraction model was trained in the following manner: Obtain the first sample audio data without state tags from the first sample transformer; Based on the first sample audio data, a training dataset is determined, wherein the training dataset includes sample audio data of at least two preset sampling rates corresponding to the first sample audio data; For each preset sampling rate sample audio data in the training dataset, feature extraction is performed on the sample audio data at the preset sampling rate to obtain the sample audio features of the sample audio data at the preset sampling rate; The sample audio features at each of the preset sampling rates are input into the initial encoder to obtain the predicted coding features corresponding to each of the preset sampling rates output by the initial encoder; Based on the predictive coding features of the first sample audio data at different preset sampling rates, the first loss information is determined. Based on the first loss information, the voiceprint extraction model is determined.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the transformer state identification method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the transformer state identification method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice recognition method and device, and computer storage medium

    CN110459205A

  • Processing method and device of audio fingerprint feature extraction model and computer equipment

    CN116758936A