Model generation method, anomaly detection method, device and electronic device

By training a generator network with audio features and adjusting it using a discriminator, the method addresses inefficiencies in manual inspection and improves anomaly detection in power system equipment, enhancing detection accuracy and efficiency.

CN114400019BActive Publication Date: 2025-07-15VOICEAI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111666960.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-07-15
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The abnormal detection method of traditional power operation equipment requires manual inspection, which is inefficient, and the deep neural network performs poorly when there is insufficient abnormal data.

Method used

By obtaining a training data set including normal audio information and abnormal audio information, the training generator network is trained to form an initial abnormality detection model, and adjust it using the discriminator network to form a target abnormality detection model.

Benefits of technology

Automatic abnormal detection is realized, detection efficiency is improved, and detection accuracy is maintained in the absence of abnormal audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114400019B_ABST
    Figure CN114400019B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a model generation method, an anomaly detection method, an apparatus, and an electronic device. The method includes: obtaining a training data set, where the training data set includes first audio features of multiple audio information of a target device, and the multiple audio information includes normal audio information and abnormal audio information; training a generator network to be trained through the training data set to use the converged generator network to be trained as an initial anomaly detection model; adjusting the initial anomaly detection model through a discriminator network to use the adjusted generator network as a target anomaly detection model. By the above method, it is possible to input the first audio feature of the device to be detected into the target anomaly detection model to perform anomaly detection on the device to be detected, saving manpower and improving the anomaly detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more particularly, to a model generation method, an anomaly detection method, an apparatus, and an electronic device. Background Art

[0002] With the continuous growth of power demand, the power system has increasingly higher requirements for the economy and reliability of power operation equipment. Due to the action of long-term load and the influence of natural environment (temperature, air pressure, humidity, pollution, etc.), the aging and wear of power operation equipment will occur, resulting in a gradual reduction in the performance and reliability of power operation equipment and potential safety hazards. Therefore, it is very necessary to monitor and detect the operating state of power operation equipment.

[0003] However, in the related power operation equipment detection methods, manual inspection is required, resulting in low efficiency. Summary of the Invention

[0004] In view of the above problems, the present application provides a model generation method, an anomaly detection method, an apparatus, an electronic device, and a storage medium to improve the above problems.

[0005] In a first aspect, the present application provides a model generation method applied to an electronic device. The method includes: obtaining a training data set, where the training data set includes first audio features of a plurality of audio information of a target device, and the plurality of audio information includes normal audio information and abnormal audio information; training a generator network to be trained through the training data set to use the converged generator network to be trained as an initial anomaly detection model; and adjusting the initial anomaly detection model through a discriminator network to use the adjusted generator network as a target anomaly detection model.

[0006] In a second aspect, the present application provides an anomaly detection method applied to an electronic device. The method includes: obtaining an audio to be detected; performing frame division, windowing, and fast Fourier transform on the audio to be detected to obtain a first audio feature corresponding to the audio to be detected, where the first audio feature is a spectrogram corresponding to the audio to be detected; and inputting the first audio feature into the target anomaly detection model obtained by the above method to obtain a detection result output by the target anomaly detection model.

[0007] In a third aspect, the present application provides a model generation device that runs on an electronic device. The device includes: a data set acquisition unit configured to acquire a training data set, where the training data set includes first audio features of multiple audio messages of a target device, and the multiple audio messages include normal audio messages and abnormal audio messages; an initial anomaly detection model acquisition unit configured to train a generator network to be trained through the training data set, and use the converged generator network to be trained as an initial anomaly detection model; and a target anomaly detection model acquisition unit configured to adjust the initial anomaly detection model through a discriminator network, and use the adjusted generator network as a target anomaly detection model.

[0008] In a fourth aspect, the present application provides an anomaly detection device that runs on an electronic device. The device includes: a to-be-detected audio acquisition unit configured to acquire an audio to be detected; a first audio feature acquisition unit configured to frame, window, and perform fast Fourier transform on the audio to be detected to obtain a first audio feature corresponding to the audio to be detected, where the first audio feature is a spectrogram corresponding to the audio to be detected; and a detection result acquisition unit configured to input the first audio feature into the target anomaly detection model obtained by the above method to obtain a detection result output by the target anomaly detection model.

[0009] In a fifth aspect, the present application provides an electronic device, including one or more processors and a memory; one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the above method.

[0010] In a sixth aspect, the present application provides a computer-readable storage medium, in which program code is stored, and when the program code runs, the above method is executed.

[0011] A model generation method, an anomaly detection method, a device, an electronic device, and a storage medium provided by the present application, after obtaining a training data set of first audio features including normal audio information and abnormal audio information, train a generator network to be trained through the training data set to use the converged generator network to be trained as an initial anomaly detection model, and then adjust the initial anomaly detection model through a discriminator network to use the adjusted generator network as a target anomaly detection model. By the above method, after training the target anomaly detection model through the first audio features of normal audio information and abnormal audio information, during the process of anomaly detection of the device to be detected, the first audio features of the device to be detected can be input into the target anomaly detection model to perform anomaly detection on the device to be detected, saving manpower and improving the efficiency of anomaly detection. Moreover, by adjusting the initial anomaly detection model through the discriminator network, the target anomaly detection model can have better performance even in the case of lack of abnormal audio training data, that is, it can have a higher accuracy rate in discriminating normal audio and abnormal audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0013] Figure 1 The flowchart of a model generation method proposed in an embodiment of the present application is shown;

[0014] Figure 2 Shows the present application Figure 1 The flowchart of an embodiment of S110 in the present application;

[0015] Figure 3 The schematic diagram of the process of a method for obtaining an audio data set proposed in the present application is shown;

[0016] Figure 4 The schematic diagram of the process of a method for obtaining first audio features proposed in the present application is shown;

[0017] Figure 5 The schematic diagram of a generator network to be trained proposed in the present application is shown;

[0018] Figure 6 The flowchart of a model generation method proposed in another embodiment of the present application is shown;

[0019] Figure 7 The flowchart of a model generation method proposed in yet another embodiment of the present application is shown;

[0020] Figure 8 Shows a schematic diagram of an anomaly detection model to be trained proposed by the present application;

[0021] Figure 9 Shows a flowchart of an anomaly detection method proposed by an embodiment of the present application;

[0022] Figure 10 Shows a structural block diagram of a model generation device proposed by an embodiment of the present application;

[0023] Figure 11 Shows a structural block diagram of an anomaly detection device proposed by an embodiment of the present application;

[0024] Figure 12 Shows a structural block diagram of an electronic device proposed by the present application;

[0025] Figure 13 Is a storage unit for storing or carrying program codes for implementing a parameter acquisition method according to an embodiment of the present application. Detailed implementation manners

[0026] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts shall fall within the protection scope of the present application.

[0027] With the continuous growth of power demand, the power system has increasing requirements for the economy and reliability of power operation equipment. Due to the action of long-term load and the influence of natural environment (temperature, air pressure, humidity, pollution, etc.), it will cause the aging and wear of power operation equipment, thus gradually reducing the performance and reliability of power operation equipment and posing potential safety hazards. Therefore, it is very necessary to monitor and detect the operation state of power operation equipment.

[0028] The inventors found in the research on the anomaly detection of related power operation equipment that the traditional method for anomaly detection of power operation equipment requires manual investigation and has low efficiency. The method for anomaly detection of power operation equipment based on a deep neural network requires a large amount of anomaly data to train network parameters. However, in the actual production environment, the number of anomaly data is small, resulting in poor performance of the deep neural network.

[0029] Therefore, the inventors propose a model generation method, an anomaly detection method, a device, and an electronic device in the present application. After obtaining a training data set including the first audio features of normal audio information and abnormal audio information, the to-be-trained generator network is trained by the training data set to use the converged to-be-trained generator network as an initial anomaly detection model, and then the initial anomaly detection model is adjusted by the discriminator network to use the adjusted generator network as a target anomaly detection model. By the above method, after the target anomaly detection model is trained by the first audio features of normal audio information and abnormal audio information, during the process of performing anomaly detection on the device to be detected, the first audio features of the device to be detected can be input into the target anomaly detection model to perform anomaly detection on the device to be detected, saving manpower and improving the anomaly detection efficiency. Moreover, by adjusting the initial anomaly detection model by the discriminator network, the target anomaly detection model can also have good performance in the case of lack of abnormal audio training data, that is, it can have a high accuracy rate in distinguishing normal audio and abnormal audio.

[0030] Please refer to Figure 1 , a model generation method provided by the present application, which is applied to an electronic device, and the method includes:

[0031] S110: Obtain a training data set, where the training data set includes the first audio features of multiple audio information of a target device, and the multiple audio information includes normal audio information and abnormal audio information.

[0032] In an embodiment of the present application, the target device may be a device selected for anomaly detection. Among them, the target device may be an electric power operation device, such as: a generator, a motor, a transformer, etc.

[0033] Among them, as Figure 2 shown, obtaining the training data set includes:

[0034] S111: Obtain multiple audio information of the target device.

[0035] Among them, as a way, the sound generated by the target device during operation due to its own internal structure or hardware conditions can be sampled multiple times through an audio acquisition device (such as a tape recorder, etc.), and the multiple sampled audio can be used as multiple audio information of the target device. Among them, the multiple audio information can include normal audio information (the audio when the target device is operating normally) and abnormal audio information (the audio when the target device has an abnormality, for example: the sound when a transformer has overcurrent due to an external short circuit, or overloading due to the load exceeding the rated capacity for a long time, etc.). Exemplarily, when the target device is a transformer, multiple audio during normal operation and abnormal conditions of the transformer are sampled through an audio acquisition device, and multiple audio information with a sampling rate of 16 kHz and a sampling accuracy of 16 bits can be obtained.

[0036] It should be noted that the sampling rate and sampling accuracy are generally related to the hardware conditions of the audio acquisition device. The higher the sampling rate, the more sampling points can be indicated that the audio acquisition device collects per second. For example, when the sampling rate is 16 kHz, it can be indicated that the audio acquisition device collects 16,000 sampling points per second; the higher the sampling accuracy, the larger the representation range of each sampling point data. For example, when the sampling accuracy is 16 bits, it can be indicated that the range represented by each sampling point data is between -32768(2 (16-1) ) and +32767.

[0037] As Figure 3 shown, after the audio acquisition device collects multiple audio information of the target device, the multiple audio information of the target device can be stored and labeled, so as to use the labeled multiple audio information as an audio data set. Exemplarily, the label of the collected normal audio information can be set to 1, and the label of the collected abnormal audio information can be set to 0.

[0038] It should be noted that there are multiple ways to label the multiple audio information of the target device. It can be manual labeling, or automatic labeling using a classification model, or manual calibration after automatic labeling using a classification model.

[0039] S112: Frame, window, and perform fast Fourier transform on the audio information to obtain the spectrogram corresponding to the audio information.

[0040] Among them, as a way, the audio data set including multiple audio information of the target device can be preprocessed and the result of the preprocessing can be stored as a training data set. The training data set includes the spectrograms corresponding to multiple audio information respectively, so as to convert the audio information into image information and input it into a deep neural network for training. Among them, the preprocessing can include frame, window, and fast Fourier transform, as Figure 4As shown, the labeled audio can be first subjected to frame segmentation and windowing, and then the fast Fourier transform is performed on each window to obtain the spectrogram corresponding to the labeled audio. Exemplarily, when the duration of each audio collected by the audio acquisition device is 1 s, the sampling rate is 16 kHz, the frame length is 25 ms, the frame shift is 10 ms, and the window is a Hanning window, each audio information corresponds to 16,000 sampling point data, and each frame in each audio information corresponds to 400 sampling point data. When the window length is the same as the frame length, after multiplying each frame of data by the Hanning window function, 400 windowed sampling point data can be obtained. After performing a 512-point fast Fourier transform on the windowed sampling point data, 512 spectral lines corresponding to one frame of data can be obtained. After completing this calculation, the Hanning window can be moved backward by 10 ms (160 sampling point data) for the next calculation until all frames are calculated, and the spectrogram corresponding to each audio information can be obtained.

[0041] As another approach, an audio data set including multiple audio information of the target device can be processed in real time through frame segmentation, windowing, and fast Fourier transform to obtain the spectrograms corresponding to the above-mentioned multiple audio information respectively, so as to convert the audio information into image information and input it into the deep neural network for training in real time.

[0042] It should be noted that the frame length, frame shift, window function, and the number of points of the fast Fourier transform can be determined according to actual requirements. Exemplarily, considering the real-time performance of the anomaly detection model and the resolution of the audio information, the number of points of the fast Fourier transform can be set to 512, so that rich audio information can be extracted and the model can maintain a relatively fast calculation speed.

[0043] Furthermore, it should be noted that in actual production and life, the situation of abnormal operation of power operation equipment is relatively rare, so that the number of spectrograms corresponding to normal audio in the training data set may be more than the number of spectrograms corresponding to abnormal audio. Therefore, the training data set in the embodiments of the present application can be unbalanced.

[0044] S113: Use the spectrogram as the first audio feature of the audio information.

[0045] Among them, as a way, multiple spectrograms of the target device can be used as the first audio features of multiple audio information of the target device respectively.

[0046] S120: Train the generator network to be trained through the training data set, and use the converged generator network to be trained as the initial anomaly detection model.

[0047] Among them, the generator network to be trained (Generator, G) can be used to generate a second audio feature that conforms to the distribution of the first audio feature based on the first audio feature, and perform feature extraction on the second audio feature. In the embodiments of the present application, the generator network to be trained may include a first audio feature reconstruction network and a feature extraction network. Among them, the first audio feature reconstruction network may include a first encoder, an LSTM (Long Short-Term Memory), and a decoder, and the feature extraction network may include a second encoder. As a way, the first audio feature can be input into the first encoder, and the first audio feature can be encoded through a non-linear transformation to obtain a low-dimensional feature of the first audio feature; then the encoded feature (the low-dimensional feature of the first audio feature) can be input into the LSTM. Since the LSTM has good ability to extract time information and the audio features in the present application are related to time, the LSTM can be used to further extract the encoded feature to obtain a more effective latent representation (the low-dimensional feature of the first audio feature), so that the generator network to be trained can learn the features of normal audio and abnormal audio in the time dimension respectively, thereby improving the discrimination ability of the generator network to be trained for normal audio and abnormal audio; then the latent representation can be input into the decoder, and the latent representation can be decoded through an inverse mapping to reconstruct the first audio feature. At this time, the reconstructed first audio feature can be called the second audio feature, where the first audio feature and the second audio feature have the same size; after obtaining the second audio feature, the second audio feature can be input into the second encoder to perform further feature extraction and dimensionality reduction on the second audio feature.

[0048] Optionally, the generator network to be trained may further include a fully connected layer. After the fully connected layer, the softmax activation function can be used to output the abnormal detection result of the audio information corresponding to the first audio feature. By training the generator network to be trained with a training data set including multiple first audio features, a converged generator network to be trained can be obtained, and the converged generator network to be trained can be used as an initial abnormal detection model.

[0049] Optionally, in order to reduce the loss of feature information, in the generator network to be trained, a residual connection can be adopted between the intermediate features of the same size of the first encoder and the decoder. Exemplarily, as Figure 5 shown, in the first encoder part, the first audio feature can obtain an intermediate feature after a 3×3 convolution operation. Similarly, in the decoder part, the output of the LSTM can obtain an output feature after a 3×3 deconvolution operation. Adding the output feature to the intermediate feature of the first encoding can obtain an intermediate feature of the decoder, and the intermediate feature of the decoder has the same size as the intermediate feature of the first encoder.

[0050] It should be noted that the structures of the first encoder and the second encoder may be the same. Exemplarily, as Figure 5 shown, the first encoder and the second encoder of the generator network to be trained may each include 3 two-dimensional convolutional layers of 3×3, and the decoder may include 3 two-dimensional transposed convolutional layers of 3×3. In addition, the generator network to be trained may further include an LSTM layer and a fully connected layer.

[0051] Furthermore, it should be noted that the network depths of the encoder, decoder, LSTM layer, and fully connected layer in the generator network to be trained and the parameters corresponding to each layer of the network (e.g., the sizes of the convolutional layer and the transposed convolutional layer, etc.) can be flexibly set according to different target devices, the size of the first audio feature, etc. Optionally, in order to improve the performance of the model, an attention mechanism can be introduced into the generator network to be trained.

[0052] S130: Adjust the initial anomaly detection model through the discriminator network to use the adjusted generator network as the target anomaly detection model.

[0053] Among them, the discriminator network (Discriminator, D) can be used to determine whether the distributions of the first audio feature and the second audio feature in the initial anomaly detection model are consistent. In the embodiments of the present application, the discriminator network may include a third encoder. Optionally, the network structure of the third encoder may be the same as that of the first encoder or the second encoder.

[0054] As a way, since the training purpose of the generator network in the initial anomaly detection model can be to generate a second audio feature that conforms to the distribution of the first audio feature (true distribution), and the training purpose of the discriminator network can be to correctly determine whether the distributions of the first audio feature and the second audio feature are consistent, so adjusting the network parameters (e.g., weights, etc.) of the initial anomaly detection model through the discriminator network can enable the generator network in the initial anomaly detection model to learn the data distributions of the first audio features corresponding to normal audio information and abnormal audio information respectively during the process of competing with the discriminator network, so that the adjusted generator network can be used as the target anomaly detection model to perform anomaly detection on the input audio features.

[0055] As another approach, it is possible to determine whether to adjust the initial anomaly detection model through the discriminator network based on the proportion of anomalous audio information in the training dataset. If the proportion of anomalous audio information is less than a threshold value, the initial anomaly detection model is adjusted through the discriminator network, and the adjusted generator network is used as the target anomaly detection model. Exemplarily, assume that the threshold value is A and the proportion of anomalous audio information in the training dataset is B. If B is less than A, it indicates that the training samples of anomalous audio information are insufficient, which may cause the initial anomaly detection model to fail to learn the data distribution of the first audio features corresponding to the anomalous audio information, resulting in the initial anomaly detection model being unable to accurately distinguish between normal audio and anomalous audio. At this time, the network parameters (such as weights, etc.) of the initial anomaly detection model can be adjusted through the discriminator network, so that the generator network in the initial anomaly detection model can learn the data distributions of the first audio features corresponding to normal audio information and anomalous audio information respectively during the process of competing with the discriminator network, so that the adjusted generator network can be used as the target anomaly detection model to perform anomaly detection on the input audio features.

[0056] A model generation method provided in this embodiment, after obtaining a training dataset including the first audio features of normal audio information and anomalous audio information, trains the generator network to be trained through this training dataset, and uses the converged generator network to be trained as the initial anomaly detection model, and then adjusts the initial anomaly detection model through the discriminator network to use the adjusted generator network as the target anomaly detection model. Through the above method, after obtaining the target anomaly detection model by training with the first audio features of normal audio information and anomalous audio information, during the process of performing anomaly detection on the device to be detected, the first audio features of the device to be detected can be input into the target anomaly detection model to perform anomaly detection on the device to be detected, saving manpower and improving the anomaly detection efficiency. Moreover, by adjusting the initial anomaly detection model through the discriminator network, the target anomaly detection model can also have good performance in the case of a lack of anomalous audio training data, that is, it can have a high accuracy rate in discriminating between normal audio and anomalous audio.

[0057] Please refer to Figure 6 , a model generation method provided in this application, is applied to an electronic device, and the method includes:

[0058] S210: Obtain a training dataset, where the training dataset includes the first audio features of multiple audio information of a target device, and the multiple audio information includes normal audio information and anomalous audio information.

[0059] S220: Input the training dataset into the generator network to be trained, and obtain the output of the generator network to be trained.

[0060] Among them, as a way, the first audio features corresponding to the normal audio information and the abnormal audio information can be input into the generator network to be trained, and the output of the generator network to be trained is obtained. In this way, there can be multiple normal audio information and multiple abnormal audio information, and each normal or abnormal audio information corresponds to a first audio feature.

[0061] S230: Train the generator network to be trained based on the output, the first loss function, and the second loss function, so as to use the converged generator network to be trained as the initial anomaly detection model, where the first loss function is the absolute value of the difference between the output results of the first encoder and the second encoder, and the second loss function is the absolute value of the difference between the first audio feature and the second audio feature, and the second audio feature is the output result of the decoder.

[0062] Among them, in the embodiments of the present application, the first loss function can be used to minimize the distance between the output features of the first encoder and the second encoder in the generator network to be trained (Generator, G), so that the generator network to be trained can learn the respective coding feature distributions of the normal audio information and the abnormal audio information. As a way, the first loss function can be the absolute value of the difference between the output results of the first encoder and the second encoder, and the calculation formula of the first loss function is as follows:

[0063] Loss_g1 = ‖z1 - z2‖

[0064] Among them, z1 can represent the output result of the first encoder, and z2 can represent the output result of the second encoder.

[0065] Furthermore, in the embodiments of the present application, the second loss function can be used to minimize the distance between the first audio feature and the second audio feature in the generator network to be trained (Generator, G), so that the generator network to be trained can learn the respective texture feature distributions of the normal audio information and the abnormal audio information. As a way, the second loss function can be the absolute value of the difference between the first audio feature and the output result of the decoder, and the calculation formula of the second loss function is as follows:

[0066] Loss_g2 = ‖x - G(x)‖

[0067] Among them, x can represent the first audio feature, and G(x) can represent the output result of the decoder.

[0068] As a way, the weighted sum of the first loss function and the second loss function can be used as the loss function of the generator network to be trained. Based on the output of the generator network to be trained and the loss function of the generator network to be trained, the generator network to be trained is trained to use the converged generator network to be trained as the initial anomaly detection model. The calculation formula of the loss function of the generator network to be trained is as follows:

[0069] Loss_G = xLoss_g1 + yLoss_g2

[0070] Wherein, the sum of x and y is 1, and the values of x and y can be set based on experience or obtained through training as trainable parameters of the generator network to be trained.

[0071] S240: Adjust the initial anomaly detection model through the discriminator network to use the adjusted generator network as the target anomaly detection model.

[0072] A model generation method provided in this embodiment enables, through the above method, that after training the target anomaly detection model through the first audio features of normal audio information and abnormal audio information, during the process of anomaly detection of the device to be detected, the first audio features of the device to be detected can be input into the target anomaly detection model to perform anomaly detection on the device to be detected, saving manpower and improving the anomaly detection efficiency. Moreover, by adjusting the initial anomaly detection model through the discriminator network, the target anomaly detection model can also have good performance in the case of lack of abnormal audio training data, that is, it can have a high accuracy rate in the discrimination of normal audio and abnormal audio. And in this embodiment, the generator network to be trained is trained through the output of the generator network to be trained, the first loss function and the second loss function to use the converged generator network to be trained as the initial anomaly detection model, so that the initial anomaly detection model can obtain the respective feature distribution situations of normal audio information and abnormal audio information, improving the discrimination ability of the initial anomaly detection model for normal audio and abnormal audio.

[0073] Please refer to Figure 7 , a model generation method provided in this application, is applied to an electronic device, and the method includes:

[0074] S310: Obtain a training data set, where the training data set includes the first audio features of multiple audio information of the target device, and the multiple audio information includes normal audio information and abnormal audio information.

[0075] S320: Train the generator network to be trained through the training data set to use the converged generator network to be trained as the initial anomaly detection model.

[0076] S330: Obtain the anomaly detection model to be trained, where the anomaly detection model to be trained includes the initial anomaly detection model and the third encoder.

[0077] Among them, as a way, as Figure 8 shown, the anomaly detection model to be trained may include a first encoder, a decoder, a second encoder, and a third encoder. Among them, the third encoder can be used as a discriminator network (Discriminator, D), and the input of the third encoder can be the first audio feature and the output of the decoder (the second audio feature).

[0078] Optionally, the structures of the first encoder, the second encoder, and the third encoder may be the same.

[0079] S340: Input the training data set into the anomaly detection model to be trained to obtain the output of the anomaly detection model to be trained.

[0080] Among them, as a way, multiple first audio features corresponding to normal audio information and abnormal audio information can be input into the anomaly detection model to be trained to obtain the output of the anomaly detection model to be trained.

[0081] S350: Adjust the anomaly detection model to be trained based on the output, the first loss function, the second loss function, and the third loss function to obtain a converged anomaly detection model to be trained. Among them, the third loss function is the absolute value of the difference between the third audio feature and the fourth audio feature. The third audio feature is the output result of the third encoder corresponding to the first audio feature, and the fourth audio feature is the output result of the third encoder corresponding to the second audio feature.

[0082] Among them, in the embodiments of the present application, the third loss function can be used to minimize the distance between the discriminator network output features corresponding to the first audio feature and the discriminator network output features corresponding to the second audio feature, so that the anomaly detection model to be trained can learn features that can deceive the discriminator network (Discriminator, D), that is, the discriminator network cannot confirm whether the second audio feature is a generated feature. As a way, the third loss function can be the absolute value of the difference between the third audio feature and the fourth audio feature, and the calculation formula of the third loss function is as follows:

[0083] Loss_d = ‖D(x) - D(G(x))‖

[0084] Among them, D(x) can represent the third audio feature, and the third audio feature can be the output result obtained by inputting the first audio feature into the discriminator network; D(G(x)) can represent the fourth audio feature, and the fourth audio feature can be the output result obtained by inputting the second audio feature into the discriminator network.

[0085] As a way, the weighted sum of the first loss function, the second loss function, and the third loss function can be used as the loss function of the anomaly detection model to be trained. The generator network to be trained is trained through the output of the anomaly detection model to be trained and the loss function of the anomaly detection model to be trained, so as to use the converged generator network to be trained as the initial anomaly detection model. The calculation formula of the loss function of the generator network to be trained is as follows:

[0086] Loss=xLoss_g1+yLoss-g2+zLoss_d

[0087] Wherein, the sum of x, y, and z is 1. The values of x, y, and z can be set based on experience or obtained through training as the trainable parameters of the anomaly detection model to be trained.

[0088] S360: Use the generator network in the converged anomaly detection model to be trained as the target anomaly detection model.

[0089] The model generation method provided in this embodiment enables, through the above method, that after training the target anomaly detection model through the first audio features of normal audio information and abnormal audio information, during the process of anomaly detection of the device to be detected, the first audio features of the device to be detected can be input into the target anomaly detection model to perform anomaly detection on the device to be detected, saving manpower and improving the anomaly detection efficiency. Moreover, by adjusting the initial anomaly detection model through the discriminator network, the target anomaly detection model can also have good performance in the case of lack of abnormal audio training data, that is, it can have a high accuracy rate in the discrimination of normal audio and abnormal audio. And, in this embodiment, by using the discriminator network to discriminate the authenticity of the audio features of the generator network in the anomaly detection model to be trained (the label of the first audio feature is true, and the label of the second audio feature is false), the ability of the generator network to obtain normal audio features can be improved, thereby increasing the difference between the first audio features corresponding to normal audio and abnormal audio respectively after feature extraction through the generator network, and further improving the performance of the generator network in the anomaly detection model to be trained, that is, the performance of the target anomaly detection model.

[0090] Please refer to Figure 9 , an anomaly detection method provided in this application, which is applied to an electronic device. The method includes:

[0091] S410: Obtain the audio to be detected.

[0092] Among them, the audio to be detected can be the sound emitted by power operation equipment (such as generators, motors, transformers, etc.) during operation, and this sound can be emitted due to the internal structure or hardware conditions of the power equipment itself. As a way, the audio to be detected can be periodically obtained through an audio acquisition device, so that real-time detection of the power operation equipment can be carried out, so that when the power operation equipment has an abnormality, it can be discovered and maintained in time to avoid the occurrence of potential safety hazards. Exemplarily, the audio to be detected can be obtained through the audio acquisition device every 2s.

[0093] S420: Frame, window, and perform fast Fourier transform on the audio to be detected to obtain a first audio feature corresponding to the audio to be detected, where the first audio feature is a spectrogram corresponding to the audio to be detected.

[0094] In the embodiments of the present application, the spectrogram can be a two-dimensional image, and the size of the spectrogram can be related to the number of points of the fast Fourier transform and the number of frames after the audio to be detected is framed and windowed. Among them, the first-dimensional size of the spectrogram can be obtained through the formula: number of fast Fourier transform points / 2 + 1, where 1 can represent the DC component; the second-dimensional size of the spectrogram can be obtained through the formula: (sampling rate × audio duration - sampling rate × frame length) / (sampling rate × frame shift) + 1, where the units of the frame length and the frame shift are s. Exemplarily, when the sampling rate is 16 kHz, the frame length is 25 ms, the frame shift is 10 ms, and the number of points of the fast Fourier transform is 512, a spectrogram of 257×198 can be obtained for the audio to be detected with a duration of 2 s.

[0095] S430: Input the first audio feature into the target anomaly detection model and obtain the detection result output by the target anomaly detection model.

[0096] Among them, as a way, the spectrogram corresponding to the audio to be detected can be input into the target anomaly detection model, and the target anomaly detection model can output whether the audio to be detected is an abnormal audio. If it is an abnormal audio, it indicates that the power operation equipment corresponding to the audio to be detected has an abnormality, and it is necessary to troubleshoot the faults of the power equipment; if it is a normal audio, it indicates that the power operation equipment corresponding to the audio to be detected is in a normal working state.

[0097] An anomaly detection method provided in this embodiment enables, through the above method, during the process of anomaly detection of the device to be detected, the first audio feature of the device to be detected can be input into the target anomaly detection model to perform anomaly detection on the device to be detected, saving manpower and improving the anomaly detection efficiency.

[0098] Please refer to Figure 10 , a model generation device 600 provided in the present application, runs on an electronic device, and the device 600 includes:

[0099] A dataset acquisition unit 610 is configured to acquire a training dataset, where the training dataset includes first audio features of multiple audio messages of a target device, and the multiple audio messages include normal audio messages and abnormal audio messages.

[0100] An initial anomaly detection model acquisition unit 620 is configured to train a generator network to be trained by using the training dataset, so as to use the converged generator network to be trained as an initial anomaly detection model.

[0101] A target anomaly detection model acquisition unit 630 is configured to adjust the initial anomaly detection model by using a discriminator network, so as to use the adjusted generator network as a target anomaly detection model.

[0102] Wherein, as a way, the dataset acquisition unit 610 is specifically configured to acquire multiple audio messages of the target device; perform frame segmentation, windowing, and fast Fourier transform on the audio messages to obtain a spectrogram corresponding to the audio messages; and use the spectrogram as the first audio feature of the audio messages.

[0103] As a way, the generator network includes a first audio feature reconstruction network and a feature extraction network. The first audio feature reconstruction network includes a first encoder, an LSTM, and a decoder. The feature extraction network includes a second encoder. The initial anomaly detection model acquisition unit 620 is specifically configured to input the training dataset into the generator network to be trained to obtain an output of the generator network to be trained; and train the generator network to be trained based on the output, a first loss function, and a second loss function to obtain an initial anomaly detection model, where the first loss function is the absolute value of the difference between the output result of the first encoder and the output result of the second encoder, and the second loss function is the absolute value of the difference between the first audio feature and a second audio feature, and the second audio feature is the output result of the decoder.

[0104] As a way, the discriminator network includes a third encoder. The target anomaly detection model acquisition unit 630 is specifically configured to acquire a to-be-trained anomaly detection model, where the to-be-trained anomaly detection model includes the initial anomaly detection model and the third encoder; input the training data set into the to-be-trained anomaly detection model to obtain the output of the to-be-trained anomaly detection model; adjust the to-be-trained anomaly detection model based on the output, the first loss function, the second loss function, and the third loss function to obtain a converged to-be-trained anomaly detection model, where the third loss function is the absolute value of the difference between the third audio feature and the fourth audio feature, the third audio feature is the output result of the third encoder corresponding to the first audio feature, and the fourth audio feature is the output result of the third encoder corresponding to the second audio feature; use the generator network in the converged to-be-trained anomaly detection model as the target anomaly detection model.

[0105] Optionally, the first encoder, the second encoder, and the third encoder have the same structure.

[0106] Please refer to Figure 11 , an anomaly detection device 800 provided by the present application, which runs on an electronic device. The device 800 includes:

[0107] A detected audio acquisition unit 810, configured to acquire a to-be-detected audio.

[0108] A first audio feature acquisition unit 820, configured to frame, window, and perform fast Fourier transform on the to-be-detected audio to obtain a first audio feature corresponding to the to-be-detected audio, where the first audio feature is a spectrogram corresponding to the to-be-detected audio.

[0109] A detection result acquisition unit 830, configured to input the first audio feature into the target anomaly detection model to acquire a detection result output by the target anomaly detection model.

[0110] Wherein, as a way, the detected audio acquisition unit 810 is specifically configured to periodically acquire the to-be-detected audio.

[0111] Next, a description will be given of an electronic device provided by the present application in conjunction with Figure 12

[0112] Please refer to Figure 12, based on the above model generation method, anomaly detection method, and apparatus, another electronic device 100 capable of executing the foregoing model generation method and anomaly detection method is further provided in an embodiment of the present application. The electronic device 100 includes one or more (only one is shown in the figure) processors 102 and a memory 104 that are coupled to each other. Among them, the memory 104 stores a program that can execute the content in the foregoing embodiments, and the processor 102 can execute the program stored in the memory 104.

[0113] Among them, the processor 102 may include one or more processing cores. The processor 102 connects various parts within the entire electronic device 100 using various interfaces and lines, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 104, and by calling data stored in the memory 104, it executes various functions of the electronic device 100 and processes data. Optionally, the processor 102 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 102 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the modem is used to process wireless communications. It can be understood that the above modem may not be integrated into the processor 102 and may be implemented separately through a communication chip.

[0114] The memory 104 may include random access memory (RAM) and may also include read-only memory. The memory 104 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the terminal 100 (such as phone book, audio and video data, chat record data, etc.).

[0115] Please refer to Figure 13, which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable storage medium 1000, and the program code can be called by a processor to execute the method described in the above method embodiment.

[0116] The computer-readable storage medium 1000 can be an electronic memory such as a flash memory, EEPROM (electrically erasable programmable read-only memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 800 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 1000 has a storage space for the program code 1010 that executes any method step in the above method. These program codes can be read from or written into one or more computer program products. The program code 1010 can be compressed in a suitable form, for example.

[0117] In summary, for a model generation method, an anomaly detection method, an apparatus, and an electronic device provided by the present application, after obtaining a training data set including first audio features of normal audio information and abnormal audio information, a generator network to be trained is trained through the training data set to use the converged generator network to be trained as an initial anomaly detection model, and then the initial anomaly detection model is adjusted through a discriminator network to use the adjusted generator network as a target anomaly detection model. By the above method, after training a target anomaly detection model through the first audio features of normal audio information and abnormal audio information, during the process of performing anomaly detection on a device to be detected, the first audio features of the device to be detected can be input into the target anomaly detection model to perform anomaly detection on the device to be detected, saving manpower and improving the anomaly detection efficiency. Moreover, by adjusting the initial anomaly detection model through the discriminator network, the target anomaly detection model can have better performance even in the case of a lack of abnormal audio training data, that is, it can have a high accuracy rate in distinguishing normal audio and abnormal audio.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A model generation method, characterized in that, Applied to an electronic device, the method includes: Obtain a training data set, where the training data set includes first audio features of multiple audio information of a target device, and the multiple audio information includes normal audio information and abnormal audio information; Train a generator network to be trained through the training data set to use the converged generator network to be trained as an initial anomaly detection model; wherein, the generator network to be trained includes a first audio feature reconstruction network and a second encoder connected in series, and the first audio feature reconstruction network includes a first encoder, an LSTM, and a decoder connected in series; the first encoder includes n layers of two-dimensional convolutional layers, the decoder includes n layers of two-dimensional transposed convolutional layers, and there is a residual connection between the output of the k-th layer of two-dimensional convolutional layers in the first encoder and the output of the (n - k + 1)-th layer of two-dimensional transposed convolutional layers in the decoder, 1 ≤ k ≤ n, and n and k are integers; Adjust the initial anomaly detection model through a discriminator network to use the adjusted generator network as a target anomaly detection model.

2. The method according to claim 1, wherein The step of training the generator network to be trained through the training data set to use the converged generator network to be trained as an initial anomaly detection model includes: Input the training data set into the generator network to be trained to obtain the output of the generator network to be trained; Train the generator network to be trained based on the output, a first loss function, and a second loss function to use the converged generator network to be trained as an initial anomaly detection model, wherein the first loss function is the absolute value of the difference between the output results of the first encoder and the second encoder, and the second loss function is the absolute value of the difference between the first audio feature and a second audio feature, and the second audio feature is the output result of the decoder.

3. The method according to claim 2, wherein, The discriminator network includes a third encoder. The step of adjusting the initial anomaly detection model through the discriminator network to use the adjusted generator network as a target anomaly detection model includes: Obtain an anomaly detection model to be trained, where the anomaly detection model to be trained includes the initial anomaly detection model and the third encoder; Input the training data set into the anomaly detection model to be trained to obtain the output of the anomaly detection model to be trained; Adjust the anomaly detection model to be trained based on the output, the first loss function, the second loss function, and a third loss function to obtain a converged anomaly detection model to be trained, wherein the third loss function is the absolute value of the difference between a third audio feature and a fourth audio feature, the third audio feature is the output result of the third encoder corresponding to the first audio feature, and the fourth audio feature is the output result of the third encoder corresponding to the second audio feature; Use the generator network in the converged anomaly detection model to be trained as a target anomaly detection model.

4. The method according to claim 3, characterized in that, The first encoder, the second encoder, and the third encoder have the same structure.

5. The method according to claim 1, wherein The obtaining of the training data set, where the training data set includes first audio features of multiple audio messages of a target device, and the multiple audio messages include normal audio messages and abnormal audio messages, includes: Obtain multiple audio messages of the target device; Perform frame segmentation, windowing, and fast Fourier transform on the audio messages to obtain a spectrogram corresponding to the audio messages; Use the spectrogram as the first audio feature of the audio messages.

6. An anomaly detection method, characterized in that, When applied to an electronic device, the method includes: Obtain an audio message to be detected; Perform frame segmentation, windowing, and fast Fourier transform on the audio message to be detected to obtain a first audio feature corresponding to the audio message to be detected, where the first audio feature is a spectrogram corresponding to the audio message to be detected; Input the first audio feature into a target anomaly detection model obtained by any one of claims 1-5 to obtain a detection result output by the target anomaly detection model.

7. The method according to claim 6, characterized in that When applied to an electronic device, the obtaining of the audio message to be detected includes: Periodically obtain an audio message to be detected.

8. A model generation device, characterized in that, When running on an electronic device, the apparatus includes: A data set obtaining unit, configured to obtain a training data set, where the training data set includes first audio features of multiple audio messages of a target device, and the multiple audio messages include normal audio messages and abnormal audio messages; An initial anomaly detection model obtaining unit, configured to train a generator network to be trained through the training data set, and use the converged generator network to be trained as an initial anomaly detection model; where the generator network to be trained includes a first audio feature reconstruction network and a second encoder connected in series, the first audio feature reconstruction network includes a first encoder, an LSTM, and a decoder connected in series; the first encoder includes n layers of two-dimensional convolutional layers, the decoder includes n layers of two-dimensional transposed convolutional layers, and there is a residual connection between the output of the kth layer of two-dimensional convolutional layers in the first encoder and the output of the (n-k+1)th layer of two-dimensional transposed convolutional layers in the decoder, 1≤k≤n, and n and k are integers; A target anomaly detection model obtaining unit, configured to adjust the initial anomaly detection model through a discriminator network, and use the adjusted generator network as a target anomaly detection model.

9. An anomaly detection device, characterized in that, When running on an electronic device, the apparatus includes: An audio message to be detected obtaining unit, configured to obtain an audio message to be detected; A first audio feature obtaining unit, configured to perform frame segmentation, windowing, and fast Fourier transform on the audio message to be detected to obtain a first audio feature corresponding to the audio message to be detected, where the first audio feature is a spectrogram corresponding to the audio message to be detected; A detection result obtaining unit, configured to input the first audio feature into a target anomaly detection model obtained by any one of claims 1-5 to obtain a detection result output by the target anomaly detection model.

10. An electronic device, characterized in that, Includes one or more processors and a memory; One or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs are configured to execute any one of the methods of claims 1-7.

11. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, wherein the method according to any one of claims 1-7 is executed when the program code runs.

Citation Information

Patent Citations

  • Audio abnormality detection method based on confrontation network generation

    CN109461458A