An audio detection method, detection device, and storage medium
By using a data reconstruction model based on normal sound data training, the reconstruction error of the audio characteristics to be tested is solved, and the problem of high time cost in the prior art is realized quickly identifying the anomalies of audio data.
Patent Information
- Application Number
- CN202211566590.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-07
AI Technical Summary
The existing anomaly sound detection algorithm requires additional acquisition of large amounts of abnormal sound data for training, resulting in high time cost and it is difficult to efficiently identify bad audio in live broadcasts on the Internet.
A data reconstruction model based on normal sound data is adopted to calculate the reconstruction error of the audio characteristics to be measured and the target audio characteristics to be measured, so as to determine whether the audio data is abnormal, and avoid additional acquisition of abnormal sound data.
When detecting audio data, it is not necessary to spend a lot of time to collect abnormal sound data, which can quickly identify the abnormality of audio data and improve detection efficiency.
Smart Images

Figure CN116434772B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of data processing, and in particular, to an audio detection method, a detection device, and a storage medium. Background Art
[0002] In recent years, the network live broadcast industry has developed rapidly and has become an important social medium reflecting the development of participatory culture. In the network live broadcast industry, there are generally three different content production methods: user-generated content (UGC), professionally-generated content (PGC), and professional user-generated content (PUGC). The content is diverse, from text to text + pictures, and then to text + pictures + videos + live broadcasts. The content is rich in form and shows explosive growth. However, the uneven quality of the content also brings great challenges to the platform in terms of content review.
[0003] The forms of expression of bad content are diverse. In addition to using pictures as a carrier, it will also be spread in the form of audio, hiding a large amount of bad audio under normal pictures, attempting to avoid being detected by the detection system. The existing methods mainly focus on detecting abnormal sounds in scenarios such as network live broadcasts, and the task of identifying whether the sound emitted from an object is a normal sound or an abnormal sound. Abnormal sound detection mainly includes sound event detection and anomaly detection; anomaly detection refers to the problem of finding data that does not conform to the expected behavior pattern in the data. In different application fields, these non-conforming patterns are usually called anomalies, and anomaly detection is widely used in various applications.
[0004] The existing abnormal sound detection algorithms generally input abnormal sound data into an algorithm model. The algorithm model extracts the abnormal sound characteristics of the abnormal sound data. When performing detection, it determines whether the current audio data is abnormal sound data based on the characteristics of the current audio data and the abnormal sound characteristics. However, the database corresponding to the algorithm model generally stores normal sound data, and the algorithm model needs to collect an additional sufficient amount of abnormal sound data for training during detection, resulting in a relatively high time cost in the implementation process of the existing abnormal sound detection algorithms. Summary of the Invention
[0005] Embodiments of the present application provide an audio detection method, a detection device, and a storage medium, which can detect audio data without spending a relatively high time cost to collect abnormal sound data additionally.
[0006] Embodiments of the present application provide an audio detection method, including:
[0007] Obtain the audio data to be measured and the first data reconstruction model; the first data reconstruction model is trained based on the audio features of normal sound data, and the normal sound data is pre-stored audio data that conforms to the preset content rules; when the audio data input to the first data reconstruction model is the normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is less than the preset value, and when the audio data input to the first data reconstruction model is not the normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is greater than the preset value;
[0008] Obtain the audio features to be measured of the audio data to be measured according to the preset rules;
[0009] Input the audio features to be measured into the first data reconstruction model to obtain the reconstructed target audio features;
[0010] Calculate the reconstruction error between the audio features to be measured and the target audio features, and determine whether the audio data to be measured is abnormal.
[0011] Further, the obtaining of the first data reconstruction model includes:
[0012] Obtain the initial data reconstruction model and the audio feature training set of the first normal sound data;
[0013] Input the audio feature training set of the first normal sound data into the initial data reconstruction model to obtain the output audio features;
[0014] Based on the output audio features and the audio feature training set of the first normal sound data, determine the loss value;
[0015] Based on the loss value, determine whether the initial data reconstruction model converges;
[0016] If it does not converge, obtain the audio feature training set of the second normal sound data and input the audio feature training set of the second normal sound data into the initial data reconstruction model;
[0017] If it converges, the training is completed to obtain the first data reconstruction model.
[0018] Further, the audio features to be measured include: multiple frames of audio vectors; the first data reconstruction model includes: a feature encoder and a feature decoder;
[0019] The inputting the audio features to be measured into the first data reconstruction model to obtain the reconstructed target audio features includes:
[0020] Input multiple frames of the audio vectors into the feature encoder to obtain the intermediate variables of the multiple frames of the audio vectors;
[0021] Input the intermediate variable into the feature decoder to obtain the target audio vector after reconstruction of each frame of the audio vector.
[0022] Further, calculating the reconstruction error between the audio feature to be measured and the target audio feature, and determining whether the audio data to be measured is abnormal includes:
[0023] Calculate the absolute value of the difference between each frame of the audio vector and the corresponding target audio vector to obtain a plurality of error values;
[0024] Take the average of the plurality of error values to obtain an average error value, and use the average error value as the reconstruction error between the audio feature to be measured and the target audio feature;
[0025] If the average error value is greater than a preset error threshold, determine that the audio data to be measured is abnormal sound data;
[0026] If the average error value is less than the preset error threshold, determine that the audio data to be measured is normal sound data.
[0027] Further, after obtaining the audio feature to be measured of the audio data to be measured according to a preset rule, before inputting the audio feature to be measured into the first data reconstruction model to obtain the reconstructed target audio feature, the method further includes:
[0028] Perform data enhancement processing on the audio feature to be measured of the audio data to be measured to obtain an enhanced audio feature;
[0029] The step of inputting the audio feature to be measured into the first data reconstruction model to obtain the reconstructed target audio feature includes:
[0030] Input the enhanced audio feature into the first data reconstruction model to obtain the reconstructed target audio feature.
[0031] Further, if the reconstruction error between the audio feature to be measured and the target audio feature is less than a preset threshold, the method further includes:
[0032] Obtain the identification information of the audio data to be measured and a second data reconstruction model, where the second data reconstruction model is trained based on the identification features of the normal sound data; when the identification feature input to the second data reconstruction model is the identification feature of the normal sound data, the reconstruction error between the reconstructed identification feature of the second data reconstruction model and the input identification feature is less than a preset value, and when the identification feature input to the second data reconstruction model is not the identification feature of the normal sound data, the reconstruction error between the reconstructed identification feature of the second data reconstruction model and the input identification feature is greater than the preset value;
[0033] Obtain the to-be-measured identification feature of the identification information of the audio data to be measured according to a preset rule;
[0034] Input the to-be-measured identification feature into the second data reconstruction model to obtain a reconstructed target identification feature;
[0035] Determine whether the audio data to be measured is abnormal sound data according to the relative situation between the identification information of the normal sound data and the identification information corresponding to the target identification feature.
[0036] Further, the determining whether the audio data to be measured is abnormal sound data according to the relative situation between the identification information of the normal sound data and the identification information corresponding to the target identification feature includes:
[0037] Convert the target identification feature into target identification information;
[0038] If there is identification information in the identification information of the normal sound data that is the same as the target identification information, determine that the audio data to be measured is normal sound data;
[0039] If there is no identification information in the identification information of the normal sound data that is the same as the target identification information, determine that the audio data to be measured is abnormal sound data.
[0040] Further, if the reconstruction error between the audio feature to be measured and the target audio feature is less than a preset threshold, the method further includes:
[0041] Obtain the identification information of the audio data to be measured and a preset list of identification information, where the list of identification information includes: the identification information of the normal sound data and the identification information of other audio data that conforms to the preset content rules;
[0042] If the identification information of the audio data to be measured exists in the preset list of identification information, determine that the audio data to be measured is normal sound data;
[0043] If the identification information of the to-be-detected audio data does not exist in the preset list of identification information, it is determined that the to-be-detected audio data is abnormal sound data.
[0044] The embodiment of the present application also provides an audio detection device, including:
[0045] A first acquisition unit, configured to acquire to-be-detected audio data and a first data reconstruction model; the first data reconstruction model is trained based on the audio of normal sound data, and the normal sound data is pre-stored audio data that conforms to a preset content rule; when the audio data input to the first data reconstruction model is the normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is less than a preset value, and when the audio data input to the first data reconstruction model is not the normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is greater than the preset value;
[0046] A second acquisition unit, configured to acquire the to-be-detected audio features of the to-be-detected audio data according to a preset rule;
[0047] An input unit, configured to input the to-be-detected audio features into the first data reconstruction model to obtain the reconstructed target audio features;
[0048] A determination unit, configured to calculate the reconstruction error between the to-be-detected audio features and the target audio features, and determine whether the to-be-detected audio data is abnormal.
[0049] The embodiment of the present application also provides an audio detection device, including:
[0050] A central processing unit, a memory, and an input / output interface;
[0051] The memory is a transient storage memory or a persistent storage memory;
[0052] The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to perform the above method.
[0053] The embodiment of the present application also provides a computer-readable storage medium, including instructions, which when run on a computer, cause the computer to execute the above method.
[0054] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0055] In the embodiments of the present application, the to-be-detected audio data is acquired, and a first data reconstruction model trained based on the audio features of normal sound data is obtained; the normal sound data is audio data stored in advance that conforms to the preset content rules; the to-be-detected audio features of the to-be-detected audio data are obtained according to the preset rules; the to-be-detected audio features are input into the first data reconstruction model to obtain the reconstructed target audio features; the reconstruction error between the to-be-detected audio features and the target audio features is calculated to determine whether the to-be-detected audio data is abnormal. It can be seen that in the embodiments of the present application, the data reconstruction model trained based on the normal sound data is used to obtain the target audio features, and by calculating the reconstruction error between the to-be-detected audio features and the target audio features of the to-be-detected audio data, it is determined whether the to-be-detected audio data is abnormal. When detecting audio data, it is possible to determine whether the to-be-detected audio data is abnormal by using the pre-stored normal sound data, without spending a large amount of time cost to additionally collect abnormal sound data. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0057] Figure 1 It is a communication network structure diagram of an audio detection device disclosed in the embodiments of the present application;
[0058] Figure 2 It is a flowchart of an audio detection disclosed in the embodiments of the present application;
[0059] Figure 3 It is a flowchart of another audio detection disclosed in the embodiments of the present application;
[0060] Figure 4 It is a schematic structural diagram of an autoencoder disclosed in the embodiments of the present application;
[0061] Figure 5 It is a flowchart of another audio detection disclosed in the embodiments of the present application;
[0062] Figure 6 It is a reconstruction flowchart of an autoencoder disclosed in the embodiments of the present application;
[0063] Figure 7 It is a schematic diagram of an audio detection device disclosed in the embodiments of the present application;
[0064] Figure 8 It is a system architecture diagram of an audio detection device disclosed in the embodiments of the present application;
[0065] Figure 9Schematic diagram of an audio detection device disclosed in an embodiment of the present application. Detailed implementation manners
[0066] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.
[0067] In the following description, reference is made to "a specific implementation manner" or "an embodiment" or similar expressions, which describe a subset of all possible embodiments. However, it can be understood that "a specific implementation manner" or "an embodiment" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict. In the following description, the term "a plurality of" refers to at least two.
[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0069] Existing audio detection devices are mainly used to detect whether there are abnormalities in audio, which can be songs or videos, and no specific limitation is made here. Generally, an audio detection device can perform audio detection on the audio data of an audio player to determine whether the audio data is abnormal, such as Figure 1As shown, the audio detection device 101 is connected to the audio player 102. Before the audio player 102 plays the audio, the audio detection device 101 detects whether the audio data of the to-be-played audio is abnormal. If the detection is abnormal, the playback operation of the audio player 102 is interrupted. If the detection is normal, the audio player 102 continues to execute the playback operation. It can be understood that the volume detection device 101 can be connected to one or more audio players 102, and no specific limitation is made here; this connection can be a wired network connection or a wireless network connection, and no specific limitation is made here; the file format of this audio can be mp3, m4a or wav, and no specific limitation is made here, that is, the audio player 102 can be an MP3 player, a WMA player or an MP4 player, and no specific limitation is made here. The audio detection device 101 can be assembled with the audio player 102 in the same audio playback device, or can be not assembled in the same audio playback device, and no specific limitation is made here. The audio detection device 101 can also periodically collect the audio data in the audio player 102 and detect the audio data. The existing audio detection device 101 detects the audio data through an abnormal sound detection algorithm. This abnormal sound detection algorithm generally inputs the abnormal sound data into an algorithm model. The algorithm model extracts the abnormal sound characteristics of the abnormal sound data. When performing detection, it is determined whether the current audio data is abnormal sound data according to the characteristics of the current audio data and the abnormal sound characteristics.
[0070] However, the database corresponding to the algorithm model generally stores normal sound data. When the algorithm model performs detection, it needs to additionally collect a sufficient amount of abnormal sound data for training, resulting in a relatively high time cost in the implementation process of the existing abnormal sound detection algorithm. Therefore, the embodiments of the present application provide an audio detection method, which can detect audio data without spending a relatively high time cost to additionally collect abnormal sound data. As Figure 2 shown, the specific steps are as follows:
[0071] 201. Obtain the to-be-detected audio data and the first data reconstruction model.
[0072] In the embodiments of the present application, the audio detection device can obtain the to-be-detected audio data. The to-be-detected audio data can be the audio data that the audio player is playing or about to play. This audio data is generally stored in the form of a data signal. The audio detection device is connected to the audio player and can collect the audio data of the audio player or receive the audio data sent by the audio player, and no specific limitation is made here.
[0073] The audio detection device can also obtain a first data reconstruction model, which is trained based on the audio features of normal sound data. The normal sound data is pre-stored audio data that conforms to the preset content rules. It can be understood that the pre-stored audio data can be the audio data stored in the database of the audio detection device or the audio data stored in the database of the audio player, and the audio player is connected to the audio detection device; when the audio detection device detects audio data, it can obtain the pre-stored audio data from the audio detection device itself or from the audio player. Specifically, no limitation is made here.
[0074] The conformity to the preset content rules means that the audio data conforms to the corresponding playback scenario. For example, in the KTV live broadcast scenario, the audio data corresponding to singing voices, speaking voices, and ambient sounds are audio data that conform to the preset content rules. In the voice call scenario, the audio data corresponding to human voices and ambient sounds are audio data that conform to the preset content rules. The audio detection device or the audio player will pre-store this audio data that conforms to the preset content rules.
[0075] The audio features of the normal sound data refer to the feature expressions of the normal sound data, such as audio vectors or audio matrices. Specifically, no limitation is made here. The first data reconstruction model has a specific function. When the audio data input into the first data reconstruction model is normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is less than the preset value. That is, when the difference between the audio features input into the first data reconstruction model and the audio features of the normal sound data is small, the difference between the audio features reconstructed by the first data reconstruction model and the input audio features is small; when the audio data input into the first data reconstruction model is not normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is greater than the preset value. That is, when the difference between the audio features input into the first data reconstruction model and the audio features of the normal sound data is large, the difference between the audio features reconstructed by the first data reconstruction model and the input audio features is large. It can be understood that the reconstruction error being less than the preset value (small difference and small variation) also includes the case where the two audio features are the same. The reconstruction error between the audio features being less than the preset value means that the audio features are relatively similar, that is, the feature distributions of the two audio features are relatively close. When the audio feature is an audio vector, the feature distribution refers to the value of each dimension in the multi-dimensional audio vector. The feature distributions being close means that there are corresponding dimension values that are the same in the two multi-dimensional audio vectors, and the number of the same values reaches the preset threshold. The preset threshold can be 3 or 4. Specifically, no limitation is made here. The difference between the audio features can be represented by cross-entropy, KL divergence, and JS divergence. Specifically, no limitation is made here.
[0076] Specifically, when the Jensen-Shannon divergence of two audio features is less than a preset value, it can be determined that the difference is small. When the Jensen-Shannon divergence of two audio features is greater than or equal to the preset value, it can be determined that the difference is large. The preset value can be 0.4 or 0.5, and no specific limitation is made here.
[0077] 202. Obtain the audio features to be measured of the audio data to be measured according to a preset rule.
[0078] The audio detection device can obtain the audio features to be measured of the audio data to be measured according to a preset rule. It can input the audio data to be measured into an audio preprocessing tool to extract the audio features to be measured of the audio data to be measured. The audio preprocessing tool can be a librosa extractor or a python_speech_features extractor, and no specific limitation is made here. It can be understood that before using normal sound data to train the first data reconstruction model, the audio features of the normal sound data can be extracted by using the audio preprocessing tool first, and then the audio features of the normal sound data can be used to train the first data reconstruction model.
[0079] 203. Input the audio features to be measured into the first data reconstruction model to obtain the reconstructed target audio features.
[0080] After obtaining the audio features to be measured of the audio data to be measured, the audio detection device can input the audio features to be measured of the audio data to be measured into the first data reconstruction model to obtain the reconstructed target audio features. When training the first data reconstruction model, the feature distribution of the audio features of the normal sound data can be recorded. When the audio features to be measured of the audio data to be measured are input, the feature distribution of the audio features to be measured of the audio data to be measured is recognized, and the relationship between the feature distribution of the audio features of the normal sound data and the feature distribution of the audio features to be measured of the audio data to be measured is compared. If the feature distribution of the audio features of the input audio data has been recorded during the training process of the first data reconstruction model, the audio features of the reconstructed audio data are more similar to the audio features of the input audio data. That is, when the feature distribution of the audio features of the normal sound data is relatively close to the feature distribution of the audio features to be measured of the audio data to be measured, the difference between the target audio features reconstructed and output by the first data reconstruction model and the audio features to be measured of the input audio data to be measured is small, that is, the reconstruction error is small. When the feature distribution of the audio features of the normal sound data is quite different from the feature distribution of the audio features to be measured of the audio data to be measured, the difference between the target audio features reconstructed and output by the first data reconstruction model and the audio features to be measured of the input audio data to be measured is large, that is, the reconstruction error is large.
[0081] 204. Calculate the reconstruction error of the audio features to be measured and the target audio features, and determine whether the audio data to be measured is abnormal.
[0082] The audio detection device can calculate the reconstruction errors of the audio features to be measured and the target audio features according to the relative situation between the audio features to be measured of the audio data to be measured and the target audio features, and determine whether the audio data to be measured is abnormal. When the reconstruction errors of the audio features to be measured and the target audio features are small, that is, when the difference between the target audio features reconstructed and output by the first data reconstruction model and the audio features to be measured of the input audio data to be measured is small, it is determined that the audio data to be measured is normal sound data; when the reconstruction errors of the audio features to be measured and the target audio features are large, that is, when the difference between the target audio features reconstructed and output by the first data reconstruction model and the audio features to be measured of the input audio data to be measured is large, it is determined that the audio data to be measured is abnormal sound data.
[0083] In the embodiments of the present application, the audio data to be measured and the first data reconstruction model trained based on the audio features of the normal sound data are obtained; the normal sound data is the audio data that conforms to the preset content rules and is stored in advance; the audio features to be measured of the audio data to be measured are obtained according to the preset rules; the audio features to be measured are input into the first data reconstruction model to obtain the reconstructed target audio features; the reconstruction errors of the audio features to be measured and the target audio features are calculated to determine whether the audio data to be measured is abnormal. It can be seen that the embodiments of the present application use the data reconstruction model trained based on the normal sound data to obtain the target audio features, and determine whether the audio data to be measured is abnormal by calculating the reconstruction errors of the audio features to be measured and the target audio features of the audio data to be measured. When detecting the audio data, it is only necessary to use the pre-stored normal sound data to determine whether the audio data to be measured is abnormal, without spending more time costs to collect abnormal sound data additionally.
[0084] To make the audio detection method in the embodiments of the present application clearer, the method embodiments shown as Figure 2 will be described in detail below. As shown in Figure 3 , specifically as follows:
[0085] 301. Obtain the first data reconstruction model according to the initial data reconstruction model and the audio feature training set.
[0086] The audio detection device can obtain the first data reconstruction model trained based on the audio features of the normal sound data, that is, obtain the first data reconstruction model according to the initial data reconstruction model and the audio feature training set of the normal sound data.
[0087] Specifically, the audio detection device can obtain an initial data reconstruction model and an audio feature training set of the first normal sound data. The audio detection device can obtain the initial data reconstruction model through network connection or pre-storage. Specifically, it is not limited here. The audio feature training set of the first normal sound data includes audio features of multiple normal sound data. Specifically, the audio detection device can extract audio data of a preset duration from the audio files corresponding to multiple normal sound data to obtain multiple audio data. This audio data can be understood as normal sound samples. The preset duration can be 10 seconds or 15 seconds. Specifically, it is not limited here. The number of the multiple audio files can be 900 or 1000. Specifically, it is not limited here. To ensure the accuracy of training, the number of audio files should not be too small. The audio detection device can convert the multiple audio data into audio features of multiple normal sound data to obtain the audio feature training set of normal sound data.
[0088] Next, input the audio feature training set of the first normal sound data into the initial data reconstruction model to obtain the output audio features. In the embodiment of the present application, the training process of the first data reconstruction model is iterative. In each round of iteration, input the audio features of the normal sound data into the first data reconstruction model, and the output audio features corresponding to the iteration process can be obtained. It can be understood that the audio features of the normal sound data input into the first data reconstruction model in each iteration are different, and the audio features output by the first data reconstruction model in each iteration are also different. The audio detection device can determine the loss value based on the output audio features and the audio feature training set of the first normal sound data. Specifically, the mean square error formula can be used as the loss function to obtain the loss value. Determine whether the initial data reconstruction model converges based on the loss value. Specifically, when the loss value is less than the preset threshold, it can be determined that the initial data reconstruction model converges. The preset threshold can be understood as when the loss value is represented by the JS divergence, the preset threshold can be 0.4 or 0.5. Specifically, it is not limited here. If the initial data reconstruction model does not converge, obtain the audio feature training set of the second normal sound data and input the audio feature training set of the second normal sound data into the initial data reconstruction model. It can be understood that the normal sound data corresponding to the audio feature training set of the first normal sound data and the audio feature training set of the second normal sound data are not the same audio data. That is, if the initial data reconstruction model does not converge, continue to input the audio features of other normal sound data into the initial data reconstruction model for training. If it converges, the training is completed to obtain the first data reconstruction model.
[0089] It can be understood that after the first data reconstruction model is obtained upon completion of training, a test set containing multiple voice samples can be used to test the performance of the first data reconstruction model. The multiple voice samples include: normal voice data and abnormal voice data. During the test, when the voice sample input to the first data reconstruction model is the audio feature of normal voice data, if the difference between the audio feature output by the first data reconstruction model and the input audio feature is small, that is, the reconstruction error is small, it is determined that the performance of the first data reconstruction model is excellent; if the difference between the audio feature output by the first data reconstruction model and the input audio feature is large, it is determined that the performance of the first data reconstruction model is poor. When the voice sample input to the first data reconstruction model is the audio feature of abnormal voice data, if the difference between the audio feature output by the first data reconstruction model and the input audio feature is small, it is determined that the performance of the first data reconstruction model is poor; if the difference between the audio feature output by the first data reconstruction model and the input audio feature is large, it is determined that the performance of the first data reconstruction model is excellent. When it is determined that the performance of the first data reconstruction model is poor, the first data reconstruction model needs to be continuously trained.
[0090] 302. Obtain the audio data to be measured, and extract the audio vector of each frame in the audio data to be measured through an audio feature extraction tool.
[0091] It can be understood that the audio feature of the audio data can be an audio vector. After the audio detection device obtains the audio data to be measured, it can extract the audio vector of each frame in the audio data to be measured through an audio feature extraction tool. Specifically, the Mel feature of each frame in the audio data to be measured can be extracted through an audio preprocessing tool. The Mel feature is generally close to the human ear's auditory perception, and an N*M-dimensional Mel spectrogram can be obtained, where N is the number of frames of the audio data to be measured, and M is the dimension of the Mel spectrogram. Specifically, the audio data to be measured can be downsampled based on a preset frequency, and the short-time Fourier transform (STFT) of a preset window is applied to transform the waveform obtained by downsampling, and then through a Mel-scaled filter bank, a Mel spectrogram with M dimensions and N frames is obtained. The Mel spectrogram is normalized to zero mean and unit variance to obtain multiple frames of audio vectors.
[0092] It should be noted that the sequence relationship between steps 301 and 302 is not limited here.
[0093] 303. Perform data augmentation processing on the audio vectors of the audio data to be measured to obtain enhanced audio vectors.
[0094] After the audio detection device extracts the audio vectors of each frame in the audio data to be measured through the audio feature extraction tool, it can perform data enhancement processing on the audio vectors of the audio data to be measured to obtain enhanced audio vectors. It can be understood that Mixup (a data enhancement algorithm) and SpecAugment (a speech recognition data enhancement algorithm) can be used to perform data enhancement processing on the audio vectors of the audio data to be measured to increase the number of audio vectors of the audio data to be measured. Among them, Mixup can randomly generate new audio vectors by linearly weighted summation for different combinations of audio vectors. The labels of the new audio vectors are different from those of the combined audio vectors. The new audio vector can be expressed as:
[0095]
[0096]
[0097] Among them, (x i , y i ), (x j , y j ) are two audio vectors, α is the mixing ratio, that is, the mixing weight of the audio vectors. α can take any value from 0 to 1. x represents the audio vector, and y represents the label of the audio vector.
[0098] SpecAugment enhancement can be directly applied to the Mel spectrogram corresponding to the audio vector. The enhancement methods include masking of frequency step size, time step size, and time warping. Among them, masking of frequency step size and time step size means randomly masking consecutive channels on the frequency axis or consecutive frames on the time axis of the Mel spectrogram. Generally, a mask is used to randomly mask consecutive frames. Each mask, whether horizontal or vertical, has two parameters: the starting position S and the length L. Once the mask direction is selected, all values in the rows (or columns) within the range from S to S + L in the Mel spectrogram are set to 0. And time warping is the deformation of the time-frequency sequence on the time axis of the Mel spectrogram.
[0099] It should be noted that step 303 can be executed or not; in order to improve the accuracy of the data, preferably, step 303 is executed to obtain the enhanced audio vector, and the enhanced audio vector in step 304 is input into the first data reconstruction model.
[0100] 304. Input multiple frames of audio vectors into the first data reconstruction model to obtain the reconstructed target audio vector.
[0101] After obtaining multiple frames of audio vectors of the audio data to be measured, the audio detection device may input the multiple frames of audio vectors into the first data reconstruction model to obtain the reconstructed target audio vectors. The first data reconstruction model may be an autoencoder based on Transformer, and Transformer is a neural network with self-attention layers. The autoencoder includes: a feature encoder (Transformer Encoder) and a feature decoder (Transformer Decoder). As Figure 4 shown, the feature encoder includes a linear layer, a positional encoding layer, N Transformer blocks (TransformerBlock), and a linear layer with an activation function unit (ReLU); the feature decoder includes a linear layer and N Transformer blocks. Among them, the feature encoder further includes: a linear layer provided with a bottleneck layer, and the feature decoder further includes: a linear layer for expanding the vector; the Transformer block includes: a normalization layer (LeyerNorm), a fully connected layer (Dropout), a linear layer that compresses the dimension of the vector into one-half (1 / 2dim), a linear layer with an activation function unit, a linear layer that expands the dimension of the vector to twice (2dim), and a multi-head attention layer (MultiHeadAttenion).
[0102] When inputting multiple frames of audio vectors into the autoencoder, first input the multiple frames of audio vectors into the feature encoder to obtain intermediate variables of the multiple frames of audio vectors, that is, the feature encoder encodes the high-dimensional audio vectors into low-dimensional latent variables, so that the autoencoder records the feature distribution of the latent variables; after obtaining the intermediate variables, input the intermediate variables into the feature decoder to obtain the reconstructed target audio vectors for each frame of audio vectors, that is, the feature decoder restores the latent variables to the initial dimension according to the feature distribution pre-learned during training to obtain the reconstructed target audio vectors. Specifically, when training the autoencoder by inputting multi-dimensional audio vectors, specific values of a preset dimension can be recorded during training. The preset dimension can be the second dimension or the third dimension. During reconstruction, if the original value of the preset dimension in the multi-dimensional audio vectors input into the autoencoder is greater than the specific value, that is, the original value of the preset dimension is restored to the specific value after reconstruction. When the original value corresponding to the preset dimension in the multi-dimensional audio vectors input into the autoencoder is less than the specific value, that is, the original value of the preset dimension remains unchanged after reconstruction. In this way, the value of the reconstructed preset dimension is different from the specific value of the input preset dimension. When there are more preset dimensions with differences, the difference between the reconstructed multi-dimensional audio vectors and the input multi-dimensional audio vectors is larger, and when there are fewer preset dimensions with differences, the difference between the reconstructed multi-dimensional audio vectors and the input multi-dimensional audio vectors is smaller.
[0103] 305. Determine whether the audio data to be measured is abnormal sound data according to the relative situation between each frame of audio vector in the audio data to be measured and the corresponding target audio vector.
[0104] The audio detection device can determine whether the audio data to be measured is abnormal sound data according to the relative situation between each frame of audio vector in the audio data to be measured and the corresponding target audio vector. Specifically, the absolute value of the difference between each frame of audio vector and the corresponding target audio vector can be calculated to obtain multiple error values. For example, if the audio vector of the audio data to be measured input to the autoencoder at the t-th frame is x t , and the corresponding output target audio vector is The reconstruction error e t Can be calculated as: Average the multiple error values to obtain an average error value, and judge whether it is abnormal sound data according to the average error value; if the average error value is greater than the preset error threshold, determine that the audio data to be measured is abnormal sound data; if the average error value is less than the preset error threshold, determine that the audio data to be measured is normal sound data. The preset error threshold can be 0.4 or 0.5, and no specific limitation is made here.
[0105] Furthermore, if the reconstruction errors of the audio feature to be measured and the target audio feature are less than the preset threshold, that is, when the difference value between the audio feature of the audio data to be measured and the target audio feature is less than the preset threshold, in order to improve the accuracy of audio detection, the identification information of the audio data can be used to further determine whether the audio data to be measured is abnormal sound data, such as Figure 5 As shown, the specific steps are as follows:
[0106] 501. Obtain the identification information of the audio data to be measured and the second data reconstruction model.
[0107] The audio detection device can obtain the identification information of the audio data to be measured. This identification information can be the ID or label of the sender of the audio data to be measured. Specifically, it is not limited here. Preferably, this identification information can be the user ID that sends the audio data to be measured. The audio detection device can also obtain a second data reconstruction model trained based on the identification features of normal sound data. The identification features of normal sound data refer to the feature expressions of the identification information of normal sound data. Preferably, this identification feature is the identification vector of the identification information. When the identification feature input to the second data reconstruction model is the identification feature of normal sound data, the reconstruction error between the reconstructed identification feature of the second data reconstruction model and the input identification feature is less than a preset value. When the identification feature input to the second data reconstruction model is not the identification feature of normal sound data, the reconstruction error between the reconstructed identification feature of the second data reconstruction model and the input identification feature is greater than the preset value. Among them, the process of training the second data reconstruction model based on the identification features of normal sound data is similar to the above step 301, and will not be elaborated here specifically.
[0108] 502. Obtain the to-be-measured identification feature of the identification information of the audio data to be measured according to a preset rule.
[0109] The audio detection device can obtain the to-be-measured identification feature of the identification information of the audio data to be measured according to a preset rule. Specifically, the user ID can be converted into a unique identification vector (embedding) through an embedding layer or through an FM algorithm. It is not limited here specifically.
[0110] 503. Input the to-be-measured identification feature into the second data reconstruction model to obtain the reconstructed target identification feature.
[0111] After obtaining the to-be-measured identification feature of the identification information of the audio data to be measured, the audio detection device can input the to-be-measured identification feature into the second data reconstruction model to obtain the reconstructed target identification feature. Specifically, the identification vector can be input into an autoencoder, and the autoencoder outputs the reconstructed target identification vector. It can be understood that the above first data reconstruction model and this second data reconstruction model can be the same data reconstruction model or different data reconstruction models. It is not limited here specifically. When they are the same data reconstruction model, as Figure 6 shown, where Z refers to the intermediate variable encoded by the feature encoder. During training, the audio vector of the audio data and the identification vector of the identification information can be used to train the same autoencoder. At this time, the autoencoder has two sub-modules that respectively output the reconstructed target audio vector and the reconstructed target identification vector. Among them, the process of inputting the identification vector into the autoencoder and the autoencoder outputting the reconstructed target identification vector is similar to the above step 304, and will not be elaborated here specifically.
[0112] 504. Determine whether the audio data to be tested is abnormal sound data based on the relative situation between the identification information of the normal sound data and the identification information corresponding to the target identification feature.
[0113] The audio detection device can determine whether the audio data to be tested is abnormal sound data based on the relative situation between the identification information of the normal sound data and the identification information corresponding to the target identification feature. Specifically, the target identification feature can be converted into target identification information; it can be understood that the target identification vector can be converted into a user ID. When converting, if the value of a certain dimension in the target identification vector is the same as the value of the corresponding dimension of the identification vector of the normal sound data, the converted user ID is the same as the user ID of the normal sound data. If there is identification information in the identification information of the normal sound data that is the same as the target identification information, it is determined that the audio data to be tested is normal sound data; if there is no identification information in the identification information of the normal sound data that is the same as the target identification information, it is determined that the audio data to be tested is abnormal sound data; that is, when there is a match between the multiple user IDs of the normal sound data used for training and the converted user ID, it is determined that the audio data to be tested is normal sound data.
[0114] Further, using the identification information to further determine whether the audio data to be tested is abnormal sound data can also be determined through the following steps:
[0115] 505. Obtain the identification information of the audio data to be tested and the preset list of identification information.
[0116] The audio detection device can obtain the identification information of the audio data to be tested and the preset list of identification information; among them, the identification information of the normal sound data and the identification information of other audio data that conforms to the preset content rules. It can be understood that the identification information of other audio data that conforms to the preset content rules, that is, other audio data except the audio detection device and the audio player connected to the audio detection device. When it is determined that this other audio data is normal sound data, the identification information of this other audio data can be added to the list of identification information.
[0117] 506. Determine whether the audio data to be tested is abnormal sound data based on the relative situation between the identification information of the audio data to be tested and the preset list of identification information.
[0118] The audio detection device can determine whether the audio data to be tested is abnormal sound data based on the relative situation between the identification information of the audio data to be tested and the preset list of identification information. Specifically, if the identification information of the audio data to be tested exists in the preset list of identification information, it is determined that the audio data to be tested is normal sound data; if the identification information of the audio data to be tested does not exist in the preset list of identification information, it is determined that the audio data to be tested is abnormal sound data.
[0119] The embodiment of the present application further provides an audio detection device, as Figure 7 shown, including:
[0120] A first acquisition unit 701, configured to acquire audio data to be measured and a first data reconstruction model; the first data reconstruction model is trained based on normal sound data, and the normal sound data is pre-stored audio data that conforms to a preset content rule; when the audio data input to the first data reconstruction model is the normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is less than a preset value, and when the audio data input to the first data reconstruction model is not the normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is greater than the preset value;
[0121] A second acquisition unit 702, configured to acquire the audio features to be measured of the audio data to be measured according to a preset rule;
[0122] An input unit 703, configured to input the audio features to be measured into the first data reconstruction model to obtain the reconstructed target audio features;
[0123] A determination unit 704, configured to calculate the reconstruction error between the audio features to be measured and the target audio features, and determine whether the audio data to be measured is abnormal.
[0124] Further, as Figure 8 shown, the audio detection device can be divided into a storage module, an abnormal sound detection module, an audio acquisition module, a computer host, an audio data display module, and a reminder module; the storage module can store normal sound data and identification information of the normal sound data, and the abnormal sound detection module includes Figure 8 shown multiple units. The computer host can control the audio acquisition module to acquire the audio data to be measured of the audio player at a preset sampling rate, and transmit the audio data to be measured to the abnormal sound detection module. The audio data display module can display the difference value between the audio features to be measured and the target audio features of the audio data to be measured. When the difference value is greater than a preset threshold, that is, when the reconstruction error is large, the computer host controls the reminder module to issue an abnormal reminder for manual review.
[0125] The embodiment of the present application further provides an audio detection device 900, as Figure 9 shown, including:
[0126] A central processing unit 901, a memory 902, and an input / output interface 903;
[0127] The memory 902 is a transient storage memory or a persistent storage memory;
[0128] The central processing unit 901 is configured to communicate with the memory 902 and execute instruction operations in the memory 902 to perform the above audio detection method.
[0129] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0130] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0131] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0132] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0133] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.
Claims
1. An audio detection method, characterized in that Including: Obtaining audio data to be measured and a first data reconstruction model; the first data reconstruction model is trained based on the audio features of normal sound data, and the normal sound data is pre-stored audio data that conforms to a preset content rule; when the audio data input to the first data reconstruction model is the normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is less than a preset value, and when the audio data input to the first data reconstruction model is not the normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is greater than the preset value; Obtaining the audio features to be measured of the audio data to be measured according to a preset rule; Inputting the audio features to be measured into the first data reconstruction model to obtain the reconstructed target audio features; Calculating the reconstruction error between the audio features to be measured and the target audio features to determine whether the audio data to be measured is abnormal; The obtaining of the first data reconstruction model includes: Obtaining an initial data reconstruction model and a training set of audio features of first normal sound data; Inputting the training set of audio features of the first normal sound data into the initial data reconstruction model to obtain the output audio features; Determining a loss value based on the output audio features and the training set of audio features of the first normal sound data; Determining whether the initial data reconstruction model converges based on the loss value; If it does not converge, obtaining a training set of audio features of second normal sound data and inputting the training set of audio features of the second normal sound data into the initial data reconstruction model; If it converges, the training is completed to obtain the first data reconstruction model.
2. The audio detection method according to claim 1, wherein The audio features to be measured include: multiple frames of audio vectors; the first data reconstruction model includes: a feature encoder and a feature decoder; The inputting the audio features to be measured into the first data reconstruction model to obtain the reconstructed target audio features includes: Inputting multiple frames of the audio vectors into the feature encoder to obtain intermediate variables of the multiple frames of the audio vectors; Inputting the intermediate variables into the feature decoder to obtain the reconstructed target audio vectors for each frame of the audio vectors.
3. The audio detection method according to claim 2, wherein The calculating the reconstruction error between the audio features to be measured and the target audio features to determine whether the audio data to be measured is abnormal includes: Calculating the absolute value of the difference between each frame of the audio vectors and the corresponding target audio vectors to obtain multiple error values; Taking the average of the multiple error values to obtain an average error value, and using the average error value as the reconstruction error between the audio features to be measured and the target audio features; If the average error value is greater than a preset error threshold, determining that the audio data to be measured is abnormal sound data; If the average error value is less than the preset error threshold, determining that the audio data to be measured is normal sound data.
4. The audio detection method according to claim 1, wherein After obtaining the audio features to be measured of the audio data to be measured according to a preset rule and before inputting the audio features to be measured into the first data reconstruction model to obtain the reconstructed target audio features, the method further includes: Perform data augmentation processing on the audio features to be measured of the audio data to be measured, and obtain the enhanced audio features; The step of inputting the audio features to be measured into the first data reconstruction model to obtain the reconstructed target audio features includes: Input the enhanced audio features into the first data reconstruction model to obtain the reconstructed target audio features.
5. The audio detection method according to claim 1, characterized in that, If the reconstruction error between the audio features to be measured and the target audio features is less than a preset threshold, the method further includes: Obtain the identification information of the audio data to be measured and a second data reconstruction model, where the second data reconstruction model is trained based on the identification features of the normal sound data; when the identification feature input into the second data reconstruction model is the identification feature of the normal sound data, the reconstruction error between the reconstructed identification feature of the second data reconstruction model and the input identification feature is less than a preset value, and when the identification feature input into the second data reconstruction model is not the identification feature of the normal sound data, the reconstruction error between the reconstructed identification feature of the second data reconstruction model and the input identification feature is greater than a preset value; Obtain the to-be-measured identification feature of the identification information of the audio data to be measured according to a preset rule; Input the to-be-measured identification feature into the second data reconstruction model to obtain the reconstructed target identification feature; Determine whether the audio data to be measured is abnormal sound data according to the relative situation between the identification information of the normal sound data and the identification information corresponding to the target identification feature.
6. The audio detection method according to claim 5, characterized in that, The step of determining whether the audio data to be measured is abnormal sound data according to the relative situation between the identification information of the normal sound data and the identification information corresponding to the target identification feature includes: Convert the target identification feature into target identification information; If there is identification information in the identification information of the normal sound data that is the same as the target identification information, determine that the audio data to be measured is normal sound data; If there is no identification information in the identification information of the normal sound data that is the same as the target identification information, determine that the audio data to be measured is abnormal sound data.
7. The audio detection method according to claim 1, wherein If the reconstruction error between the audio features to be measured and the target audio features is less than a preset threshold, the method further includes: Obtain the identification information of the audio data to be measured and a preset list of identification information, where the list of identification information includes: the identification information of the normal sound data and the identification information of other audio data that conforms to the preset content rules; If the identification information of the audio data to be measured exists in the preset list of identification information, determine that the audio data to be measured is normal sound data; If the identification information of the audio data to be measured does not exist in the preset list of identification information, determine that the audio data to be measured is abnormal sound data.
8. An audio detection device, characterized in that, Includes: A first acquisition unit, configured to acquire audio data to be measured and a first data reconstruction model; the first data reconstruction model is trained based on audio of normal sound data, and the normal sound data is pre-stored audio data conforming to a preset content rule; when the audio data input to the first data reconstruction model is the normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is less than a preset value, and when the audio data input to the first data reconstruction model is not the normal sound data, the reconstruction error between the audio features reconstructed by the first data reconstruction model and the input audio features is greater than the preset value; A second acquisition unit, configured to acquire measured audio features of the audio data to be measured according to a preset rule; An input unit, configured to input the measured audio features into the first data reconstruction model to obtain reconstructed target audio features; A determination unit, configured to calculate the reconstruction error between the measured audio features and the target audio features, and determine whether the audio data to be measured is abnormal; Specifically, the first acquisition unit is configured to acquire an initial data reconstruction model and an audio feature training set of first normal sound data; input the audio feature training set of the first normal sound data into the initial data reconstruction model to obtain output audio features; Determine a loss value based on the output audio features and the audio feature training set of the first normal sound data; Based on the loss value, determine whether the initial data reconstruction model converges; if not, acquire an audio feature training set of second normal sound data, and input the audio feature training set of the second normal sound data into the initial data reconstruction model; if it converges, the training is completed to obtain the first data reconstruction model.
9. An audio detection device, characterized in that, Comprising: A central processing unit, a memory, and an input / output interface; The memory is a transient storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Including instructions, when the instructions run on a computer, causing the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Abnormal sound detection method and abnormal sound detection device
CN111354366A
Sound anomaly detection method and device, computer equipment and storage medium
CN113470695A