Sound event detection method and device and computer readable storage medium
By converting spectral features into universal audio features without physical meaning on the end-side device and using the GRU network for sound event detection, the technical problem of the limited number of sound event detection model categories in the existing technology is solved, and the detection of multiple sound events on the end-side device is realized.
Patent Information
- Application Number
- CN202510879443.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-03
AI Technical Summary
Existing sound event detection models are limited by the computing resources of end-side devices, making it difficult to effectively expand the number of sound event detection categories while maintaining computing power requirements. This results in a decrease in detection accuracy and recall rate, and is unable to meet the needs of complex applications.
An audio feature extraction model is used to convert spectral features into a general audio feature space without physical meaning, and reasoning analysis is performed through a multi-sound event classification model. The GRU network is used to capture time-dependent features and long short-term memory network technologies are used, combined with general audio features for sound event detection, to achieve detection of multiple sound events.
With the limited computing resources of the end-side devices, it can effectively increase the types of detectable sound events, maintain high accuracy and balance for each sound event, and meet actual application needs.
Smart Images

Figure CN120748439A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of sound detection technology, and in particular to a sound event detection method, a sound event detection device, and a computer-readable storage medium. Background Art
[0002] With the development of artificial intelligence (AI) technology, AI algorithms have been widely used in the field of intelligent surveillance. By deploying AI algorithms in edge devices such as IP cameras, various functions such as pedestrian detection and tracking, face detection and recognition, voice communication and noise reduction, echo cancellation, and sound event detection (SED) are realized. SED, in particular, can identify specific sound events in a scene. For example, in home surveillance, cameras can monitor the ambient audio in real time and send timely notifications to the homeowner upon detecting a preset sound event, such as an unusual sound.
[0003] However, due to the limited computing resources of end-devices, existing sound event detection models can only detect one or two sound events. Therefore, effectively expanding the number of sound event detection categories while maintaining affordable computing power on end-devices has become a pressing issue. Summary of the Invention
[0004] In view of the above problems, the embodiments of the present application provide a sound event detection method, a sound event detection device and a computer-readable storage medium, which are used to solve the problem that most sound event detection models in the prior art can only detect one or two sound events.
[0005] According to one aspect of an embodiment of the present application, a sound event detection method is provided, the method comprising: obtaining spectral features of audio to be detected, wherein the dimensions of the spectral features include the number of frequency points and the number of consecutive frames; inputting the spectral features into an audio feature extraction model to obtain universal audio features, wherein the universal audio features are feature vectors for indicating audio information; inputting the universal audio features into a multi-sound event classification model to obtain a first predicted probability of each sound event in a preset plurality of categories of sound events; and determining the presence of each sound event in the audio to be detected based on the first predicted probability.
[0006] In an optional manner, the audio feature extraction model includes at least one convolutional neural network block and at least one fully connected block cascaded in sequence, the spectral feature is a Mel spectrum feature, and the number of frequency points corresponds to the number of Mel filters.
[0007] In an optional manner, after inputting the spectral features into the audio feature extraction model to obtain the general audio features, the method further includes: inputting the general audio features continuously output by the audio feature extraction model into a buffer in sequence, wherein the buffer capacity is L features, L is a positive integer greater than 1, and when the current moment is time t and t>L-1, the buffer stores the general audio features at time t and the previous L-1 moments.
[0008] In an optional embodiment, the preset multiple categories of sound events include a first sound event, and the method further includes: when it is determined that the first sound event exists in the audio to be detected, inputting the spectral characteristics of the audio to be detected into a first sound event classification model corresponding to the first sound event to obtain a second predicted probability of the first sound event; and verifying the presence of the first sound event in the audio to be detected based on the second predicted probability.
[0009] In an optional manner, the inputting of the general audio features into a multi-sound event classification model to obtain a first predicted probability of each sound event in a preset multi-category sound event includes: inputting all the general audio features in the buffer at the current moment into the multi-sound event classification model to obtain a first predicted probability of each sound event in a preset multi-category sound event at the current moment.
[0010] In an optional manner, determining the existence of each sound event in the audio to be detected based on the first predicted probability includes: performing a first judgment on the first predicted probability at L consecutive moments to obtain L first judgment results, wherein the L moments include moment t and L-1 moments thereafter, and the first judgment result is whether each sound event exists or does not exist in the audio to be detected at each moment; based on the L first judgment results, determining the existence of each sound event in the audio to be detected at moment t.
[0011] In an optional manner, the preset multiple categories of sound events include a first sound event, and the first judgment result for the first sound event is whether the first sound event exists or not in the audio to be detected; determining the existence of each sound event in the audio to be detected at time t based on the L first judgment results includes: judging whether the number of target judgment results in the L first judgment results for the first sound event reaches a preset threshold, wherein the target judgment result is that the first sound event exists in the audio to be detected; if the number of target judgment results in the L first judgment results for the first sound event reaches a preset threshold, it is determined that the first sound event exists in the audio to be detected at time t.
[0012] In an optional manner, the parameters of the audio feature extraction model are obtained as follows: loading the parameters of a pre-trained audio feature extraction model and randomly initializing the parameters of a multi-sound event classification model to be trained; inputting the spectral features of the audio sample data into the pre-trained audio feature extraction model, and inputting the output results of the pre-trained audio feature extraction model into the multi-sound event classification model to be trained, wherein the audio sample data corresponds to the preset multiple categories of sound events; training the multi-sound event classification model to be trained by gradient back propagation, and fine-tuning the parameters of the pre-trained audio feature extraction model.
[0013] According to another aspect of an embodiment of the present application, a sound event detection device is provided, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the sound event detection method of any of the above embodiments.
[0014] According to another aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the sound event detection method of any of the above embodiments is implemented.
[0015] The embodiment of the present application first uses an audio feature extraction model to convert the spectral features in the audio data into a non-physical feature space to obtain a universal audio feature. The universal audio features are then inferred and analyzed through a multi-sound event classification model to predict the probability of occurrence of different sound events, and based on this, the existence of each sound event is judged. Since the input to the multi-sound event classification model is universal audio features that have no physical meaning, these features have a balanced characterization capability for a variety of sound events. Therefore, when detecting a variety of sound events, it is possible to maintain a high and balanced detection accuracy for each sound event, meeting the needs of practical applications. In addition, the design of the audio feature extraction model and the multi-sound event classification model does not need to be too complicated to ensure that the types of detectable sound events are effectively increased under the limited computing resources of the end-side device.
[0016] The above description is only an overview of the technical solution of the embodiment of the present application. In order to more clearly understand the technical means of the embodiment of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiment of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are only used to illustrate the embodiments and are not to be considered as limiting the present application. In addition, the same reference symbols are used to represent the same components throughout the drawings. In the drawings:
[0018] Figure 1A schematic diagram showing an application scenario of the sound event detection method provided in an embodiment of the present application is shown;
[0019] Figure 2 A schematic structural diagram of a camera device provided in an embodiment of the present application is shown;
[0020] Figure 3 A schematic structural diagram of a camera device provided in another embodiment of the present application is shown;
[0021] Figure 4 A schematic diagram of a parameter fine-tuning process of an audio feature extraction model provided in an embodiment of the present application is shown;
[0022] Figure 5 A flow chart of a sound event detection method according to an embodiment of the present application is shown;
[0023] Figure 6 A schematic flow chart showing a sound event detection method according to another embodiment of the present application is shown;
[0024] Figure 7 The test results of the sound event detection method according to the embodiment of the present application are shown. DETAILED DESCRIPTION
[0025] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited to the embodiments set forth herein.
[0026] With the development of AI technology, AI algorithms have been widely used in the field of intelligent monitoring. AI algorithms can be deployed not only on cloud devices such as servers, but also on edge devices. Edge devices refer to various computing devices located at edge nodes, such as smartphones, tablets, smart watches, smart speakers, smart homes, routers, and cameras. Computing and resource storage on edge devices make data processing and transmission more efficient and faster.
[0027] Deploying AI algorithms in cameras enables a variety of functions, including pedestrian detection and following, face detection and recognition, voice communication and noise reduction, echo cancellation, and sound event detection (SED). SED, in particular, identifies specific sound events within a scene. For example, in home surveillance, cameras can monitor ambient audio in real time and, upon detecting a pre-defined sound event, such as an unusual sound, promptly notify the homeowner.
[0028] However, due to the limitations of computing resources on the end-side devices, existing sound event detection models are often designed to be relatively simple. A common technique is for the end-side device to extract audio signal features, calculate the Log Mel Spectrum (LMS) and input it into a Convolutional Neural Network (CNN), which then outputs the probability of each sound event to determine the occurrence of the event. This simple model has low computing power requirements, but its representational capabilities are limited when processing multiple sound events, resulting in a decrease in the harmonic mean of precision and recall (F1 Score), an increase in missed alerts and false alarms, and difficulty meeting the needs of complex applications. Currently, most products on the market can only detect one or two sound events.
[0029] Faced with the challenge of expanding the number of sound event detection categories, increasing AI model complexity increases model size and computing power requirements, further increasing the chip load on edge devices, where computing resources are already limited. Therefore, how to effectively expand the number of sound event detection categories while maintaining affordable computing power on edge devices has become a pressing issue.
[0030] To solve this problem, an embodiment of the present application proposes a sound event detection method. By extracting general audio features (GAF) that have no physical meaning, sound event detection is performed based on the general audio features, thereby effectively increasing the types of detectable sound events with limited computing resources of the terminal device.
[0031] Figure 1 The following is a schematic diagram illustrating an application scenario of the sound event detection method provided in an embodiment of the present application. In this embodiment, the sound event detection method is applied to a camera device 1. In other embodiments, the sound event detection method provided in an embodiment of the present application can also be applied to other devices, such as devices with audio acquisition capabilities. Of course, the device may also not have an audio acquisition function and may acquire audio transmitted by other devices and then perform sound event detection on the audio.
[0032] Please continue reading Figure 1 , camera device 1 establishes a communication connection with electronic device 2 via network 3. Camera device 1 may be a security surveillance camera, an IP camera, or other video surveillance equipment. Electronic device 2 may be a touchscreen phone, a smartphone, a tablet computer, a portable electronic device, or other electronic device, also commonly referred to as a client. Network 3 includes, but is not limited to, one or more of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a 4G / 5G network, Wi-Fi, Bluetooth, and a peer-to-peer (P2P) communication network.
[0033] The camera device 1 and the electronic device 2 may each include one or more processors, which may be a microcontroller unit (MCU), a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the present embodiment, without limitation herein. The one or more processors may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs, without limitation herein.
[0034] The camera device 1 is installed in a monitoring area (such as a home, office, shopping mall, etc.), captures the monitoring video in the monitoring area and sends the captured video to the electronic device 2 for the user to browse via the network 3. The monitoring video captured by the camera device 1 may include audio data.
[0035] Figure 2 FIG. 1 shows a schematic structural diagram of the camera device 1 provided in an embodiment of the present application. Figure 2 As shown, the camera device 1 may include a processor 11, a memory 12, and a computer program 13 stored in the memory 12. The processor 11 executes the computer program 13 to implement the sound event detection method of the embodiment of the present application.
[0036] In some embodiments, the processor 11 is an MCU, and the camera device 1 further includes other functional components. Figure 3 FIG. 1 shows a structural diagram of a camera device 1 provided in another embodiment of the present application. Figure 3 As shown, the camera device 1 may include an MCU 21 , an audio acquisition unit 22 , an encoding chip 23 and a memory 12 .
[0037] The MCU 21 is used to execute computer programs, specifically, to perform the relevant steps of the sound event detection method embodiments provided herein. The MCU 21 has low power consumption, and its highly integrated design reduces the need for external components, thereby reducing the cost and size of the camera device 1. In other embodiments, other control units such as a CPU and an ASIC may be used to implement the functions of the MCU 21.
[0038] The audio acquisition unit 22 can be a microphone or microphone array, which is used to capture sound waves in the monitored area and convert them into audio electrical signals. The camera device 1 can also include an image acquisition unit, which is composed of an image sensor (such as a CMOS or CCD sensor) and related signal processing circuits. The image acquisition unit captures light in the monitored area and converts it into video electrical signals.
[0039] The encoding chip 23 is used to convert the audio electrical signal captured by the audio acquisition unit 22 into a digital signal and compress and encode it to obtain an audio code stream. The encoding chip 23 also converts the video electrical signal captured by the image acquisition unit into a digital signal and compress and encode it to obtain a video code stream.
[0040] Memory 12 is used to store audio and video streams. Memory 12 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk drive. RAM may use double data rate synchronous dynamic random access memory (DDR SDRAM). The audio and video streams can also be transmitted to a larger storage device such as a network video recorder (NVR) or a cloud server for storage and user viewing.
[0041] The camera device 1 may also include other units, such as a sensing unit, a communication unit, etc., which are not listed here one by one. The sensing unit is used to detect targets in the monitoring area, such as cars, people, animals, etc. The sensing unit can be a passive infrared (PIR) sensor, a motion detection (MD) module, a human detection module, etc. Taking the PIR sensor as an example, if someone enters the monitoring area, the sensing unit can sense the infrared heat source radiated by the human body. The communication unit is used to wirelessly transmit audio and video data to the external electronic device 2 or the cloud server via the network 3.
[0042] The camera device 1 deploys an audio feature extraction model and a multi-sound event classification model. Using these two models, sound event detection is performed in two stages. In the first stage, the audio feature extraction model converts the spectral features in the audio data into a non-physical feature space to obtain universal audio features. In the second stage, the multi-sound event classification model performs inference analysis on the universal audio features to predict the probability of occurrence of different sound events.
[0043] The audio feature extraction model deployed in the camera device 1 can adopt a pre-trained model (such as a convolutional neural network pre-trained on a large audio dataset) without the need for further training.
[0044] The audio feature extraction model may include at least one convolutional neural network block (CNN block) and at least one fully connected block, which are cascaded in sequence. The CNN block is used to extract local features, and the fully connected block is used to integrate the local features extracted by the CNN block into global features and map them from a high-dimensional feature space to a low-dimensional feature space.
[0045] Taking the audio feature extraction model including three CNN blocks and two fully connected blocks as an example, Table 1 is the architecture of the audio feature extraction model of this embodiment.
[0046] Table 1 Audio feature extraction model architecture
[0047]
[0048]
[0049] As shown in Table 1, blocks 1, 2, and 3 have the same structure. Convolutional layers (Conv2d) are used to extract local features. Batch normalization layers (BatchNorm2d) standardize the data before activation to accelerate training and improve model stability. ReLU activation functions are used to introduce nonlinearity and alleviate the vanishing gradient problem. Maxpooling layers (MaxPool2d) are used to reduce the size of feature maps while preserving important features. Tensor resize (Permute) in block 4 changes the order of tensor dimensions to facilitate processing by the subsequent fully connected layers. In block 5, local features are integrated using a fully connected layer (Linear) and an activation function (ReLU) to obtain a global feature representation, which is then mapped to the final output space. This audio feature extraction model progressively reduces the dimensionality of the audio data through multiple convolutional layers, batch normalization layers, activation functions, and maxpooling layers, while extracting more abstract, higher-level features. Finally, general audio features are obtained through fully connected layers.
[0050] In Table 1, regarding (number of input channels, number of output channels), taking (1,64) of the first Conv2d layer in block 1 as an example, (1,64) means that the number of feature channels input to the first Conv2d layer is 1. After processing by the first Conv2d layer, the number of feature channels output increases to 64. "-" in (number of input channels, number of output channels) means that the number of output channels does not change compared to the number of input channels.
[0051] Regarding (kernel size, padding size, stride), using the example of (3,1,1) for the first Conv2d layer in block 1, (3,1,1) indicates a 3×3 kernel size for the first Conv2d convolution operation, a padding value of 1 at the outer edges of the input features, and a sliding row of 1 at each convolution operation. A value of "-" for (kernel size, padding size, stride) indicates no convolution operation is performed on that layer. A padding value of "-" indicates no padding is performed on that layer.
[0052] It should be noted that the universal audio feature has no physical meaning. It is an abstract feature generated by a neural network. Unlike the amplitude spectrum or spectrogram commonly seen in audio, the universal audio feature itself has no special physical meaning. It is a highly abstract vector used to represent universal audio features. In other words, the universal audio feature has a balanced representation capability for a variety of sound events. It does not tend to represent voiced features such as human voices, or unvoiced features, nor does it tend to represent natural white noise, etc. It is equivalent to obtaining the mean value that represents all sound events. Therefore, by detecting a variety of sound events through universal audio features, it is possible to maintain a high and balanced detection accuracy for each sound event, meeting the needs of practical applications.
[0053] The multi-sound event classification model deployed in the camera device 1 can be a model obtained by training a neural network. The trained neural network can be a gated recurrent unit (GRU), a recurrent neural network (RNN), a long-short term memory network (LSTM), and other networks. The above networks can capture the time series features and long-term dependencies in the audio signal, thereby effectively identifying the dynamic changes of the voice. A multi-sound event classification model means that the model can be used to identify multiple sound events, rather than just a single type of sound event. Table 2 is the architecture of a multi-sound event classification model.
[0054] Table 2 Multi-sound event classification model architecture
[0055] Layer Index Layer Name Input / Output_Unit 1 GRU1 D / 2 2 GRU2 D / 2 3 Linear (D / 2,C) 4 Sigmoid -
[0056] As shown in Table 2, both the first and second layers are GRUs. The GRU is a variant of the RNN that addresses the vanishing and exploding gradient problems that standard RNNs often encounter when dealing with long-term dependencies. It introduces a gating mechanism to control the flow of information, enabling the network to better capture and utilize long-term dependencies. The GRU primarily consists of an update gate, a reset gate, candidate hidden states, and hidden state updates. The update gate determines when the hidden state should be updated, controlling the balance between new input information and the previous hidden state. Its output is a value between 0 and 1 that weights the previous hidden state. The reset gate determines when the hidden state should be reset, controlling the influence of the previous hidden state. Its output is also a value between 0 and 1 that weights the previous hidden state. The candidate hidden state combines the current input with the previous hidden state weighted by the reset gate to generate a new candidate hidden state. The hidden state update combines the update gate and the candidate hidden state to generate a new hidden state, which is the weighted sum of the previous and candidate hidden states. The multi-sound event classification model architecture constructed using the GRU features a simple structure, few parameters, and high computational efficiency. Moreover, GRU can effectively handle long-term dependency problems, avoid gradient vanishing and gradient exploding, and is superior to LSTM in the sound event detection task in the embodiment of the present application.
[0057] The multi-sound event classification model shown in Table 2 uses the first GRU layer to capture the temporal dependencies of the input's general audio features, extracting useful information. The second GRU layer then continues to extract and abstract features. Stacking two GRU layers enhances the model's ability to learn complex patterns. The linear layer maps the output of GRU2 to a new dimensional space, preparing for classification, where C represents the number of predicted sound event types. Finally, the sigmoid activation function converts the output of the linear layer into a probability value (ranging between 0 and 1) to generate a likelihood score for each category.
[0058] During the training of the neural network with the architecture shown in Table 2, the weights and biases of the network are updated through the back-propagation algorithm to minimize the loss function, thereby obtaining a trained multi-sound event classification model.
[0059] When training a multiple sound event classification model, the parameters of a pre-trained audio feature extraction model can be directly loaded, and the general audio features output by the audio feature extraction model can be used to train the multiple sound event classification model.
[0060] The audio feature extraction model can use a pre-trained model, or it can be fine-tuned on the basis of the pre-trained model according to the sound event data related to the type of sound event to be detected by the sound event detection method of the embodiment of the present application. Figure 4 FIG. 4 shows a flow chart of parameter fine-tuning of the audio feature extraction model provided in an embodiment of the present application. Figure 4 As shown, in some embodiments, the parameters of the audio feature extraction model are obtained by the following steps:
[0061] S110 , loading parameters of a pre-trained audio feature extraction model, and randomly initializing parameters of a multi-sound event classification model to be trained.
[0062] S120: Inputting the spectrum features of the audio sample data into a pre-trained audio feature extraction model, and inputting the output result of the pre-trained audio feature extraction model into a multiple sound event classification model to be trained.
[0063] The audio sample data corresponds to multiple preset sound event types. For example, the sound event detection method of the present embodiment needs to detect sound events including a baby crying, a cat meowing, a dog barking, a car horn, an alarm, and glass breaking. The audio sample data used to train the multi-sound event classification model includes the aforementioned six sound event types.
[0064] Spectral features can be mel spectrum features, amplitude spectrum features, phase spectrum features, and power spectrum features. The dimensions of spectral features include the number of frequency bins and the number of consecutive frames, that is, the spectral features are two-dimensional time-frequency matrices. Each row of this matrix represents a frequency bin, each column represents a time frame, and each element represents the signal energy at a specific time frame and a specific frequency bin. If the spectral features are mel spectrum features, amplitude spectrum features, or power spectrum features, then each element in the two-dimensional time-frequency matrix represents an energy value. For example, the amplitude spectrum represents the energy value in the short-time Fourier transform (STFT) domain, while the mel spectrum represents the energy value in the mel domain. If the spectral features are phase spectrum features, then each element in the two-dimensional time-frequency matrix represents the phase offset of the frequency component in the time domain.
[0065] S130 , training the multi-sound event classification model to be trained through gradient back propagation, and fine-tuning the parameters of the pre-trained audio feature extraction model.
[0066] When training the multi-sound event classification model using the six types of sound event data described above, the gradient of the pre-trained audio feature extraction model is also enabled. Backpropagation of the gradient not only trains the multi-sound event classification model but also fine-tunes and updates the parameters of the pre-trained audio feature extraction model. The loss function used for training the multi-sound event classification model can be binary cross entropy (BCE) loss.
[0067] Through the above method, on the basis of using the pre-trained audio feature extraction model, the parameters of the audio feature extraction model are fine-tuned in combination with the sample data corresponding to the type of sound event that actually needs to be detected, so that the audio feature extraction model finally applied is more suitable for the sound event that needs to be detected.
[0068] As mentioned above, since universal audio features have no physical meaning, they have balanced representation capabilities for multiple sound events and are not biased towards any one sound event. Even if the parameters of the audio feature extraction model are fine-tuned in the above manner, it is equivalent to making a small offset based on the mean of all sound events extracted by the pre-trained audio feature extraction model, and will not be too far from the mean. On the basis of ensuring effective detection of multiple sound event types involved in the sample data, the detection accuracy of these sound events is improved.
[0069] It is understandable that the above training process can be carried out on a server with strong computing power, and multiple sound event classification models can be trained. After completing the parameter fine-tuning of the audio feature extraction model, the audio feature extraction model and the multiple sound event classification model can be deployed to the device that needs to perform sound event detection, such as the aforementioned camera device 1.
[0070] After the camera device 1 deploys the audio feature extraction model and the multi-sound event classification model, it can perform sound event detection on the surrounding environment. Figure 5 FIG. 1 shows a flow chart of a sound event detection method according to an embodiment of the present application, which can be applied to the processor 11 or the MCU 21 in the camera device 1. Figure 5 As shown, the method includes the following steps:
[0071] S200: Acquire frequency spectrum features of the audio to be detected.
[0072] The camera 1 collects audio signals from the monitored area through the audio acquisition unit 22 and extracts the spectral characteristics of the audio signals. The spectral characteristics include the number of frequency points. Because audio signals have temporal characteristics, the spectral characteristics also include the number of consecutive frames.
[0073] In a possible implementation, the spectral feature is a Mel spectrum feature, which is extracted from the audio signal collected by the camera 1. First, the time domain audio signal is defined using formula (1):
[0074] X=(x1,x2,...,x n ),n∈Z + (1)
[0075] Wherein, X represents the time domain audio signal, and n represents the sampling point number.
[0076] Perform a short-time Fourier transform (STFT) on the time-domain audio signal, and then calculate the amplitude spectrum M(t,f) using equation (2):
[0077] M(t,f)=|STFT(X)| (2)
[0078] Where t and f represent the frame and frequency indexes respectively.
[0079] Add a Mel-Filter Bank (MFB) to the amplitude spectrum and take its logarithm to obtain the Mel-spectrum feature LMS shown in formula (3):
[0080] LMS=10·log10(MFB(M(t,f))) (3)
[0081] Here, MFB(·) represents a Mel filter bank, a digital signal processing algorithm used to convert the spectrum of an audio signal from the Hertz (Hz) frequency domain to the Mel frequency domain. It simulates the frequency perception characteristics of the human auditory system through a series of triangular filters. These filters are calculated using mathematical formulas and implemented on a computer or embedded system. The Mel spectrum feature LMS is a two-dimensional time-frequency matrix, where each element is a logarithmic energy value.
[0082] Of course, the spectrum characteristics may also be amplitude spectrum characteristics, phase spectrum characteristics, power spectrum characteristics, etc., which are not listed one by one in this application.
[0083] S220: Input the spectrum features into an audio feature extraction model to obtain universal audio features.
[0084] In this step, the spectral features are mapped to a non-physical feature space through an audio feature extraction model to obtain a feature vector indicating audio information, which is also a universal audio feature.
[0085] It is important to understand that when directly using spectral features that characterize the physical characteristics of audio to predict the probability of sound events, although the prediction accuracy of sound events with significant specific physical characteristics is high, this leads to a decrease in the prediction accuracy of other sound events that do not have the significance of the physical characteristics. Due to this imbalance in prediction accuracy, it is impossible to meet the diverse needs of sound event detection. For example, the logarithmic Mel spectrum feature is a set of features constructed based on the nonlinear perception of frequency by the human ear. It mainly emphasizes the low-frequency part and compresses the high-frequency part. It has wide applicability for human voice detection. However, if it is used for the detection of multiple sound events at the same time, the detection accuracy of sound events other than human voice is poor, and there are false positives or missed positives.
[0086] Unlike the implementation method of directly inputting Mel-spectrogram features into the classification network for sound event classification, the embodiment of the present application only inputs Mel-spectrogram features as initial features into the first-stage audio feature extraction model (GAF), and does not directly input them into the multiple sound event classification model. The audio feature extraction model extracts general audio features from the Mel-spectrogram features. General audio features are abstract features generated by neural networks. Unlike the amplitude spectra or spectrograms commonly used in the field of audio signal processing, general audio features have no physical meaning and are highly abstract vectors. Their expression is:
[0087]
[0088] Where f1(·) is the mapping expression of the audio feature extraction model, and σ lms are the global mean and standard deviation of the Mel spectrum features. The mel spectrum features are globally normalized so that each mel spectrum feature has zero mean and unit variance to eliminate the dimension effect, make the feature distribution uniform, reduce the model's sensitivity to outliers, and improve the model convergence speed and model performance.
[0089] The dimensions of the mel spectrum features input to the audio feature extraction model include (n_mels, n_frames), where n_mels is the number of mel filters and n_frames is the number of consecutive frames. After the audio feature extraction model extracts the features, the input feature dimension changes from (n_mels, n_frames) to (D,), where D is the output feature dimension of the audio feature extraction model, that is, the process of converting the audio feature from a two-dimensional time-frequency matrix to a one-dimensional feature vector. In the embodiment of the present application, n_mels = 64, frames = 96, D = 512, and the sampling rate fs = 16000 (i.e., 16kHz).
[0090] S240: Input the general audio features into a multi-sound event classification model to obtain a first predicted probability of each sound event in the preset multiple categories of sound events.
[0091] The multi-sound event classification model is defined as f2(·), and the probability value vector P output by the multi-sound event classification model is:
[0092] P=f2(GAF)={p i |p i ≥0 and p i ≤1,i=1,2,...,C} (5)
[0093] Where C is the number of sound event types to be detected, p i is the probability of the i-th sound event. If there are six types of sound events to be detected, then for the audio signal to be detected, the multi-sound event classification model outputs the probability value of each type of sound event, a total of six probability values.
[0094] S260: Determine the existence of each sound event in the audio to be detected based on the first prediction probability.
[0095] According to the first predicted probability p i , the prediction result of the i-th sound event in the audio signal to be detected for:
[0096]
[0097] where θ i ∈Θ,Θ={θ i |θ i >0.5andθ i <1, i=1, 2, ..., C} is the probability threshold for determining whether the i-th sound event exists, and Θ is the set of all C thresholds for detection events, which is user-selectable and adjustable. For example, electronic device 2, which is communicatively connected to camera device 1, provides a graphical user interface through which the user can select the type of sound event to detect and adjust the threshold for the event to be detected in Θ. Setting the set of thresholds for detection events to be user-adjustable can effectively tailor to the user's personalized preferences.
[0098] like If is 1, it means that there is the i-th sound event in the audio signal to be detected; if If it is 0, it means that there is no i-th sound event in the audio signal to be detected.
[0099] The sound event detection method provided in the embodiment of the present application first uses an audio feature extraction model to convert the spectral features in the audio data into a non-physical feature space to obtain a universal audio feature. Then, the universal audio features are inferred and analyzed through a multi-sound event classification model to predict the probability of occurrence of different sound events, and the existence of each sound event is judged accordingly. Since the universal audio features that are input to the multi-sound event classification model have no physical meaning, these features have a balanced characterization capability for a variety of sound events, and do not tend to characterize voiced features such as human voices, or unvoiced features, nor do they tend to characterize natural white noise, etc., which is equivalent to obtaining the mean value that characterizes all sound events. Therefore, when detecting multiple sound events, it is possible to maintain a high and balanced detection accuracy for each sound event, meeting the needs of practical applications. In addition, the design of the audio feature extraction model and the multi-sound event classification model does not need to be too complicated, ensuring that the types of detectable sound events are effectively increased under the limited computing resources of the end-side device.
[0100] In some embodiments, after step S220, the following steps are further included:
[0101] The general audio features continuously output by the audio feature extraction model are input into the buffer in sequence.
[0102] The buffer capacity is L features, where L is a positive integer greater than 1. The buffer uses a first-in, first-out (FIFO) mechanism. That is, when the current time is t and t>L-1, the buffer stores the common audio features from time t and the previous L-1 times. Storing common audio features in the buffer smoothes the data flow and avoids data loss caused by inconsistencies between the audio acquisition rate and the processing rate.
[0103] The input features of the audio feature extraction model are (n_mels, n_frames), and the output features are (D,), that is, the two-dimensional Mel spectrum features are changed into one-dimensional general audio features. If the one-dimensional features are directly input into the multi-sound event classification model, it will be difficult to obtain the temporal characteristics of the audio. In the embodiment of the present application, the sampling rate is 16kHz, and n_frames=96 means that the Mel spectrum features will be input into the audio feature extraction model after nearly 1 second. The signal feature length of 1 second is also short for the sound event detection task. Usually, the feature time length of the sound event detection task is more than 3 seconds. In addition, the real-time requirements of the sound event detection task are lower than those of the voice call task. Therefore, in order to obtain better classification results, in some embodiments, multiple continuous frames are used as input features of the multi-sound event classification model to obtain better temporal correlation. Specifically, after the general audio features continuously output by the audio feature extraction model are sequentially input into the buffer, all the general audio features in the buffer at the current moment are input into the multi-sound event classification model to obtain the first predicted probability of each sound event in the preset multiple categories of sound events at the current moment.
[0104] After the audio feature extraction model outputs the general audio features, it enters the buffer and waits. The buffer length is recorded as L, and the buffer B t for:
[0105]
[0106] Among them, norm(·) is the global normalization in formula (4), that is, It represents the Mel spectrum features from t*n_frames to (t+1)*n_frames time frames.
[0107] According to formula (7), buffer zone B t is a time series feature sequence with a dimension of (L, D), which conforms to the input format of the multi-sound event classification model. When t>L-1, the length of the buffer remains unchanged, but it is updated in a first-in-first-out queue mode. Taking L=5 as an example, B t The first five time series are B4 = {d0, d1, d2, d3, d4}, and the sixth time series becomes B5 = {d1, d2, d3, d4, d5}, and so on. t The five general audio features in the input are used as input to the multiple sound event classification model.
[0108] Then at time t, the predicted probability of the i-th sound event p t,i for:
[0109] p t,i =f2(B t+L-1 ) (8)
[0110] Among them, the input of f2 is buffer B t All features in , that is, the features of time point t and the previous L-1 time points, a total of L time series features.
[0111] The prediction result of the i-th sound event in this audio signal for:
[0112]
[0113] Among them, θ i The definition of is the same as in formula (6).
[0114] In some embodiments, the above-mentioned time series prediction results are further summarized and adjusted through post-processing. Figure 6 FIG. 1 is a flow chart showing a method for detecting a sound event according to another embodiment of the present invention. Figure 6 As shown, the method includes the following steps:
[0115] S300: Acquire frequency spectrum features of the audio to be detected.
[0116] S320: Input the spectrum features into an audio feature extraction model to obtain universal audio features.
[0117] S330: Input the common audio features continuously output by the audio feature extraction model into the buffer in sequence.
[0118] S340: Input all common audio features in the current buffer into the multi-sound event classification model to obtain a first predicted probability of each sound event in the preset multiple types of sound events at the current moment.
[0119] S360: Perform a first judgment on the first prediction probabilities at L consecutive moments to obtain L first judgment results.
[0120] The L moments include moment t and L-1 moments thereafter, and the first judgment result is whether each sound event exists or does not exist in the audio to be detected at each moment.
[0121] S380: Determine the existence of each sound event in the audio to be detected at time t based on the L first judgment results.
[0122] The preset multiple types of sound events include a first sound event. The detection of the first sound event will be described below by taking the first sound event as an example.
[0123] The first judgment result for the first sound event is whether the first sound event exists in the audio to be detected or not, then step S360 obtains L first judgment results for whether the first sound event exists in the audio to be detected or not at each moment in L consecutive moments. Step S380 further includes:
[0124] S381: Determine whether the number of target judgment results in the L first judgment results for the first sound event reaches a preset threshold.
[0125] Among them, the target judgment result is that the first sound event exists in the audio to be detected.
[0126] S382: If the number of target judgment results in the L first judgment results for the first sound event reaches a preset threshold, it is determined that the first sound event exists in the audio to be detected at time t.
[0127] As shown in equation (8), starting from d4, the multi-sound event classification model performs L inferences for each output of the audio feature extraction model, and the L inference results are counted through post-processing in steps S360 and S380. After step S360 obtains L first judgment results, the overall situation of these L first judgment results is counted using equation (10).
[0128]
[0129] For example, if If it is 1, it means that the buffer B at time t t There is the i-th sound event in this audio signal; if If it is 0, it means that the buffer B at time t t There is no i-th sound event in this audio signal. By using formula (10), we can count L sum.
[0130] In step S381 and step S382, it is determined whether there is a sound event i at time t by the following formula:
[0131]
[0132] Among them, φ i ∈Φ,Φ={φ i |φ i ≥1andφ i ≤L,i=1,2,...,C} is the post-processing threshold for determining whether the i-th sound event exists, and Φ is the set of post-processing thresholds for all C events to be detected. This set is similar to Θ. The user can also select the type of sound event to be detected and adjust the post-processing threshold.
[0133] The following is a specific example of how to summarize and adjust the timing results through post-processing in the embodiment of the present application. Taking L=5 as an example, starting from the fifth moment (the feature output by the audio feature extraction model corresponds to d4), the buffer feature is:
[0134] B4={d0,d1,d2,d3,d4}
[0135] B5={d1,d2,d3,d4,d5}
[0136] B6={d2,d3,d4,d5,d6}
[0137] B7={d3,d4,d5,d6,d7}
[0138] B8={d4,d5,d6,d7,d8}
[0139] After the inference of the multi-sound event classification model, from B4 to B8, each inference gives the probability of formula (8), then:
[0140]
[0141] Then use formula (9) to convert p t,i And the probability threshold θ in formula (8) i Compare and get get:
[0142]
[0143] Then apply formula (10), from arrive Statistics show:
[0144]
[0145] Therefore, the final statistical result is: in the five inferences of the multi-sound event classification model, how many times are the events judged to exist? Then compare the result with the post-processing threshold according to formula (11). If it exceeds the post-processing threshold φ i , then it is considered that event i exists at time t=4, otherwise sound event i does not exist at time t=4, as shown in the following formula:
[0146]
[0147] This embodiment filters the output time series results of the multi-sound event classification model through post-processing, thereby improving the accuracy of the detection results while allowing a certain delay.
[0148] In some embodiments, if it is determined that a certain type of sound event exists in the audio to be detected at time t, the prediction result can be further verified twice. The following is a detailed description using the first sound event as an example. After step S380, the sound event detection method provided in the embodiment of the present application further includes the following steps:
[0149] S390: When it is determined that a first sound event exists in the audio to be detected, the frequency spectrum features of the audio to be detected are input into a first sound event classification model corresponding to the first sound event to obtain a second predicted probability of the first sound event.
[0150] S400: Verify the existence of the first sound event in the audio to be detected based on the second predicted probability.
[0151] In the above method, the first sound event classification model is a model specifically used to identify the first sound event. The first sound event classification model can be a model obtained by training a neural network, and the neural network can be a GRU, RNN, LSTM or other network, such as the network in Table 2 above. The difference from the training of multiple sound event classification models is that, since the data input into the first sound event classification model is spectral features rather than general audio features, the training process of the first sound event classification model does not need to be combined with the aforementioned audio feature extraction model; in addition, the audio sample data used for model training includes positive samples containing the first sound event and negative samples that do not contain the first sound event, that is, the audio sample data used for model training is samples collected or produced for the first sound event, rather than samples collected or produced for all preset multiple categories of sound events. The training process of the first sound event classification model can be similar to the training process of the conventional classification model, which will not be repeated here.
[0152] If it is determined based on the first predicted probability that a first sound event exists in the audio to be detected, the spectral features of the audio to be detected are further directly input into the first sound event classification model for secondary detection. The secondary detection obtains a second predicted probability of the first sound event, and then the presence of the first sound event in the audio to be detected is verified based on the second predicted probability. If it is determined based on the second predicted probability that the first sound event exists in the audio to be detected, then it is determined that the first sound event exists in the audio to be detected, and at this time a notification of the presence of the first sound event can be sent to the user. If it is determined based on the second predicted probability that the first sound event does not exist in the audio to be detected, then the first predicted probability is considered inaccurate, and no notification is sent to the user.
[0153] Because the first sound event classification model is trained specifically for the first sound event and is specifically designed to identify the first sound event, it can accurately identify the first sound event. When the multi-sound event classification model determines that the first sound event exists, further verifying the presence of the first sound event using the first sound event classification model can improve the accuracy of the prediction of the first sound event.
[0154] In some embodiments, a sound event classification model corresponding to each of the multiple sound event categories is deployed in the camera device 1. For example, a first sound event classification model corresponds to a first sound event, a second sound event classification model corresponds to a second sound event... The training samples of each sound event classification model are samples collected or produced for that type of sound event. On this basis, one or more sound events set by the user can be secondary verified using the above method, which is conducive to meeting the user's personalized customization needs for secondary verification and improving the accuracy of the detection results of the sound events set by the user that require secondary verification.
[0155] In some embodiments, based on the deployment of sound event classification models corresponding to each of the above-mentioned preset multiple types of sound events in the camera device 1, all sound events can also be subjected to secondary verification by default, thereby improving the accuracy of all sound event detection results.
[0156] Figure 7 Figure 2 shows the test results of the sound event detection method of the embodiment of the present application. Figure 7 As shown, in a home indoor noise environment, the sound event detection method of the embodiment of the present application is used to detect the test sample with five sound events added. The audio feature extraction model deployed in the camera device used for the experiment adopts the model of the architecture shown in Table 1, and the multi-sound event classification model adopts the model of the architecture shown in Table 2. The sampling rate of the test sample is fs=16000 (i.e., 16kHz), and the feature of the input audio feature extraction model is the logarithmic Mel spectrum feature. The dimension of the logarithmic Mel spectrum feature includes (n_mels, n_frames), n_mels=64, frames=96, and the output feature dimension of the audio feature extraction module is D=512.
[0157] Figure 7In the figure, different colors represent different sound events. The first row shows the mel-spectrogram features of the audio signal, the second row shows the true event labels, and the third to seventh rows show the predicted sounds of a baby cry, a car horn, a cat meow, a dog bark, and glass break, respectively. The detection results shown in this figure are those without post-processing. For example, box B1 in the fourth row is a predicted car horn event, which is consistent with the true label of the event at the corresponding position in the second row. Therefore, the sample corresponding to this prediction is a true positive (TP), that is, a sample correctly predicted by the model as positive. Box B2 in the sixth row is a predicted dog barking event, which is inconsistent with the true label of the event at the corresponding position in the second row. No event is labeled at this position, so the sample corresponding to this prediction is a false positive (FP), that is, a sample incorrectly predicted by the model as positive. Box B3 in the second row is the true label of the car horn event, but no car horn event is detected at the corresponding position in the car horn event prediction result in the fourth row. Therefore, the sample corresponding to the true label of this event is a false negative (FN), that is, a sample incorrectly predicted by the model as negative. As can be seen from this, there are very few false positives in the test results, and the small number of missed negatives is because the sound events used in the test only last 1-3 seconds. In real scenarios, most of the above sound events last for a long time (for example, a baby's cry usually does not last only one time, and a cat or dog barks several times). Therefore, the prediction results in real scenarios will be slower than those in real scenarios. Figure 7 The shown is more accurate.
[0158] Based on the TP, FP, and FN in the detection results, the precision and recall of each sound event are calculated respectively. Then, the F1 Score of each sound event is calculated according to the precision and recall of each sound event. The calculation formulas are (13), (14), and (15), respectively:
[0159]
[0160] The calculation shows that the average F1 score of the five sound events is above 0.9. This shows that the method provided by the embodiment of the present application can accurately detect multiple types of sound events without burdening the terminal device with computing resources.
[0161] An embodiment of the present application provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned sound event detection method embodiment is implemented.
[0162] An embodiment of the present application provides a computer program, which can be executed by a processor to implement the above-mentioned sound event detection method embodiment.
[0163] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned sound event detection method embodiment is implemented.
[0164] In the several embodiments provided in this application, if any function is implemented in the form of a software function module / unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the technical solution of this application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or other electronic device) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store computer program code.
[0165] The algorithm or demonstration provided here are not inherently relevant to any particular computer, virtual system or other equipment. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing this type of system. In addition, the present application embodiment is not directed to any specific programming language yet. It should be understood that various programming languages can be utilized to realize the content of the present application described here, and the above description of specific languages is for the purpose of disclosing the best mode of implementation of the present application.
[0166] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In claims that list several means, several units or modules of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names. The steps in the above embodiments should not be understood as limiting the order of execution unless otherwise specified.
[0167] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A sound event detection method, characterized in that: The method comprises: Acquire a spectrum feature of the audio to be detected, wherein the dimensions of the spectrum feature include the number of frequency points and the number of consecutive frames; Inputting the spectral features into an audio feature extraction model to obtain a general audio feature, where the general audio feature is a feature vector used to indicate audio information; Inputting the universal audio feature into a multi-sound event classification model to obtain a first predicted probability of each sound event in a preset multi-class sound event; The existence of each sound event in the audio to be detected is determined based on the first predicted probability.
2. The method according to claim 1, characterized in that The audio feature extraction model includes at least one convolutional neural network block and at least one fully connected block cascaded in sequence, the spectral feature is a Mel spectrum feature, and the number of frequency points corresponds to the number of Mel filters.
3. The method according to claim 2, characterized in that After inputting the spectral features into the audio feature extraction model to obtain the general audio features, the method further includes: The general audio features continuously output by the audio feature extraction model are input into the buffer in sequence, wherein the buffer capacity is L features, L is a positive integer greater than 1, and when the current moment is t and t>L-1, the buffer stores the general audio features at time t and the previous L-1 moments.
4. The method according to claim 1, wherein The preset multiple types of sound events include a first sound event, and the method further includes: When it is determined that the first sound event exists in the audio to be detected, inputting the spectral features of the audio to be detected into a first sound event classification model corresponding to the first sound event to obtain a second predicted probability of the first sound event; Based on the second predicted probability, the existence of the first sound event in the audio to be detected is verified.
5. The method according to claim 3, characterized in that The method of inputting the universal audio features into a multi-sound event classification model to obtain a first predicted probability of each sound event in a preset multi-category sound event includes: inputting all universal audio features in the buffer at the current moment into the multi-sound event classification model to obtain a first predicted probability of each sound event in a preset multi-category sound event at the current moment.
6. The method according to claim 5, characterized in that Determining the presence of each sound event in the audio to be detected based on the first predicted probability includes: Performing a first judgment on the first predicted probabilities at L consecutive moments to obtain L first judgment results, where the L moments include moment t and L-1 moments thereafter, and the first judgment results are whether each sound event exists in the audio to be detected at each moment; Based on the L first judgment results, the existence of each sound event in the audio to be detected at time t is determined.
7. The method according to claim 6, characterized in that The preset multiple types of sound events include a first sound event, and the first judgment result for the first sound event is whether the first sound event exists in the audio to be detected; The determining, based on the L first judgment results, the presence of each sound event in the audio to be detected at time t includes: Determining whether the number of target judgment results among the L first judgment results for the first sound event reaches a preset threshold, wherein the target judgment result is that the first sound event exists in the audio to be detected; If the number of target judgment results in the L first judgment results for the first sound event reaches a preset threshold, it is determined that the first sound event exists in the audio to be detected at time t.
8. The method according to claim 1, characterized in that The parameters of the audio feature extraction model are obtained as follows: Load the parameters of the pre-trained audio feature extraction model and randomly initialize the parameters of the multi-sound event classification model to be trained; Inputting the spectral features of the audio sample data into the pre-trained audio feature extraction model, and inputting the output results of the pre-trained audio feature extraction model into the multi-sound event classification model to be trained, wherein the audio sample data corresponds to the preset multiple types of sound events; The multi-sound event classification model to be trained is trained by gradient back propagation, and the parameters of the pre-trained audio feature extraction model are fine-tuned.
9. A sound event detection device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.