Sound event detection method

By introducing the frequency dynamic convolution module of the attention mechanism and the pre-trained Transfomer audio teacher and student model in sound event detection, the problem of insufficient feature extraction and low detection accuracy in the scenes of overlapping multiple sound sources in the prior art is solved, and higher detection accuracy and robustness are achieved.

CN119993202AActive Publication Date: 2025-05-13NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202510107923.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing sound event detection technology has low detection accuracy in complex scenarios such as feature extraction and multi-sound source overlap, ignoring the unique prior knowledge of sound events in acoustic scenes. The translation invariance and fixed receptive field of convolutional neural networks limit the feature capture capability.

Method used

A method for sound event detection is proposed, combining the frequency dynamic convolution module of attention mechanism (LSKFDY-CNN) and the pre-trained Transfomer-based audio teacher student model (ATST), adaptively adjust the receptive field, enhance the multi-scale feature capture ability, and make full use of spatial information.

Benefits of technology

It improves the accuracy and robustness of sound event detection, overcomes the defects of translation invariance and fixed receptive fields when extracting features by convolutional neural networks, and enhances the detection performance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993202A_ABST
    Figure CN119993202A_ABST
Patent Text Reader

Abstract

The invention provides a sound event detection method, and belongs to the technical field of sound event detection, and the method comprises the steps: obtaining a to-be-detected audio signal; preprocessing the audio signal to be detected; performing data enhancement on the audio signal after data preprocessing; performing feature extraction on the audio signal after data enhancement; performing context information extraction on the extracted features to obtain first branch features; inputting the audio signal after data enhancement into a pre-trained Transfomer-based audio teacher-student model to obtain a second branch feature; and splicing the first branch feature and the second branch feature, and inputting the spliced first branch feature and the spliced second branch feature into a classifier to obtain a sound event detection result of the to-be-detected audio signal. According to the invention, enough background information corresponding to different sound events can be obtained, and the multi-scale feature capturing capability of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of sound event detection, and in particular relates to a method for sound event detection. Background Art

[0002] As an important sub-discipline of machine hearing, Sound Event Detection (SED) is one of the main tasks in acoustic scene analysis. It aims to determine the categories of sound events in the input audio and their corresponding start and offset times (timestamps) in the audio frequency band. At present, SED-related technologies have been widely used in smart home, equipment fault monitoring, autonomous driving and other fields. Sound event detection helps to monitor the environment in real time and detect abnormal sounds in time, which is particularly important for scenarios such as smart transportation, industrial fault monitoring, and smart home.

[0003] The use of convolutional recurrent neural network (CRNN) to achieve SED tasks is very extensive and is the current mainstream model. Its structure is composed of cascaded convolutional neural networks (CNN) and recurrent neural networks (RNN), where CNN is responsible for feature extraction and RNN is responsible for obtaining context information. Guo Fanghong (patent application number: CN116798448A) et al. based on the CRNN network, used hyperparameterized convolutional layers for feature extraction, reduced model parameters and computational complexity, and improved performance. Xie Zongxia et al. (patent application number: CN114881212A) used a dual-branch feature extraction and fusion method to improve the accuracy of sound event detection by improving the distinction between sound event classes. However, current technologies all ignore the unique prior knowledge presented by sound events in acoustic scenes. The lack of spatial information leads to insufficient feature extraction, which limits the performance of sound event detection in multi-sound source scenes.

[0004] From the current research status at home and abroad, we can see that sound event detection technology has a certain research foundation. However, there are still problems that need to be improved in terms of feature extraction and low detection accuracy in complex scenes with multiple overlapping sound sources. The main problems are as follows:

[0005] (1) The shortcomings of CNN’s translation invariance and fixed receptive field make it difficult for it to capture the subtle relationship between different frequency dimensions in the spectrogram, resulting in limited detection accuracy.

[0006] (2) The existing sound event detection network ignores the unique prior knowledge presented by sound events in acoustic scenes, resulting in missed detections and false detections in overlapping sound event detections.

[0007] Existing models usually need to convert audio into images when extracting audio signal features using CNN, and extract signal time-frequency features on the spectrogram. However, unlike images, even if the high-frequency part and the low-frequency part show the same features on the audio spectrogram, they may not be the same acoustic event. This is the biggest difference between spectrograms and natural images. Traditional CNNs do not take this into account. In recent years, there have been a variety of methods that combine attention mechanisms with convolutional blocks to solve this problem, such as SENET (Squeeze-and-Excitation Networks) and CBAM (Convolutional Block Attention Module). However, since the spectrogram of the audio signal is not static, the attention weight is required to be constant, and these network architectures generally obtain dynamic attention weights by changing the image, resulting in limited feature acquisition accuracy and affecting the accuracy of SED. For this reason, some researchers have proposed frequency dynamic convolution, which generates a convolution kernel that is adaptive to the input frequency component by weighting the frequency dimension, achieving good results. However, the receptive field range of the adaptive convolution kernel is still fixed, ignoring the powerful and valuable prior knowledge of the sound event in the acoustic scene, resulting in the loss of feature information. The reason is that the duration and feature complexity of different sound events are very different. For example, alarms and ringtones are usually single, clear sound events with obvious features, and their detection requires relatively little contextual information; while the recognition of cat and dog sounds requires considering their species, activity status and environment, so the features are relatively complex. Different breeds of dogs have different barking or chirping styles, and may be affected by the surrounding environment and emotional state. The features change complexly over time. Recognizing such events depends on extensive contextual information because the surrounding environment can provide valuable clues about their location, duration, environmental conditions, etc. The fixed-size convolution receptive field will cause the network to only focus on the local features of the sound event. When complex or overlapping sound events appear in the scene, the detection performance will decrease. Therefore, in order to improve the accuracy and robustness of SED in complex scenes, it is necessary to study a convolution kernel with a large receptive field range and strong adaptability.

[0008] (3) The pre-trained model BEATs commonly used in existing research is a patch-wise organized (block-level) input. The extracted features are more similar to those of CNN and lack attention to global features. Therefore, the integration of the two limits the performance of sound event detection. Summary of the invention

[0009] In view of the shortcomings of the prior art, this application proposes a method for detecting sound events, which enables the model to obtain sufficient background information corresponding to different sound events, thereby improving the model's ability to capture multi-scale features. In addition, ATST is a frame-wise input, which compensates for the shortcomings of CNN in extracting local features and enhances global characteristics. Therefore, the benefits from CNN are higher, the method of this application is more compatible, and the detection effect has been greatly improved.

[0010] In a first aspect, the present application proposes a method for detecting a sound event, comprising:

[0011] Acquire the audio signal to be detected;

[0012] Preprocessing the audio signal to be detected;

[0013] Perform data enhancement on the audio signal after data preprocessing;

[0014] Extract features from the audio signal after data enhancement;

[0015] Extract context information from the extracted features to obtain the first branch features;

[0016] The data-augmented audio signal is input into the pre-trained Transformer-based audio teacher-student model to obtain the second branch features;

[0017] The first branch feature and the second branch feature are concatenated and input into the classifier to obtain the sound event detection result of the audio signal to be detected.

[0018] The preprocessing of the audio signal to be detected includes:

[0019] Resampling the audio signal to be detected;

[0020] The resampled signal is subjected to Fourier transform to obtain a two-dimensional spectrum diagram.

[0021] The step of performing data enhancement on the audio signal after data preprocessing includes:

[0022] Randomly shift the time frame of the audio signal after data preprocessing along the time axis to obtain a first enhancement result;

[0023] The preprocessed audio signal is randomly mixed using a mixing parameter to obtain a second enhancement result;

[0024] Selecting a time window to be masked according to the mask, shielding the selected time window for consecutive time frames so that a corresponding portion of the preprocessed audio signal is empty, and obtaining a third enhancement result;

[0025] The first enhancement result, the second enhancement result and the third enhancement result are taken as the data enhancement results of the original audio signal.

[0026] The feature extraction of the audio signal after data enhancement includes:

[0027] The audio signal after data enhancement is input into the CNN module for preliminary feature extraction to obtain the initial features;

[0028] The initial features are input into the frequency dynamic convolution module of multiple attention mechanisms for in-depth feature extraction to obtain in-depth feature extraction results.

[0029] The audio signal after data enhancement is input into the CNN module for preliminary feature extraction to obtain initial features, including:

[0030] The audio signal after data preprocessing is input into the CNN layer for feature extraction;

[0031] Input the output of the CNN layer to the first normalization layer for normalization;

[0032] The normalized result is input into the first GLU layer for linear processing;

[0033] Input the result after linear processing to the first pooling layer for pooling processing;

[0034] The pooled result is regularized to obtain a regularized result, which is used as the initial feature.

[0035] The initial features are input into the frequency dynamic convolution modules of multiple attention mechanisms for in-depth feature extraction, and in-depth feature extraction results are obtained, including:

[0036] Step S4.2.1: performing frequency dynamic convolution on the initial features;

[0037] Step S4.2.2: Input the frequency dynamic convolution result into the second normalization layer for normalization processing, and use the normalization result as the first feature;

[0038] Step S4.2.3: Input the normalized results into the first convolution kernel and the second convolution kernel respectively, and obtain the first convolution kernel output and the second convolution kernel output respectively;

[0039] Step S4.2.4: Input the first convolution kernel output into the first convolution layer to obtain the first convolution result;

[0040] Step S4.2.5: Input the second convolution kernel output into the second convolution layer to obtain a second convolution result;

[0041] Step S4.2.6: concatenating the first convolution result and the second convolution result to obtain a first concatenated result;

[0042] Step S4.2.7: input the first splicing result into the second pooling layer for pooling processing;

[0043] Step S4.2.8: Input the pooling processing result to the third convolutional layer;

[0044] Step S4.2.9: Input the output of the third convolutional layer into the Sigmoid activation function to obtain the first attention weight and the second attention weight;

[0045] Step S4.2.10: multiply the first convolution result by the first attention weight to obtain a first multiplication result;

[0046] Step S4.2.11: multiply the second convolution result by the second attention weight to obtain a second multiplication result;

[0047] Step S4.2.12: concatenating the first multiplication result and the second multiplication result to obtain a second concatenated result;

[0048] Step S4.2.13: multiply the second concatenation result by the first feature to obtain a dynamic convolution structure extraction feature result;

[0049] Step S4.2.14: Input the dynamic feature structure extraction feature result into the second GLU layer to obtain the linear processing result;

[0050] Step S4.2.15: The linear processing result is pooled through the third pooling layer to obtain the first in-depth feature extraction result;

[0051] Step S4.2.16: Use the first in-depth feature extraction result as the initial feature and return to step S4.2.1 until the preset number of repetitions is reached, and use the last in-depth feature extraction result as the final in-depth feature extraction result.

[0052] The frequency dynamic convolution is used to perform frequency dimension segmentation and weighting on the initial features, and the calculation formula is as follows:

[0053]

[0054] G k =π k (f,x)

[0055] y k (t,f)=W k *X(t,f)+B

[0056] Among them, Y(t,f,x) is the frequency dynamic convolution result, t is time, f is frequency, k is the kth base kernel, N is the total number of base kernels, x is the input data, π k is the attention function of the kth base kernel, y k (t,f) is the output of the kth base kernel, G k is the attention weight of the kth base kernel, W k is the input weight matrix, B is the bias matrix, and X(t,f) is the initial feature.

[0057] The extracting context information from the extracted features to obtain the first branch features includes:

[0058] Input the extracted features into the BGRU module;

[0059] The output of the BGRU module is input into the attention module, and the output of the attention module is used as the first branch feature.

[0060] In a second aspect, the present application proposes a computer program product, including a computer program or instructions, which implement the method for detecting a sound event when executed by a processor.

[0061] Beneficial effects:

[0062] This application proposes a method for sound event detection. In order to capture the deep features of different sound events, the model adaptively adjusts the receptive field, and fully utilizes spatial information, a frequency dynamic convolution module of the attention mechanism (i.e., the LSKFDY-CNN convolution structure) is proposed. On the one hand, more frequency information of the sound event is retained, and on the other hand, a wide and adaptive receptive range is provided for the model. By dynamically adjusting the receptive field, the model can obtain enough background information corresponding to different sound events, thereby improving the model's multi-scale feature capture capability. It overcomes the defects of existing convolutions that the convolutional neural network is prone to lose the relationship between features and unique contextual background information when extracting features due to translation invariance and fixed receptive field.

[0063] In order to effectively utilize the pre-trained model and further improve the performance of sound event detection, this application integrates the pre-trained Transfomer-based Audio Teacher Student Model (ATST model) with the frequency dynamic convolution module of the attention mechanism to obtain more discriminative features and achieve efficient feature extraction when there is limited strong label data. The ATST model is very good at generating frame-level audio representations, with more refined features and obvious linear separability. The LSKFDY-CNN convolution structure can benefit greatly from it. It compensates for the shortcomings of CNN in extracting local features and enhances the global characteristics. Therefore, the combination of the features extracted by the two further improves the detection effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 A flow chart of a method for detecting sound events according to an embodiment of the present invention;

[0065] Figure 2 A schematic diagram of a flow chart of a sound event detection method based on a dynamic convolution structure according to an embodiment of the present invention;

[0066] Figure 3 A schematic diagram of the structure of the LSKFDYCNN module of an embodiment of the present invention;

[0067] Figure 4 The structural stability analysis of the model of the embodiment of the present invention, wherein (a) is a schematic diagram of PSDS1, and (b) is a schematic diagram of PSDS2;

[0068] Figure 5 Visual comparison of model prediction results in actual scenarios of the embodiments of the present invention, where (a) is the prediction result in a single-tone scenario and (b) is the prediction result in a multi-tone scenario. DETAILED DESCRIPTION

[0069] The specific implementation of the present application is further described in detail below in conjunction with the drawings and examples.

[0070] This application proposes a method for sound event detection, namely, a method for sound event detection based on Large Selective Kernel Frequency Dynamic Convolutional Neural Network (LSKFDY-CNN), which aims to improve the accuracy of sound event detection and enhance the robustness of the sound event detection model of the system in acoustic scenes of different complexity. This application builds a new convolution structure, which combines FDY (FrequencyDynamic) convolution and Large Selective Kernel (LSK) attention mechanism for feature extraction. The FDY convolution block is responsible for obtaining the feature differences of the spectrograms in different frequency bands. The LSK module can capture the background spatial information of different events and events of different durations by dynamically adjusting the receptive field of the feature extraction backbone, and adaptively extract contextual information of different scales.

[0071] On this basis, the ATST model (ATST: Audio Representation Learning with Teacher-Student Transformer, audio teacher-student model based on Transformer) is introduced to extract embedded features, realize efficient feature extraction with limited data, and ultimately effectively improve the performance of sound detection. It overcomes the shortcomings of insufficient sound feature extraction in existing networks and the inability to effectively utilize powerful and valuable prior knowledge of different events.

[0072] Example 1

[0073] This embodiment proposes a method for detecting a sound event, such as Figure 1 As shown, including:

[0074] Step S1: obtaining an audio signal to be detected;

[0075] Step S2: preprocessing the audio signal to be detected;

[0076] Step S3: Perform data enhancement on the audio signal after data preprocessing

[0077] Step S4: extracting features from the audio signal after data enhancement;

[0078] Step S5: extracting context information from the extracted features to obtain first branch features;

[0079] Step S6: input the data-enhanced audio signal into the pre-trained Transformer-based audio teacher-student model to obtain the second branch feature;

[0080] Step S7: splicing the first branch feature and the second branch feature and inputting them into the classifier to obtain the sound event detection result of the audio signal to be detected.

[0081] In step S2, the preprocessing of the audio signal to be detected includes:

[0082] Step S2.1: resampling the audio signal to be detected;

[0083] Step S2.2: Perform Fourier transform on the resampled signal to obtain a two-dimensional spectrum diagram.

[0084] In this embodiment, the audio signal to be detected of the input system is preprocessed and converted from a one-dimensional audio signal to a two-dimensional spectrogram. The specific preprocessing process is: the 44.1kHz audio signal is resampled to 16kHz, and the audio signal is divided into frames with 2,048 samples each and a jump length of 256 samples. Each time frame undergoes a 2048-point fast Fourier transform (FFT), and then undergoes a 128-dimensional Mel filter bank analysis, and finally obtains a 128-dimensional Log-Mel spectrogram.

[0085] In step S3, the data enhancement is performed on the audio signal after data preprocessing, including:

[0086] Step S3.2: Perform data enhancement on the preprocessed original audio signal and use the enhanced result as a training sample, including:

[0087] Step S3.2.1: randomly shifting the preprocessed audio signal along the time axis by a time frame to obtain a first enhancement result;

[0088] Step S3.2.2: randomly mix the preprocessed audio signal using a mixing parameter to obtain a second enhancement result;

[0089] Step S3.2.3: Select the time window to be masked according to the mask, and mask the selected time window in consecutive time frames so that the corresponding part of the preprocessed audio signal is empty, thereby obtaining a third enhancement result;

[0090] Step S3.2.4: Use the first enhancement result, the second enhancement result and the third enhancement result as the data enhancement results of the historical audio signal.

[0091] In this embodiment, data enhancement is performed on the data after preprocessing. In the present invention, three different data enhancement techniques, namely frame shifting, data mixing and time masking, are used to expand the training data. Among them, frame shifting is to randomly move the time frame along the time axis to generate new training samples; the mixing method is to randomly mix the selected data samples with mixing parameters to obtain new samples; time masking is to first select the time window to be masked according to the strong mask, and then shield the selected time window with continuous time frames, so that the corresponding part on the spectrum diagram is empty, and new training data is obtained. The use of data enhancement technology can enhance the diversity of training data, overcome the problem of small data volume to a certain extent, and at the same time help the model better identify the characteristics of different sound events and the subtle differences between them, and enhance the robustness and detection accuracy of the model.

[0092] In step S4, the audio signal after data enhancement is subjected to feature extraction, such as Figure 2 As shown, including:

[0093] Step S4.1: Input the data-enhanced audio signal into the CNN module for preliminary feature extraction to obtain initial features, including:

[0094] Step S4.1.1: Input the audio signal after data preprocessing into the CNN layer for feature extraction;

[0095] Step S4.1.2: Input the output of the CNN layer to the first normalization layer for normalization;

[0096] Step S4.1.3: Input the normalized result to the first GLU layer for linear processing;

[0097] Step S4.1.4: input the result after linear processing to the first pooling layer for pooling processing;

[0098] Step S4.1.5: Regularize the pooled result to obtain a regularized result, and use the regularized result as the initial feature.

[0099] In this embodiment, the audio signal after data preprocessing is subjected to CNN to extract shallow features, and then passes through a normalization layer, a first GLU layer and a first pooling layer to obtain initial features.

[0100] Step S4.2: Input the initial features into the frequency dynamic convolution module of multiple attention mechanisms for deep feature extraction, and obtain the deep feature extraction results, such as Figure 3 As shown, including:

[0101] Step S4.2.1: performing frequency dynamic convolution on the initial features;

[0102] Step S4.2.2: Input the frequency dynamic convolution result into the second normalization layer for normalization, and use the normalization result as the first feature;

[0103] Step S4.2.3: Input the normalized results into the first convolution kernel and the second convolution kernel respectively, and obtain the first convolution kernel output and the second convolution kernel output respectively;

[0104] Step S4.2.4: Input the first convolution kernel output into the first convolution layer to obtain the first convolution result;

[0105] Step S4.2.5: Input the second convolution kernel output into the second convolution layer to obtain a second convolution result;

[0106] Step S4.2.6: concatenating the first convolution result and the second convolution result to obtain a first concatenated result;

[0107] Step S4.2.7: input the first splicing result into the second pooling layer for pooling processing;

[0108] Step S4.2.8: Input the pooling processing result to the third convolutional layer;

[0109] Step S4.2.9: Input the output of the third convolutional layer into the Sigmoid activation function to obtain the first attention weight and the second attention weight;

[0110] Step S4.2.10: multiply the first convolution result by the first attention weight to obtain a first multiplication result;

[0111] Step S4.2.11: multiply the second convolution result by the second attention weight to obtain a second multiplication result;

[0112] Step S4.2.12: concatenating the first multiplication result and the second multiplication result to obtain a second concatenated result;

[0113] Step S4.2.13: multiply the second concatenation result by the first feature to obtain a dynamic convolution structure extraction feature result;

[0114] Step S4.2.14: Input the dynamic feature structure extraction feature result into the second GLU layer to obtain the linear processing result;

[0115] Step S4.2.15: The linear processing result is pooled through the third pooling layer to obtain the first deep feature extraction result; Step S4.2.16: The first deep feature extraction result is used as the initial feature, and the process returns to step S4.2.1 until the preset number of repetitions is reached, and the last deep feature extraction result is used as the final deep feature extraction result.

[0116] In this embodiment, the model architecture is as follows: Figure 3As shown. LSKFDY-CNN is used to extract the features of the spectrogram, including frequency features, spatial context information and high-level semantic information. LSKFDY-CNN is the core module proposed by the present invention. The module consists of two parts: the FDY convolution block and the LSK module. The FDY convolution block generates an adaptive frequency convolution kernel by dividing and weighting the feature frequency dimension, which weakens the translation invariance of CNN on the frequency axis and obtains sufficient frequency information, making the convolution operation more reasonable for audio features; the LSK module uses its adaptive receptive field to enhance and fuse the feature map according to the differences in the features of different sound events, pays more attention to spatial information, and provides reasonable context information for different event types. The adaptive selection mechanism of LSKFDY can effectively weight the features processed by a series of large kernels and then merge them in space. The weights of these kernels are dynamically determined based on the input, allowing the model to adaptively use different large kernels (i.e., the first convolution kernel and the second convolution kernel) and adjust the receptive field of each target in the space as needed. The first convolution kernel k 1 With the second convolution kernel k 2 , two large convolution kernels of different sizes are decomposed from a larger large kernel, and the dilated convolution method is used. The decomposed large convolution kernel adopts the dilated convolution method, for example, a (3×3) large convolution kernel with a dilation rate of (2×2) is used to replace the (5×5) convolution. The advantage of this is that it alleviates the quadratic increase in computational cost caused by the large kernel size of ordinary convolution. The dilated convolution quickly expands the receptive field through the dilation factor, which can improve the speed and efficiency of the model. Through the first convolution kernel k 1 With the second convolution kernel k 2 After that, through a series of operations such as convolutional dimension reduction, pooling, activation function, and feature map concatenation, the features of different scales extracted by the two large cores are assigned corresponding weights, thus completing the selection mechanism. The detailed principles and processes are as follows:

[0117] FDYCNN calculates the weight of each base kernel by dividing the frequency dimension of the spectrogram features and obtains the output of each base kernel:

[0118] y k (t,f)=W k *X(t,f)+B

[0119] Among them, t and f represent time and frequency, k represents the kth base kernel, W represents the weight matrix, and B represents the bias matrix. Then, the attention weights of the k base kernels are adaptively generated:

[0120] G k =π k (f,x)

[0121] Finally, the weighted sum is calculated for the output results:

[0122]

[0123] Among them, Y(t,f,x) is the frequency dynamic convolution result, that is, the output of the frequency dynamic layer after aggregating multiple base kernels, t is time, f is frequency, k represents the kth base kernel, N represents the total number of base kernels, x is the input data, π k is the attention function of the kth base kernel, y k (t,f) is the output of the kth base kernel, G k is the attention weight of the kth base kernel, W k is the input weight matrix, B is the bias matrix, and X(t,f) is the initial feature.

[0124] The detailed workflow of the LSK module is as follows: The large convolution kernels of the two branches are used to extract features respectively. In the experiment, the sizes of the two selective large kernels are 5×5 (i.e., the first convolution kernel k 1 ) and 25×25 (i.e., the second convolution kernel k 2 ), and then the size of the feature map is reduced to 1 dimension through (1×1) convolution, that is, X 1 , X 2 ∈R 1×F×F , concatenate the features of two different scales to get X. In order to allow information interaction between different spatial descriptors, the spatial maximum pooling and average features are connected, and after spatial pooling, we get:

[0125] Agg=(P max (X)+P avg (X)

[0126] Among them, P max and P avg They represent the maximum pooling and average pooling operations respectively, and Agg represents the features after spatial pooling.

[0127] Then, the sigmoid function is used to process the vector and map its value to between 0 and 1 to obtain the attention weights Wi of the two branch feature maps:

[0128] W i =sigmoid(F 2→c (Agg)

[0129] Among them, W i Represents the attention weights of the feature maps of different branches, including: the first attention weight W 1 And the second attention weight W 2 .

[0130] Finally, the attention weight is multiplied by the original feature map to obtain the weighted feature map, and the final output F(X) is obtained after feature fusion.

[0131]

[0132] in, is the i-th convolution result, including: the first convolution result And the second convolution result In this embodiment, N is 2 and X' is the first feature.

[0133] In step S5, context information is extracted from the extracted features to obtain first branch features, including:

[0134] Step S5.1: Input the extracted features into the BGRU module;

[0135] Step S5.2: Input the output of the BGRU module to the attention module, and use the input of the attention module as the first branch feature.

[0136] In this embodiment, the extracted features are input into the BGRU module to fit the context information, and then the feature map is enhanced through a single LSK module.

[0137] Step S6: The pre-trained Transfomer audio teacher-student model performs feature extraction on the data-enhanced samples.

[0138] The pre-trained Transformer-based audio teacher-student model (i.e., pre-trained model ATST) is integrated with the backbone network as another branch to extract more discriminative features from the audio, helping the original model to obtain high-level features containing rich semantic information with limited data. The audio feature dimension extracted by ATST is (250×1×768), and the audio feature dimension extracted by LSKFDY-CNN is (250×1×256). After concatenating the two, a feature with 1024 channels is obtained.

[0139] In step S7, the first branch feature and the second branch feature are concatenated and input into a classifier to obtain a sound event detection result of the audio signal to be detected.

[0140] In this embodiment, the features extracted by the two branches are concatenated and sent to the extracted feature input classifier, and the time-distributed fully connected layer (FC) and activation function sigmoid are applied to predict the existence probability of the sound event along the time axis. Finally, the predicted probabilities are aggregated and the sound event category and timestamp information are output.

[0141] This embodiment uses the indoor sound event detection data set (DESED), which is currently the most mainstream data set. This embodiment is used to evaluate the performance of the SED system by calculating the polyphonic sound event detection score (PSDS) and F1 score. F1 score is the harmonic mean of precision (P) and recall (R). The PSDS index is obtained by the normalized area under the ROC curve surrounded by a series of coordinates. A higher PSDS value indicates better model performance. It is more robust to label subjectivity and can better understand data deviations and classification stability across sound classes. In the present invention, PSDS1 and PSDS2 are calculated separately. The former focuses on the rapid detection of the start time in the timestamp information of the sound event, and the latter focuses on the model's ability to distinguish sound categories.

[0142] (1) It satisfies the requirement of accurate detection of sound events in acoustic scenes with multiple sound sources. The invention is simple to implement, has high detection accuracy, is highly compatible with other sound processing tasks, and meets the application requirements. In order to verify the superiority of the proposed structure, it is compared with the mainstream model with and without the pre-trained model, and the results are shown in Tables 1 and 2.

[0143] Table 1 Performance comparison of baseline and different dynamic convolutional structure models

[0144]

[0145] Table 2 Comparison of the effects of pre-trained model ATST feature embedding model

[0146]

[0147] From Table 1 and Table 2, it can be seen that the four performance indicators of the LSKFDY-CNN proposed in this paper are higher than other similar systems, which shows the effectiveness of the LSKFDY convolution. At the same time, the network performance is further improved after adding the pre-trained model ATST in Table 2. At the same time, it further illustrates the effectiveness of the LSKFDY structure. In addition, compared with BEATs, ATST is better at generating frame-level audio representations, with more refined features and obvious linear separability. Therefore, the LSKFDY structure proposed by us is more compatible and complementary with the pre-trained model ATST. ATST compensates for the characteristics of CNN extracting local features and enhances the global characteristics. The proposed LSKFDY-CNN structure can obtain enough information from itself and the environment for sound events with different feature complexities. The large receptive field characteristics of the convolution can enhance the ability to distinguish long and short time series events. The selective mechanism can focus on the important features related to the corresponding events. Through selective attention, the model can identify and emphasize the key features related to the event, thereby improving the detection accuracy.

[0148] (2) The model is more stable and has better performance. In order to test the dynamic structural stability of the LSKFDY-CNN network proposed in this paper, we recorded the PSDS index of 12 experiments and compared it with FDYCNN. The results are shown in the figure below. Figure 4 shown.

[0149] The evaluation experiment results show that the stability of the PSDS1 model we proposed has been significantly enhanced compared to FDYCNN; although the stability of the PSDS2 indicator has not been significantly improved, the effect of each experiment is better than that of FDYCNN, and there is still a significant improvement in performance. The reason is that events with different time-frequency pattern complexity obtain their corresponding prior information, and the model is less affected by data diversity and small data volume, so the accuracy is improved and the detection speed is accelerated, making the PSDS1 indicator more stable.

[0150] (3) High detection accuracy in multi-sound overlapping scenarios. It can effectively model overlapping sound events and provide accurate classification and timestamp detection. The test results for actual scenarios are as follows: Figure 5 shown.

[0151] Figure 5 (a) and (b) intuitively show the sound category and event activity range output by different models on the test set in two acoustic scenarios: monophonic and polyphonic. The true label shows the true sound category and event activity range of each time frame, which provides a standard for evaluating the prediction accuracy of other models. From the results, it can be seen that compared with the existing mainstream model structure, the prediction results of our proposed model structure are closest to the true label, and the classification characteristics and timestamp prediction also achieve the best detection effect. The sound category can be predicted more accurately in most time frames with a low misjudgment rate. Our model is more sensitive in detecting the start and end of audio events, and has the best detection effect for overlapping sounds. In the monophonic acoustic scene, the true label is almost perfectly matched over the entire time series. In the complex acoustic scene with multiple overlapping sounds, the category and timestamp information can also be successfully detected, indicating the effective modeling of overlapping sound events, which can achieve high detection accuracy even in complex environments with multiple sound sources.

[0152] Embodiment 2:

[0153] This embodiment provides a computer program product, including a computer program or instructions, which implement the method for detecting a sound event when executed by a processor.

[0154] Based on such understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a computer program product.

[0155] The various embodiments in the present application are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

[0156] The protection scope of the present application is not limited to the above-mentioned embodiments. Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the scope and spirit of the present disclosure. If these changes and modifications fall within the scope of the claims of the present disclosure and their equivalents, the intention of the present disclosure also includes these changes and modifications.

Claims

1. A method for detecting a sound event, characterized in that: include: Acquire the audio signal to be detected; Preprocessing the audio signal to be detected; Perform data enhancement on the audio signal after data preprocessing; Extract features from the audio signal after data enhancement; Extract context information from the extracted features to obtain the first branch features; The data-augmented audio signal is input into the pre-trained Transformer-based audio teacher-student model to obtain the second branch features; The first branch feature and the second branch feature are concatenated and input into the classifier to obtain the sound event detection result of the audio signal to be detected.

2. The method for detecting a sound event according to claim 1, characterized in that: The preprocessing of the audio signal to be detected includes: Resampling the audio signal to be detected; The resampled signal is subjected to Fourier transform to obtain a two-dimensional spectrum diagram.

3. The method for detecting a sound event according to claim 1, characterized in that: The step of performing data enhancement on the audio signal after data preprocessing includes: Randomly shift the time frame of the audio signal after data preprocessing along the time axis to obtain a first enhancement result; The preprocessed audio signal is randomly mixed using a mixing parameter to obtain a second enhancement result; Selecting a time window to be masked according to the mask, shielding the selected time window for consecutive time frames so that a corresponding portion of the preprocessed audio signal is empty, and obtaining a third enhancement result; The first enhancement result, the second enhancement result and the third enhancement result are taken as the data enhancement results of the original audio signal.

4. The method for detecting a sound event according to claim 1, characterized in that: The feature extraction of the audio signal after data enhancement includes: The audio signal after data enhancement is input into the CNN module for preliminary feature extraction to obtain the initial features; The initial features are input into the frequency dynamic convolution module of multiple attention mechanisms for in-depth feature extraction to obtain in-depth feature extraction results.

5. The method for detecting a sound event according to claim 4, characterized in that: The audio signal after data enhancement is input into the CNN module for preliminary feature extraction to obtain initial features, including: The audio signal after data preprocessing is input into the CNN layer for feature extraction; Input the output of the CNN layer to the first normalization layer for normalization; The normalized result is input into the first GLU layer for linear processing; Input the result after linear processing to the first pooling layer for pooling processing; The pooled result is regularized to obtain a regularized result, which is used as the initial feature.

6. The method for detecting a sound event according to claim 4, characterized in that: The initial features are input into the frequency dynamic convolution modules of multiple attention mechanisms for in-depth feature extraction, and in-depth feature extraction results are obtained, including: Step S4.2.1: performing frequency dynamic convolution on the initial features; Step S4.2.2: Input the frequency dynamic convolution result into the second normalization layer for normalization processing, and use the normalization result as the first feature; Step S4.2.3: Input the normalized results into the first convolution kernel and the second convolution kernel respectively, and obtain the first convolution kernel output and the second convolution kernel output respectively; Step S4.2.4: Input the first convolution kernel output into the first convolution layer to obtain the first convolution result; Step S4.2.5: Input the second convolution kernel output into the second convolution layer to obtain a second convolution result; Step S4.2.6: concatenating the first convolution result and the second convolution result to obtain a first concatenated result; Step S4.2.7: input the first splicing result into the second pooling layer for pooling processing; Step S4.2.8: Input the pooling processing result to the third convolutional layer; Step S4.2.9: Input the output of the third convolutional layer into the Sigmoid activation function to obtain the first attention weight and the second attention weight; Step S4.2.10: multiply the first convolution result by the first attention weight to obtain a first multiplication result; Step S4.2.11: multiply the second convolution result by the second attention weight to obtain a second multiplication result; Step S4.2.12: concatenating the first multiplication result and the second multiplication result to obtain a second concatenated result; Step S4.2.13: multiply the second concatenation result by the first feature to obtain a dynamic convolution structure extraction feature result; Step S4.2.14: Input the dynamic feature structure extraction feature result into the second GLU layer to obtain the linear processing result; Step S4.2.15: The linear processing result is pooled through the third pooling layer to obtain the first in-depth feature extraction result; Step S4.2.16: Use the first in-depth feature extraction result as the initial feature and return to step S4.2.1 until the preset number of repetitions is reached, and use the last in-depth feature extraction result as the final in-depth feature extraction result.

7. The method for detecting a sound event according to claim 6, characterized in that: The frequency dynamic convolution is used to perform frequency dimension segmentation and weighting on the initial features, and the calculation formula is as follows: G k =π k (f,x) y k (t,f)=W k *X(t,f)+B Among them, Y(t,f,x) is the frequency dynamic convolution result, t is time, f is frequency, k is the kth base kernel, N is the total number of base kernels, x is the input data, π k is the attention function of the kth base kernel, y k (t,f) is the output of the kth base kernel, G k is the attention weight of the kth base kernel, W k is the input weight matrix, B is the bias matrix, and X(t,f) is the initial feature.

8. The method for detecting a sound event according to claim 1, characterized in that: The extracting context information from the extracted features to obtain the first branch features includes: Input the extracted features into the BGRU module; The output of the BGRU module is input into the attention module, and the output of the attention module is used as the first branch feature.

9. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, a method for detecting a sound event as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Sound event detection method based on double-branch discriminant feature neural network

    CN114881212A

  • Sound event detection method based on lightweight CRNN model

    CN116798448A

  • Audio data processing method and device and storage medium

    CN110503970A

  • Multi-sound-source localization and detection method based on global-local feature recalibration

    CN117612557A

  • Sound event detection method based on cross-model two-stage training

    CN117877516A

Cited By

  • Sound signal periodic feature extraction method, network model training method, storage medium and equipment

    CN121617417A