Sound event detection method and system for collaborative prompting of class distribution and time sequence context
By introducing a method of class distribution and timing context collaborative prompting in the sound event detection system, using the synergy between global distribution and local timing prompting modules to fine-tune the audio pre-training model, the challenges of traditional methods in detection accuracy and model generalization capabilities are solved, and more efficient knowledge transfer and performance improvement are achieved.
Patent Information
- Application Number
- CN202510211179.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Traditional sound event detection methods have significant challenges in detection accuracy and model generalization capabilities, especially in complex scenarios, where it is difficult to effectively utilize the potential performance of audio pre-trained models.
The method of synergistic hint between class distribution and timing context is adopted, and the synergy between the global distribution prompt module and the local timing prompt module is fine-tuned to the audio pre-training model to improve the detection capability of the sound event detection system.
It significantly improves the audio classification and positioning performance of the sound event detection system, enhances the modeling ability of the audio pre-trained model to global distribution information and local timing information, is better than the full fine-tuning method, and reduces the complexity of the model.
Smart Images

Figure CN120048284A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sound event detection in artificial intelligence technology, and particularly relates to a sound event detection method and system that collaboratively prompt class distribution and temporal context. Background Art
[0002] Sound event detection is a key task in the fields of speech processing and audio analysis, and its purpose is to identify the category and start and end times of sound events in audio segments. This technology has important value in multiple practical application scenarios. For example, in a home environment, it can identify events such as smoke alarm sounds, glass breaking sounds, doorbell sounds, etc., providing real-time anomaly alerts and automated responses; in the field of human-computer interaction, this technology can enhance the environmental perception ability of voice assistants, enabling them to detect important sound events and thus make more intelligent responses; in the industrial field, this technology can be used for equipment status monitoring and fault diagnosis.
[0003] Traditional sound event detection methods usually rely on a large amount of high-quality manually labeled data for supervised learning. However, the manual labeling process is not only time-consuming and laborious, but also difficult to meet the needs of diverse data in complex scenarios. To address the challenge of insufficient data, researchers have proposed semi-supervised learning methods that utilize weakly labeled data and unlabeled data, significantly improving the performance of sound event detection systems. However, these methods still face significant challenges in terms of detection accuracy and model generalization ability. In recent years, with the rapid development of deep learning technology, audio-based pre-trained models have shown significant advantages in sound event detection tasks. These pre-trained models can effectively learn and extract high-level feature representations in speech signals through training on large-scale datasets. Mainstream sound event detection systems usually use audio pre-trained models in audio tagging tasks as feature extractors. For example, in the DCASE2023 challenge, the baseline method first introduced the audio pre-trained model BEATs for feature extraction, significantly improving the performance of the sound event detection system. However, this direct transfer method fails to fully exploit the potential performance of audio pre-trained models. In addition, although fully fine-tuning audio pre-trained models can make them better adapt to sound event detection tasks, their high computational resources and storage costs severely limit their application in practical scenarios.
[0004] Prompt optimization methods have gained extensive attention and applications in the fields of natural language processing and computer vision in recent years. This method achieves efficient knowledge transfer by freezing the parameters of the pre-trained model and introducing learnable prompt vectors. However, existing prompt optimization methods mainly focus on the modeling of global features and often neglect the extraction of local temporal features. In the sound event detection task, local temporal features are crucial for the precise localization of sound event boundaries, while global features contribute to the accurate identification of event categories. However, traditional prompt optimization methods are difficult to effectively model these two types of features simultaneously, resulting in the overall performance of sound event detection not reaching the optimal level. Summary of the Invention
[0005] Aiming at the deficiencies in the prior art, the present invention provides a sound event detection method and system that collaboratively prompt class distribution and temporal context. This method uses a collaborative prompt framework for class distribution and temporal context to fine-tune the audio pre-trained model, thereby achieving efficient knowledge transfer. Specifically, a global distribution prompt module is used to model the global class distribution information of the audio sequence, and combined with a local temporal prompt module to explore the local temporal features of the audio sequence, realizing the joint optimization of the audio pre-trained model, and significantly improving the audio classification and localization performance of the sound event detection system.
[0006] The present invention achieves the above technical objectives through the following technical means.
[0007] Sound event detection method that collaboratively prompts class distribution and temporal context:
[0008] Convert the original audio signal into a sequence of signal frames. The pre-trained model branch extracts Mel filter bank features, and the downstream model branch extracts Mel spectrogram features;
[0009] In the pre-trained model branch, a collaborative combination of a local temporal prompt module and a global distribution prompt module is introduced in each layer of the pre-trained model to jointly optimize the pre-trained model. The pre-trained model outputs audio sequence features with enhanced local temporal features and the global distribution prompt module;
[0010] The downstream model branch inputs the Mel spectrogram feature sequence into a convolutional neural network for feature extraction, and performs feature concatenation with the audio sequence features generated by the pre-trained model, and then performs temporal modeling through a recurrent neural network;
[0011] The prediction head processes the output of the downstream model to obtain the frame-level prediction probability, that is, the localization result of the sound event; subsequently, an attention pooling operation is performed on the frame-level prediction probability to generate the sentence-level prediction probability of the downstream model; for the pre-trained model, the global distribution prompt module is processed through a linear layer to obtain the sentence-level prediction probability of the pre-trained model;
[0012] Finally, decision-level fusion is performed on the sentence-level prediction probabilities of the downstream model and the sentence-level prediction probabilities of the pre-trained model to obtain the classification results of the sound events.
[0013] Furthermore, the collaborative combination of the local temporal cue module and the global distribution cue module is as follows: specifically, design a local temporal cue module that interacts with the Mel filter bank feature sequence, design a global distribution cue module and embed it at the front end of the interacted feature sequence.
[0014] Even further, the local temporal cue module is as follows:
[0015]
[0016] where P local represents the local temporal cue vector, represents the local temporal cue vector of the k-th frame, T represents the length of the Mel filter bank feature sequence, d represents the Mel filter bank feature dimension, represents a set of vectors with dimension d, represents a positive integer.
[0017] Even further, the local temporal cue module interacts with the Mel filter bank feature sequence. Specifically, the local temporal cue vector is added to the Mel filter bank sequence A frame by frame to obtain the interacted feature sequence A':
[0018]
[0019] Even further, the global distribution cue module is embedded at the front end of the interacted feature sequence to construct the feature sequence A'' = [P global , A'], where P global represents the global distribution cue vector.
[0020] Even further, the collaborative combination of the local temporal cue module and the global distribution cue module is as follows:
[0021]
[0022] where L i represents the i-th layer of the Transformer encoder, and A i represent the global distribution cue module and the audio sequence features output by the i-th layer respectively, and N represents the number of layers of the Transformer encoder.
[0023] Even further, feature concatenation is performed with the audio sequence features generated by the pre-trained model. Specifically:
[0024]
[0025] Among them, X concat represents the concatenated feature representation, and CNN output represents the features extracted by the convolutional neural network, represents the audio sequence features after adaptive average pooling, and dim = -1 represents the last dimension of the feature dimension.
[0026] Furthermore, decision-level fusion is performed on the sentence-level prediction probabilities of the downstream model and the pre-trained model:
[0027]
[0028] Among them, α is an adjustable hyperparameter, represents the sentence-level prediction probability of the downstream model, represents the sentence-level prediction probability of the pre-trained model.
[0029] A sound event detection system that collaboratively cues class distribution and temporal context includes:
[0030] A signal preprocessing module that converts the original audio signal into a sequence of signal frames;
[0031] An acoustic feature extraction module that extracts Mel filter bank features in the pre-trained model branch and Mel spectrogram features in the downstream model branch;
[0032] An acoustic model modeling module that uses any mainstream audio pre-trained model in the pre-trained model branch and any mainstream deep learning model in the downstream model branch;
[0033] A model prediction module that respectively obtains the sentence-level prediction probabilities of the downstream model and the pre-trained model, and performs fusion to obtain the classification result of the sound event.
[0034] In the above technical solution, the benchmark model of the audio pre-trained model is BEATs, and the benchmark model of the downstream model is a convolutional recurrent neural network.
[0035] The beneficial effects of the present invention are:
[0036] (1) For the first time, the present invention introduces a prompt optimization method from both global and local perspectives into the sound event detection task, and innovatively proposes a collaborative prompt mechanism of class distribution and temporal context to fine-tune the audio pre-trained model. Through the collaborative effect of the global distribution prompt module and the newly designed local temporal prompt module, this mechanism improves the detection ability of the sound time detection system in complex audio environments. Specifically, the global distribution prompt module is embedded at the front end of the audio sequence to capture and model the global event distribution information; at the same time, the local temporal prompt module interacts with the audio sequence, mainly used to capture local temporal dynamic features. By embedding the global distribution prompt module and the local temporal prompt module layer by layer inside the audio pre-trained model and combining the attention mechanism, bidirectional interaction between global distribution information and local temporal information is achieved, effectively enhancing the audio pre-trained model's ability to model global distribution information and improving its recognition ability for complex audio events.
[0037] (2) The present invention designs an efficient fine-tuning method for the audio pre-trained model, breaking through the limitations of traditional prompt optimization methods in the fine-tuning process of the audio pre-trained model, and effectively solving problems such as coarse feature extraction granularity and loss of temporal information existing in the traditional prompt optimization method during the fine-tuning process of the audio pre-trained model. Especially in the sound event detection task, this problem makes it difficult for the model to accurately identify and locate the boundaries of sound events. The present invention optimizes the audio pre-trained model from the perspective of local time series, and innovatively designs a local temporal prompt module. This module captures the fine-grained temporal features of the audio signal by fusing the features of local temporal prompt vectors frame by frame, significantly enhancing the local temporal feature extraction ability of the audio pre-trained model in the sound event detection task, and enhancing the audio pre-training's ability to accurately identify and locate the boundaries of sound events, thereby promoting the knowledge transfer of the audio pre-trained model in the sound event detection task.
[0038] (3) The global distribution prompt module and the local temporal prompt module of the present invention are designed to be lightweight, plug-and-play, and are easy to efficiently fine-tune any audio pre-trained model to complete the sound event detection task. The present invention achieves the best results on multiple widely used audio pre-trained models and is superior to the full fine-tuning method, significantly reducing the model complexity while greatly improving the performance. In addition, the present invention greatly reduces the complexity of model deployment, enabling the sound event detection technology to be more conveniently applied to actual scenarios, providing an efficient and low-resource consumption solution for fields such as intelligent monitoring, environmental perception, and medical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a framework diagram of the sound event detection system based on the collaborative prompt of class distribution and temporal context according to the present invention;
[0040] Figure 2(a) shows the comparison of the visualization results of the sound event detection according to the present invention Figure 1 ;
[0041] Figure 2(b) shows the second comparison diagram of the visualization results of the sound event detection according to the present invention Detailed implementation manners
[0042] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but the protection scope of the present invention is not limited thereto
[0043] As Figure 1As shown, the sound event detection system based on collaborative prompting of class distribution and temporal context in the present invention follows the audio signal data stream transmission process, which is successively signal preprocessing → acoustic feature extraction → acoustic model modeling → model prediction. In the signal preprocessing stage, the system converts the high-dimensional and complex original audio signal into a sequence of signal frames with lower dimensions and continuous time series for subsequent processing. In the acoustic feature extraction stage, the system adopts a two-branch architecture, where the pre-trained model branch extracts Mel-filter Bank features, and the downstream model branch extracts Mel-spectrogram features. The acoustic feature extraction process helps to initially filter out redundant information, thereby improving the modeling efficiency of the acoustic model. In the acoustic model modeling stage, any mainstream audio pre-trained model can be adopted for the pre-trained model branch. The benchmark model of the audio pre-trained model is BEATs (Bidirectional Encoder representations from Audio Transformers, bidirectional encoder representation based on audio Transformer). BEATs consists of 12 layers of Transformer encoders. The downstream model branch can be any mainstream deep learning model, and the benchmark model of the downstream model is a convolutional recurrent neural network (CRNN), which is specifically composed of 7 convolutional blocks and 2 layers of recurrent neural networks. In the pre-trained model branch, the present invention innovatively integrates the global distribution prompting module and the local temporal prompting module into the input layer of the Transformer encoder. The pre-trained model branch outputs the global distribution prompting module containing global class distribution information and audio features that enhance local temporal features. In the downstream model branch, first, the Mel-spectrogram feature sequence is input into the convolutional network for feature extraction, and then it is concatenated with the audio features generated by the pre-trained model. The fused features then undergo temporal modeling through the recurrent neural network. In the model prediction stage, the output of the downstream model (i.e., the recurrent neural network) is processed by the prediction head to obtain the frame-level prediction probability, that is, the localization result of the sound event. Subsequently, the Attention Pooling module is used to calculate the sentence-level prediction probability of the downstream model, and the global distribution prompting module of the pre-trained model branch is processed through a linear layer to obtain the sentence-level prediction probability of the pre-trained model. Finally, decision-level fusion is performed based on the sentence-level prediction probabilities of the two branches to obtain the classification result of the sound event.
[0044] This invention is based on the DCASE2023 benchmark system and uses the DESED dataset, which contains 1,578 weakly labeled data, 10,000 synthetic data, and 14,412 unlabeled audio segments. In addition, 3,470 strongly labeled samples from AudioSet are used. All the data in the above dataset are 10-second audio signals. Moreover, the dataset contains ten types of target sound events in a home scenario, specifically "alarm", "blender", "cat meowing", "dishes", "dog barking", "electric shaver", "frying", "running water", "talking", and "vacuum cleaner".
[0045] This system architecture consists of a pre-trained model branch and a downstream model branch. In the pre-trained model branch, a global distribution prompt module and a local temporal prompt module are introduced to fine-tune the parameters of the audio pre-trained model, thereby effectively completing the knowledge transfer process. The specific process is as follows:
[0046] 1) First, define the audio feature (i.e., Mel filter bank feature) sequence A extracted by the audio pre-trained model, where T represents the length of the audio feature sequence and d represents the audio feature dimension, specifically expressed as:
[0047]
[0048] where a k represents the d-dimensional feature vector of the k-th frame, represents the vector set with dimension d, represents a positive integer.
[0049] 2) In the pre-trained model branch, taking the input of the first layer of the Transformer encoder as an example, define the local temporal prompt module, whose length and feature dimension are the same as the input audio feature sequence A, specifically expressed as:
[0050]
[0051] where, represents the local temporal prompt vector of the k-th frame, aiming to capture the detailed local temporal information in the audio feature sequence. Add the local temporal prompt vector to the audio feature sequence A frame by frame to obtain the interacted feature sequence A':
[0052]
[0053] To further model the global distribution information of the audio sequence, define the global distribution prompt module, specifically expressed as:
[0054]
[0055] where, Denote the l-th global distribution prompt vector as, and M represents the length of the global distribution prompt module. In the present invention, experiments show that the best effect is achieved when M = 10, so M = 10 is set. Embed the global distribution prompt module into the front end of the interacted feature sequence A′ to construct a new feature sequence A″, and use it as the input of the first-layer Transformer encoder, which is specifically expressed as:
[0056] A″ = [P global , A′] (5)
[0057] 3) The audio pre-training model consists of multiple layers of Transformer encoders. Therefore, a combination of a new local temporal prompt module and a global distribution prompt module is introduced in each layer to jointly optimize the audio pre-training model. The specific process is as follows:
[0058]
[0059] Where: L i represents the i-th layer of the Transformer encoder, and A i respectively represent the global distribution prompt module and the audio sequence features output by the i-th layer. N represents the number of layers of the Transformer encoder. In the present invention, N = 12. The output of the audio pre-training model consists of two parts: the global distribution prompt module which contains rich global distribution information; and the audio sequence features A N with significantly enhanced local temporal information.
[0060] 4) In the downstream model branch, define the input audio sequence B of the downstream model, whose length is T′ and the feature dimension is d′, which is specifically expressed as:
[0061]
[0062] Where, b k represents the d′-dimensional feature vector of the k-th frame.
[0063] The convolutional neural network (CNN) is used to model local spatial-frequency domain information and simultaneously perform downsampling to reduce the computational complexity. In the present invention, the convolutional neural network consists of 7 convolutional blocks, and each convolutional block includes a convolutional layer (Conv2D), a batch normalization layer (Batch Normalization), a gated linear unit (GLU), and an average pooling layer (AvgPool2D). Taking the i-th convolutional block as an example, its calculation process is as follows:
[0064] B i ′ -1 = BatchNorm(Conv(B i-1 )) (8)
[0065] B i = Pooling.(W i gate ·B′ i-1 )·σ(W i sigmoid ·B′ i-1 ) / (9)
[0066] where Conv(·) represents the convolution operation, BatchNorm(·) represents batch normalization, W i gate and W i sigmoid both represent the linear transformation weights of the gated linear unit, σ(·) represents the Sigmoid function, B i represents the output of the i-th convolutional block, B i ′ -1 represents the intermediate quantity after convolution and normalization.
[0067] The overall feature extraction operation can be simplified as:
[0068] CNN output = CNN(B) (10)
[0069] 5) In the feature fusion stage, fuse the audio sequence feature A N (step 3) with the feature CNN output (step 4) extracted by the convolutional neural network. However, due to the mismatch problem of the two features in the time dimension, adaptive average pooling is used to perform temporal alignment on A N to make its time step consistent with that of CNN output :
[0070]
[0071] where T CNN represents the time step of CNN output .
[0072] After completing the time alignment, concatenate the features of CNN output and the adjusted A N along the feature dimension (the last dimension, i.e., dim = -1) to fuse the information of the two:
[0073]
[0074] where X concat represents the fused feature representation.
[0075] 6) To perform temporal modeling on the fused features, a recurrent neural network (RNN) is constructed using a two-layer bidirectional gated recurrent unit (BiGRU). BiGRU can capture both forward and backward temporal dependencies simultaneously, thereby improving the sequence modeling ability. Its processing process can be expressed as:
[0076] H = BiGRU(X concat ) (13)
[0077] 7) In the model prediction stage, for the downstream model, a prediction head is used to calculate the frame-level prediction probability for H (step 6) extracted by the recurrent neural network, which is the localization result of the system:
[0078]
[0079] Then, an attention pooling operation is applied to the frame-level prediction probability to obtain the sentence-level prediction probability of the downstream model:
[0080]
[0081] where: W is a trainable attention weight matrix, and softmax(·) is an operator.
[0082] For the audio pre-trained model, the global distribution hint module (step 3) is processed through a linear layer to calculate its sentence-level prediction probability:
[0083]
[0084] Then, the sentence-level prediction probability of the downstream model is fused with the sentence-level prediction probability obtained from the global distribution hint module through decision fusion to obtain the final classification result of the system:
[0085]
[0086] where: α is an adjustable hyperparameter used to control the weights of the downstream model and the audio pre-trained model in the final decision. In the present invention, α is set to 0.5, that is, equal contributions are given to both.
[0087] 8) This study adopts the Mean Teacher framework, where the event localization result and the event classification result obtained in step 7) are predicted by the student model, denoted as Similarly, the prediction results of the Teacher Model are denoted as (y′, Y′), where: y′ represents the event localization prediction result of the Teacher Model; Y′ represents the event classification prediction result of the Teacher Model, which is obtained by making a weighted decision on the sentence-level prediction probability of the downstream model and the sentence-level prediction probability of the audio pre-training model.
[0088] To optimize the model performance, the training objective of the system is to minimize the joint loss function This loss function consists of the supervised loss and the unsupervised loss as follows:
[0089]
[0090] where (y, Y) is the true value of the true labeled data, ω(t) is the joint regularization weight at the t-th step during the training process, and it increases gradually during the training process.
[0091] Specifically, the supervised loss measures the error between the prediction results of the student model and the true annotation, including:
[0092]
[0093] where, is the event localization loss, which is used to measure the deviation between the predicted event location and the true location; is the event classification loss, which is used to measure the error between the predicted class and the true class.
[0094] The unsupervised loss constrains the consistency between the student model and the teacher model, and is defined as follows:
[0095]
[0096] where the event localization consistency loss constrains the event location prediction of the student model to be consistent with the prediction of the teacher model; the event classification consistency loss constrains the classification prediction of the student model to be consistent with the prediction of the teacher model.
[0097] The global distribution prompting module and the local temporal prompting module are two sets of independent and learnable parameterized modules. During the feature extraction process of the audio pre-training model, they respectively model the global class distribution information and the local temporal information, so as to optimize the feature representation of the audio pre-training model and achieve knowledge transfer, thereby enhancing the overall recognition ability of the system. In addition, the global distribution prompting module and the local temporal prompting module are lightweight, efficient and flexible in structure, and can be embedded into any audio pre-training model to achieve efficient knowledge transfer to complete the sound event detection task while maintaining a low computational overhead.
[0098] To verify the effectiveness of the present invention, a visual analysis and comparative experiment of the event localization performance were carried out on the DCASE2023 standard evaluation set. Specifically, the trained sound event detection system based on the collaborative prompting of class distribution and temporal context was used to perform a visual analysis of the localization results for randomly selected test samples (as shown in Figures 2(a) and (b)). To evaluate the contribution of each module, the global distribution prompting module and the local temporal prompting module were removed through ablation experiments to systematically observe the changes in the localization performance. The visualization results were presented in the form of three rows of subgraphs, corresponding to: the true annotation value (the first row of subgraphs), the localization results after removing the global distribution prompting module and the local temporal prompting module (the second row of subgraphs), and the localization results of the complete model (the third row of subgraphs). The horizontal axis represents time, and the vertical axis represents the event categories labeled "C1 to C10". The events corresponding to these categories are: "alarm", "blender", "cat meowing", "dishes", "dog barking", "electric shaver", "frying", "running water", "talking", "vacuum cleaner". Based on the visualization analysis results, the sound event detection system using the collaborative prompting mechanism of class distribution and temporal context can effectively maintain the integrity of the event localization boundary while significantly reducing the false detection rate, achieving the optimal sound event detection performance.
[0099] The described embodiments are the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Without departing from the essential content of the present invention, any obvious improvements, substitutions or variations that those skilled in the art can make all belong to the protection scope of the present invention.
Claims
1. A sound event detection method based on class distribution and temporal context collaborative prompting, characterized in that: The original audio signal is converted into a signal frame sequence, the pre-trained model branch extracts the Mel filter bank features, and the downstream model branch extracts the Mel spectrum features; In the pre-trained model branch, each layer of the pre-trained model introduces a synergistic combination of a local temporal prompt module and a global distribution prompt module to jointly optimize the pre-trained model, wherein the pre-trained model outputs an audio sequence feature enhanced with a local temporal feature and a global distribution prompt module; The downstream model branch inputs the Mel spectrum feature sequence into the convolutional neural network for feature extraction, and concatenates it with the audio sequence features generated by the pre-trained model, and then performs time series modeling through the recurrent neural network. The prediction head processes the output of the downstream model to obtain the frame-level prediction probability, that is, the positioning result of the sound event; then, the frame-level prediction probability is subjected to an attention pooling operation to generate the sentence-level prediction probability of the downstream model; for the pre-trained model, the global distribution prompt module is processed through a linear layer to obtain the sentence-level prediction probability of the pre-trained model; Finally, the sentence-level prediction probability of the downstream model and the sentence-level prediction probability of the pre-trained model are fused at the decision level to obtain the classification result of the sound event.
2. The method for detecting sound events based on the collaborative prompting of class distribution and temporal context according to claim 1, characterized in that: The synergistic combination of the local temporal prompt module and the global distribution prompt module, specifically: designing a local temporal prompt module, the local temporal prompt module interacts with the Mel filter bank feature sequence, and designing a global distribution prompt module and embedding it into the front end of the interactive feature sequence.
3. The method for detecting sound events based on the collaborative prompting of class distribution and temporal context according to claim 2, characterized in that: The local timing prompt module is: Among them, P local represents the local temporal cue vector, represents the local temporal cue vector of the kth frame, T represents the length of the Mel filter bank feature sequence, d represents the Mel filter bank feature dimension, represents a vector set of dimension d, Represents a positive integer.
4. The method for detecting sound events by collaboratively prompting class distribution and temporal context according to claim 3, characterized in that: The local temporal prompt module interacts with the Mel filter bank feature sequence, specifically adding the local temporal prompt vector frame by frame to the Mel filter bank sequence A to obtain the interactive feature sequence A. ′ :
5. The method for detecting sound events by collaboratively prompting class distribution and temporal context according to claim 4, characterized in that: The global distribution prompt module is embedded into the front end of the feature sequence after interaction to construct the feature sequence A ″ =[P global ,A ′ ], where P global Represents the global distribution hint vector.
6. The method for detecting sound events by collaboratively prompting class distribution and temporal context according to claim 5, characterized in that: The synergistic combination of the local temporal prompt module and the global distribution prompt module is as follows: Among them, L i represents the Transformer encoder of the i-th layer, and A i They represent the global distribution prompt module and audio sequence features output by the i-th layer respectively, and N represents the number of layers of the Transformer encoder.
7. The method for detecting sound events by collaboratively prompting class distribution and temporal context according to claim 6, characterized in that: Perform feature concatenation with the audio sequence features generated by the pre-trained model, specifically: Among them, X concat Represents the concatenated feature representation, CNN output represents the features extracted by the convolutional neural network, Represents the audio sequence features after adaptive average pooling, and dim=-1 represents the last dimension of the feature dimension.
8. The method for detecting sound events by collaboratively prompting class distribution and temporal context according to claim 1, characterized in that: Perform decision-level fusion on the sentence-level prediction probability of the downstream model and the sentence-level prediction probability of the pre-trained model: Among them, α is an adjustable hyperparameter, represents the sentence-level prediction probability of the downstream model, Represents the sentence-level prediction probability of the pre-trained model.
9. A system for implementing the method for detecting sound events with coordinated prompting of class distribution and temporal context as described in any one of claims 1 to 8, characterized in that: include: A signal preprocessing module converts the original audio signal into a signal frame sequence; Acoustic feature extraction module, the pre-trained model branch extracts Mel filter bank features, and the downstream model branch extracts Mel spectrum features; Acoustic model modeling module, the pre-trained model branch adopts any mainstream audio pre-trained model, and the downstream model branch adopts any mainstream deep learning model; The model prediction module obtains the sentence-level prediction probability of the downstream model and the sentence-level prediction probability of the pre-trained model respectively, and fuses them to obtain the classification result of the sound event.
10. The system according to claim 9, characterized in that The baseline model for the audio pre-training model is BEATs, and the baseline model for the downstream model is a convolutional recurrent neural network.
Citation Information
Patent Citations
Weak supervision sound event detection method and system based on adaptive hierarchical aggregation
CN114974303A
Audio detection model training method, audio detection method and related device
CN117059075A
Multi-sound-source localization and detection method based on global-local feature recalibration
CN117612557A
Model training and scene recognition method and apparatus, device, and medium
WO2023056889A1
Method and system for weakly-supervised sound event detection by using self-adaptive hierarchical aggregation
WO2023221237A1
Cited By
Human voice data collection method and device, equipment and storage medium
CN121171265A