Lightweight sound detection model training and sound event detection method, device and equipment
Through the teacher-student joint training model, the number of parameters and calculation complexity of the sound detection model is reduced, and the problem of deploying the sound event detection model on devices with limited computing capabilities is solved, achieving efficient and fast and lightweight sound detection.
Patent Information
- Application Number
- CN202510244635.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-13
AI Technical Summary
When existing sound event detection models are deployed on devices with limited computing power, the parameters are too large to be effective.
The lightweight sound detection model training method is adopted. Through the teacher-student joint training model, the student model only uses part of the teacher feature extraction module to reduce the amount of parameters and calculation complexity.
It realizes a lightweight sound detection model that runs quickly on resource-constrained devices, while retaining high detection performance, improving the application efficiency and inference speed of the model.
Smart Images

Figure CN120148486A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of artificial intelligence technology and sound event detection, and particularly to a method, device, equipment and medium for training a lightweight sound detection model and sound event detection. Background Art
[0002] With the rapid development of machine learning and neural network technologies, artificial intelligence has gradually penetrated into all aspects of people's daily lives, and intelligent voice technology is one of the outstanding ones. This technology covers multiple aspects such as sound event detection (SED), speech recognition, and speech enhancement, bringing great convenience to people's lives. Among them, sound event detection is a technology that uses audio signal processing and deep learning technology to detect specific sound event categories in audio signals and accurately locate the start and end times of these events. It can identify various sound events such as baby cries, cat meows, snoring, music, and electric toothbrush sounds, so it has been widely used in many fields such as environmental monitoring, smart home, and autonomous driving.
[0003] In recent years, the introduction of deep neural networks, especially neural network architectures such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), has greatly improved the accuracy of sound event detection. These deep learning models can automatically learn features from raw audio signals, avoiding the cumbersome and limitations of traditional manual feature extraction. Currently, common sound event detection models include CNN-based models, RNN- or long short-term memory (LSTM)-based models, and hybrid network structures such as convolutional recurrent neural networks (CRNNs). These models have achieved remarkable results in sound event detection tasks, promoting the further development of sound event detection technology.
[0004] However, although deep learning models have achieved remarkable results in sound event detection, they still face some challenges in practical applications. Especially in sound event detection in the home scenario, due to the limited computing power of embedded devices or mobile devices, large-scale deep learning models cannot be deployed like servers or the cloud. Therefore, how to achieve model lightweight while ensuring detection performance has become an urgent problem to be solved. Summary of the Invention
[0005] The present application provides a method, device, equipment and medium for training a lightweight sound detection model and detecting sound events, which are used to solve the problem that the high-performance deep learning model adopted by existing sound event detection has too many parameters and cannot be deployed on devices with limited computing power.
[0006] In a first aspect, the present application provides a method for training a lightweight sound detection model, and the method includes:
[0007] Obtain an audio sample set and an original teacher-student joint training model;
[0008] Wherein, the audio sample set contains audio samples, the embedded features corresponding to the audio samples, the true event detection labels of the audio samples, and the true audio classification labels of the audio samples;
[0009] The original teacher-student joint training model includes a teacher model and a student model;
[0010] The teacher model includes a teacher feature extraction module and a teacher task processing module, and the teacher feature extraction module includes N cascaded convolutional layers, where N is an integer greater than 1;
[0011] The student model includes a student feature extraction module and a student task processing module, and the student feature extraction module is the first M convolutional layers of the teacher feature extraction module, where M is a positive integer less than N;
[0012] The initializations of the teacher task processing module and the student task processing module are the same;
[0013] Based on the audio sample set, perform iterative training on the original teacher-student joint training model to determine a lightweight sound detection model based on the trained target student feature extraction module and target student task processing module;
[0014] During each iterative training process:
[0015] Input any audio sample into the teacher feature extraction module, and through the M convolutional layers, obtain the first audio feature extracted by the student feature extraction module, and through the N convolutional layers, obtain the second audio feature extracted by the teacher feature extraction module;
[0016] Through the student task processing module, based on the first audio feature, obtain the student event detection prediction result and the student audio classification prediction result;
[0017] Through the teacher task processing module, based on the second audio sample and the embedded features of the audio sample, obtain the teacher event detection prediction result and the teacher audio classification prediction result;
[0018] Based on the target loss function, adjust the parameters of the original teacher-student joint training model for the current iteration to determine the trained student model as the lightweight voice detection model; wherein, the target loss function includes the following losses: the first loss between the teacher event detection prediction result and the student event detection prediction result respectively and the true event detection label, the second loss between the teacher audio classification prediction result and the student audio classification prediction result respectively and the true audio classification label, the third loss between the student event detection prediction result and the teacher event detection prediction result, and the fourth loss between the student audio classification prediction result and the teacher audio classification prediction result.
[0019] In a second aspect, the present application further provides a voice event detection method based on the above-mentioned model, and the method includes:
[0020] Obtain the audio to be detected;
[0021] Based on the audio to be detected, obtain the event detection result and the audio classification result in the audio to be detected through the pre-trained lightweight voice detection model.
[0022] In a third aspect, the present application further provides a training device for a lightweight voice detection model, and the device includes:
[0023] An acquisition unit, configured to acquire an audio sample set and an original teacher-student joint training model;
[0024] Wherein, the audio sample set contains audio samples, the embedded features corresponding to the audio samples, the true event detection labels of the audio samples, and the true audio classification labels of the audio samples;
[0025] The original teacher-student joint training model includes a teacher model and a student model;
[0026] The teacher model includes a teacher feature extraction module and a teacher task processing module, and the teacher feature extraction module includes N cascaded convolutional layers, where N is an integer greater than 1;
[0027] The student model includes a student feature extraction module and a student task processing module, and the student feature extraction module is the first M convolutional layers of the teacher feature extraction module, where M is a positive integer less than N;
[0028] The initializations of the teacher task processing module and the student task processing module are the same;
[0029] A training unit for iteratively training the original teacher-student joint training model based on the audio sample set to determine a lightweight voice detection model based on the trained target student feature extraction module and target student task processing module;
[0030] During each iterative training process:
[0031] Input any audio sample into the feature extraction module, extract the first audio feature through the student feature extraction module, and extract the second audio feature through the teacher feature extraction module;
[0032] Through the student task processing module, based on the first audio feature, obtain the student event detection prediction result and the student audio classification prediction result;
[0033] Through the teacher task processing module, based on the second audio sample and the embedded feature of the audio sample, obtain the teacher event detection prediction result and the teacher audio classification prediction result;
[0034] Based on the target loss function, adjust the parameters of the original teacher-student joint training model in the current iteration to determine the trained student model as a lightweight voice detection model; wherein, the target loss function includes the following losses: the first loss between the teacher event detection prediction result and the student event detection prediction result and the true event detection label, the second loss between the teacher audio classification prediction result and the student audio classification prediction result and the true audio classification label, the third loss between the student event detection prediction result and the teacher event detection prediction result, and the fourth loss between the student audio classification prediction result and the teacher audio classification prediction result.
[0035] In a fourth aspect, the present application further provides a voice event detection device based on the above-mentioned model, and the device includes:
[0036] An acquisition module for acquiring the audio to be detected;
[0037] A processing module for obtaining the event detection result and the audio classification result in the audio to be detected based on the audio to be detected through the pre-trained lightweight voice detection model.
[0038] In a fifth aspect, the present application provides a computer device, and the computer device includes a processor, and the processor is used to implement the steps of the training method of the lightweight voice detection model as described above or the steps of the voice event detection method as described above when executing a computer program stored in a memory.
[0039] In a sixth aspect, the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps of the training method of the lightweight sound detection model as described above, or implements the steps of the sound event detection method as described above.
[0040] The beneficial effects of the present application are as follows:
[0041] 1. Since the student model only uses a part of the teacher feature extraction module, its parameter quantity and computational complexity are greatly reduced, enabling the student model to run quickly on resource-constrained devices, thus achieving model lightweighting. At the same time, the initialization of the student task processing module and the teacher task processing module is the same, which means that at the beginning of training, the student task processing module and the teacher task processing module have the same parameters, providing a basis for subsequent knowledge distillation and model training, enabling the student model to learn effective knowledge from the teacher model.
[0042] 2. Through the original teacher-student joint training model provided by the present application, the teacher model and the student model can be co-trained. During the task completion process, the teacher model and the student model are not independent, but jointly complete the feature learning and optimization in the task through the shared M convolutional layers and their mutual cooperation, enhancing the learning ability of feature representation during training. This enables the student model to significantly reduce the computational overhead while retaining high performance, improving the application efficiency and inference speed of the model.
[0043] 3. Through the above method, a model can be trained to have both the functions of sound event detection and audio classification. It can not only improve the multi-dimensional understanding and temporal accuracy of the model for audio, but also enhance the generalization ability, computational efficiency, and robustness of the model. This multi-task learning design enables the model to more flexibly handle complex scenarios, improve performance, and optimize the utilization of training resources, especially having significant advantages when dealing with multi-event audio that requires precise positioning.
[0044] 4. The first loss, second loss, third loss, and fourth loss in the target loss function required for training measure the performance of the model from different perspectives. The first loss and the second loss focus on the difference between the prediction results of the original teacher-student joint training model and the true labels. By minimizing these two losses, the original teacher-student joint training model can learn the correct event detection and classification rules. The core idea of the third loss and the fourth loss stems from the knowledge distillation concept, that is, using the output of the teacher model as the target and calculating the difference between the outputs of the teacher model and the student model, so as to better learn the knowledge contained in the teacher model, realizing efficient knowledge transfer between the teacher model and the student model, optimizing the training process of the student model, and ultimately improving the performance of the student model.
[0045] 5. Through the above method, the detection performance of a larger and complete teacher model can be distilled into a smaller student model. The student model can gradually reduce the output difference from the teacher model during the training process, thereby better learning the knowledge contained in the teacher model, achieving efficient knowledge transfer between the teacher model and the student model, optimizing the training process of the student model, and without the participation of an additional second model. While ensuring the sound event detection performance of the student model, the size of the student model is reduced to facilitate the completion of the sound event detection task under limited resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0047] Figure 1 Schematic diagram of the training process of a lightweight sound detection model provided by an embodiment of the present application;
[0048] Figure 2 Schematic diagram of the training flow of a lightweight sound detection model provided by an embodiment of the present application;
[0049] Figure 3 Schematic diagram of the structure of an original teacher-student joint training model provided by an embodiment of the present application;
[0050] Figure 4 Experimental result diagram of the training method of the lightweight sound detection model provided by an embodiment of the present application;
[0051] Figure 5 Schematic diagram of the process of a sound event detection provided by an embodiment of the present application;
[0052] Figure 6 Schematic diagram of the structure of a training device for a lightweight sound detection model provided by an embodiment of the present application;
[0053] Figure 7 Schematic diagram of the structure of a sound event detection device provided by an embodiment of the present application;
[0054] Figure 8 Schematic diagram of the structure of a computer device provided by an alternative embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of them. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0056] In recent years, deep neural networks have achieved remarkable results in the task of sound event detection, promoting the further development of sound event detection technology. However, although deep learning models have achieved remarkable results in sound event detection, they still face some challenges in practical applications. Especially in the sound event detection in the home scenario, due to the limited computing power of embedded devices or mobile devices, large-scale deep learning models cannot be deployed like servers or the cloud. Therefore, how to achieve model lightweight while ensuring detection performance has become an urgent problem to be solved.
[0057] To solve this problem, the knowledge distillation technique (Knowledge Distillation, KD) is introduced into the sound event detection task. Knowledge distillation is a method of learning important knowledge from a large model (teacher model) and transferring this knowledge to a smaller model (student model) to improve its performance. In the sound event detection task, first, a high-performance teacher model is trained to detect sound events, and then the prediction output of the teacher model is used as soft labels to guide the learning of the student model. This method can improve the accuracy and robustness of the student model in the sound event detection task while achieving model lightweight.
[0058] However, the application of knowledge distillation in the sound event detection task also faces some challenges. First, the teacher and student models cannot be optimized simultaneously. The student model can only learn through the soft labels of the teacher model, which limits the diversity of distilled knowledge and the learning effect of the student model. Second, the selection of the distillation strategy is crucial. Different distillation methods may lead to different performance results. An inappropriate strategy may cause the student model to fail to effectively learn the key knowledge of the teacher model and cannot achieve the optimal performance. In addition, since the training involves two models, the computational overhead and training time are relatively long, which may be unacceptable in an environment with limited resources.
[0059] In summary, in the research process of sound event detection algorithms, there are mainly two major problems: resource constraints and challenges in the application of knowledge distillation. How to achieve model lightweight while ensuring detection performance, and how to design a more reasonable distillation strategy to optimize the knowledge distillation effect and improve the sound event detection performance have become the focus of current research.
[0060] To solve the above problems, the present application provides a method, apparatus, device, and medium for training a lightweight sound detection model and detecting sound events.
[0061] Embodiment 1:
[0062] The present application provides a method for training a lightweight sound detection model. Figure 1 FIG. is a schematic diagram of the training process of a lightweight sound detection model provided by an embodiment of the present application. The process includes:
[0063] S101: Obtain an audio sample set and an original teacher-student joint training model; wherein, the audio sample set contains audio samples, the embedding features corresponding to the audio samples, the true event detection labels of the audio samples, and the true audio classification labels of the audio samples; the original teacher-student joint training model includes a teacher model and a student model; the teacher model includes a teacher feature extraction module and a teacher task processing module, the teacher feature extraction module includes N cascaded convolutional layers, and N is an integer greater than 1; the student model includes a student feature extraction module and a student task processing module, the student feature extraction module is the first M convolutional layers of the teacher feature extraction module, and M is a positive integer less than N; the initializations of the teacher task processing module and the student task processing module are the same.
[0064] In the present application, the method for training the lightweight sound detection model is applied to a computer device, which can be an intelligent terminal, such as a computer, a robot, etc., or a server, such as an application server, a business server, etc.
[0065] In order to obtain a high-performance lightweight sound detection model, the present application pre-collects an audio sample set and conducts model training based on this, and finally obtains a trained lightweight sound detection model. The audio sample set is the basic data source for the entire training process, containing a large number of audio samples with various formats, such as common WAV, MP3, etc.
[0066] In one example, to enable the trained model to have good generalization ability and accurately identify sound events in various scenarios, the construction of the audio sample set needs to be extensive and diverse. For example, in terms of the scenario dimension, it can cover common scenarios and special scenarios in daily life. Common daily life scenarios include kitchen cooking sounds, living room TV playing sounds, bedroom alarm clock sounds, etc. in a home environment; public environment scenarios include mall noise, subway station announcements and train running sounds, school class bell sounds and student reading sounds, etc.; special scenarios involve factory workshop machine roars, hospital medical equipment operation sounds and ambulance alarm sounds, etc. In terms of the sound type dimension, it can include various different types of sounds. Voice types include human voices of different genders, ages, and accents, covering forms such as daily conversations, speeches, singing, etc.; natural sound types include wind sounds, rain sounds, bird chirping sounds, ocean wave sounds, etc.; artificial sound types can include car horn sounds, electrical appliance operation sounds, alarm sounds, etc. Comprehensively collecting these different types of sounds can enable the model to learn richer sound features, thereby improving its recognition ability in different scenarios.
[0067] In terms of the number of samples, to ensure the sufficiency of model training, sufficient samples should be collected as much as possible for each scenario and sound type. For common scenarios and sound types, more samples can be collected to enhance the model's learning and adaptation abilities; for relatively rare scenarios and sound types, a certain number of samples should also be collected as much as possible to avoid recognition difficulties when the model encounters these situations.
[0068] Among them, the audio samples can be obtained from one or more of the following sources: obtained from public datasets (such as ESC–50, UrbanSound8K, FreeSound, etc.), obtained from audio data collected in the actual working environment, and purchased from professional sample collection companies.
[0069] In some possible implementation manners, after collecting the original audio samples, a series of preprocessing operations can be performed on them to facilitate subsequent model processing and reduce the computational amount. Among them, the preprocessing operations can include, but are not limited to, one or more of the following: unifying the audio format, unifying the sampling rate, unifying the bit rate, and extracting acoustic features (such as Mel Frequency Cepstral Coefficients (MFCC), Linear Predictive Cepstral Coefficients (LPCC), spectral features, etc.).
[0070] In this application, for each acquired audio sample, accurate annotation is also required. The annotation content includes the embedded features corresponding to the audio sample, the real event detection label, and the real audio classification label. The embedded feature is a vector representation obtained by performing deep feature extraction on the audio. It contains key feature information in the audio, and this key feature information can be further used for various downstream tasks, such as sound audio classification and audio annotation. For example, the embedded features of the audio sample are obtained through the BEATs model pre-trained on the Audioset dataset, etc. The real event detection label is used to mark the start and end time points of each sound event in the audio sample. The real audio classification label, on the other hand, tags each segment or the whole of the audio sample to indicate what types of sound events are included in the audio sample.
[0071] In order to reduce the size of the model while maintaining good detection ability in a resource-constrained environment, this application provides an original teacher-student joint training model. The original teacher-student joint training model includes a teacher model and a student model. The teacher model contains a teacher feature extraction module and a teacher task processing module. The teacher feature extraction module is used to extract deep and complex audio features from the input audio, and the teacher task processing module is used to perform downstream tasks based on the audio features extracted by the teacher feature extraction module, so as to achieve event detection and classification of the audio. The teacher feature extraction module contains N cascaded convolutional layers. Among them, N is an integer greater than 1. Through these N cascaded convolutional layers, the teacher model can extract audio features. Through the progressive layering of multiple convolutional layers, the teacher model can gradually abstract high-level semantic features from the low-level features of the audio, so as to perform event detection and classification more accurately. The student model also includes a student feature extraction module and a student task processing module. The student feature extraction module is used to extract deep and complex audio features from the input audio, and the student task processing module is used to perform downstream tasks based on the audio features extracted by the student feature extraction module, so as to achieve event detection and classification of the audio. The student feature extraction module is the first M convolutional layers of the teacher feature extraction module. Among them, M is a positive integer less than N. Since the student model only uses a part of the teacher feature extraction module, its number of parameters and computational complexity are greatly reduced, enabling the student model to run quickly on resource-constrained devices, thus realizing the lightweight of the model. At the same time, the initialization of the student task processing module and the teacher task processing module is the same, which means that at the beginning of training, the student task processing module and the teacher task processing module have the same parameters, facilitating subsequent knowledge distillation and model training. For example, both the student task processing module and the teacher task processing module are multi-layer recurrent neural networks (such as bidirectional gated recurrent units (BGRU)), etc., which provides a basis for subsequent knowledge distillation and model training, enabling the student model to learn effective knowledge from the teacher model.
[0072] Through this original teacher-student joint training model, the teacher model and the student model can be co-trained. In the process of completing the task, the teacher model and the student model are not independent. Instead, through the shared M convolutional layers and their mutual cooperation, they jointly complete the feature learning and optimization in the task, enhancing the learning ability of feature representation during the training process. This enables the student model to significantly reduce the computational overhead while maintaining high performance, improving the application efficiency and inference speed of the model.
[0073] S102: Based on the audio sample set, perform iterative training on the original teacher-student joint training model to determine a lightweight voice detection model based on the trained target student feature extraction module and target student task processing module; in each iterative training process: Input any audio sample into the teacher feature extraction module, and through the M convolutional layers, obtain the first audio feature extracted by the student feature extraction module, and through the N convolutional layers, obtain the second audio feature extracted by the teacher feature extraction module; Through the student task processing module, based on the first audio feature, obtain the student event detection prediction result and the student audio classification prediction result; Through the teacher task processing module, based on the second audio sample and the embedded feature of the audio sample, obtain the teacher event detection prediction result and the teacher audio classification prediction result; Based on the target loss function, adjust the parameters of the original teacher-student joint training model in the current iteration to determine the trained student model as the lightweight voice detection model; wherein, the target loss function includes the following losses: the first loss between the teacher event detection prediction result and the student event detection prediction result and the true event detection label, the second loss between the teacher audio classification prediction result and the student audio classification prediction result and the true audio classification label, the third loss between the student event detection prediction result and the teacher event detection prediction result, and the fourth loss between the student audio classification prediction result and the teacher audio classification prediction result.
[0074] After obtaining the audio sample set and the original teacher-student joint training model based on the above embodiments, the original teacher-student joint training model can be iteratively trained through this audio sample set. In each iterative training process, multiple steps are required. The following describes any iterative process:
[0075] Obtain any audio sample from the audio sample set, and input the audio sample into the teacher feature extraction module in the original teacher-student joint training model. That is, input the audio sample into the student feature extraction module and the teacher processing module simultaneously. Through multiple convolutional layers included in the teacher feature extraction module, the input audio sample is processed layer by layer. At the Mth convolutional layer, the extracted audio feature (denoted as the first audio feature) can be output, and this first audio feature is the audio feature extracted by the student extraction module. At the Nth convolutional layer, the extracted audio feature (denoted as the second audio feature) can be output, and this second audio feature is the audio feature extracted by the teacher extraction module.
[0076] Input the obtained first audio feature into the student task processing module in the original teacher-student joint training model. Through this student task processing module, calculations are performed based on the first audio feature, and the student event detection prediction result and the student audio classification prediction result can be obtained.
[0077] Meanwhile, input the obtained second audio feature and the embedding feature of the audio sample into the teacher task processing module in the original teacher-student joint training model. The teacher task processing module combines the second audio feature and the embedding feature of the audio sample to obtain the teacher event detection prediction result and the teacher audio classification prediction result. Among them, the embedding feature of the audio sample can provide additional semantic information for the teacher model to help the teacher model make more accurate predictions.
[0078] In an example, the teacher model further includes an alignment and splicing module, and the alignment and splicing module is used to align the second audio feature output by the teacher feature extraction module with the embedding feature in the time dimension; splice the aligned second audio feature and the aligned embedding feature in the feature dimension to obtain a spliced feature vector;
[0079] Through the teacher task processing module, based on the second audio sample and the embedding feature of the audio sample, obtaining the teacher event detection prediction result and the teacher audio classification prediction result includes:
[0080] Through the teacher task processing module, based on the spliced feature vector, obtain the teacher event detection prediction result and the teacher audio classification prediction result.
[0081] Considering that after obtaining the second audio feature output by the teacher feature extraction module and the embedding feature corresponding to the audio sample, there may be a situation where they are inconsistent in the time dimension, so it is necessary to perform an alignment operation in the time dimension. Based on this, in this application, the teacher model may further include an alignment and splicing module, which is used to align the second audio feature output by the teacher feature extraction module with the embedding feature in the time dimension; splice the aligned second audio feature and the aligned embedding feature in the feature dimension to obtain a spliced feature vector. Among them, the alignment operation can adopt any one of the following: interpolation, downsampling, adaptive average pooling. For example, for the second audio feature output by the teacher feature extraction module, the alignment and splicing module inputs it into an adaptive average pooling layer, and this adaptive average pooling layer will process the second audio feature according to the set output time dimension T. Specifically, the adaptive average pooling will divide the second audio feature in the time dimension, calculate the average value of the feature values in each divided area, so as to obtain the aligned second audio feature with the time dimension of T. Similarly, for the embedding feature of the audio sample, it is also input into another adaptive average pooling layer (the output time dimension of this layer is also set to T) to adjust the embedding feature in the time dimension and obtain the aligned embedding feature with the time dimension of T. After completing the alignment in the time dimension, next, the aligned second audio feature and the aligned embedding feature are spliced in the feature dimension, so that the spliced feature vector contains richer audio feature information. The second audio feature mainly captures the time domain and frequency domain features of the audio, while the embedding feature contains the semantic information of the audio. Through splicing, the advantages of these two features can be fully utilized to provide more comprehensive information for subsequent event detection and classification tasks, thereby improving the prediction accuracy of the teacher model. Then the alignment and splicing module inputs the spliced feature vector into the teacher task processing module. The teacher task processing module can obtain the teacher event detection prediction result and the teacher audio classification prediction result based on the spliced feature vector.
[0082] Through the above method, a model can be trained to have two functions of sound event detection (SED) and audio classification (Audio-tagging) at the same time, which can not only improve the multi-dimensional understanding and time accuracy of the model for audio, but also enhance the generalization ability, computational efficiency and robustness of the model. This design of multi-task learning can make the model more flexible to handle complex scenarios, improve performance, and optimize the utilization of training resources, especially having significant advantages when dealing with multi-event audio that requires precise positioning.
[0083] After obtaining the student event detection prediction result, the student audio classification prediction result, the teacher event detection prediction result, and the teacher audio classification prediction result based on the above embodiments, the loss value can be determined according to the target loss function. Based on this loss value, the parameters of the original teacher-student joint training model in the current iteration are updated.
[0084] Among them, the target loss function includes the following losses: the first loss (such as binary cross-entropy loss, mean squared error loss, etc.) between the teacher event detection prediction result and the student event detection prediction result and the true event detection label, and the second loss (such as binary cross-entropy loss, mean squared error loss, etc.) between the teacher audio classification prediction result and the student audio classification prediction result and the true audio classification label. In addition, it also includes the third loss (such as mean squared error loss, Kullback-Leibler divergence, cosine similarity loss, etc.) between the student event detection prediction result and the teacher event detection prediction result, and the fourth loss (such as mean squared error loss, Kullback-Leibler divergence, cosine similarity loss, etc.) between the student audio classification prediction result and the teacher audio classification prediction result.
[0085] The first loss, the second loss, the third loss, and the fourth loss in the target loss function measure the performance of the model from different perspectives. The first loss and the second loss focus on the difference between the prediction results of the original teacher-student joint training model and the true labels. By minimizing these two losses, the original teacher-student joint training model can learn the correct event detection and classification rules.
[0086] The core idea of the third loss and the fourth loss stems from the knowledge distillation concept, that is, using the output of the teacher model as the target and calculating the difference between the outputs of the teacher model and the student model to guide the learning process of the student model. Specifically, the output of the teacher model contains rich knowledge information, covering high-level feature representations and learned complex patterns. The student model gradually approaches the prediction result of the teacher model by mimicking the output of the teacher model. When calculating the loss, the self-knowledge distillation loss function is adopted here as an index to measure the inconsistency between the outputs of the teacher model and the student model. This method can effectively quantify the gap between the output of the student model and the output of the teacher model, and update the parameters of the student model through backpropagation to reduce this difference. By minimizing the third loss and the fourth loss, the student model can gradually reduce the output difference from the teacher model during the training process, thereby better learning the knowledge contained in the teacher model, achieving efficient knowledge transfer between the teacher model and the student model, optimizing the training process of the student model, and finally improving the performance of the student model.
[0087] Since the audio sample set for training the original teacher-student joint training model contains a large number of audio samples, and the above operations are performed on each audio sample, when the preset convergence condition is met, the training of the original teacher-student joint training model is completed, and thus a trained student model is obtained. The trained student model is determined as the lightweight sound detection model.
[0088] In each iteration, an optimization algorithm (such as stochastic gradient descent, Adam, etc.) can be used to update the parameters of the model according to the gradient of the objective loss function. The choice of the optimization algorithm will affect the training speed and convergence effect of the model, and different optimization algorithms have different characteristics and applicable scenarios. For example, the stochastic gradient descent algorithm is simple to calculate, but the convergence speed may be slow; the Adam algorithm combines the advantages of momentum and adaptive learning rate and can converge faster in most cases.
[0089] Among them, meeting the preset convergence condition can be that the sum of the loss values corresponding to the audio samples in the current iteration audio sample set reaches the minimum value or tends to be stable, or the number of iterations for training the original teacher-student joint training model reaches the set maximum number of iterations, etc. It can be flexibly set in specific implementation and is not specifically limited here.
[0090] As a possible implementation manner, when training the model, the audio samples in the audio sample set can be divided into training samples, validation samples, and test samples. First, the original teacher-student joint training model is trained based on the training samples, then the reliability of the above-trained teacher-student joint training model is verified based on the validation samples, and the performance of the trained student model is tested based on the test samples.
[0091] Through the above method, the detection performance of the larger complete teacher model can be distilled into a smaller student model. The student model can gradually reduce the output difference with the teacher model during the training process, so as to better learn the knowledge contained in the teacher model, realizing the efficient knowledge transfer between the teacher model and the student model, optimizing the training process of the student model, and without the participation of an additional second model. While ensuring the sound event detection performance of the student model, the size of the student model is reduced to facilitate the completion of the sound event detection task under limited resources.
[0092] The beneficial effects of this application are as follows:
[0093] 1. Since the student model only uses a part of the teacher feature extraction module, its parameter quantity and computational complexity are greatly reduced, enabling the student model to run quickly on resource-constrained devices, thus achieving model lightweighting. At the same time, the initialization of the student task processing module is the same as that of the teacher task processing module, which means that at the beginning of training, the student task processing module and the teacher task processing module have the same parameters, providing a basis for subsequent knowledge distillation and model training, and enabling the student model to learn effective knowledge from the teacher model.
[0094] 2. Through the original teacher-student joint training model provided by this application, the teacher model and the student model can be co-trained. During the task completion process, the teacher model and the student model are not independent, but jointly complete the feature learning and optimization in the task through the shared M convolutional layers and their mutual cooperation, enhancing the learning ability of feature representation during training. This enables the student model to significantly reduce the computational overhead while retaining high performance, improving the application efficiency and inference speed of the model.
[0095] 3. Through the above method, a model can be trained to have two functions of sound event detection and audio classification at the same time, which can not only improve the multi-dimensional understanding and temporal accuracy of the model for audio, but also enhance the generalization ability, computational efficiency and robustness of the model. This design of multi-task learning can make the model more flexible in dealing with complex scenarios, improve performance, and optimize the utilization of training resources, especially having significant advantages when dealing with multi-event audio that requires precise positioning.
[0096] 4. The first loss, the second loss, the third loss and the fourth loss in the target loss function required for training measure the performance of the model from different perspectives. The first loss and the second loss focus on the difference between the prediction results of the original teacher-student joint training model and the true labels. By minimizing these two losses, the original teacher-student joint training model can learn the correct event detection and classification rules. The core idea of the third loss and the fourth loss comes from the knowledge distillation concept, that is, using the output of the teacher model as the target and calculating the difference between the outputs of the teacher model and the student model, so as to better learn the knowledge contained in the teacher model, realizing the efficient knowledge transfer between the teacher model and the student model, optimizing the training process of the student model, and ultimately improving the performance of the student model.
[0097] 5. Through the above method, the detection performance of a larger and complete teacher model can be distilled into a smaller student model. The student model can gradually reduce the output difference from the teacher model during the training process, thereby better learning the knowledge contained in the teacher model, achieving efficient knowledge transfer between the teacher model and the student model, optimizing the training process of the student model, and without the participation of an additional second model. While ensuring the sound event detection performance of the student model, the size of the student model is reduced to facilitate the completion of the sound event detection task under limited resources.
[0098] Embodiment 2:
[0099] In order to accurately improve the event detection performance of the student model, based on the above embodiment, in this application, the third loss for the audio sample is determined in the following manner:
[0100] Obtain the sum of squared differences between the student event detection prediction results and the teacher event detection prediction results for each frame and each event category;
[0101] Determine the product between the number of frames of the audio sample and the total number of pre-configured event categories;
[0102] Determine the third loss according to the quotient between the sum of squared differences and the product.
[0103] In this application, the student event detection prediction results and the teacher event detection prediction results can be represented by probability values or binary values. For example, in binary representation, "1" may indicate that a certain event is detected, and "0" indicates that the event is not detected; while the probability value represents the likelihood of detecting an event within the range of [0,1]. The following uses a specific example to illustrate the method for determining the third loss of the audio sample provided in this application:
[0104] Assume that the shape of the student event detection prediction results is [T, C], where T represents the number of frames of the audio sample, and C represents the total number of pre-configured event categories; the shape of the teacher event detection prediction results is also [T, C]. This means that for each frame of audio, there are prediction values for C event categories.
[0105] For each position (t, j) in the i-th audio sample, where t represents the frame number (0 < t ≤ T) and j represents the event category (0 < j ≤ C), calculate the difference between the student event detection prediction result y(i,t,j) and the teacher event detection prediction result and then square the difference. Then, sum the squares of the differences at all positions to obtain the sum of squared differences.
[0106] Obtain the product between the number of frames F of the i-th audio sample and the total number C of pre-configured event categories.
[0107] Then, determine the third loss according to the quotient between the sum of squared differences and the product. For example, directly use the determined quotient as the third loss. Another example is to process the determined quotient through a pre-configured mathematical algorithm (such as regularization, linear transformation, etc.) and determine the processed value as the third loss.
[0108] The goal of the third loss is to minimize the per-frame per-category error of the model, that is, to ensure that the predicted values of the student model for each event category in each frame can match the predicted values of the teacher model as accurately as possible.
[0109] In a possible implementation manner, the third loss of the audio sample can be expressed by the following formula:
[0110]
[0111] where i represents the i-th audio sample, T represents the number of frames of the audio sample, C represents the total number of event categories, y(i,t,c) represents the predicted result of the student event detection at the t-th frame and the c-th sound time of the i-th audio sample, represents the predicted result of the teacher event detection at the t-th frame and the c-th sound time of the i-th audio sample.
[0112] Example 3:
[0113] In order to further accurately improve the audio classification performance of the student model, based on the above embodiments, in this application, the fourth loss for the audio sample is determined in the following manner:
[0114] Determine the square of the difference between the student audio classification prediction result and the teacher audio classification prediction result.
[0115] In this application, the student audio classification prediction result and the teacher audio classification prediction result can be represented by probability values or binary values. For example, in a simple binary classification problem, the prediction result may be 0 or 1, indicating that the audio belongs to a certain class or does not belong to a certain class; in a multi-classification problem, it may be a binary vector, and each element represents whether it belongs to the corresponding class. The probability value represents the likelihood of belonging to the sound event class within the range of [0,1]. The following uses a specific example to illustrate the method for determining the fourth loss of the audio sample provided in this application:
[0116] Suppose the student audio classification prediction result is represented as a vector R = {r 1 , r 2 , r 3 …r C}, and the teacher audio classification prediction result is represented as a vector where C is the total number of event categories of the sound events.
[0117] For the i-th audio sample, calculate the square of the difference between the vector R and each pair of corresponding elements in. Then sum the squares of all the differences. Based on the determined sum, determine the fourth loss of the i-th audio sample. For example, directly use the determined sum as the fourth loss. Another example is to process the determined sum through a pre-configured mathematical algorithm (such as regularization, linear transformation, etc.), and determine the processed value as the fourth loss.
[0118] In a possible implementation manner, the fourth loss of the audio sample can be represented by the following formula:
[0119]
[0120] where i represents the i-th audio sample, R i represents the student audio classification prediction result of the i-th audio sample, represents the teacher audio classification prediction result, and L 4 represents the fourth loss.
[0121] Example 4:
[0122] In order to further accurately improve the overall performance of the student model, based on the above embodiments, in this application, the target loss function is represented by the following formula:
[0123]
[0124] where L total represents the target loss function, Q represents the weight values corresponding to different losses, represents the first loss between the student event detection prediction result and the true event detection label, represents the first loss between the teacher event detection prediction result and the true event detection label, represents the second loss between the student audio classification prediction result and the true audio classification label, represents the second loss between the teacher audio classification prediction result and the true audio classification label, and L 3 represents the third loss, and L 4 represents the fourth loss.
[0125] In the training of a lightweight sound detection model, the target loss function contains multiple loss terms, namely the first loss of the student model, the first loss of the teacher model, the second loss of the student model, the second loss of the teacher model, the third loss, and the fourth loss. To more flexibly control the model training process and improve model performance, corresponding weights can be assigned to different loss terms. These weights can be adjusted according to actual needs, and they can be completely different, completely the same, or partially different. For example, and are both the first numerical values, and Q 3 and Q 4 are both the second numerical values, and the first numerical value and the second numerical value are different. Considering that different loss terms have different roles in model training. For example, the first loss and the second loss mainly measure the difference between the model prediction result and the true label, which helps the model learn the correct event detection and classification rules; the third loss and the fourth loss focus on the student model learning from the teacher model, making the prediction result of the student model closer to that of the teacher model. By assigning weights to different loss terms, the influence degree of each loss term on model parameter update can be flexibly adjusted according to specific application scenarios and training objectives.
[0126] Exemplarily, in some application scenarios with extremely high requirements for real-time performance, more attention may be paid to the fast learning and response capabilities of the student model. At this time, relatively large weights can be assigned to the first loss and the second loss of the student model. For example, the weight of the first loss of the student model is set to 0.3, and the weight of the second loss of the student model is set to 0.3. While the weights of the first loss and the second loss of the teacher model can be appropriately reduced, set to 0.1 and 0.1 respectively. The third loss and the fourth loss are used for knowledge distillation to help the student model learn the knowledge of the teacher model, and their weights can both be set to 0.1. Such a weight assignment method can enable the model to focus on optimizing the performance of the student model during training, making it reach a relatively high detection and classification accuracy faster to meet the real-time requirements.
[0127] Another exemplarily, if the application scenario emphasizes more on the overall generalization ability of the model and hopes that the model can have good performance in various different sound environments and event types, then it can be considered to assign the same weight to all loss terms, all set to 1. This method can ensure that each loss term has the same influence during the model parameter update process, enabling the model to learn knowledge in all aspects evenly, avoiding a decline in performance in other aspects due to over-focusing on a certain loss term, thereby improving the generalization ability of the model and enabling it to play a stable role in different actual scenarios.
[0128] In another example, in some specific application scenarios, it may occur that the importance of some loss terms is relatively close, while the difference in importance from other loss terms is relatively large. For example, in a scenario where high accuracy in sound event classification is required, but certain real-time requirements for event detection are also needed, the second loss of the student model and the second loss of the teacher model can be regarded as a group because they are mainly related to the classification task. The same and relatively large weights are assigned to these two loss terms, such as both set to 1. The first loss and the third loss of the student model are related to event detection and the student model learning detection knowledge from the teacher model, and their weights are both set to 0.2. The first loss and the fourth loss of the teacher model are of relatively lower importance, and their weights are both set to 0.1. Through this partial different weight assignment method, the training direction of the model can be more accurately guided according to the requirements of the application scenario.
[0129] When determining the weight assignment scheme, one or more of the following methods can be adopted:
[0130] 1. Experimental comparison method. First, based on the preliminary analysis of the application scenario, several different weight assignment schemes are formulated. Then, using the same audio sample set and the original teacher-student joint training model, multiple rounds of training are carried out according to different weight schemes respectively. After each round of training, the performance of the model is evaluated using the validation set, and the evaluation metrics can include the accuracy, recall rate of event detection, the accuracy, F1 value, etc. of event classification. By comparing the performance of the model on the validation set under different schemes, the weight assignment scheme with the best performance is selected as the final scheme.
[0131] 2. Dynamic adjustment during training. In the initial stage of training, the model's understanding of the data is still relatively limited, and it may be necessary to pay more attention to the matching degree between the model and the true labels. Therefore, the weights of the first loss and the second loss of the student model and the teacher model can be appropriately increased. As the training progresses, when the model already has a certain foundation, the weights of the third loss and the fourth loss can be gradually increased to allow the student model to learn more knowledge from the teacher model and further improve the performance. For example, a threshold of the number of training rounds can be set. When the number of training rounds is less than this threshold, the weight of the first loss of the student model is 0.3. After the number of training rounds exceeds this threshold, its weight is adjusted to 0.2, and at the same time, the weight of the third loss is adjusted from 0.1 to 0.2. This way of dynamically adjusting weights can enable the model to better adapt to the learning needs at different training stages and improve the training effect.
[0132] In a possible implementation, if the weights corresponding to the third loss and the fourth loss remain unchanged throughout the training process, it may cause the model to overly rely on the knowledge of the teacher model, resulting in overfitting. Therefore, the weights corresponding to the third loss and the fourth loss can be dynamically adjusted only to a certain extent to avoid this situation, enabling the model to fully utilize its own learning ability while learning the knowledge of the teacher model and improving the generalization ability of the model. Exemplarily, when 3 and 4 are both the same self-knowledge distillation weight parameter value, the self-knowledge distillation weight parameter value is represented by the following formula:
[0133]
[0134] where α is the self-knowledge distillation weight parameter value, K represents the constant maximum value of the loss weight between the teacher and the student, β is an exponential parameter used to control the speed and amplitude of weight adjustment, s is the current training step, G represents the number of rounds of warm-up training, and Z is the length of each training round. K is an empirical value that determines the maximum influence that the self-knowledge distillation loss can achieve in the target loss function. If K takes a larger value, then in the appropriate training stage, the influence of the self-knowledge distillation loss on the model parameter update will be greater, and the model will pay more attention to the student model learning from the teacher model; conversely, if K takes a smaller value, the influence of the self-knowledge distillation loss is relatively limited. β affects the rate and range of change of the weight with the training steps. When β is larger, the change of the weight with the training steps will be more drastic, that is, the speed of weight adjustment is faster and the amplitude is larger; when β is smaller, the change of the weight will be relatively gentle, and both the adjustment speed and amplitude will become smaller. For example, in the initial stage of training, if β is larger, the weight may rise or fall rapidly; if β is smaller, the change of the weight will be relatively slow. s reflects the current training stage of the model. As the training progresses, the value of s continuously increases. In the initial stage of training, s is smaller, which will cause the value of the exponential function to be larger, and the self-knowledge distillation weight parameter value may be close to K; as the training step s increases, the value of the exponential function will become smaller, and the self-knowledge distillation weight parameter value will also gradually decrease. G and Z affect the relative position of s in the entire training process. For example, if G or Z increases, then the entire training cycle becomes longer, and the proportion of the same training step s in the entire training process will become smaller, thereby affecting the change of the self-knowledge distillation weight parameter value.
[0135] It should be noted that the values of the above parameters can be determined in various ways. For example, they can be determined through empirical values, through grid search, etc., and no specific limitation is made here.
[0136] 3. Adjust the weights in combination with the convergence of the model. During the training process, monitor the changes in the target loss function and the specific values of each loss term in real time. If it is found that the decline rate of a certain loss term is too slow during the training process, it indicates that the model's learning effect in this aspect is not good, and the weight of this loss term can be appropriately increased to strengthen the optimization in this aspect; conversely, if a certain loss term has quickly converged to a small value, it indicates that the model has achieved good results in this aspect, and its weight can be appropriately reduced to avoid overfitting caused by over-optimization. For example, when it is found that the second loss of the student model has a small decline amplitude in consecutive multiple rounds of training, its weight can be increased from 0.2 to 0.3 to prompt the model to pay more attention to the accuracy of event classification.
[0137] Example 5:
[0138] The training method of the lightweight sound detection model provided by the present application will be described below through specific examples. Figure 2 It is a schematic diagram of the training process of a lightweight sound detection model provided by an embodiment of the present application. The process includes:
[0139] S201: Collect the audio samples included in the audio sample set, and perform annotation processing on the collected audio samples to obtain the embedding features of the audio samples, the true event detection labels, and the true audio classification labels of the audio samples.
[0140] S202: Obtain the original teacher-student joint training model.
[0141] Figure 3 It is a schematic diagram of the structure of an original teacher-student joint training model provided by an embodiment of the present application. The original teacher-student joint training model includes a teacher model and a student model. The teacher model includes a teacher feature extraction module and a teacher task processing module. The teacher feature extraction module includes N cascaded convolutional layers, where N is an integer greater than 1. The student model includes a student feature extraction module and a student task processing module. The student feature extraction module is the first M convolutional layers of the teacher feature extraction module, where M is a positive integer less than N. The initializations of the teacher task processing module and the student task processing module are the same.
[0142] Based on the audio sample set obtained in S201, perform iterative training on the original teacher-student joint training model obtained in S202 to determine the lightweight sound detection model based on the trained target student feature extraction module and target student task processing module. The following will describe each iterative training process:
[0143] S203: Obtain any audio sample in the audio sample set and input it into the teacher feature extraction module of the original teacher-student joint training model.
[0144] S204: Obtain the first audio feature extracted by the student feature extraction module through the first M convolutional layers in the teacher feature extraction module, and obtain the second audio feature extracted by the teacher feature extraction module through N convolutional layers.
[0145] S205: Through the student task processing module in the original teacher-student joint training model, based on the first audio feature, obtain the student event detection prediction result and the student audio classification prediction result, and execute S208.
[0146] S206: Through the alignment and splicing module included in the teacher model in the original teacher-student joint training model, align the second audio feature output by the teacher feature extraction module with the embedded feature in the time dimension, and splice the aligned second audio feature and the aligned embedded feature in the feature dimension to obtain a spliced feature vector.
[0147] S207: Through the teacher task processing module in the original teacher-student joint training model, based on the spliced feature vector, obtain the teacher event detection prediction result and the teacher audio classification prediction result.
[0148] S208: Determine the first loss between the teacher event detection prediction result and the student event detection prediction result and the true event detection label respectively.
[0149] S209: Determine the second loss between the teacher audio classification prediction result and the student audio classification prediction result and the true audio classification label respectively.
[0150] In a possible implementation manner, both the first loss and the second loss determined in S208 and S209 are binary cross-entropy losses. Exemplarily, the formula for determining the binary cross-entropy loss of any audio sample is expressed as follows:
[0151] L 二元交叉熵 =-h i log(p i )+(1-h i )log(1-p i )
[0152] Wherein, L 二元交叉熵 represents the binary cross-entropy loss, h i is the true label of the i-th audio sample, taking values of 0 or 1; p i is the probability value of the i-th audio sample in the model prediction result, with the value range between (0, 1), and N is the total number of audio samples.
[0153] S210: Determine the third loss between the student event detection prediction result and the teacher event detection prediction result.
[0154] In a possible implementation, the third loss for the audio sample is determined as follows:
[0155] Obtain the sum of squared differences between the student event detection prediction results and the teacher event detection prediction results for each frame and each event category;
[0156] Determine the product between the number of frames of the audio sample and the total number of pre-configured event categories;
[0157] Determine the third loss according to the quotient between the sum of squared differences and the product.
[0158] S211: Determine the fourth loss between the student audio classification prediction result and the teacher audio classification prediction result.
[0159] In a possible implementation, the fourth loss for the audio sample is determined as follows:
[0160] Determine the square of the difference between the student audio classification prediction result and the teacher audio classification prediction result.
[0161] S212: Obtain the weight values corresponding to different losses.
[0162] In the training of the lightweight sound detection model, the target loss function includes multiple loss terms, namely the first loss of the student model, the first loss of the teacher model, the second loss of the student model, the second loss of the teacher model, the third loss, and the fourth loss. To more flexibly control the training process of the model and improve the model performance, corresponding weights can be assigned to different loss terms. These weights can be adjusted according to actual needs, and they can be completely different, completely the same, or partially different.
[0163] In a possible implementation, when both the third loss Q 3 and the fourth loss Q 4 are the same self-knowledge distillation weight parameter value, the self-knowledge distillation weight parameter value is represented by the following formula:
[0164]
[0165] Among them, α is the self-knowledge distillation weight parameter value, K represents the constant maximum value of the loss weight between the teacher and the student, β is the exponential parameter used to control the speed and amplitude of weight adjustment, s is the current training step number, G represents the number of rounds of warm-up training, and Z is the length of each training round. K is an empirical value that determines the maximum influence that the self-knowledge distillation loss can achieve in the target loss function. If K takes a larger value, then in the appropriate training stage, the influence of the self-knowledge distillation loss on the model parameter update will be greater, and the model will pay more attention to the student model learning from the teacher model; conversely, if K takes a smaller value, the influence of the self-knowledge distillation loss is relatively limited. β affects the rate and range of change of the weight with the training step number. When β is larger, the change of the weight with the training step number will be more drastic, that is, the speed of weight adjustment is faster and the amplitude is larger; when β is smaller, the change of the weight will be relatively gentle, and both the adjustment speed and amplitude will become smaller. For example, in the initial stage of training, if β is larger, the weight may rise or fall rapidly; while if β is smaller, the change of the weight will be relatively slow. s reflects the current training stage of the model. As the training progresses, the value of s increases continuously. In the initial stage of training, s is smaller, which will cause the value of the exponential function to be larger, and the self-knowledge distillation weight parameter value may be close to K; as the training step number s increases, the value of the exponential function will become smaller, and the self-knowledge distillation weight parameter value will also gradually decrease. G and Z affect the relative position of s in the entire training process. For example, if G or Z increases, then the entire training cycle becomes longer, and the proportion of the same training step number s in the entire training process will become smaller, thereby affecting the change of the self-knowledge distillation weight parameter value.
[0166] S213: Determine the target loss value through the target loss function based on different losses and their respectively corresponding weight values.
[0167] In a possible implementation manner, the target loss function is represented by the following formula:
[0168]
[0169] Among them, L total represents the target loss function, Q represents the weight values respectively corresponding to different losses, represents the first loss between the student event detection prediction result and the true event detection label, represents the first loss between the teacher event detection prediction result and the true event detection label, represents the second loss between the student audio classification prediction result and the true audio classification label, represents the second loss between the teacher audio classification prediction result and the true audio classification label, L 3 represents the third loss, L 4Denote the fourth loss. For example,
[0170]
[0171] S214: Adjust the parameters of the original teacher-student joint training model for the current iteration based on the target loss value.
[0172] Since the audio sample set for training the original teacher-student joint training model contains a large number of audio samples, perform the above operations S203 to S213 on each audio sample to obtain the target loss value corresponding to each audio sample. Based on the sum of the target loss values, adjust the parameters of the original teacher-student joint training model for the current iteration. When the preset convergence condition is met, the training of the original teacher-student joint training model is completed, thereby obtaining the trained student model, and determining the trained student model as the lightweight sound detection model.
[0173] After obtaining the lightweight sound detection model in this application, the model performance can be evaluated using the event-based F1 score (Event-based-F1), segment-based F1 score (Segment-based-F1), intersection-based F1 score (Intersection-based-F1), and multi-threshold score - polyphonic sound event detection scores (PSDS1 and PSDS2). Among them, the event-based F1 score (Event-based-F1) is used to measure the sound event detection performance. In addition to using multiple sets of operating points, PSDS also sets other parameters for easy adjustment according to different applications to meet various user experience requirements.
[0174] The F1 score (F1-Score) is a statistical metric used to evaluate the performance of a classification model. It takes into account both the precision and recall of the classification model, and is particularly suitable for tasks with class imbalance or when both precision and recall need to be considered comprehensively. The F1 score can be regarded as a weighted average of the model's precision and recall. Its maximum value is 1, and its minimum value is 0. The closer the value is to 1, the better the model performance. Its calculation method can be briefly described as:
[0175]
[0176] Among them, the calculation method of precision is:
[0177]
[0178] The calculation method of recall is:
[0179]
[0180] Among them, TP (True Positive) is the number of samples that the model correctly predicts as the positive class; FP (False Positive) is the number of samples that are incorrectly predicted as the positive class; FN (False Negative) is the number of samples that are incorrectly predicted as the negative class.
[0181] The calculation method of the event-based F1 score uses a complete event as the calculation unit, and this method focuses on the accuracy at the event level in the detection results. In this evaluation method, only when the category of a detected event is correct and its start time and end time are close enough to the reference event, the event is considered correct (True Positive, TP).
[0182] The fragment-based F1 score uses a certain length fragment after cutting as a complete calculation unit, and this method divides the time axis into fixed-length fragments. In each fragment, only whether the target event exists is judged, and the start and end times of the event are not strictly considered. If the reference and predicted events overlap within the fragment, they are considered a match. In the present invention, a 1-second length segment is selected as the fragment-based measurement method.
[0183] The intersection-based F1 score focuses on the overlap situation of the reference event and the predicted event in time. As long as the detected event has sufficient overlap with the reference event in time, it is considered a match, and the time of the specific overlap part is usually defined by proportion or time threshold.
[0184] PSDS1 focuses on the detection speed of the model. When the system detects an event, it needs to respond quickly (such as triggering an alarm, adapting to a home automation system, etc.), that is, it attaches importance to the recall rate; PSDS2 means that when the response time is not so important, the system accurately detects the category of the event, that is, it must avoid confusion between classes and pays more attention to the precision rate. The calculation method of PSDS can be briefly described as:
[0185]
[0186] where e max is the maximum eFPR (effective false positive rate), and r(e) is the ROC curve calculated at multiple operating points.
[0187] Preprocess and extract features from the test set audio data, and send the data to be tested into the lightweight sound event detection model, then the sound events contained in the data to be tested can be predicted. Calculate three F1 scores, PSDS1 and PSDS2 according to the probability values of each event category predicted by the model in each frame.
[0188] Figure 4This is the experimental result graph of the training method for the lightweight sound detection model provided by the embodiments of the present application. Among them, System 1 is the official baseline system of DCASE 2023 Task4 A, and System 2 is the training method for the lightweight sound detection model provided by the embodiments of the present application. From Figure 4 It can be seen that compared with the official baseline system of DCASE 2023 Task4 A, the method provided by the present application has an increase of 3.35 percentage points in Event-based F1, that is, a relative increase of 7.8%. The Intersection-based F1 increases by 4.53 percentage points, with a relative increase of 7.1%. The Segment-based F1 increases by 4.29 percentage points, with a relative increase of 6.2%. The PSDS1 increases by 0.92 percentage points, with a relative increase of 2.5%. The PSDS2 increases by 1.84 percentage points, with a relative increase of 3.2%. The experimental results show that the method proposed by the present application has an obvious effect on improving the detection performance. The model structure and the target loss function in the training of the lightweight sound detection model provided by the present application have an obvious improvement in the event detection performance, that is, it meets the expected effect.
[0189] Example 6:
[0190] The present application also provides a sound event detection method based on the model trained according to any of the above embodiments. Figure 5 This is a schematic diagram of the process of sound event detection provided by the embodiments of the present application. The process includes:
[0191] S501: Obtain the audio to be detected.
[0192] S502: Through the pre-trained lightweight sound detection model, based on the audio to be detected, obtain the event detection result and the audio classification result in the audio to be detected.
[0193] The sound event detection method provided by the present application is applied to a computer device, which can be an intelligent device or a server. Among them, the computer device for performing sound event detection in the present application can be the same as or different from the computer device for training the lightweight sound detection model.
[0194] In a possible implementation manner, generally, the training of the lightweight sound detection model is carried out offline. After obtaining the lightweight sound detection model, the lightweight sound detection model can be saved to the computer device for sound event detection.
[0195] In an actual application scenario, when a computer device receives an audio to be detected, it will quickly call the saved lightweight sound detection model. Among them, the audio to be detected can be collected by the computer device itself or received from other collection devices. The lightweight sound detection model is trained based on any of the above embodiments. The lightweight sound detection model has learned rich feature representations and knowledge from the training process through the teacher model. Therefore, it can efficiently process audio data and output accurate event detection results and audio classification results in a short time. For computer devices with limited resources, the lightweight sound detection model can quickly complete the detection task without occupying too much computing resources and memory. For example, in a smart home system, as a smart device, when the smart speaker detects the sound in the surrounding environment, with the help of the lightweight sound detection model, it can quickly determine whether it is the normal conversation sound of family members, the running sound of electrical appliances, or possible abnormal alarm sounds and other events, and accurately classify these sounds, such as classifying the sounds into categories such as human voices, mechanical sounds, or natural environment sounds, so as to provide a basis for subsequent automatic control or information feedback.
[0196] It should be noted that the training process of the lightweight sound detection model has been described in the above Embodiments 1-5, and will not be specifically elaborated here.
[0197] The computer device inputs the obtained audio to be detected into the lightweight sound detection model. Through the processing of the lightweight sound detection model, the event detection result and audio classification result of the audio to be detected can be obtained, that is, the type of sound event included in the audio to be detected and the time period when the event occurs are predicted. Finally, the detected sound event category and its corresponding time boundary are output. Among them, the audio to be detected may contain one or more sound events.
[0198] In one example, before inputting the audio to be detected into the lightweight sound detection model, some preprocessing operations can be performed to ensure that the format and quality of the audio meet the requirements of the model. Common preprocessing steps include, but are not limited to, one or more of the following: audio format unification, sampling rate unification, bit rate unification, acoustic feature extraction (such as Mel Frequency Cepstral Coefficients (MFCC), Linear Predictive Cepstral Coefficients (LPCC), spectral features, etc.).
[0199] Embodiment 7:
[0200] Based on the same inventive concept, the present application also provides a training device for a lightweight sound detection model. Figure 6 FIG. is a schematic structural diagram of a training device for a lightweight sound detection model provided by an embodiment of the present application. The device includes:
[0201] An acquisition unit 61, configured to acquire an audio sample set and an original teacher-student joint training model;
[0202] Wherein, the audio sample set includes audio samples, the embedded features corresponding to the audio samples, the true event detection labels of the audio samples, and the true audio classification labels of the audio samples;
[0203] The original teacher-student joint training model includes a teacher model and a student model;
[0204] The teacher model includes a teacher feature extraction module and a teacher task processing module, and the teacher feature extraction module includes N cascaded convolutional layers, where N is an integer greater than 1;
[0205] The student model includes a student feature extraction module and a student task processing module, and the student feature extraction module is the first M convolutional layers of the teacher feature extraction module, where M is a positive integer less than N;
[0206] The initializations of the teacher task processing module and the student task processing module are the same;
[0207] A training unit 62, configured to iteratively train the original teacher-student joint training model based on the audio sample set, so as to determine a lightweight sound detection model based on the trained target student feature extraction module and target student task processing module;
[0208] During each iterative training process:
[0209] Input any audio sample into the feature extraction module, extract first audio features through the student feature extraction module, and extract second audio features through the teacher feature extraction module;
[0210] Through the student task processing module, based on the first audio features, obtain a student event detection prediction result and a student audio classification prediction result;
[0211] Through the teacher task processing module, based on the second audio sample and the embedded features of the audio sample, obtain a teacher event detection prediction result and a teacher audio classification prediction result;
[0212] Based on the target loss function, adjust the parameters of the original teacher-student joint training model in the current iteration to determine the trained student model as a lightweight voice detection model; wherein, the target loss function includes the following losses: the first loss between the teacher event detection prediction result and the student event detection prediction result respectively and the true event detection label, the second loss between the teacher audio classification prediction result and the student audio classification prediction result respectively and the true audio classification label, the third loss between the student event detection prediction result and the teacher event detection prediction result, and the fourth loss between the student audio classification prediction result and the teacher audio classification prediction result.
[0213] The training device of the lightweight voice detection model in this embodiment is presented in the form of functional modules. Here, the module refers to an application specific integrated circuit (ASIC), a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0214] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments 1-5 above, and will not be elaborated here.
[0215] Embodiment 8:
[0216] Based on the same inventive concept, the present application also provides a voice event detection device for a model trained by any of the methods in Embodiments 1-5. Figure 7 FIG. is a schematic structural diagram of a voice event detection device provided by an embodiment of the present application. The device includes:
[0217] An acquisition module 71, configured to acquire an audio to be detected;
[0218] A processing module 72, configured to obtain an event detection result and an audio classification result in the audio to be detected based on the audio to be detected through a pre-trained lightweight voice detection model.
[0219] The voice event detection device in this embodiment is presented in the form of functional modules. Here, the module refers to an application specific integrated circuit (ASIC), a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0220] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding Embodiment 6, and will not be elaborated here.
[0221] Embodiment 9:
[0222] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present application. As Figure 8 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 8 In
[0223] Figure 8 , a single processor 10 is taken as an example.
[0224] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.
[0225] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.
[0226] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid state drive; the memory 20 may further include a combination of the above types of memory.
[0227] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 8 Taking connection via a bus as an example.
[0228] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device, such as touch screen, keypad, mouse, trackpad, touchpad, pointing stick, one or more mouse buttons, trackball, joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., LED), and a haptic feedback device (e.g., vibration motor), etc. The above display device includes but is not limited to liquid crystal display, light emitting diode, display and plasma display. In some alternative embodiments, the display device may be a touch screen.
[0229] Embodiment 10:
[0230] Based on the above embodiments, the embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor is caused to execute the following steps:
[0231] Obtain an audio sample set and an original teacher-student joint training model;
[0232] Wherein, the audio sample set contains audio samples, the embedding features corresponding to the audio samples, the true event detection labels of the audio samples, and the true audio classification labels of the audio samples;
[0233] The original teacher-student joint training model includes a teacher model and a student model;
[0234] The teacher model includes a teacher feature extraction module and a teacher task processing module. The teacher feature extraction module includes N cascaded convolutional layers, where N is an integer greater than 1;
[0235] The student model includes a student feature extraction module and a student task processing module. The student feature extraction module is the first M convolutional layers of the teacher feature extraction module, where M is a positive integer less than N;
[0236] The initializations of the teacher task processing module and the student task processing module are the same;
[0237] Based on the audio sample set, iteratively train the original teacher-student joint training model to determine a lightweight voice detection model based on the trained target student feature extraction module and target student task processing module;
[0238] During each iterative training process:
[0239] Input any audio sample into the teacher feature extraction module, and through the M convolutional layers, obtain the first audio feature extracted by the student feature extraction module, and through the N convolutional layers, obtain the second audio feature extracted by the teacher feature extraction module;
[0240] Through the student task processing module, based on the first audio feature, obtain the student event detection prediction result and the student audio classification prediction result;
[0241] Through the teacher task processing module, based on the second audio sample and the embedded feature of the audio sample, obtain the teacher event detection prediction result and the teacher audio classification prediction result;
[0242] Based on the target loss function, adjust the parameters of the original teacher-student joint training model in the current iteration to determine the trained student model as a lightweight voice detection model; wherein, the target loss function includes the following losses: the first loss between the teacher event detection prediction result and the student event detection prediction result and the true event detection label, the second loss between the teacher audio classification prediction result and the student audio classification prediction result and the true audio classification label, the third loss between the student event detection prediction result and the teacher event detection prediction result, and the fourth loss between the student audio classification prediction result and the teacher audio classification prediction result.
[0243] Since the principle of the above computer-readable storage medium for solving problems is similar to the training method of the lightweight voice detection model, the implementation of the above computer-readable storage medium can refer to Embodiments 1-5 of the method, and the repeated parts will not be elaborated.
[0244] Embodiment 11:
[0245] Based on the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor is caused to execute the following steps when executing:
[0246] Obtain the audio to be detected;
[0247] Based on the lightweight sound detection model completed through pre-training, and based on the audio to be detected, obtain the event detection result and audio classification result in the audio to be detected.
[0248] Since the principle of the above computer-readable storage medium for solving problems is similar to that of the sound event detection method, the implementation of the above computer-readable storage medium can refer to Embodiment 6 of the method, and the repeated parts will not be elaborated here.
[0249] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A training method for a lightweight sound detection model, characterized in that: The method comprises: Get the audio sample set and the original teacher-student joint training model; The audio sample set includes audio samples, embedded features corresponding to the audio samples, real event detection labels of the audio samples, and real audio classification labels of the audio samples; The original teacher-student joint training model includes a teacher model and a student model; The teacher model includes a teacher feature extraction module and a teacher task processing module, and the teacher feature extraction module includes N convolutional layers connected in series, where N is an integer greater than 1; The student model includes a student feature extraction module and a student task processing module, wherein the student feature extraction module is the first M convolutional layers of the teacher feature extraction module, where M is a positive integer less than N; The initialization of the teacher task processing module and the student task processing module is the same; Based on the audio sample set, iteratively train the original teacher-student joint training model to determine a lightweight sound detection model based on the trained target student feature extraction module and the target student task processing module; During each training iteration: Input any audio sample into the teacher feature extraction module, obtain the first audio feature extracted by the student feature extraction module through the M convolutional layers, and obtain the second audio feature extracted by the teacher feature extraction module through the N convolutional layers; Obtaining, through the student task processing module, a student event detection prediction result and a student audio classification prediction result based on the first audio feature; Obtaining, through the teacher task processing module, a teacher event detection prediction result and a teacher audio classification prediction result based on the second audio sample and the embedded features of the audio sample; Based on the target loss function, the parameters of the original teacher-student joint training model of the current iteration are adjusted to determine the trained student model as a lightweight sound detection model; wherein the target loss function includes the following losses: a first loss between the teacher event detection prediction result, the student event detection prediction result and the real event detection label, respectively, a second loss between the teacher audio classification prediction result, the student audio classification prediction result and the real audio classification label, respectively, a third loss between the student event detection prediction result and the teacher event detection prediction result, and a fourth loss between the student audio classification prediction result and the teacher audio classification prediction result.
2. The method according to claim 1, characterized in that The third loss for the audio sample is determined as follows: Obtain the sum of squares of differences between the student event detection prediction result and the teacher event detection prediction result in each frame and each event category; Determine the product between the number of frames of the audio sample and the total number of preconfigured event categories; The third loss is determined according to a quotient between the sum of squared differences and the product.
3. The method according to claim 1, characterized in that The fourth loss for the audio sample is determined as follows: The square of the difference between the student audio classification prediction result and the teacher audio classification prediction result is determined.
4. The method according to claim 1, characterized in that The objective loss function is expressed by the following formula: Among them, L total represents the target loss function, Q represents the weight values corresponding to different losses, represents the first loss between the student event detection prediction result and the real event detection label, represents the first loss between the teacher event detection prediction result and the true event detection label, represents the second loss between the student audio classification prediction result and the true audio classification label, represents the second loss between the teacher audio classification prediction result and the real audio classification label, L3 represents the third loss, and L4 represents the fourth loss.
5. The method according to claim 4, characterized in that In the case where both Q3 and Q4 have the same self-knowledge distillation weight parameter value, the self-knowledge distillation weight parameter value is expressed by the following formula: Among them, α is the weight parameter value of the self-knowledge distillation, K represents the constant maximum value of the loss weight between the teacher and the student, β is an exponential parameter used to control the speed and amplitude of weight adjustment, s is the current training step number, G represents the number of warm-up training rounds, and Z is the length of each training round.
6. The method according to claim 1, characterized in that The teacher model further includes an alignment and splicing module, which is used to align the second audio feature output by the teacher feature extraction module with the embedded feature in a time dimension; splice the aligned second audio feature with the aligned embedded feature in a feature dimension to obtain a spliced feature vector; Obtaining a teacher event detection prediction result and a teacher audio classification prediction result based on the second audio sample and the embedded features of the audio sample through the teacher task processing module, including: Through the teacher task processing module, the teacher event detection prediction result and the teacher audio classification prediction result are obtained based on the spliced feature vector.
7. A method for detecting sound events based on a model trained by the method according to any one of claims 1 to 6, characterized in that: The method comprises: Get the audio to be detected; Through the pre-trained lightweight sound detection model, based on the audio to be detected, the event detection result and the audio classification result in the audio to be detected are obtained.
8. A training device for a lightweight sound detection model, characterized in that: The device comprises: An acquisition unit, used to acquire an audio sample set and an original teacher-student joint training model; The audio sample set includes audio samples, embedded features corresponding to the audio samples, real event detection labels of the audio samples, and real audio classification labels of the audio samples; The original teacher-student joint training model includes a teacher model and a student model; The teacher model includes a teacher feature extraction module and a teacher task processing module, and the teacher feature extraction module includes N convolutional layers connected in series, where N is an integer greater than 1; The student model includes a student feature extraction module and a student task processing module, wherein the student feature extraction module is the first M convolutional layers of the teacher feature extraction module, where M is a positive integer less than N; The initialization of the teacher task processing module and the student task processing module is the same; A training unit, configured to iteratively train the original teacher-student joint training model based on the audio sample set, so as to determine a lightweight sound detection model based on the trained target student feature extraction module and the target student task processing module; During each training iteration: Input any audio sample into the feature extraction module, extract the first audio feature through the student feature extraction module, and extract the second audio feature through the teacher feature extraction module; Obtaining, through the student task processing module, a student event detection prediction result and a student audio classification prediction result based on the first audio feature; Obtaining, through the teacher task processing module, a teacher event detection prediction result and a teacher audio classification prediction result based on the second audio sample and the embedded features of the audio sample; Based on the target loss function, the parameters of the original teacher-student joint training model of the current iteration are adjusted to determine the trained student model as a lightweight sound detection model; wherein the target loss function includes the following losses: a first loss between the teacher event detection prediction result, the student event detection prediction result and the real event detection label, respectively, a second loss between the teacher audio classification prediction result, the student audio classification prediction result and the real audio classification label, respectively, a third loss between the student event detection prediction result and the teacher event detection prediction result, and a fourth loss between the student audio classification prediction result and the teacher audio classification prediction result.
9. A sound event detection device based on a model trained by the method according to any one of claims 1 to 6, characterized in that: The device comprises: An acquisition module, used to acquire the audio to be detected; The processing module is used to obtain event detection results and audio classification results in the audio to be detected based on the audio to be detected through a pre-trained lightweight sound detection model.
10. A computer device, characterized in that: The computer device includes a processor, which is used to implement the steps of the training method of the lightweight sound detection model as described in any one of claims 1 to 6 when executing the computer program stored in the memory, or to implement the steps of the sound event detection method as described in claim 7.