Acoustic event detection method, device, electronic device and storage medium

By using multi-model fusion and weighting processing methods in acoustic event detection, the problems of domain imbalance and multi-class event overlap are solved, and the detection effect and system universality and generalization are improved.

CN114627861BActive Publication Date: 2025-05-16SHANGHAI NORMAL UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210200026.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-01
Publication Date
2025-05-16
Estimated Expiration
2042-03-01

AI Technical Summary

Technical Problem

In the acoustic event detection task, due to the complex real environment data, domain imbalance and overlapping of multiple events, the model model is difficult and the detection performance is paranoid and imbalanced.

Method used

A method of acoustic event detection is proposed, which is preprocessed by obtaining target audio and inputting it into the first acoustic event detection model trained by the original training sample and the second acoustic event detection model trained by the high-quality training sample adjustment. Through multi-model fusion, adaptive weighting determines the acoustic event category of each sound segment in the target audio.

Benefits of technology

Through multi-model fusion and weighting processing, the detection effect of acoustic event detection is improved, the model's deviation of detection capabilities for different categories is reduced, and the universality and generalization of the system are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627861B_ABST
    Figure CN114627861B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose an acoustic event detection method, device, electronic device and storage medium. A specific implementation of the method includes: obtaining target audio; preprocessing the target audio; inputting the preprocessed target audio into a first acoustic event detection model trained with original training samples and a second acoustic event detection model trained with high-quality training samples; and determining the acoustic event category of each sound segment in the target audio based on the output of the first acoustic event detection model and the second acoustic event detection model. This implementation provides an acoustic event detection mechanism based on multiple models, which improves the detection effect of acoustic event detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technology, and in particular, to acoustic event detection methods, devices, electronic devices, and storage media. Background Art

[0002] With the rapid development of artificial intelligence technology and deep neural networks and the rise of artificial intelligence applications, intelligent voice technology has gradually been widely used in people's production and life, including sound scene classification, sound event classification, abnormal acoustic event detection and other aspects. Among them, acoustic event detection technology imitates the ability of humans to identify acoustic events, and uses audio signal processing and deep learning technology to complete the recognition and classification of acoustic events, such as speaking, alarm sounds, car engine sounds, bird calls, etc. Acoustic event detection (AED) refers to predicting the category of acoustic events occurring in audio clips and identifying the start and offset timestamps of these events. Acoustic event detection can be applied to many fields, such as smart homes, health monitoring systems, unmanned driving, multimedia retrieval, and speech recognition in complex scenarios.

[0003] However, in the acoustic event detection task, since the data collected from the real environment is very complex, there are domain imbalance problems and overlapping of multiple types of events in most cases, which makes it difficult to model the acoustic event detection system and it is difficult to achieve the accuracy required for practical applications. In addition, the detection performance of a single model is biased, and there are obvious differences in the detection capabilities of different categories, resulting in unbalanced detection results, which has a serious impact on the universality and generalization of the system. Summary of the invention

[0004] The embodiments of the present application provide an acoustic event detection method, device, electronic device, and storage medium.

[0005] In a first aspect, some embodiments of the present application provide an acoustic event detection method, which includes: obtaining target audio; preprocessing the target audio; inputting the preprocessed target audio into a first acoustic event detection model trained with original training samples and a second acoustic event detection model trained with high-quality training samples; and determining the acoustic event category of each sound segment in the target audio based on the outputs of the first acoustic event detection model and the second acoustic event detection model.

[0006] In some embodiments, the high-quality training samples include training samples obtained by screening through the following steps: obtaining separated audio segments included in the original audio data and first type information corresponding to each separated audio segment through a pre-trained speech separation model; labeling the separated audio segments through a pre-trained acoustic event detection model to obtain second type information corresponding to each separated audio segment; screening out separated audio segments whose corresponding first type information is the same as the second type information; determining the screened separated audio segments as high-quality training samples, and determining the first type information or the second type information corresponding to the screened separated audio segments as labels of the high-quality training samples.

[0007] In some embodiments, filtering out separated audio segments whose corresponding first type information is identical to the second type information includes: filtering out separated audio segments whose corresponding first type information is identical to the second type information, and whose first type information or the second type information is of the target event type.

[0008] In some embodiments, the method also includes the step of adjusting the training of the second acoustic event detection model, and the step of adjusting the training of the second acoustic event detection model includes: superimposing the separated audio segments in the high-quality training samples to form a mixed audio; inputting the mixed audio into the third acoustic event detection model to obtain a first prediction result; inputting the separated audio segments in the high-quality training samples into the fourth acoustic event detection model in sequence, and generating corresponding second prediction results in sequence; adjusting the weight parameters of the weighted first prediction result and the second prediction result and the parameters of the fourth acoustic event detection model according to the first prediction result and the second prediction result and the labels of the high-quality training samples.

[0009] In some embodiments, the second acoustic event detection model includes a fourth acoustic event detection model with adjusted parameters, and / or a fifth acoustic event detection model obtained by weighting the third acoustic event detection model and the fourth acoustic event detection model with adjusted parameters according to the adjusted weight parameters.

[0010] In some embodiments, the acoustic event category of each sound segment in the target audio is determined based on the output of the first acoustic event detection model and the second acoustic event detection model, including: adaptively weighted fusion of the output results of the first acoustic event detection model and the second acoustic event detection model based on the class discrimination of the first acoustic event detection model and the second acoustic event detection model; and determining the acoustic event category of each sound segment in the target audio based on the output results after weighted fusion.

[0011] In a second aspect, some embodiments of the present application provide an acoustic event detection device, which includes: an acquisition unit, configured to acquire target audio; a preprocessing unit, configured to preprocess the target audio; an input unit, configured to input the preprocessed target audio into a first acoustic event detection model trained with original training samples and a second acoustic event detection model trained with high-quality training samples; a determination unit, configured to determine the acoustic event category of each sound segment in the target audio based on the outputs of the first acoustic event detection model and the second acoustic event detection model.

[0012] In some embodiments, the device also includes a screening unit, which is configured to: obtain separated audio segments included in the original audio data and first type information corresponding to each separated audio segment through a pre-trained speech separation model; label the separated audio segments through a pre-trained acoustic event detection model to obtain second type information corresponding to each separated audio segment; screen out separated audio segments whose corresponding first type information is the same as the second type information; determine the screened separated audio segments as high-quality training samples, and determine the first type information or the second type information corresponding to the screened separated audio segments as labels of high-quality training samples.

[0013] In some embodiments, the screening unit is further configured to screen out separated audio segments whose corresponding first type information is the same as the second type information, and the first type information or the second type information is of the target event type.

[0014] In some embodiments, the device also includes an adjustment unit, which is configured to: superimpose the separated audio segments in the high-quality training samples to form a mixed audio; input the mixed audio into the third acoustic event detection model to obtain a first prediction result; input the separated audio segments in the high-quality training samples into the fourth acoustic event detection model in sequence, and generate corresponding second prediction results in sequence; adjust the weight parameters of the first prediction result and the second prediction result and the parameters of the fourth acoustic event detection model according to the first prediction result and the second prediction result and the labels of the high-quality training samples.

[0015] In some embodiments, the second acoustic event detection model includes a fourth acoustic event detection model with adjusted parameters, and / or a fifth acoustic event detection model obtained by weighting the third acoustic event detection model and the fourth acoustic event detection model with adjusted parameters according to the adjusted weight parameters.

[0016] In some embodiments, the determination unit is further configured to: adaptively weight and fuse the output results of the first acoustic event detection model and the second acoustic event detection model based on the class discrimination of the first acoustic event detection model and the second acoustic event detection model; and determine the acoustic event category of each sound segment in the target audio according to the output results after weighted fusion.

[0017] In a third aspect, some embodiments of the present application provide a device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above in the first aspect.

[0018] In a fourth aspect, some embodiments of the present application provide a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method described above in the first aspect.

[0019] The acoustic event detection method, device, electronic device and storage medium provided in the embodiments of the present application obtain target audio; preprocess the target audio; input the preprocessed target audio into a first acoustic event detection model trained with original training samples and a second acoustic event detection model trained with high-quality training samples; and determine the acoustic event category of each sound segment in the target audio according to the output of the first acoustic event detection model and the second acoustic event detection model, thereby providing a multi-model-based acoustic event detection mechanism and improving the detection effect of acoustic event detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0021] Figure 1 are some exemplary system architecture diagrams to which the present application may be applied;

[0022] Figure 2 is a flow chart of an embodiment of an acoustic event detection method according to the present application;

[0023] Figure 3 is a flow chart of a modeling method in an optional implementation of the acoustic event detection method of the present application;

[0024] Figure 4 is a schematic diagram of a multi-model score fusion method in an optional implementation of the acoustic event detection method of the present application;

[0025] Figure 5 is a schematic structural diagram of an embodiment of an acoustic event detection device according to the present application;

[0026] Figure 6 It is a structural diagram of a computer system of a server or terminal suitable for implementing some embodiments of the present application. DETAILED DESCRIPTION

[0027] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0028] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0029] Figure 1 An exemplary system architecture 100 is shown to which an embodiment of an acoustic event detection method or an acoustic event detection apparatus of the present application can be applied.

[0030] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0031] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various client applications can be installed on terminal devices 101, 102, 103, such as voice recognition applications, smart speaker applications, Internet of Things applications, search applications, etc.

[0032] Terminal devices 101, 102, 103 may be hardware or software. When terminal devices 101, 102, 103 are hardware, they may be various electronic devices, including but not limited to smart speakers, smart phones, tablet computers, laptop computers, desktop computers, etc. When terminal devices 101, 102, 103 are software, they may be installed in the electronic devices listed above. They may be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is made here.

[0033] The server 105 may be a server that provides various services, such as a background server that provides support for applications installed on the terminal devices 101, 102, and 103. The server 105 may obtain the target audio uploaded by the terminal devices 101, 102, and 103; preprocess the target audio; input the preprocessed target audio into a first acoustic event detection model trained with original training samples and a second acoustic event detection model trained with high-quality training samples; and determine the acoustic event category of each sound segment in the target audio based on the outputs of the first acoustic event detection model and the second acoustic event detection model.

[0034] It should be noted that the acoustic event detection method provided in the embodiment of the present application can be executed by the server 105 or by the terminal devices 101, 102, 103. Accordingly, the acoustic event detection device can be set in the server 105 or in the terminal devices 101, 102, 103.

[0035] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0036] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0037] Continue to refer Figure 2 , shows a process 200 of an embodiment of an acoustic event detection method according to the present application. The acoustic event detection method comprises the following steps:

[0038] Step 201, obtaining target audio.

[0039] In this embodiment, the acoustic event detection method execution body (eg Figure 1 The server or terminal shown in the figure can first obtain the target audio, and the target audio can include the audio for which the acoustic event detection is to be performed.

[0040] Step 202: pre-process the target audio.

[0041] In this embodiment, the above-mentioned execution entity can pre-process the target audio obtained in step 201. The pre-processing may include operations such as pre-emphasis, framing, windowing, endpoint detection, i.e., feature extraction. The specific pre-processing method can be selected according to actual needs.

[0042] Step 203: input the preprocessed target audio into a first acoustic event detection model trained with original training samples and a second acoustic event detection model trained with high-quality training samples.

[0043] In this embodiment, the above-mentioned execution entity can input the target audio preprocessed in step 202 into the first acoustic event detection model trained by the original training samples and the second acoustic event detection model adjusted and trained by the high-quality training samples. The original training samples may include data collected in a real environment. The original training data has not been processed by screening, etc., and contains environmental noise, other non-target events, and multiple target events, etc. mixed together, which is extremely unfriendly to model training. High-quality training samples may include samples processed by separation, screening and other operations. High-quality training samples can be obtained through existing sample sets, or can be obtained using trained models, such as speech separation models. If the second acoustic event detection model is trained using separated audio segments, the preprocessed target audio can first be obtained by the speech separation model to obtain separated audio segments, and then the obtained separated audio segments can be input into the second acoustic event detection model.

[0044] In some optional implementations of this embodiment, the high-quality training samples include training samples screened through the following steps: obtaining separated audio segments included in the original audio data and the first type information corresponding to each separated audio segment through a pre-trained speech separation model; labeling the separated audio segments through a pre-trained acoustic event detection model to obtain the second type information corresponding to each separated audio segment; screening out separated audio segments whose corresponding first type information and second type information are the same; determining the screened separated audio segments as high-quality training samples, and determining the first type information or the second type information corresponding to the screened separated audio segments as labels of the high-quality training samples. This implementation effectively solves the impact of event overlap in training data on system modeling.

[0045] In this implementation, speech separation technology can be used for acoustic event detection tasks. First, by separating the mixed sound signal, separated data containing only a single category of audio is obtained, and then each separated data can be used to train the AED system. The detection effect obtained on the separated signal is usually more accurate than the detection effect obtained on the mixed signal. However, using separated audio to assist in modeling the acoustic event detection system is not a simple problem. The separated sound may contain erroneous or inaccurate information. For example, the separated audio data usually contains non-target events and other background noise, which in turn may affect the training of the acoustic event detection model.

[0046] For the original audio data, after passing through the speech separation model, a separated audio segment containing only a single category of events can be obtained, and the output of the speech separation model can include weak label data (classification results) and strong label data (timestamp). Among them, the classification results may include target events, non-target events, background noise, etc. After the separation operation, an original mixed audio segment is separated to obtain multiple separated segments, which contain target events, non-target events and background noise related to the original mixed audio segment. It should be noted that in order to obtain accurate separation results, the number of sound sources output by the separation system can be set by counting the maximum number of events that appear in a single audio segment of the development data set label data set, so as to ensure the accuracy and completeness of the separated audio segment as accurately as possible. As an example, according to the statistical results in the label data, the number of separation sources set can be 6. At this point, multiple relatively clean sound segments containing a single event type (target event, non-target event, background noise, etc.) have been obtained.

[0047] The second type of information corresponding to each separated audio segment obtained by the pre-trained acoustic event detection model for labeling the separated audio segments can be regarded as a pseudo-label mark. The structure of the pre-trained acoustic event detection model is similar to that of the speech separation model, except that the number of categories output by the classifier is different. Because there are other non-target events or background noise in the sound segments generated by the separation system, these data segments will also obtain corresponding pseudo-labels during the labeling period. In order to more accurately mark the label information for the separated audio segments, the number of predictions of the classifier can be redefined for the model. As an example, the "other" class can be added, and the model can eventually generate 11 categories of classification results (10 categories of target events plus the "other" class). It should be noted that during the labeling process, only the classification results (weak labels) predicted by the model can be retained as pseudo-labels.

[0048] However, the classifier that marks the pseudo-label also has its own shortcomings, which may lead to incorrect classification, so the pseudo-label it generates cannot be fully trusted. At the same time, the separated audio segment with the classification label "other" is not the required sound segment. In order to obtain high-quality separated audio segments to assist the AED system, the separated sound segments can be screened with high confidence. This method proposes a data screening method based on original labels. For the separated audio segments, by referring to the original label (Ground truth), that is, the label output by the speech separation model including the first type of information, and the pseudo-label (Pseudolabel), that is, the label output by the pre-trained acoustic event detection model including the second type of information, a combination of the method can screen out more accurate separation data.

[0049] In some optional implementations of this embodiment, screening out the separated audio segments whose corresponding first type information is identical to the second type information includes: screening out the separated audio segments whose corresponding first type information is identical to the second type information and whose first type information or second type information is of the target event type. This implementation effectively solves the influence of event overlap and background noise interference in the training data on system modeling.

[0050] In this implementation, the original label can be compared with the pseudo-label corresponding to the separated audio segment, the separated audio segment whose pseudo-label is consistent with the original label is retained, and the separated audio segment whose pseudo-label is different from the original label is deleted. At the same time, the separated audio segment whose pseudo-label is "other" will also be deleted by the screening model (Sc selection module). After screening, only the same category information as that contained in the original data is retained. Although the complete correctness of the screened data cannot be guaranteed, the errors caused by erroneous separation data and other background classes can be effectively reduced. In addition, for the sound segment with strong labels, the corresponding strong label information is still retained in the final label, which greatly improves the effectiveness of the pseudo-label information in the separated audio segment.

[0051] In some optional implementations of the present embodiment, the method also includes the step of adjusting the training of the second acoustic event detection model, and the step of adjusting the training of the second acoustic event detection model includes: superimposing the separated audio segments in the high-quality training samples to form a mixed audio; inputting the mixed audio into the third acoustic event detection model to obtain a first prediction result; inputting the separated audio segments in the high-quality training samples into the fourth acoustic event detection model in sequence, and generating corresponding second prediction results in sequence; adjusting the weight parameters of the weighted first prediction result and the second prediction result and the parameters of the fourth acoustic event detection model according to the first prediction result and the second prediction result and the labels of the high-quality training samples.

[0052] In this implementation, the third acoustic event detection model and the fourth acoustic event detection model can be two independently pre-trained AED models. Although the AED system trained by the normal training method has been able to achieve good detection performance, the model is affected by the background noise disturbance in the development set data, the interference of non-target events, and the overlap of multiple events during the training process. Its detection performance still has room for further improvement and optimization. Relatively clean separated audio clips without event overlap interference can be obtained through high-quality training samples. Using them to re-fine-tune the AED system can further enhance the detection capability of the AED system.

[0053] See also Figure 3 , Figure 3It is a flow chart of the modeling method in an optional implementation of the acoustic event detection method of the present application. The specific method of adjusting the training can be divided into three steps. First, based on high-quality training samples, two independently pre-trained AED models, namely the third acoustic event detection model and the fourth acoustic event detection model, are used. In the first branch, all separated sound clips are superimposed to form a new mixed audio vector, and then the new audio vector is sent to the pre-trained third acoustic event detection model, but it should be noted that all parameters of the branch network are no longer involved in the update training; secondly, in another branch, the separated sound clips are sent to the pre-trained fourth acoustic event detection model in turn, and the corresponding prediction results are generated in turn, and then multiple groups of prediction results are added and combined as the final prediction results of the branch. In this branch, the model uses each separated audio clip for fine-tuning training; finally, the prediction results of the third acoustic event detection model in the first branch and the prediction results of the fourth acoustic event detection model in the second branch are weighted averaged to obtain the final prediction result of the system. The weight of the weighted average can be learned during the training process, and is defined as:

[0054] P final =ψ·P SED +(1-ψ)·P F-SED

[0055] Where P final is the final prediction of the joint training system, P SED and P F-SED represent the prediction results of the third acoustic event detection model and the fourth acoustic event detection model respectively, and ψ is the weight parameter to be learned.

[0056] In some optional implementations of the present embodiment, the second acoustic event detection model includes a fourth acoustic event detection model with adjusted parameters, and / or a fifth acoustic event detection model obtained by weighting the third acoustic event detection model and the fourth acoustic event detection model with adjusted parameters according to the adjusted weight parameters. Through fine-tuning training, the model is readjusted according to the characteristic information of the single event sample, which can effectively improve the impact of the multi-event overlap problem on the model detection capability. At the same time, by combining the original model, that is, the third acoustic event detection model, with the prediction results of the weighted mixed data, the system's detection capability for the mixed data to be tested can be retained, thereby improving the performance of the entire acoustic event detection system.

[0057] Step 204: Determine the acoustic event category of each sound segment in the target audio according to the outputs of the first acoustic event detection model and the second acoustic event detection model.

[0058] In this embodiment, the execution subject may determine the acoustic event category of each sound segment in the target audio according to the output of the first acoustic event detection model and the second acoustic event detection model in step 203. As an example, the execution subject may pre-set the weighted weights of the outputs of the first acoustic event detection model and the second acoustic event detection model, and may also select to use the output of the first acoustic event detection model or the second acoustic event detection model according to the characteristics of the target audio, or the weighted weights of the outputs of the first acoustic event detection model and the second acoustic event detection model according to the characteristics of the target audio.

[0059] In some optional implementations of the present embodiment, the acoustic event category of each sound segment in the target audio is determined according to the output of the first acoustic event detection model and the second acoustic event detection model, including: adaptively weighted fusion of the output results of the first acoustic event detection model and the second acoustic event detection model based on the class discrimination of the first acoustic event detection model and the second acoustic event detection model; and determining the acoustic event category of each sound segment in the target audio according to the output results after weighted fusion. The first acoustic event detection model and the second acoustic event detection model each have their own detection characteristics and have different detection capabilities for different categories. By selecting the detection category that the model is "good at" for weighting based on the evaluation index of class discrimination (F1 score), and fusing multiple groups of model scores, better detection capabilities can be obtained.

[0060] Compared with using a single model, combining the prediction results of multiple models usually produces better results, but the quality of the fusion method of multi-model prediction results will greatly affect the final detection performance. For AED tasks, a score fusion method based on class discrimination can be used to adaptively weight the detection performance of each category according to different models.

[0061] Taking the second acoustic event detection model including the fourth acoustic event detection model and the fifth acoustic event detection model after adjusting parameters as an example, according to the three groups of single models, namely the first acoustic event detection model, the fourth acoustic event detection model and the fifth acoustic event detection model, the F1 score of event detection is and its frame-level predicted posterior probability P cn , weighted fusion is performed on the results of the three groups of models.

[0062] Further references Figure 4, which shows a schematic diagram of a multi-model score fusion method in an optional implementation of the acoustic event detection method according to the present application; first, model 1 can be the fourth acoustic event detection model fine-tuned by separating data. This section mainly considers the model obtained after fine-tuning training with the assistance of screened high-quality separated audio clips; Model 2 and Model 3 can be the first acoustic event detection model and the fifth acoustic event detection model. These two groups of models can be tested using the original mixed data, which can effectively avoid the errors caused by separating data during the test process. In the test phase, the prediction results of these three independent models are fused and calculated to finally obtain the detection results of the system.

[0063] According to the characteristics of AED tasks, a score fusion method based on class discrimination can be established. cn Defined as the posterior probability of category c predicted by model n, the proportion of the same category in each model in the fusion result is calculated based on the F1 score (Class-wise F1-Score) of each model for each type of event, and the posterior probabilities of different models are combined in a weighted manner to obtain the final prediction score. The weighted fusion method based on class discrimination is as follows:

[0064]

[0065] Where N is the number of single models, is an adjustable scalar parameter, and the final frame-level score It can be used for model evaluation. By calculating the weights according to the category differentiation method, the advantages of each model in predicting the current category are fully considered. The final fusion result can also retain the best detection results of different models for each category, thereby improving the overall performance of the system. Through the multi-model score fusion strategy based on class differentiation, the detection performance of different models for each category is adaptively weighted and fused, further integrating the advantages of the models and improving the detection capability of the system.

[0066] Acoustic event detection uses event-based F1 score (Event-based-F1) and multi-threshold-based score-polyphonic acoustic event detection score (PSDS1 and PSDS2). Event-based F1 score (Event-based-F1) is used to measure the performance of acoustic event detection. In addition to using multiple sets of operating points, PSDS also sets other parameters to facilitate adjustment according to different applications to meet various user experience requirements.

[0067] F1-Score is an indicator used in statistics to measure the accuracy of classification models. It takes into account both the precision and recall of the classification model. The F1 score can be regarded as a weighted average of the model's precision and recall, with a maximum value of 1 and a minimum value of 0. Its calculation method can be briefly described as:

[0068]

[0069] PSDS1 means that the system needs to react quickly when detecting an event (e.g. triggering an alarm, adapting to a home automation system, etc.), which means that recall is important; PSDS2 means that the system must avoid confusion between classes, but reaction time is not that important, which means that precision is important. The calculation method of PSDS can be briefly described as:

[0070]

[0071] Among them, e max is the maximum eFPR (effective false positive rate), and r(e) is the ROC curve calculated under multiple operating points.

[0072] The audio data of the test set is preprocessed and feature extracted, and the acoustic features of the test data are used as the student model to predict the acoustic events contained in the test data. The predicted acoustic event category is the event category corresponding to the maximum probability in the output score of the student model, and the F1 score, PSDS1 and PSDS2 are calculated based on the output results of the test data.

[0073] The method provided in the above embodiment of the present application is superior to the model that does not use the training method in terms of event-based F1 score. The score fusion method proposed in the present application has improved Event-based F1, PSDS1, and PSDS2, which shows that the method proposed in the present invention has a significant effect on improving detection performance.

[0074] Further references Figure 5 As an implementation of the methods shown in the above figures, the present application provides an embodiment of an acoustic event detection device, which is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0075] like Figure 5As shown, the acoustic event detection device 500 of this embodiment includes: an acquisition unit 501, a preprocessing unit 502, an input unit 503, and a determination unit 504. The acquisition unit is configured to acquire the target audio; the preprocessing unit is configured to preprocess the target audio; the input unit is configured to input the preprocessed target audio into the first acoustic event detection model trained with the original training samples and the second acoustic event detection model trained with the high-quality training samples; the determination unit is configured to determine the acoustic event category of each sound segment in the target audio according to the outputs of the first acoustic event detection model and the second acoustic event detection model.

[0076] In this embodiment, the specific processing of the acquisition unit 501, the preprocessing unit 502, the input unit 503, and the determination unit 504 of the acoustic event detection device 500 can refer to Figure 2 This corresponds to step 201, step 202, step 203 and step 204 in the embodiment.

[0077] In some optional implementations of the present embodiment, the device also includes a screening unit, which is configured to: obtain separated audio segments included in the original audio data and first type information corresponding to each separated audio segment through a pre-trained speech separation model; label the separated audio segments through a pre-trained acoustic event detection model to obtain second type information corresponding to each separated audio segment; screen out separated audio segments whose corresponding first type information is the same as the second type information; determine the screened separated audio segments as high-quality training samples, and determine the first type information or the second type information corresponding to the screened separated audio segments as labels of high-quality training samples.

[0078] In some optional implementations of this embodiment, the screening unit is further configured to: screen out the separated audio segments whose corresponding first type information is the same as the second type information, and the first type information or the second type information is of the target event type.

[0079] In some optional implementations of this embodiment, the device also includes an adjustment unit, which is configured to: superimpose the separated audio segments in the high-quality training samples to form a mixed audio; input the mixed audio into the third acoustic event detection model to obtain a first prediction result; input the separated audio segments in the high-quality training samples into the fourth acoustic event detection model in sequence, and generate corresponding second prediction results in sequence; adjust the weight parameters of the first prediction result and the second prediction result and the parameters of the fourth acoustic event detection model according to the first prediction result and the second prediction result and the labels of the high-quality training samples.

[0080] In some optional implementations of this embodiment, the second acoustic event detection model includes a fourth acoustic event detection model with adjusted parameters, and / or a fifth acoustic event detection model obtained by weighting the third acoustic event detection model and the fourth acoustic event detection model with adjusted parameters according to the adjusted weight parameters.

[0081] In some optional implementations of this embodiment, the determination unit is further configured to: adaptively weighted fuse the output results of the first acoustic event detection model and the second acoustic event detection model based on the class discrimination of the first acoustic event detection model and the second acoustic event detection model; and determine the acoustic event category of each sound segment in the target audio according to the output results after weighted fusion.

[0082] The device provided by the above-mentioned embodiments of the present application obtains target audio; preprocesses the target audio; inputs the preprocessed target audio into a first acoustic event detection model trained with original training samples and a second acoustic event detection model trained by adjusting high-quality training samples; and determines the acoustic event category of each sound segment in the target audio according to the output of the first acoustic event detection model and the second acoustic event detection model, thereby providing a multi-model-based acoustic event detection mechanism and improving the detection effect of acoustic event detection.

[0083] Reference below Figure 6 , which shows a schematic diagram of the structure of a computer system 600 suitable for implementing a server or terminal of an embodiment of the present application. Figure 6 The server or terminal shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0084] like Figure 6 As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0085] The following components can be connected to the I / O interface 605: an input section 606 including such as a keyboard, a mouse, etc.; an output section 607 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage section 608 as needed.

[0086] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable medium described in the present application can be a computer-readable signal medium or a computer-readable medium or any combination of the above two. The computer-readable medium can be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, an apparatus or a device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, an apparatus or a device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0087] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as C or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0088] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0089] The units involved in the embodiments described in the present application may be implemented by software or by hardware. The units described may also be provided in a processor, for example, may be described as: a processor including an acquisition unit, a preprocessing unit, an input unit, and a determination unit. The names of these units do not, in some cases, constitute limitations on the units themselves, for example, the acquisition unit may also be described as "a unit configured to acquire target audio".

[0090] As another aspect, the present application also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently without being assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the device, the device: obtains the target audio; preprocesses the target audio; inputs the preprocessed target audio into the first acoustic event detection model trained with the original training samples and the second acoustic event detection model trained with the high-quality training samples; determines the acoustic event category of each sound segment in the target audio according to the output of the first acoustic event detection model and the second acoustic event detection model.

[0091] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.

Claims

1. A method for detecting an acoustic event, comprising: Get the target audio; Preprocessing the target audio; Inputting the preprocessed target audio into a first acoustic event detection model trained with original training samples and a second acoustic event detection model trained with high-quality training samples; Determining the acoustic event category of each sound segment in the target audio according to the outputs of the first acoustic event detection model and the second acoustic event detection model; The high-quality training samples include training samples obtained by screening through the following steps: Obtaining separated audio segments included in the original audio data and first type information corresponding to each separated audio segment through a pre-trained speech separation model; Labeling the separated audio segments by using a pre-trained acoustic event detection model to obtain second type information corresponding to each separated audio segment; Filter out separated audio segments corresponding to the same first type of information and second type of information; The screened separated audio segment is determined as the high-quality training sample, and the first type information or the second type information corresponding to the screened separated audio segment is determined as a label of the high-quality training sample.

2. The method according to claim 1, wherein: The step of screening out the separated audio segments corresponding to the same first type of information and second type of information includes: The separated audio segments whose corresponding first type information and second type information are the same and whose first type information or second type information is of the target event type are screened out.

3. The method according to any one of claims 1 to 2, wherein: The method further comprises the step of adjusting and training the second acoustic event detection model, wherein the step of adjusting and training the second acoustic event detection model comprises: superimposing the separated audio segments in the high-quality training samples to form mixed audio; Inputting the mixed audio into a third acoustic event detection model to obtain a first prediction result; Inputting the separated audio segments in the high-quality training samples into the fourth acoustic event detection model in sequence, and generating corresponding second prediction results in sequence; According to the first prediction result, the second prediction result and the label of the high-quality training sample, the weight parameters of the weighted weighting of the first prediction result and the second prediction result and the parameters of the fourth acoustic event detection model are adjusted.

4. The method according to claim 3, wherein: The second acoustic event detection model includes a fourth acoustic event detection model with adjusted parameters, and / or a fifth acoustic event detection model obtained by weighting the third acoustic event detection model and the fourth acoustic event detection model with adjusted parameters according to the adjusted weight parameters.

5. The method according to claim 1, wherein: The step of determining the acoustic event category of each sound segment in the target audio according to the outputs of the first acoustic event detection model and the second acoustic event detection model includes: Adaptively weighting and fusing output results of the first acoustic event detection model and the second acoustic event detection model based on class distinctions of the first acoustic event detection model and the second acoustic event detection model; The acoustic event category of each sound segment in the target audio is determined according to the output result after weighted fusion.

6. An acoustic event detection device, comprising: An acquisition unit, configured to acquire target audio; A preprocessing unit, configured to preprocess the target audio; An input unit configured to input the preprocessed target audio into a first acoustic event detection model trained with original training samples and a second acoustic event detection model trained with high-quality training samples; a determining unit, configured to determine the acoustic event category of each sound segment in the target audio according to the outputs of the first acoustic event detection model and the second acoustic event detection model; The device further comprises a screening unit, wherein the screening unit is configured to: Obtaining separated audio segments included in the original audio data and first type information corresponding to each separated audio segment through a pre-trained speech separation model; Labeling the separated audio segments by using a pre-trained acoustic event detection model to obtain second type information corresponding to each separated audio segment; Filter out separated audio segments corresponding to the same first type of information and second type of information; The screened separated audio segment is determined as the high-quality training sample, and the first type information or the second type information corresponding to the screened separated audio segment is determined as a label of the high-quality training sample.

7. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 5.

8. A computer readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Audio event detection model training method and device

    CN113724740A