Squeal Detection Method, Device, Medium, and Computing Device

By preprocessing the audio signal and cascade neural network detection, the accuracy and real-time problems of howling detection in conference scenes are solved, the error detection rate is reduced, and the call quality is improved.

CN114067837BActive Publication Date: 2025-05-27HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111347480.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-15
Publication Date
2025-05-27
Estimated Expiration
2041-11-15

AI Technical Summary

Technical Problem

In the conference scenarios of voice communication and multimedia communication, the howling phenomenon caused by equipment problems and environmental problems is complex and changeable, making it difficult to accurately measure, resulting in high error detection rates, which seriously affects call quality and user experience.

Method used

A howling detection method is used to obtain the audio signal to be detected for preprocessing, extract audio feature data, and use the trained convolutional neural network and recurrent neural network cascade model for detection, to obtain the detection results.

Benefits of technology

It improves the accuracy of detection of nonlinear howling, reduces the false detection rate, improves call quality and user experience, and ensures real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067837B_ABST
    Figure CN114067837B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a howling detection method, apparatus, medium, and computing device. The method includes: obtaining an audio signal to be detected, and preprocessing the audio signal to be detected to obtain audio feature data to be detected; obtaining a detection result based on the audio feature data to be detected through a target howling detection model; the detection result at least includes howling attribute information; the target howling detection model is obtained by training a first howling detection model based on sample audio signals collected in different communication scenarios and corresponding howling annotation information; the first howling detection model includes a convolutional neural network and a recurrent neural network cascaded in sequence. The present disclosure can detect howls of different attributes, reduce the false detection rate, and improve the call quality and the experience of participants in a meeting.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] This section aims to provide background or context for the embodiments of the present disclosure stated in the claims. The description herein is not admitted to be prior art merely by including it in this section.

[0003] In the fields of voice communication and multimedia communication, in a conference call or multimedia conference scenario, due to a large number of participants, howling phenomena caused by equipment problems or environmental problems are very likely to occur, seriously affecting the quality of the conference call. Therefore, for a specific scenario, how to accurately identify howling in voice becomes the key to determining the quality of the conference call.

[0004] However, in existing conference call scenarios, there are phenomena such as complex and variable network transmission, call environment, changing positions of participating devices, and frequency response differences between participating devices, etc., making the generated howling complex and variable, with non-linear characteristics, difficult to accurately measure, resulting in a high false detection rate, seriously affecting the call quality and the subjective experience of participants. Summary of the Invention

[0005] Embodiments of the present disclosure provide a howling detection method, a howling detection device, a medium, and a computing device.

[0006] In the first aspect of the embodiments of the present disclosure, a howling detection method is provided, including:

[0007] Obtain a to-be-detected audio signal, and preprocess the to-be-detected audio signal to obtain to-be-detected audio feature data;

[0008] Obtain a detection result through a target howling detection model based on the to-be-detected audio feature data; the detection result at least includes howling attribute information;

[0009] Wherein, the target howling detection model is obtained by training a first howling detection model based on sample audio signals collected in different communication scenarios and corresponding howling annotation information; the first howling detection model includes a convolutional neural network and a recurrent neural network cascaded in sequence.

[0010] In some embodiments of the present disclosure, based on the foregoing solution, the preprocessing the to-be-detected audio signal to obtain to-be-detected audio feature data includes:

[0011] Resample the to-be-detected audio signal to normalize the to-be-detected audio signal to a specified sampling rate;

[0012] Perform frame segmentation on the normalized to-be-detected audio signal;

[0013] Extract features from one frame of the segmented to-be-detected audio signal to obtain to-be-detected audio feature data.

[0014] In some embodiments of the present disclosure, based on the foregoing solution, the training of the first howling detection model includes:

[0015] Based on the convolutional neural network, perform convolutional processing on the feature data of the sample audio signal to output a first feature vector, where the first feature vector contains temporal information;

[0016] Based on the recurrent neural network, perform temporal feature learning on the first feature vector to output a second feature vector;

[0017] Perform focusing processing on the second feature vector to output a howling attribute probability distribution vector, where the howling attribute probability distribution vector represents the probabilities corresponding to each howling attribute;

[0018] Based on the probability of the howling attribute corresponding to the feature data of the sample audio signal and the howling annotation information, use a loss function to determine the target loss information;

[0019] According to the target loss information, adjust the parameters of the first howling detection model.

[0020] In some embodiments of the present disclosure, based on the foregoing solution, the training of the first howling detection model further includes:

[0021] Prune the neurons and their connection relationships in the first howling detection model based on a specified measurement criterion; and / or

[0022] Quantize the parameters of the first howling detection model based on a specified quantization criterion.

[0023] In some embodiments of the present disclosure, based on the foregoing solution, the howling annotation information includes one or more of whether there is howling, howling type information, and howling level information.

[0024] In some embodiments of the present disclosure, based on the foregoing solution, the step of using a loss function to determine the target loss information based on the probability of the howling attribute corresponding to the feature data of the sample audio signal and the howling annotation information includes at least one of the following:

[0025] When the howling annotation information includes whether there is howling, use a weighted binary cross-entropy loss function to determine the target loss information;

[0026] When the howling annotation information includes whether there is howling and howling type information, use a weighted binary cross-entropy loss function for binary classification to determine the first loss information corresponding to whether there is howling and use a first loss function corresponding to multi-classification to determine the second loss information corresponding to the howling type information, and obtain the target loss information based on the first loss information and the second loss information;

[0027] When the howling annotation information includes whether there is howling and howling level information, a weighted binary cross-entropy loss function for binary classification is used to determine the first loss information corresponding to whether there is howling, and a second loss function corresponding to multi-classification is used to determine the third loss information corresponding to the howling level information, and the target loss information is obtained based on the first loss information and the third loss information;

[0028] When the howling annotation information includes whether there is howling, howling type information, and howling level information, a weighted binary cross-entropy loss function for binary classification is used to determine the first loss information corresponding to whether there is howling, a first loss function corresponding to multi-classification is used to determine the second loss information corresponding to the howling type, and a second loss function corresponding to multi-classification is used to determine the third loss information corresponding to the howling level information, and the target loss information is obtained based on the first loss information, the second loss information, and the third loss information.

[0029] In some embodiments of the present disclosure, based on the foregoing solution, the convolutional processing of the feature data of the sample audio signal by the convolutional neural network to output a first feature vector includes:

[0030] For different howling types, the convolutional layer parameters of the convolutional neural network are discarded to make the target howling detection model applicable to the detection of different howling types.

[0031] In some embodiments of the present disclosure, based on the foregoing solution, the training of the first howling detection model includes:

[0032] Performing parameter regularization and / or parameter normalization on the model parameters of the first howling detection model.

[0033] In some embodiments of the present disclosure, based on the foregoing solution, the obtaining of the audio feature data to be detected by performing feature extraction on a frame of the audio signal to be detected after frame division includes:

[0034] Performing windowing processing on a frame of the audio signal to be detected after frame division;

[0035] Performing a fast Fourier transform on the windowed audio signal to be detected to obtain a corresponding frequency-domain signal to be detected;

[0036] Filtering the frequency-domain signal to be detected through a corresponding frequency-domain filter;

[0037] Converting the filtered frequency-domain signal to the logarithmic domain to obtain the audio feature data to be detected.

[0038] In some embodiments of the present disclosure, based on the foregoing solution, the method further includes: obtaining the sample audio signal by collecting audio signals played by a microphone and / or different performance communication devices.

[0039] In some embodiments of the present disclosure, based on the foregoing solution, after the target whistling detection model obtains a detection result based on the audio feature data to be detected, the method further includes:

[0040] Comparing the number of frames of the audio signal corresponding to the specified whistling attribute information in the detection result with a preset threshold;

[0041] Determining the whistling attribute information of the audio signal to be detected according to the comparison result.

[0042] In some embodiments of the present disclosure, based on the foregoing solution, after the target whistling detection model obtains a detection result based on the audio feature data to be detected, the method further includes:

[0043] For a specific frame signal, averaging the detection results of the frame signal in consecutive specified detections, and comparing the processed result with a posterior threshold;

[0044] Determining the whistling attribute information of the audio signal to be detected according to the comparison result.

[0045] In the second aspect of the embodiments of the present invention, a whistling detection device is provided, including:

[0046] A preprocessing module, configured to obtain an audio signal to be detected, and preprocess the audio signal to be detected to obtain audio feature data to be detected;

[0047] A detection module, configured to obtain a detection result based on the audio feature data to be detected through a target whistling detection model; the detection result at least includes whistling attribute information;

[0048] Wherein, the target whistling detection model is obtained by training a first whistling detection model based on sample audio signals collected in different communication scenarios and corresponding whistling annotation information; the first whistling detection model includes a cascaded convolutional neural network and a recurrent neural network in sequence.

[0049] In some embodiments of the present disclosure, based on the foregoing solution, the preprocessing module includes:

[0050] A resampling module, configured to resample the audio signal to be detected, so that the audio signal to be detected is normalized to a specified sampling rate;

[0051] A framing module, configured to perform framing processing on the normalized audio signal to be detected;

[0052] A feature extraction module, configured to extract features from a frame of audio signal to be detected after frame division to obtain audio feature data to be detected.

[0053] In some embodiments of the present disclosure, based on the foregoing solution, the apparatus further includes a training module, and the training module is configured to:

[0054] Based on the convolutional neural network, perform convolutional processing on the feature data of the sample audio signal, and output a first feature vector, where the first feature vector contains temporal information;

[0055] Based on the recurrent neural network, perform temporal feature learning on the first feature vector, and output a second feature vector;

[0056] Perform focusing processing on the second feature vector, and output a howling attribute probability distribution vector, where the howling attribute probability distribution vector represents the probabilities corresponding to each howling attribute;

[0057] Based on the probability of the howling attribute corresponding to the feature data of the sample audio signal and the howling annotation information, use a loss function to determine target loss information;

[0058] According to the target loss information, adjust the parameters of the first howling detection model.

[0059] In some embodiments of the present disclosure, based on the foregoing solution, the training module is further configured to:

[0060] Prune the neurons and their connection relationships in the first howling detection model based on a specified measurement criterion; and / or

[0061] Quantize the parameters of the first howling detection model based on a specified quantization criterion.

[0062] In some embodiments of the present disclosure, based on the foregoing solution, the howling annotation information includes one or more of whether it contains howling, howling type information, and howling level information.

[0063] In some embodiments of the present disclosure, based on the foregoing solution, the training module is further configured to:

[0064] When the howling annotation information includes whether it contains howling, use a weighted binary cross-entropy loss function to determine the target loss information;

[0065] When the howling annotation information includes whether it contains howling and howling type information, use a weighted binary cross-entropy loss function for binary classification to determine the first loss information corresponding to whether it contains howling and use a first loss function corresponding to multi-classification to determine the second loss information corresponding to the howling type information, and obtain the target loss information based on the first loss information and the second loss information;

[0066] When the howling annotation information includes whether there is howling and howling level information, a weighted binary cross-entropy loss function for binary classification is used to determine the first loss information corresponding to whether there is howling, and a second loss function corresponding to multi-classification is used to determine the third loss information corresponding to the howling level information, and the target loss information is obtained based on the first loss information and the third loss information;

[0067] When the howling annotation information includes whether there is howling, howling type information, and howling level information, a weighted binary cross-entropy loss function for binary classification is used to determine the first loss information corresponding to whether there is howling, a first loss function corresponding to multi-classification is used to determine the second loss information corresponding to the howling type, and a second loss function corresponding to multi-classification is used to determine the third loss information corresponding to the howling level information, and the target loss information is obtained based on the first loss information, the second loss information, and the third loss information.

[0068] In some embodiments of the present disclosure, based on the foregoing solution, the training module is further configured to:

[0069] For different howling types, perform dropout processing on the convolutional layer parameters of the convolutional neural network so that the target howling detection model is applicable to the detection of different howling types.

[0070] In some embodiments of the present disclosure, based on the foregoing solution, the training module is configured to:

[0071] Perform parameter regularization and / or parameter normalization on the model parameters of the first howling detection model.

[0072] In some embodiments of the present disclosure, based on the foregoing solution, the feature extraction module includes:

[0073] A windowing sub-module for windowing a frame of audio signal to be detected after frame segmentation;

[0074] A transformation sub-module for performing a fast Fourier transform on the windowed audio signal to be detected to obtain a corresponding frequency-domain signal to be detected;

[0075] A filtering sub-module for filtering the frequency-domain signal to be detected through a corresponding frequency-domain filter;

[0076] A conversion sub-module for converting the filtered frequency-domain signal to the logarithmic domain to obtain audio feature data to be detected.

[0077] In some embodiments of the present disclosure, based on the foregoing solution, the device further includes:

[0078] A sample acquisition module, configured to acquire the sample audio signal by collecting audio signals played by a microphone and / or communication devices with different performances.

[0079] In some embodiments of the present disclosure, based on the foregoing solution, the apparatus further includes a first post-processing module, and the first post-processing module is configured to:

[0080] Compare the number of frames of the audio signal corresponding to the specified howling attribute information in the detection result with a preset threshold;

[0081] Determine the howling attribute information of the audio signal to be detected according to the comparison result.

[0082] In some embodiments of the present disclosure, based on the foregoing solution, the apparatus further includes a second post-processing module, and the second post-processing module is configured to:

[0083] For a specific frame signal, perform averaging processing on the detection results of the frame signal in consecutive specified detections, and compare the processing result with a posterior threshold;

[0084] Determine the howling attribute information of the audio signal to be detected according to the comparison result.

[0085] In a third aspect of the embodiments of the present invention, a medium is provided, on which a program is stored, and when the program is executed by a processor, the howling detection method described in the above embodiments is implemented.

[0086] In a fourth aspect of the embodiments of the present invention, a computing device is provided, including: a processor and a memory, the memory stores executable instructions, and the processor is configured to call the executable instructions stored in the memory to execute the howling detection method described in the above embodiments.

[0087] According to the howling detection method provided by the embodiments of the present disclosure, preprocessing is performed on the audio signal to be detected to obtain the audio feature data to be detected; the detection result is obtained based on the audio feature data to be detected through the target howling detection model; the detection result includes howling attribute information. On the one hand, by training the first howling detection model with sample audio signals in different communication scenarios, the trained target howling detection model is robust to various howling scenarios, improving the detection accuracy of non-linear howling; on the other hand, the detection result pays more attention to the attribute characteristics of the howling itself, can detect howling with different attributes, reduce the false detection rate, and improve the call quality and the experience of conference participants. In addition, by using a cascaded convolutional neural network and a recurrent neural network as the howling detection network, the real-time performance of the detection can be ensured while ensuring the detection accuracy. Description of the Drawings

[0088] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understandable. In the drawings, several embodiments of the present disclosure are shown by way of illustration and not limitation, wherein:

[0089] Figure 1 Schematically shows a block diagram of a feedback system according to an embodiment of the present disclosure;

[0090] Figure 2 Schematically shows a schematic diagram of a typical howling scenario according to an embodiment of the present disclosure;

[0091] Figure 3 Schematically shows a howling spectrogram according to an embodiment of the present disclosure, wherein (a) is a single-frequency howling spectrogram and (b) is a multi-frequency continuous howling spectrogram;

[0092] Figure 4 Schematically shows a schematic diagram of an acoustic loop in a real-time communication scenario according to an embodiment of the present disclosure;

[0093] Figure 5 Schematically shows a howling spectrogram in a real-time communication scenario according to an embodiment of the present disclosure;

[0094] Figure 6 Schematically shows a flowchart of voice processing in a real-time communication scenario according to an embodiment of the present disclosure;

[0095] Figure 7 Schematically shows a schematic diagram of a howling detection method process according to an embodiment of the present disclosure;

[0096] Figure 8 Schematically shows the structure and processing flowchart of a target howling detection model according to an embodiment of the present disclosure;

[0097] Figure 9 Schematically shows a flowchart of the training process implementation of a first howling detection model according to an embodiment of the present disclosure;

[0098] Figure 10 Schematically shows a flowchart of the implementation of a howling detection method according to an embodiment of the present disclosure;

[0099] Figure 11 Schematically shows a block diagram of a howling detection device structure according to an embodiment of the present disclosure;

[0100] Figure 12 Schematically shows a schematic diagram of the structure of a storage medium suitable for implementing the embodiments of the present invention;

[0101] Figure 13A schematic structural diagram of a computing device suitable for implementing the embodiments of the present invention is shown.

[0102] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed implementation manners

[0103] The principles and spirit of the present disclosure will be described below with reference to several exemplary implementation manners. It should be understood that these implementation manners are provided only to enable those skilled in the art to better understand and then implement the present disclosure, and do not limit the scope of the present disclosure in any way. On the contrary, these implementation manners are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.

[0104] Those skilled in the art know that the implementation manners of the present disclosure can be implemented as a system, a device, an equipment, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0105] According to the implementation manners of the present disclosure, a howling detection method, a howling detection device, a medium, and a computing device are provided.

[0106] The main terms mentioned in the present disclosure are explained as follows:

[0107] Convolutional neural network: In machine learning, it is a feedforward neural network in which artificial neurons can respond to surrounding units. A convolutional neural network includes a convolutional layer and a pooling layer. A convolutional neural network includes a one-dimensional convolutional neural network, a two-dimensional convolutional neural network, and a three-dimensional convolutional neural network. In the embodiments of the present application, it may refer to a two-dimensional convolutional neural network including several convolutional modules, and each convolutional module includes a convolutional layer and a pooling layer.

[0108] Recurrent neural network: It is a class of recursive neural networks that take sequential data as input, perform recursion in the evolution direction of the sequence, and all nodes are connected in a chain.

[0109] Loss function: It is a function that maps the values of a random event or its related random variables to non-negative real numbers to represent the "risk" or "loss" of the random event. In applications, the loss function is usually associated with a learning criterion and an optimization problem, that is, the model is solved and evaluated by minimizing the loss function. For example, in machine learning, the loss function is used for parameter estimation of the model, and the loss value obtained based on the loss function can be used to describe the difference between the predicted value and the actual value of the model. Common loss functions include mean square error loss function, cross-entropy loss function, etc.

[0110] 3A processing refers to the collective term for AGC, AEC, and NS. Among them, AGC (Automatic Gain Control) represents automatic gain control, which is used to automatically adjust the receiving volume of the microphone so that the participants can receive a certain volume level, and there will be no problem of sudden changes in sound volume when the distance between the speaker and the microphone changes. AEC (Acoustic Echo Cancellation) represents acoustic echo cancellation. Based on the correlation between the speaker signal and the multipath echo generated by it, a voice model of the far-end signal is established, and it is used to estimate the echo and continuously modify the coefficients of the filter to make the estimated value closer to the real echo. Then, the echo estimate value is subtracted from the input signal of the microphone to achieve the purpose of echo cancellation. AEC also compares the input of the microphone with the past values of the speaker to eliminate the acoustic echo of multiple reflections with extended delays. According to the amount of the past output values of the speaker stored in the memory, AEC can cancel echoes of various delays. NS (Noise Suppression) represents noise suppression, which is used to detect background noise with a fixed frequency and eliminate background noise, such as the sound of fans and air conditioners, and automatically filter it out to present clear voices of the participants.

[0111] Howling: It is caused by the self-excitation of energy due to problems such as too close distance between the sound source and the sound amplification device. Howling is a kind of feedback sound. Its essence is that the feedback system is in an unstable state. The stability of the feedback system can be judged by using the Nyquist stability criterion through the open-loop transfer function of the system. A typical block diagram of a feedback system is as Figure 1 shown. Figure 1 In it, R(s) is the input signal of the system, C(s) is the output signal of the system, G(s) is the forward transfer function of the system, and H(s) is the feedback transfer function of the system. Therefore, the open-loop transfer function of the system can be deduced as: H(s)·G(s). Based on this, the stability of the system can be judged according to the Nyquist diagram or Bode diagram of the open-loop transfer function.

[0112] Here, a judgment basis for the unstable state of the system can be given, that is, when the feedback system meets the following two conditions, the system will be in an unstable state.

[0113] (1) The feedback signal and the input signal are in phase;

[0114] (2) The feedback loop is a positive feedback, that is, the corresponding open-loop gain is greater than 1.

[0115] In an acoustic scenario, howling is likely to occur when a sound feedback closed loop is formed. For example, as Figure 2As shown, the microphone picks up sound and the speaker plays sound. At this time, the signal played by the speaker is picked up by the microphone again, thus generating an acoustic loop. Common howling scenarios include KTV, auditorium meetings, headphones (such as noise-canceling headphones), etc. In these scenarios, the system often howls by itself, and the acoustic characteristics shown are mostly single-frequency or multi-frequency continuous howling. The spectrogram of its howling is as shown in Figure 3 shown. Figure 3 In (a) of Figure 3 it is a single-frequency howling spectrogram, and (b) is a multi-frequency continuous howling spectrogram. In the figure, the abscissa is time, with the unit of second, and the ordinate is frequency, with the unit of Hz. From Figure 3 (a), it can be seen that there is a relatively bright straight line between 600 - 700 Hz, and this straight line indicates the existence of continuous howling here. From Figure 3 (b), it can be seen that there are 7 relatively bright straight lines in the figure, indicating the existence of continuous howling at the corresponding seven frequency points. From Summary of the Invention

[0117] For the real-time communication RTC (Real Time Communication) scenario, for example, the scenario where two communication devices participate in a meeting at the same time. When two mobile phones join the meeting at the same time, if the mobile phones are relatively close to each other and an acoustic loop is generated, howling is likely to occur, as shown in Figure 4 shown:

[0118] When the user speaks, mobile phones A and B collect signals at the same time. For mobile phone A, the collected signal is path 1 (solid line), then it goes through process 2, is transmitted to network 3, and is transmitted to mobile phone B through the network; mobile phone B plays the sound. At this time, due to the close distance between A and B, mobile phone A will collect the sound played by mobile phone B, that is, path 5. At this time, it can be found that the solid line paths 2 - 5 form a loop. Similarly, for mobile phone B, the dotted line paths 2 - 5 also form a loop.

[0119] In the actual communication link, each communication device has its own 3A processing. Due to the existence of network transmission, environmental uncertainty, changes in device location, frequency response differences of devices, etc., the influence of these factors on the audio signal is non-linear, and it is impossible to use the traditional quantitative measurement transfer function to determine howling. And these non-linear factors will also cause many characteristics different from traditional howling scenarios, such as the intermittency of howling, multi-frequency point howling, howling frequency point movement, howling frequency point spread, etc. In addition, such as the uncertainty of AEC residue, the noise reduction module will process a part of the howling signal, etc., which will all cause these howling characteristics different from traditional howling. The spectrogram of the howling characteristics in a typical RTC scenario is as shown in Figure 5 shown. From Figure 5As can be seen, there is no obvious entire bright bar in the figure. Only relatively bright local areas exist, and the relatively bright areas scatter around, indicating the presence of frequency point diffusion here; and intermittent bright bars appear in the figure, and the positions of the bright bars move up or down, indicating the presence of intermittent whistling and frequency point movement; in addition, there are multiple bright bands in the figure, indicating multi-frequency point whistling. From Figure 5 it can be known that in the RTC scenario, the whistling signal has characteristics such as intermittency, multi-frequency points, frequency point movement, and frequency point diffusion. Compared with Figure 3 it is obvious that Figure 5 the whistling phenomenon in is more difficult to distinguish. That is to say, the whistling detection in such a complex scenario is more difficult.

[0120] The RTC whistling scenario with a 3A processing process description is as shown in Figure 6 . From Figure 6 it can be seen that the sound processing process can be divided into acoustic upstream signal processing and acoustic downstream signal processing. Both processing processes include 3A processing, namely acoustic echo cancellation, noise suppression, and automatic gain control, indicating that 3A processing is crucial for the communication quality in the real-time communication process. Therefore, this disclosure selects communication devices with different 3A processing capabilities for sample data collection to increase the robustness of the model to different devices.

[0121] Figure 6 Among them, the noise reduction processing in the 3A processing has a greater impact on the call quality. The noise tracking of noise reduction may track the whistling as noise and eliminate it to a certain extent. However, due to the still existing system acoustic loop, the whistling will still be generated again due to external excitation; on the other hand, if the noise reduction cannot completely eliminate the whistling, only a part of it is eliminated, intermittent, large and small whistling will be generated. At the same time, other non-linear processing, etc. will affect the phase amplitude characteristics of the system, causing phenomena such as changes and diffusion of the frequency points of the whistling. In addition, due to the frequency response differences in the collection and playback of different devices, the transfer functions of the device itself systems are inconsistent, so the whistling frequency points and characteristics generated by different communication devices (such as mobile phones) are also different.

[0122] The current whistling detection method is: training a whistling detection model based on machine learning through a large number of training samples to perform whistling detection, and the detection result is generally with whistling or without whistling. However, for a general machine learning process, it will introduce a large amount of computational workload and occupy a large amount of resources of the device, making it difficult to ensure the real-time performance of the detection. In addition, this simple binary classification detection method is difficult to apply to complex communication scenarios, especially for non-linear audio features, it is difficult to accurately detect, resulting in a high false detection rate and seriously affecting the call quality and user experience.

[0123] To this end, the present disclosure provides a howling detection method, which can be applied to a terminal device or a server, such as a cloud server or an application server. By obtaining an audio signal to be detected and preprocessing the audio signal to be detected to obtain audio feature data to be detected; obtaining a detection result based on the audio feature data to be detected through a target howling detection model; the detection result includes howling attribute information. On the one hand, by training a first howling detection model with sample audio signals in different communication scenarios, the trained target howling detection model has robustness to various howling scenarios, improving the detection accuracy of non-linear howling; on the other hand, the detection result pays more attention to the attribute characteristics of the howling itself, can detect howlings with different attributes, reduce the false detection rate, and improve the call quality and the experience of participants. In addition, by using a cascaded convolutional neural network and a recurrent neural network as the howling detection network, the real-time performance of the detection can be ensured while ensuring the detection accuracy.

[0124] Exemplary Method

[0125] The preferred embodiments of the present disclosure will be described below with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention. And without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0126] The following refers to Figure 7 to describe the howling detection method according to an exemplary embodiment of the present disclosure, including steps S710-S720.

[0127] Step S710: Obtain an audio signal to be detected, and preprocess the audio signal to be detected to obtain audio feature data to be detected.

[0128] In some example embodiments, the audio signal to be detected may be an audio signal collected in real time by a communication device, which is continuous-frame time-domain data. In order to perform howling detection, it is necessary to preprocess the continuous-frame time-domain data to obtain the time-domain features or frequency-domain features of each frame of data, such as log Mel spectrum, MFCC (Mel-Frequency Cepstral Coefficients), MGC (Mel cepstral coefficients), sub-band spectral energy, etc., that is, obtain audio feature data to be detected. The preprocessing of the data needs to be designed according to the actual situation, and factors such as the size of the detection model and the processing time need to be considered.

[0129] In some example embodiments, preprocessing the audio signal to be detected to obtain audio feature data to be detected includes:

[0130] Resample the audio signal to be detected to normalize it to a specified sampling rate. In the present exemplary embodiment, resampling of the audio signal may include upsampling and downsampling, corresponding to interpolation and decimation operations, or a combination of upsampling and downsampling, which is specifically determined according to actual business requirements. In the field of speech recognition, downsampling is generally used to reduce the data volume. For example, resample the audio signal to be detected to normalize it to a specified 16 kHz.

[0131] Perform frame segmentation on the normalized audio signal to be detected. In the present exemplary embodiment, segment the continuous-frame audio signal after normalization, set the frame length, and obtain multiple audio frame signals to be detected with the same frame length. For example, select a frame length of 40 ms for frame segmentation.

[0132] Extract features from one frame of the segmented audio signal to be detected to obtain audio feature data to be detected. Extract features from each frame of the segmented audio signal to be detected to obtain corresponding feature data, which can be time-domain feature data or frequency-domain feature data. This exemplary embodiment does not make special limitations on this.

[0133] In some exemplary embodiments, extracting features from one frame of the segmented audio signal to be detected to obtain audio feature data to be detected includes:

[0134] Perform windowing on one frame of the segmented audio signal to be detected; windowing is to intercept a signal of finite length for subsequent processing. A fixed window length and overlap length can be selected for windowing, and the window function for windowing can be selected as needed. For example, rectangular window, triangular window, Gaussian window, etc. In this exemplary embodiment, a rectangular window is selected for windowing, with a window length of 80 ms and an overlap length of 40 ms, that is, there is an overlap length of 40 ms between the signals intercepted by two consecutive windowings.

[0135] Perform a fast Fourier transform on the windowed audio signal to be detected to obtain the corresponding frequency-domain signal to be detected. Convert the signal from the time domain to the frequency domain through a fast Fourier transform to obtain the corresponding frequency-domain signal.

[0136] Filter the frequency-domain signal to be detected through a corresponding frequency-domain filter. When the selected feature data is a log Mel spectrum, filter through a Mel filter.

[0137] Convert the filtered frequency-domain signal to the logarithmic domain to obtain audio feature data to be detected. When the selected feature data is not in the logarithmic domain, this step is omitted. In this exemplary embodiment, the selected feature data is logarithmic domain features, such as log Mel spectrum, so it is necessary to convert the filtered data to the logarithmic domain. In addition, the dimension of the feature data can also be set to avoid excessive data processing volume of the model.

[0138] Step 720: Obtain a detection result based on the audio feature data to be detected through the target whistling detection model; the detection result at least includes whistling attribute information.

[0139] The target whistling detection model can be obtained by training a first whistling detection model based on sample audio signals collected in different communication scenarios and corresponding whistling annotation information. The first whistling detection model includes a convolutional neural network and a recurrent neural network cascaded in sequence. Different communication scenarios can include a quiet environment, a noisy environment, a whistling scenario, a non-whistling scenario, a single-device communication scenario, a multi-device communication scenario, and so on.

[0140] In some exemplary embodiments, the whistling annotation information may include one or more of whether there is whistling, whistling type information, and whistling level information. The whistling type information may include one or more of continuous whistling, intermittent whistling, single-frequency-point whistling, multi-frequency-point whistling, and diffusive whistling. The whistling level information can divide whistling into several levels according to the strength of the whistling, for example, strong, medium, weak, etc., and can also make a more detailed division of the whistling level according to requirements. This example does not make special limitations on this.

[0141] The first whistling detection model includes a convolutional neural network and a recurrent neural network cascaded in sequence. The first whistling detection model can be a CRNN, that is, a convolutional recurrent neural network. For example, as Figure 8 shown, it is the cascaded diagram of the structure of the first whistling detection model in this example. The model includes an input layer, a convolutional neural network unit, a recurrent neural network unit, a fully connected unit, and an output layer. Among them, the input layer is used to input log Mel spectrogram. The convolutional neural network unit includes three convolutional layers, and a pooling layer is set after each convolutional layer. After the convolutional kernel of each convolutional layer, there are also processes such as regularization (Relu), batch normalization (BN), and parameter dropout. In this example, parameter dropout processing is added to each convolutional layer to make the target whistling detection model applicable to the detection of different whistling types, that is, it is robust to different whistling types. In addition, it can also prevent overfitting. The recurrent neural network unit includes two cascaded bidirectional long short-term memory networks, and its number of channels is the same as that of the last convolutional layer. The fully connected unit is used to integrate the features extracted by each neuron to obtain global features. The output layer is a single-channel fully connected layer for outputting the whistling attribute probability distribution. Figure 8 The output example in is the output for the whistling attribute. Figure 8 The model parameters and the output data size of each layer are shown in Table 1.

[0142] Table 1 Structure of the First Whistling Detection Model

[0143]

[0144]

[0145] In some embodiments, as Figure 8 shown, the first howling detection model may include an input layer, a convolutional neural network, a recurrent neural network, a fully connected unit, and an output layer. Among them, the Log Mel spectrum of the input sample audio data is input in the input layer, and its size is 1×32×60 (number of channels × number of frames × number of dimensions). The convolutional neural network includes three cascaded convolutional modules. The first convolutional module includes a 3×3 convolutional kernel, a batch normalization layer, and a 1×5 average pooling layer, and its output size is 16×32×12; the second convolutional module includes a 3×3 convolutional kernel, a batch normalization layer, and a 1×4 average pooling layer, and its output size is 32×32×3; the third convolutional module includes a 3×3 convolutional kernel, a batch normalization layer, and a 1×3 average pooling layer, and its output size is 32×32×1; each convolutional module adopts parameter regularization processing, and dropout is added after each convolutional module. The recurrent neural network includes two cascaded bidirectional long short-term memory networks, and the number of channels of each bidirectional long short-term memory network is 32, and the output size is 32×32×1. The fully connected unit may include a fully connected linear layer with 32 channels, and dropout is added after it; the output size of each bidirectional long short-term memory network is 32×32×1. The output layer may be a single-channel linear layer, and its output size is 1×32×1, that is, a vector composed of 0 or / and 1.

[0146] In some embodiments, the output howling attributes further include howling type and / or howling level. It is necessary to extract features from the sample data corresponding to the howling type and / or howling level, and splice the extracted feature data (such as the dimension is 32*1) with the feature data extracted from the sample data corresponding to whether there is howling ( Figure 8 the Log Mel spectrum therein, with the dimension of 32*60) to form input data (with the dimension of 32*61). Correspondingly, the output channels also need to be added. For example, when the input includes whether there is howling and the howling type, the number of channels of the output layer corresponds to two, one channel outputs whether there is howling, and one channel outputs the howling type.

[0147] The training of the first howling detection model may include processes such as sample data collection, sample data annotation, and model training. Each process is described in detail below.

[0148] (1) Sample data collection:

[0149] In order to obtain audio data in complex communication scenarios, sample data collection is carried out in the actual communication scenario in this example. The communication scenario takes into account signals, devices, environments, and scenarios, etc.

[0150] In some embodiments, the sample audio signal is obtained by collecting audio signals played by a microphone and / or a communication device with different performances. Specifically, the collected signal may include an input signal collected by a microphone of the communication device and transmitted to the 3A processing module after analog-to-digital conversion, and may also include a signal played by the communication device, which may include voice, music, noise, ambient sound, and some special sounds, such as bells, bird calls, whistles, etc.

[0151] In some embodiments, the communication devices collected involve the frequency response characteristics of each device and the diversity of the 3A processing algorithms of the adapted devices, which has an important impact on the robustness of the device-related algorithms. Therefore, it is necessary to select devices with different performance, styles, and processing capabilities and 3A processing algorithms for sample collection.

[0152] In some embodiments, the collected environment may cover quiet, noisy, and other environments with different signal-to-noise ratios and reverberation.

[0153] In some embodiments, the collected scenarios may include scenarios with howling and without howling, including single device joining a conference, multiple devices joining a conference, devices being in different physical locations, and the like.

[0154] In the above examples, the diversity of communication scenarios and devices can ensure the robustness of the target howling detection model finally trained, so that it can have high detection accuracy for nonlinear signals in complex communication scenarios.

[0155] (II) Sample Data Labeling

[0156] In order to cope with the nonlinear characteristics of complex communication scenarios, howling can be classified. In addition to the traditional categories corresponding to whether or not there is howling, howling can be divided into more detailed categories based on the characteristics of the howling itself, such as continuous howling, intermittent howling, single-frequency howling, multi-frequency howling, diffuse howling, etc.

[0157] Sample data annotation refers to annotating the howling categories of the collected sample data. Since the howling in the RTC scenario has various characteristics, refined and diverse howling annotations are required. For example, frame-level annotation is performed according to whether there is howling. For complex scenarios, such as intermittent annotations, they are also marked as howling. Further, a general classification can be made according to the characteristics of howling, such as continuous howling, intermittent howling, single-frequency howling, multi-frequency howling, diffusive howling, etc., and each category is annotated separately. The howling level can also be separately annotated according to the strength of howling. In addition, the capabilities of the 3A processing algorithms of different devices can be graded, such as the 3A processing intensity being divided into different levels such as strong, medium, and weak. The howling detection model trained with the above diverse and detailed annotation information can detect the howling type and howling level information in addition to whether there is howling information, enabling subsequent targeted suppression of each type of howling, reducing the false detection rate, and improving the howling suppression effect.

[0158] (III) Model Training

[0159] This can be achieved through the following process: First, initialize the parameters of the model; the learning rate and the maximum number of training epochs can also be set. Then, input the feature data of the sample audio data with howling annotation information into the initialized model to train the model. Based on the stochastic gradient descent method, use the loss function to backpropagate and update the model parameters until the model converges or reaches the specified number of training epochs, and stop the training process to obtain the target howling detection model.

[0160] Reference Figure 9 , the training process of each epoch inside the model includes:

[0161] Step S910, based on the convolutional neural network, perform convolutional processing on the feature data of the sample audio signal, and output a first feature vector, where the first feature vector contains temporal information. In the present exemplary embodiment, the convolutional neural network includes multiple convolutional layers, and each convolutional layer performs convolutional operation processing on the feature data of the sample audio signal through a convolutional kernel to extract the local features of the sample audio signal and obtain the first feature vector.

[0162] Step S920, based on the recurrent neural network, perform temporal feature learning on the first feature vector, and output a second feature vector; in the present exemplary embodiment, the recurrent neural network includes an input layer, an output layer, and multiple hidden layers, and the input of each hidden layer also includes the output of the previous hidden layer, and in this way, learn the temporal features of the first feature vector and output the second feature vector.

[0163] Step S930: Perform focusing processing on the second feature vector and output a howling attribute probability distribution vector, where the howling attribute probability distribution vector represents the probabilities corresponding to each howling attribute. In the present exemplary embodiment, the second feature vector can be focused through a fully connected layer to integrate the features extracted by the convolutional neural network and the recurrent neural network.

[0164] Step S940: Based on the probability of the howling attribute corresponding to the feature data of the sample audio signal and the howling annotation information, use a loss function to determine the target loss information. In the present exemplary embodiment, the target loss information is determined by calculating the error information between the probability of the howling attribute corresponding to the feature data of the sample audio signal and the corresponding howling annotation information and using a loss function.

[0165] In the present exemplary embodiment, the loss function can be determined by one of the following methods.

[0166] Embodiment 1: When the howling annotation information includes whether there is howling, use a weighted binary cross-entropy loss function to determine the target loss information. In this embodiment, the calculation result of the weighted binary cross-entropy loss function can be used as the target loss information.

[0167] Embodiment 2: When the howling annotation information includes whether there is howling and howling type information, use a weighted binary cross-entropy loss function for binary classification to determine the first loss information corresponding to whether there is howling and use a first loss function corresponding to multi-classification to determine the second loss information corresponding to the howling type information, and obtain the target loss information based on the first loss information and the second loss information. In this embodiment, the first loss information and the second loss information can be subjected to a relevant operation, and the result of the relevant operation can be used as the target loss information. This relevant operation can be direct summation, or weighted summation of the first loss information and the second loss information respectively and then summation, or multiplication, etc. This example does not make special limitations on this.

[0168] Embodiment 3: When the howling annotation information includes whether there is howling and howling level information, use a weighted binary cross-entropy loss function for binary classification to determine the first loss information corresponding to whether there is howling and use a second loss function corresponding to multi-classification to determine the third loss information corresponding to the howling level information, and obtain the target loss information based on the first loss information and the third loss information. In this embodiment, the first loss information and the third loss information can be subjected to a relevant operation, and the result of the relevant operation can be used as the target loss information. This relevant operation can be direct summation, or weighted summation of the first loss information and the third loss information respectively and then summation, or multiplication, etc. This example does not make special limitations on this.

[0169] Embodiment 4: When the whistling annotation information includes whether there is whistling, whistling type information, and whistling level information, a weighted binary cross-entropy loss function for binary classification is used to determine the first loss information corresponding to whether there is whistling, a first loss function corresponding to multi-classification is used to determine the second loss information corresponding to the whistling type, and a second loss function corresponding to multi-classification is used to determine the third loss information corresponding to the whistling level information. Based on the first loss information, the second loss information, and the third loss information, the target loss information is obtained. In this embodiment, the first loss information, the second loss information, and the third loss information can be subjected to relevant operations, and the result of the relevant operations is used as the target loss information. The relevant operations can be direct summation, or weighted summation after weighting the first loss information, the second loss information, and the third loss information respectively, or multiplication, etc. This example does not make special limitations on this.

[0170] For example, in view of the problem of reducing the false detection rate, which is different from a general binary classification problem, the present disclosure can use a weighted binary cross-entropy loss function to guide the trends of the detection rate and the false detection rate. Among them, the weighted binary cross-entropy loss function is:

[0171] L = -α·p·logq - (2 - α)·(1 - p)·log(1 - q);

[0172] In the above formula, L represents the loss function, α is the weight value, and its value ranges between (0, 2]. Generally, the smaller the value of α, the lower the false detection rate; p is the true value, and q is the predicted value output by the model.

[0173] The first loss function corresponding to multi-classification can adopt the mean square error loss function or the cross-entropy loss function. The second loss function corresponding to multi-classification can adopt the mean square error loss function or the cross-entropy loss function. The first loss function and the second loss function can be the same or different. This example does not make special limitations on this.

[0174] Step S950: Adjust the parameters of the first whistling detection model according to the target loss information. In this exemplary embodiment, the random gradient descent method is used for backpropagation to guide the update of the parameters of the first whistling detection model.

[0175] In some embodiments, model compression and model quantization can be performed during model training. Specifically, model compression is to prune the neurons and their connection relationships in the first howling detection model based on a specified metric. Model compression uses methods such as model pruning and model quantization. The purpose of model pruning is to remove unimportant neurons in the model, and pruning methods based on different criteria such as thresholds or energies can be adopted, such as Level Pruner, L1 Pruner, L2 Pruner, etc., which can significantly improve the inference speed of the model. Model quantization is to quantize the parameters of the first howling detection model based on a specified quantization criterion. For example, by quantizing 32-bit floating-point parameters into low-bit parameters, the purpose of reducing the model size and improving the inference speed can be achieved, such as static quantization, dynamic quantization, QAT (Quantization Aware Training), etc. At the same time, in order to ensure the accuracy of the model, parameter fine-tuning operations can be performed while model pruning and model quantization. The present disclosure further reduces the size and overhead of the model through model pruning and model quantization to meet the real-time requirements.

[0176] In some embodiments, considering the real-time requirements, 32 frames of input, that is, 1.28s of data, are used for task processing, instead of the 10s of data commonly used in traditional sound event detection tasks.

[0177] In some embodiments, parameter regularization and / or parameter normalization are performed on the model parameters of the first howling detection model to increase the robustness of the model.

[0178] In some embodiments, instead of directly using the output of the model as the final detection result, post-processing is performed on the model output. The final detection result is determined according to the result of the post-processing.

[0179] Embodiment 1: Compare the number of frames of the audio signal corresponding to the specified howling attribute information in the detection result with a preset threshold; according to the comparison result, determine the howling attribute information of the audio signal to be detected. For example, process and output every 32 frames. If the number of frames is greater than the preset threshold, it is considered howling and 1 is output, otherwise 0 is output.

[0180] Embodiment 2: For a specific frame signal, average the detection results of this frame signal in consecutive specified detections, and compare the processed result with a posterior threshold; according to the comparison result, determine the howling attribute information of the audio signal to be detected. For example, for a specific frame signal, when the buffered signal is shifted frame by frame, in fact, this frame signal will appear in the detection results of 32 consecutive times. After averaging the 32 detection results, then judge whether this frame is howling according to the posterior threshold. If it is greater than the threshold, 1 is output, otherwise 0 is output.

[0181] In some embodiments, after the output of the howling result, subsequent smoothing processing can also be performed on the model output to filter out short-term howling and the like.

[0182] The howling detection result of the present disclosure can be applied to subsequent howling suppression and other processing, such as muting the system microphone and / or speaker; the system reminds the user that howling occurs and performs muting processing; as a label, starts howling suppression processing, etc. The howling suppression processing can be a signal processing means, such as a notch filter, an adaptive filter, etc., or a neural network processing method. For example, a noise reduction network, RNN-noise (artificial neural network noise reduction), DS-Net (dense scale neural network), etc.

[0183] Next, in conjunction with Figure 10 The specific process of a howling detection method according to an embodiment of the present application will be introduced.

[0184] Refer to Figure 10 , the specific process of the howling detection method includes the following steps:

[0185] Step S1001, obtain the audio signal to be detected; in this example, the audio signal to be detected can be obtained through the microphone of the communication device.

[0186] Step S1002, resample the audio signal to be detected to normalize the audio signal to be detected to a specified sampling rate; in this example, downsampling can be used to normalize the audio signal to be detected to 16 kHz.

[0187] Step S1003, perform frame division processing on the normalized audio signal to be detected; in this example, the frame length can be set, and the audio signal to be detected is intercepted in sequence according to the frame length to obtain the audio signal to be detected frame by frame.

[0188] Step S1004, perform feature extraction on a frame of the audio signal to be detected after frame division to obtain the audio feature data to be detected. The feature extraction process in this example can be: perform windowing processing on a frame of the audio signal to be detected after frame division; then perform a fast Fourier transform on the windowed audio signal to be detected to obtain the corresponding frequency domain signal to be detected; filter the frequency domain signal to be detected through a corresponding frequency domain filter; convert the filtered frequency domain signal to the logarithmic domain to obtain the audio feature data to be detected.

[0189] For example, the audio signal data stream to be detected is downsampled to 16 kHz, and frames of 40 ms are used for log-domain Mel spectrum feature extraction, with a dimension of 60, and are cached. The data feature cache has a total of 32 frames, constituting the audio feature data to be detected of (32×60).

[0190] Step S1005, sample audio signals are collected under different communication scenarios, and the sample audio is labeled with howling attributes; in this exemplary embodiment, different communication scenarios may include quiet environment, noisy environment, howling scenario, non-howling scenario, single-device communication scenario, multi-device communication scenario, and so on. Howling attributes may include whether there is howling (such as with howling or without howling), howling type (such as continuous howling, intermittent howling, single-frequency howling, multi-frequency howling, and diffusive howling), and howling level (such as strong, medium, weak), etc.

[0191] Step S1006, the first howling detection model is trained using the sample audio signals and the corresponding howling annotation information. Specifically, during the training process, parameter regularization and parameter normalization are performed, and at the same time, a dropout module (dropout processing) is added within each convolutional layer. Model compression and model quantization processing can also be performed during or after the training process to reduce the model size and computational complexity.

[0192] Step S1007, a target howling detection model is obtained. In this exemplary embodiment, the target howling detection model can be obtained offline or through an online preloading process.

[0193] Step S1008, the audio feature data to be detected is input into the target howling detection model, and the howling attribute probability of each frame is output; step S1009 or step S1010 is executed. In this exemplary embodiment, the howling detection process of the audio feature data to be detected is an online detection process. To ensure real-time performance, 32 frames of input, that is, 1.28s of data features, are used for detection in this example, and the output is a vector composed of the howling attribute probabilities of each frame of audio feature data.

[0194] Step S1009, the number of frames of the audio signal corresponding to the specified howling attribute information in the detection results of each input multi-frame data is compared with a preset threshold; according to the comparison result, the howling attribute information of the audio signal to be detected is determined. In this exemplary embodiment, the input multi-frame feature data (such as 32 frames) can be detected and output once, and the threshold decision is made based on the number of frames detected in this segment. If it is greater than the preset threshold, it is considered howling and 1 is output; otherwise, 0 is output.

[0195] Step S1010: For a specific frame signal, average the detection results of the frame signal in consecutive specified detections, and compare the processed result with a posterior threshold; based on the comparison result, determine the howling attribute information of the audio signal to be detected. In this exemplary embodiment, frame-by-frame processing output can be performed. For a specific frame signal, when the buffered signal is shifted frame by frame, actually this frame signal will appear in the detection results of consecutive multiple times (such as 32 times). After averaging the 32 detection results, then compare the average value with the posterior threshold to determine whether this frame is howling. If the average value is greater than the posterior threshold, output 1; otherwise, output 0.

[0196] Step S1011: Determine the final detection result according to Step S1009 or Step S1010.

[0197] According to the howling detection method of this embodiment, by collecting sample data in different communication scenarios and adding howling type and howling level information to the annotation information at the same time, the obtained target howling detection model after training has a good detection rate for non-linear audio features in complex scenarios, and can perform more detailed howling attribute detection for the attributes of the howling itself, reducing the false detection rate and improving the call quality and user experience. In addition, by designing the model training process and adding parameter dropout processing in each convolutional layer, the model has good detection performance for different howling types. Further, by model compression and model quantization to reduce the model size and overhead, combined with the selection of the input data size, the model detection process can meet the real-time requirements of communication. And the post-processing process of the model detection results ensures the detection accuracy.

[0198] Exemplary Apparatus

[0199] It should be noted that for the howling detection method provided in the embodiments of the present disclosure, the execution subject can be a howling detection device, or a control module in the howling detection device for executing the loaded howling detection method. In the embodiments of the present disclosure, the howling detection device is taken as an example for executing the loaded howling detection method to illustrate the howling detection method provided in the embodiments of the present disclosure. Next, refer to Figure 11 Describe the howling detection device of the exemplary embodiment of the present disclosure.

[0200] Figure 11 The block diagram of a howling detection device according to an embodiment of the present invention is schematically shown.

[0201] Refer to Figure 11 As shown, a howling detection device 1100 according to an embodiment of the present invention includes:

[0202] A preprocessing module 1110, configured to obtain an audio signal to be detected and preprocess the audio signal to be detected to obtain audio feature data to be detected;

[0203] A detection module 1120, configured to obtain a detection result based on the audio feature data to be detected through a target whistling detection model; the detection result at least includes whistling attribute information;

[0204] Wherein, the target whistling detection model is obtained by training a first whistling detection model based on sample audio signals collected in different communication scenarios and corresponding whistling annotation information; the first whistling detection model includes a cascaded convolutional neural network and a recurrent neural network in sequence.

[0205] In some embodiments of the present disclosure, based on the foregoing solution, the preprocessing module 1110 includes:

[0206] A resampling module, configured to resample the audio signal to be detected so that the audio signal to be detected is normalized to a specified sampling rate;

[0207] A framing module, configured to perform framing processing on the normalized audio signal to be detected;

[0208] A feature extraction module, configured to extract features from a frame of the framed audio signal to be detected to obtain audio feature data to be detected.

[0209] In some embodiments of the present disclosure, based on the foregoing solution, the apparatus 1100 further includes a training module, and the training module is configured to:

[0210] Based on the convolutional neural network, perform convolutional processing on the feature data of the sample audio signal, and output a first feature vector, where the first feature vector includes temporal information;

[0211] Based on the recurrent neural network, perform temporal feature learning on the first feature vector, and output a second feature vector;

[0212] Perform focusing processing on the second feature vector, and output a whistling attribute probability distribution vector, where the whistling attribute probability distribution vector represents the probabilities corresponding to each whistling attribute;

[0213] Based on the probability of the whistling attribute corresponding to the feature data of the sample audio signal and the whistling annotation information, use a loss function to determine target loss information;

[0214] According to the target loss information, adjust the parameters of the first whistling detection model.

[0215] In some embodiments of the present disclosure, based on the foregoing solution, the training module is further configured to:

[0216] Pruning neurons and their connection relationships in the first howling detection model based on a specified measurement criterion; and / or

[0217] Quantizing the parameters of the first howling detection model based on a specified quantization criterion.

[0218] In some embodiments of the present disclosure, based on the foregoing solution, the howling annotation information includes one or more of whether there is howling, howling type information, and howling level information.

[0219] In some embodiments of the present disclosure, based on the foregoing solution, the training module is further configured to:

[0220] When the howling annotation information includes whether there is howling, a weighted binary cross-entropy loss function is used to determine the target loss information;

[0221] When the howling annotation information includes whether there is howling and howling type information, a binary weighted binary cross-entropy loss function is used to determine the first loss information corresponding to whether there is howling, and a first loss function corresponding to multi-classification is used to determine the second loss information corresponding to the howling type information, and the target loss information is obtained based on the first loss information and the second loss information;

[0222] When the howling annotation information includes whether there is howling and howling level information, a binary weighted binary cross-entropy loss function is used to determine the first loss information corresponding to whether there is howling, and a second loss function corresponding to multi-classification is used to determine the third loss information corresponding to the howling level information, and the target loss information is obtained based on the first loss information and the third loss information;

[0223] When the howling annotation information includes whether there is howling, howling type information, and howling level information, a binary weighted binary cross-entropy loss function is used to determine the first loss information corresponding to whether there is howling, a first loss function corresponding to multi-classification is used to determine the second loss information corresponding to the howling type, and a second loss function corresponding to multi-classification is used to determine the third loss information corresponding to the howling level information, and the target loss information is obtained based on the first loss information, the second loss information, and the third loss information.

[0224] In some embodiments of the present disclosure, based on the foregoing solution, the training module is further configured to:

[0225] For different howling types, perform dropout processing on the convolutional layer parameters of the convolutional neural network so that the target howling detection model is applicable to the detection of different howling types.

[0226] In some embodiments of the present disclosure, based on the foregoing solution, the training module is configured to:

[0227] Perform parameter regularization and / or parameter normalization on the model parameters of the first whistling detection model.

[0228] In some embodiments of the present disclosure, based on the foregoing solution, the feature extraction module includes:

[0229] A windowing sub-module for windowing a frame of audio signal to be detected after frame division;

[0230] A transformation sub-module for performing a fast Fourier transform on the windowed audio signal to be detected to obtain the corresponding frequency domain signal to be detected;

[0231] A filtering sub-module for filtering the frequency domain signal to be detected through a corresponding frequency domain filter;

[0232] A conversion sub-module for converting the filtered frequency domain signal to the logarithmic domain to obtain the audio feature data to be detected.

[0233] In some embodiments of the present disclosure, based on the foregoing solution, the device 1100 further includes:

[0234] A sample acquisition module for acquiring the sample audio signal by collecting the audio signals played by a microphone and / or communication devices with different performances.

[0235] In some embodiments of the present disclosure, based on the foregoing solution, the device 1100 further includes a first post-processing module, and the first post-processing module is used for:

[0236] Compare the number of frames of the audio signal corresponding to the specified whistling attribute information in the detection result with a preset threshold;

[0237] Determine the whistling attribute information of the audio signal to be detected according to the comparison result.

[0238] In some embodiments of the present disclosure, based on the foregoing solution, the device 1100 further includes a second post-processing module, and the second post-processing module is used for:

[0239] For a specific frame signal, average the detection results of the frame signal in consecutive specified detections, and compare the processing result with a posterior threshold;

[0240] Determine the whistling attribute information of the audio signal to be detected according to the comparison result.

[0241] Exemplary Medium

[0242] After introducing the method of the exemplary embodiment of the present invention, next, the medium of the exemplary embodiment of the present invention will be described.

[0243] In some possible embodiments, with reference to Figure 12 , various aspects of the present invention can also be implemented as a medium 1200 storing program code, which, when executed by a processor of a device, is used to implement the steps in the howling detection method according to various exemplary embodiments of the present invention described in the "Exemplary Method" section above of this specification.

[0244] Specifically, when the processor of the device executes the program code, it is used to implement the following steps:

[0245] Obtain an audio signal to be detected, and preprocess the audio signal to be detected to obtain audio feature data to be detected;

[0246] Obtain a detection result through a target howling detection model based on the audio feature data to be detected; the detection result at least includes howling attribute information;

[0247] Wherein, the target howling detection model is obtained by training a first howling detection model based on sample audio signals collected in different communication scenarios and corresponding howling annotation information; the first howling detection model includes a convolutional neural network and a recurrent neural network cascaded in sequence.

[0248] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above howling detection method belong to the same concept. For the details not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above howling detection method.

[0249] It should be noted that: the above medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0250] A readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which readable program code is carried. Such a propagated data signal can take various forms, including but not limited to: electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0251] The program code contained on the readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.

[0252] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0253] Exemplary Computing Device

[0254] After introducing the methods, media, and devices of the exemplary embodiments of the present disclosure, next, a computing device according to another exemplary embodiment of the present disclosure will be introduced.

[0255] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, method, or program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to herein as "circuits", "modules", or "systems".

[0256] The following refers to Figure 13 to describe the computing device 1300 according to such an embodiment of the present invention. Figure 13 The shown computing device 1300 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0257] As Figure 13As shown, the computing device 1300 is presented in the form of a general-purpose computing device. The components of the computing device 1300 may include, but are not limited to: at least one of the above-mentioned processing units 1310, at least one of the above-mentioned storage units 1320, and a bus 1330 that connects different system components (including the storage unit 1320 and the processing unit 1310).

[0258] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1310, so that the processing unit 1310 executes the steps according to various exemplary embodiments of the present invention described in the "Exemplary Method" section of the present specification above.

[0259] The storage unit 1320 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 13201 and / or a cache storage unit 13202, and may further include a read-only storage unit (ROM) 13203.

[0260] The storage unit 1320 may also include a program / utility 13204 having a set (at least one) of program modules 13205. Such program modules 13205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0261] The bus 1330 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.

[0262] The computing device 1300 can also communicate with one or more external devices (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the computing device 1300, and / or communicate with any device that enables the computing device 1300 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the display unit 1340 and the input / output (I / O) interface 1350 connected to the display unit 1340. Moreover, the computing device 1300 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1360. As shown in the figure, the network adapter 1360 communicates with other modules of the computing device 1300 through the bus 1330. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the computing device 1300, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0263] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0264] The above is a schematic solution of a computing device 1300 in this embodiment. It should be noted that the technical solution of the computing device 1300 and the technical solution of the above-mentioned howling detection method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above-mentioned howling detection method.

[0265] It should be noted that although several modules or sub-modules of the howling detection device are mentioned in the above detailed description, this division is only exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0266] Moreover, although the operations of the method of the present invention are described in a specific order in the drawings, this is not a requirement or implication that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0267] Although the spirit and principles of the present invention have been described with reference to several specific embodiments, it should be understood that the present invention is not limited to the specific embodiments invented, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. This division is only for convenience of expression. The present invention is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A howling detection method, characterized in that, it includes: Obtain the audio signal to be detected, and preprocess the audio signal to be detected to obtain the audio feature data to be detected; Obtain the detection result through the target howling detection model based on the audio feature data to be detected; the detection result at least includes howling attribute information; Among them, the target howling detection model is obtained by training the first howling detection model based on the sample audio signals collected in different communication scenarios and the corresponding howling annotation information; the first howling detection model includes a convolutional neural network and a recurrent neural network cascaded in sequence; The training of the first howling detection model includes: Based on the convolutional neural network, perform convolutional processing on the feature data of the sample audio signal, and output the first feature vector, and the first feature vector contains timing information; Based on the recurrent neural network, perform timing feature learning on the first feature vector, and output the second feature vector; Perform focusing processing on the second feature vector, and output the howling attribute probability distribution vector, and the howling attribute probability distribution vector represents the probabilities corresponding to each howling attribute; Based on the probability of the howling attribute corresponding to the feature data of the sample audio signal and the howling annotation information, use a loss function to determine the target loss information; According to the target loss information, adjust the parameters of the first howling detection model.

2. The howling detection method according to claim 1, characterized in that, The preprocessing of the audio signal to be detected to obtain the audio feature data to be detected includes: Resample the audio signal to be detected to normalize the audio signal to be detected to a specified sampling rate; Perform frame segmentation on the normalized audio signal to be detected; Extract features from one frame of the segmented audio signal to be detected to obtain the audio feature data to be detected.

3. The howling detection method according to claim 1, characterized in that, The training of the first howling detection model further includes: Based on a specified measurement criterion, prune the neurons and their connection relationships in the first howling detection model; and / or Quantize the parameters of the first howling detection model based on a specified quantization criterion.

4. The howling detection method according to claim 1, characterized in that, The howling annotation information includes one or more of whether there is howling, howling type information, and howling level information.

5. The howling detection method according to claim 1, characterized in that, The step of determining the target loss information by using a loss function based on the probability of the howling attribute corresponding to the feature data of the sample audio signal and the howling annotation information includes at least one of the following: When the howling annotation information includes whether there is howling, use a weighted binary cross-entropy loss function to determine the target loss information; When the howling annotation information includes whether there is howling and howling type information, a weighted binary cross-entropy loss function for binary classification is used to determine the first loss information corresponding to whether there is howling, and a first loss function corresponding to multi-classification is used to determine the second loss information corresponding to the howling type information, and the target loss information is obtained based on the first loss information and the second loss information; When the howling annotation information includes whether there is howling and howling level information, a weighted binary cross-entropy loss function for binary classification is used to determine the first loss information corresponding to whether there is howling, and a second loss function corresponding to multi-classification is used to determine the third loss information corresponding to the howling level information, and the target loss information is obtained based on the first loss information and the third loss information; When the howling annotation information includes whether there is howling, howling type information and howling level information, a weighted binary cross-entropy loss function for binary classification is used to determine the first loss information corresponding to whether there is howling, a first loss function corresponding to multi-classification is used to determine the second loss information corresponding to the howling type, and a second loss function corresponding to multi-classification is used to determine the third loss information corresponding to the howling level information, and the target loss information is obtained based on the first loss information, the second loss information and the third loss information.

6. The howling detection method according to claim 1, wherein, The convolutional processing of the feature data of the sample audio signal based on the convolutional neural network to output a first feature vector includes: For different howling types, the convolutional layer parameters of the convolutional neural network are discarded to make the target howling detection model applicable to the detection of different howling types.

7. The howling detection method according to claim 1, wherein, The training of the first howling detection model includes: Performing parameter regularization and / or parameter normalization on the model parameters of the first howling detection model.

8. The howling detection method according to claim 2, wherein, The obtaining of the audio feature data to be detected by performing feature extraction on a frame of the audio signal to be detected after frame division includes: Performing windowing processing on a frame of the audio signal to be detected after frame division; Performing fast Fourier transform on the windowed audio signal to be detected to obtain the corresponding frequency domain signal to be detected; Performing filtering processing on the frequency domain signal to be detected through a corresponding frequency domain filter; Converting the filtered frequency domain signal to the logarithmic domain to obtain the audio feature data to be detected.

9. The howling detection method according to claim 1, wherein, The method further includes: obtaining the sample audio signal by collecting the audio signals played by a microphone and / or different performance communication devices.

10. The howling detection method according to claim 1, wherein, After obtaining the detection result based on the audio feature data to be detected by the target howling detection model, the method further includes: Comparing the number of frames of the audio signal corresponding to the specified howling attribute information in the detection result with a preset threshold; Determining the howling attribute information of the audio signal to be detected according to the comparison result.

11. The howling detection method according to claim 1, It is characterized in that after the target whistling detection model obtains a detection result based on the audio feature data to be detected, the method further includes: For a specific frame signal, average the detection results of the frame signal in consecutive specified detections, and compare the processed result with a posterior threshold; According to the comparison result, determine the whistling attribute information of the audio signal to be detected.

12. A whistling detection device It is characterized in that including: A preprocessing module for obtaining an audio signal to be detected and preprocessing the audio signal to be detected to obtain audio feature data to be detected; A detection module for obtaining a detection result based on the audio feature data to be detected through a target whistling detection model; the detection result at least includes whistling attribute information; Among them, the target whistling detection model is obtained by training a first whistling detection model based on sample audio signals collected in different communication scenarios and corresponding whistling annotation information; the first whistling detection model includes a cascaded convolutional neural network and a recurrent neural network in sequence; A training module for performing convolutional processing on the feature data of the sample audio signal based on the convolutional neural network, outputting a first feature vector, where the first feature vector contains timing information; based on the recurrent neural network, performing timing feature learning on the first feature vector, outputting a second feature vector; performing focusing processing on the second feature vector, outputting a whistling attribute probability distribution vector, where the whistling attribute probability distribution vector represents the probabilities corresponding to each whistling attribute; based on the probability of the whistling attribute corresponding to the feature data of the sample audio signal and the whistling annotation information, determining target loss information using a loss function; adjusting the parameters of the first whistling detection model according to the target loss information.

13. The whistling detection device according to claim 12 It is characterized in that The preprocessing module includes: A resampling module for resampling the audio signal to be detected to normalize the audio signal to be detected to a specified sampling rate; A framing module for framing the normalized audio signal to be detected; A feature extraction module for extracting features from a framed audio signal to be detected to obtain audio feature data to be detected.

14. The whistling detection device according to claim 12 It is characterized in that The training module is further configured to: Prune the neurons and their connection relationships in the first whistling detection model based on a specified measurement criterion; and / or Quantize the parameters of the first whistling detection model based on a specified quantization criterion.

15. The whistling detection device according to claim 12 It is characterized in that The whistling annotation information includes one or more of whether it contains whistling, whistling type information, and whistling level information.

16. The whistling detection device according to claim 12 It is characterized in that The training module is further configured to: When the whistling annotation information includes whether it contains whistling, use a weighted binary cross-entropy loss function to determine the target loss information; When the howling annotation information includes whether there is howling and howling type information, a weighted binary cross-entropy loss function for binary classification is used to determine the first loss information corresponding to whether there is howling, and a first loss function corresponding to multi-classification is used to determine the second loss information corresponding to the howling type information, and the target loss information is obtained based on the first loss information and the second loss information; When the howling annotation information includes whether there is howling and howling level information, a weighted binary cross-entropy loss function for binary classification is used to determine the first loss information corresponding to whether there is howling, and a second loss function corresponding to multi-classification is used to determine the third loss information corresponding to the howling level information, and the target loss information is obtained based on the first loss information and the third loss information; When the howling annotation information includes whether there is howling, howling type information and howling level information, a weighted binary cross-entropy loss function for binary classification is used to determine the first loss information corresponding to whether there is howling, a first loss function corresponding to multi-classification is used to determine the second loss information corresponding to the howling type, and a second loss function corresponding to multi-classification is used to determine the third loss information corresponding to the howling level information, and the target loss information is obtained based on the first loss information, the second loss information and the third loss information.

17. The howling detection device according to claim 12, wherein, the training module is further configured to: For different howling types, perform dropout processing on the convolutional layer parameters of the convolutional neural network, so that the target howling detection model is applicable to the detection of different howling types.

18. The howling detection device according to claim 12, wherein, the training module is configured to: Perform parameter regularization and / or parameter normalization on the model parameters of the first howling detection model.

19. The howling detection device according to claim 13, wherein, the feature extraction module includes: a windowing sub-module, configured to perform windowing processing on a frame of audio signal to be detected after frame division; a transformation sub-module, configured to perform fast Fourier transform on the audio signal to be detected after windowing processing to obtain a corresponding frequency-domain signal to be detected; a filtering sub-module, configured to perform filtering processing on the frequency-domain signal to be detected through a corresponding frequency-domain filter; a conversion sub-module, configured to convert the frequency-domain signal after filtering processing to the logarithmic domain to obtain audio feature data to be detected.

20. The howling detection device according to claim 12, wherein, the device further includes: a sample acquisition module, configured to acquire the sample audio signal by collecting audio signals played by a microphone and / or different performance communication devices.

21. The howling detection device according to claim 12, wherein, the device further includes a first post-processing module, and the first post-processing module is configured to: Compare the number of frames of the audio signal corresponding to the specified howling attribute information in the detection result with a preset threshold; Determine the howling attribute information of the audio signal to be detected according to the comparison result.

22. The howling detection device according to claim 12, wherein, the device further includes a second post-processing module, and the second post-processing module is configured to: For a specific frame signal, the detection results of the frame signal in consecutive specified detections are averaged, and the processed result is compared with a posterior threshold; According to the comparison result, the howling attribute information of the audio signal to be detected is determined.

23. A medium, characterized in that, a program is stored thereon, and when the program is executed by a processor, the method described in any one of claims 1 to 11 is implemented.

24. A computing device, characterized in that, comprising: a processor and a memory, the memory stores executable instructions, and the processor is configured to call the executable instructions stored in the memory to execute the method described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Acoustic amplification system howling point detection method based on neural network

    CN111526469A