Method and device for detecting howling voice signal, electronic equipment and storage medium

By performing frequency filtering and mapping on the voice signal collected by the microphone to determine the frequency, and combining the amplitude change to detect howling, the problem of low howling detection efficiency in the existing technology is solved, and fast and accurate howling detection is achieved.

CN116543793BActive Publication Date: 2025-12-30XIAOMI TECH (WUHAN) CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310524220.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-05
Publication Date
2025-12-30
Estimated Expiration
2043-05-05

AI Technical Summary

Technical Problem

The existing technology has low efficiency in howling detection because frequency domain analysis requires a long detection time, which cannot meet real-time requirements.

Method used

By acquiring a speech time-domain signal of a set duration collected by a microphone, performing frequency filtering processing, a single-frequency speech time-domain signal is obtained, and the frequency is determined according to a set mapping relationship. In response to the frequency matching with the target howling frequency point, howling speech signal detection is performed using amplitude changes.

Benefits of technology

By tracking a single frequency signal in the time domain and combining amplitude changes to identify howling, detection efficiency is improved, the inefficiency of frequency domain frame processing is avoided, and fast and accurate howling detection is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543793B_ABST
    Figure CN116543793B_ABST
Patent Text Reader

Abstract

The application provides a howling voice signal detection method and device, electronic equipment and storage medium. The method comprises the following steps: performing frequency filtering processing on a plurality of sampling time first voice time domain signals collected by a microphone to obtain a single frequency second voice time domain signal at each sampling time, and determining the frequency of the second voice time domain signal at each sampling time. The single frequency signal at each sampling time is tracked in the time domain, so that the second voice time domain signal tracked has both time domain information and frequency domain information. In response to the fact that the frequency of the second voice time domain signal at each sampling time and a target howling frequency point are matched, the second voice time domain signal at a plurality of sampling times is detected according to the change of the amplitude of the second voice time domain signal at a plurality of sampling times. The same effect is achieved in the time domain and the frequency domain without processing in the frequency domain, and the efficiency of howling voice signal detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a method, apparatus, electronic device and storage medium for detecting howling speech signals. Background Technology

[0002] Feedback noise can occur in any scenario where an audio loop exists, and this feedback severely affects the normal use of voice services, causing great discomfort to customers.

[0003] In related technologies, howling detection is generally performed in the time domain and frequency domain separately. Frequency domain analysis requires Fourier transform (FFT). In order to improve the accuracy of detection, the number of FFT points needs to be relatively large, which makes the length of each frame longer. Therefore, it requires a long detection time and has low detection efficiency. Summary of the Invention

[0004] This application aims to at least partially address one of the technical problems in the related art.

[0005] Therefore, this application proposes a method, apparatus, electronic device, and storage medium for detecting howling voice signals, in order to reduce detection time and improve detection efficiency.

[0006] One embodiment of this application proposes a method for detecting howling voice signals, including:

[0007] Acquire a first speech time-domain signal of a set duration collected by a microphone; the set duration includes multiple sampling times;

[0008] Frequency filtering is performed on the first speech time-domain signal at each of the sampling times to obtain a second speech time-domain signal with a single frequency at each of the sampling times;

[0009] Based on the established mapping relationship, the frequency of the second speech time-domain signal at each of the sampling times is determined;

[0010] In response to the fact that the frequency of the second speech time-domain signal at each of the sampling times matches the target howling frequency, howling speech signal detection is performed on the second speech time-domain signal at the multiple sampling times based on the amplitude of the second speech time-domain signal at the multiple sampling times.

[0011] Another embodiment of this application proposes a device for detecting howling voice signals, comprising:

[0012] The acquisition module is used to acquire a first speech time-domain signal of a set duration collected by the microphone; the set duration includes multiple sampling times.

[0013] The processing module is used to perform frequency filtering processing on the first speech time-domain signal at each of the sampling times to obtain a second speech time-domain signal with a single frequency at each of the sampling times.

[0014] The determining module is used to determine the frequency of the second speech time-domain signal at each of the sampling times according to the set mapping relationship;

[0015] The detection module is used to detect howling voice signals in response to the fact that the frequency of the second speech time-domain signal at each of the sampling times matches the target howling frequency point, based on the amplitude of the second speech time-domain signal at the multiple sampling times.

[0016] Another embodiment of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the foregoing aspect.

[0017] Another embodiment of this application proposes a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the foregoing aspect.

[0018] Another embodiment of this application proposes a computer program product having a computer program stored thereon, which, when executed by a processor, implements the method described in the foregoing aspect.

[0019] The method, apparatus, electronic device, and storage medium for detecting howling voice signals proposed in this application acquire a first speech time-domain signal of a set duration collected by a microphone. Frequency filtering is performed on the first speech time-domain signal at each sampling moment of the set duration to obtain a second speech time-domain signal with a single frequency at each sampling moment. Based on a set mapping relationship, the frequency of the second speech time-domain signal at each sampling moment is determined. Since the frequency of the second speech time-domain signal at each sampling moment matches the target howling frequency, a preliminary determination of the howling voice signal is achieved. Furthermore, howling voice signal detection is performed on the second speech time-domain signal at multiple sampling moments based on the amplitude changes of the second speech time-domain signal at multiple sampling moments. This achieves the tracking of the single frequency signal at each sampling moment in the time domain, ensuring that the tracked second speech time-domain signal has both time-domain and frequency-domain information. Furthermore, by tracking the amplitude of the second speech time-domain signal at each moment, the same effect of tracking in both the time and frequency domains is achieved, while eliminating the need to process each frame signal in the frequency domain, thus improving the efficiency of howling voice signal detection.

[0020] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0021] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0022] Figure 1 A flowchart illustrating a method for detecting howling voice signals provided in an embodiment of this application;

[0023] Figure 2 A flowchart illustrating another method for detecting howling voice signals provided in an embodiment of this application;

[0024] Figure 3 This is a schematic diagram of an adaptive filter performing signal processing according to an embodiment of this application;

[0025] Figure 4 A schematic diagram of the structure of a device for detecting howling voice signals provided in an embodiment of this application;

[0026] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0027] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0028] The following description, with reference to the accompanying drawings, describes a method, apparatus, electronic device, and storage medium for detecting howling voice signals according to embodiments of this application.

[0029] In related technologies, howling detection algorithms typically perform trend tracking in both the time and frequency domains. Since howling sounds experience a rapid increase in volume and their energy is concentrated at a specific frequency, after processing the signal in frames, the energy of several consecutive frames is first statistically analyzed in the time domain to check for a significant increasing trend. If such a trend is found, the frequency domain is then analyzed. If the amplitude of a particular frequency point suddenly increases and is significantly greater than that of other frequency points, howling is identified. Frequency domain analysis requires Fourier Transform (FFT). Since speech signals contain multiple single-frequency points, detection requires high frequency resolution, resulting in a large number of FFT points and a long frame length. For example, if the signal sampling frequency is 48kHz, each frame would be at least 2048 bytes long. Detecting five frames to determine the presence of howling would require over 200 milliseconds of detection time, which is time-consuming and inefficient.

[0030] To address this, this application proposes a method for detecting howling speech signals. The method involves acquiring a first speech time-domain signal collected by a microphone, filtering the first speech time-domain signal to obtain a second speech time-domain signal with the same frequency, matching the frequency of the second speech time-domain signal with a target frequency, and detecting howling speech signals based on the amplitude of the second speech time-domain signal. In this application, speech signal detection is performed in the time domain. When the frequency of the filtered second speech time-domain signal containing the same frequency matches the howling frequency, the method determines whether the second speech time-domain signal is howling speech based on the change in its amplitude. Since the speech signal is not analyzed in the frequency domain, the method avoids the low processing efficiency caused by the long frame length in single-frequency signal analysis, thus improving processing efficiency.

[0031] Figure 1 This is a flowchart illustrating a method for detecting howling voice signals provided in an embodiment of this application.

[0032] The execution subject of the feedback voice signal detection method in this application embodiment is a feedback voice signal detection device. The device can be installed in an electronic device, which can be a terminal device or a server. The terminal device can be a smartphone, smartwatch, tablet, and smart wearable device, etc. The server can be a cloud server or a local server cluster. This embodiment does not limit the specific implementation.

[0033] like Figure 1 As shown, the method may include the following steps:

[0034] Step 101: Acquire the first speech time-domain signal of a set duration collected by the microphone.

[0035] The set duration includes multiple sampling times.

[0036] In this embodiment, the microphone collects speech signals using a set sampling frequency to obtain continuous first speech time-domain signals of a set duration. For example, the sampling frequency is 48kHz. Based on the sampling frequency, the sampling time can be determined, i.e., a first speech time-domain signal is collected every 0.02 milliseconds. The collected speech signal is a time-domain signal, and the collection duration is a set duration, which can be set according to requirements, for example, 10 milliseconds.

[0037] Step 102: Perform frequency filtering on the first speech time-domain signal at each sampling time to obtain the second speech time-domain signal at a single frequency at each sampling time.

[0038] In this embodiment, the first speech time-domain signal collected by the microphone contains signals of multiple frequencies, including possible howling speech signals. By filtering the first speech time-domain signal at each sampling moment, i.e. using a frequency discrimination algorithm to filter out speech time-domain signals other than the set frequency and retaining the speech time-domain signal of the set frequency, the resulting filtered second speech time-domain signal at that sampling moment is of a single frequency, wherein the set frequency can be the maximum frequency.

[0039] Step 103: Determine the frequency of the second speech time-domain signal at each sampling time according to the set mapping relationship.

[0040] In this embodiment of the application, when the second speech time-domain signal of a single frequency is obtained by filtering at each sampling time, it is necessary to determine which specific frequency the single frequency is. Therefore, for the second speech time-domain signal of a single frequency at each sampling time, a set mapping relationship is used to map and obtain the frequency of the second speech time-domain signal at each sampling time.

[0041] As an example, the mapping relationship is set as an inverse cosine operation, the second speech time-domain signal at time n is x(n), and the frequency of the second speech time-domain signal at time n is f. n , where f n The following relationship must be satisfied:

[0042] f n = acos(x(n)).

[0043] Step 104: In response to the fact that the frequency of the second speech time-domain signal at each sampling time matches the target howling frequency, howling speech signal detection is performed on the second speech time-domain signal at multiple sampling times based on the amplitude of the second speech time-domain signal at multiple sampling times.

[0044] Among them, the target howling frequency point is a frequency point of the howling voice signal. For example, the frequency point of the howling voice signal corresponds to a frequency of 500Hz, and the frequency point of 1000Hz is also the frequency of the howling frequency point. In other words, the frequency point that has a multiple relationship with the frequency point of the howling voice signal is also the howling frequency point.

[0045] In this embodiment, for each sampling time, the frequency of the second speech time-domain signal at that sampling time is matched with the target howling frequency. One implementation method is to subtract the frequency of the target howling frequency from the frequency of the second speech time-domain signal at that sampling time. If the difference is greater than zero and less than or equal to a set threshold, then the frequency of the second speech time-domain signal at that sampling time is considered to match the target howling frequency. Similarly, it can be determined whether the frequency of the second speech time-domain signal at other sampling times matches the target howling frequency. If the frequency of the second speech time-domain signal at each sampling time matches the target howling frequency, then the second speech time-domain signal of a set duration is likely a howling speech signal.

[0046] Furthermore, the amplitude changes of the second speech time-domain signal at multiple sampling times are tracked. As one implementation method, if the amplitude of the second speech time-domain signal sampled at multiple sampling times continues to increase, and if the amplitude is greater than the set amplitude threshold, then the second speech time-domain signal of the set duration is determined to be a howling speech signal; otherwise, the second speech time-domain signal of the set duration is not a howling speech signal.

[0047] In the method for detecting howling voice signals in this application embodiment, a first voice time-domain signal of a set duration is acquired by a microphone. Frequency filtering processing is performed on the first voice time-domain signal at each sampling moment of the set duration to obtain a second voice time-domain signal with a single frequency at each sampling moment. According to a set mapping relationship, the frequency of the second voice time-domain signal at each sampling moment is determined. Since the frequency of the second voice time-domain signal at each sampling moment matches the target howling frequency, the initial determination of the howling voice signal is achieved. Furthermore, howling voice signal detection is performed on the second voice time-domain signal at multiple sampling moments based on the changes in the amplitude of the second voice time-domain signal at multiple sampling moments. This achieves the tracking of the single frequency signal at each sampling moment in the time domain, so that the tracked second voice time-domain signal has both time-domain and frequency-domain information. Furthermore, the amplitude of the second voice time-domain signal at each moment is tracked, achieving the same effect of tracking in both the time and frequency domains. At the same time, it is not necessary to process each frame signal in the frequency domain, thus improving the efficiency of howling voice signal detection.

[0048] Figure 2 A flowchart illustrating another method for detecting howling voice signals provided in this application embodiment is shown below. Figure 2 As shown, the method includes the following steps:

[0049] Step 201: Acquire the first speech time-domain signal of a set duration collected by the microphone.

[0050] The set duration includes the cycles of multiple howling voice signals.

[0051] The period of the howling voice signal is determined based on the target howling frequency. For example, if the target howling frequency is 500Hz, then the period of one howling voice signal is 1000 / 500ms = 2ms (milliseconds). The period of multiple howling voice signals is, for example, 5 periods, which means the set duration is 10ms.

[0052] Step 202: For each cycle, the first speech time-domain signal at each sampling moment within the cycle is input into the adaptive filter corresponding to each sampling moment for frequency filtering processing to obtain the second speech time-domain signal at a single frequency at each sampling moment. Each cycle contains multiple sampling moments. Taking a microphone sampling frequency of 48kHz as an example, with a sampling period of 0.02 milliseconds, a 2-millisecond cycle includes 100 sampling moments.

[0053] In this embodiment, for any sampling time in each cycle, the adaptive filter corresponding to that sampling time is obtained by updating its parameters based on the residual signal of the adaptive filter corresponding to the previous sampling time. By continuously updating the adaptive filter, the accuracy of the adaptive filter is improved.

[0054] If the sampling time is the first sampling time, the residual signal of the previous sampling time is the set residual signal, which is determined based on the speech signal and prior experience. If the sampling time is not the first sampling time, the residual signal of the previous sampling time is determined based on the difference between the first and second speech time-domain signals of the previous sampling time, and the parameters of the adaptive filter of the previous sampling time are adjusted according to the residual signal, where the parameters of the adaptive filter are the coefficients of the adaptive filter.

[0055] As an example, such as Figure 3 As shown, for example, the input adaptive filter H k (z) represents the original speech signal captured by the microphone at time n, i.e., the first speech time-domain signal is x(n). The adaptive filter will use a frequency discrimination algorithm to perform frequency filtering processing, and output the second speech time-domain signal with a single frequency at time n as x(n). k (n), x k If (n) is an estimated value of the carrier frequency containing a sinusoidal signal, then the residual signal at time n satisfies the following relationship:

[0056] ε k (n)=x(n)-x k (n);

[0057] Here, 'k' indicates the identifier of the howling frequency when there are multiple howling frequencies.

[0058] It should be understood that when there are multiple howling frequencies, the howling frequency to be detected is the target howling frequency. The methods for identifying the target howling frequency from a speech time-domain signal are all the same, and will not be described in detail in this embodiment.

[0059] Step 203: Determine the frequency of the second speech time-domain signal at each sampling time according to the set mapping relationship.

[0060] The explanations in the foregoing embodiments are the same, and will not be repeated here.

[0061] Step 204: In response to the matching of the frequency of the second speech time-domain signal and the target howling frequency at each sampling time, determine the peak amplitude of the second speech time-domain signal in each period, as well as the time information of each peak amplitude.

[0062] In this embodiment, the second speech time-domain signal at each sampling moment in each cycle carries amplitude information, thus determining the amplitude value of the second speech time-domain signal at each sampling moment. Comparing multiple amplitude values ​​determines the peak amplitude, which is the maximum amplitude value. Similarly, the peak amplitude of the second speech time-domain signal in other cycles can be determined.

[0063] The timing information of the amplitude peaks in each period can be determined in the following way:

[0064] One implementation method is to determine the time sequence of the amplitude peaks in each cycle based on the time order of multiple cycles, and use this time sequence as the time information for each amplitude peak. For example, if there are 5 cycles, they are labeled 1, 2, 3, 4, and 5 in chronological order. Therefore, the time order of the amplitude peaks in the 5 cycles can be determined as 1, 2, 3, 4, and 5. A larger number indicates a later occurrence.

[0065] As a second implementation method, the time information of the amplitude peak of each cycle is the corresponding sampling time.

[0066] Step 205: Sort the multiple amplitude peaks according to the time information of the multiple amplitude peaks.

[0067] In this embodiment of the application, the amplitude peaks in multiple periods are sorted according to the time information of the amplitude peaks in multiple periods, for example, sorted in chronological order, i.e., ascending order, such as P1, P2, P3, P4 and P5.

[0068] Step 206: Based on the sorting, determine at least one difference between adjacent amplitude peaks.

[0069] In this embodiment of the application, based on the sorting, at least one difference is determined by subtracting adjacent amplitude peaks in the sorting, or by subtracting amplitude peaks with a set interval to determine at least one difference.

[0070] As an example, multiple amplitude peaks are ordered chronologically as P1, P2, P3, P4, and P5. The differences between these amplitude peaks are determined based on adjacent amplitude peaks in the order. This results in the following four differences:

[0071] Difference 1 = P2 - P1;

[0072] Difference 2 = P3 - P2;

[0073] Difference 3 = P4 - P3;

[0074] Difference 5 = P5 - P4.

[0075] Step 207: Compare at least one difference with a set threshold.

[0076] In this embodiment of the application, each difference is compared with a set threshold, wherein the set threshold indicates the change between adjacent amplitude peaks, for example, 2dB.

[0077] Step 208: In response to at least one difference being greater than a set threshold, the second speech time-domain signal at multiple sampling times is determined to be a howling speech signal.

[0078] If each difference is greater than a set threshold, i.e. P2-P1>Thr, P3-P2>Thr, P4-P3>Thr, and P5-P4>Thr, then the second speech time-domain signal at multiple sampling times is considered to be a howling speech signal, meaning that the first speech time-domain signal will howl.

[0079] Step 209: In response to any difference being less than or equal to a set threshold, determine that the second speech time-domain signal at multiple sampling times is not a howling speech signal.

[0080] If any difference is less than or equal to a set threshold, for example, if any difference is less than or equal to a set threshold, or if two differences are both less than or equal to a set threshold, then the second speech time-domain signal at multiple sampling times is determined not to be a howling speech signal, that is, the first speech time-domain signal will not howl.

[0081] In this embodiment, by tracking the amplitude peak of a single frequency across multiple cycles, the potential for howling is identified based on the difference between the amplitude peaks of a single frequency in adjacent cycles. Compared to directly comparing the amplitude value with the corresponding threshold, this method can predict whether howling will occur in advance. This is because if the amplitude peak is directly compared with the corresponding threshold, howling will have already occurred when the amplitude peak meets the howling requirement, making it impossible to predict in advance for howling suppression and resulting in a poor user experience.

[0082] Step 210: In response to the mismatch between the frequency of the second speech time-domain signal at at least one sampling time and the target howling frequency, determine that the second speech time-domain signal at multiple sampling times is not a howling speech signal.

[0083] In this embodiment of the application, if the frequency of the second speech time-domain signal at at least one sampling time does not match the target howling frequency, it is considered that the frequency of the second speech time-domain signal at at least one sampling time does not meet the requirement of the target howling frequency, and it is considered that the first speech time-domain signal will not howl.

[0084] In the method for detecting howling speech signals in this application embodiment, single-frequency signals at each sampling time are tracked in the time domain, so that the tracked second speech time-domain signal has both time-domain and frequency-domain information. Furthermore, the amplitude peaks of single frequencies across multiple cycles are tracked. The difference between the amplitude peaks of single frequencies in adjacent cycles is used to identify whether howling will occur. Compared to directly comparing the amplitude value with a corresponding threshold, this method can predict whether howling will occur in advance for howling suppression. This is because directly comparing the amplitude value with the corresponding threshold means that howling has already occurred when the amplitude peak meets the howling requirement, making it impossible to predict and suppress howling in advance, resulting in a poor user experience. Therefore, this application achieves the same effect of tracking in both the time and frequency domains by only performing detection in the time domain, thus improving the efficiency of howling speech signal detection.

[0085] To achieve the above embodiments, this application also proposes a device for detecting howling voice signals.

[0086] Figure 4 This is a schematic diagram of a device for detecting howling voice signals provided in an embodiment of this application.

[0087] like Figure 4 As shown, the device may include:

[0088] The acquisition module 41 is used to acquire a first speech time-domain signal of a set duration collected by the microphone; the set duration includes multiple sampling times.

[0089] The processing module 42 is used to perform frequency filtering processing on the first speech time-domain signal at each of the sampling times to obtain a second speech time-domain signal with a single frequency at each of the sampling times.

[0090] The determining module 43 is used to determine the frequency of the second speech time-domain signal at each of the sampling times according to the set mapping relationship.

[0091] The detection module 44 is used to detect howling voice signals in the second speech time-domain signals at the multiple sampling times, based on the amplitude of the second speech time-domain signals at the multiple sampling times, in response to the fact that the frequency of the second speech time-domain signal at each of the sampling times matches the target howling frequency.

[0092] Furthermore, in one implementation of this application embodiment, the set duration includes the periods of multiple howling voice signals; the processing module 42 is specifically used for:

[0093] For each of the aforementioned cycles, the first speech time-domain signal at each sampling moment within the cycle is input into the adaptive filter corresponding to each sampling moment for frequency filtering processing to obtain the second speech time-domain signal at a single frequency at each sampling moment.

[0094] In one implementation of this application embodiment, the detection module 44 is specifically used for:

[0095] Determine the peak amplitude of the second speech time-domain signal within each of the said periods, and the time information of each of the said peak amplitudes;

[0096] The multiple amplitude peaks are sorted according to their time information;

[0097] Based on the sorting, at least one difference between adjacent amplitude peaks is determined;

[0098] Each of the at least one difference values ​​is compared with a set threshold.

[0099] In response to the fact that at least one difference is greater than the set threshold, the second speech time-domain signal at the plurality of sampling times is determined to be a howling speech signal.

[0100] In one implementation of this application embodiment, the detection module 44 is further configured to:

[0101] In response to at least one of the differences being less than or equal to the set threshold, it is determined that the second speech time-domain signal at the plurality of sampling times is not a howling speech signal.

[0102] In one implementation of this application, the apparatus further includes:

[0103] The update module is used to obtain the residual signal of the adaptive filter corresponding to the previous sampling time for each sampling time in the period; update the parameters of the adaptive filter according to the residual signal of the previous sampling time to obtain the adaptive filter corresponding to the sampling time.

[0104] In one implementation of this application embodiment, the detection module 44 is further configured to:

[0105] In response to a mismatch between the frequency of the second speech time-domain signal and the target howling frequency at at least one of the sampling times, it is determined that the second speech time-domain signal at the plurality of sampling times is not a howling speech signal.

[0106] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and will not be repeated here.

[0107] In the feedback speech signal detection device of this application embodiment, single-frequency signals at each sampling time are tracked in the time domain, so that the tracked second speech time-domain signal has both time-domain and frequency-domain information. Furthermore, the amplitude peaks of single frequencies across multiple cycles are tracked. Based on the difference between the amplitude peaks of single frequencies in adjacent cycles, feedback is identified. Compared to directly comparing the amplitude value with a corresponding threshold, feedback can be predicted in advance for suppression. This is because directly comparing the amplitude value with a threshold means that feedback has already occurred when the amplitude peak meets the feedback requirement, making it impossible to predict and suppress feedback in advance, resulting in a poor user experience. Therefore, this application achieves the same effect of tracking in both the time and frequency domains by only performing detection in the time domain, thus improving the efficiency of feedback speech signal detection.

[0108] To implement the above embodiments, this application also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the foregoing method embodiments.

[0109] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing method embodiments.

[0110] To implement the above embodiments, this application also proposes a computer program product having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the foregoing method embodiments.

[0111] Figure 5This is a block diagram of an electronic device provided in an embodiment of this application. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0112] Reference Figure 5 The electronic device 800 may include one or more of the following components: processing component 802, memory 804, power component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.

[0113] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0114] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of this data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0115] Power component 806 provides power to various components of electronic device 800. Power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.

[0116] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0117] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0118] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0119] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0120] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0121] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0122] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of an electronic device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0123] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0124] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0125] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0126] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0127] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0128] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0129] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0130] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method of detecting a howling voice signal, characterized by, The method comprises: acquiring a first voice time domain signal collected by a microphone for a set time length; the set time length comprises a plurality of sampling time points; performing frequency filtering processing on the first voice time domain signal of each sampling time point to obtain a second voice time domain signal of a single frequency of each sampling time point; determining the frequency of the second voice time domain signal of each sampling time point according to a set mapping relationship; in response to the frequency of the second voice time domain signal of each sampling time point and a target howling frequency point both matching, performing howling voice signal detection on the second voice time domain signal of the plurality of sampling time points according to the amplitude of the second voice time domain signal of the plurality of sampling time points.

2. The method of claim 1, wherein, The set time length comprises a plurality of periods of howling voice signals; the frequency filtering processing on the first voice time domain signal of each sampling time point to obtain a second voice time domain signal of a single frequency of each sampling time point comprises: for each period, inputting the first voice time domain signal of each sampling time point in the period into an adaptive filter corresponding to each sampling time point to perform frequency filtering processing to obtain a second voice time domain signal of a single frequency of each sampling time point.

3. The method of claim 2, wherein, The howling voice signal detection on the second voice time domain signal of the plurality of sampling time points according to the amplitude of the second voice time domain signal of the plurality of sampling time points comprises: determining the amplitude peak value of the second voice time domain signal in each period and the time information of each amplitude peak value; sorting the plurality of amplitude peak values according to the time information of the plurality of amplitude peak values; determining at least one difference value between adjacent amplitude peak values based on the sorting; comparing the at least one difference value with a set threshold value respectively; in response to the at least one difference value being greater than the set threshold value, determining that the second voice time domain signal of the plurality of sampling time points is a howling voice signal.

4. The method of claim 3, wherein, The method further comprises: in response to at least one of the difference values being less than or equal to the set threshold value, determining that the second voice time domain signal of the plurality of sampling time points is not a howling voice signal.

5. The method of claim 2, wherein, The method further comprises: for each sampling time point in the period, acquiring a residual signal corresponding to a previous sampling time point of the sampling time point of the adaptive filter; updating the parameters of the adaptive filter according to the residual signal of the previous sampling time point to obtain an adaptive filter corresponding to the sampling time point.

6. The method of any one of claims 1-5, wherein, The method comprises: in response to the frequency of the second voice time domain signal of at least one sampling time point not matching the target howling frequency point, determining that the second voice time domain signal of the plurality of sampling time points is not a howling voice signal.

7. An apparatus for detecting a howling voice signal, characterized by comprising: The method comprises: an acquisition module, configured to acquire a first voice time domain signal collected by a microphone for a set time length; the set time length comprises a plurality of sampling time points; a processing module, configured to perform frequency filtering processing on the first voice time domain signal of each sampling time point to obtain a second voice time domain signal of a single frequency of each sampling time point; a determination module, configured to determine the frequency of the second voice time domain signal of each sampling time point according to a set mapping relationship; The detection module is configured to, in response to the frequency of the second speech time domain signal at each sampling time and the target howling frequency point being matched, detect the howling speech signal according to the amplitude of the second speech time domain signal at the plurality of sampling times.

8. The apparatus of claim 7, wherein, The set time length includes a plurality of periods of howling speech signals; and the processing module is specifically configured to: For each period, the first speech time domain signal at each sampling time in the period is input into a corresponding adaptive filter at the sampling time to perform frequency filtering processing, so as to obtain the second speech time domain signal of a single frequency at each sampling time.

9. An electronic device, comprising: A computer program product includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method according to any one of claims 1-6 when executing the program.

10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Voice call data detection method, device, storage medium and mobile terminal

    CN108449496A

  • Digital hearing aid squeal detection and suppression algorithm based on filter bank and hardware implementation method

    CN112954576A