A method for detecting noise in an audio signal and related electronic devices

By performing time-frequency conversion and vibration position detection on the audio signal, and using power threshold judgment, the noise problem when the speaker plays a specific audio signal is solved, improving the accuracy and user experience of noise detection.

CN114464212BActive Publication Date: 2025-07-04XIAN GLORY TERMINAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111007794.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-30
Publication Date
2025-07-04
Estimated Expiration
2041-08-30

AI Technical Summary

Technical Problem

In the prior art, electronic devices are prone to generate noise when playing specific audio signals through speakers, especially audio signals with uneven energy distribution such as piano sounds, resulting in damage to the auditory experience.

Method used

By converting the audio signal time-frequency, detecting the starting position, and determining whether a noise is generated based on the starting position and power value, using the power threshold of the spectrum signal to determine whether the audio signal is a noise signal, saving computing resources and improving detection accuracy.

Benefits of technology

It improves the accuracy of audio signal noise detection, reduces the possibility of electronic devices generating noise when playing audio, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114464212B_ABST
    Figure CN114464212B_ABST
Patent Text Reader

Abstract

The present application provides a method for detecting noise in an audio signal and related electronic devices. The method includes: obtaining N frames of first spectrum signals, where the N frames of first spectrum signals are signals obtained by performing time-frequency transformation on a first audio input signal; detecting a first starting position of the first audio input signal in the time domain based on the N frames of first spectrum signals; determining M frames of first target spectrum signals from the N frames of first spectrum signals based on the first starting position; and determining whether the first audio input signal is an audio signal generating noise based on the power values of the first target spectrum signals. In the above embodiment, since noise is easily generated during the starting process of the audio input signal, by detecting the starting position of the audio input signal and performing noise detection on a part of the audio signal corresponding to the starting position, the accuracy of detecting noise in the audio signal is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio signal detection, and in particular to a method for detecting noise in an audio signal and related electronic devices. Background Art

[0002] When electronic devices such as mobile phones and tablet computers play audio through speakers, due to the limitations of the structural design of the speakers, the sound quality of the electronic devices played through the speakers is far inferior to that of high-fidelity (HIFI) speakers and professional speakers. When an electronic device plays a certain type of specific audio signal (with uneven energy distribution in the frequency spectrum, mainly concentrated in the mid-low frequency, and strong transient nature of the concentrated energy) through the speaker, a lot of noise will be generated, seriously affecting the user's auditory experience. For this type of specific audio signal, it is called the first type of audio signal. The piano sound is a common first type of audio signal. When an electronic device plays music / video with piano sound through the speaker, a "burst-like" noise that fluctuates dynamically with the piano notes may be generated, that is, when a certain piano note appears, a "hissing" noise may be superimposed, which will seriously damage the original piano timbre and even mask the piano fundamental tone in severe cases. In addition, due to the differences in audio signal encoding and decoding on third-party music platforms, most audio sources have varying degrees of noise. Therefore, how to accurately detect whether there is noise in an audio signal is an issue that technicians are increasingly concerned about. Summary of the Invention

[0003] An embodiment of this application provides a method for detecting noise in an audio signal, which solves the problem of low accuracy in detecting noise in an audio signal.

[0004] In a first aspect, an embodiment of this application provides a method for detecting noise in an audio signal, which is characterized by including: obtaining N frames of first spectral signals; the N frames of first spectral signals are N frames of signals obtained by performing time-frequency transformation on a first audio input signal; detecting a first starting position of the first audio input signal in the time domain based on the N frames of first spectral signals; determining M frames of first target spectral signals from the N frames of first spectral signals based on the first starting position; and determining whether the first audio input signal is an audio signal generating noise based on the power values of the first target spectral signals. In the above embodiment, since noise is easily generated during the starting process of the audio input signal, by detecting the starting position of the audio input signal and performing noise detection on a part of the audio signal corresponding to the starting position, the accuracy of detecting noise in the audio signal is higher.

[0005] In combination with the first aspect, in a possible implementation manner, it is determined whether the first audio input signal is an audio signal generating noise based on the power value of the first target spectral signal, specifically including: determining whether there is a situation where the power value of the first target spectral signal is greater than a first threshold; if so, determining that the first audio input signal is an audio signal generating noise; if not, determining that the first audio input signal is an audio signal not generating noise. In this way, the electronic device can determine whether the audio input signal of the startup part is an audio signal generating noise, and the accuracy of the electronic device for noise detection is higher.

[0006] In combination with the first aspect, in a possible implementation manner, N frames of second spectral signals are obtained; the N frames of second spectral signals are N frames of signals obtained by performing time-frequency transformation on the second audio input signal, and the first audio input signal is the original signal of the second input audio signal; the second startup position of the second audio input signal in the time domain is detected based on the N frames of second spectral signals; M frames of second target spectral signals are determined from the N frames of second spectral signals based on the second startup position; it is determined whether the first audio input signal is an audio signal generating noise based on the power value of the first target spectral signal, specifically including: determining whether there is a situation where the power value of the i-th frame of the second signal is greater than the power value of the i-th frame of the first signal, where the second signal is a signal in the second target spectral signal with a power value greater than a second threshold, and the first signal is a signal in the first target spectral signal with a power value greater than the second threshold; if so, determining that the first audio input signal is an audio signal generating noise; if not, determining that the first audio input signal is not an audio signal generating noise. In the above embodiment, the second audio input signal can be the first audio input signal after being played back through the electronic device speaker. By comparing the power values of the frequency points with power values greater than the second threshold in the startup part of the second audio input signal with the power values of the frequency points with power values greater than the second threshold in the startup part of the first audio input signal, it can be determined whether non-linear distortion occurs during the playback of the first audio input signal through the speaker, thereby generating noise. Through the above method for audio noise detection, the detection result is more accurate.

[0007] In combination with the first aspect, in a possible implementation manner, detecting the first starting position of the first audio input signal in the time domain based on N frames of first spectral signals specifically includes: calculating the first difference value of each frame of the first spectral signal according to the average power of each frame of the first spectral signal; performing inter-frame interpolation on the first difference value of each frame of the first spectral signal to obtain a first audio signal; downsampling the first audio signal to obtain a second audio signal; calculating the first starting dynamic threshold of the second audio signal; and determining the first starting position based on the second audio signal and the first starting dynamic threshold. In this way, the electronic device can determine the starting position of the first audio input signal in the time domain, which is beneficial for the electronic device to determine the part related to starting in the first audio input signal, so that the electronic device only performs spectral noise analysis on the signal related to starting during the process of detecting background noise, rather than performing spectral noise analysis on the entire audio input signal, thereby saving a large amount of computing resources of the electronic device.

[0008] In combination with the first aspect, in a possible implementation manner, calculating the first starting dynamic threshold of the second audio signal specifically includes: calculating the first starting dynamic threshold according to the formula Thresh(m)′ = Delta.Median(Y(m)′); where Thresh(m)′ is the first starting dynamic threshold of the second audio signal at the m-th sampling point in the time domain, Y(m)′ is the power value of the second audio signal at the m-th sampling point in the time domain, and Delta is the first constant coefficient. In this way, it is beneficial for the electronic device to determine the starting position of the first audio input signal in the time domain based on the calculated first starting dynamic threshold and the second audio signal, so that the electronic device can determine the partial signal related to starting in the first audio input signal, enabling the electronic device to only perform noise analysis on this part of the signal, thereby achieving the purpose of detecting background noise and saving a large amount of computing resources of the electronic device.

[0009] In combination with the first aspect, in a possible implementation manner, determining the first starting position based on the second audio signal and the first starting dynamic threshold specifically includes: calculating the first power difference of the second audio signal according to the formula Diff(m)′ = Y(m)′ - Thresh(m)′; Diff(m)′ is the first power difference of the second audio signal at the m-th sampling point in the time domain, Y(m)′ is the power value of the second audio signal at the m-th sampling point in the time domain, and Thresh(m)′ is the first starting dynamic threshold of the second audio signal at the m-th sampling point in the time domain; determining the maximum value of the first power difference; obtaining the starting occurrence time t of the second audio signal based on the maximum value; i ; based on t i calculating the first starting position (t i - t1 to t i+(t2), where t1 is the first starting oscillation time length and t2 is the second starting oscillation time length. In this way, the electronic device can determine the part of the signal related to the starting oscillation in the first audio input signal according to the first starting oscillation position, so that the electronic device can only perform noise analysis on this part of the signal, thereby achieving the purpose of detecting background noise and saving a large amount of computing resources of the electronic device.

[0010] Combined with the first aspect, in a possible implementation manner, the first target spectral signal is the first spectral signal with sampling points within the first starting oscillation position.

[0011] Combined with the first aspect, in a possible implementation manner, detecting the second starting oscillation position of the second audio input signal in the time domain based on N frames of second spectral signals specifically includes: calculating the second difference value of each frame of the second spectral signal according to the average power of each frame of the second spectral signal; performing inter-frame interpolation on the second difference values of each frame of the second spectral signal to obtain a third audio signal; downsampling the third audio signal to obtain a fourth audio signal; calculating the second starting oscillation dynamic threshold of the fourth audio signal; and determining the second starting oscillation position based on the fourth audio signal and the second starting oscillation dynamic threshold. In this way, the electronic device can determine the starting oscillation position of the second audio input signal in the time domain, which is beneficial for the electronic device to determine the part related to the starting oscillation in the second audio input signal, so that the electronic device only performs spectral noise analysis on the signal related to the starting oscillation during the background noise detection process, rather than performing spectral noise analysis on the entire audio input signal, thereby saving a large amount of computing resources of the electronic device.

[0012] Combined with the first aspect, in a possible implementation manner, calculating the second starting oscillation dynamic threshold of the fourth audio signal specifically includes: calculating the second starting oscillation dynamic threshold according to the formula Thresh(m)″ = Delta · Median(Y(m)″); where Thresh(m)″ is the second starting oscillation dynamic threshold of the m-th sampling point in the time domain of the fourth audio signal, Y(m)″ is the power value of the m-th sampling point in the time domain of the fourth audio signal, and the Delta is the first constant coefficient. In this way, it is beneficial for the electronic device to determine the starting oscillation position of the second audio input signal in the time domain based on the calculated second starting oscillation dynamic threshold and the fourth audio signal, so that the electronic device can determine the part of the signal related to the starting oscillation in the second audio input signal, so that the electronic device can only perform noise analysis on this part of the signal, thereby achieving the purpose of detecting background noise and saving a large amount of computing resources of the electronic device.

[0013] In combination with the first aspect, in a possible implementation manner, the second oscillation start position is determined based on the fourth audio signal and the second oscillation start dynamic threshold, which specifically includes: calculating the second power difference of the fourth audio signal according to the formula Diff(m)″ = Y(m)″ - Thresh(m)″; Diff(m)″ is the second power difference of the fourth audio signal at the m-th sampling point in the time domain, Y(m)″ is the power value of the fourth audio signal at the m-th sampling point in the time domain, and Thresh(m)″ is the second oscillation start dynamic threshold of the fourth audio signal at the m-th sampling point in the time domain; determining the maximum value of the second power difference; obtaining the oscillation start occurrence time t of the fourth audio signal based on this maximum value j ; based on the t j calculate the second oscillation start position (t j -t1 to t j +t2), where t1 is the second oscillation start time length and t2 is the second oscillation start time length. In this way, the electronic device can determine the part of the signal related to the oscillation start in the fourth audio input signal according to the second oscillation start position, so that the electronic device can only perform noise analysis on this part of the signal, thereby achieving the purpose of detecting background noise and saving a large amount of computing resources of the electronic device.

[0014] In combination with the first aspect, in a possible implementation manner, the second target spectrum signal is the second spectrum signal with sampling points within the second oscillation start position.

[0015] In a second aspect, an embodiment of the present application provides an electronic device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to cause the electronic device to perform: obtaining N frames of first spectrum signals; the N frames of first spectrum signals are N frames of signals obtained by performing time-frequency transformation on the first audio input signal; detecting the first oscillation start position of the first audio input signal in the time domain based on the N frames of first spectrum signals; determining M frames of first target spectrum signals from the N frames of first spectrum signals based on the first oscillation start position; determining whether the first audio input signal is an audio signal that generates background noise based on the power value of the first target spectrum signal.

[0016] In combination with the second aspect, in a possible implementation manner, the one or more processors are further used to call the computer instructions to cause the electronic device to perform: determining whether there is a power value of the first target spectrum signal greater than a first threshold; if so, determining that the first audio input signal is an audio signal that generates background noise; if not, determining that the first audio input signal is an audio signal that does not generate background noise.

[0017] In combination with the second aspect, in a possible implementation manner, the one or more processors are further configured to call the computer instructions to cause the electronic device to perform: obtaining N frames of second spectrum signals; the N frames of second spectrum signals are N frames of signals obtained by performing time-frequency transformation on a second audio input signal, and the first audio input signal is the original signal of the second input audio signal; detecting a second starting position of the second audio input signal in the time domain based on the N frames of second spectrum signals; determining M frames of second target spectrum signals from the N frames of second spectrum signals based on the second starting position; determining whether the first audio input signal is a noise-generating audio signal based on the power value of the first target spectrum signal, specifically including: determining whether there is a power value of the i-th frame of the second signal greater than the power value of the i-th frame of the first signal, where the second signal is a signal in the second target spectrum signal with a power value greater than a second threshold, and the first signal is a signal in the first target spectrum signal with a power value greater than the second threshold; if so, determining that the first audio input signal is a noise-generating audio signal; if not, determining that the first audio input signal is not a noise-generating audio signal.

[0018] In combination with the second aspect, in a possible implementation manner, the one or more processors are further configured to call the computer instructions to cause the electronic device to perform: detecting a first starting position of the first audio input signal in the time domain based on the N frames of first spectrum signals, specifically including: calculating a first difference value of each frame of the first spectrum signal according to the average power value of each frame of the first spectrum signal; performing inter-frame interpolation on the first difference values of each frame of the first spectrum signal to obtain a first audio signal; downsampling the first audio signal to obtain a second audio signal; calculating a first starting dynamic threshold of the second audio signal; and determining the first starting position based on the second audio signal and the first starting dynamic threshold.

[0019] In combination with the second aspect, in a possible implementation manner, the one or more processors are further configured to call the computer instructions to cause the electronic device to perform: calculating the first starting dynamic threshold of the second audio signal, specifically including: calculating the first starting dynamic threshold according to the formula Thresh(m)′ = Delta·Median(Y(m)′); where Thresh(m)′ is the first starting dynamic threshold of the second audio signal at the m-th sampling point in the time domain, Y(m)′ is the power value of the second audio signal at the m-th sampling point in the time domain, and Delta is a first constant coefficient.

[0020] In combination with the second aspect, in a possible implementation manner, the one or more processors are further configured to call the computer instructions to cause the electronic device to perform: determining a first startup position based on a second audio signal and a first startup dynamic threshold, specifically including: calculating a first power difference value of the second audio signal according to the formula Diff(m)' = Y(m)' - Thresh(m)'; Diff(m)' is the first power difference value of the second audio signal at the m-th sampling point in the time domain, Y(m)' is the power value of the second audio signal at the m-th sampling point in the time domain, and Thresh(m)' is the first startup dynamic threshold of the second audio signal at the m-th sampling point in the time domain; determining a maximum value of the first power difference value; obtaining a startup occurrence time t of the second audio signal based on the maximum value i ; based on t i calculate the first startup position (t i - t1 to t i + t2), where t1 is the first startup time length and t2 is the second startup time length.

[0021] In combination with the second aspect, in a possible implementation manner, the one or more processors are further configured to call the computer instructions to cause the electronic device to perform: detecting a second startup position of a second audio input signal in the time domain based on N frames of second spectrum signals, specifically including: calculating a second difference value of each frame of the second spectrum signal according to the average power of each frame of the second spectrum signal; performing inter-frame interpolation on the second difference values of each frame of the second spectrum signal to obtain a third audio signal; downsampling the third audio signal to obtain a fourth audio signal; calculating a second startup dynamic threshold of the fourth audio signal; determining the second startup position based on the fourth audio signal and the second startup dynamic threshold.

[0022] In combination with the second aspect, in a possible implementation manner, the one or more processors are further configured to call the computer instructions to cause the electronic device to perform: calculating a second startup dynamic threshold of the fourth audio signal, specifically including: calculating the second startup dynamic threshold according to the formula Thresh(m)'' = Delta · Median(Y(m)''), where Thresh(m)'' is the second startup dynamic threshold of the fourth audio signal at the m-th sampling point in the time domain, Y(m)'' is the power value of the fourth audio signal at the m-th sampling point in the time domain, and Delta is a first constant coefficient.

[0023] In combination with the second aspect, in a possible implementation manner, the one or more processors are further configured to call the computer instructions to cause the electronic device to execute: determining a second startup position based on a fourth audio signal and a second startup dynamic threshold, specifically including: calculating a second power difference of the fourth audio signal according to the formula Diff(m)″ = Y(m)″ - Thresh(m)″; Diff(m)″ is the second power difference of the fourth audio signal at the m-th sampling point in the time domain, Y(m)″ is the power value of the fourth audio signal at the m-th sampling point in the time domain, and Thresh(m)″ is the second startup dynamic threshold of the fourth audio signal at the m-th sampling point in the time domain; determining a maximum value of the second power difference; obtaining a startup occurrence time t of the fourth audio signal based on the maximum value j ; based on the t j calculate a second startup position (t j - t1 to t j + t2), where t1 is the second startup time length and t2 is the second startup time length.

[0024] In a third aspect, an embodiment of the present application provides an electronic device, including: a touch screen, a camera, one or more processors, and one or more memories; the one or more processors are coupled to the touch screen, the camera, and the one or more memories, and the one or more memories are configured to store computer program code, and the computer program code includes computer instructions. When the one or more processors execute the computer instructions, the electronic device is caused to execute the method as described in the first aspect or any one of the possible implementation manners of the first aspect.

[0025] In a fourth aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device. The chip system includes one or more processors, and the processor is configured to call computer instructions to cause the electronic device to execute the method as described in the first aspect or any one of the possible implementation manners of the first aspect.

[0026] In a fifth aspect, an embodiment of the present application provides a computer program product containing instructions. When the computer program product runs on an electronic device, the electronic device is caused to execute the method as described in the first aspect or any one of the possible implementation manners of the first aspect.

[0027] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions. When the instructions run on an electronic device, the electronic device is caused to execute the method as described in the first aspect or any one of the possible implementation manners of the first aspect. Description of the Drawings

[0028] Figure 1 is a schematic hardware structure diagram of an electronic device 100 provided by an embodiment of the present application;

[0029] Figure 2 It is a schematic diagram of the startup process of an audio signal provided by an embodiment of the present application;

[0030] Figures 3A - 3E It is a schematic diagram of a scenario for detecting noise in an audio signal provided by an embodiment of the present application;

[0031] Figure 4 It is a flowchart of a method for detecting noise in an audio signal provided by an embodiment of the present application;

[0032] Figure 5 It is a schematic diagram of the process for determining the first target spectral signal provided by an embodiment of the present application;

[0033] Figure 6A It is a spectrogram of the first spectral signal provided by an embodiment of the present application;

[0034] Figure 6B It is a spectrogram of the first spectral signal after power weighting provided by an embodiment of the present application;

[0035] Figure 7A It is an amplitude-time domain diagram of the first audio input signal provided by an embodiment of the present application;

[0036] Figure 7B It is a power-time domain diagram of the first audio input signal provided by an embodiment of the present application;

[0037] Figure 8A It is a power-time domain diagram of the first original audio signal provided by an embodiment of the present application;

[0038] Figure 8B It is a relationship diagram between the first difference value of the first original audio signal and the time domain provided by an embodiment of the present application;

[0039] Figure 9 It is a Diff curve diagram of the second audio signal provided by an embodiment of the present application;

[0040] Figures 10A - 10I It is a schematic diagram of another scenario for detecting noise in an audio signal provided by an embodiment of the present application;

[0041] Figure 11 It is a flowchart of another method for detecting noise in an audio signal provided by an embodiment of the present application;

[0042] Figure 12 It is a schematic diagram of the process for determining the second target spectral signal provided by an embodiment of the present application. Detailed implementation manners

[0043] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The mention of "embodiment" in this article means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. Those skilled in the art can explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0044] The terms "first", "second", "third", etc. in the specification and claims of the present application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a series of steps or units are included, or optionally, steps or units not listed are further included, or optionally, other steps or units inherent to these processes, methods, products, or devices are further included.

[0045] Only parts related to the present application are shown in the accompanying drawings, rather than all the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but there can also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0046] The terms "component", "module", "system", "unit", etc. used in this specification are used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a unit can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or distributed between two or more computers. In addition, these units can be executed from various computer-readable media storing various data structures. A unit can communicate, for example, through local and / or remote processes according to a signal having one or more data packets (such as data from a second unit interacting with another unit between a local system, a distributed system, and / or a network. For example, the Internet interacting with other systems through a signal).

[0047] The structure of the electronic device 100 will be introduced below. Please refer to Figure 1 , Figure 1 which is a schematic diagram of the hardware structure of the electronic device 100 provided in an embodiment of the present application.

[0048] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0049] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0050] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0051] The wireless communication function of the electronic device 100 can be implemented by Antenna 1, Antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0052] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: Antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0053] The mobile communication module 150 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by Antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves through Antenna 1 and radiate it out. In some embodiments, at least some functional modules of the mobile communication module 150 can be disposed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 can be disposed in the same device.

[0054] The wireless communication module 160 can provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), BLE broadcast, global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via Antenna 2, frequency-modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signals to be transmitted from the processor 110, frequency-modulate and amplify them, and convert them into electromagnetic waves through Antenna 2 and radiate them out.

[0055] The electronic device 100 realizes the display function through the GPU, the display screen 194, the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0056] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.

[0057] The electronic device 100 can realize the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, the application processor, etc.

[0058] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera sensor. The optical signal is converted into an electrical signal, and the camera sensor transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP may be set in the camera 193.

[0059] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0060] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc.

[0061] The electronic device 100 can implement audio functions through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and the application processor, etc. Such as music playback, recording, etc.

[0062] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0063] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or hands-free calls through the speaker 170A.

[0064] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the voice can be received by bringing the receiver 170B close to the human ear.

[0065] The microphone 170C, also known as the "microphone", "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to implement functions such as collecting sound signals, noise reduction, and also identifying the sound source and implementing a directional recording function, etc.

[0066] The pressure sensor 180A is used to sense pressure signals and can convert pressure signals into electrical signals. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194.

[0067] The barometric pressure sensor 180C is used to measure barometric pressure. In some embodiments, the electronic device 100 calculates the altitude based on the barometric pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.

[0068] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip case.

[0069] The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to functions such as horizontal and vertical screen switching and pedometer applications.

[0070] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access to application locks, fingerprint photography, fingerprint answering of incoming calls, etc.

[0071] The touch sensor 180K, also known as the "touch panel". The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also known as the "touch screen". The touch sensor 180K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100 at a different position from the display screen 194.

[0072] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire the vibration signals of the vibrating bone mass of the human vocal tract.

[0073] When the electronic device plays audio through the built-in speaker, due to the size limitation of the device, the size of the speaker is relatively small, and the allowable membrane vibration amplitude is small. When the volume of the external audio is too large, it will cause the membrane vibration amplitude of the speaker to exceed the maximum value, resulting in easy distortion of the sound when playing at a high volume and generating a hissing-like noise. For those audio signals with uneven energy distribution in the frequency domain and mainly concentrated in the mid-low frequency and strong transient energy concentration (such as piano sounds), when played through the speaker, it is more likely to generate noise. Therefore, how to accurately detect whether the audio signal generates noise so as to process the audio signal that is about to generate noise is an issue that current technicians are increasingly concerned about.

[0074] To solve the problem of insufficient accuracy in detecting background noise in audio signals, an embodiment of the present application provides a method for detecting audio signals. By identifying and marking the start and decay parts of an audio signal and analyzing the spectral energy of the harmonic / noise components of the signal between the start and decay, if the spectral energy is greater than or equal to a first threshold, it is determined that the audio signal is an audio signal generating background noise, so that an electronic device can process the audio signal to avoid generating background noise when the electronic device plays the audio signal through a speaker. As Figure 2 shown, the change process of an audio signal from a minimum energy value to a maximum energy value is called the start of the audio signal. During the start process of the audio signal, the energy of the audio signal gradually increases. When the energy of the audio signal exceeds the maximum limit of hardware devices such as speakers, it will cause non-linear distortion of the audio signal, thereby generating noise.

[0075] Next, the application scenarios of the method for detecting background noise in the audio signal will be introduced in combination with Figures 3A - 3E .

[0076] As Figure 3A shown, the main interface of the electronic device includes multiple application icons. When the electronic device detects an input operation (e.g., double-click) on the audio detection icon 30, the electronic device starts the audio detection software and displays an audio detection interface 31 as shown in Figure 3B .

[0077] As Figure 3B shown, the audio detection interface 31 includes an add audio icon 311. When the electronic device detects an input operation (e.g., single-click) on the add audio icon 311, the electronic device displays an audio addition interface 32 as shown in Figure 3C .

[0078] As Figure 3C shown, the audio addition interface 312 is used to add the audio to be detected. The audio addition interface includes multiple pieces of local song information ( Figure 3C only lists songs 1, 2, 3, and 4), and the duration of each song (e.g., song 1 is 50s, song 2 is 2 minutes and 50 seconds, song 3 is 10 seconds, and song 4 is 1 minute and 25 seconds). When the electronic device detects that the user checks song 1 and detects a single-click operation on the confirm control 321, the electronic device displays a detection interface 33 as shown in Figure 3D .

[0079] As Figure 3D shown, the detection interface 33 is used to display whether the current audio detection software is detecting whether there is background noise in song 1. After the audio detection software finishes the detection, the electronic device displays a detection report interface 34 as shown in Figure 3E .

[0080] As shown Figure 3E in the detection report interface 34, it includes the audio information indicating that Song 1 has background noise, the time point when Song 1 has background noise, and the power value / energy value corresponding to Song 1 at this time point.

[0081] Next, in conjunction with the accompanying drawings, the specific process of the electronic device for detecting whether the audio signal has background noise will be described. The electronic device includes an audio detection module. Please refer to Figure 4 , Figure 4 which is a flowchart of a method for detecting background noise of an audio signal provided by an embodiment of the present application. The specific process is as follows:

[0082] Step S401: The audio detection module acquires a first audio input signal.

[0083] Specifically, the first audio input signal is an audio signal in an audio application. For example, in the above Figures 3A - 3E embodiment, the audio detection module can acquire the audio input signal stored in the local memory of the electronic device, or can acquire the audio input signal online through an audio application / browser or other applications. The manner in which the audio detection module acquires the first audio input signal is only an example in the embodiment of the present application and is not limited.

[0084] Step S402: The audio detection module performs frame division processing on the first audio input signal to obtain N frames of first original audio signals.

[0085] Specifically, after the audio detection module acquires the first audio input signal, it samples the first audio input signal to obtain multiple sampling points. In this way, the first audio input signal can be converted from a time-domain analog signal to a time-domain digital signal. Then, the audio detection module divides the audio input signal into N frames of first original audio signals with equal lengths.

[0086] Step S403: The audio detection module performs time-frequency transformation on the N frames of first original audio signals to obtain N frames of first spectrum signals.

[0087] Specifically, the audio detection module can perform time-frequency transformation on the N frames of first original audio signals by using discrete Fourier transform (DFT), or can use fast Fourier transform (FFT), or can also use short-time Fourier transform (STFT). The embodiment of the present application does not limit this. The signal obtained through time-frequency transformation is the first spectrum signal.

[0088] In some embodiments, the electronic device may divide each frame of the first original audio signal into a spectral signal corresponding to L frequency points through a 2L-point FFT. Wherein, L is an integer power of 2, and the value of L is determined by the computing power of the electronic device. The greater the computing and processing power of the electronic device, the greater the value of L. For example, each frame of the first original audio signal may be divided into a first spectral signal corresponding to 1024 frequency points through a 2048-point DFT, and then the first spectral signal may be represented by an array, and the array includes 1024 elements. Any element is used to represent a frequency point, which includes two values, one value represents the frequency (HZ) of the first spectral signal corresponding to the frequency point, and the other value represents the power of the first spectral signal corresponding to the frequency point (in units of dB). It should be understood that in addition to the array, the electronic device may also express the first spectral signal in other ways, such as a matrix, etc., and the embodiments of the present application do not limit this.

[0089] Step S404: The audio detection module detects the first starting position of the first audio input signal in the time domain based on the N frames of the first spectral signals, and determines M frames of first target spectral signals from the N frames of the first spectral signals based on the first starting position.

[0090] Specifically, the first starting position is the starting time range of the first audio input signal in the time domain. The specific process of the audio detection module detecting the starting position of the N frames of the first spectral signals and determining M frames of first target spectral signals from the N frames of the first spectral signals is as Figure 5 shown:

[0091] S501. The audio detection module performs power weighting on each frame of the first spectral signal in the frequency domain to obtain the first spectral signal after power weighting.

[0092] Specifically, the purpose of the audio detection module performing power weighting on each frame of the spectral signal is to amplify the energy of the high frequency, which is beneficial to the detection of transient components in the audio signal and the detection of the starting position. Let X(n, k) represent the spectral signal obtained through time-frequency transformation, where n represents the nth frame of the spectral signal, the value range of n is 1 to N, k is the kth frequency point of the nth frame of the spectral signal, and X(n, k) is used to characterize the power value of the kth frequency point of the nth frame of the spectral signal (the amplitude corresponding to the kth frequency point in the frequency domain). The audio detection module can weight the power value corresponding to each frequency point on each frame of the first spectral signal through formula (1), and formula (1) is as follows:

[0093] X′(n, k) = k · abs(X(n, k)) 2 (1)

[0094] Among them, X′(n, k) is the power value after power weighting of the k-th frequency point of the first spectral signal of the n-th frame, and abs(X(n, k)) 2 Take the absolute value of the power value corresponding to the k-th frequency point of the first spectral signal of the n-th frame. As can be seen from formula (1), for frequency points with larger frequencies (larger k values), the corresponding power values are amplified by a larger multiple after weighting.

[0095] Exemplarily, in combination with Figures 6A - 6B The process of power weighting the audio detection module is described in detail: If Figure 6A is the spectrogram of the first spectral signal X(1, k), assuming that X(1, k) has four frequency points (k = 1, 2, 3, 4), the frequencies corresponding to these four frequency points are f1 to f4 respectively, and the power values corresponding to these four frequency points are 1 dB, 2 dB, 3 dB, and 4 dB respectively. Then, as can be seen from formula (1), the weighted power corresponding to the first frequency point is 1 * 1 = 1 dB, and the weighted power corresponding to the second frequency point is 2 * 2 2 = 8 dB, the weighted power corresponding to the third frequency point is 3 * 3 2 = 27 dB, the weighted power corresponding to the fourth frequency point is 4 * 4 2 = 64 dB, and X′(1, k) is as Figure 6B shown. The power values corresponding to high-frequency points are significantly increased. Since the transient components of the audio signal energy are generally concentrated in the high frequency, therefore, amplifying the high-frequency energy in the frequency domain is beneficial to detecting the transient components, and thus beneficial to detecting the starting position of the audio signal.

[0096] S502. The audio detection module calculates the power average value of the first spectral signal after power weighting for each frame.

[0097] Specifically, the electronic device can calculate the power average value of the spectral signal after power weighting for each frame according to formula (2). Formula (2) is as follows:

[0098]

[0099] Among them, L is the frequency point of the first spectral signal of each frame, and X″(n, k) is the power average value corresponding to the k-th frequency point of the first spectral signal of the n-th frame. In this way, the audio detection module can calculate the power average values of N frames of the first spectral signal through the above formula (2), and the time corresponding to the power average value is the middle time of each frame of the audio signal.

[0100] Exemplarily, in combination with Figures 7A - 7B the above process is described. As Figure 7AThe figure shows the amplitude-time domain diagram of the first audio input signal. The horizontal axis represents the time of the audio input signal, and the vertical axis represents the amplitude of the audio input signal. The audio detection module frames the audio input signal to obtain N frames of audio signals. Assuming the length of each frame of audio signal is t, the time corresponding to the first frame of audio signal is 0 to t. Then, the intermediate time corresponding to the first frame of audio signal is Similarly, the intermediate time of the second frame of audio signal is When the audio detection module calculates the average power value of each frame of the first spectral signal in the frequency domain, the power-time domain diagram as shown in Figure 7B can be obtained. As shown in Figure 7B , the power value corresponding to the intermediate time of each frame of audio signal is the average power value corresponding to that frame of audio signal. In this way, the power of the first audio input signal can be combined with the time domain.

[0101] S503. The audio detection module performs a first-order difference calculation based on the average power value of each frame of the first spectral signal to obtain the first difference value of each frame of the first spectral signal.

[0102] Specifically, after calculating the average power value of each frame of the first spectral signal, a first-order difference calculation is performed based on this average power value. Exemplarily, in combination with Figures 8A - 8B the calculation process of the first-order difference is described in detail. As shown in Figure 8A , if there are 9 frames of the first original audio signals, each frame of the first original audio signal corresponds to an average power value. It is set that the difference value of the first frame of the first original audio signal is 0, the difference value of the second frame of the first original audio signal is the difference between the average power value of the second frame of the first original audio signal and the average power value of the first frame of the first original audio signal, the difference value of the third frame of the first original audio signal is the difference between the average power value of the third frame of the first original audio signal and the average power of the second frame of the first original audio signal, and so on. In this way, the first difference value D(n) of each frame of the first original audio signal as shown in Figure 8B can be calculated.

[0103] S504. The audio detection module performs inter-frame interpolation based on the first difference value of each frame of the first original audio signal to obtain the first audio signal.

[0104] Specifically, after the audio detection module obtains the first audio input signal, it samples the input signal, converting the first audio input signal from an analog signal in the time domain to a digital signal in the time domain. The number of sampling points is equal to the product of the sampling frequency and the duration of the audio input signal. For example, if the duration of the first audio input signal is 10 seconds and the sampling frequency is 48 KHZ, then the number of sampling points is 48,000. Thus, it can be seen that the number of sampling points is generally much larger than the number of frames of the framed audio signal. To improve the accuracy of the audio detection module in detecting the starting position of the first audio input signal, the audio detection module can perform inter-frame interpolation calculations based on the power difference values of each frame of the first original audio signal, so that each sampling point of the first audio input signal in the time domain corresponds to a power value, and the signal obtained by inter-frame difference is the first audio signal. In the time domain, each sampling point of the first audio signal corresponds to a power value, thus combining the energy in the frequency domain of the first original audio signal with the time sampling points in the time domain, so that the subsequent audio detection module can detect the starting time range according to the energy change of the first audio signal in the time domain. The specific process of the audio detection module performing inter-frame interpolation is as follows: By using the first-order function fitting method, the difference values corresponding to N frames of the first original audio signal are interpolated to the time domain sampling points L of the entire audio signal, thereby obtaining the spectral power difference value corresponding to each sampling point.

[0105] S505. The audio detection module downsamples the first audio signal to obtain a second audio signal.

[0106] Specifically, the sampling frequency of the first audio signal obtained by inter-frame interpolation is much lower than that of the first audio input signal. To improve the calculation efficiency, the audio detection module needs to perform downsampling processing on the first audio signal in the time domain. For example, usually the sampling frequency of music digital audio signals can reach 48 KHZ. However, the sampling frequency of the second audio signal obtained by inter-frame interpolation based on the first difference value D(n) is much lower than 48 KHZ. If the audio detection module detects the starting position of the first audio signal in the time domain based on 48 KHZ, even if the length of the first audio signal is 1 s, there are 48,000 sampling points, which means that the audio detection module needs to detect these 48,000 sampling points, increasing the calculation amount of the audio detection module. Therefore, it is necessary to perform downsampling processing on the first audio signal to reduce unnecessary sampling points without affecting the correctness of the detection result, thereby reducing the calculation pressure on the audio detection module.

[0107] The audio detection module first performs low-pass filtering on the first audio signal, and then downsamples the first audio signal. The method of downsampling is as follows: increase the sampling frequency to reduce the sampling points of the first audio signal. The audio detection module can perform downsampling on the first audio signal according to formula (3), and formula (3) is shown as follows:

[0108] Y(m)′=Y(Kl) (3)

[0109] Where Y(m)′ is the power value corresponding to the m-th sampling point of the second audio signal in the time domain, K is the sampling multiple, and Y(Kl) is the power value corresponding to the K·l-th sampling point of the first audio signal in the time domain.

[0110] S506. The audio detection module calculates the first startup dynamic threshold of the second audio signal.

[0111] Specifically, the audio detection module can calculate the first startup dynamic threshold according to formula (4), and formula (4) is shown as follows:

[0112] Thresh(m)′=Delta·Median(Y(m)′) (4)

[0113] Where Delta is the first constant coefficient, which can be obtained based on historical data, or based on empirical values, or based on experimental data, and the embodiments of the present application do not make limitations. Thresh(m)′ is the first startup dynamic threshold of the second audio signal at the m-th sampling point in the time domain, and Median(Y(m)′) is used to represent the median calculation of the power values of the m sampling points of the second audio signal and one or more adjacent sampling points.

[0114] S507. The audio detection module calculates the first power difference between the second audio signal and the first startup dynamic threshold.

[0115] Specifically, the audio detection module can obtain the power deviation value of each sampling point of the second audio signal according to formula (5), and formula (5) is shown as follows:

[0116] Diff(m)′=Y(m)′-Thresh(m)′ (5)

[0117] Where Diff(m)′ is the first power difference of the second audio signal at the m-th sampling point in the time domain.

[0118] S508. The audio detection module determines M frames of first target spectrum signals from the N frames of first spectrum signals based on the first power difference.

[0119] Specifically, the audio detection module can detect the peak points of the Diff to obtain the sampling point positions of several local maxima, and the corresponding sampling points are the time points t when the oscillation starts. i , and determine the oscillation start time range (t i -t1 to t i +t2), the M-frame first spectral signals of the sampling points within the oscillation start time range, the time point t1 when the oscillation occurs, and other information. Among them, the oscillation start time range (t i -t1 to t i +t2) is the first oscillation start position of the first audio input signal in the time domain, and the M-frame first spectral signals of the sampling points within the oscillation start time range are the first target spectral signals.

[0120] Exemplarily, Figure 9 the Diff(t) curve obtained by fitting the Diff(m) of the second audio signal. Points A and B are the two peak points of Diff(t), and the times of the corresponding sampling points are 11 ms and 21 ms respectively. If t1 is 2 ms and t2 is 1 ms, then the audio detection module can determine that the oscillation start time ranges are 9 ms to 12 ms and 19 ms to 22 ms respectively. If the signal length corresponding to each frame of the first spectral signal in the time domain is 5 ms, then the audio detection module determines that the second frame, the third frame, the fourth frame, and the fifth frame of the first spectral signals are the first target spectral signals, and records information such as the time points (11 ms, 21 ms) when the oscillation occurs.

[0121] Among them, t1 and t2 can be obtained from historical data, can be obtained from empirical values, or can be obtained from experimental data. The embodiments of the present application do not limit this.

[0122] Step S405: The audio detection module performs noise analysis on the M-frame first target spectral signals.

[0123] Specifically, after determining the first target spectral signals, the audio detection module performs noise analysis on each frame of the first target spectral signals in the frequency domain. If there is noise in the first target spectral signals, the audio detection module determines that the first audio input signal is an audio signal generating noise. Otherwise, the audio detection module determines that the first input audio signal is an audio signal not generating noise.

[0124] The specific process of the audio detection module performing noise analysis on the first target spectral signal is as follows: The audio detection module determines in the frequency domain whether there is a frequency point corresponding to a power value greater than or equal to the first threshold. If the determination is yes, the audio detection module determines that the first target spectral signal of this frame is a spectral signal generating noise, the power value corresponding to this frequency point is the spectral energy of the noise, and the noise component is separated. If the determination is no, the audio detection module determines that the first target spectral signal of this frame is a spectral signal not generating noise, and then performs noise analysis on the first target spectral signal of the next frame.

[0125] Among them, the first threshold can be obtained based on historical experience values or experimental data, and the embodiments of the present application do not make limitations.

[0126] In some embodiments, the electronic device can select a large number of audio signals not generating noise as training samples based on characteristics such as the frequency domain characteristics, timbre, and tonality value of the signal, and train the audio detection module to calculate the first threshold, so that when the electronic device inputs an audio input signal, the audio detection module can generate the first threshold according to the frequency domain characteristics, timbre, tonality value, etc. of the audio input signal, and determine whether the target spectral signal generates noise based on the first threshold.

[0127] In this way, the first threshold calculated by the electronic device is more accurate and comprehensive, enabling the electronic device to perform noise analysis on the target spectral signal more accurately.

[0128] In some other embodiments, the electronic device can select a large number of audio signals with noise and marked with the first threshold as training samples based on characteristics such as the frequency domain characteristics, timbre, and tonality value of the audio signal, and train the audio detection module to calculate the first threshold, so that when the electronic device inputs an audio input signal, the audio detection module can generate the first threshold according to the frequency domain characteristics, timbre, tonality value, etc. of the audio input signal, and determine whether the target spectral signal generates noise based on the first threshold.

[0129] In this way, the first threshold calculated by the electronic device is more accurate and comprehensive, enabling the electronic device to perform noise analysis on the target spectral signal more accurately.

[0130] Step S406: The audio detection module outputs the first detection information.

[0131] Specifically, in the case where the audio detection module determines that the first audio input signal is an audio signal generating noise, the first detection information is used to indicate that the first audio input signal is an audio signal generating noise, and the first detection information includes: the time point when the oscillation occurs, the spectral energy of the harmonic part corresponding to the first target spectral signal, and the spectral energy corresponding to the noise part. Exemplarily, the first detection information can be the above Figure 3EThe indication information in the detection report interface 34 in the embodiment.

[0132] In the case where the audio detection module determines that the first audio input signal is an audio signal without generating noise, the first detection information is used to indicate that the first audio input signal is an audio signal without generating noise, and the first detection information may include: the time point when the oscillation occurs, and the spectral energy of the harmonic part corresponding to the first target spectral signal.

[0133] In the embodiment of the present application, the audio detection module detects the starting oscillation time range of the first audio input signal, selects M frames of audio signals related to the starting oscillation process, and performs noise analysis on the spectral signals corresponding to these M frames of audio signals, so as to determine whether the audio input signal is an audio signal generating noise. Since audio signals are prone to generate noise during the starting oscillation process, therefore, it is only necessary to detect whether there is spectral noise in the frequency domain signal within the starting oscillation time range of the audio signal to determine whether the audio input signal generates noise, without having to detect the entire audio input signal, thus saving a large amount of computing resources of the audio detection module. Through the method for detecting audio signal noise in the above embodiment, while saving the computing resources of the audio detection module, the accuracy of detecting audio signal noise by the audio detection module is improved.

[0134] The above Figure 4 The embodiment introduces a flowchart of a method for detecting audio signal noise. Next, in combination with Figures 10A - 10I , another scenario for detecting audio signal noise is introduced.

[0135] As Figure 10A shown, the main interface of the electronic device 100 includes a plurality of application icons. When the electronic device 100 detects an input operation (for example, double-click) on the audio detection icon 40, the electronic device 100 starts the audio detection software and displays an audio detection interface 41 as Figure 10B shown.

[0136] As Figure 10B shown, the audio detection interface 41 includes an audio collection control 411. When the electronic device 100 detects an input operation (for example, single-click) on the audio collection control 411, in response to this input operation, the electronic device 100 turns on the microphone and displays an audio detection interface 42 as Figure 3C shown.

[0137] As Figure 10C shown, the audio detection interface 42 is used to prompt the user that the microphone of the electronic device has been turned on and prompt the user whether to collect audio. When the electronic device 100 detects an input operation (for example, single-click) on the "Yes" control 421, in response to this input operation, the electronic device 100 starts to collect audio and displays as Figure 10DThe audio detection interface 43 shown

[0138] As Figure 10D shown, the microphone of the electronic device 100 is collecting the audio signal of Song 1 played by the speaker of the electronic device 200. At the same time, the audio detection interface 43 displayed by the electronic device 100 is used to prompt the user that the electronic device is currently collecting audio. After the audio collection of the electronic device 100 is completed, the displayed audio detection interface is as Figure 10E shown, the audio detection interface 44

[0139] As Figure 10E shown, the audio detection interface 44 includes a prompt message for prompting the user that the audio collection is completed and whether to detect whether there is noise in the collected audio. When the electronic device 100 detects an input operation (for example, a click) on the "Yes" control 441, in response to this operation, the electronic device 100 displays the audio detection interface as Figure 10F shown, the audio detection interface 45

[0140] As Figure 10F shown, the audio detection interface 45 includes a prompt message for prompting the user to select an audio detection method. The audio detection interface 45 includes a "Single Audio Detection" control and an "Audio Comparison Detection" control 451. When the electronic device 100 detects an input operation (for example, a click) on the "Single Audio Detection" control, in response to this input operation, the electronic device 100 performs a noise detection on the audio signal of Song 1 collected by its microphone. When the electronic device 100 detects an input operation (for example, a click) on the "Audio Comparison Detection" control 451, in response to this input operation, the electronic device 100 displays the audio selection interface as Figure 10G shown, the audio selection interface 46

[0141] As Figure 10G shown, the audio selection interface 46 is used to add the original audio. The audio selection interface 46 includes multiple local song information ( Figure 10G only Song 1, Song 2, Song 3, and Song 4 are listed), and the duration of each song (such as Song 1 is 50s, Song 2 is 2 minutes and 50 seconds, Song 3 is 10 seconds, and Song 4 is 1 minute and 25 seconds). When the electronic device 100 detects that the user checks Song 1 and detects a click operation on the OK control 461, the electronic device 100 displays the detection interface as Figure 10H shown, the detection interface 47

[0142] As Figure 10H shown, when the detection interface 47 is used to display whether the current audio detection software is detecting whether there is noise in Song 1. After the audio detection software finishes the detection, the electronic device displays the detection report interface as Figure 10I shown, the detection report interface 48

[0143] As Figure 10IAs shown, the detection report interface 48 includes audio information indicating that Song 1 has noise, the time point when Song 1 has noise, and the power value / energy value corresponding to Song 1 at that time point.

[0144] Next, in combination with Figure 11 , the specific process of noise detection for the audio signal in the above Figures 10A - 10I will be described. Please refer to Figure 11 , Figure 11 which is a flowchart of another method for detecting noise in an audio signal provided by an embodiment of the present application. The specific process is as follows:

[0145] Step S1101: The microphone of the electronic device collects a first sound signal and converts the first sound signal into a second audio input signal.

[0146] Specifically, the microphone converts the collected first sound signal into an electrical signal, and the electrical signal is an analog signal in the time domain. The microphone samples the analog signal in the time domain to obtain a second input audio signal. Exemplarily, the first sound signal can be the sound signal of Song 1 played by the speaker of the electronic device 200 in the above Figure 10D . The sound signal is a signal obtained by converting the audio signal (electrical signal) of Song 1 by the speaker of the electronic device 200, and the audio signal of Song 1 is the first audio input signal.

[0147] Step S1102: The audio detection module obtains the first audio input signal and the second audio input signal.

[0148] Specifically, the first audio input signal is an audio signal that has not been processed by the speaker and the first audio input signal is the original signal of the second audio input signal. Exemplarily, the first audio input signal can be the audio signal of Song 1 in the above Figure 10G , and the Song 1 stored in the electronic device 100 is the same song as the Song 1 played back by the speaker of the electronic device 200 in the above Figure 10D . Since the Song 1 in the electronic device 100 has not been processed and played by the speaker, the audio signal of Song 1 in the electronic device 100 is the original signal of the audio signal of Song 1 collected by the microphone of the electronic device 100.

[0149] Step S1103: The audio detection module performs frame splitting on the first audio input signal to obtain N frames of first original audio signals.

[0150] Step S1104: The audio detection module performs time-frequency transformation on the N frames of first original audio signals to obtain N frames of first spectrum signals.

[0151] Step S1105: The audio detection module detects the first starting position of the first audio input signal in the time domain based on the N frames of first spectrum signals, and determines M frames of first target spectrum signals from the N frames of first spectrum signals based on the first starting position.

[0152] For the relevant descriptions of steps S1104 to S1105, please refer to the relevant descriptions of steps S402 to S404, which will not be elaborated here.

[0153] Step S1106: The audio detection module determines the first signal in each frame of the first target spectrum signals based on the second threshold.

[0154] Specifically, in each frame of the first target spectrum signals, the audio detection module determines the signal whose power value corresponding to the frequency point is greater than the second threshold as the first signal, so as to obtain M frames of first signals.

[0155] Among them, the second threshold can be obtained from historical data, or from empirical values, or from experimental data. The embodiments of the present application do not limit this.

[0156] Step S1107: The audio detection module performs frame division processing on the second audio input signal to obtain N frames of second original audio signals.

[0157] Step S1108: The audio detection module performs time-frequency transformation on the N frames of second original audio signals to obtain N frames of second spectrum signals.

[0158] For the relevant descriptions of steps S1107 to S1108, please refer to the relevant descriptions of steps S402 to S403, which will not be elaborated here.

[0159] Step S1109: The audio detection module detects the second starting position of the second audio input signal in the time domain based on the N frames of second spectrum signals, and determines M frames of second target spectrum signals from the N frames of second spectrum signals based on the second starting position.

[0160] Specifically, the second starting position is the starting time range of the second audio input signal in the time domain. The specific process of the audio detection module performing starting position detection on the N frames of second spectrum signals and determining M frames of second target spectrum signals from the N frames of second spectrum signals is as Figure 12 shown:

[0161] S1201. The audio detection module performs power weighting on each frame of the second spectrum signals in the frequency domain to obtain the second spectrum signals after power weighting.

[0162] S1202. The audio detection module calculates the power average value of each frame of the second spectrum signals after power weighting.

[0163] S1203. The audio detection module performs a first-order difference calculation based on the average power value of each frame of the second spectrum signal to obtain the second difference value of each frame of the second spectrum signal.

[0164] S1204. The audio detection module performs frame interpolation based on the second difference value of each frame of the second original audio signal to obtain a third audio signal.

[0165] S1205. The audio detection module downsamples the third audio signal to obtain a fourth audio signal.

[0166] S1206. The audio detection module calculates the second startup dynamic threshold of the fourth audio signal.

[0167] For S1201 to S1206, please refer to the above S501 to S506 and will not be elaborated here.

[0168] Specifically, the audio detection module can calculate the second startup dynamic threshold according to formula (6), and formula (6) is as follows:

[0169] Thresh(m)″=Delta·Median(Y(m)″) (6)

[0170] Where Delta is the first constant coefficient, and the Delta can be obtained based on historical data, or based on empirical values, or based on experimental data. The embodiments of the present application do not make limitations. Thresh(m)″ is the second startup dynamic threshold of the fourth audio signal at the mth sampling point in the time domain, and Median(Y(m)″) is used to represent the median calculation of the power values of the mth sampling point of the fourth audio signal and one or more adjacent sampling points.

[0171] S1207. The audio detection module calculates the second power difference between the fourth audio signal and the second startup dynamic threshold.

[0172] Specifically, the audio detection module can obtain the power deviation value of each sampling point of the fourth audio signal according to formula (7), and formula (7) is as follows:

[0173] Diff(m)″=Y(m)″-Thresh(m)″ (7)

[0174] Where Diff(m)″ is the second power difference of the fourth audio signal at the mth sampling point in the time domain.

[0175] S1208. The audio detection module determines M frames of second target spectrum signals from the N frames of second spectrum signals based on the second power difference.

[0176] For S1208, please refer to the above S508 and will not be elaborated here.

[0177] Step S1110: The audio detection module determines a second signal in each frame of the second target spectral signal based on a second threshold.

[0178] Specifically, in each frame of the second target spectral signal, the audio detection module determines, as the second signal, a signal whose power value corresponding to a frequency point is greater than the second threshold, thereby obtaining M frames of second signals.

[0179] Step S1111: The audio detection module determines whether there is a power value of the second signal in the i-th frame that is greater than the power value of the first signal in the i-th frame.

[0180] Step S1112: If the determination is yes, the audio detection module determines that the first audio input signal is an audio signal generating noise, and outputs second detection information.

[0181] Wherein, the second detection information is used to indicate that the first audio input signal is an audio signal generating noise, and the second detection information includes: the time point when oscillation starts, the spectral energy of the harmonic part corresponding to the first target spectral signal and the spectral energy of the noise part, the spectral energy of the harmonic part corresponding to the second target spectral signal and the spectral energy of the noise part, and other information.

[0182] Step S1113: If the determination is no, the audio detection module determines that the first audio input signal is an audio signal not generating noise, and outputs third detection information.

[0183] Wherein, the third detection information is used to indicate that the first audio input signal is an audio signal not generating noise, and the third detection information may include: the time point when oscillation starts, the spectral energy of the harmonic part corresponding to the first target spectral signal, the spectral energy of the harmonic part corresponding to the second target spectral signal, and other information. Exemplarily, the first detection information may be the indication information in the detection report interface 48 in the above Figure 10I embodiment.

[0184] In the embodiment of the present application, the second audio input signal may be the first audio input signal after being played back through the electronic device speaker. By comparing the oscillation part of the second audio input signal and the power values of the frequency points whose power values are greater than the second threshold with the oscillation part of the first audio input signal and the power values of the frequency points whose power values are greater than the second threshold, it can be determined whether the first audio input signal has undergone non-linear distortion during the playback process through the speaker, thereby generating noise. By performing audio noise detection through the above method, the detection result is more accurate.

[0185] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk).

[0186] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware with a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: various media such as ROM or random access memory RAM, magnetic disk, or optical disc that can store program codes.

[0187] In summary, the above are only embodiments of the technical solutions of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made based on the disclosure of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for detecting noise in an audio signal, characterized in that, Including: Obtain N frames of first spectral signals; the N frames of first spectral signals are N frames of signals obtained by performing time-frequency transformation on a first audio input signal; Detect a first starting position of the first audio input signal in the time domain based on the N frames of first spectral signals; Determine M frames of first target spectral signals from the N frames of first spectral signals based on the first starting position; Judge whether the first audio input signal is an audio signal generating noise based on the power values of the first target spectral signals; The detecting the first starting position of the first audio input signal in the time domain based on the N frames of first spectral signals specifically includes: Calculate a first difference value of each frame of first spectral signals according to the average power of each frame of first spectral signals; Perform inter-frame interpolation on the first difference values of each frame of first spectral signals to obtain a first audio signal; Downsample the first audio signal to obtain a second audio signal; Calculate a first starting dynamic threshold of the second audio signal; Determine the first starting position based on the second audio signal and the first starting dynamic threshold.

2. The method according to claim 1, wherein The judging whether the first audio input signal is an audio signal generating noise based on the power values of the first target spectral signals specifically includes: Judge whether there is a power value of a first target spectral signal greater than a first threshold; If so, determine that the first audio input signal is an audio signal generating noise; If not, determine that the first audio input signal is an audio signal not generating noise.

3. The method according to claim 1, characterized in that, The calculating the first starting dynamic threshold of the second audio signal specifically includes: According to the formula = Delta Median ( ), calculate the first vibration state threshold value Among them, the is the first startup dynamic threshold of the m -th sampling point of the second audio signal in the time domain, and the is the power value of the m -th sampling point of the second audio signal in the time domain, and the Delta is the first constant coefficient.

4. The method according to claim 3, wherein The determining the first starting position based on the second audio signal and the first starting dynamic threshold specifically includes: According to the formula = - , calculate the first power difference of the second audio signal; the is the first power difference of the second audio signal at the m th sampling point in the time domain, the is the power value of the second audio signal at the m th sampling point in the time domain, and the is the first starting vibration state threshold of the second audio signal at the m th sampling point in the time domain; Determine a maximum value of a first power difference; Obtain the starting oscillation time of the second audio signal based on the maximum value ; Based on the above-mentioned calculate the first starting position ( - ~ + ), where the is the first starting time length, and the is the second starting time length.

5. The method according to any one of claims 1-4, characterized in that, The first target spectral signals are first spectral signals having sampling points within the first starting position.

6. A method for detecting noise in an audio signal, characterized in that, Including: Obtain N frames of first spectral signals; The N frames of first spectral signals are N frames of signals obtained by performing time-frequency transformation on a first audio input signal; Detect a first starting position of the first audio input signal in the time domain based on the N frames of first spectral signals; Determine M frames of first target spectral signals from the N frames of first spectral signals based on the first starting position; Judge whether the first audio input signal is an audio signal generating noise based on the power values of the first target spectral signals; The method further includes: Obtain N frames of second spectral signals; the N frames of second spectral signals are N frames of signals obtained by performing time-frequency transformation on a second audio input signal, the first audio input signal being the original signal of the second audio input signal; the second audio input signal being the first audio input signal after being played back through a speaker; Detect a second starting position of the second audio input signal in the time domain based on the N frames of second spectral signals; Determine M frames of second target spectral signals from the N frames of second spectral signals based on the second starting position; The judging whether the first audio input signal is an audio signal generating noise based on the power values of the first target spectral signals specifically includes: Determine whether there is a i power value of the second signal in the i frame that is greater than the power value of the first signal in the frame, where the second signal is a signal in the second target spectrum signal with a power value greater than a second threshold, and the first signal is a signal in the first target spectrum signal with a power value greater than the second threshold; If it is yes, determine that the first audio input signal is the audio signal generating noise; If it is no, determine that the first audio input signal is not the audio signal generating noise.

7. The method according to claim 6, wherein The detecting the first starting position of the first audio input signal in the time domain based on the N-frame first spectrum signals specifically includes: Calculating a first difference value of each frame of the first spectrum signals according to the power average value of each frame of the first spectrum signals; Performing inter-frame interpolation on the first difference values of each frame of the first spectrum signals to obtain a first audio signal; Downsampling the first audio signal to obtain a second audio signal; Calculating a first starting dynamic threshold of the second audio signal; Determining the first starting position based on the second audio signal and the first starting dynamic threshold.

8. The method according to claim 7, wherein The calculating the first starting dynamic threshold of the second audio signal specifically includes: According to the formula = Delta Median ( ), calculate the first vibration state threshold Among them, the is the first starting vibration state threshold of the m -th sampling point of the second audio signal in the time domain, and the is the power value of the second audio signal at the m -th sampling point in the time domain, and the Delta is the first constant coefficient.

9. The method according to claim 7, wherein The determining the first starting position based on the second audio signal and the first starting dynamic threshold specifically includes: According to the formula = - , calculate the first power difference of the second audio signal; the is the first power difference of the second audio signal at the m th sampling point in the time domain, the is the power value of the second audio signal at the m th sampling point in the time domain, and the is the first startup dynamic threshold of the second audio signal at the m th sampling point in the time domain; Determining a maximum value of a first power difference; Obtain the starting oscillation occurrence time of the second audio signal based on the maximum value ; Based on the above-mentioned calculate the first starting position ( - ~ + ), where the is the first starting time length, and the is the second starting time length.

10. The method according to any one of claims 6-9, characterized in that, The first target spectrum signal is the first spectrum signal having sampling points within the first starting position.

11. The method according to claim 6, wherein The detecting the second starting position of the second audio input signal in the time domain based on the N-frame second spectrum signals specifically includes: Calculating a second difference value of each frame of the second spectrum signals according to the power average value of each frame of the second spectrum signals; Performing inter-frame interpolation on the second difference values of each frame of the second spectrum signals to obtain a third audio signal; Downsampling the third audio signal to obtain a fourth audio signal; Calculating a second starting dynamic threshold of the fourth audio signal; Determining the second starting position based on the fourth audio signal and the second starting dynamic threshold.

12. The method according to claim 11, wherein, The calculating the second starting dynamic threshold of the fourth audio signal specifically includes: According to the formula = Delta Median ( ), calculate the second starting vibration state threshold value Among them, the is the second startup dynamic threshold of the m th sampling point of the fourth audio signal in the time domain. The is the power value of the m th sampling point of the fourth audio signal in the time domain. The Delta is the first constant coefficient.

13. The method according to claim 11, wherein The determining the second starting position based on the fourth audio signal and the second starting dynamic threshold specifically includes: According to the formula = - , calculate the second power difference of the fourth audio signal; the is the second power difference of the fourth audio signal at the m th sampling point in the time domain, the is the power value of the fourth audio signal at the m th sampling point in the time domain, and the is the second starting vibration state threshold of the fourth audio signal at the m th sampling point in the time domain; Determining a maximum value of a second power difference; Obtain the starting oscillation time of the fourth audio signal based on the maximum value ; Based on the above-mentioned Calculate the second starting position ( - ~ + ), where the is the second starting time length, and the is the second starting time length.

14. The method according to any one of claims 6-9 and 11-13, characterized in that, The second target spectrum signal is the second spectrum signal having sampling points within the second starting position.

15. An electronic device, characterized in that, Including: A memory, a processor and a touch screen; wherein: The touch screen is used for displaying content; The memory is used for storing a computer program, and the computer program includes program instructions; The processor is used for calling the program instructions to enable the electronic device to execute the method according to any one of claims 1-14.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-14 is implemented.

17. A computer program product comprising instructions, characterized in that, When the computer program product runs on an electronic device, the electronic device is enabled to execute the method according to any one of claims 1-14.

Citation Information

Patent Citations

  • Noise detection method and device and electronic equipment

    CN112163117A

  • Noise detection method and terminal

    CN112908347A