Audio processing method and apparatus, electronic device, and storage medium

By converting the audio signal to the frequency domain and using the difference between the filter and the reference value to determine the noise floor information, the problem of inaccurate noise floor measurement in the prior art is solved, and accurate measurement of the noise floor of audio devices is achieved.

CN119851683BActive Publication Date: 2026-05-08BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2023-10-18
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing technologies, when measuring the noise floor of audio devices through audio compression, it is impossible to accurately determine the specific magnitude of the noise floor current sound, and the reference value is not set precisely enough, resulting in inaccurate test results.

Method used

The audio signal is converted to the frequency domain, and the frequency domain information of a specific frequency band is preserved by a filter to determine the reference value. The noise floor information is then accurately determined based on the difference between the amplitude value of the frequency component and the reference value.

Benefits of technology

It enables accurate measurement of the noise floor of audio devices, reduces the influence of subjective human assessment, and ensures the accuracy of test results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851683B_ABST
    Figure CN119851683B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio processing method and device, electronic equipment and storage medium, the method comprising: converting an audio signal of a target audio device collected to a frequency domain to obtain first frequency domain information of the audio signal; wherein the first frequency domain information comprises amplitude values of different frequency components of the audio signal; determining a reference value according to the amplitude values; determining noise floor information of the audio signal according to differences between the amplitude values of different frequency components and the reference value; and the noise floor information is used to reflect a noise floor processing capability of the target audio device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio detection technology, and in particular to an audio processing method and apparatus, electronic device and storage medium. Background Technology

[0002] Existing technologies measure audio device noise floor through audio compression, which only measures whether the audio device's noise floor current exceeds a certain benchmark. The presence of noise floor is determined when a signal exceeds the benchmark. However, this method only confirms the existence of noise floor, not its specific magnitude. Furthermore, the benchmark in this method is manually set and lacks precision; when determining whether a signal exceeds the benchmark, it is prone to missing detections due to excessively low noise floor values, thus affecting the test results. Summary of the Invention

[0003] This disclosure provides an audio processing method and apparatus, an electronic device, and a storage medium.

[0004] According to a first aspect of the present disclosure, an audio processing method is provided, comprising:

[0005] Based on the audio signal of the target audio device being collected and converted to the frequency domain, the first frequency domain information of the audio signal is obtained; wherein, the first frequency domain information includes: the amplitude values ​​of different frequency components of the audio signal;

[0006] Based on the amplitude value, determine the reference value;

[0007] The noise floor information of the audio signal is determined based on the difference between the amplitude value of different frequency components and the reference value; the noise floor information is used to reflect the noise floor processing capability of the target audio device.

[0008] Based on the above scheme, the audio signal from the acquired target audio device is converted to the frequency domain to obtain the first frequency domain information of the audio signal, including:

[0009] The audio signal is converted from the time domain to the frequency domain to generate second frequency domain information;

[0010] The third frequency domain information is obtained by retaining the frequency domain information of the first frequency band in the second frequency domain information through the first filter;

[0011] The first frequency domain information is obtained by retaining the frequency domain information of the second frequency band in the third frequency domain information through the second filter; the second frequency band is a sub-band of the first frequency band.

[0012] Based on the above scheme, the step of converting the audio signal from the time domain to the frequency domain to generate second frequency domain information includes:

[0013] The audio signal is divided into frames in the time domain to obtain multiple audio frames;

[0014] Execute the window function on the audio frame;

[0015] The audio frames after executing the window function are converted to the frequency domain to generate second frequency domain information.

[0016] Based on the above scheme, determining the reference value according to the amplitude value includes:

[0017] Determine the confidence level;

[0018] Using the confidence level as a probability, the amplitude value in the first frequency domain information is estimated in intervals to determine the range of the amplitude value;

[0019] The range of amplitude values ​​is defined as the confidence interval;

[0020] The baseline value is determined based on the confidence interval.

[0021] Based on the above scheme, the noise floor information includes at least:

[0022] Noise floor value in decibels, and / or, noise floor frequency range.

[0023] Based on the above scheme, determining the noise floor information of the audio signal according to the difference between the amplitude values ​​of different frequency components and the reference value includes:

[0024] Determine the difference between the amplitude value of different frequency components in the first frequency domain information and the reference value;

[0025] When the difference exceeds a preset threshold, the difference is determined as the noise floor value in decibels in the audio signal;

[0026] And / or,

[0027] When the difference exceeds a preset threshold, the frequency component corresponding to the amplitude value corresponding to the difference is determined as the noise floor frequency range.

[0028] According to a second aspect of the present disclosure, an audio processing apparatus is provided, the apparatus comprising:

[0029] The first acquisition module is used to convert the acquired audio signal from the target audio device to the frequency domain to obtain the first frequency domain information of the audio signal; wherein, the first frequency domain information includes: the amplitude values ​​of different frequency components of the audio signal;

[0030] The first determining module is used to determine the reference value based on the amplitude value;

[0031] The second determining module is used to determine the background noise information of the audio signal based on the difference between the amplitude value of different frequency components and the reference value; the background noise information is used to reflect the background noise processing capability of the target audio device.

[0032] Based on the above scheme, the first acquisition module is specifically used for:

[0033] The audio signal is converted from the time domain to the frequency domain to generate second frequency domain information;

[0034] The third frequency domain information is obtained by retaining the frequency domain information of the first frequency band in the second frequency domain information through the first filter;

[0035] The first frequency domain information is obtained by retaining the frequency domain information of the second frequency band in the third frequency domain information through the second filter; the second frequency band is a sub-band of the first frequency band.

[0036] Based on the above scheme, the first acquisition module is further used for:

[0037] The audio signal is divided into frames in the time domain to obtain multiple audio frames;

[0038] Execute the window function on the audio frame;

[0039] The audio frames after executing the window function are converted to the frequency domain to generate second frequency domain information.

[0040] Based on the above scheme, the first determining module is specifically used for:

[0041] Determine the confidence level;

[0042] Using the confidence level as a probability, the amplitude value in the first frequency domain information is estimated in intervals to determine the range of the amplitude value;

[0043] The range of amplitude values ​​is defined as the confidence interval;

[0044] The baseline value is determined based on the confidence interval.

[0045] Based on the above scheme, the noise floor information includes at least:

[0046] Noise floor value in decibels, and / or, noise floor frequency range.

[0047] Based on the above scheme, the second determining module is specifically used for:

[0048] Determine the difference between the amplitude value of different frequency components in the first frequency domain information and the reference value;

[0049] When the difference exceeds a preset threshold, the difference is determined as the noise floor value in decibels in the audio signal;

[0050] And / or,

[0051] When the difference exceeds a preset threshold, the frequency component corresponding to the amplitude value corresponding to the difference is determined as the noise floor frequency range.

[0052] A third aspect of the present disclosure provides an electronic device, comprising:

[0053] Memory used to store processor-executable instructions;

[0054] The processor is connected to the memory;

[0055] The processor is configured to execute the audio processing method described above.

[0056] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a computer's processor, the computer is able to perform the audio processing method described above.

[0057] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:

[0058] As can be seen from the above embodiments, the technical solution provided by this disclosure converts the collected audio signal of the target audio device into the frequency domain to obtain first frequency domain information including the amplitude values ​​of different frequency components of the audio signal; a reference value is determined by the amplitude value in the first frequency domain information of the audio signal, reducing the problem of inaccurate noise measurement results caused by subjective human evaluation to set the noise floor reference value; at the same time, the noise floor information of the audio signal is determined according to the difference between the amplitude value of different frequency components in the frequency domain information and the reference value. In this way, not only can all the noise floor present in the target audio device be determined, but also the specific magnitude of the noise floor of the target audio device can be accurately determined.

[0059] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0060] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0061] Figure 1 This is a schematic diagram illustrating a conventional noise floor determination method according to an exemplary embodiment;

[0062] Figure 2 This is a flowchart illustrating an audio processing method according to an exemplary embodiment;

[0063] Figure 3 This is a flowchart illustrating another audio processing method according to an exemplary embodiment;

[0064] Figure 4 This is a flowchart illustrating another audio processing method according to an exemplary embodiment;

[0065] Figure 5 This is a schematic diagram illustrating the result of third frequency domain information according to an exemplary embodiment;

[0066] Figure 6 This is a schematic diagram illustrating the result of a first frequency domain information according to an exemplary embodiment;

[0067] Figure 7 This is a schematic diagram illustrating a benchmark value result according to an exemplary embodiment;

[0068] Figure 8 This is a schematic diagram illustrating a noise floor result according to an exemplary embodiment;

[0069] Figure 9 This is a schematic diagram illustrating an application scenario of an audio processing method according to an exemplary embodiment;

[0070] Figure 10 This is a flowchart illustrating an audio processing method according to an exemplary embodiment;

[0071] Figure 11 This is a structural block diagram of an audio processing apparatus according to an exemplary embodiment;

[0072] Figure 12 This is a block diagram illustrating the structural composition of an audio processing apparatus according to an exemplary embodiment. Detailed Implementation

[0073] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses consistent with some aspects of this disclosure as detailed in the appended claims.

[0074] The terminology used in this embodiment of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the embodiments of the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of the invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0075] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of the present invention, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of embodiments of the present invention, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0076] Currently, audio performance testing for True Wireless Stereo (TWS) earbuds is typically based on Audio Pression. According to the Bluetooth transmission protocol of TWS earbuds, audio performance testing is usually divided into call testing (Hands-free Profile, HFP) and Bluetooth Audio Distribution Profile (A2DP).

[0077] In some implementations, audio compression is typically used to measure headphone noise floor, such as... Figure 1 As shown, this method only measures whether the background noise current exceeds a certain baseline. It determines the presence of background noise in the headphones when a signal exceeds the baseline. However, this method can only determine if the background noise exceeds a certain range, not the specific magnitude of the background noise current. This method of determining whether the signal exceeds the baseline is prone to missing detections due to excessively low background noise values, affecting the test results and consequently the quality assessment of the audio equipment. Furthermore, the baseline in the above method is usually manually set and is not precise.

[0078] like Figures 2 to 8 As shown in the embodiments of this disclosure, an audio processing method is provided, the method comprising:

[0079] 101: Based on the audio signal of the target audio device being collected, the audio signal is converted to the frequency domain to obtain the first frequency domain information of the audio signal; wherein, the first frequency domain information includes: the amplitude values ​​of different frequency components of the audio signal;

[0080] 102: Determine the reference value based on the amplitude value;

[0081] 103: Determine the noise floor information of the audio signal based on the difference between the amplitude value of different frequency components and the reference value; the noise floor information is used to reflect the noise floor processing capability of the target audio device.

[0082] For example, the target audio device can be headphones, such as Bluetooth headphones or wired headphones; for another example, the target audio device can also be a speaker, amplifier, or microphone; in this embodiment, the target audio device is headphones.

[0083] Noise floor generally refers to all interference in a system that is generated, inspected, measured, or recorded, regardless of the presence or absence of a signal. In industrial or environmental noise measurement, it refers to the ambient noise outside the noise source being measured. For example, the noise floor can refer to the inherent electronic or electromagnetic noise in the target audio device, and the magnitude of the noise floor in decibels is one of the standards for measuring the quality of the target audio device.

[0084] In this embodiment, the larger the background noise value in the background noise information, the worse the background noise processing capability of the target audio device; the smaller the background noise value, the better the background noise processing capability of the target audio device.

[0085] The technical solution provided in this disclosure converts the collected audio signal of the target audio device into the frequency domain to obtain first frequency domain information including the amplitude values ​​of different frequency components of the audio signal. A reference value is determined by the amplitude values ​​in the first frequency domain information, reducing the problem of inaccurate noise measurement results caused by subjective human evaluation to set the noise floor reference value. At the same time, the noise floor information of the audio signal is determined based on the difference between the amplitude values ​​of different frequency components in the frequency domain information and the reference value. In this way, not only can all the noise floor present in the target audio device be determined, but also the specific magnitude of the noise floor of the target audio device can be accurately determined.

[0086] For example, such as Figure 3 As shown, step 101 above includes:

[0087] 201: Convert the audio signal from the time domain to the frequency domain to generate second frequency domain information;

[0088] 202: The third frequency domain information is obtained by retaining the frequency domain information of the first frequency band in the second frequency domain information through the first filter;

[0089] 203: The first frequency domain information is obtained by retaining the frequency domain information of the second frequency band in the third frequency domain information through the second filter; the second frequency band is a sub-band of the first frequency band.

[0090] In some embodiments, in step 201 above, the audio signal is a time-domain signal, and any method that can convert the audio signal from the time domain to the frequency domain is within the protection scope of this embodiment.

[0091] For example, the first filter can be a bandpass filter, which retains the frequency domain information of the first frequency band in the second frequency domain information to obtain the third frequency domain information. Figure 5 This is a schematic diagram illustrating the result of a third frequency domain information in this embodiment.

[0092] For example, the first frequency band can be any frequency range from 0Hz to 48000Hz; preferably, in this embodiment, the first frequency band can be from 700Hz to 16000Hz; here, the first frequency band is the sound frequency range that can be heard by the human ear; the third frequency domain information is obtained through the first filter, and the third frequency domain information is the second frequency domain information of the frequency band corresponding to the first frequency band; in this way, the background noise of the target audio device can be better determined when the target audio device plays audio signals in the frequency range that can be heard by the human ear.

[0093] For example, the second filter is a high-pass filter, such as a Butterworth filter, a Gaussian high-pass filter, etc.

[0094] For example, the second frequency band may be less than or equal to the first frequency band; in this embodiment, the second frequency band is equal to the first frequency band, and the second frequency band is also 700Hz to 16000Hz.

[0095] Specifically, the third frequency domain information is passed through a high-pass filter to retain the frequency domain information of the second frequency band in the third frequency domain information to obtain the first frequency domain information; the high-pass filter is used to remove the low-frequency envelope in the third frequency domain information and retain the high-frequency details. Figure 6 This is a schematic diagram illustrating the result of a first frequency domain information in this embodiment.

[0096] It should be noted that before step 203, the third frequency domain information can be passed through a Hilbert filter to extract the frequency domain information corresponding to the high-frequency signal in the third frequency domain information. In this way, the amplitude value in the third frequency domain information obtained by passing through the Hilbert filter can be smoother, thereby improving the quality of the third frequency domain information.

[0097] The high-frequency information extracted from the third frequency domain information is used as the input information for step 203, and step 203 is executed.

[0098] In some embodiments, such as Figure 4 As shown, step 201 above includes:

[0099] 2011: The audio signal is divided into frames in the time domain to obtain multiple audio frames;

[0100] 2012: Execute the window function on the audio frame;

[0101] 2013: Convert the audio frames after executing the window function to the frequency domain to generate second frequency domain information.

[0102] For example, the audio signal is acquired based on a certain sampling rate; here, the sampling rate can be 48000HZ, 16000HZ, 8000HZ or 4000HZ, etc., and the sampling rate can be dynamically adjusted.

[0103] In step 2011 above, the frame length of the frame is determined; for example, the frame length of the frame can be 16ms, 64ms, 120ms or 1000ms, etc.; preferably, the frame length can be 1000ms, i.e. 1s.

[0104] In the time domain, based on the frame length, the acquired audio signal is divided into frames to obtain multiple audio frames with a frame length of 1 second.

[0105] For example, the window length of the window function is less than the frame length; in this embodiment, preferably, the window length can be 500ms, i.e., 0.5s.

[0106] By using a window length of 0.5s to window multiple audio frames with a frame length of 1s and executing the window function, spectral leakage can be reduced.

[0107] Windowing can reduce the error between the framed signal and the original signal, making the framed signal continuous and ensuring that each frame exhibits the characteristics of a periodic function. For example, window functions such as the Hamming window or the Hanning window can be used.

[0108] For example, in step 2013 above, the audio frame after the window function is executed can be converted to the frequency domain by Fourier transform to generate second frequency domain information.

[0109] Here, the windowed audio frames are arranged in frame order using Fourier transform to obtain the second frequency domain information.

[0110] In some embodiments, step 102 above includes:

[0111] Determine the confidence level;

[0112] Using the confidence level as a probability, the amplitude value in the first frequency domain information is estimated in intervals to determine the range of the amplitude value;

[0113] The range of amplitude values ​​is defined as the confidence interval;

[0114] The baseline value is determined based on the confidence interval.

[0115] For example, the confidence level can be 90%, 80%, or 95%. Preferably, in this embodiment, the confidence level is 80%.

[0116] In some embodiments, the confidence level may be a preset confidence level.

[0117] Specifically, with a confidence level of 80%, the amplitude value in the first frequency domain information is estimated by interval estimation, and the determined amplitude value range is represented by [A, B]; the amplitude value range is determined as the confidence interval, that is, the confidence interval is also [A, B].

[0118] Let X represent the benchmark value, then the benchmark value ; Figure 7 This is a schematic diagram illustrating a benchmark value result according to an embodiment of the present disclosure.

[0119] In this embodiment of the disclosure, the noise reference value is determined by the frequency domain information of the audio signal played by the target audio device, which solves the problem that the accuracy of the subsequent determination of the noise floor reference value is affected by the noise reference value being set manually.

[0120] In some embodiments, in step 103 above, the noise floor information includes at least: a noise floor decibel value, and / or a noise floor frequency range.

[0121] In some embodiments, step 103 above includes:

[0122] Determine the difference between the amplitude value of different frequency components in the first frequency domain information and the reference value;

[0123] When the difference exceeds a preset threshold, the difference is determined as the noise floor value in decibels in the audio signal;

[0124] And / or,

[0125] When the difference exceeds a preset threshold, the frequency component corresponding to the amplitude value corresponding to the difference is determined as the noise floor frequency range.

[0126] For example, the preset threshold can be a default value or can be dynamically adjusted according to the application scenario or the noise floor standard value of the target audio device; for example, in one embodiment, the preset threshold can be 5dB.

[0127] Specifically, the difference between the amplitude value of different frequency components in the first frequency domain information and the reference value is represented by Delta.

[0128] The absolute values ​​of the amplitude values ​​of different frequency components in the first frequency domain information are subtracted from the reference value X. When the difference Delta exceeds the preset threshold, the difference Delta is determined as the noise floor value in decibels of the audio signal, and the frequency of the amplitude value corresponding to the difference Delta is determined as the noise floor frequency. In this way, the noise floor value in decibels and the frequency range in which the noise floor is located in the target audio device can be accurately determined. Figure 8 As shown, Figure 8 This is a schematic diagram illustrating a noise floor result in an embodiment of this disclosure.

[0129] In one embodiment, the method further includes: summing the noise floor decibel values ​​in the determined noise floor information to determine the total noise floor decibel value of the target audio device.

[0130] The technical solution provided in this disclosure converts the acquired audio signal of the target audio device into the frequency domain to obtain frequency domain information including the amplitude values ​​of different frequency components of the audio signal. A reference value is determined by the amplitude value in the frequency domain information of the audio signal, reducing the problem of inaccurate noise measurement results caused by subjective human evaluation to set the noise floor reference value. At the same time, the noise floor information of the audio signal is determined based on the difference between the amplitude value of different frequency components in the frequency domain information and the reference value. Thus, not only is it determined whether the target audio device has noise floor, but also the specific magnitude of the noise floor of the target audio device is accurately determined.

[0131] The following specific embodiments further illustrate the audio processing method and its application scenarios provided in this disclosure.

[0132] Figure 9 This is a schematic diagram illustrating an application scenario of the audio processing method provided in this embodiment; the audio processing method is used in... Figure 9 The test environment shown includes a test terminal 11, a sound card 12, a mobile terminal 13, TWS earphones 14, and a low-noise artificial ear 15.

[0133] The TWS earphones 14 are connected to the mobile terminal 13 via Bluetooth. In a relatively quiet environment, the low-noise artificial ear 15 collects the audio signal played by the TWS earphones 14. The audio signal collected by the artificial ear 15 is converted from analog to digital by the sound card 12 to generate a digital signal. The digital signal is then input to the test terminal 11 for calculation and analysis. By ensuring the test environment is quiet, the ambient noise in the collected audio signal can be reduced.

[0134] Figure 10 This is a flowchart illustrating an audio processing method according to an exemplary embodiment; the specific flow of the method may include the following steps:

[0135] Step 1: Acquire the audio signal played by the target audio device;

[0136] Specifically, the audio signal played by TWS earphones is collected through a low-noise artificial ear.

[0137] Step 2: Perform a short-time Fourier transform on the audio signal to generate second frequency domain information;

[0138] Specifically, the audio signal is processed by performing a short-time Fourier transform on the audio signal, converting the audio signal from the time domain to the frequency domain, and generating second frequency domain information.

[0139] Step 2.1: Divide the audio signal into frames in the time domain to obtain multiple audio frames;

[0140] For example, the frame length of the segment can be 16ms, 64ms, 120ms or 1000ms, etc.; preferably, the frame length can be 1000ms, i.e. 1s.

[0141] Specifically, in the time domain, the acquired audio signal is divided into frames based on the frame length to obtain multiple audio frames with a frame length of 1 second.

[0142] Step 2.2: Execute the window function on the audio frame;

[0143] For example, the window length of the window function is less than the frame length; in this embodiment, preferably, the window length can be 500ms, i.e., 0.5s.

[0144] Specifically, the audio signal is overlapped and framed by using a window length smaller than the frame length; multiple audio frames with a frame length of 1s are windowed with a window length of 0.5s and a window function is executed; in this way, the error between the framed signal and the original signal can be reduced, making the framed signal continuous and each frame exhibiting the characteristics of a periodic function.

[0145] Step 2.3: Convert the audio frames after the window function is executed to the frequency domain using Fourier transform to generate second frequency domain information.

[0146] Step 3: Perform bandpass filtering on the second frequency domain information to generate the third frequency domain information.

[0147] Specifically, the second frequency domain information is bandpass filtered by a bandpass filter, and the frequency domain information of the first frequency band in the second frequency domain information is retained to obtain the third frequency domain information;

[0148] For example, the first frequency band is 700Hz to 16000Hz; here, the first frequency band is the range of sound frequencies that can be heard by the human ear; the third frequency domain information is obtained through the first filter, and the third frequency domain information is the second frequency domain information of the frequency band corresponding to the first frequency band; in this way, the background noise of the target audio device can be better determined when the target audio device plays audio signals in the frequency range that can be heard by the human ear.

[0149] Step 4: Extract the frequency domain information corresponding to the high-frequency signal in the third frequency domain information using a Hilbert filter to generate the fourth frequency domain information.

[0150] Here, the Hilbert filter can make the amplitude values ​​in the third frequency domain information smoother, thus improving the quality of the fourth frequency domain information.

[0151] Step 5: Obtain the first frequency domain information by retaining the frequency domain information of the second frequency band in the fourth frequency domain information through a high-pass filter; the second frequency band is a sub-band of the first frequency band.

[0152] For example, the second frequency band may be less than or equal to the first frequency band; in this embodiment, the second frequency band is equal to the first frequency band, and the second frequency band is 700Hz to 16000Hz.

[0153] Here, a high-pass filter is used to remove the low-frequency envelope from the fourth frequency domain information, while retaining the high-frequency details.

[0154] Step 6: Determine the baseline value.

[0155] Step 6.1: Determine the confidence level;

[0156] Here, the confidence level can be a preset confidence level or it can be dynamically adjusted;

[0157] For example, the confidence level can be 90%, 80%, or 95%. Preferably, in this embodiment, the confidence level is 80%.

[0158] Step 6.2: Using the confidence level as a probability, perform interval estimation on the amplitude value in the first frequency domain information to determine the confidence interval;

[0159] Specifically, with a confidence level of 80%, the amplitude value in the first frequency domain information is estimated by interval estimation, and the determined amplitude value range is represented by [A, B]; the amplitude value range is determined as the confidence interval, that is, the confidence interval is also [A, B].

[0160] Step 6.3: Determine the benchmark value based on the confidence interval;

[0161] Specifically, the benchmark value is represented by X, and the average value of the confidence interval is used as the benchmark value. .

[0162] Step 7: Determine the noise floor information;

[0163] Step 7.1: Determine the difference between the amplitude value of different frequency components in the first frequency domain information and the reference value;

[0164] Specifically, the difference between the amplitude values ​​of different frequency components in the first frequency domain information and the reference value X is determined; the difference is represented by Delta.

[0165] Step 7.2: When the difference exceeds a preset threshold, the difference is determined as the noise floor decibel value in the audio signal; and / or, when the difference exceeds a preset threshold, the frequency component corresponding to the amplitude value corresponding to the difference is determined as the noise floor frequency range;

[0166] For example, the preset threshold can be a default value or can be dynamically adjusted according to the application scenario or the noise floor standard value of the target audio device; for example, in this embodiment, the preset threshold can be 5dB.

[0167] Specifically, when the difference Delta exceeds 5 dB, the difference Delta is determined as the noise floor value in the audio signal in decibels, and the frequency of the amplitude value corresponding to the difference Delta is determined as the noise floor frequency. In this way, the noise floor value in the target audio device and its frequency range can be accurately determined.

[0168] Step 8: Determine the total noise floor value in decibels of the target audio device;

[0169] Specifically, the noise floor decibel values ​​in the determined noise floor information are summed to determine the total noise floor decibel value of the target audio device.

[0170] The technical solution provided in this disclosure converts the acquired audio signal of the target audio device into the frequency domain to obtain frequency domain information including the amplitude values ​​of different frequency components of the audio signal. A reference value is determined by the amplitude value in the frequency domain information of the audio signal, reducing the problem of inaccurate noise measurement results caused by subjective human evaluation to set the noise floor reference value. At the same time, the noise floor information of the audio signal is determined based on the difference between the amplitude value of different frequency components in the frequency domain information and the reference value. Thus, not only is it determined whether the target audio device has noise floor, but also the specific magnitude of the noise floor of the target audio device is accurately determined.

[0171] Figure 11 This is a structural block diagram of an audio processing apparatus according to an exemplary embodiment. (Refer to...) Figure 10 The audio processing device 300 may include:

[0172] The first acquisition module 301 is used to convert the acquired audio signal from the target audio device to the frequency domain to obtain the first frequency domain information of the audio signal; wherein, the first frequency domain information includes: the amplitude values ​​of different frequency components of the audio signal;

[0173] The first determining module 302 is used to determine a reference value based on the amplitude value;

[0174] The second determining module 303 is used to determine the background noise information of the audio signal based on the difference between the amplitude value of different frequency components and the reference value; the background noise information is used to reflect the background noise processing capability of the target audio device.

[0175] In some embodiments, the first acquisition module 301 is specifically used for:

[0176] The audio signal is converted from the time domain to the frequency domain to generate second frequency domain information;

[0177] The third frequency domain information is obtained by retaining the frequency domain information of the first frequency band in the second frequency domain information through the first filter;

[0178] The first frequency domain information is obtained by retaining the frequency domain information of the second frequency band in the third frequency domain information through the second filter; the second frequency band is a sub-band of the first frequency band.

[0179] In some embodiments, the first acquisition module 301 further includes:

[0180] The audio signal is divided into frames in the time domain to obtain multiple audio frames;

[0181] Execute the window function on the audio frame;

[0182] The audio frames after executing the window function are converted to the frequency domain to generate second frequency domain information.

[0183] In some embodiments, the first determining module 302 is specifically used for:

[0184] Determine the confidence level;

[0185] Using the confidence level as a probability, the amplitude value in the first frequency domain information is estimated in intervals to determine the range of the amplitude value;

[0186] The range of amplitude values ​​is defined as the confidence interval;

[0187] The baseline value is determined based on the confidence interval.

[0188] In some embodiments, the noise floor information includes at least:

[0189] Noise floor value in decibels, and / or, noise floor frequency range.

[0190] In some embodiments, the second determining module 303 is specifically used for:

[0191] Determine the difference between the amplitude value of different frequency components in the first frequency domain information and the reference value;

[0192] When the difference exceeds a preset threshold, the difference is determined as the noise floor value in decibels in the audio signal;

[0193] And / or,

[0194] When the difference exceeds a preset threshold, the frequency component corresponding to the amplitude value corresponding to the difference is determined as the noise floor frequency range.

[0195] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0196] In an exemplary embodiment, the first acquisition module 301, the first determination module 302, and the second determination module 303 may be implemented by one or more central processing units (CPUs), graphics processing units (GPUs), baseband processors (BPs), application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.

[0197] Figure 12This is a block diagram illustrating an audio processing apparatus 400 according to an exemplary embodiment. For example, apparatus 400 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0198] Reference Figure 12 The device 400 may include one or more of the following components: a processing component 402, a memory 404, a power supply component 406, a multimedia component 408, an audio component 410, an input / output (I / O) interface 412, a sensor component 414, and a communication component 416.

[0199] Processing component 402 typically controls the overall operation of device 400, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 402 may include one or more processors 420 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 402 may include one or more modules to facilitate interaction between processing component 402 and other components. For example, processing component 402 may include a multimedia module to facilitate interaction between multimedia component 408 and processing component 402.

[0200] Memory 404 is configured to store various types of data to support the operation of device 400. Examples of such data include instructions for any application or method operating on device 400, contact data, phonebook data, messages, pictures, videos, etc. Memory 404 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0201] Power supply component 406 provides power to various components of device 400. Power supply component 406 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 400.

[0202] Multimedia component 408 includes a screen that provides an output interface between the device 400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 408 includes a front-facing camera and / or a rear-facing camera. When the device 400 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0203] Audio component 410 is configured to output and / or input audio signals. For example, audio component 410 includes a microphone (MIC) configured to receive external audio signals when device 400 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 404 or transmitted via communication component 416. In some embodiments, audio component 410 also includes a speaker for outputting audio signals.

[0204] I / O interface 412 provides an interface between processing component 402 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0205] Sensor assembly 414 includes one or more sensors for providing status assessments of various aspects of device 400. For example, sensor assembly 414 may detect the on / off state of device 400, the relative positioning of components such as the display and keypad of device 400, changes in the position of device 400 or a component of device 400, the presence or absence of user contact with device 400, the orientation or acceleration / deceleration of device 400, and temperature changes of device 400. Sensor assembly 414 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 414 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 414 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0206] Communication component 416 is configured to facilitate wired or wireless communication between device 400 and other devices. Device 400 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 416 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 416 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0207] In an exemplary embodiment, the apparatus 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0208] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions, which can be executed by a processor 420 of the device 400 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0209] After the instruction is executed by the processor, it can perform the following operations:

[0210] Based on the audio signal of the target audio device being collected and converted to the frequency domain, the first frequency domain information of the audio signal is obtained; wherein, the first frequency domain information includes: the amplitude values ​​of different frequency components of the audio signal;

[0211] Based on the amplitude value, determine the reference value;

[0212] The noise floor information of the audio signal is determined based on the difference between the amplitude value of different frequency components and the reference value; the noise floor information is used to reflect the noise floor processing capability of the target audio device.

[0213] Understandably, the conversion of the audio signal from the acquired target audio device to the frequency domain to obtain the first frequency domain information of the audio signal includes:

[0214] The audio signal is converted from the time domain to the frequency domain to generate second frequency domain information;

[0215] The third frequency domain information is obtained by retaining the frequency domain information of the first frequency band in the second frequency domain information through the first filter;

[0216] The first frequency domain information is obtained by retaining the frequency domain information of the second frequency band in the third frequency domain information through the second filter; the second frequency band is a sub-band of the first frequency band.

[0217] Understandably, the step of converting the audio signal from the time domain to the frequency domain to generate second frequency domain information includes:

[0218] The audio signal is divided into frames in the time domain to obtain multiple audio frames;

[0219] Execute the window function on the audio frame;

[0220] The audio frames after executing the window function are converted to the frequency domain to generate second frequency domain information.

[0221] Understandably, determining the reference value based on the amplitude value includes:

[0222] Determine the confidence level;

[0223] Using the confidence level as a probability, the amplitude value in the first frequency domain information is estimated in intervals to determine the range of the amplitude value;

[0224] The range of amplitude values ​​is defined as the confidence interval;

[0225] The baseline value is determined based on the confidence interval.

[0226] Understandably, the noise floor information includes at least:

[0227] Noise floor value in decibels, and / or, noise floor frequency range.

[0228] Understandably, determining the noise floor information of the audio signal based on the difference between the amplitude values ​​of different frequency components and the reference value includes:

[0229] Determine the difference between the amplitude value of different frequency components in the first frequency domain information and the reference value;

[0230] When the difference exceeds a preset threshold, the difference is determined as the noise floor value in decibels in the audio signal;

[0231] And / or,

[0232] When the difference exceeds a preset threshold, the frequency component corresponding to the amplitude value corresponding to the difference is determined as the noise floor frequency range.

[0233] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims.

[0234] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. An audio processing method, characterized in that, include: Based on the audio signal of the target audio device being collected and converted to the frequency domain, the first frequency domain information of the audio signal is obtained; wherein, the first frequency domain information includes: the amplitude values ​​of different frequency components of the audio signal; Based on the amplitude value, determine the reference value; The step of determining the benchmark value based on the amplitude value includes: determining a confidence level; using the confidence level as a probability, performing interval estimation on the amplitude value in the first frequency domain information to determine the amplitude value range; determining the amplitude value range as a confidence interval; and determining the benchmark value based on the confidence interval. The noise floor information of the audio signal is determined based on the difference between the amplitude value of different frequency components and the reference value; the noise floor information is used to reflect the noise floor processing capability of the target audio device.

2. The method according to claim 1, characterized in that, The audio signal from the target audio device is converted to the frequency domain to obtain the first frequency domain information of the audio signal, including: The audio signal is converted from the time domain to the frequency domain to generate second frequency domain information; The third frequency domain information is obtained by retaining the frequency domain information of the first frequency band in the second frequency domain information through the first filter; The first frequency domain information is obtained by retaining the frequency domain information of the second frequency band in the third frequency domain information through the second filter; the second frequency band is a sub-band of the first frequency band.

3. The method according to claim 2, characterized in that, The step of converting the audio signal from the time domain to the frequency domain to generate second frequency domain information includes: The audio signal is divided into frames in the time domain to obtain multiple audio frames; Execute the window function on the audio frame; The audio frames after executing the window function are converted to the frequency domain to generate second frequency domain information.

4. The method according to claim 1, characterized in that, The noise floor information includes at least: Noise floor value in decibels, and / or, noise floor frequency range.

5. The method according to claim 4, characterized in that, Determining the noise floor information of the audio signal based on the difference between the amplitude values ​​of different frequency components and the reference value includes: Determine the difference between the amplitude value of different frequency components in the first frequency domain information and the reference value; When the difference exceeds a preset threshold, the difference is determined as the noise floor value in decibels in the audio signal; And / or, When the difference exceeds a preset threshold, the frequency component corresponding to the amplitude value corresponding to the difference is determined as the noise floor frequency, so as to determine the noise floor frequency range in which the noise floor is located based on the noise floor frequency.

6. An audio processing apparatus, characterized in that, The device includes: The first acquisition module is used to convert the acquired audio signal from the target audio device to the frequency domain to obtain the first frequency domain information of the audio signal; wherein, the first frequency domain information includes: the amplitude values ​​of different frequency components of the audio signal; A first determining module is configured to determine a benchmark value based on the amplitude value; wherein, determining the benchmark value based on the amplitude value includes: determining a confidence level; using the confidence level as a probability, performing interval estimation on the amplitude value in the first frequency domain information to determine the amplitude value range; determining the amplitude value range as a confidence interval; and determining the benchmark value based on the confidence interval. The second determining module is used to determine the background noise information of the audio signal based on the difference between the amplitude value of different frequency components and the reference value; the background noise information is used to reflect the background noise processing capability of the target audio device.

7. The apparatus according to claim 6, characterized in that, The first acquisition module is specifically used for: The audio signal is converted from the time domain to the frequency domain to generate second frequency domain information; The third frequency domain information is obtained by retaining the frequency domain information of the first frequency band in the second frequency domain information through the first filter; The first frequency domain information is obtained by retaining the frequency domain information of the second frequency band in the third frequency domain information through the second filter; the second frequency band is a sub-band of the first frequency band.

8. The apparatus according to claim 7, characterized in that, The first acquisition module is also specifically used for: The audio signal is divided into frames in the time domain to obtain multiple audio frames; Execute the window function on the audio frame; The audio frames after executing the window function are converted to the frequency domain to generate second frequency domain information.

9. The apparatus according to claim 6, characterized in that, The noise floor information includes at least: Noise floor value in decibels, and / or, noise floor frequency range.

10. The apparatus according to claim 9, characterized in that, The second determining module is specifically used for: Determine the difference between the amplitude value of different frequency components in the first frequency domain information and the reference value; When the difference exceeds a preset threshold, the difference is determined as the noise floor value in decibels in the audio signal; And / or, When the difference exceeds a preset threshold, the frequency component corresponding to the amplitude value corresponding to the difference is determined as the noise floor frequency, so as to determine the noise floor frequency range in which the noise floor is located based on the noise floor frequency.

11. An electronic device, characterized in that, include: Memory used to store processor-executable instructions; The processor is connected to the memory; The processor is configured to perform the audio processing method provided in any one of claims 1 to 5.

12. A non-transitory computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of a computer, the computer is able to perform any one of the audio processing methods as claimed in claims 1 to 5.

Citation Information

Patent Citations

  • Audio signal processing method and device and storage medium

    CN111968662A

  • Estimation method and device of wireless frequency spectrum environment background noise, electronic equipment and storage medium

    CN112003627A