Audio processing method and device and electronic equipment
By using a combination of a high signal-to-noise ratio (SNR) microphone and multiple low SNR microphones in a mobile terminal, the amplitude information of the high SNR microphone is used to repair the recording defects of the low SNR microphone, thus solving the problem of recording saturation distortion in mobile terminals, improving recording quality and user experience, while reducing cost and power consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-01
Smart Images

Figure CN121967587A_ABST
Abstract
Description
Audio processing methods, devices and electronic equipment Technical Field
[0001] This application relates to the field of mobile terminal technology, and more specifically, to an audio processing method, apparatus, and electronic device. Background Technology
[0002] Recording is an essential basic function of smartphones and other terminal devices, and it is widely used in voice calls, sound collection, and other occasions. Therefore, the quality of device recording directly affects the user experience of terminal devices.
[0003] The recording function of terminal devices is mainly achieved through audio acquisition via microphones. However, due to the low acoustic overload point (AOP) of mobile terminal microphones, and the combined effects of microphone pipe and circuit analog gain configuration, the audio recording results of mobile terminal microphones are poor when recording high sound pressure levels. This leads to recording saturation distortion, resulting in a series of problems such as distortion and noise, which in turn reduces the recording quality and efficiency of the terminal device and affects the user experience. Summary of the Invention
[0004] This application proposes an audio processing method, apparatus, and electronic device.
[0005] In a first aspect, this application provides an audio processing method applied to an electronic device, the electronic device including multiple audio acquisition components, the multiple audio acquisition components including a first audio acquisition component and a second audio acquisition component, wherein the signal-to-noise ratio of the second audio acquisition component is higher than that of the first audio acquisition component, the method including: acquiring a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component; if it is detected that the first audio signal does not meet preset usage conditions, acquiring phase information of the first audio signal and amplitude information of the second audio signal; and synthesizing a target audio signal based on the phase information of the first audio signal and the amplitude information of the second audio signal.
[0006] Secondly, this application also provides an audio processing apparatus applied to an electronic device. The electronic device includes multiple audio acquisition components, including a first audio acquisition component and a second audio acquisition component, wherein the signal-to-noise ratio (SNR) of the second audio acquisition component is higher than that of the first audio acquisition component. The apparatus includes: an acquisition unit, a detection unit, and a synthesis unit. The acquisition unit is used to acquire a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component. The detection unit is used to acquire phase information of the first audio signal and amplitude information of the second audio signal if it detects that the first audio signal does not meet preset usage conditions. The synthesis unit is used to synthesize a target audio signal based on the phase information of the first audio signal and the amplitude information of the second audio signal.
[0007] Thirdly, this application also provides an electronic device comprising multiple audio acquisition components, the multiple audio acquisition components including a first audio acquisition component and a second audio acquisition component, wherein the signal-to-noise ratio of the second audio acquisition component is higher than that of the first audio acquisition component; and an audio processing module connected to the first audio acquisition component and the second audio acquisition component respectively, the audio processing module being used to execute the above method.
[0008] The audio processing method, apparatus, and electronic device provided in this application have a second audio acquisition component with a higher signal-to-noise ratio (SNR) than the first audio acquisition component. They acquire a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component. If the first audio signal is detected as not meeting preset usage conditions, they acquire the phase information of the first audio signal and the amplitude information of the second audio signal. Based on the phase information of the first audio signal and the amplitude information of the second audio signal, a target audio signal is synthesized. Therefore, when using a high SNR audio acquisition component, and when using both low and high SNR audio acquisition components to acquire audio, if the audio acquired by the low SNR audio acquisition component does not meet the usage conditions, the target audio signal is synthesized using the phase of the low SNR signal and the amplitude of the high SNR signal, replacing the signal that would normally be output from the low SNR audio acquisition component.
[0009] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 shows a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0012] Figure 2 shows a schematic diagram of the landscape and portrait screen states of an electronic device provided in an embodiment of this application;
[0013] Figure 3 shows a schematic diagram of an audio acquisition component of an electronic device provided in an embodiment of this application;
[0014] Figure 4 shows a flowchart of an audio processing method provided in an embodiment of this application;
[0015] Figure 5 shows a schematic diagram illustrating the positions of the left channel audio acquisition component and the right channel audio acquisition component provided in an embodiment of this application;
[0016] Figure 6 shows a schematic diagram illustrating the positions of the left channel audio acquisition component and the right channel audio acquisition component provided in another embodiment of this application;
[0017] Figure 7 shows a schematic diagram illustrating the positions of the left channel audio acquisition component and the right channel audio acquisition component provided in another embodiment of this application;
[0018] Figure 8 shows a flowchart of an audio processing method provided in another embodiment of this application;
[0019] Figure 9 shows a structural block diagram of an electronic device provided in an embodiment of this application;
[0020] Figure 10 shows a structural block diagram of an electronic device provided in yet another embodiment of this application;
[0021] Figure 11 shows a structural block diagram of an electronic device provided in another embodiment of this application;
[0022] Figure 12 shows a structural block diagram of an electronic device provided in yet another embodiment of this application;
[0023] Figure 13 shows a block diagram of an audio processing apparatus provided in an embodiment of this application;
[0024] Figure 14 illustrates a storage unit according to an embodiment of the present application for storing or carrying program code implementing the method according to an embodiment of the present application. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without inventive effort are within the scope of protection of the present application.
[0026] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0027] The recording function of terminal devices is mainly achieved through audio acquisition via microphones. However, due to the low acoustic overload point (AOP) of mobile terminal microphones, and the combined effects of microphone pipe and circuit analog gain configuration, mobile terminal microphones produce poor audio recording results when recording high sound pressure levels (e.g., the sounds of bass, drum kit, kick drum, and other low-frequency percussion instruments in a concert setting). This leads to recording saturation distortion, resulting in a series of problems such as distortion and noise, which in turn reduces the recording quality and efficiency of the terminal device and affects the user experience.
[0028] Typically, when recording with mobile devices, two microphones are used to record the left and right channels separately to improve the stereo effect. The inventors discovered that current solutions for AOP (Aspect-Oriented Programming) in the left and right channels rely on increasing the number of high signal-to-noise ratio (SNR) microphones, meaning each channel uses a separate high SNR microphone. This results in higher costs, larger motherboard footprint, more layout limitations, and greater complexity.
[0029] Therefore, this application provides an audio processing method that can improve the quality of the acquired audio signal while reducing the number of high signal-to-noise ratio microphones, thus meeting the usage requirements.
[0030] Before introducing this method, we will first introduce the hardware environment of the electronic device of this application.
[0031] It should be noted that the electronic device can be a smartphone, tablet computer, e-reader, or other device capable of running applications. In the embodiments of this application, a smartphone is used as an example to illustrate the various embodiments of this application; however, the electronic device is not limited to smartphones.
[0032] As one implementation method, the screen orientation of an electronic device is defined as having four orientations: a first portrait orientation, a first landscape orientation, a second portrait orientation, and a second landscape orientation. As shown in Figure 1, the screen of the electronic device includes a first short side 201, a second short side 202, a first long side 203, and a second long side 204. It can be seen that the first short side 201 and the second short side 202 are positioned opposite each other, and the first long side 203 and the second long side 204 are positioned opposite each other. In this embodiment, the first short side 201 can be the top edge of the corresponding electronic device, and the second short side 202 can be the bottom edge of the corresponding electronic device. As shown in Figure 1, the top status bar of the electronic device is displayed at the position of the first short side 201.
[0033] In addition, to facilitate the explanation of the above four states, four different directions can be defined, as shown in Figure 2. These four different directions include direction F1, direction F2, direction F3, and direction F4. Figure 2(a) shows the first portrait screen state of the electronic device. Direction F1 refers to the direction from the first long side 203 to the second long side 204, and direction F2 refers to the direction from the first short side 201 to the second short side 202. When the user is facing the screen of the electronic device in the first portrait screen state, with the screen as a reference, the first short side 201 is located above the screen, the second short side 202 is located below the screen, the first long side 203 is located on the left side of the screen, and the second long side 204 is located on the right side of the screen.
[0034] Figure 2(b) shows the first landscape orientation of the electronic device. The direction from the second long side 204 to the first long side 203 is named direction F3. With the user facing the screen of the electronic device in the first landscape orientation, the second long side 204 is located above the screen, the first long side 203 is located below the screen, the first short side 201 is located on the left side of the screen, and the second short side 202 is located on the right side of the screen.
[0035] Figure 2(c) shows the second portrait screen state of the electronic device. The direction from the second short side 202 to the first short side 201 is named direction F4. When the user is facing the screen of the electronic device in the second portrait screen state, with the screen as a reference, the second short side 202 is located above the screen, the first short side 201 is located below the screen, the second long side 204 is located on the left side of the screen, and the first long side 203 is located on the right side of the screen.
[0036] Figure 2(d) shows the second landscape state of the electronic device. With the user facing the screen of the electronic device in the second landscape state, the first long side 203 is located above the screen, the second long side 204 is located below the screen, the second short side 202 is located on the left side of the screen, and the first short side 201 is located on the right side of the screen.
[0037] As can be seen, the left, right, top, and bottom edges of the screen are different in different landscape and portrait modes. The first landscape mode and the second landscape mode are reversed, that is, the second landscape mode is the state after rotating the first landscape mode by 180 degrees. Similarly, the first portrait mode and the second portrait mode are reversed.
[0038] In one implementation, the electronic device is equipped with multiple audio acquisition components, i.e., multiple microphones, which are distributed in different preset areas of the electronic device. These preset areas can be at least one of a top area, a bottom area, and a back area. The top area refers to the top housing area of the electronic device, i.e., the area of the aforementioned first short side 201; the bottom area refers to the bottom housing area of the electronic device, i.e., the area of the aforementioned second short side 202; and the back area refers to the back cover area of the electronic device. For example, a back microphone can be placed near the rear camera to better capture sound during video recording, reduce interference with other components (such as speakers), and minimize obstruction of the back microphone by the phone case.
[0039] Among the multiple audio acquisition components, there are high signal-to-noise ratio (SNR) and low SNR components, meaning there are high SNR microphones and low SNR microphones. SNR measures the ratio between signal strength and background noise intensity, usually expressed in decibels (dB). In microphones, high and low SNR represent the clarity and background interference levels when capturing sound, respectively. High SNR microphones can capture clear, strong audio signals in relatively low background noise environments. Their signal strength is significantly higher than the noise level, for example, above 60 dB. Low SNR microphones capture a lower signal-to-noise ratio, possibly below 30 dB. This means that the background noise is relatively strong, which may affect the clarity of the audio.
[0040] It is understood that, in the embodiments of this application, the number of audio acquisition components and the installation position of each audio acquisition component are not limited. However, in some embodiments, the number of audio acquisition components may be four, and these four audio acquisition components (microphones) are distributed in the aforementioned preset area of the electronic device. For example, two microphones are provided in the top area, one microphone is provided in the bottom area, and one microphone is provided in the back area. Alternatively, one microphone may be provided in the top area, two microphones in the bottom area, and one microphone in the back area.
[0041] For example, as shown in FIG3, the multiple audio acquisition components of the electronic device are a first microphone 301, a second microphone 302, a third microphone 303 and a fourth microphone 304, wherein the first microphone 301 is disposed in the top area of the electronic device, the second microphone 302 and the third microphone 303 are disposed in the bottom area of the electronic device, and the fourth microphone 304 is disposed in the back area of the electronic device.
[0042] It should be noted that the perspective shown in Figure 3 is the user's perspective when looking at the screen of the electronic device. Therefore, the dashed line corresponding to the fourth microphone 304 indicates that the fourth microphone 304 is located in the back area of the electronic device. In addition, the circular patterns for each microphone shown in Figure 3 are used to roughly show the distribution area of each microphone, and do not limit the specific shape, specific position or specific size of the microphone.
[0043] As mentioned above, high signal-to-noise ratio (SNR) microphones can indeed improve sound reception quality and solve the microphone AOP (Aspect-Oriented Programming) problem. However, high SNR microphones are too expensive. Therefore, the electronic device provided in this application aims to solve the AOP problem in the recording process with fewer high SNR microphones.
[0044] In the embodiments of this application, among the multiple audio acquisition components of the electronic device, only one audio acquisition component may be a high signal-to-noise ratio microphone, while the other audio acquisition components are low signal-to-noise ratio microphones.
[0045] As shown in Figure 4, this application embodiment provides an audio processing method applied to the aforementioned electronic device. The electronic device includes multiple audio acquisition components, including a first audio acquisition component and a second audio acquisition component. The signal-to-noise ratio (SNR) of the second audio acquisition component is higher than that of the first audio acquisition component; that is, the second audio acquisition component can be a high SNR microphone, and the first audio acquisition component can be a low SNR microphone. Specifically, the method includes steps S401 to S403.
[0046] S401: Obtain the first audio signal acquired by the first audio acquisition component and the second audio signal acquired by the second audio acquisition component.
[0047] In this embodiment of the application, the first audio signal acquired by the first audio acquisition component is named the first audio signal, and the audio signal acquired by the second audio acquisition component is named the second audio signal. It can be understood that the first and second audio acquisition components acquire audio in the same scenario. That is, during this audio acquisition process, both the first and second audio acquisition components participate in audio recording, and they acquire audio almost synchronously. "Almost synchronously" means that when the electronic device starts recording, it synchronously triggers each audio acquisition component participating in the audio acquisition process to perform audio acquisition operations. However, due to possible delays in instructions or other reasons, there may be slight differences in the start times of audio acquisition by each audio acquisition component. These differences are negligible, and therefore, the first and second audio acquisition components can be considered to be acquiring audio synchronously.
[0048] In one implementation, the first audio acquisition component is a low signal-to-noise ratio (SNR) microphone, and the second audio acquisition component is a high SNR microphone. That is, the SNR of the second audio acquisition component is higher than that of the first audio acquisition component. The first audio acquisition component can be an audio acquisition component used to record the audio signal corresponding to the left or right channel during this audio acquisition process, and the second audio acquisition component can also be an audio acquisition component used to record the audio signal corresponding to the left or right channel. Of course, the channels recorded by the second audio acquisition component and the first audio acquisition component are different. For example, if the first audio acquisition component is used to record the audio signal corresponding to the left channel, then the second audio acquisition component is used to record the audio signal corresponding to the right channel, or if the first audio acquisition component is used to record the audio signal corresponding to the right channel, then the second audio acquisition component is used to record the audio signal corresponding to the left channel.
[0049] Alternatively, the second audio acquisition component may not be used to record left or right channel audio signals, but rather serves as a reference signal so that the audio signals used for recording both left and right channels can use this reference signal to correct defects in the recorded audio signals. In this case, if the second audio acquisition component is not used to record left or right channel audio signals, the audio processing method provided in this application will be executed for both the audio components used to record left channel signals and the audio components used to record right channel signals. That is, there are two first audio acquisition components: a left channel audio acquisition component and a right channel audio acquisition component, both of which serve as first audio acquisition components to execute the audio processing method of this application.
[0050] To describe the first and second audio acquisition components in the embodiments of this application, the selection method of the left and right channel audio acquisition components applied in this audio acquisition process is first explained. If the screen orientation is landscape, two audio acquisition components are sequentially determined along a first direction of the electronic device, serving as the left and right channel audio acquisition components respectively, wherein the first direction is the direction between the top and bottom regions. If the screen orientation is portrait, two audio acquisition components are sequentially determined along a second direction of the electronic device, serving as the left and right channel audio acquisition components respectively, wherein the second direction is the direction between the first and second sides of the electronic device.
[0051] For example, the left channel audio acquisition component and the right channel audio acquisition component are illustrated with the structure shown in Figure 3.
[0052] As shown in Figure 5, when the electronic device is used in the first landscape mode, the left channel audio acquisition component 501 is located at the top of the electronic device, i.e., the first short side 201, and the right channel audio acquisition component 502 is located at the bottom of the electronic device, i.e., the second short side 202. The left channel audio acquisition component 501 is located on the left side of the screen, and the right channel audio acquisition component 502 is located on the right side of the screen. Referring to Figure 3, it can be seen that the left channel audio acquisition component 501 is the first microphone 301, and the right channel audio acquisition component 502 is the third microphone 303. Among them, there are two microphones distributed on the bottom edge, namely the second microphone 302 and the third microphone 303. Since the third microphone 303 is closer to the upper part of the screen (i.e., the second long side 204) than the second microphone 302 in the first landscape mode, based on the approximate holding area of the user in the landscape mode, the third microphone 303 is selected as the right channel audio acquisition component 502, which is less likely to be blocked by the user's hand compared to the second microphone 302.
[0053] As shown in Figure 6, when the electronic device is used in the second landscape mode, the left channel audio acquisition component 601 is located at the bottom of the electronic device, i.e., the second short side 202, and the right channel audio acquisition component 602 is located at the top of the electronic device, i.e., the first short side 201. Referring to Figure 3, it can be seen that the left channel audio acquisition component 601 is the second microphone 302, and the right channel audio acquisition component 602 is the first microphone 301. Similarly, there are two microphones distributed at the bottom, namely the second microphone 302 and the third microphone 303. Since the second microphone 302 is closer to the top of the screen than the third microphone 303 in the second landscape mode, the second microphone 302 is chosen as the left channel audio acquisition component 601 to reduce the probability of the microphone being blocked by the user's hand.
[0054] As shown in Figure 7, when the electronic device is used in the first portrait screen state, the left channel audio acquisition component 701 and the right channel audio acquisition component 702 are located on the left side of the screen (i.e., the first long side 203) and the right channel audio acquisition component 702 is located on the right side of the screen (i.e., the second long side 204). Referring to Figure 3, it can be seen that the left channel audio acquisition component 701 is the second microphone 302 and the right channel audio acquisition component 702 is the third microphone 303.
[0055] In addition, since the second vertical screen state is not the user's usual usage mode, no illustration is provided. Compared with the first vertical screen state, the left and right audio acquisition components in the second usage state are still located in the bottom area of the electronic device, which is the opposite of the distribution in the first vertical screen state. That is, the left audio acquisition component is the third microphone 303, and the right audio acquisition component is the second microphone 302.
[0056] Therefore, based on the foregoing description, there are two possible ways to configure the first audio acquisition component and the second audio acquisition component.
[0057] In the first configuration, there are two audio acquisition components: a left-channel audio acquisition component and a right-channel audio acquisition component. That is, the first audio acquisition component can be either a top or bottom audio acquisition component, and the second audio acquisition component is a rear audio acquisition component. Taking landscape mode as an example, with two first audio acquisition components—a top microphone and a bottom microphone—and a rear microphone, the top and rear microphones are combined to execute the audio processing method of this application, and the bottom and rear microphones are also combined to execute the audio processing method of this application.
[0058] The second type involves a first audio acquisition component and a second audio acquisition component, one of which is a left channel audio acquisition component and the other is a right channel audio acquisition component. It is assumed that the first audio acquisition component is a top microphone and the second audio acquisition component is a bottom microphone.
[0059] Both methods are applicable to the audio processing method of this application. The difference lies in the following: In the first method, both the left and right channel audio acquisition components are low signal-to-noise ratio (SNR) microphones, while the high SNR microphone is a rear microphone. In some embodiments, the electronic device has only one high SNR microphone, i.e., a rear microphone. In the second method, one of the left and right channel audio acquisition components is a high SNR microphone, and the other is a low SNR microphone. The first method allows for a more balanced audio quality in the audio signals acquired by the left and right channel audio acquisition components. Because both have low SNR, it avoids the listener perceiving one channel as clearer while the other is blurry, thus affecting the overall stereo experience. The second method reduces the number of audio acquisition components involved in the audio acquisition process, lowering the power consumption of the electronic device. In some embodiments, the second method can also reduce the number of audio acquisition components in the electronic device, thus reducing the cost of the electronic device.
[0060] It is understood that the first and second methods can be set based on actual usage needs and are not limited thereto. However, for the sake of explaining the embodiments, this application may assume the first method is used, that is, assume the second audio acquisition component is a rear microphone, the first audio acquisition component is a top microphone or a bottom microphone, the top microphone serves as the acquisition component for the left channel audio, and the bottom microphone serves as the acquisition component for the right channel audio. It is understood that this is not a limitation of this application.
[0061] It is understood that in the embodiments of this application, the first audio acquisition component may not be used for acquiring left channel audio or right channel audio, but may also be used for mono audio acquisition. For example, in some scenarios, mono mode is selected to highlight human voices. In mono mode, the reason why a high signal-to-noise ratio microphone is not always used is that the position of the high signal-to-noise ratio microphone of the electronic device is relatively fixed. For example, a high signal-to-noise ratio microphone is set in the back area of the electronic device. Users may choose different microphones in mono mode to perform recording operations based on different usage habits. Therefore, users do not have to use the high signal-to-noise ratio microphone in a fixed way, giving users more choices. At the same time, the microphone in mono mode can be corrected based on the high signal-to-noise ratio microphone to ensure sound quality.
[0062] S402: If it is detected that the first audio signal does not meet the preset usage conditions, obtain the phase information of the first audio signal and the amplitude information of the second audio signal.
[0063] It is understood that the preset usage conditions refer to the pre-set conditions for evaluating the quality of the first audio signal acquired by the first audio acquisition component. The first audio signal may not meet the preset usage conditions in the following situations: poor sound quality of the first audio signal, for example, excessive background noise or electrical noise, signal distortion (such as the risk of sound distortion); unsuitable dynamic range of the first audio signal, for example, the volume is too low or too high; incomplete recording, for example, there is a lack of audio in some time periods.
[0064] In this embodiment of the application, the first audio signal not meeting the preset usage conditions means that the first audio signal has the risk of distortion, which will be explained in detail in the following embodiments.
[0065] Understandably, if the first audio signal meets the preset usage conditions, indicating that it can be directly used as the left or right channel, then the first audio signal can be used as the target audio signal. Therefore, the target audio signal can be understood as the channel signal output by the electronic device in this audio acquisition operation, which can be considered to correspond to the first audio acquisition component.
[0066] S403: Based on the phase information of the first audio signal and the amplitude information of the second audio signal, a target audio signal is synthesized.
[0067] If the first audio signal is found to not meet the preset usage conditions, it indicates that the first audio signal should not be used as a left channel signal or a right channel signal. Therefore, the first audio signal needs to be repaired. The repair method is to obtain the phase information of the first audio signal and the amplitude information of the second audio signal, and then synthesize the target audio signal based on the phase information and amplitude information.
[0068] Specifically, a Fourier transform is performed on the first audio signal to obtain the frequency domain information corresponding to the first audio signal, and the phase information corresponding to the first audio signal is extracted from the frequency domain information. Similarly, a Fourier transform is performed on the second audio signal to obtain the frequency domain information corresponding to the second audio signal, and the amplitude information corresponding to the second audio signal is extracted from the frequency domain information. A new frequency domain signal is generated based on the phase information and amplitude information, and then an inverse Fourier transform is performed on the frequency domain signal to obtain the target audio signal.
[0069] It is understood that the target audio signal can be used as either left or right channel audio, depending on the function of the first audio acquisition component. For example, if the first audio acquisition component is used to acquire the left channel signal, then the target audio signal is used as the left channel audio; if the first audio acquisition component is used to acquire the right channel signal, then the target audio signal is used as the right channel audio. In other words, this embodiment utilizes the clearer amplitude information of the audio signal acquired by a high signal-to-noise ratio microphone, while retaining the phase characteristics of the left or right channel acquired by a low signal-to-noise ratio microphone, thereby generating a new audio signal, reducing the possibility of distortion and clipping, and preserving the phase characteristics of the left or right channel.
[0070] Therefore, when using a high signal-to-noise ratio (SNR) audio acquisition component, and when using both low and high SNR audio acquisition components to acquire audio, if the audio acquired by the low SNR audio acquisition component does not meet the usage conditions, the target audio signal is synthesized using the phase of the low SNR signal and the amplitude of the high SNR signal to replace the signal that would normally be output from the low SNR audio acquisition component.
[0071] Please refer to Figure 8. This application embodiment provides an audio processing method applied to the aforementioned electronic device. Specifically, the method includes: S801 to S806.
[0072] S801: Obtain the first audio signal acquired by the first audio acquisition component and the second audio signal acquired by the second audio acquisition component.
[0073] S802: Obtain the signal strength of the first audio signal.
[0074] S803: Determine whether there is a risk of distortion in the first audio signal based on the signal strength.
[0075] Crackling (or distortion) refers to the distortion of sound during audio signal transmission or playback. This typically manifests as a loss of clarity and sound quality, resulting in a blurry, harsh, or unnatural sound. Crackling mainly includes overload distortion, frequency distortion, time-domain distortion, and nonlinear distortion. Overload distortion occurs when the intensity of the audio signal exceeds the processing range of a device (such as a microphone, amplifier, or speaker), causing the waveform to be clipped, resulting in distortion. In this case, the top of the audio waveform is flattened, producing a discordant sound. Frequency distortion refers to the excessive amplification or attenuation of certain frequency components, causing a change in timbre. For example, excessively strong low frequencies may make the sound muddy, while attenuation of high frequencies may make the sound dull. Time-domain distortion refers to improper processing of the audio signal over time, such as delays and echoes, which can also lead to crackling. Nonlinear distortion occurs when a device's response is not linear during audio processing, potentially causing the amplification or attenuation of certain frequencies, thus inducing distortion.
[0076] Compared to high signal-to-noise ratio (SNR) microphones, low SNR microphones are more prone to audio distortion. Specifically, low SNR microphones capture audio signals with a lower ratio of background noise. During recording, background noise mixes with the useful audio signal, which may force the gain of the low SNR audio signal to be increased to capture a clearer signal. This can easily lead to overload distortion in the captured audio, further increasing the risk of distortion. In other words, to overcome the excessive noise of low SNR microphones, users typically need to increase the gain of the useful signal. When the gain is too high, the audio signal can easily reach the device's limit, leading to overload and distortion, and ultimately, distortion.
[0077] Therefore, after the first audio acquisition component acquires the first audio signal, since the first audio acquisition component is a microphone with a low signal-to-noise ratio, it is necessary to determine whether there is a risk of distortion in the audio acquired by the first audio acquisition component, so as to avoid a poor listening experience for the user when playing the audio later.
[0078] Understandably, based on the foregoing, distortion may be caused by overload distortion of the audio signal, meaning the audio signal is clipped due to excessive intensity, leading to distortion and thus making it prone to distortion. Therefore, the method for determining the risk of distortion could be to obtain the signal strength of the first audio signal and determine whether the first audio signal has a risk of distortion based on the signal strength of the first audio signal.
[0079] Specifically, by setting the maximum value of the audio signal supported by the electronic device as a specified threshold, the implementation method for determining whether the first audio signal has a risk of distortion based on the signal strength of the first audio signal can include at least the following two detection methods:
[0080] The first detection method sets a value less than a specified threshold, denoted as the first specified value. This first specified value can be used as a preset value. If the signal strength of the first audio signal is greater than the preset value, it is determined that the first audio signal has a risk of distortion. If the signal strength of the first audio signal is less than or equal to the preset value, it is also determined that the first audio signal has a risk of distortion. For example, the first specified value can be a smaller value than the maximum value. Assuming the maximum value is 0dB, the first specified value can be -3dB or -6dB.
[0081] The second detection method is similar to the first. Specifically, it obtains the absolute value of the difference between the signal strength of the first audio signal and a specified threshold, using this as the specified difference. If the specified difference is less than a preset value, it obtains the phase information of the first audio signal and the amplitude information of the second audio signal. For example, the preset value can be a small value; if the specified difference is less than the preset value, it indicates that the signal strength of the first audio signal is very close to or has already reached the specified threshold.
[0082] It should be noted that the signal strength of the first audio signal refers to the signal strength, i.e., amplitude, of the digital audio signal after the first audio signal acquired by the first audio acquisition component is converted into a digital audio signal. Therefore, based on the difference between the input digital audio signal and the maximum value of 0dB, if the input signal is very close to 0dB or is already 0dB, it is considered to have a risk of distortion.
[0083] Understandably, this specified threshold (e.g., 0dB) as a maximum value refers to the maximum value that an electronic device typically represents, i.e., the maximum volume that a digital system can represent. Exceeding 0dB can lead to clipping and distortion. Therefore, when the audio signal acquired by the first audio acquisition component is determined to be equal to or close to 0dB, it indicates that it may have been clipped, and there is a high probability that it will cause distortion during later playback.
[0084] It should be noted that during the actual testing process, the test result of whether there is a risk of sound distortion may not be given. However, the test result of whether the signal strength of the first audio signal is greater than the preset value or whether the specified difference is less than the preset value can be given. The test result can indirectly determine whether there is a risk of sound distortion.
[0085] For example, if the specified difference is determined to be less than a preset value, the operation of obtaining the phase information of the first audio signal and the amplitude information of the second audio signal, as well as the subsequent synthesis operation, are directly executed.
[0086] S804: Obtain the phase information of the first audio signal and the amplitude information of the second audio signal.
[0087] S805: Based on the phase information of the first audio signal and the amplitude information of the second audio signal, a target audio signal is synthesized.
[0088] S806: Use the first audio signal as the target audio signal.
[0089] In other words, if it is determined that the first audio signal has a risk of distortion, for example, if the specified difference is less than a preset value, then S804 is executed; if it is determined that the first audio signal does not have a risk of distortion, for example, if the specified difference is greater than or equal to a preset value, then S806 is executed.
[0090] Therefore, in this embodiment of the application, when there is a risk of distortion in the audio signal collected by the low signal-to-noise ratio microphone, a new signal, namely the target audio signal, can be obtained by synthesizing the amplitude information of the high signal-to-noise ratio audio signal and the phase of the audio signal collected by the low signal-to-noise ratio microphone. This new signal serves as the final output signal of the low signal-to-noise ratio microphone.
[0091] Furthermore, as mentioned above, in the field of left and right channel audio acquisition, there are two first audio acquisition components: a left channel audio acquisition component and a right channel audio acquisition component. That is, the audio acquisition process is in left and right channel mode, so the corresponding first audio signal can include both left and right channel audio signals. Assuming the second audio acquisition component is neither a left nor a right channel audio acquisition component, the aforementioned audio processing method can include: acquiring the initial left channel audio signal acquired by the left channel audio acquisition component, the initial right channel audio signal acquired by the right channel audio acquisition component, and the second audio signal acquired by the second audio acquisition component; if the initial left channel audio signal is detected to not meet preset usage conditions, the initial left channel audio signal is acquired... The system obtains the phase information of the initial left channel audio signal and the amplitude information of the second audio signal. Based on these two information, a left channel output audio signal is synthesized. If the initial left channel audio signal meets a preset usage condition, it is used as the left channel output audio signal. If the initial right channel audio signal does not meet the preset usage condition, the system obtains the phase information of the initial right channel audio signal and the amplitude information of the second audio signal. Based on these two information, a right channel output audio signal is synthesized. If the initial right channel audio signal meets the preset usage condition, it is used as the right channel output audio signal. Therefore, a high signal-to-noise ratio microphone can be used to solve the distortion problem of the left and right channel audio signals.
[0092] Furthermore, in the case of left and right channel mode during this audio acquisition process, if one of the first and second audio acquisition components is a left channel audio acquisition component and the other is a right channel audio acquisition component, then the first audio signal acquired by the first audio acquisition component and the second audio signal acquired by the second audio acquisition component are acquired. The second audio signal acquired by the second audio acquisition component is used as the output audio signal of the corresponding channel of the second audio acquisition component. If the first audio signal is detected to not meet the preset usage conditions, the phase information of the first audio signal and the amplitude information of the second audio signal are acquired. Based on the phase information of the first audio signal and the amplitude information of the second audio signal, a target audio signal is synthesized and used as the output audio signal of the corresponding channel of the first audio acquisition component. If the first audio signal is detected to meet the preset usage conditions, the first audio signal is used as the output audio signal of the corresponding channel of the first audio acquisition component. Therefore, even when only one of the left and right channels has a high signal-to-noise ratio audio acquisition component, the distortion problem of the audio signals in the left and right channels can be solved.
[0093] Here, the number of the first audio acquisition components is one, that is, the audio acquisition process is in mono mode and the above method can be executed. It can not only solve the distortion problem in mono mode, but also the second audio acquisition component can be the only high signal-to-noise ratio microphone of the electronic device. Users can freely choose any microphone with a low microphone signal-to-noise ratio other than the high signal-to-noise ratio microphone value as the first audio acquisition component without worrying about distortion problem or causing excessive cost.
[0094] For example, the above audio processing method will be further explained below with reference to the hardware structure of the electronic device performing the above audio processing method, specifically for the acquisition of the left channel signal and the right channel signal.
[0095] Figure 9 shows a structural block diagram of an electronic device provided in an embodiment of this application. The electronic device 100 can be a smartphone, tablet computer, e-reader, or other electronic device capable of running applications. The electronic device 100 in this application may include one or more of the following components: an audio processing module 101, an audio acquisition component 102, and one or more applications. The one or more applications may be stored in a memory and configured to be executed by the audio processing module 101. The one or more applications are configured to perform the methods described in the foregoing method embodiments. The implementation of the audio processing module 101 and the audio acquisition component 102 can be referred to the foregoing content and will not be repeated here.
[0096] As shown in Figure 10, both the first audio acquisition component and the second audio acquisition component are connected to the audio processing module, which executes the method described above. As shown in Figure 10, the first audio acquisition component is connected to the audio processing module via a first analog-to-digital converter (ADC1), and the second audio acquisition component is connected to the audio processing module via a second analog-to-digital converter (ADC2). In other words, the signal processed by the audio processing module is the digital signal obtained after analog-to-digital conversion of the audio signals acquired by the first and second audio acquisition components. That is, both the aforementioned first and second audio signals are digital audio signals.
[0097] As shown in Figure 11, the audio processing module 1100 includes an amplitude comparator 1101 and a signal replacement 1102. The amplitude comparator 1101 is connected to the first audio acquisition component and the signal replacement 1102. The amplitude comparator 1101 acquires the signal strength of the first audio signal. If it is determined based on the signal strength that the first audio signal has a risk of distortion, it outputs a first specified signal to the signal replacement 1102. The signal replacement 1102 is connected to the amplitude comparator 1101, the first audio acquisition component, and the second audio acquisition component. The signal replacement is used to acquire the phase information of the first audio signal and the amplitude information of the second audio signal when the first specified signal is received. Based on the phase information of the first audio signal and the amplitude information of the second audio signal, a target audio signal is synthesized.
[0098] It should be noted that the amplitude comparator can output a variety of different signals to inform the signal replacement of different detection results. Furthermore, as long as there is a risk of distortion in the first audio signal, the replacement operation is performed, that is, the target audio signal is synthesized based on the phase information of the first audio signal and the amplitude information of the second audio signal.
[0099] Specifically, in addition to detecting whether the first audio signal is distorted, the amplitude comparator also detects whether the second signal is distorted. Although the replacement operation is performed when the first audio signal is distorted, the detection result can be informed to the signal replacement device so that the signal replacement device can perform some special processing in some scenarios.
[0100] Specifically, assuming the first designated signal includes a first signal and a second signal, if it is determined based on the strength of the first signal that the first audio signal has a risk of distortion and based on the strength of the second signal that the second audio signal does not have a risk of distortion, the first signal is output to the signal replacement device; if it is determined based on the strength of the first signal that the first audio signal has a risk of distortion and based on the strength of the second signal that the second audio signal has a risk of distortion, the second signal is output to the signal replacement device. Since the function of the first designated signal is to trigger the signal replacement device to perform a replacement operation, the signal replacement device will perform the replacement operation when it detects either the first signal or the second signal. Of course, since there is a difference between the first signal and the second signal, the signal replacement device can determine whether the second audio signal has a risk of distortion based on the first and second signals, knowing that the first audio signal has a risk of distortion. In scenarios where it is necessary to remind the user, it can issue a corresponding reminder based on the first and second signals, or inform the backend processing module, so that when the amplitude comparator sends the second signal, the backend processing module can perform a repair operation on the target audio signal output by the signal replacement device to repair the distortion in the signal. The specific operation is not limited here.
[0101] In addition, the amplitude comparator can also output a second specified signal. Specifically, the amplitude comparator is also used to output a second specified signal to the signal replacement if it is determined based on the signal strength that the first audio signal does not have a risk of distortion. The signal replacement is also used to take the first audio signal as the target audio signal when the second specified signal is received.
[0102] As mentioned earlier, in some scenarios, specifically in the field of left and right channel audio acquisition, the target audio signal includes both the left and right channel audio signals. As shown in Figure 12, a first audio acquisition component is used to acquire either the left or right channel audio signal. There are two first audio acquisition components: a left channel audio acquisition component and a right channel audio acquisition component. There are at least two audio processing modules: a first audio processing module and a second audio processing module. The first audio processing module is connected to the left and second audio acquisition components. If the audio signal from the left channel audio acquisition component is detected to not meet preset usage conditions, it synthesizes the left channel audio signal based on the phase information of the left channel audio acquisition component and the amplitude information of the second audio signal. The second audio processing module is connected to the right and second audio acquisition components. If the audio signal from the right channel audio acquisition component is detected to not meet preset usage conditions, it synthesizes the right channel audio signal based on the phase information of the right channel audio acquisition component and the amplitude information of the second audio signal. The left channel audio acquisition component is located in the top area of the electronic device, the right channel audio acquisition component is located in the bottom area of the electronic device, and the second audio acquisition component is located in the rear area of the electronic device.
[0103] Therefore, by detecting and replacing the signal, the left channel output signal is guaranteed to have no distortion or less distortion, while spatial information is preserved, achieving the effect of two high signal-to-noise ratio microphones with a single high signal-to-noise ratio microphone.
[0104] Please refer to Figure 13, which shows a structural block diagram of an audio processing device provided in an embodiment of this application. This device is applied to an electronic device, which includes multiple audio acquisition components, including a first audio acquisition component and a second audio acquisition component, wherein the signal-to-noise ratio (SNR) of the second audio acquisition component is higher than that of the first audio acquisition component. Specifically, the device may include: an acquisition unit 1301, a detection unit 1302, and a synthesis unit 1303.
[0105] The acquisition unit 1301 is used to acquire the first audio signal acquired by the first audio acquisition component and the second audio signal acquired by the second audio acquisition component.
[0106] The detection unit 1302 is used to acquire the phase information of the first audio signal and the amplitude information of the second audio signal if the first audio signal is detected to not meet the preset usage conditions.
[0107] Furthermore, the detection unit 1302 is also used to obtain the signal strength of the first audio signal; if it is determined based on the signal strength that the first audio signal has a risk of distortion, then the phase information of the first audio signal and the amplitude information of the second audio signal are obtained.
[0108] Furthermore, the detection unit 1302 is also used to obtain the absolute value of the difference between the signal strength of the first audio signal and a specified threshold, as the specified difference; if the specified difference is less than a preset value, the phase information of the first audio signal and the amplitude information of the second audio signal are obtained.
[0109] Furthermore, the detection unit 1302 is also configured to use the first audio signal as the target audio signal if the first audio signal is detected to meet the preset usage conditions.
[0110] The synthesis unit 1303 is used to synthesize a target audio signal based on the phase information of the first audio signal and the amplitude information of the second audio signal.
[0111] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0112] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0113] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0114] Please refer to Figure 14, which shows a structural block diagram of a computer-readable medium provided in an embodiment of this application. The computer-readable medium 1400 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0115] Computer-readable medium 1400 may be an electronic storage device such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, computer-readable medium 1400 includes non-volatile computer-readable storage medium. Computer-readable medium 1400 has storage space for program code 1410 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. Program code 1410 may be compressed, for example, in a suitable form.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An audio processing method, characterized in that, An electronic device is applied to the present invention, the electronic device including multiple audio acquisition components, the multiple audio acquisition components including a first audio acquisition component and a second audio acquisition component, wherein the signal-to-noise ratio of the second audio acquisition component is higher than that of the first audio acquisition component. The method includes: acquiring a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component; if it is detected that the first audio signal does not meet preset usage conditions, acquiring phase information of the first audio signal and amplitude information of the second audio signal; and synthesizing a target audio signal based on the phase information of the first audio signal and the amplitude information of the second audio signal.
2. The method according to claim 1, characterized in that, If the first audio signal is detected to not meet the preset usage conditions, the step of obtaining the phase information of the first audio signal and the amplitude information of the second audio signal includes: obtaining the signal strength of the first audio signal; if it is determined based on the signal strength that the first audio signal has a risk of distortion, then obtaining the phase information of the first audio signal and the amplitude information of the second audio signal.
3. The method according to claim 2, characterized in that, If it is determined based on the signal strength that the first audio signal has a risk of distortion, then the phase information of the first audio signal and the amplitude information of the second audio signal are obtained, including: obtaining the absolute value of the difference between the signal strength of the first audio signal and a specified threshold, as the specified difference; if the specified difference is less than a preset value, then the phase information of the first audio signal and the amplitude information of the second audio signal are obtained.
4. The method according to claim 1, characterized in that, Also includes: If the first audio signal is detected to meet the preset usage conditions, the first audio signal is used as the target audio signal.
5. The method according to any one of claims 1-4, characterized in that, The first audio acquisition component is used to acquire the left channel audio signal or the right channel audio signal, wherein the target audio signal is the left channel audio signal or the right channel audio signal.
6. An audio processing apparatus, characterized in that, An electronic device is applied to an electronic device, the electronic device including multiple audio acquisition components, the multiple audio acquisition components including a first audio acquisition component and a second audio acquisition component, wherein the signal-to-noise ratio of the second audio acquisition component is higher than that of the first audio acquisition component. The device includes: an acquisition unit, configured to acquire a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component; a detection unit, configured to acquire phase information of the first audio signal and amplitude information of the second audio signal if the first audio signal is detected to not meet preset usage conditions; and a synthesis unit, configured to synthesize a target audio signal based on the phase information of the first audio signal and the amplitude information of the second audio signal.
7. An electronic device, characterized in that, include: Multiple audio acquisition components, including a first audio acquisition component and a second audio acquisition component, wherein the signal-to-noise ratio of the second audio acquisition component is higher than that of the first audio acquisition component; An audio processing module is connected to the first audio acquisition component and the second audio acquisition component, respectively, and the audio processing module is used to perform the method as described in any one of claims 1-5.
8. The electronic device according to claim 7, characterized in that, The audio processing module includes an amplitude comparator and a signal substitute. The amplitude comparator is connected to the first audio acquisition component and the signal substitute. The amplitude comparator is used to obtain the signal strength of the first audio signal. If it is determined based on the signal strength that the first audio signal has a risk of distortion, a first specified signal is output to the signal substitute. The signal substitute is connected to the amplitude comparator, the first audio acquisition component, and the second audio acquisition component. The signal substitute is used to obtain the first audio signal acquired by the first audio acquisition component and the second audio signal acquired by the second audio acquisition component. Upon receiving the first specified signal, the signal substitute obtains the phase information of the first audio signal and the amplitude information of the second audio signal. Based on the phase information of the first audio signal and the amplitude information of the second audio signal, a target audio signal is synthesized.
9. The electronic device according to claim 8, characterized in that, The first designated signal includes a first signal and a second signal; the amplitude comparator is also connected to the second audio acquisition component; the amplitude comparator is further configured to: acquire the first audio signal acquired by the first audio acquisition component and the second audio signal acquired by the second audio acquisition component, acquire the first signal strength of the first audio signal and the second signal strength of the second audio signal; if it is determined based on the first signal strength that the first audio signal has a risk of distortion and based on the second signal strength that the second audio signal does not have a risk of distortion, output the first signal to the signal replacement device; if it is determined based on the first signal strength that the first audio signal has a risk of distortion and based on the second signal strength that the second audio signal has a risk of distortion, output the second signal to the signal replacement device.
10. The electronic device according to claim 8, characterized in that: The amplitude comparator is further configured to output a second specified signal to the signal replacement unit if it is determined based on the signal strength that the first audio signal does not have a risk of distortion; the signal replacement unit is further configured to use the first audio signal as the target audio signal upon receiving the second specified signal.
11. The electronic device according to claim 7, characterized in that, The target audio signal includes a left channel audio signal and a right channel audio signal. The first audio acquisition component is used to acquire the left channel audio signal or the right channel audio signal. There are two first audio acquisition components, namely a left channel audio acquisition component and a right channel audio acquisition component. There are at least two audio processing modules, namely a first audio processing module and a second audio processing module. The first audio processing module is connected to the left channel audio acquisition component and the second audio acquisition component, and is used to synthesize a left channel audio signal based on the phase information of the left channel audio acquisition component and the amplitude information of the second audio signal if the audio signal of the left channel audio acquisition component is detected to not meet the preset usage conditions; the second audio processing module is connected to the right channel audio acquisition component and the second audio acquisition component, and is used to synthesize a right channel audio signal based on the phase information of the right channel audio acquisition component and the amplitude information of the second audio signal if the audio signal of the right channel audio acquisition component is detected to not meet the preset usage conditions.
12. The electronic device according to claim 11, characterized in that, The left channel audio acquisition component is located in the top area of the electronic device, the right channel audio acquisition component is located in the bottom area of the electronic device, and the second audio acquisition component is located in the rear area of the electronic device.