Audio processing method and device and electronic equipment
By generating a common basic signal to replace the audio signal from the faulty microphone, the problem of recording interruption or low volume caused by microphone malfunction was solved, thus ensuring the continuity and quality of recording.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-01
AI Technical Summary
When the microphone fails to record properly for some reason, the recording may be silent or very quiet, affecting the user's normal experience.
By acquiring audio signals from multiple audio acquisition components, a common base signal is generated. When an abnormal audio acquisition component is detected, this common base signal is used as the target audio signal of the abnormal component, ensuring the continuity and quality of the audio acquisition process.
When the microphone malfunctions, ensure recording continuity and audio quality, avoid recording interruptions or low volume, and improve user experience.
Smart Images

Figure CN121967589A_ABST
Abstract
Description
Audio processing methods, devices and electronic equipment Technical Field
[0001] This application relates to the field of mobile terminal technology, and more specifically, to an audio processing method, apparatus, and electronic device. Background Technology
[0002] Recording is an essential function of smartphones and other terminal devices, widely used in voice calls and sound capture. Therefore, the quality of the recording directly impacts the user experience. The recording function of terminal devices primarily relies on the microphone for audio capture. However, when the microphone malfunctions for various reasons, the recording may be silent or very quiet, affecting the user's experience. Summary of the Invention
[0003] This application proposes an audio processing method, apparatus, and electronic device.
[0004] In a first aspect, this application provides an audio processing method applied to an electronic device, the electronic device including multiple audio acquisition components, the multiple audio acquisition components including a first audio acquisition component and a second audio acquisition component, the method including: during audio acquisition, acquiring a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component; obtaining a common base signal based on the first audio signal and the second audio signal; if it is determined that the first audio acquisition component is abnormal based on the first audio signal and the second audio signal, using the common base signal as the target audio signal corresponding to the first audio acquisition component.
[0005] Secondly, this application also provides an audio processing apparatus applied to an electronic device. The electronic device includes multiple audio acquisition components, including a first audio acquisition component and a second audio acquisition component. The apparatus includes: an acquisition unit, a synthesis unit, and an execution unit. The acquisition unit is used to acquire, during the audio acquisition process, a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component. The synthesis unit is used to obtain a common base signal based on the first audio signal and the second audio signal. The execution unit is used to, if it is determined based on the first audio signal and the second audio signal that the first audio acquisition component is abnormal, use the common base signal as the target audio signal corresponding to the first audio acquisition component.
[0006] Thirdly, this application also provides an electronic device, including: a plurality of audio acquisition components, the plurality of audio acquisition components including a first audio acquisition component and a second audio acquisition component; an audio processing module, respectively connected to the first audio acquisition component and the second audio acquisition component, the audio processing module being used to perform the above method.
[0007] The audio processing method, apparatus, electronic device, and computer-readable medium provided in this application, during the audio acquisition process, acquire a first audio signal acquired by a first audio acquisition component and a second audio signal acquired by a second audio acquisition component; obtain a common base signal based on the first audio signal and the second audio signal; if it is determined that the first audio acquisition component is malfunctioning based on the first audio signal and the second audio signal, use the common base signal as the target audio signal corresponding to the first audio acquisition component. Therefore, when it is determined that the first audio acquisition component is malfunctioning, the common base signal obtained based on the first audio signal of the first audio acquisition component and the second audio signal acquired by the second audio acquisition component is used as the target audio signal corresponding to the first audio acquisition component, avoiding the first audio acquisition component being unable to obtain audio or obtaining audio of poor quality due to malfunction.
[0008] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 shows a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0011] Figure 2 shows a schematic diagram of the landscape and portrait screen states of an electronic device provided in an embodiment of this application;
[0012] Figure 3 shows a schematic diagram of an audio acquisition component of an electronic device provided in an embodiment of this application;
[0013] Figure 4 shows a flowchart of an audio processing method provided in an embodiment of this application;
[0014] Figure 5 shows a schematic diagram illustrating the positions of the left channel audio acquisition component and the right channel audio acquisition component provided in an embodiment of this application;
[0015] Figure 6 shows a schematic diagram illustrating the positions of the left channel audio acquisition component and the right channel audio acquisition component provided in another embodiment of this application;
[0016] Figure 7 shows a schematic diagram illustrating the positions of the left channel audio acquisition component and the right channel audio acquisition component provided in another embodiment of this application;
[0017] Figure 8 shows a flowchart of an audio processing method provided in another embodiment of this application;
[0018] Figure 9 shows a structural block diagram of an electronic device provided in an embodiment of this application;
[0019] Figure 10 shows a structural block diagram of an electronic device provided in yet another embodiment of this application;
[0020] Figure 11 shows a structural block diagram of an electronic device provided in another embodiment of this application;
[0021] Figure 12 shows a structural block diagram of an electronic device provided in yet another embodiment of this application;
[0022] Figure 13 shows a block diagram of an audio processing apparatus provided in an embodiment of this application;
[0023] Figure 14 illustrates a storage unit according to an embodiment of the present application for storing or carrying program code implementing the method according to an embodiment of the present application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without inventive effort are within the scope of protection of the present application.
[0025] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0026] Recording is an essential function of smartphones and other terminal devices, widely used in voice calls and sound capture. Therefore, the quality of the recording directly impacts the user experience. The recording function of terminal devices primarily relies on the microphone for audio capture. However, if the microphone malfunctions for various reasons, the recording may be silent or very quiet, affecting the user experience. For example, if a user is recording video on an electronic device in landscape mode and using the top and bottom microphones for audio capture, they may easily block the microphone holes, resulting in silent or low-volume recordings.
[0027] Therefore, this application provides an audio processing method that, in the event of a microphone malfunction, replaces the audio captured by the malfunctioning microphone with audio captured by other microphones participating in the audio acquisition operation, thereby ensuring that normal audio output can still be sent to the backend even when the microphone malfunctions.
[0028] Before introducing this method, we will first introduce the hardware environment of the electronic device of this application.
[0029] It should be noted that the electronic device can be a smartphone, tablet computer, e-reader, or other device capable of running applications. In the embodiments of this application, a smartphone is used as an example to illustrate the various embodiments of this application; however, the electronic device is not limited to smartphones.
[0030] As one implementation method, the screen orientation of an electronic device is defined as having four orientations: a first portrait orientation, a first landscape orientation, a second portrait orientation, and a second landscape orientation. As shown in Figure 1, the screen of the electronic device includes a first short side 201, a second short side 202, a first long side 203, and a second long side 204. It can be seen that the first short side 201 and the second short side 202 are positioned opposite each other, and the first long side 203 and the second long side 204 are positioned opposite each other. In this embodiment, the first short side 201 can be the top edge of the corresponding electronic device, and the second short side 202 can be the bottom edge of the corresponding electronic device. As shown in Figure 1, the top status bar of the electronic device is displayed at the position of the first short side 201.
[0031] In addition, to facilitate the explanation of the above four states, four different directions can be defined, as shown in Figure 2. These four different directions include direction F1, direction F2, direction F3, and direction F4. Figure 2(a) shows the first portrait screen state of the electronic device. Direction F1 refers to the direction from the first long side 203 to the second long side 204, and direction F2 refers to the direction from the first short side 201 to the second short side 202. When the user is facing the screen of the electronic device in the first portrait screen state, with the screen as a reference, the first short side 201 is located above the screen, the second short side 202 is located below the screen, the first long side 203 is located on the left side of the screen, and the second long side 204 is located on the right side of the screen.
[0032] Figure 2(b) shows the first landscape orientation of the electronic device. The direction from the second long side 204 to the first long side 203 is named direction F3. With the user facing the screen of the electronic device in the first landscape orientation, the second long side 204 is located above the screen, the first long side 203 is located below the screen, the first short side 201 is located on the left side of the screen, and the second short side 202 is located on the right side of the screen.
[0033] Figure 2(c) shows the second portrait screen state of the electronic device. The direction from the second short side 202 to the first short side 201 is named direction F4. When the user is facing the screen of the electronic device in the second portrait screen state, with the screen as a reference, the second short side 202 is located above the screen, the first short side 201 is located below the screen, the second long side 204 is located on the left side of the screen, and the first long side 203 is located on the right side of the screen.
[0034] Figure 2(d) shows the second landscape state of the electronic device. With the user facing the screen of the electronic device in the second landscape state, the first long side 203 is located above the screen, the second long side 204 is located below the screen, the second short side 202 is located on the left side of the screen, and the first short side 201 is located on the right side of the screen.
[0035] As can be seen, the left, right, top, and bottom edges of the screen are different in different landscape and portrait modes. The first landscape mode and the second landscape mode are reversed, that is, the second landscape mode is the state after rotating the first landscape mode by 180 degrees. Similarly, the first portrait mode and the second portrait mode are reversed.
[0036] In one implementation, the electronic device is equipped with multiple audio acquisition components, i.e., multiple microphones, which are distributed in different preset areas of the electronic device. These preset areas can be at least one of a top area, a bottom area, and a back area. The top area refers to the top housing area of the electronic device, i.e., the area of the aforementioned first short side 201; the bottom area refers to the bottom housing area of the electronic device, i.e., the area of the aforementioned second short side 202; and the back area refers to the back cover area of the electronic device. For example, a back microphone can be placed near the rear camera to better capture sound during video recording, reduce interference with other components (such as speakers), and minimize obstruction of the back microphone by the phone case.
[0037] It is understood that, in the embodiments of this application, the number of audio acquisition components and the installation position of each audio acquisition component are not limited. However, in some embodiments, the number of audio acquisition components may be four, and these four audio acquisition components (microphones) are distributed in the aforementioned preset area of the electronic device. For example, two microphones are provided in the top area, one microphone is provided in the bottom area, and one microphone is provided in the back area. Alternatively, one microphone may be provided in the top area, two microphones in the bottom area, and one microphone in the back area.
[0038] For example, as shown in FIG3, the multiple audio acquisition components of the electronic device are a first microphone 301, a second microphone 302, a third microphone 303 and a fourth microphone 304, wherein the first microphone 301 is disposed in the top area of the electronic device, the second microphone 302 and the third microphone 303 are disposed in the bottom area of the electronic device, and the fourth microphone 304 is disposed in the back area of the electronic device.
[0039] It should be noted that the perspective shown in Figure 3 is the user's perspective when looking at the screen of the electronic device. Therefore, the dashed line corresponding to the fourth microphone 304 indicates that the fourth microphone 304 is located in the back area of the electronic device. In addition, the circular patterns for each microphone shown in Figure 3 are used to roughly show the distribution area of each microphone, and do not limit the specific shape, specific position or specific size of the microphone.
[0040] As shown in Figure 4, this application embodiment provides an audio processing method applied to the above-mentioned electronic device. The electronic device includes multiple audio acquisition components, including a first audio acquisition component and a second audio acquisition component. The method includes steps S401 to S403.
[0041] S401: During the audio acquisition process, acquire the first audio signal acquired by the first audio acquisition component and the second audio signal acquired by the second audio acquisition component.
[0042] In this embodiment of the application, the first audio signal acquired by the first audio acquisition component is named the first audio signal, and the audio signal acquired by the second audio acquisition component is named the second audio signal. It can be understood that the first and second audio acquisition components acquire audio in the same scenario. That is, during this audio acquisition process, both the first and second audio acquisition components participate in audio recording, and they acquire audio almost synchronously. "Almost synchronously" means that when the electronic device starts recording, it synchronously triggers each audio acquisition component participating in the audio acquisition process to perform audio acquisition operations. However, due to possible delays in instructions or other reasons, there may be slight differences in the start times of audio acquisition by each audio acquisition component. These differences are negligible, and therefore, the first and second audio acquisition components can be considered to be acquiring audio synchronously.
[0043] It should be noted that the audio acquisition process can be in mono mode or left and right channel mode.
[0044] In one implementation, the left and right channel mode can refer to a first audio acquisition component and a second audio acquisition component, one used for left channel audio acquisition and the other for right channel audio acquisition. That is, the first audio acquisition component can be an audio acquisition component used to record the audio signal corresponding to the left or right channel during this audio acquisition process, and the second audio acquisition component can also be an audio acquisition component used to record the audio signal corresponding to the left or right channel. Of course, the channels recorded by the second audio acquisition component and the first audio acquisition component are different. For example, if the first audio acquisition component is used to record the audio signal corresponding to the left channel, then the second audio acquisition component is used to record the audio signal corresponding to the right channel, or if the first audio acquisition component is used to record the audio signal corresponding to the right channel, then the second audio acquisition component is used to record the audio signal corresponding to the left channel.
[0045] Alternatively, the second audio acquisition component may not be used to record the left or right channel audio signal, but rather serves as a reference signal so that the audio signals used for recording the left and right channels can be used to correct defects in the recorded audio signals. In this case, if the second audio acquisition component is not used to record the left or right channel audio signal, the audio processing method provided in this application will be executed for both the audio acquisition component used to record the left channel signal and the audio acquisition component used to record the right channel signal. That is, there are two first audio acquisition components: a left channel audio acquisition component and a right channel audio acquisition component, both of which serve as first audio acquisition components to execute the audio processing method of this application.
[0046] Therefore, the number of first audio acquisition components can be one or two. That is, the number of first audio acquisition components can be one, in which case the first audio acquisition component and the second audio acquisition component are used to acquire the left channel audio and the right channel audio respectively. Alternatively, there can be two first audio acquisition components, which are used to acquire the left channel audio and the right channel audio respectively. There is no limitation on this.
[0047] To describe the first and second audio acquisition components in the embodiments of this application, the selection method of the left and right channel audio acquisition components applied in this audio acquisition process is first explained. If the screen orientation is landscape, two audio acquisition components are sequentially determined along a first direction of the electronic device, serving as the left and right channel audio acquisition components respectively, wherein the first direction is the direction between the top and bottom regions. If the screen orientation is portrait, two audio acquisition components are sequentially determined along a second direction of the electronic device, serving as the left and right channel audio acquisition components respectively, wherein the second direction is the direction between the first and second sides of the electronic device.
[0048] For example, the left channel audio acquisition component and the right channel audio acquisition component are illustrated with the structure shown in Figure 3.
[0049] As shown in Figure 5, when the electronic device is used in the first landscape mode, the left channel audio acquisition component 501 is located at the top of the electronic device, i.e., the first short side 201, and the right channel audio acquisition component 502 is located at the bottom of the electronic device, i.e., the second short side 202. The left channel audio acquisition component 501 is located on the left side of the screen, and the right channel audio acquisition component 502 is located on the right side of the screen. Referring to Figure 3, it can be seen that the left channel audio acquisition component 501 is the first microphone 301, and the right channel audio acquisition component 502 is the third microphone 303. Among them, there are two microphones distributed on the bottom edge, namely the second microphone 302 and the third microphone 303. Since the third microphone 303 is closer to the upper part of the screen (i.e., the second long side 204) than the second microphone 302 in the first landscape mode, based on the approximate holding area of the user in the landscape mode, the third microphone 303 is selected as the right channel audio acquisition component 502, which is less likely to be blocked by the user's hand compared to the second microphone 302.
[0050] As shown in Figure 6, when the electronic device is used in the second landscape mode, the left channel audio acquisition component 601 is located at the bottom of the electronic device, i.e., the second short side 202, and the right channel audio acquisition component 602 is located at the top of the electronic device, i.e., the first short side 201. Referring to Figure 3, it can be seen that the left channel audio acquisition component 601 is the second microphone 302, and the right channel audio acquisition component 602 is the first microphone 301. Similarly, there are two microphones distributed at the bottom, namely the second microphone 302 and the third microphone 303. Since the second microphone 302 is closer to the top of the screen than the third microphone 303 in the second landscape mode, the second microphone 302 is chosen as the left channel audio acquisition component 601 to reduce the probability of the microphone being blocked by the user's hand.
[0051] As shown in Figure 7, when the electronic device is used in the first portrait screen state, the left channel audio acquisition component 701 and the right channel audio acquisition component 702 are located on the left side of the screen (i.e., the first long side 203) and the right channel audio acquisition component 702 is located on the right side of the screen (i.e., the second long side 204). Referring to Figure 3, it can be seen that the left channel audio acquisition component 701 is the second microphone 302 and the right channel audio acquisition component 702 is the third microphone 303.
[0052] In addition, since the second vertical screen state is not the user's usual usage mode, no illustration is provided. Compared with the first vertical screen state, the left and right audio acquisition components in the second usage state are still located in the bottom area of the electronic device, which is the opposite of the distribution in the first vertical screen state. That is, the left audio acquisition component is the third microphone 303, and the right audio acquisition component is the second microphone 302.
[0053] Therefore, based on the foregoing description, there are two possible configuration methods for the first and second audio acquisition components for the left and right channel modes.
[0054] In the first configuration, there are two audio acquisition components: a left-channel audio acquisition component and a right-channel audio acquisition component. That is, the first audio acquisition component can be either a top or bottom audio acquisition component, and the second audio acquisition component is a rear audio acquisition component. Taking landscape mode as an example, with two first audio acquisition components—a top microphone and a bottom microphone—and a rear microphone, the top and rear microphones are combined to execute the audio processing method of this application, and the bottom and rear microphones are also combined to execute the audio processing method of this application.
[0055] The second type involves a first audio acquisition component and a second audio acquisition component, one of which is a left channel audio acquisition component and the other is a right channel audio acquisition component. It is assumed that the first audio acquisition component is a top microphone and the second audio acquisition component is a bottom microphone.
[0056] It is understood that the first and second methods can be set based on actual usage needs and are not limited thereto. However, for the sake of explaining the embodiments, this application may assume the first method is used, that is, assume the second audio acquisition component is a rear microphone, the first audio acquisition component is a top microphone or a bottom microphone, the top microphone serves as the acquisition component for the left channel audio, and the bottom microphone serves as the acquisition component for the right channel audio. It is understood that this is not a limitation of this application.
[0057] In addition, in the embodiments of this application, the first audio acquisition component may not be used for the acquisition of left channel audio or right channel audio, but may also be used for mono audio acquisition. For example, in some scenarios, mono mode is selected to highlight human voices.
[0058] S402: Obtain a common base signal based on the first audio signal and the second audio signal.
[0059] It is understandable that, since the audio acquisition components participating in this audio acquisition process are divided into a first audio acquisition component and a second audio acquisition component, without limiting the number of the first and second audio acquisition components, the common basic signal can be understood as the common basic signal obtained from the audio signals acquired by each microphone participating in this audio acquisition process. As one implementation method, the common basic signal can be obtained by weighted summation of the first and second audio signals.
[0060] S403: If it is determined that the first audio acquisition component is abnormal based on the first audio signal and the second audio signal, the common basic signal is used as the target audio signal corresponding to the first audio acquisition component.
[0061] It should be noted that abnormalities in audio acquisition components can typically include sound quality abnormalities, functional abnormalities, and physical abnormalities. Sound quality abnormalities refer to distortion: sound distortion, excessive background noise, and inability to capture certain frequencies of sound. Functional abnormalities can include no microphone output, intermittent malfunctions, and inability to capture sound properly. Physical abnormalities can include physical damage, such as damage, cracks, or water ingress, as well as insufficient or unstable microphone power supply.
[0062] There are many reasons for the aforementioned anomalies, including obstruction and physical damage. In this application, we assume the anomaly is caused by obstruction, i.e., a blocked microphone hole. For example, when a user holds an electronic device, their hand may block the microphone's opening, preventing it from recording sound properly and leading to problems such as sound distortion, no sound, or extremely low volume. The method provided in this application can solve the audio recording defects caused by a blocked microphone hole.
[0063] Taking the example of an abnormality referring to a blocked hole, it is understandable that when selecting a second audio acquisition component from multiple audio acquisition components of an electronic device, an audio acquisition component that is not easily blocked will be selected as the second audio acquisition component. For example, the microphone on the back of the electronic device will be selected, as it is not easily blocked by the user's hand when the electronic device is used to record video. Of course, the second audio acquisition component can also be a microphone in other locations, which is not limited here.
[0064] It is understandable that the audio acquired by the malfunctioning audio acquisition component will differ from that acquired by the normal audio acquisition component, for example, in terms of poorer sound quality and lower signal strength. Therefore, by analyzing the first audio signal based on the first and second audio signals, it is possible to determine whether the first audio signal is abnormal, such as whether it is blocked.
[0065] Specifically, for the same sound source, especially a distant sound source, when multiple microphones simultaneously receive audio signals, the correlation between the audio signals collected by each microphone is significant. If the audio signals from highly correlated microphones show a significantly smaller difference in input amplitude between one microphone and the others, the algorithm determines it to be a blocked state; otherwise, it is considered an unblocked state. For example, if the correlation coefficient between the first and second audio signals is greater than a first specified threshold, and the amplitude of the first audio signal is less than the amplitude of the second audio signal (e.g., the amplitude of the first audio signal is less than the amplitude of the second audio signal and the difference between them is greater than a second specified threshold), then the first audio signal is determined to be blocked, i.e., abnormal. The correlation between two audio signals refers to the degree of similarity or interdependence between the two signals in time and frequency, and this correlation is commonly quantified using a correlation coefficient (such as the Pearson correlation coefficient). A correlation coefficient greater than the first specified threshold indicates that the audio signals collected by the first and second audio signals originate from the same scene; the amplitude of the first audio signal being less than the amplitude of the second audio signal and the difference between them being greater than the second specified threshold means that the amplitude of the first audio signal is significantly lower than that of the second audio signal.
[0066] If the first audio acquisition component is determined to be faulty, the common basic signal is used as the target audio signal corresponding to the faulty first audio acquisition component. The target audio signal can be regarded as the output audio of the first audio acquisition component. Therefore, when the first audio acquisition component is determined to be faulty, the common basic signal obtained based on the first audio signal of the first audio acquisition component and the second audio signal acquired by the second audio acquisition component is used as the target audio signal corresponding to the first audio acquisition component. This avoids the first audio acquisition component being unable to obtain audio or obtaining audio of poor quality due to the faulty first audio acquisition component.
[0067] Taking the aforementioned left and right channel mode as an example, assuming there are two first audio acquisition components, namely a left channel audio acquisition component and a right channel audio acquisition component, and the number of second audio acquisition components is not limited, then the first audio signal acquired by the left channel audio acquisition component is the initial left channel audio signal, and the first audio signal acquired by the right channel audio acquisition component is the initial right channel audio signal. The implementation method described above is as follows:
[0068] During audio acquisition, the initial left-channel audio signal acquired by the left-channel audio acquisition component, the initial right-channel audio signal acquired by the right-channel audio acquisition component, and the second audio signal acquired by the second audio acquisition component are obtained. A common base signal is obtained based on the initial left-channel audio signal, the initial right-channel audio signal, and the second audio signal. If the left-channel audio acquisition component malfunctions, the common base signal is used as the target audio signal corresponding to the left-channel audio acquisition component, i.e., the left-channel output audio signal. If the right-channel audio acquisition component malfunctions, the common base signal is used as the target audio signal corresponding to the right-channel audio acquisition component, i.e., the right-channel output audio signal.
[0069] It should be noted that the method for performing anomaly detection on the left channel audio acquisition component can be determined based on the initial left channel audio signal and the second audio signal. The method for performing anomaly detection on the right channel audio acquisition component can also be determined based on the initial right channel audio signal and the second audio signal. Alternatively, if the correlation coefficients of the initial left channel audio signal, the initial right channel audio signal, and the second audio signal are all greater than a first specified threshold, the audio acquisition component corresponding to the audio signal with significantly reduced amplitude is identified as the abnormal audio acquisition component. The implementation method for significantly reduced amplitude can refer to the aforementioned embodiments, and will not be repeated here.
[0070] Please refer to Figure 8. This application embodiment provides an audio processing method applied to the aforementioned electronic device. Specifically, the method includes: S801 to S805.
[0071] S801: During the audio acquisition process, acquire the first audio signal acquired by the first audio acquisition component and the second audio signal acquired by the second audio acquisition component.
[0072] S802: Obtain a common basic signal based on the first audio signal and the second audio signal.
[0073] S803: Determine whether the first audio acquisition component is abnormal based on the first audio signal and the second audio signal.
[0074] In this embodiment, "abnormality of the first audio acquisition component" means that the first audio acquisition component is suspected of being blocked, also known as "blocked hole". If the first audio acquisition component is abnormal, S804 is executed; if the first audio acquisition component is not abnormal, S805 is executed. The specific detection method can be referred to the foregoing embodiments, and will not be repeated here.
[0075] S804: Use the common basic signal as the target audio signal corresponding to the first audio acquisition component.
[0076] Understandably, although the amplitude of the audio signal acquired by the first audio acquisition component is smaller when it malfunctions, the amplitude of the second audio signal acquired by the second audio acquisition component is larger because the second audio acquisition component is acquiring audio normally. Therefore, using the common basic signal as the target audio signal (output signal) corresponding to the first audio acquisition component ensures that the audio signal output by the first audio acquisition component is relatively normal, reducing the impact of the first audio acquisition component's malfunction on audio acquisition. It is also understood that this target audio signal can be used as the output signal corresponding to the audio acquisition process; that is, when the electronic device performs an audio acquisition operation, it will use the target audio signal as the output signal corresponding to that audio operation.
[0077] S805: The target audio signal corresponding to the first audio acquisition component is obtained by synthesizing the common basic signal and the first audio signal.
[0078] In other words, assuming the first audio acquisition component is functioning normally, a fused signal is synthesized based on the common base signal and the first audio signal, and this fused signal is used as the target audio signal corresponding to the first audio acquisition component. This fusion can be achieved by weighting the common base signal and the first audio signal, where the first weight corresponding to the common base signal is less than the second weight corresponding to the first audio signal.
[0079] In addition, the acquisition of the fused signal can include at least three methods, namely the first fusion method, the second fusion method, and the third fusion method.
[0080] For this first fusion method, if it is determined that the first audio acquisition component is not abnormal based on the first audio signal and the second audio signal, a preset operation is performed on the first audio signal; the common basic signal and the first audio signal after the preset operation are combined to obtain a target audio signal corresponding to the first audio acquisition component.
[0081] For the second fusion method, if it is determined that the first audio acquisition component is not abnormal based on the first audio signal and the second audio signal, a preset operation is performed on the common basic signal; the first audio signal and the common basic signal after the preset operation are combined to obtain the target audio signal corresponding to the first audio acquisition component.
[0082] For the third fusion method, if it is determined that the first audio acquisition component is not abnormal based on the first audio signal and the second audio signal, the first audio signal and the common basic signal are synthesized to obtain a synthesized signal; a preset operation is performed on the synthesized signal to obtain a target audio signal corresponding to the first audio acquisition component.
[0083] As can be seen, all three fusion methods described above execute a preset operation, the difference being the object being operated on. In the first fusion method, the object is the first audio signal; in the second fusion method, it is the common base signal; and in the third fusion method, it is the synthesized signal. However, the fused signal obtained by these three fusion methods, which is the target audio signal corresponding to the first audio acquisition component, all include the common base signal.
[0084] Of course, it is understandable that, when the first audio acquisition component is not malfunctioning, the first audio signal can also be used as the target audio signal corresponding to the first audio acquisition component. However, in the embodiments of this application, when the first audio acquisition component is not malfunctioning, the above-mentioned fused signal is used as the target audio signal corresponding to the first audio acquisition component. That is to say, the first audio acquisition component always contains the common basic signal component, regardless of whether it is malfunctioning or not. It should be noted that this does not mean that the common basic signal contains the same audio content, but rather that the common basic signal generation operation will be performed regardless of whether the first audio acquisition component is malfunctioning or not, and the target audio signal corresponding to the first audio acquisition component will always contain the common basic signal.
[0085] The advantage of this is that it ensures that the target audio signal output by the first audio acquisition component always includes the common fundamental signal throughout the entire audio acquisition operation. This reduces the abruptness to the user caused by the missing or low amplitude of the first audio signal acquired by the first audio acquisition component when it malfunctions. In other words, since the common fundamental signal maintains a certain energy percentage when the first audio acquisition component is not malfunctioning, it can still output a sound with a certain loudness and maintain uninterrupted sound even if the first audio acquisition component malfunctions.
[0086] It should be noted that the preset operation includes at least one of signal attenuation operation and sound field expansion operation.
[0087] Specifically, the signal attenuation operation includes identifying the common portion of the first audio signal and the common base signal as the target portion, and performing a signal attenuation operation on the target portion of the first audio signal. The common target portion refers to the shared portion of the first audio signal and the common base signal. For example, this common portion can be the same frequency domain portion of both; that is, the same frequency domain portion of the first audio signal and the common base signal is identified as the target frequency domain. Then, the signal attenuation operation is performed on the audio portion of the target frequency domain of the audio signal to be attenuated. The audio signal to be attenuated corresponds to the three different fusion methods mentioned above, namely, the first audio signal, the common base signal, and the synthesized signal.
[0088] Taking the first audio signal as an example, when the first audio acquisition component is detected to be non-abnormal, the same target part of the first audio signal and the common base signal is determined, and a signal attenuation operation is performed on the target part of the first audio signal to obtain the first audio signal after signal attenuation. The common base signal and the first audio signal after signal attenuation are combined to obtain the target audio signal corresponding to the first audio acquisition component.
[0089] Understandably, assuming the first audio acquisition component is functioning correctly, when fusing the first audio signal and the common base signal, in order to avoid excessive volume and distortion or poor listening experience for the user due to amplitude accumulation, the common portion of the first audio signal and the common base signal is attenuated. This can reduce the aforementioned accumulation problem. The amplitude of this signal attenuation can be greater than 12dB. Alternatively, the amplitude of the audio portion corresponding to the common portion in the first audio signal can be completely removed, i.e., 100% removal.
[0090] As one implementation method, to enhance the spatial feel of the sound, a sound field expansion operation can be added. For example, this sound field expansion operation can include converting a mono audio signal to stereo. For instance, if the aforementioned method is applied to mono mode, the sound field expansion operation can be a stereo conversion operation, that is, converting a mono audio signal to stereo. Channel balance, delay, and frequency processing can be used to create a sense of space. Additionally, the sound field expansion operation can also include reverb and delay, i.e., adding reverb to simulate the reflection and diffusion of sound in different environments, increasing the depth and richness of the audio, and using delay effects to create sound echoes, thereby enhancing the spatial feel. Furthermore, the sound field expansion operation can also use phase processing techniques to increase the stereo feel and depth of the sound field.
[0091] In this embodiment, the sound field expansion operation can be phase expansion, i.e., analyzing the phase difference of the sound source and further expanding the phase difference to improve the sense of space. Specifically, if there are two first audio acquisition components, namely a left channel audio acquisition component and a right channel audio acquisition component, the electronic device can simultaneously acquire the initial left channel audio signal and the initial right channel audio signal, and then determine the phase difference between them. By performing phase expansion operations on the initial left channel audio signal and the initial right channel audio signal respectively, the phase difference between them can be increased, thereby improving the sense of space in the sound.
[0092] It should be noted that if the preset operation includes signal attenuation operation and sound field expansion operation, the execution order of the two operations is not limited. In one embodiment, the signal attenuation operation can be performed first and then the sound field expansion operation can be performed.
[0093] Taking the aforementioned left and right channel mode as an example, assuming there are two first audio acquisition components, namely a left channel audio acquisition component and a right channel audio acquisition component, the implementation method described above is as follows:
[0094] During audio acquisition, the initial left-channel audio signal acquired by the left-channel audio acquisition component, the initial right-channel audio signal acquired by the right-channel audio acquisition component, and the second audio signal acquired by the second audio acquisition component are obtained. A common base signal is obtained based on the initial left-channel audio signal, the initial right-channel audio signal, and the second audio signal. If the left-channel audio acquisition component malfunctions, the common base signal is used as the target audio signal corresponding to the left-channel audio acquisition component, i.e., the left-channel output audio signal. If the left-channel audio acquisition component is not malfunctioning, a preset operation is performed on the initial left-channel audio signal, and the common base signal and the initial left-channel audio signal after the preset operation are combined to obtain the left-channel output audio signal. If the right-channel audio acquisition component malfunctions, the common base signal is used as the target audio signal corresponding to the right-channel audio acquisition component, i.e., the right-channel output audio signal. If the right-channel audio acquisition component is not malfunctioning, a preset operation is performed on the initial right-channel audio signal, and the common base signal and the initial right-channel audio signal after the preset operation are combined to obtain the right-channel output audio signal.
[0095] Regarding the left and right channel modes mentioned above, assuming there is one first audio acquisition component, then one of the first audio acquisition component and the other of the second audio acquisition component is the left channel audio acquisition component and the other is the right channel audio acquisition component. The specific application process can be found in the previous content, and will not be repeated here.
[0096] For example, the above audio processing method will be further explained below with reference to the hardware structure of the electronic device performing the above audio processing method, specifically for the acquisition of the left channel signal and the right channel signal.
[0097] Figure 9 shows a structural block diagram of an electronic device provided in an embodiment of this application. The electronic device 100 can be a smartphone, tablet computer, e-reader, or other electronic device capable of running applications. The electronic device 100 in this application may include one or more of the following components: an audio processing module 101, an audio acquisition component 102, and one or more applications. The one or more applications may be stored in a memory and configured to be executed by the audio processing module 101. The one or more applications are configured to perform the methods described in the foregoing method embodiments. The implementation of the audio processing module 101 and the audio acquisition component 102 can be referred to the foregoing content and will not be repeated here.
[0098] As shown in Figure 10, both the first audio acquisition component and the second audio acquisition component are connected to the audio processing module 101, which executes the above-described method. As shown in Figure 10, the first audio acquisition component is connected to the audio processing module 101 via a first analog-to-digital converter (ADC1), and the second audio acquisition component is connected to the audio processing module 101 via a second analog-to-digital converter (ADC2). In other words, the signal processed by the audio processing module 101 is the digital signal obtained after analog-to-digital conversion of the audio signal acquired by the first and second audio acquisition components. This means that both the first and second audio signals are digital audio signals.
[0099] As shown in Figure 11, the audio processing module 1100 includes a detection module 1101, a sound mixing module 1103, and a common signal module 1102. The common signal module 1102 is connected to the first audio acquisition component and the second audio acquisition component, respectively, and is used to acquire a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component, and to obtain a common base signal based on the first and second audio signals. The detection module 1101 is connected to the first audio acquisition component and the second audio acquisition component, respectively, and is used to perform anomaly detection on the first audio acquisition component based on the first and second audio signals, and send the detection result to the sound mixing module 1103. The sound mixing module 1103 is connected to the first audio acquisition component, the detection module 1101, and the common signal module 1102, respectively. The sound mixing module 1103 is used to, when the detection result indicates an anomaly in the first audio acquisition component, use the common base signal as the target audio signal corresponding to the first audio acquisition component.
[0100] In other words, the detection module 1101 not only sends the detection result to the common signal module 1102, but also sends the first audio signal and the second audio signal to the common signal module 1102. As one implementation, the detection result corresponds to a control signal, which includes a first signal and a second signal. If the first audio acquisition component is detected to be abnormal, the first signal is sent to the common signal module 1102; if the first audio acquisition component is detected to be normal, the second signal is sent to the common signal module 1102. Then, the common signal module sends the control signal and the common basic signal to the sound mixing module 1103. If the sound mixing module 1103 determines that the control signal is the first signal and therefore the detection result indicates that the first audio acquisition component is abnormal, then it uses the common basic signal as the target audio signal corresponding to the first audio acquisition component. If the sound mixing module 1103 determines that the control signal is the second signal and therefore the detection result indicates that the first audio acquisition component is normal, then it synthesizes the target audio signal corresponding to the first audio acquisition component based on the common basic signal and the first audio signal.
[0101] As mentioned earlier, in some scenarios, specifically in the field of left and right channel audio acquisition, the target audio signal includes both the left and right channel audio signals. As shown in Figure 12, the first audio acquisition component is used to acquire either the left or right channel audio signal. There are two first audio acquisition components: a left channel audio acquisition component and a right channel audio acquisition component. The detection module 1101 is connected to the left channel audio acquisition component, the right channel audio acquisition component, and the second audio acquisition component, respectively. Based on the second audio signal and the initial left channel audio signal acquired by the left channel audio acquisition component, it obtains the detection result of the anomaly detection of the left channel audio acquisition component and sends it to the sound mixing module. Similarly, based on the second audio signal and the initial right channel audio signal acquired by the right channel audio acquisition component, it obtains the detection result of the anomaly detection of the right channel audio acquisition component and sends it to the sound mixing module.
[0102] The sound mixing module includes a left channel mixing module 1201 and a right channel mixing module 1202.
[0103] The left channel mixing module 1201 is connected to the left channel audio acquisition component and the common signal module 1102, and is used to use the common basic signal as the left channel output audio signal corresponding to the left channel audio acquisition component when it is determined that the left channel audio acquisition component is abnormal.
[0104] The right channel mixing module 1202 is connected to the right channel audio acquisition component and the common signal module 1102, and is used to use the common basic signal as the right channel output audio signal corresponding to the right channel audio acquisition component when it is determined that the right channel audio acquisition component is abnormal.
[0105] In one implementation, the left channel audio acquisition component is located in the top area of the electronic device, i.e., the top microphone; the right channel audio acquisition component is located in the bottom area of the electronic device, i.e., the bottom microphone; and the second audio acquisition component is located in the rear area of the electronic device, i.e., the rear microphone.
[0106] Therefore, through the above embodiments, even if either the top or bottom microphone is blocked, or both are blocked simultaneously, as long as the rear microphone is not blocked, sound output from both left and right channels will continue uninterrupted, greatly reducing the sudden changes caused by the blockage.
[0107] Additionally, it should be noted that the above-mentioned audio acquisition process, i.e. the application scenarios of the various embodiments of this application, can be video recording processes, and further, can be scenarios where video recording operations are performed in the landscape state of an electronic device. Of course, the specifics are not limited here.
[0108] Please refer to Figure 13, which shows a structural block diagram of an audio processing device 1300 provided in an embodiment of this application. The device is applied to an electronic device, which includes multiple audio acquisition components, including a first audio acquisition component and a second audio acquisition component. The device may include: an acquisition unit 1301, a synthesis unit 1302, and an execution unit 1303.
[0109] The acquisition unit 1301 is used to acquire, during the audio acquisition process, a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component.
[0110] Synthesis unit 1302 is used to obtain a common base signal based on the first audio signal and the second audio signal.
[0111] The execution unit 1303 is configured to, if it is determined that the first audio acquisition component is abnormal based on the first audio signal and the second audio signal, use the common basic signal as the target audio signal corresponding to the first audio acquisition component.
[0112] Furthermore, the execution unit 1303 is also configured to, if it is determined that the first audio acquisition component is not abnormal based on the first audio signal and the second audio signal, synthesize the target audio signal corresponding to the first audio acquisition component based on the common base signal and the first audio signal.
[0113] Furthermore, the execution unit 1303 is also used to perform a preset operation on the first audio signal; and to synthesize the common basic signal and the first audio signal after performing the preset operation to obtain a target audio signal corresponding to the first audio acquisition component.
[0114] Furthermore, the signal attenuation operation includes identifying the same portion of the first audio signal and the common base signal as the target portion, and performing a signal attenuation operation on the target portion of the first audio signal.
[0115] Furthermore, there are two first audio acquisition components, namely a left channel audio acquisition component and a right channel audio acquisition component. The first audio signal acquired by the left channel audio acquisition component is the initial left channel audio signal, and the first audio signal acquired by the right channel audio acquisition component is the initial right channel audio signal. The target audio signal corresponding to the left channel audio acquisition component is the left channel output audio signal, and the target audio signal corresponding to the right channel audio acquisition component is the right channel output audio signal.
[0116] Furthermore, the synthesis unit 1302 is also used to obtain a common basic signal by weighted fusion of the initial left channel audio signal, the initial right channel audio signal, and the second audio signal.
[0117] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0118] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0119] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0120] Please refer to Figure 14, which shows a structural block diagram of a computer-readable medium provided in an embodiment of this application. The computer-readable medium 1400 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0121] Computer-readable medium 1400 may be an electronic storage device such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, computer-readable medium 1400 includes non-volatile computer-readable storage medium. Computer-readable medium 1400 has storage space for program code 1410 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. Program code 1410 may be compressed, for example, in a suitable form.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An audio processing method, characterized in that, An electronic device is used, the electronic device including multiple audio acquisition components, the multiple audio acquisition components including a first audio acquisition component and a second audio acquisition component, the method including: during the audio acquisition process, acquiring a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component; obtaining a common base signal based on the first audio signal and the second audio signal; if it is determined that the first audio acquisition component is abnormal based on the first audio signal and the second audio signal, using the common base signal as the target audio signal corresponding to the first audio acquisition component.
2. The method according to claim 1, characterized in that, Also includes: If it is determined that the first audio acquisition component is not abnormal based on the first audio signal and the second audio signal, the target audio signal corresponding to the first audio acquisition component is synthesized based on the common base signal and the first audio signal.
3. The method according to claim 2, characterized in that, The method for synthesizing a target audio signal corresponding to the first audio acquisition component based on a common basic signal and a first audio signal includes: performing a preset operation on the first audio signal; and synthesizing the common basic signal and the first audio signal after performing the preset operation to obtain a target audio signal corresponding to the first audio acquisition component.
4. The method according to claim 3, characterized in that, The preset operation includes at least one of signal attenuation operation and sound field expansion operation.
5. The method according to claim 4, characterized in that, The signal attenuation operation includes identifying the common portion of the first audio signal and the common base signal as the target portion, and performing a signal attenuation operation on the target portion of the first audio signal.
6. The method according to any one of claims 1-5, characterized in that, The first audio acquisition component consists of two components: a left channel audio acquisition component and a right channel audio acquisition component. The first audio signal acquired by the left channel audio acquisition component is the initial left channel audio signal, and the first audio signal acquired by the right channel audio acquisition component is the initial right channel audio signal. The target audio signal corresponding to the left channel audio acquisition component is the left channel output audio signal, and the target audio signal corresponding to the right channel audio acquisition component is the right channel output audio signal.
7. The method according to claim 6, characterized in that, The common base signal is obtained based on the first audio signal and the second audio signal, including: obtaining the common base signal by weighted fusion of the initial left channel audio signal, the initial right channel audio signal and the second audio signal.
8. An audio processing apparatus, characterized in that, An electronic device is used, the electronic device including multiple audio acquisition components, the multiple audio acquisition components including a first audio acquisition component and a second audio acquisition component, the device including: an acquisition unit, used to acquire a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component during the audio acquisition process; a synthesis unit, used to obtain a common base signal based on the first audio signal and the second audio signal; and an execution unit, used to, if it is determined based on the first audio signal and the second audio signal that the first audio acquisition component is abnormal, use the common base signal as the target audio signal corresponding to the first audio acquisition component.
9. An electronic device, characterized in that, include: Multiple audio acquisition components, including a first audio acquisition component and a second audio acquisition component; An audio processing module is connected to the first audio acquisition component and the second audio acquisition component, respectively, and the audio processing module is used to perform the method as described in any one of claims 1-7.
10. The electronic device according to claim 9, characterized in that, The audio processing module includes a sound mixing module, a common signal module, and a detection module. The common signal module is connected to the first audio acquisition component and the second audio acquisition component, respectively, and is used to acquire a first audio signal acquired by the first audio acquisition component and a second audio signal acquired by the second audio acquisition component, and obtain a common basic signal based on the first audio signal and the second audio signal. The detection module is connected to the first audio acquisition component and the second audio acquisition component, respectively, and is used to perform anomaly detection on the first audio acquisition component based on the first audio signal and the second audio signal, and send the detection result to the sound mixing module. The sound mixing module is connected to the first audio acquisition component, the detection module and the common signal module respectively. The sound mixing module is used to take the common basic signal as the target audio signal corresponding to the first audio acquisition component when the detection result is that the first audio acquisition component is abnormal.
11. The electronic device according to claim 10, characterized in that, The first audio acquisition component consists of two parts: a left channel audio acquisition component and a right channel audio acquisition component. The detection module is connected to the left channel audio acquisition component, the right channel audio acquisition component, and the second audio acquisition component. It is used to obtain the detection result of the anomaly detection of the left channel audio acquisition component based on the second audio signal and the initial left channel audio signal acquired by the left channel audio acquisition component, and send it to the sound mixing module. It also obtains the detection result of the anomaly detection of the right channel audio acquisition component based on the second audio signal and the initial right channel audio signal acquired by the right channel audio acquisition component, and sends it to the sound mixing module. The sound mixing module includes a left channel mixing module and a right channel mixing module; the left channel mixing module is connected to the left channel audio acquisition component and the common signal module, and is used to use the common basic signal as the left channel output audio signal corresponding to the left channel audio acquisition component when it is determined that the left channel audio acquisition component is abnormal. The right channel mixing module is connected to the right channel audio acquisition component and the common signal module, and is used to use the common basic signal as the right channel output audio signal corresponding to the right channel audio acquisition component when it is determined that the right channel audio acquisition component is abnormal.
12. The electronic device according to claim 11, characterized in that, The left channel audio acquisition component is located in the top area of the electronic device, the right channel audio acquisition component is located in the bottom area of the electronic device, and the second audio acquisition component is located in the rear area of the electronic device.