Audio signal processing methods, devices, terminal equipment and storage media
By using multiple audio acquisition units to acquire and superimpose audio signals in terminal devices and filtering noise, the problem of audio signal interference in multi-sound-source environments is solved, thereby improving the quality of audio signals and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING XIAOMI MOBILE SOFTWARE CO LTD
- Filing Date
- 2023-08-14
- Publication Date
- 2026-05-26
Smart Images

Figure CN119497017B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of audio signal processing technology, and in particular to an audio signal processing method, apparatus, terminal device and storage medium. Background Technology
[0002] With the development of audio technology, more and more terminal devices have audio signal processing functions. Terminal devices with audio signal processing functions usually have microphones, which can collect audio signals from the environment in which the terminal device is located.
[0003] In certain application scenarios, when multiple sound sources exist in the environment where the terminal device is located, the microphone will collect various sounds from the current environment. For example, if multiple people are talking in the current scene, the microphone can collect the voices of each person. Some sounds may be required according to business needs, while others may not be required for this business scenario. Summary of the Invention
[0004] This disclosure provides an audio signal processing method, apparatus, terminal device, and storage medium.
[0005] A first aspect of this disclosure provides an audio signal processing method applied to a terminal device. The method includes: acquiring a first audio signal through a first acquisition unit of the terminal device and acquiring a second audio signal through a second acquisition unit of the terminal device; wherein the distance difference between the first acquisition unit and the second acquisition unit and a first sound source is within a first range; the distance difference between the first acquisition unit and the second acquisition unit and the second sound source is greater than the maximum value of the first range, and the first sound source is a target acquisition object; superimposing the first audio signal and the second audio signal in the time domain to obtain a superimposed audio signal; filtering out noise audio signals with amplitudes less than a preset value from the superimposed audio signal to obtain a target audio signal of the target acquisition object.
[0006] In one embodiment, the terminal device includes a rectangular frame and a back cover; the rectangular frame includes a set of parallel long sides and a set of parallel short sides; the long sides and the short sides are connected end to end; when the terminal device is in landscape mode, the sensing surfaces of the first sensing unit and the second sensing unit are located at different positions on the same long side; when the terminal device is in portrait mode, the sensing surfaces of the first sensing unit and the second sensing unit are located on different sides of the rectangular frame; when the terminal device is in a tilted or vertical position, the sensing surface of the first sensing unit is located on the rectangular frame, and the sensing surface of the second sensing unit is located on the back cover.
[0007] In one embodiment, the method further includes: determining a collection scenario; setting audio processing parameters of the terminal device according to the current collection scenario; wherein the audio processing parameters are at least used to determine the collection unit and / or collection parameters for audio collection in the collection scenario;
[0008] The acquisition scenario includes a first scenario; wherein the first scenario is a scenario of audio acquisition from a single sound source; in the first scenario, the audio processing parameters are used at least to determine the audio acquired by the first acquisition unit and the second acquisition unit and to filter it.
[0009] In one embodiment, the audio processing parameters corresponding to the first scenario include: the preset value; first direction information, used to indicate the sound acquisition direction of the first acquisition unit and the second acquisition unit; and first noise reduction parameters, used for ambient sound noise reduction in the first scenario.
[0010] In one embodiment, the acquisition scenario further includes a second scenario, wherein, in the second scenario, audio signals from multiple sound sources are acquired through at least one audio acquisition unit in the terminal device.
[0011] In one embodiment, the audio processing parameters corresponding to the second scene include at least one of the following: acquisition unit information, used to determine the acquisition unit for acquiring multiple sound sources; second direction information, used to acquire the acquisition direction of the multiple sound sources; distance information, used to determine the acquisition distance of the sound sources; and second noise reduction parameters, used for ambient sound noise reduction in the second scene.
[0012] In one embodiment, the acquisition scenario further includes a third scenario, wherein the third scenario is used to acquire audio based on an audio acquisition peripheral independent of the terminal device.
[0013] In one embodiment, the audio processing parameters corresponding to the second scenario include at least one of the following: a first indication information for indicating the use of an audio acquisition peripheral connected to the terminal device to acquire audio signals; a third noise reduction parameter for indicating the disabling of the environmental noise reduction function when using the acquisition unit on the terminal device to acquire audio; and a second indication information for indicating the acquisition unit of the terminal device to acquire ambient sound.
[0014] In one embodiment, determining the acquisition scenario includes: determining the acquisition scenario based on input operations performed on the terminal device; or, determining the acquisition scenario based on whether there is an audio acquisition peripheral connected to the terminal device.
[0015] A third aspect of this disclosure provides an audio signal processing apparatus, comprising:
[0016] The acquisition module is used to acquire a first audio signal through a first acquisition unit and a second audio signal through a second acquisition unit; wherein the distance difference between the first acquisition unit and the second acquisition unit and the first sound source is within a first range; the distance difference between the first acquisition unit and the second acquisition unit and the second sound source is greater than the maximum value of the first range, and the first sound source is the target acquisition object; the superposition module is used to superimpose the first audio signal and the second audio signal in the time domain to obtain a superimposed audio signal; the filtering module is used to filter out noise audio signals with amplitudes less than a preset value in the superimposed audio signal to obtain the target audio signal of the target acquisition object.
[0017] A third aspect of this disclosure provides a terminal device, including:
[0018] A processor and a memory for storing executable instructions that can run on the processor, wherein: when the processor runs the executable instructions, the executable instructions perform the method described in any of the above embodiments.
[0019] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the method described in any of the above embodiments.
[0020] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0021] The audio signal processing method in this embodiment acquires different audio signals through different acquisition units. For example, a first acquisition unit acquires a first audio signal, and a second acquisition unit acquires a second audio signal. After two sound sources emit sound, the two acquisition units can respectively acquire the audio signals emitted by the two sound sources. Since the distance difference between the two acquisition units and the first sound source is within a first range, the audio signals of the first sound source acquired by the two acquisition units have a small difference in the time domain. After superimposing the audio signals of the first sound source acquired by the two acquisition units, the audio signal can be enhanced and the amplitude increased. Since the distance difference between the two acquisition units and the second sound source is greater than the maximum value of the first range, the audio signals of the second sound source acquired by the two acquisition units have a large difference in the time domain, and the amplitude of the superimposed signal will not increase significantly. According to a preset value, noise signals with amplitudes smaller than the preset value in the superimposed audio signal can be filtered out. These noise signals include the audio signal emitted by the second sound source. In this way, the audio signal of the second sound source is filtered out from the audio signals acquired by the two acquisition units, and the audio signal emitted by the first sound source is obtained. By reducing the interference of the sound emitted by the second sound source on the audio signal emitted by the desired first sound source, the target audio signal of the target acquisition object is obtained. This reduces noise in the target audio signal, improves the quality of the target audio signal, and thus enhances the user experience.
[0022] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0024] Figure 1 This is a schematic diagram illustrating an audio signal processing method according to an exemplary embodiment;
[0025] Figure 2 This is a schematic diagram illustrating the distribution of an audio signal acquisition unit according to an exemplary embodiment;
[0026] Figure 3 This is a schematic diagram illustrating the distribution of another audio signal acquisition unit according to an exemplary embodiment;
[0027] Figure 4 This is a schematic diagram illustrating the determination of an application scenario according to an exemplary embodiment;
[0028] Figure 5 This is a schematic diagram illustrating the determination of a data acquisition scenario according to an exemplary embodiment;
[0029] Figure 6This is an example of a method based on an exemplary embodiment. Figure 5 A schematic diagram of the interface for determining the data collection scenario;
[0030] Figure 7 This is a schematic diagram of an audio signal processing apparatus according to an exemplary embodiment;
[0031] Figure 8 This is a schematic diagram illustrating an audio signal processing method according to an exemplary embodiment;
[0032] Figure 9 This is a schematic diagram illustrating an audio signal processing method according to an exemplary embodiment;
[0033] Figure 10 This is a block diagram illustrating a terminal device according to an exemplary embodiment. Detailed Implementation
[0034] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses consistent with some aspects of this disclosure as detailed in the appended claims.
[0035] Voice over Internet Protocol (VoIP) calls have gradually replaced SIM card-based calls, becoming the primary method of information transmission in daily life and work, such as online meetings and instant messaging. The quality of VoIP calls is a crucial evaluation criterion for the call quality of mobile phones or tablets. However, in real-world scenarios, call quality is often affected by ambient noise. For example, when a user participates in an online meeting, their voice is transmitted to the receiving end along with surrounding noise and other people's voices, resulting in a noisy experience for the receiving user and degrading the call experience.
[0036] Typically, audio acquisition units (such as microphones) on terminal devices such as tablets and mobile phones can perform 360-degree omnidirectional sound pickup. However, when used in noisy environments, other sounds in the environment and the sound that is expected to be acquired will be picked up by the microphone of the terminal device. Even if the noise reduction algorithm of the terminal device suppresses ambient sounds, other sounds in the environment will still be retained, causing the sound that is expected to be acquired to be interfered with or even covered by other sounds in the environment.
[0037] For example, when a user participates in a video conference at home, the microphone picks up sound from all directions (360 degrees) while the user is speaking. The sound of other family members speaking nearby is also picked up by the microphone and transmitted to the receiving end. This results in a noisy sound received by the receiving end, affecting the user's ability to obtain the sound they need, reducing call quality, and also reducing the user's privacy.
[0038] refer to Figure 1 This diagram illustrates an audio signal processing method. This method can be executed in at least a terminal device with an audio signal acquisition unit. The terminal device can include mobile terminal devices and fixed terminal devices, meaning the executing entity of this method can include at least mobile terminal devices and fixed terminal devices. Mobile terminal devices can include mobile phones, tablets, and wearable devices.
[0039] The audio signal processing method includes:
[0040] S100: A first audio signal is acquired through the first acquisition unit of the terminal device, and a second audio signal is acquired through the second acquisition unit of the terminal device; wherein, the distance difference between the first acquisition unit and the second acquisition unit and the first sound source is within a first range; the distance difference between the first acquisition unit and the second acquisition unit and the second sound source is greater than the maximum value of the first range, and the first sound source is the target acquisition object.
[0041] S200: Superimpose the first audio signal and the second audio signal in the time domain to obtain the superimposed audio signal.
[0042] S300: Filters out noise audio signals with amplitudes less than a preset value from the superimposed audio signal to obtain the target audio signal of the target acquisition object.
[0043] In the terminal device of this embodiment, the terminal device has multiple audio signal acquisition units, referred to simply as acquisition units, and the positions of each acquisition unit are different. For example, it may include 3, 4, or 5 acquisition units, etc. The acquisition units may be located on the edge of the terminal device or on the back of the terminal device, i.e., on the back cover. The terminal device may be a terminal device with a display screen, and the back of the terminal device may be the side opposite to the display screen.
[0044] For the S100, the audio signal acquisition unit may include a microphone, which may also be referred to as a pickup.
[0045] For example, acquisition unit 1 and acquisition unit 2 are located on the first edge of the terminal device, acquisition unit 3 is located on the second edge of the terminal device, and acquisition unit 4 is located on the rear shell of the terminal device. The first edge can be perpendicular to the second edge, or the first edge can be parallel to the second edge.
[0046] refer to Figure 2 This is a schematic diagram of the distribution of acquisition units. The acquisition units are located in... Figure 2 The omnidirectional microphone is used to indicate this. Figure 2 This is a front view of the terminal device, where the display screen is located. Omnidirectional microphones mic1 and mic2 are located at different positions on the same bezel of the terminal device. The terminal device has a rectangular bezel, and omnidirectional microphone mic4 is located on the other bezel. The two bezels are perpendicular and are two sides of the rectangular bezel of the terminal device.
[0047] For example, the acquisition surfaces of omnidirectional microphone mic1 and omnidirectional microphone mic2 are located on the same frame.
[0048] refer to Figure 3 This is a schematic diagram showing the distribution of another type of acquisition unit. The acquisition unit is located in... Figure 3 The omnidirectional microphone is used to indicate this. Figure 3 This is a rear view of the terminal device, i.e., a schematic diagram of the back cover. The omnidirectional microphone mic3 is located on the back cover.
[0049] For example, the acquisition surface of the omnidirectional microphone mic3 is located on the rear cover.
[0050] Figure 2 and Figure 3 The document also shows a speaker, a camera (such as a front-facing camera), and a USB port.
[0051] In one embodiment, Figure 2 and Figure 3 This is a diagram illustrating the terminal device in landscape orientation. Figure 2 Rotating 90 degrees clockwise creates a diagram illustrating the vertical screen orientation of the terminal device. Figure 3 Rotating 90 degrees counterclockwise creates a diagram illustrating the vertical screen orientation of the terminal device.
[0052] In one embodiment, the terminal device is a tablet computer or a mobile phone. For example, Figure 2 and Figure 3 This is a schematic diagram of a tablet computer.
[0053] There is a preset distance between the two different acquisition units, and the preset distance can be determined according to the design of the terminal equipment. For example, Figure 2 and Figure 3 The omnidirectional microphones mic1 and mic2 shown have a preset distance between them, and mic1 and mic2 can divide the first frame into three equal parts. The omnidirectional microphone mic4 can be located in the middle of the second frame. The first frame is perpendicular to the second frame.
[0054] For example, the first acquisition unit can be Figure 2 and Figure 3 The omnidirectional microphone mic1 in the middle, the second acquisition unit can be Figure 2 and Figure 3 The omnidirectional microphone mic2 is included.
[0055] The sound emitted by each sound source is collected by two of the multiple audio signal acquisition units in the terminal device. These two acquisition units are not fixed and can be determined according to the actual use scenario.
[0056] For example, the terminal device in Figure 2 In the shown state, audio signals emitted by the first sound source can be acquired using omnidirectional microphones mic1 and mic2, respectively, and audio signals emitted by the second sound source can be acquired using omnidirectional microphones mic1 and mic2, respectively. Omnidirectional microphone mic1 can be used as the first acquisition unit, and omnidirectional microphone mic2 can be used as the second acquisition unit.
[0057] For example, omnidirectional microphones mic1 and mic4 can respectively collect audio signals from various sound sources. Omnidirectional microphone mic1 can be used as the first acquisition unit, and omnidirectional microphone mic4 can be used as the second acquisition unit.
[0058] For example, terminal devices in Figure 3 In the state shown, any one of the omnidirectional microphones mic1, mic2, and mic4 can be used as the first acquisition unit, and mic3 can be used as the second acquisition unit.
[0059] The audio signal acquired by the first acquisition unit is recorded as the first audio signal, and the audio signal acquired by the second acquisition unit is recorded as the second audio signal.
[0060] For example, the second sound source is located at Figure 2 If a second sound source emits sound 1 from either the left or right side of the terminal device shown, both the first acquisition unit (e.g., omnidirectional microphone mic1) and the second acquisition unit (e.g., omnidirectional microphone mic2) can acquire this sound 1. The sound 1 acquired by the first audio signal is denoted as the first audio signal, and the sound 1 acquired by the second audio signal is denoted as the second audio signal.
[0061] Since the distances between the first acquisition unit and the second acquisition unit and the second sound source are different, the time it takes for sound 1 to be transmitted to the first acquisition unit and the second acquisition unit are different, and the time it takes for the first acquisition unit and the second acquisition unit to acquire sound 1 are also different.
[0062] For example, the direction of the first sound source Figure 2If the front of the terminal device is shown and the distance between it and the omnidirectional microphones mic1 and mic2 is equal, then the sound 2 emitted by the first sound source will be transmitted to the omnidirectional microphones mic1 and mic2 in the same amount of time, and the omnidirectional microphones mic1 and mic2 will also collect the sound 2 in the same amount of time.
[0063] For example, sound source 2 can be the user of the terminal device.
[0064] When a sound source emits a sound, each acquisition unit on the terminal device can capture that sound. Similarly, when different sound sources emit sounds, each acquisition unit can capture the sounds emitted by each of those sources. The sound emitted by each sound source is within the acquisition range of each acquisition unit.
[0065] In one embodiment, there are two sound sources: a first sound source and a second sound source. Audio signals emitted by these two sound sources are acquired by a first acquisition unit and a second acquisition unit. The first sound source is located at a distance from both the first and second acquisition units, and the difference between these distances falls within a first range. The first sound source is the target of the acquisition, and the final audio signal obtained is the audio signal emitted by the first sound source. The second sound source is also located at a distance from both the first and second acquisition units, and the difference between these distances is greater than the maximum value of the first range.
[0066] The first range can be determined based on actual usage requirements; it can be a distance range measured in centimeters or millimeters. Specific numerical values are not limited.
[0067] A first audio signal can be acquired through the first acquisition unit, and a second audio signal can be acquired through the second acquisition unit. The first and second audio signals can be audio signals from the same sound source. For example, it can be determined whether the first and second audio signals are audio signals from the same sound source based on the spectral analysis of the audio signals.
[0068] For example, the first audio signal and the second audio signal can be audio signals emitted by the first sound source, and the first audio signal and the second audio signal can also be audio signals emitted by the second sound source.
[0069] After acquiring the first audio signal and the second audio signal, the spectral information of the first audio signal in the time domain and the spectral information of the second audio signal in the time domain can be determined.
[0070] The greater the distance difference between the sound source and the two acquisition units, the greater the difference in the time domain between the audio signals emitted by the sound source acquired by the two acquisition units, and the farther apart the amplitude peaks are. Conversely, the smaller the distance difference between the sound source and the two acquisition units, the smaller the difference in the time domain between the audio signals emitted by the sound source acquired by the two acquisition units, and the closer the amplitude peaks are.
[0071] For S200, after obtaining the first audio signal and the second audio signal, the first audio signal and the second audio signal are superimposed in the time domain to obtain a superimposed audio signal. When the spectral information of the first audio signal and the second audio signal differs greatly in the time domain, the amplitude overlap of the first audio signal and the second audio signal is low, and the amplitude of the superimposed audio signal may be small, resulting in a small enhancement effect.
[0072] When the spectral information of the first and second audio signals differs little in the time domain, the amplitudes of the first and second audio signals overlap significantly, resulting in a larger amplitude of the superimposed audio signal and a greater enhancement effect. The amplitude of the superimposed audio signal includes the maximum amplitude.
[0073] For the S300, the superimposed audio signal is filtered using a preset value. Audio signals with amplitudes smaller than the preset value are treated as noise signals and filtered to obtain the target audio signal. This filters out sounds from sources with large distance differences between the sound source and the two acquisition units, such as sounds from sources on the left, right, or rear of the terminal device, thus obtaining the audio signal emitted by the target acquisition object. This reduces noise in the obtained audio signal and improves the quality of the target audio signal.
[0074] Because the distance difference between the two acquisition units and the first sound source is within a certain range, the audio signals from the first sound source acquired by the two acquisition units have a small difference in the time domain. Superimposing these audio signals enhances the audio signal and increases its amplitude. However, because the distance difference between the two acquisition units and the second sound source is greater than the maximum value of the first range, the audio signals from the second sound source acquired by the two acquisition units have a larger difference in the time domain, and the amplitude of the superimposed signal does not increase significantly. Based on a preset value, noise signals with amplitudes smaller than the preset value in the superimposed audio signal can be filtered out. This noise signal includes the audio signal emitted by the second sound source. This filters out the audio signal of the second sound source from the audio signals received by the two acquisition units, resulting in the audio signal emitted by the first sound source. Reducing the interference of the second sound source on the desired audio signal from the first sound source yields the target audio signal for the target acquisition object. This reduces noise in the target audio signal, improves its quality, and thus enhances the user experience.
[0075] In one embodiment, the terminal device includes a rectangular frame and a rear shell. The rectangular frame includes a set of parallel long sides and a set of parallel short sides; the long sides and short sides are connected end to end.
[0076] refer to Figure 2 and Figure 3 The largest rectangle shown is the border, consisting of two parallel long sides and two parallel short sides. The long and short sides meet at their ends to form the rectangular border. For example, two can be grouped together.
[0077] refer to Figure 2 and Figure 3 When the terminal device is in landscape mode, the acquisition surfaces of the first and second acquisition units are located at different positions along the same long side. For example, Figure 2 The topmost border shown is the long border. The acquisition surface of the first acquisition unit and the second acquisition unit are located at different positions on the top border, that is, the omnidirectional microphone mic1 and the omnidirectional microphone mic2 are located on the same long border.
[0078] For example, omnidirectional microphones mic1 and omnidirectional microphone mic2 can also be located on the bottom long edge border.
[0079] For example, omnidirectional microphone mic1 and omnidirectional microphone mic2 can also have one located on the long side border and the other on the short side border.
[0080] For example, in real-world application scenarios, such as when using instant messaging software for voice or video communication, the terminal device can be in landscape orientation.
[0081] In one embodiment, when the terminal device is in portrait orientation, the acquisition surface of the first acquisition unit and the acquisition surface of the second acquisition unit are located on different sides of the rectangular frame.
[0082] For example, the acquisition surface of the first acquisition unit is located on the long side border, and the acquisition surface of the second acquisition unit is located on the short side border. Figure 2 The omnidirectional microphone mic1 is located on the long side border, and the omnidirectional microphone mic4 is on the short side border.
[0083] For example, the acquisition surface of the first acquisition unit is located on the first long side border, and the acquisition surface of the second acquisition unit is located on the other long side border.
[0084] In one embodiment, when the terminal device is tilted or in a vertical position, the acquisition surface of the first acquisition unit is located on the rectangular frame, and the acquisition surface of the second acquisition unit is located on the back cover.
[0085] For example, the acquisition surface of the first acquisition unit is located on the border of any long side, and the acquisition surface of the second acquisition unit is located on the rear shell. Figure 3 The omnidirectional microphone mic1 is located on the long side frame, and the omnidirectional microphone mic3 is located on the back cover.
[0086] In one embodiment, reference Figure 4 This is a schematic diagram illustrating the determination of an application scenario. The method further includes:
[0087] S10, Determine the data collection scenario;
[0088] S20, based on the current acquisition scenario, set the audio processing parameters of the terminal device; wherein, the audio processing parameters are at least used to determine the acquisition unit and / or acquisition parameters for audio acquisition in the acquisition scenario.
[0089] The acquisition scenarios include the first scenario; the first scenario is a scenario in which audio is acquired from a single sound source.
[0090] In the first scenario, the audio processing parameters are used at least to determine the audio acquired by the first acquisition unit and the second acquisition unit and to filter it.
[0091] In S10, the data collection scenario is determined, including:
[0092] The acquisition scenario is determined based on the input operation performed on the terminal device; or, the acquisition scenario is determined based on whether there is an audio acquisition peripheral connected to the terminal device.
[0093] Different input operations determine different acquisition scenarios, and different acquisition scenarios correspond to different audio processing parameters. When the audio processing parameters are different, the acquisition method of the audio signal acquired by the acquisition unit is different, the processing method of the acquired audio signal is also different, and the acquisition unit for acquiring the audio signal will also be different. The acquisition method can be determined according to the acquisition parameters; when the acquisition parameters are different, the acquisition method can be different.
[0094] The acquisition scenario can also be determined based on whether the terminal device is connected to an audio acquisition peripheral. If an audio acquisition peripheral is connected, the acquisition scenario related to the acquisition of audio signals by the audio acquisition peripheral can be determined, such as the scenario of acquiring audio signals through the audio acquisition peripheral.
[0095] refer to Figure 5 This is a schematic diagram of an interface for defining a data collection scenario. (Refer to...) Figure 6 For based on Figure 5 A schematic diagram of the interface for determining the data collection scenario.
[0096] Figure 5The diagram illustrates the interface of a VoIP-based application on a terminal device, such as an application with conferencing capabilities, including a conferencing toolbox interface that includes call noise reduction controls. This interface is displayed after the application is launched. For example, the conferencing toolbox could also be an application itself, with an interface as shown below. Figure 5 As shown. The result can be obtained based on the operation performed on the call noise reduction control. Figure 6 The interface shown.
[0097] Figure 6 for Figure 5 A detailed schematic diagram of the call noise reduction interface is shown. Figure 6 The interface for defining the acquisition scene is shown, including single-person scenes, multi-person scenes, microphone noise reduction scenes, and speaker / headphone noise reduction scenes. For example... Figure 6 As shown, the call noise reduction interface includes controls for enabling / disabling microphone noise reduction and scene controls, which can determine the acquisition scene based on input operations. The first scene can include a single-person scene, which is a scene for audio acquisition from a single sound source. In the single-person scene, the microphone picks up the sound directly in front of the screen.
[0098] After selecting a single-person acquisition scenario, in the first scenario, the audio processing parameters are used to determine at least the audio acquired by the first acquisition unit and the second acquisition unit and to filter it. The target audio signal can be determined according to S100 to S300.
[0099] In one embodiment, the audio processing parameters corresponding to the first scene include at least one of the following:
[0100] Preset values for filtering superimposed audio signals;
[0101] The first direction information is used to indicate the sound acquisition direction of the first acquisition unit and the second acquisition unit;
[0102] The first noise reduction parameter is used for ambient sound noise reduction in the first scenario.
[0103] The acquisition unit information is used to determine the first acquisition unit and the second acquisition unit.
[0104] The preset value can be determined according to the actual scenario, and can be an integer or a decimal in decibels.
[0105] The first directional information can include directions 30 degrees to the left and right of the front of the terminal device, indicating a 60-degree range within which the sound is collected. Of course, it can also be other directions, indicating angles deviating from the front.
[0106] The first noise reduction parameter may include noise reduction-related parameters such as maximum noise reduction depth, noise reduction coefficient, noise reduction intensity, and noise reduction radius.
[0107] The information collected can be the identification information of the first and second collection units.
[0108] In one embodiment, the data collection scenario further includes:
[0109] In the second scenario, audio signals from multiple sound sources are acquired through at least one audio acquisition unit in the terminal device.
[0110] The second scenario differs from the first scenario in that it uses one or more acquisition units to collect audio signals emitted by multiple sound sources.
[0111] For example, according to the action on Figure 6 The multi-person scene operations shown on the interface confirm that the scene being collected is the second scene. The second scene includes... Figure 6 The scene shown is a multi-person scenario.
[0112] For example, in the second scenario, it may not be necessary to execute S100 to S300.
[0113] In one embodiment, the audio processing parameters corresponding to the second scene include at least one of the following:
[0114] Acquisition unit information, used to determine the acquisition unit for acquiring multiple sound sources;
[0115] The second direction information is used to determine the acquisition direction for sound acquisition from the multiple sound sources.
[0116] Distance information is used to determine the sampling distance to the sound source;
[0117] The second noise reduction parameter is used for ambient sound noise reduction in the second scenario. Refer to the content included in the first noise reduction parameter for details.
[0118] The acquisition unit information can be the identification information of the acquisition unit that acquires multiple sound sources.
[0119] The second direction information can be similar to the first direction information, but the specific direction indicated can be different. The second direction information can indicate the audio signals emitted by multiple sound sources within a 360-degree radius around the acquisition terminal device, with the acquisition direction being 360 degrees.
[0120] Distance information can be determined based on the performance parameters of the acquisition unit and / or call requirements, such as an area with a radius of 5 meters centered on the terminal device. The maximum acquisition distance is 5 meters.
[0121] Audio signals can be acquired based on the audio processing parameters corresponding to the second scene.
[0122] In one embodiment, the data collection scenario further includes:
[0123] The third scenario is used to acquire audio based on an audio acquisition peripheral that is independent of the terminal device.
[0124] refer to Figure 6 The headphone noise cancellation acquisition mode shown is the third acquisition scenario when the speaker / headphone noise cancellation switch is turned on. The third scenario differs from the first and second scenarios in that it does not acquire audio signals through the acquisition unit on the terminal device, but rather through an audio acquisition peripheral independent of the terminal device.
[0125] For example, audio signals around the earphones can be collected through wireless earphones that are independent of the terminal device, and then the earphones can send the collected audio signals to the terminal device.
[0126] In one embodiment, the audio processing parameters corresponding to the second scene include at least one of the following:
[0127] The first instruction information is used to instruct the use of an audio acquisition peripheral connected to the terminal device to acquire audio signals;
[0128] The third noise reduction parameter is used to indicate whether the environmental noise reduction function is turned off when using the acquisition unit on the terminal device to acquire audio.
[0129] The second instruction is used to instruct the terminal device's acquisition unit to collect ambient sound.
[0130] When the audio processing parameters corresponding to the second scenario include the first indication information, it can be determined whether to use an audio acquisition peripheral connected to the terminal device to acquire audio signals based on the first indication information.
[0131] The third noise reduction parameter can be referenced from the content included in the first noise reduction parameter.
[0132] When the audio processing parameters corresponding to the second scene include the second instruction information, ambient sound can be collected through the acquisition unit of the terminal device according to the second instruction information.
[0133] The first instruction information and the second instruction information can coexist, and audio signals can be acquired simultaneously through audio acquisition peripherals and acquisition units.
[0134] In one embodiment, reference Figure 7 This is a schematic diagram of an audio signal processing device, which includes:
[0135] Acquisition module 1 is used to acquire a first audio signal through a first acquisition unit and acquire a second audio signal through a second acquisition unit; wherein the distance difference between the first acquisition unit and the second acquisition unit and the first sound source is within a first range; the distance difference between the first acquisition unit and the second acquisition unit and the second sound source is greater than the maximum value of the first range, and the first sound source is the target acquisition object;
[0136] Superposition module 2 is used to superimpose the first audio signal and the second audio signal in the time domain to obtain a superimposed audio signal;
[0137] The filtering module 3 is used to filter out noise audio signals with amplitudes less than a preset value in the superimposed audio signal to obtain the target audio signal of the target acquisition object.
[0138] In one embodiment, the terminal device includes a rectangular frame and a back cover; the rectangular frame includes: a set of parallel long sides and a set of parallel short sides; the long sides and the short sides are connected end to end;
[0139] When the terminal device is in landscape orientation, the acquisition surface of the first acquisition unit and the acquisition surface of the second acquisition unit are located at different positions on the same long side;
[0140] When the terminal device is in portrait orientation, the acquisition surface of the first acquisition unit and the acquisition surface of the second acquisition unit are located on different sides of the rectangular frame;
[0141] When the terminal device is tilted or in a vertical position, the acquisition surface of the first acquisition unit is located on the rectangular frame, and the acquisition surface of the second acquisition unit is located on the back cover.
[0142] In one embodiment, the apparatus further includes:
[0143] The scene determination module is used to determine the data collection scene;
[0144] The parameter determination module is used to set the audio processing parameters of the terminal device according to the current acquisition scenario; wherein, the audio processing parameters are at least used to determine the acquisition unit and / or acquisition parameters for audio acquisition in the acquisition scenario;
[0145] The acquisition scenario includes a first scenario; wherein the first scenario is a scenario for acquiring audio from a single sound source.
[0146] In the first scenario, the audio processing parameters are at least used to determine the audio collected by the first acquisition unit and the second acquisition unit and to filter it.
[0147] In one embodiment, the audio processing parameters corresponding to the first scene include at least one of the following:
[0148] The preset value;
[0149] First direction information is used to indicate the sound acquisition direction of the first acquisition unit and the second acquisition unit;
[0150] The first noise reduction parameter is used for ambient sound noise reduction in the first scenario.
[0151] In one embodiment, the data acquisition scenario further includes:
[0152] In the second scenario, at least one audio acquisition unit in the terminal device acquires audio signals from multiple sound sources.
[0153] In one embodiment, the audio processing parameters corresponding to the second scene include at least one of the following:
[0154] Acquisition unit information, used to determine the acquisition unit for acquiring multiple sound sources;
[0155] The second direction information is used to determine the acquisition direction for sound acquisition from the multiple sound sources.
[0156] Distance information is used to determine the sampling distance to the sound source;
[0157] The second noise reduction parameter is used for ambient sound noise reduction in the second scenario.
[0158] In one embodiment, the data acquisition scenario further includes:
[0159] The third scenario is used to acquire audio based on an audio acquisition peripheral independent of the terminal device.
[0160] In one embodiment, the audio processing parameters corresponding to the second scene include at least one of the following:
[0161] The first instruction information is used to instruct the use of an audio acquisition peripheral connected to the terminal device to acquire audio signals;
[0162] The third noise reduction parameter is used to indicate that the environmental noise reduction function is turned off when using the acquisition unit on the terminal device to perform audio acquisition.
[0163] The second instruction information is used to instruct the acquisition unit of the terminal device to acquire ambient sound.
[0164] In one embodiment, the scene determination module is further configured to:
[0165] The data collection scenario is determined based on the input operation performed on the terminal device; or,
[0166] The acquisition scenario is determined based on whether there is an audio acquisition peripheral device that is connected to the terminal device.
[0167] It should be noted that the terms "first" and "second" in the embodiments of this disclosure are for ease of description and distinction only, and have no other specific meaning.
[0168] In one embodiment, Figure 8 and Figure 9 This diagram illustrates two different audio signal processing methods. The call toolkit is a VoIP-based application, and ADSP is an advanced digital signal processor. When the VoIP app initiates a call, the default pickup mode is a multi-user scenario. It sends audio signal processing parameters, such as the `remote_record_mode` parameter, to the AudioHAL (Audio Hardware Abstraction Layer). Upon receiving these parameters, AudioHAL processes the audio signal, based on... Figure 2 and Figure 3 The microphone combination shown is used to configure the VoIP uplink audio path and the corresponding audio signal processing parameters. Figure 8 and Figure 9 (The algorithm in the text).
[0169] For example, during a meeting, if a user needs to suppress surrounding noise, they can click [the button]. Figure 7 In the single-player scene shown, the audio signal processing parameters for the single-player scene will be sent, such as the remote_record_mode parameter. After receiving the audio signal processing parameters, AudioHAL will adjust the microphone combination, turn on the microphone on the back of the tablet, and modify the audio signal processing parameters so that the tablet only picks up sound from a certain angle range in front, and suppresses sound from the back and surroundings.
[0170] The terminal's gyroscope senses whether the tablet is currently in landscape or portrait mode. In a single-user scenario, data from four microphones is sent to the voice algorithm for processing. The principle of this algorithm can be referenced from S100 to S300, and algorithms with the same principle can be included within this algorithm. If the current mode is landscape, the sound from the sides and back of the tablet is suppressed based on the data from omnidirectional microphones mic1, mic2, and mic3. If the current mode is portrait, only the sound from the longer side where omnidirectional microphones mic1 and mic2 are located is suppressed. In this case, the audio data from the shorter side omnidirectional microphone mic4 can also be read to suppress the sound from the other side, thus optimizing the sound pickup effect in portrait mode for single-user scenarios.
[0171] Figure 10 This is a block diagram illustrating a terminal device according to an exemplary embodiment. For example, the terminal device may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0172] Reference Figure 10The terminal device may include one or more of the following components: processing component 802, memory 804, power component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.
[0173] Processing component 802 typically controls the overall operation of the terminal device, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0174] Memory 804 is configured to store various types of data to support operation on the terminal device. Examples of this data include instructions for any application or method operating on the terminal device, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0175] Power component 806 provides power to various components of the terminal device. Power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the terminal device.
[0176] Multimedia component 808 includes a screen that provides an output interface between a terminal device and a user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the terminal device is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0177] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when the terminal device is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0178] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0179] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of the terminal device. For example, sensor assembly 814 can detect the on / off state of the terminal device, the relative positioning of components such as the display and keypad of the terminal device, changes in the position of the terminal device or a component of the terminal device, the presence or absence of user contact with the terminal device, the orientation or acceleration / deceleration of the terminal device, and temperature changes of the terminal device. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0180] Communication component 816 is configured to facilitate wired or wireless communication between the terminal device and other devices. The terminal device can access wireless networks based on communication standards, such as Wi-Fi, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0181] In an exemplary embodiment, the terminal device may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0182] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0183] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An audio signal processing method, characterized in that, The method includes: A first audio signal is acquired through a first acquisition unit of the terminal device, and a second audio signal is acquired through a second acquisition unit of the terminal device; wherein the distance difference between the first acquisition unit and the second acquisition unit and the first sound source is within a first range; the distance difference between the first acquisition unit and the second acquisition unit and the second sound source is greater than the maximum value of the first range, and the first sound source is the target acquisition object; The first audio signal and the second audio signal are superimposed in the time domain to obtain the superimposed audio signal; The target audio signal of the target acquisition object is obtained by filtering out the noise audio signal with an amplitude smaller than a preset value from the superimposed audio signal.
2. The method according to claim 1, characterized in that, The terminal device includes a rectangular frame and a back cover; the rectangular frame includes a set of parallel long sides and a set of parallel short sides; the long sides and the short sides are connected end to end; When the terminal device is in landscape orientation, the acquisition surface of the first acquisition unit and the acquisition surface of the second acquisition unit are located at different positions on the same long side; When the terminal device is in portrait orientation, the acquisition surface of the first acquisition unit and the acquisition surface of the second acquisition unit are located on different sides of the rectangular frame; When the terminal device is tilted or in a vertical position, the acquisition surface of the first acquisition unit is located on the rectangular frame, and the acquisition surface of the second acquisition unit is located on the back cover.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Determine the data collection scenario; According to the acquisition scenario, the audio processing parameters of the terminal device are set; wherein, the audio processing parameters are at least used to determine the acquisition unit and / or acquisition parameters for audio acquisition in the acquisition scenario; The acquisition scenario includes a first scenario; wherein the first scenario is a scenario for acquiring audio from a single sound source. In the first scenario, the audio processing parameters are at least used to determine the audio collected by the first acquisition unit and the second acquisition unit and to filter it.
4. The method according to claim 3, characterized in that, The audio processing parameters corresponding to the first scene include at least one of the following: The preset value; First direction information is used to indicate the sound acquisition direction of the first acquisition unit and the second acquisition unit; The first noise reduction parameter is used for ambient sound noise reduction in the first scenario.
5. The method according to claim 3, characterized in that, The data collection scenarios also include: In the second scenario, at least one audio acquisition unit in the terminal device acquires audio signals from multiple sound sources.
6. The method according to claim 5, characterized in that, The audio processing parameters corresponding to the second scene include at least one of the following: Acquisition unit information, used to determine the acquisition unit for acquiring multiple sound sources; The second direction information is used to determine the acquisition direction for sound acquisition from the multiple sound sources. Distance information is used to determine the sampling distance to the sound source; The second noise reduction parameter is used for ambient sound noise reduction in the second scenario.
7. The method according to claim 3, characterized in that, The data collection scenarios also include: The third scenario is used to acquire audio based on an audio acquisition peripheral independent of the terminal device.
8. The method according to claim 7, characterized in that, The audio processing parameters corresponding to the third scene include at least one of the following: The first instruction information is used to instruct the use of an audio acquisition peripheral connected to the terminal device to acquire audio signals; The third noise reduction parameter is used to indicate that the environmental noise reduction function is turned off when using the acquisition unit on the terminal device to acquire audio. The second instruction information is used to instruct the acquisition unit of the terminal device to acquire ambient sound.
9. The method according to claim 3, characterized in that, The determination of the data collection scenario includes: The data collection scenario is determined based on the input operation performed on the terminal device; or, The acquisition scenario is determined based on whether there is an audio acquisition peripheral device that is connected to the terminal device.
10. An audio signal processing device, characterized in that, include: The acquisition module is used to acquire a first audio signal through a first acquisition unit and acquire a second audio signal through a second acquisition unit; wherein the distance difference between the first acquisition unit and the second acquisition unit and the first sound source is within a first range; the distance difference between the first acquisition unit and the second acquisition unit and the second sound source is greater than the maximum value of the first range, and the first sound source is the target acquisition object; The overlay module is used to overlay the first audio signal and the second audio signal in the time domain to obtain an overlay audio signal; The filtering module is used to filter out noise audio signals with amplitudes less than a preset value in the superimposed audio signal to obtain the target audio signal of the target acquisition object.
11. A terminal device, characterized in that, include: A processor and a memory for storing executable instructions capable of running on the processor, wherein: When the processor is used to run the executable instructions, the executable instructions perform the method described in any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by a processor, implement the method described in any one of claims 1 to 9.