Audio acquisition method, electronic device, and storage medium

By combining directional and omnidirectional audio acquisition devices in far-field voice interaction, the effective ratio of audio data is enhanced, the problem of reduced acquisition quality with a single microphone is solved, and higher quality voice signal acquisition is achieved.

CN115884038BActive Publication Date: 2026-02-17HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111156525.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2026-02-17
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

In far-field voice interaction scenarios, the quality of voice signals acquired by a single microphone is greatly affected by noise and interference, resulting in a decrease in the effective audio data ratio.

Method used

By employing at least two directional audio acquisition units and at least one omnidirectional audio acquisition unit, and by increasing the ratio of audio data that meets preset conditions in the target direction, combined with the adjustment of the weight parameters of the directional sound pressure gradient sensor and the omnidirectional pressure sensor, directional beamforming and super-directional beamforming are performed to improve the ratio of effective audio data.

Benefits of technology

The directivity of the audio acquisition device was enhanced, the effective audio data ratio in the target direction was increased, and the quality of the speech signal was improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115884038B_ABST
    Figure CN115884038B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, in particular to an audio acquisition method, an electronic device and a storage medium. The audio acquisition method of the application uses an acoustic vector sensor (AVS) array with directivity as a pickup device, which is better than an omnidirectional microphone array, to acquire voice signals in a space, wherein each AVS in the AVS array comprises an omnidirectional microphone and a directional microphone, then the weights of the voice signals acquired by the omnidirectional microphone and the directional microphone in each AVS are adjusted according to a target direction, the voice signals of each AVS after enhancement in the target direction are obtained, then a designed super-directivity beamformer is applied to the enhanced signals acquired by each AVS for further enhancement processing, and the voice signals of the entire AVS array after enhancement in the target direction are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to an audio acquisition method, electronic device, and storage medium. Background Technology

[0002] In near-field voice interaction scenarios, such as when users engage in voice chat using voice interaction devices like mobile phones, the phone typically uses a single microphone to obtain a voice signal that meets the requirements for voice recognition. However, when voice interaction scenarios evolve to far-field voice interaction, such as smart homes and smart speakers, the quality of the voice signal acquired by a single microphone deteriorates due to the greater distance between the sound source and the microphone, as well as the presence of significant noise interference in the real environment. Summary of the Invention

[0003] To address the aforementioned problems, this application provides an audio acquisition method, an electronic device, and a storage medium. The technical solution of this application is described below.

[0004] In a first aspect, this application provides an audio acquisition method applied to an electronic device, the electronic device including at least two directional audio acquisition devices, the method including: for first directional audio data acquired by the at least two directional audio acquisition devices, increasing the ratio of valid audio data in the first directional audio data acquired by the directional audio acquisition device that satisfies a first preset condition in the target direction, and determining second directional audio data.

[0005] In conjunction with the first aspect and the above possible implementations, the electronic device further includes at least one omnidirectional audio acquisition device, and the method further includes: for the first omnidirectional audio data acquired by the at least one omnidirectional audio acquisition device, increasing the ratio of valid audio data in the first omnidirectional audio data in which the acquisition direction and the target direction satisfy a second preset condition, determining second omnidirectional audio data, and the audio acquisition data includes the second omnidirectional audio data.

[0006] In one possible implementation, the audio acquisition data is obtained by superimposing second directional audio data and second omnidirectional audio data.

[0007] The first directional audio data refers to the audio data collected by the directional audio acquisition device. This audio data includes noise, interference, and valid audio data, such as noise signals and interference signals. Valid audio data refers to the audio data of the target sound source in the target direction, such as the speech signal of the target sound source. The definitions of the first omnidirectional audio data, second omnidirectional audio data, and second directional audio data are similar to those of the first directional audio data.

[0008] In one possible implementation, the effective audio data ratio refers to the signal-to-interference-plus-noise ratio (SIR / NoiseRatio) of the audio data acquired by the directional audio collector in a specific direction. This means the ratio of the target sound source audio data to the sum of interference and noise in the acquired audio data. In other words, increasing the effective audio data ratio of the first directional audio data acquired by the directional audio collector that satisfies the first preset condition in the target direction is equivalent to increasing the SIR / NoiseR of the first directional audio data acquired by the directional audio collector that satisfies the first preset condition in the target direction.

[0009] It should be noted that the superposition of the first directional audio data and the first omnidirectional audio data is not a simple addition of the first directional audio data and the first omnidirectional audio data. The superposition here should be understood as a weighted sum of the audio signals in a physical sense (see the specific calculation process in the embodiments below).

[0010] In this context, the target direction can be understood as an ideal value. In practical applications, enhancing the audio intensity of audio data in the target direction can mean enhancing audio data within a directional range that meets certain conditions. The conditions that the target direction meets the first preset condition include: the directional difference between the acquisition direction of the audio collector and the target direction is less than the first preset value. The first preset value can be set according to the specific application scenario of the audio acquisition method. Specifically, for electronic devices with a relatively small pickup range, the first preset value can be set smaller; for electronic devices with a relatively large pickup range, the first preset value can be set larger. For example, for smart speakers and commonly used microphones in conference rooms, the pickup range of a smart speaker is typically in a home setting, compared to the pickup range of a conference room microphone, which often has dozens or hundreds of participants. Therefore, the first preset value can be set smaller, such as 10°, to improve the enhancement effect of the smart speaker on the audio data and thus improve the user experience. Conversely, for commonly used microphones in conference rooms, the pickup range is relatively large, so the first preset value can be set larger, such as 20° or 30°. It should be understood that the aforementioned values ​​and settings are merely illustrative and are not intended to limit the scope of this application.

[0011] It's understandable that, theoretically, the target direction is consistent for both omnidirectional and directional audio acquisition devices. However, in practical applications, factors such as the placement of the omnidirectional and directional audio acquisition devices may cause slight differences in the relative position of the target direction with respect to the omnidirectional audio acquisition device and the target direction with respect to the directional audio acquisition device. Therefore, in one possible implementation, the first preset condition can be consistent with the second preset condition, and correspondingly, the first preset value can also be consistent with the second preset value. In another possible implementation, the first preset condition and the second preset condition can be inconsistent, and correspondingly, the first preset value can also be inconsistent. However, it's understandable that the setting principles of the first and second preset conditions and the first and second preset values ​​are consistent. Therefore, the setting method for the second preset value can refer to the setting method for the first preset value, and will not be elaborated here.

[0012] In the above method, by increasing the ratio of effective data of the first directional audio data that aligns with the target direction (or meets the first preset condition), and by increasing the ratio of effective data of the first omnidirectional audio data that aligns with the target direction (or meets the first preset condition), and then superimposing the first directional audio data and the first omnidirectional audio data, enhanced audio acquisition data in the target direction can be obtained. The audio acquisition data also includes noise, interference, and effective audio data; however, it can be understood that the ratio of effective audio data in this audio acquisition data is improved compared to the audio acquisition data without the above processing.

[0013] In conjunction with the first aspect and the possible implementations described above, in one possible implementation, the method for increasing the ratio of valid audio data in the first directional audio data collected by the directional audio acquisition device that satisfies the first preset condition with respect to the target direction includes: determining a weight parameter for the first directional audio data collected by each directional sound pressure gradient sensor according to the target direction; increasing the weight parameter of the first directional audio data collected by the directional sound pressure gradient sensor that satisfies the first preset condition with respect to the target direction, so as to increase the ratio of valid data in the first directional audio data collected by the directional sound pressure gradient sensor that satisfies the first preset condition with respect to the target direction. Specifically, determining the weight parameter of the first directional audio data collected by each directional sound pressure gradient sensor according to the target direction can be achieved by assigning a weight to each directional audio data collected by the directional sound pressure gradient sensor according to the target direction. It can be understood that in one possible implementation, the weight of each directional sound pressure gradient sensor can be set to the same value, or the weight parameter of each directional sound pressure gradient sensor can be set to a different value. This application does not limit the specific method of setting the weight parameter of the directional sound pressure gradient sensor. The specific method for assigning weight parameters to each directional sound pressure gradient sensor will be described in detail in the following specific embodiments, and will not be repeated here.

[0014] Then, by increasing the weighting parameter in the audio data collected by the directional sound pressure gradient sensor that meets certain conditions with the target direction, the ratio of effective audio data in the audio data collected by the directional gradient pressure sensor can be increased.

[0015] In one possible implementation, the weighting parameter of the first directional audio data collected by the directional sound pressure gradient sensor that satisfies the first preset condition with respect to the target direction can be understood as directional beamforming of the directional sound pressure gradient sensor. The method of directional beamforming will be described in detail in the specific embodiments below.

[0016] In conjunction with the first aspect and the possible implementations described above, in one possible implementation, the method for increasing the ratio of effective audio data in the first directional audio data acquired by the directional audio acquisition device that satisfies the first preset condition with respect to the target direction includes: determining a weight parameter for the first directional audio data acquired by each directional sound pressure gradient sensor according to the target direction; increasing the weight parameter of the audio data acquired by the directional sound pressure gradient sensor that satisfies the first preset condition with respect to the target direction by adjusting the beamforming parameters, the directional factor, and the white noise gain, thereby increasing the ratio of effective audio data to the first directional audio data acquired by the directional sound pressure gradient sensor that satisfies the first preset condition with respect to the target direction. In one possible implementation, increasing the weight parameter of the first directional audio data acquired by the directional sound pressure gradient sensor that satisfies the first preset condition with respect to the target direction can be understood as performing super-directional beamforming on the directional sound pressure gradient sensor. The super-directional beamforming method will be described in detail in the specific embodiments below. It can be understood that the beamforming parameters refer to the beam pattern or the spatial response of the beam. For a detailed explanation and explanation of the beamforming parameters, directivity factor, and white noise gain, please refer to the relevant descriptions in the specific implementation examples section below.

[0017] In conjunction with the first aspect and the possible implementations described above, in one possible implementation, the method for increasing the ratio of valid audio data in the first omnidirectional audio data acquired by the omnidirectional audio acquisition device that satisfies the second preset condition in the target direction includes: determining a weight parameter for the first omnidirectional audio data acquired by the omnidirectional pressure sensor based on the target direction; increasing the weight parameter of the first omnidirectional audio data acquired by the omnidirectional pressure gradient sensor that satisfies the second preset condition in the target direction, thereby increasing the ratio of valid data in the first omnidirectional audio data acquired by the omnidirectional pressure gradient sensor that satisfies the second preset condition in the target direction. It can be understood that in one possible implementation, the weights of the audio data acquired by the omnidirectional pressure sensor and the directional sound pressure gradient sensor can be adjusted simultaneously; the specific method will be described in detail in the specific embodiments section below.

[0018] In conjunction with the first aspect and the possible implementations described above, in one possible implementation, the method for increasing the ratio of effective audio data in the first omnidirectional audio data acquired by the omnidirectional audio acquisition device that satisfies the second preset condition in the target direction includes: determining the weighting parameters of the first omnidirectional audio data acquired by the omnidirectional pressure sensor according to the target direction; increasing the weighting parameters of the audio data acquired by the omnidirectional pressure gradient sensor that satisfies the second preset condition in the target direction by adjusting the beamforming parameters, directivity factor, and white noise gain, thereby increasing the ratio of effective audio data to the first omnidirectional audio data acquired by the omnidirectional pressure gradient sensor that satisfies the second preset condition in the target direction.

[0019] As can be seen from the above, this application utilizes an audio acquisition device that combines the advantages of both directional and omnidirectional audio acquisition devices. This makes the directionality of the audio acquisition device stronger than that of the omnidirectional audio acquisition device. Furthermore, due to the addition of the directional audio acquisition device, the entire audio acquisition device can increase the ratio of effective audio data in audio data in all directions (including the target direction). In other words, it can increase the ratio of effective audio data in audio data acquired by the directional audio acquisition device and the omnidirectional audio acquisition device that meet certain conditions in the target direction, depending on the specific location of the target direction.

[0020] It is understandable that, in one possible implementation, the aforementioned method of increasing the ratio of effective audio data in the audio data is called "speech enhancement".

[0021] It is understood that, in one possible implementation, the aforementioned multiple audio acquisition devices can be arranged in an array, for example, a linear array, a surface array, or a volume array. This application does not impose any limitations on this. In one possible implementation, the aforementioned multiple audio acquisition devices include at least one omnidirectional audio acquisition unit and at least two directional audio acquisition units. For example, the multiple audio acquisition devices include one omnidirectional audio acquisition unit and two directional audio acquisition units, wherein the audio data acquired by the omnidirectional audio acquisition unit is omnidirectional audio data, and the audio data acquired by the directional audio acquisition units is first directional audio data. In one possible implementation, the audio data can be represented as a sound signal such as a speech signal.

[0022] It is understood that the aforementioned directional audio acquisition device and omnidirectional audio acquisition device refer only to two types of acquisition devices capable of acquiring directional audio data and omnidirectional audio data, and this application does not limit the specific form of the acquisition device.

[0023] More specifically, in conjunction with the first aspect described above, in one possible implementation of the first aspect, the electronic device may include multiple sound vector sensors, with an omnidirectional pressure sensor among the sound vector sensors serving as an omnidirectional audio acquisition device, and a directional sound pressure gradient sensor among the sound vector sensors serving as a directional audio acquisition device. It can be understood that the omnidirectional pressure sensor includes an omnidirectional microphone, and the directional sound pressure gradient sensor includes a directional microphone.

[0024] In conjunction with the first aspect and its possible implementations, in one possible implementation of the first aspect, multiple acoustic vector sensors are arranged according to a preset rule, and each acoustic vector sensor includes an omnidirectional pressure sensor and at least three directional acoustic pressure gradient sensors. The preset rule refers to the multiple acoustic vector sensors being arranged in a linear array (see the relevant description in the specific embodiments below) or in other forms, such as an area array or a volume array. This application does not limit the arrangement of the multiple acoustic vector sensors. In other possible implementations, the acoustic vector sensor array of this application can adopt a hot-wire type, differential pressure type, A-format hybrid type, or other structures. This application also does not limit the specific structure of the acoustic vector sensors.

[0025] In conjunction with the first aspect and its possible implementations, in one possible implementation of the first aspect, the electronic device includes multiple acoustic vector sensors. The target direction can be determined by: using the time difference between first directional audio data and / or omnidirectional audio data acquired by at least two acoustic vector sensors, determining the phase difference between the first directional audio data and / or omnidirectional audio data acquired by at least two acoustic vector sensors; and determining the target direction based on the phase difference between the first directional audio data and / or omnidirectional audio data acquired by at least two audio acquisition devices. Wherein, the electronic device can determine the time difference using the first directional audio data acquired by at least two acoustic vector sensors, meaning the electronic device determines the time difference between the arrival times of the two first directional audio data acquired by the two acoustic vector sensors at each acoustic vector sensor. The phase difference between the first directional audio data acquired by at least two acoustic vector sensors refers to the phase difference between the arrival times of the two first directional audio data acquired by the two acoustic vector sensors at each acoustic vector sensor. According to the direction-of-arrival estimation method, the electronic device can obtain a steering vector based on the phase difference between the audio data arriving at each acoustic vector sensor. This steering vector includes the angular information of the audio data; therefore, the electronic device can determine the target direction based on the angular information of the audio data.

[0026] It is understandable that, in one possible implementation, the target direction can be a certain directional range set by the developers based on the specific application scenario of the audio acquisition method. For example, in some application scenarios, the target direction may be fixed or rarely change. In such cases, the developers can set the target direction within a certain directional range, so that the electronic device only needs to enhance the audio data within that directional range. For instance, if the electronic device is a Bluetooth headset, then when a user interacts with the Bluetooth headset via voice, the position of the user's voice command relative to the Bluetooth headset is fixed. In this case, the target direction of the Bluetooth headset can be set to a fixed directional range. It should be understood that this application does not limit the method of setting the target direction.

[0027] It is understood that the principles and concepts of the audio acquisition method of this application can also be applied to other scenarios similar to audio acquisition, such as the enhancement of antenna millimeter wave signals, etc. This application does not limit other application scenarios of this method.

[0028] It is understood that both the directional and omnidirectional audio acquisition devices mentioned above indicate that the device can acquire specific audio data, such as directional audio data and omnidirectional audio data. In other possible implementations, the directional and omnidirectional audio acquisition devices can also be replaced by other devices that can achieve the corresponding functions, and this application does not impose any restrictions on this.

[0029] Secondly, this application provides an electronic device, the electronic device including a processor, a memory, and at least two directional audio collectors; and the memory is used to store instructions executed by one or more processors of the electronic device; and the processor, one of the processors of the electronic device, is used to run the instructions to enable the electronic device to perform the following operations: for first directional audio data collected by the at least two directional audio collectors, increasing the ratio of valid audio data in the first directional audio data collected by the directional audio collector that satisfies a first preset condition in the target direction, to determine second directional audio data; determining audio data collected by the electronic device in the target direction, wherein the audio data collected includes the second directional audio data.

[0030] In conjunction with the second aspect, in one possible implementation of the second aspect, the electronic device further includes at least one omnidirectional audio acquisition device, and the method further includes: for the first omnidirectional audio data acquired by the at least one omnidirectional audio acquisition device, increasing the ratio of valid audio data in the first omnidirectional audio data in which the acquisition direction and the target direction satisfy a second preset condition, to determine the second omnidirectional audio data; and the audio acquisition data includes the second omnidirectional audio data.

[0031] In conjunction with the second aspect, in one possible implementation of the second aspect, the audio acquisition data is obtained by superimposing the second directional audio data and the second omnidirectional audio data.

[0032] In conjunction with the second aspect, in one possible implementation of the second aspect, the electronic device includes multiple acoustic vector sensors, wherein an omnidirectional pressure sensor among the acoustic vector sensors serves as an omnidirectional audio acquisition device, and a directional sound pressure gradient sensor among the acoustic vector sensors serves as a directional audio acquisition device.

[0033] In conjunction with the second aspect and the above possible implementations, in one possible implementation of the second aspect, multiple acoustic vector sensors are arranged according to a preset rule, and each acoustic vector sensor includes an omnidirectional pressure sensor and at least three directional acoustic pressure gradient sensors.

[0034] In conjunction with the second aspect and the above possible implementations, in one possible implementation of the second aspect, the method for increasing the ratio of valid audio data in the first directional audio data collected by the directional audio collector that satisfies the first preset condition in the target direction includes: determining a weight parameter for the first directional audio data collected by each directional sound pressure gradient sensor according to the target direction; increasing the weight parameter of the first directional audio data collected by the directional sound pressure gradient sensor that satisfies the first preset condition in the target direction, so as to increase the ratio of valid data in the first directional audio data collected by the directional sound pressure gradient sensor that satisfies the first preset condition in the target direction.

[0035] In conjunction with the second aspect and the above possible implementations, in one possible implementation of the second aspect, the method for increasing the ratio of effective audio data in the first directional audio data collected by the directional audio collector that satisfies the first preset condition with respect to the target direction includes: determining the weight parameter of the first directional audio data collected by each directional sound pressure gradient sensor according to the target direction; increasing the weight parameter of the audio data collected by the directional sound pressure gradient sensor that satisfies the first preset condition with respect to the target direction by adjusting the beamforming parameter, the directional factor, and the white noise gain, thereby increasing the ratio of effective audio data to the first directional audio data collected by the directional sound pressure gradient sensor that satisfies the first preset condition with respect to the target direction.

[0036] In conjunction with the second aspect and the above possible implementations, in one possible implementation of the second aspect, the method for increasing the ratio of effective audio data in the first omnidirectional audio data acquired by the omnidirectional audio acquisition device that satisfies the second preset condition in the target direction includes: determining the weight parameter of the first omnidirectional audio data acquired by the omnidirectional pressure gradient sensor according to the target direction; increasing the weight parameter of the first omnidirectional audio data acquired by the omnidirectional pressure gradient sensor that satisfies the second preset condition in the target direction, so as to increase the ratio of effective data in the first omnidirectional audio data acquired by the omnidirectional pressure gradient sensor that satisfies the second preset condition in the target direction.

[0037] In conjunction with the second aspect and the above possible implementations, in one possible implementation of the second aspect, the method for increasing the ratio of effective audio data in the first omnidirectional audio data acquired by the omnidirectional audio acquisition device that satisfies the second preset condition with respect to the target direction includes: determining the weighting parameters of the first omnidirectional audio data acquired by the omnidirectional pressure gradient sensor according to the target direction; increasing the weighting parameters of the audio data acquired by the omnidirectional pressure gradient sensor that satisfies the second preset condition with respect to the target direction by adjusting the beamforming parameters, directivity factor, and white noise gain, thereby increasing the ratio of effective audio data to the first omnidirectional audio data acquired by the omnidirectional pressure gradient sensor that satisfies the second preset condition with respect to the target direction.

[0038] In conjunction with the second aspect and the above possible implementations, in one possible implementation of the second aspect, the electronic device includes multiple acoustic vector sensors, and the target direction can be determined by the following method: using the time difference of the first directional audio data and / or omnidirectional audio data acquired by at least two acoustic vector sensors, the phase difference of the first directional audio data and / or omnidirectional audio data acquired by at least two acoustic vector sensors is determined; based on the phase difference of the first directional audio data and / or omnidirectional audio data acquired by at least two audio acquisition devices, the target direction is determined.

[0039] In conjunction with the second aspect and the above possible implementations, in one possible implementation of the second aspect, the situation in which the target direction satisfies the first preset condition includes: the directional difference between the acquisition direction pointing to the audio collector and the target direction is less than the first preset value.

[0040] In conjunction with the second aspect and the above possible implementation methods, in one possible implementation method of the second aspect, the situation in which the acquisition direction and the target direction satisfy the second preset condition includes: the directional difference between the acquisition direction and the target direction is less than the second preset value.

[0041] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program, characterized in that the computer program, when executed by a processor, implements the audio acquisition method in any possible implementation of the first aspect described above.

[0042] Fourthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the audio acquisition method in any possible implementation of the first aspect described above.

[0043] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a schematic diagram of an example voice interaction scenario provided in some embodiments;

[0046] Figure 2 This is a schematic diagram illustrating an example of using a microphone array to locate the direction of a sound source, provided in some embodiments;

[0047] Figure 3 This is a topology diagram of an example microphone array provided in some embodiments;

[0048] Figure 4 This is a schematic diagram of an example acoustic vector sensor structure provided in some embodiments;

[0049] Figure 5 This is a schematic diagram comparing the beam pattern of an omnidirectional microphone and the beam pattern of a directional microphone, provided in some embodiments.

[0050] Figure 6 This is a schematic diagram of the structure of an audio acquisition device provided in some embodiments;

[0051] Figure 7 This is a schematic diagram of an example acoustic vector sensor array arrangement provided in some embodiments;

[0052] Figure 8 This is a flowchart illustrating an example audio acquisition method provided in some embodiments;

[0053] Figure 9This is a schematic diagram comparing the desired direction with the sound source direction provided in some embodiments;

[0054] Figure 10 This is a flowchart illustrating yet another example of an audio acquisition method provided in some embodiments;

[0055] Figure 11 These are comparative schematic diagrams illustrating the use of different microphone arrays to form beams, provided in some embodiments.

[0056] Figure 12 These are schematic diagrams comparing the performance of beamforming using different microphone arrays, provided in some embodiments.

[0057] Figure 13 These are comparative schematic diagrams illustrating the use of different microphone arrays to form beams, provided in some embodiments.

[0058] Figure 14 These are schematic diagrams comparing the performance of beamforming using different microphone arrays, provided in some embodiments.

[0059] Figure 15 These are comparative schematic diagrams illustrating the use of different microphone arrays to form beams, provided in some embodiments.

[0060] Figure 16 These are schematic diagrams comparing the performance of beamforming using different microphone arrays, provided in some embodiments.

[0061] Figure 17 This is another schematic diagram of an audio acquisition device provided in some embodiments;

[0062] Figure 18 This is a schematic diagram illustrating a scenario of sound field reconstruction using enhanced speech from an audio acquisition device, provided in some embodiments.

[0063] Figure 19 This is a schematic diagram of the hardware structure of an example voice interaction device provided in some embodiments. Detailed Implementation

[0064] With the development of artificial intelligence technology, more and more terminal devices have voice interaction functions. In this embodiment, terminal devices with voice interaction functions are referred to as voice interaction devices. Voice interaction devices are equipped with a sound pickup device (e.g., a single microphone or a microphone array). Voice interaction devices can collect voice signals through the sound pickup device and perform processing such as voice recognition on the voice signals. Voice interaction devices in this embodiment include, but are not limited to, smartphones, laptops, tablets, smart speakers, smart in-vehicle devices, smart robots, smart home devices, and smart wearable devices.

[0065] This application discloses a technical solution for improving the quality of voice signals received by electronic devices. The following is an example... Figure 1 Using the user-smart speaker interaction scenario shown as an example, the audio acquisition method of this application is introduced.

[0066] like Figure 1 As shown, when user 10 issues a voice command to smart speaker 20, due to the unavoidable noise and / or interference in the surrounding environment, such as the sound interference from the TV 30 in the figure, the voice signal collected by smart speaker 20 includes user 10's voice command, as well as the noise from the TV 30 and other interference.

[0067] In order to enable the smart speaker 20 to collect high-quality voice signals from the user 10, in some embodiments of this application, the smart speaker 20 uses an omnidirectional microphone array as a pickup device to collect the voice signals from the user 10. The omnidirectional microphone array is formed by arranging multiple omnidirectional microphones according to certain rules. An omnidirectional microphone is a microphone with consistent sensitivity to voice signals in all directions. It can collect voice signals from any direction. An omnidirectional microphone is relative to a directional microphone (or a directional sound pressure gradient sensor). A directional microphone is only sensitive to voice signals in a specific direction. That is, a directional microphone can only collect voice signals in a specific direction. The smart speaker 20 uses the phase difference between the voice signals of the user 10 collected by each omnidirectional microphone in the omnidirectional microphone array to determine the direction of the user 10 relative to the microphone array and generates a beam pointing in that direction to enhance the voice signal of the user 10 from that direction.

[0068] The beamformation represents the spatial response of a microphone array oriented in a specific direction, that is, the sensitivity of the microphone array to speech signals in that direction. The direction the beam is oriented indicates the microphone array's sensitivity to speech signals in that direction. For example, in... Figure 1 If the smart speaker 20 needs to enhance the voice signal from the user 10, then the smart speaker 20 needs to generate a beam toward the user 10 in order to enhance the voice signal from the direction of the user 10.

[0069] Specifically, in some embodiments, the smart speaker 20 can form a beam using a delay and sum beamforming (DSB) algorithm, as follows:

[0070] Figure 2 The image shows the settings at... Figure 1The smart speaker 20 shown features an omnidirectional microphone array, which can be a uniform linear array composed of M omnidirectional microphones 21a (hereinafter referred to as array elements), wherein each array element is arranged at a interval d on the Z-axis. Assuming a far-field sound source (e.g., Figure 1 If the incident direction of the voice signal of user 10 is θ, meaning the angle at which the voice signal of user 10 reaches each omnidirectional microphone is θ, then the distance from which the voice signal of user 10 reaches a later array element is greater than the distance from which it reaches a previous array element by dcosθ. Correspondingly, the time it takes for the voice signal of user 10 to reach a later array element is delayed compared to the time it takes to reach a previous array element. (c refers to the speed at which sound travels through the air).

[0071] It should be noted that a far-field sound source refers to a sound source whose distance from a microphone in the microphone array is greater than the wavelength of the sound signal emitted by the source. A near-field sound source, in contrast to a far-field sound source, refers to a sound source whose distance from a microphone in the microphone array is less than the wavelength of the sound signal emitted by the source. Generally, near-field sound sources treat the sound signal as a spherical wave. That is, in a near-field sound source scenario, enhancing a speech signal in a certain direction requires considering not only the distance of the speech signal to each element of the microphone array but also the amplitude of the speech signal reaching each element. In a far-field sound source scenario, the sound signal is treated as a plane wave. That is, in specific calculations, the amplitude difference of the speech signal reaching each element can be ignored; only the time delay or phase difference between the speech signal reaching each element needs to be considered. For ease of description, the following will use a far-field sound source scenario as an example to introduce various embodiments of this application. However, it should be understood that the methods and principles corresponding to the various embodiments of this application are also applicable to near-field sound source scenarios.

[0072] Now, the voice signal from user 10 has arrived. Figure 2 Taking the first array element 21a as a reference, the time difference between each array element and the first array element is:

[0073]

[0074] Assuming the frequency of user 10's voice signal is f0, according to the phase calculation formula: ω=2πf0t, where π is pi and t represents time, the phase difference between each array element and the first array element is:

[0075]

[0076] Assuming user 10's voice signal is s(n), the signal that the entire microphone array can collect can be calculated using the following formula (3):

[0077]

[0078] By extracting s(n) from equation (3) above, we can obtain equation (4) below:

[0079]

[0080] Equation (4) above is defined as follows:

[0081]

[0082] Then equation (4) above can be written as:

[0083] X(n)=α(θ)s(n) (6)

[0084] Where X(n) represents the voice signal (i.e. the acquired signal) collected by the microphone array, α(θ) is the steering vector of the microphone array, which is used to represent the phase difference of the voice signal reaching each element of the microphone array. It contains the angular information of the voice signal, and s(n) is the voice signal of user 10.

[0085] Assume user 10's voice signal is a sine wave. As mentioned above, due to the different distances from user 10's voice signal to the microphone array, the sine waves collected by each element in the microphone array have a phase difference Δθ0. The signal collected by the entire microphone array is the superposition of the sine waves collected by each element. Therefore, to enhance the voice signal of user 10 collected by the smart speaker 20 in the direction θ, according to the principle of sine wave superposition, smaller weights are assigned to the sine waves with canceling phases, and larger weights are assigned to the sine waves with increasing phases. Then, the sine waves are superimposed to obtain the maximum superposition effect (resulting in the largest amplitude of the final collected sine wave).

[0086] The above method can enhance the voice signal of user 10 in the θ direction to some extent. However, the directivity of this microphone array composed of omnidirectional microphones is not high. For example, if the omnidirectional microphones are arranged as follows... Figure 3 (A) to Figure 3 The linear array shown in (B), or Figure 3 (C) to Figure 3 When the array is as shown in (D), the smart speaker 20 has no way to control the beam orientation on the non-array plane outside the plane where the microphone array is located. Moreover, when θ is not equal to 0° (i.e. the incident angle of the voice signal is not at the end of the direction), the ability of the microphone array to suppress noise or interference will decrease significantly with the change of θ. Furthermore, when the voice signal of user 10 is a low-frequency signal, according to the above formula (2), the phase difference between the voice signals of user 10 collected by each array element is small. Therefore, the directivity of the omnidirectional microphone array is weak and it cannot better enhance the voice signal of user 10 in the direction θ.

[0087] In some embodiments, the smart speaker 20 can also be arranged in the following manner. Figure 3 (F) or Figure 3 The microphone array shown in (G) serves as the pickup device. While this arrangement solves the aforementioned problems and enhances the directivity of the entire omnidirectional microphone array, it also increases the number of microphones required. This makes it less practical for space-constrained voice interaction devices such as smart speakers, smartphones, and headphones. It should be understood that... Figure 3 The microphone array arrangements in (A) to 3(G) are merely exemplary. The microphone array may also have other linear array, area array, or volume array arrangements, and this application does not limit this.

[0088] To address the aforementioned technical problems, this application also provides an audio acquisition device 1. This audio acquisition device 1 employs an acoustic vector sensor (AVS) array (hereinafter referred to as AVS array) with better directivity as a sound pickup device. Then, it uses the AVS array to acquire sound signals and performs directional beamforming on each AVS constituting the AVS array according to the target direction, so that each AVS in the array can initially enhance the speech signal in the target direction. Then, it determines the speech enhancement parameters of the super-directional beamformer based on the target direction for the enhanced speech signal acquired by each AVS. Finally, it uses the super-directional beamformer to filter the enhanced speech signal in the target direction acquired by the entire AVS array to obtain the final enhanced speech signal of the AVS array in the target direction.

[0089] An AVS array is a microphone array composed of multiple AVS devices. For example, Figure 5 The diagram shows the structural diagram of the AVS array in some embodiments, wherein Figure 4 (B) is Figure 4 A simplified structural diagram of (A), as follows: Figure 4 (B) shows an AVS consisting of one omnidirectional microphone and three directional microphones, combined with... Figure 4 (A) and Figure 4 (B) It can be seen that the three directional microphones are orthogonal to each other and arranged at the same point, and each directional microphone has a pickup hole for picking up voice signals. In some embodiments, the above-mentioned directional microphones are figure-eight directional microphones. It should be noted that the figure-eight directional microphone is not actually shaped like the number 8, but rather its beam is specifically shaped like the number 8, hence the name figure-eight directional microphone. The beam diagram of the figure-eight directional microphone will be discussed below. Figure 5 The beamform of the omnidirectional microphone will be compared with that of the omnidirectional microphone, but will not be described here.

[0090] like Figure 4(C) shows Figure 4 (A) top view, as shown Figure 4 As shown in (C), the X channel represents the pickup channel of the X-axis directional microphone, the Y channel represents the pickup channel of the Y-axis directional microphone, and the z channel represents the pickup channel of the Z-axis directional microphone. This can be understood as... Figure 4 (A) and Figure 4 (B) shows only one exemplary AVS structure. In other embodiments, the AVS may have different structures. Figure 4 (A) and Figure 4 The structure shown in (B) is not limited in this application.

[0091] For example, in some embodiments, the AVS may consist of one omnidirectional microphone and two figure-eight directional microphones; this application is not limited in this regard. As another example, the AVS may also consist of one omnidirectional microphone or other directional microphones; this application is not limited in this regard. Furthermore, the AVS array may be arranged as shown above. Figure 3 (A) to Figure 3 The linear array shown in (B) can also be arranged as follows: Figure 3 (C) to Figure 3 The area array shown in (E) can also be arranged as Figure 3 (F) to Figure 3 The array shown in (G) is not limited in this application.

[0092] It is understood that the specific number of AVSs in the AVS array of this application can also be 2, 4, 5 or 6, and this application does not limit the number of AVSs. Furthermore, the number of omnidirectional microphones and directional microphones in each AVS can also be other values, such as 1 omnidirectional microphone and 2 directional microphones, or 1 omnidirectional microphone and 3 directional microphones, and this application does not limit this either.

[0093] It is understood that the spacing between each AVS in the AVS array can be 3 cm, 3.5 cm, 4 cm, or other values, and this application does not impose any restrictions on this.

[0094] In some embodiments, the spacing between each AVS in the AVS array is related to the purpose of the audio acquisition device 1. For example, if the audio acquisition device 1 is used on a smartphone, the spacing between each AVS in the AVS array should not be too large because the smartphone itself has a small space. If the audio acquisition device 1 is used on a smart speaker, the spacing between each AVS array can be designed to be relatively large because the size and space of the smart speaker are larger than that of the smartphone.

[0095] The preceding text outlines the effects achievable by the audio acquisition device of this application and the general process of achieving those effects. The following section will provide a detailed description of the speech enhancement and audio acquisition method based on the audio acquisition device of this application. Before proceeding, to facilitate understanding of the principle that AVS arrays have stronger directivity compared to omnidirectional microphone arrays, the following section briefly introduces the directivity differences between omnidirectional microphones and figure-eight directional microphones when acquiring speech signals. Figure 6 The beam patterns formed by an omnidirectional microphone and a directional microphone are shown. Figure 5 (A) is the beam pattern formed by the omnidirectional microphone. Figure 5 (B) is the beam pattern formed by a directional microphone.

[0096] from Figure 5 As shown in (A), the beamform of the omnidirectional microphone is nearly circular, with its boundary at 0 dB, meaning the maximum indentation is 0 dB. This indicates that the omnidirectional microphone is sensitive to speech signals from all directions, meaning it doesn't only collect speech signals from a specific direction. Therefore, the omnidirectional microphone has poor directivity. The maximum indentation depth represents the attenuation of the speech signal in a specific direction by a directional microphone. Ideally, it should be... Figure 5 (B) For example, the maximum indentation should be -50dB, which indicates that the directional microphone does not collect voice signals in that specific direction. However, in practical applications, due to factors such as the placement of the directional microphone, it is difficult to reach the maximum indentation of -50dB.

[0097] And from Figure 5 As shown in (B), the beamform of the directional microphone clearly resembles a figure "8," meaning it is only sensitive to speech signals from specific directions. For example, in the 30° and 210° directions, the directional beam can reach 0dB, indicating that there is no attenuation of sound in these directions. However, signals from other directions are suppressed; for example, there is significant attenuation at points P and Q in the diagram.

[0098] Therefore, the AVS array combines the advantages of both omnidirectional and directional microphones, and its directivity is superior to that of a microphone array composed entirely of omnidirectional microphones. Furthermore, since the signal acquired by the AVS array is a superposition of speech signals from all directions acquired by the omnidirectional microphones and speech signals from specific directions acquired by other directional microphones, the speech signal acquired by the AVS array itself possesses a certain degree of directivity. Therefore, the AVS array only needs to be arranged in a linear array to generate a more directional beam through the directional beamforming module 12 and the super-directional beamforming module 13, thereby achieving speech enhancement in a specific direction.

[0099] Based on the above, and in conjunction with, for example Figures 6 to 16 This application describes the audio acquisition device and audio acquisition method.

[0100] like Figure 6 As shown, the audio acquisition device includes an AVS array 11, a directional beamformer 12, and a super-directional beamformer 13.

[0101] The AVS array 11 is used to acquire speech signals in space, and the directional beamformer 12 is used to perform directional beamforming on the speech signals acquired by each acoustic vector sensor in the AVS array 11, so that each acoustic vector sensor can enhance the speech signal in the target direction, thus obtaining the speech signal initially enhanced by the AVS array in the target direction. Since white noise and other noise or interference inevitably exist in the target direction, a super-directional beamformer 13 is needed to perform super-directional filtering on the speech signal output by the directional beamformer 12 to obtain the enhanced speech signal in the target direction. The specific processing procedures of the directional beamformer 12 and the super-directional beamformer 13 will be discussed in the following sections. Figure 8 This will be introduced in detail. It is understood that in some embodiments, the directional beamformer 12 and the super-directional beamformer 13 may not be set separately, that is, the functions of the directional beamformer 12 and the super-directional beamformer 13 may be integrated into a single processor, and this application does not limit this.

[0102] To further understand the implementation method of the audio acquisition method in this application, the following will be combined with... Figure 7 The AVS array shown illustrates the process by which the audio acquisition device 1 of this application implements the audio acquisition method 800. Among other things, Figure 7 The AVS array consists of 6 AVSs arranged along the X-axis. In the figure, the white circles represent omnidirectional microphones, light gray represents X-directional microphones, dark gray represents Y-directional microphones, and dark gray represents Z-directional microphones. Each AVS consists of 1 omnidirectional microphone and 3 figure-eight directional microphones. In the figure, the omnidirectional microphone and the 3 mutually orthogonal directional microphones that make up the AVS are set at the same point.

[0103] It should be understood that in practical applications, the number of AVS in the AVS array, the signal-to-noise ratio of each microphone in each AVS, the orientation of the directional microphone, and the phase and amplitude of the speech signal arriving at each AVS all need to be carefully calculated and designed in order to better enhance the speech signal in the target direction. This application does not limit the number of AVS in the AVS array or the parameter requirements that each AVS must meet.

[0104] Signal-to-noise ratio (SNR) refers to the ratio of signal to noise in an electronic device or system. Signal refers to the electronic signal from outside the electronic device that needs to be processed by it. Noise refers to irregular extra signals (or information) generated after passing through the electronic device that are not present in the original signal, and this extra signal does not change with the original signal. Amplitude consistency indicates that the amplitude of the speech signal arriving at each AVS in the microphone array should be as consistent as possible. Amplitude can also be understood as the energy of the speech signal, reflecting whether the energy of the speech signal acquired by each AVS is consistent. Phase consistency indicates that the phase difference of the speech signal arriving at each AVS is within a certain range, so that when the signals are superimposed, the speech signal in the target direction can be maximized. The orientation consistency of the directional microphones in the AVS is to ensure that the orientation of the directional microphones in each AVS is as consistent as possible, so that the speech signals acquired by the directional microphones in each AVS differ only in phase. In some embodiments, the signal-to-noise ratio of each microphone in the above-mentioned AVS is greater than 60dB, the maximum recess depth of the directional microphone is greater than -20dB, and the amplitude consistency between each AVS is ±1dB, the phase consistency is ±5°, and the orientation consistency of the directional microphone in each AVS is ±5°. This application does not impose any limitations on these aspects.

[0105] The following is combined Figure 1 The scene diagram shown and Figure 6 The structural diagram of the audio acquisition device 1 shown illustrates the audio acquisition method 800 of this application, such as... Figure 8 As shown, method 800 includes:

[0106] 801, Acquires acoustic signals via an AVS array.

[0107] In some embodiments, the audio acquisition device 1 acquires sound signals through an AVS array, wherein the sound signals can be understood to include the voice signals of the user 10, the noise of the television 30, and other interference signals in the space.

[0108] It is understandable that, since each AVS in an AVS array consists of one omnidirectional microphone and three directional microphones, the speech signal acquired by each AVS in the AVS array is actually the superposition of the omnidirectional component acquired by the omnidirectional microphone and the directional component acquired by the three directional microphones, that is, the X-axis directional component, Y-axis directional component, and Z-axis directional component. Here, the omnidirectional component refers to the speech signal acquired by the omnidirectional microphone in all directions; the X-axis directional component refers to the speech signal acquired by the X-axis directional microphone along the X-axis direction; the Y-axis directional component refers to the speech signal acquired by the Y-axis directional microphone along the Y-axis direction; and the Z-axis directional component refers to the speech signal acquired by the X-axis directional microphone along the X-axis direction.

[0109] 802. Adjust the weights of each component of the speech signal acquired by each AVS according to the target direction so that each AVS obtains a speech signal pointing in the target direction.

[0110] It is understandable that, in order for the AVS array to obtain the enhanced speech signal in the target direction, the audio acquisition device 1 can make each AVS in the AVS array obtain the enhanced speech signal in the target direction, and then superimpose the speech signals obtained by each AVS to obtain the enhanced speech signal of the AVS array in the target direction.

[0111] Among them, the target direction From the pitch angle θ s (Value range 0°~180°) and azimuth angle (Values ​​range from -180° to 180°) constitute the structure. Figure 9 A schematic diagram of the target direction is shown, from Figure 9 It can be seen from θ s The angle between the speech signal and the z-axis. The angle between the speech signal and the x-axis represents the target direction. It can point in any direction in space.

[0112] The target direction is the direction in which the voice signal enhancement is desired. In some embodiments, the target direction can be preset by the developers according to the specific application scenario of the voice interaction device, for example, using... Figure 1Taking the interaction between user 10 and smart speaker 20 as an example, developers can use the area in front of smart speaker 20 as the target direction. That is, when user 10 stands in front of smart speaker 20 and issues a voice command, smart speaker 20 will perform voice enhancement processing on the voice signal corresponding to that voice command. In other embodiments, since the location of the sound source may change, the target direction can also be determined in real time by the audio acquisition device 1 using the direction of arrival (DOA). For example, using... Figure 1 Taking the scenario of user 10 interacting with smart speaker 20 as an example, when user 10 moves to different locations, smart speaker 20 can use DOA to determine the approximate direction of user 10 at this time, and then enhance the voice signal of user 10 in that direction. Specifically, audio acquisition device 1 determines the phase difference between the voice signals acquired by each AVS based on the time difference between the voice signals acquired by each AVS, and then calculates the steering vector of the AVS array using the above formulas (3) to (5). The steering vector includes the angle information of the voice signal, and audio acquisition device 1 can determine the target direction based on the angle information of the voice signal in the steering vector.

[0113] It is understandable that, in practical applications, enhancing speech signals in the target direction can also mean enhancing speech signals in directions that meet certain conditions relative to the target direction. These conditions include directions whose directional differences from the target direction are within a certain range. For example, assuming the target direction is (30°, 60°), that is, the target direction is the direction of the AVS array's elevation angle of 30° and azimuth angle of 60°, then enhancing the speech signal in the target direction actually refers to enhancing the speech signal within the range of directions differing from the elevation angle of 30° by ±10° and the azimuth angle by ±5°.

[0114] Furthermore, since each AVS is composed of omnidirectional and directional microphones, in order for each AVS to collect speech signals in the target direction, it is necessary to adjust the weights of the speech signals collected by each microphone in each AVS in each direction. That is, to adjust the weights of the speech signals collected by the omnidirectional microphone (hereinafter referred to as the omnidirectional component), the speech signals collected by the X-axis directional microphone (hereinafter referred to as the X-axis directional component), the speech signals collected by the Y-axis directional microphone (hereinafter referred to as the Y-axis directional component), and the speech signals collected by the Z-axis directional microphone (hereinafter referred to as the Z-axis directional component).

[0115] For example, if the target direction is (0°, 0°), that is, in Figure 9On the Z-axis, the weight of the omnidirectional component acquired by the omnidirectional microphone can be set to 0.2. This results in the X-axis directional microphone acquiring an X-axis directional component of 0, the Y-axis directional microphone acquiring a Y-axis directional component of 0, and the Z-axis directional microphone acquiring a Z-axis directional component of 0.8. In other words, the weight of the directional components acquired by the three orthogonal directional microphones is increased. If the target direction is (45°, 45°), the weight of the omnidirectional component acquired by the omnidirectional microphone can be set to 0.5. This results in the X-axis directional microphone acquiring an X-axis directional component of 0.25, the Y-axis directional microphone acquiring a Y-axis directional component of 0.25, and the Z-axis directional microphone acquiring a Z-axis directional component of 0.8. This means increasing the weight of the omnidirectional component acquired by the omnidirectional microphone. It should be noted that the weights of the directional components acquired by the three directional microphones can be the same or different. When the weights of the directional components acquired by the three directional microphones are the same, the sum of the weight of the directional component acquired by any one directional microphone and the weight of the omnidirectional component acquired by the omnidirectional microphone is 1. When the weights of the directional components acquired by the three directional microphones are different, the sum of the weights of the directional components acquired by the three directional microphones and the sum of the weights of the omnidirectional component acquired by the omnidirectional microphone is 1. This application does not impose any restrictions on this.

[0116] After adjusting the weights of each component of the acquired speech signal of each AVS according to the target direction, directional beamforming is performed on each AVS to obtain the enhanced speech signal acquired by that AVS in the target direction, so as to obtain the enhanced speech signal of the AVS array in the target direction.

[0117] The specific calculation process involved in this step will be explained in detail below.

[0118] 803. Based on the voice signal output by the AVS array, determine the various parameters of the super-directional beamformer.

[0119] It is understood that the enhanced speech signal obtained by the aforementioned AVS array in the target direction includes not only the speech signal of user 10 in that target direction, but also speech signals such as noise and interference in that target direction. Therefore, a super-directional beamformer is needed to filter the enhanced speech signal output by the aforementioned AVS array in the target direction to obtain the enhanced speech signal in the target direction.

[0120] Furthermore, super-directional beamformers generally include performance evaluation parameters such as beam spatial response, white noise gain, and directivity factor. Among these, the beam spatial response represents the signal from which the AVS array receives signals from an elevation angle θ and an azimuth angle of θ. The higher the enhancement of the speech signal in the direction, the better the beam space response, indicating that the AVS array is more responsive to elevation angle θ and azimuth angle θ. The stronger the enhancement of the speech signal in a particular direction, the higher the directivity factor represents the spatial gain of the beam, i.e., the degree of energy concentration of the beam. A higher directivity factor indicates that the AVS array can extract signals more concentratedly from a certain direction. White noise gain reflects the robustness of the AVS array to noise. A higher white noise gain means that white noise has less impact on the super-directive beamformer 13. Since these parameters are used to evaluate the performance of the beamformer, it is necessary to calculate the spatial response, white noise gain, and directivity factor mentioned above, and then use these parameters to construct the super-directive beamformer 13 to process the speech signal output from the AVS array to obtain the enhanced speech signal.

[0121] The specific calculation process involved in this step will be explained in detail below.

[0122] 804. The speech signal output from the AVS array is filtered using the super-directional beamformer described above to obtain an enhanced speech signal.

[0123] After determining the parameters of the super-directional beamformer 13 in step 803, the super-directional beamformer 13 can be used to filter the speech signal output by the AVS array to obtain the speech signal enhanced by the AVS array in the target direction.

[0124] The specific calculation process involved in this step will be explained in detail below.

[0125] It should be noted that in some embodiments, the super-directional beamformer 13 can also be used directly to filter the speech signal acquired by the AVS array. That is, 802 in the above method 800 can be omitted. That is, the audio acquisition device 1 no longer preprocesses the speech signal acquired by each AVS according to the target direction so that each AVS obtains an enhanced speech signal in the target direction, and thus obtains the enhanced speech signal of the AVS array in the target direction. Instead, it directly uses the super-directional beamformer 13 to filter the speech signal acquired by the AVS array in the target direction to obtain the enhanced speech signal.

[0126] It is understandable that the above method 800 performs consistent weight adjustment on the speech signal components acquired by each AVS in the AVS array according to the target direction, then performs directional beamforming on the speech signal acquired by each AVS, and then performs super-directional beamforming on the speech signal after directional beamforming. This limits the performance of the AVS array to a certain extent, affects the directivity of the final beam formed by the AVS array, and further affects the speech enhancement effect of the AVS array. Therefore, in order to fully utilize the performance of the AVS array, in some other embodiments of this application, the enhanced user 10 speech signal can also be obtained directly by filtering the speech signal acquired by the AVS array using the super-directional beamformer 13, based on the global optimization design concept.

[0127] The following section details the method for directly filtering the speech signal acquired by the AVS array using the super-directional beamformer 13. Figure 10 As shown, implementation details that are the same as or similar to the above method 800 can be referred to the specific implementation process of the above method 800, and will not be repeated here.

[0128] Method 1000 includes:

[0129] 1001, Acquire voice signals via AVS array.

[0130] 1002. Based on the voice signal output by the AVS array, determine the various parameters of the super-directional beamformer.

[0131] 1003. The speech signal output from the AVS array is filtered using the super-directional beamformer described above to obtain an enhanced speech signal.

[0132] As can be seen from the above introduction, the audio acquisition device 1 of this application has better directivity when acquiring speech signals because it adopts an AVS array. Therefore, after processing it with a directional beamformer and a super-directional beamformer 13, it can achieve the effect of enhancing speech signals in a specific direction.

[0133] It is understood that in other embodiments of this application, the omnidirectional microphone array described above can also be used to implement the audio acquisition method 800. However, it should be understood that, as mentioned above, since linear array and area array omnidirectional microphone arrays have poor directivity, the speech enhancement effect achieved by method 800 is necessarily not as good as the speech enhancement effect achieved by the AVS array of this application using the same method. If a volume array omnidirectional microphone array is used, in order to achieve the same effect as the AVS linear array, the number of microphones used is more than that required by the AVS linear array. Therefore, its overall effect is not as good as that of the AVS array of this application.

[0134] To intuitively understand the advantages of the audio acquisition device of this application, the following uses an AVS linear array with 6 elements and an omnidirectional microphone array linear column as examples to compare and contrast the beam patterns formed by the AVS array and the omnidirectional microphone array. It can be understood that, in this field, the beamwidth of the beam pattern can indicate whether the corresponding microphone array is more directional; the smaller the beamwidth, the stronger the directionality of the corresponding microphone array, and the larger the beamwidth, the weaker the directionality of the corresponding microphone array. Specifically:

[0135] Figure 11 This is an example of a beammap provided in some embodiments, based on a uniform linear omnidirectional microphone array and an AVS array based on this application, utilizing a superdirectional beamformer. More specifically, Figure 11 (A) is the beammap of a uniform linear omnidirectional microphone array. Figure 11 (B) is the beam pattern of the AVS array of this application formed by method 800. Figure 11 (C) is the beam pattern formed by the AVS array of this application using method 900, wherein the frequency of the speech signal in space is f = 1 kHz, and the target directions of the AVS array and the omnidirectional microphone array are... The weight allocation for each acoustic vector sensor in each AVS in the AVS array is a0 = a1 = 1 / 2.

[0136] Due to the target direction This is equivalent to the desired sound source location being on the positive half of the Z-axis. Figure 11 As can be seen from (A), the beam pattern of the omnidirectional microphone linear array has weak beam directivity, while Figure 11 In (B), the beam pattern formed by the AVS linear array has a narrower beam and stronger directivity. Figure 11 (C) The beamwidth of the beam pattern formed by the AVS linear array is also narrower, and its directivity is stronger. That is, the directivity of the omnidirectional microphone array is not as strong as that of the AVS array of this application. Furthermore, for the same AVS microphone array, the directivity of the beam formed by directly using method 1000 is better than that of the beam formed by using method 800.

[0137] Because the AVS array of this application is more directional than the omnidirectional microphone array, it is understood that the AVS array of this application is more sensitive to speech signals in a specific direction, and correspondingly, it has a suppressive effect on speech signals in other directions. Therefore, the AVS array of this application can better suppress spatial noise.

[0138] To more intuitively understand the advantages of the AVS array in this application compared to omnidirectional microphone arrays—namely, stronger directivity and better suppression of spatial background noise—the following... Figure 12 It shows the corresponding above Figure 11 Directivity factor curves for each beam pattern. Figure 12 The graph shows the directivity factor (in decibels (dB)) curves for AVS arrays and omnidirectional microphone arrays.

[0139] like Figure 12 As shown, the dotted line represents the directivity factor of the linear array of the linear omnidirectional sensor, the dashed line represents the directivity factor of the beam pattern formed by the AVS linear array using method 800, and the solid line represents the directivity factor of the beam pattern formed by the AVS linear array using method 1000.

[0140] from Figure 12 As can be seen, the beam directivity factor formed by a linear omnidirectional microphone array is relatively low, and its directivity factor decreases significantly with increasing frequency. For example, at a frequency of 2kHz, the directivity factor of the linear omnidirectional microphone array is 14dB, but at a frequency of 6kHz, the directivity factor becomes 7dB.

[0141] The beams formed by AVS arrays consistently exhibit high directivity, and their directivity factor remains relatively stable as the frequency increases. For example, at a frequency of 2 kHz, the directivity factor of the beam formed by an AVS array using method 800 is 16 dB, while the directivity factor of the beam formed by an AVS array using method 1000 is 19 dB. At a frequency of 6 kHz, the directivity factor of the beam formed by an AVS array using method 800 is 13 dB, while the directivity factor of the beam formed by an AVS array using method 1000 is 20 dB.

[0142] Therefore, the directivity of the beam formed by the AVS array of this application is stronger than that of the beam formed by the omnidirectional microphone array. Moreover, the beam formed by the AVS array of this application has frequency consistency, that is, its directivity does not change significantly with frequency. Therefore, the AVS array of this application does not have high requirements for the frequency of the voice signal in the scene, and can be applied to a wider range of scenarios than the omnidirectional microphone array.

[0143] Down Figure 13 and Figure 15 The beam patterns formed by the AVS array and omnidirectional microphone array of this application are shown respectively under different target directions. These correspond to... Figure 13 and Figure 15 , Figure 14 and Figure 16 The comparison of the directivity factors of the beams formed by the AVS array and the omnidirectional microphone array of this application are shown respectively in different target directions.

[0144] The following will elaborate on this. Before proceeding, it is necessary to clarify that... Figure 11 , Figure 13 , Figure 15 The main difference lies in the target direction; for other similar or identical parts, please refer to the above to avoid repetition. Figure 11 The description. Similarly, Figure 12 , Figure 14 , Figure 16 The main difference lies in the target direction; other similar or identical parts can be found above. Figure 12 Related descriptions. Furthermore, Figures 11 to 16 The number of AVS arrays and the spacing between array elements involved are the same as described above, and will not be repeated here.

[0145] Specifically, Figure 13 For the target direction At that time, the beam pattern formed by the omnidirectional microphone array and the AVS array. More specifically, Figure 13 (A) is the beam pattern formed by the omnidirectional microphone array. Figure 13 (B) is the beam pattern generated by the AVS array using method 800. Figure 13 (C) is the beam pattern formed by the AVS array using method 1000.

[0146] from Figure 13 As can be seen in (A), the beam formed by the omnidirectional microphone array has no obvious directionality, while Figure 13 (B) Beams formed by the AVS array and comparison Figure 13 (A) It has already shown a clear directionality (narrowing). Figure 13 Beam contrast formed by the AVS array in (C) Figure 13 (A) also has obvious directivity. Moreover, the beam width formed by the AVS array is smaller than that formed by the omnidirectional microphone array, which has good directivity. Furthermore, the beam width formed by method 1000 is smaller than that formed by method 800, which has even better directivity.

[0147] Furthermore, from Figure 14 As can be seen from (A), the directivity factor of the beam formed by the linear omnidirectional microphone array is lower than that of the AVS array. Moreover, the directivity factor of the beam formed by the linear omnidirectional microphone array changes more significantly with frequency, while the directivity factor of the beam formed by the AVS array has frequency consistency, that is, it does not change much with frequency. The directivity factor of the beam formed by method 1000 is even more frequency consistent, and its directivity factor hardly changes with frequency.

[0148] Figure 15 For the target direction At that time, the beam pattern formed by the omnidirectional microphone array and the AVS array. More specifically, Figure 15 (A) is the beam pattern formed by the omnidirectional microphone array. Figure 15 (B) is the beam pattern generated by the AVS array using method 800. Figure 15 (C) is the beam pattern formed by the AVS array using method 900.

[0149] from Figure 15 As can be seen in (A), the beam formed by the omnidirectional microphone array has no obvious directionality, while Figure 15 Beam comparison of AVS array in (B) Figure 15 (A) already has a clear direction. Figure 15 Beam contrast formed by the AVS array in (C) Figure 15 (A) also has obvious directivity. Moreover, the beam width formed by the AVS array is smaller than that formed by the omnidirectional microphone array, which has good directivity. Furthermore, the beam width formed by method 1000 is smaller than that formed by method 800, which has even better directivity.

[0150] Furthermore, from Figure 16 As can be seen from (A), the directivity factor of the beam formed by the linear omnidirectional microphone array is lower than that of the AVS array. Moreover, the directivity factor of the beam formed by the linear omnidirectional microphone array changes more significantly with frequency, while the directivity factor of the beam formed by the AVS array has frequency consistency, that is, it does not change much with frequency. The directivity factor of the beam formed by method 1000 is even more frequency consistent, and its directivity factor hardly changes with frequency.

[0151] The above comparison illustrates that, in different target directions, the AVS array has better directivity than the omnidirectional microphone array. From another perspective, comparing the above... Figure 11 , Figure 13 , Figure 15 This also demonstrates that the AVS array has better beam-direction capability than the omnidirectional microphone array. In other words, because the AVS array has good directivity, it can form a beam pointing in any direction to enhance the speech signal in that direction while suppressing speech signals in other directions.

[0152] The preceding text introduced the advantages of using an AVS array in the audio acquisition device of this application compared to arrays composed of omnidirectional microphones in other embodiments. The following text, corresponding to method 800, describes the specific implementation details and calculation processes involved in the audio acquisition method implemented by the audio acquisition device of this application. Implementation details in method 1000 that are the same as or similar to those in method 800 can be found in the relevant description of method 800, and will not be repeated here. Specifically:

[0153] Corresponding to 802 above, in some embodiments, the process of enabling the AVS array to obtain the enhanced speech signal in the target direction is as follows:

[0154] First, it is understandable that in space, the sound signals collected by each AVS in the AVS array include not only the voice signal of user 10, but also sound signals such as noise and interference signals, such as signals emitted by other sound sources besides the sound source, background white noise, etc.

[0155] As mentioned earlier, each AVS consists of one omnidirectional microphone and three directional microphones. Therefore, the essence of performing a Fourier transform on the speech signal acquired by each AVS is actually performing a short-time Fourier transform on the speech signal (omnidirectional component) acquired by the omnidirectional microphone and the speech signals (X-axis component, Y-axis component, and Z-axis component) acquired by the three directional microphones in each AVS. The Fourier transform can be understood as performing a frequency domain transformation on the speech signal acquired by each AVS to obtain the frequency response of the signal in the frequency domain, where the frequency response includes the amplitude response and phase response of the signal in the frequency domain.

[0156] Now assume that the omnidirectional component, X-axis component, Y-axis component, and Z-axis component of the speech signal acquired by each AVS can be written in vector form after Fourier transform (7):

[0157]

[0158] Where y(ω) represents the speech signal acquired by each AVS, Y o (ω) is the omnidirectional component, Y i (ω), i∈{x, y, z}, are the x-axis component, y-axis component, and z-axis component, respectively; T represents the transpose operator; X(ω) represents the speech signal acquired by each AVS; V(ω) is the noise vector, defined similarly to y(ω); ω=2πf is the angular frequency, and f>0 is the frequency. θ is the pitch angle (ranging from 0° to 180°). The azimuth angle ranges from -180° to 180°.

[0159] Assuming the target direction is Apply a weighting vector (8) to the speech signal y(ω) acquired by each AVS:

[0160]

[0161] Among them, the target direction This indicates the desired direction for enhancing the speech signal, where T represents the transpose. The target direction coincides with the sound source direction. By applying a weighting vector in the target direction to each AVS, the enhanced spatial response in that target direction can be obtained. a0 and a1 (a1 = 1 - a0) are real coefficients used to adjust the beam shape in the target direction to obtain different spatial gains.

[0162] For example, assuming the target direction is (90°, 60°), the above equation (8) indicates that the audio acquisition device 1 will enhance the speech signal in the direction with a pitch angle of 90° and an azimuth angle of 60° in space.

[0163] a0 and a1 are the weighting adjustment coefficients for the omnidirectional and directional components in each AVS, used to adjust the proportion of the speech signal components collected by each microphone in each AVS in the speech signal collected by that AVS. That is, the beam shape of each AVS can be adjusted by adjusting the values ​​of a0 and a1. For example, assuming the target direction is still (90°, 60°), a0 = 1, then a1 = 0, then the above equation (8) can be written as: w(90°, 60°) = [1, 0, 0, 0], that is, the speech signal collected by the AVS at this time is only the enhanced signal of the speech signal in each direction in space collected by the omnidirectional microphone in the direction of pitch angle of 90° and azimuth angle of 60°.

[0164] Then, using the following formula (9), the target direction acquired by each AVS can be calculated. The enhanced speech signal Z(ω) above:

[0165]

[0166] At this point, the spatial response of each AVS is:

[0167]

[0168] Expanding equation (10) using equation (8) above, we can write it as:

[0169]

[0170] The outputs of the M AVS devices in the AVS array are then superimposed to obtain the output signal of the AVS array, which can be written as:

[0171] z(ω)=[Z 1(ω)Z 2(ω) …Z M(ω) ] T (12)

[0172] Substituting equations (9), (10), and (11) into equation (12), we can write:

[0173]

[0174] in, This represents the phase delay vector of the speech signal acquired by each AVS in the target direction.

[0175] J is the imaginary unit, J 2 =-1, and Where δ represents the spacing between each AVS in the AVS array.

[0176] as well as

[0177] v (ω)=[ V1 (ω V 2(ω )… V M(ω) ] T (16)

[0178] Corresponding to 803 above, in some embodiments, the spatial response of the beam is written as:

[0179]

[0180] Where H represents the conjugate transpose, h(ω) represents a linear filter of length M, M represents the number of AVSs in the AVS array, and h H (ω) is the conjugate transpose of h(ω), and d(ω, θ) and d(ω, θ) are... s The definition of ) is similar, and will not be repeated here.

[0181] White noise gain can be written as:

[0182]

[0183] Where, the value of α is:

[0184]

[0185] The directional factor can be written as:

[0186]

[0187] Among them, Γ d (ω) represents the normalized correlation matrix.

[0188]

[0189] Its (i, j)th element can be written as:

[0190]

[0191] Where, γ ij express, Substituting equation (23) into equation (22), we can derive:

[0192]

[0193] Where sin h(γ) ij The value of ) is:

[0194]

[0195] cos h(γ ij The value of ) is:

[0196]

[0197] The value of ξ1 is:

[0198]

[0199] The values ​​of ξ2 and -ξ4 are:

[0200] ξ2=-ξ4=-2a1a0 cosθ s (28)

[0201] The values ​​of ξ3 and -ξ5 are:

[0202]

[0203] Corresponding to 804 above, in some embodiments, after determining the various parameters of the super-directional beamformer 13, the super-directional beamformer will filter the speech signal output by the AVS array to obtain an enhanced speech signal in the target direction.

[0204] Specifically, the super-directing beamformer will suppress the target direction. While eliminating interference noise, the target signal in the target direction is recovered without distortion. Specifically, the distortion-free constraint in the target direction is assumed to be:

[0205]

[0206] Since it is assumed that there is no distortion in the target direction, the response of the AVS array in the target direction is... Therefore, the above equation (30) is actually:

[0207] That is, h H (ω)d(ω,θ s ) = 1.

[0208] At this point, equation (17) can be written as:

[0209]

[0210] As can be seen from equation (31) above, when the coordinate system and the position of the AVS are determined, along the azimuth angle The beam pattern of the direction depends only on the directivity of the AVS. It will not change with the filter h H (ω) changes. Therefore, beamforming that maximizes the directivity factor can be achieved by maximizing the directivity factor while constraining the main lobe direction and elevation angle of the beam with respect to θ. s This was consistently obtained. The main lobe direction refers to the actual direction the beam points after beamforming, i.e., the direction of maximum spatial response. Mathematically, the constraints on the main lobe direction and pitch angle of beamforming are:

[0211]

[0212] The left side of the equation can be derived as follows:

[0213]

[0214] In the above formula (33) Given a diagonal matrix, its (i, i)th element is:

[0215]

[0216] when The above formula can be simplified to:

[0217]

[0218] Substituting equation (34) into equation (33), the derivative constraint of equation (32) can finally be written as: h H (ω)∑ M (θ s )d(ω,θ s )=0(35)

[0219] When the target direction is the end-fire direction, i.e., the elevation angle θ s =0°, when 1≤i≤M, Maximizing the directivity factor under the distortion-free constraint of equation (30) is equivalent to solving the following optimization problem:

[0220] min h(ω) hH (ω)Γ d (ω)h(ω)subjeot to h H (ω)d(ω,θ s )=1(36)

[0221] The solution to equation (36) above is: Then, by estimating the AVS array using the filter h(ω), super-directional beamforming can be obtained.

[0222] In some embodiments, when the target direction is another direction, i.e., the pitch angle θ s ≠0°, or a derivative constraint can be added to ensure that the maximum response occurs in the target direction. Above. At this point, it is equivalent to solving the following optimization problem:

[0223] min h(ω) h H (ω)Γ d (ω)h(ω)stC H (ω,θ s h(ω)=i1 (38)

[0224] Where, C(ω, θ) s )=[d(ω,θ s )∑ M (θ s )d(ω,θ s )](39), indicating that it is an M×2 matrix, i1=[1 0] T Solving equation (38) above yields:

[0225]

[0226] In (40) above, the autocorrelation matrix Γ d (ω) can be diagonally loaded with Γ d (ω)+∈I M , where ∈≥0 is a regularization factor that controls diagonal loading to improve the robustness of the super-directive beamformer.

[0227] Finally, z(ω) is passed through a beamformer h(ω), which has built-in fixed filter coefficients, to obtain the enhanced speech signal.

[0228]

[0229] Where H represents the conjugate transpose operator.

[0230] Regarding the above The enhanced speech signal can be obtained by performing an inverse Fourier transform.

[0231] The above details the implementation of each step in method 800. It is understood that in other embodiments of this application, the above formulas may also be other formulas known to those skilled in the art that can achieve the corresponding functions, and this application does not limit them.

[0232] It is understood that after the speech signal in the target direction acquired by the AVS array is processed by the above-mentioned method 800 or 1000, an enhanced speech signal in the target direction can be obtained. However, there may still be noise or other interference signals in the speech signal. Therefore, in order to obtain a pure speech signal, the audio acquisition device 1 of this application may also include a noise suppressor 14 for noise suppression processing of the enhanced speech signal output by the super-directional beamformer 13. Among them, the noise processing mainly includes two methods: (1) performing diffusion field noise suppression on the enhanced signal in the target direction after processing by the above-mentioned method 800 or 1000; (2) performing nonlinear filtering on the enhanced signal in the target direction after processing by the above-mentioned method 800 or 1000. It is understood that in some embodiments, the above two noise processing methods can be selected to be executed or can be executed simultaneously. That is, only method (1) can be executed, only method (2) can be executed, or both methods (1) and (2) can be executed simultaneously. This application does not limit this.

[0233] The following is a detailed explanation:

[0234] (1) Suppression of diffused field noise:

[0235] For each AVS, direct sound and diffuse noise can be distinguished based on the energy relationship of the speech signal reaching each channel of the same AVS. In this embodiment, the signal of each AVS is stored in B-Format (FuMa) format, and the acquired signals of the omnidirectional channel and the three axial channels of x, y, and z represent X, Y, and Z, respectively. w X x X y X z .

[0236] When space is a perfectly diffused field, that is, when signals with the same energy but no correlation are collected from all directions of space, the relationship satisfies the following equation (42):

[0237] X w 2 =Xx 2 +X y 2 +Xz 2 (42)

[0238] When the space contains only point source noise located in the X-axis direction, the signal X acquired by the omnidirectional channel...w =X x X w z =3X x 2 (The same applies to sound sources in the y-axis or any direction in three-dimensional space).

[0239] At this point, the energy relationship of the signals acquired between the channels can be used as a basis, i.e. To determine whether the speech signal acquired by AVS at each time frequency point is point source noise or diffuse field noise, (1,3) is mapped to (0,1) through a Gaussian relation, and the diffuse field noise is filtered and suppressed accordingly.

[0240] In some embodiments, since phase information is introduced after forming an AVS array, the noise filtering coefficients of the diffused field calculated by each AVS can be balanced and smoothed, thereby achieving higher quality and lower distortion noise suppression.

[0241] In some embodiments, the implementation of the AVS device includes, but is not limited to, hot-wire type, differential pressure type, A-format hybrid type, etc.; in some embodiments, the directivity of each AVS includes, but is not limited to, first-order directivity (0≤a0≤1), which can be regressed to zero-order (omnidirectional), or it can be other higher-order directional pickup devices.

[0242] In some embodiments, the design of the super-directional beamformer 13 can be optimized in multiple steps as in Embodiment 1 (directional beamformer 12 × super-directional beamformer 13, i.e., the method 800 described above), or the design can be optimized globally for all pickup channels according to the calculation method of the super-directional beamformer 13 (without the directional beamformer step, the method 1000 described above).

[0243] (2) Nonlinear beam filtering

[0244] For each AVS, the direction of arrival at each time frequency point can be calculated based on the sound intensity vector acquired by the AVS. The sound intensity vector of each AVS can be expressed as:

[0245]

[0246] Where (f, n) represents the time frequency point with frequency f and frame number n, the orientation of this frequency point can be expressed as:

[0247]

[0248] in The real part is taken, and the difference (0-180°) between the time frequency point and the target direction is mapped to the filter coefficients (1-0) through a Gaussian function. This allows for nonlinear beamforming of the signal to achieve directional sound acquisition, suppress interference noise and reverberation from other directions, and improve signal quality.

[0249] In some embodiments, since phase information is introduced after forming an AVS array, the noise filtering coefficients of the diffused field calculated by each AVS can be balanced and smoothed, thereby achieving higher quality and lower distortion noise suppression.

[0250] It is understood that the operation of the noise suppressor 14 is not limited to noise suppression, linear filtering and sound field reconstruction of the signal in the target direction, but can also perform reverberation suppression, interference suppression and so on of the signal in the target direction. This application does not limit the function of the noise suppressor 14.

[0251] Furthermore, in some embodiments, this application does not limit the mapping relationships involved in the noise suppressor 13, such as the noise suppression of the diffuse field and the mapping of the filter coefficients in the nonlinear beam, including but not limited to Gaussian mapping, linear mapping, piecewise linear mapping, logarithmic mapping, etc.

[0252] After the above series of processing steps, the audio acquisition device 1 can output a more refined and enhanced objectified speech signal with less noise and interference. When the voice interaction device 20a plays the speech signal output by the audio acquisition device 1, it can reconstruct the sound field based on the speech signal. That is, according to the user's selection or default playback position, the speech signal output by the audio acquisition device 1 is projected onto the stereo channels according to the settings, forming a stereo speech signal and restoring or virtually reconstructing the spatial distribution of the sound source. It can be understood that stereo channels include, but are not limited to, stereo, 5.1, and 7.1 channels. Among them, 5.1 channels are stereo surround sound, and 7.1 channels have two more channels than 5.1, making it a more powerful channel system.

[0253] For example, taking the voice interaction devices 20a and 20b with the aforementioned audio acquisition device 1 as an example, where voice interaction device 20a is a microphone and voice interaction device 20b is an earphone, microphone 20a acquires the meeting recordings of participants 10a to 10d through audio acquisition device 1, and then microphone 20a sends the meeting recording to earphone 20b. When user 10 plays the meeting recording through earphone 20b, earphone 20b can reproduce the meeting recording in space, simulating the real meeting scene corresponding to the meeting recording, so that user 10 can feel the real meeting scene when playing the meeting recording through earphone 20b, thus improving the user experience.

[0254] In some embodiments, the earphone 20b can use a head-related transfer function (HRTF) to reconstruct the sound field. The HRTF describes all the acoustic information required for a speech signal to be reflected or diffracted to the listener's head and outer ear before entering the listener's (e.g., user 10) auditory system. Essentially, it is a filter. By using a position- or angle-related HRTF to filter a specific speech signal, the effect of the speech signal being played at that position or angle can be simulated.

[0255] In some embodiments, researchers can use an HRTF measurement system to collect HRTFs of different users at different angles in advance, and then use big data to obtain angle-related HRTFs that can meet the needs of ordinary users. Then, when the headphones 20b are reconstructing the sound field, the position Ω3 of the voice signal relative to the user's head angle can be determined based on the user's head angle Ω1 and the pre-set target direction Ω2 for playing the voice signal. Then, the HRTF that matches Ω3 is selected as a filter to filter the voice signal to be played, simulating the scenario where the filtered voice signal is played from the relative angle Ω3 relative to the user's head angle.

[0256] In some embodiments, the earphone 20b may also be equipped with a head posture sensor to detect the user's real-time head angle, so that when the user's head angle changes, for example, when the user turns their head to Ω4, the earphone 20b can still determine the relative angle Ω5 between the user's head position Ω4 and the target direction Ω2, and simulate the scenario where the voice signal is played from the relative angle Ω5 relative to the user's head angle. This application does not limit this.

[0257] Figure 19 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device can be a voice interaction device, including but not limited to: smartphones, laptops, tablets, smart speakers, smart in-vehicle devices, smart robots, smart home devices, smart headphones, etc. The electronic device can also be a server that communicates with the voice interaction device.

[0258] The hardware includes a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, a speaker 170A, an AVS array 170B, a sensor module 180, buttons 190, and a display screen 194, etc. The sensor module 180 may include a pressure sensor 180A, a touch sensor 180K, etc.

[0259] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware. For example, if the electronic device is a headset, it may not have a display screen 194, and the sensor module 180 may be equipped with a posture sensor to detect the direction of head movement, etc., which is not a limitation of this application.

[0260] The processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0261] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0262] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0263] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0264] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.

[0265] The charging management module 140 is used to collect charging input from the charger. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 collects input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, display 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance).

[0266] The wireless communication function of electronic devices can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and collect electromagnetic wave signals. Mobile communication module 150 can provide wireless communication solutions for electronic devices, including 2G / 3G / 4G / 5G. Wireless communication module 160 can provide wireless communication solutions for electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies.

[0267] The display screen 194 is used to display images, videos, etc. In some embodiments, the electronic device may include one or N display screens 194, where N is a positive integer greater than 1.

[0268] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The internal memory 121 can be used to store computer executable program code, which includes instructions.

[0269] The electronic device 100 can play the enhanced speech signal described above through the speaker 170A, or use the speech signal to reconstruct the sound field, etc., and this application does not limit this.

[0270] The electronic device 100 can acquire voice signals through the aforementioned AVS array and enhance the voice signals in the target direction using the processor 110.

[0271] Pressure sensor 180A can be used to detect the user's touch pressure on the electronic device to determine the type of touch operation, such as long press, short press, or heavy press. In some embodiments, pressure sensor 180A works in conjunction with touch sensor 180K. Touch sensor 180K, also known as a "touch device," can be located on display screen 194. Touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touchscreen." Touch sensor 180K is used to detect touch operations applied to or near it. For example, touch sensor 180K on touch headphones can detect user touch signals to enable play / pause functions. The touch sensor can transmit the detected touch operation to the application processor to determine the touch event type. Visual output related to the touch operation can be provided through display screen 194. Buttons 190 include power buttons, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can collect button inputs and generate key signal inputs related to user settings and function control of the electronic device.

[0272] This application also provides a computer-readable storage medium, including: a computer program stored thereon, which, when executed by a processor, implements the audio acquisition method described in any of the above method embodiments.

[0273] All or part of the steps in the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof.

[0274] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0275] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0276] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0277] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. An audio acquisition method applied to an electronic device, comprising: The electronic device comprises at least two directional audio collectors and at least one omnidirectional audio collector; The method comprises: For the first directional audio data collected by the at least two directional audio collectors, increasing the proportion of the effective audio data in the first directional audio data collected by the directional audio collector satisfying the first preset condition with the target direction to determine the second directional audio data; For the first omnidirectional audio data collected by the at least one omnidirectional audio collector, increasing the proportion of the effective audio data in the first omnidirectional audio data satisfying the second preset condition with the target direction to determine the second omnidirectional audio data; Determining the audio collection data of the electronic device in the target direction, wherein the audio collection data is obtained by superimposing the second directional audio data and the second omnidirectional audio data; and Determining the parameters of the super-directivity beamformer, constructing the super-directivity beamformer, and filtering the audio collection data based on the super-directivity beamformer to obtain the enhanced voice signal, wherein the parameters of the super-directivity beamformer comprise at least one of the following: beam spatial response, white noise gain, and directivity factor.

2. The method of claim 1, wherein, The electronic device comprises a plurality of acoustic vector sensors, wherein the omnidirectional pressure sensor in the acoustic vector sensors is used as the omnidirectional audio collector, and the directional sound pressure gradient sensor in the acoustic vector sensors is used as the directional audio collector.

3. The method of claim 2, wherein, The plurality of acoustic vector sensors are arranged according to a preset rule, and Each acoustic vector sensor comprises one omnidirectional pressure sensor and at least three directional sound pressure gradient sensors.

4. The method of claim 3, wherein, The increasing of the proportion of the effective audio data in the first directional audio data collected by the directional audio collector satisfying the first preset condition with the target direction comprises: According to the target direction, determining the weight parameter of the first directional audio data collected by each directional sound pressure gradient sensor; Increasing the weight parameter of the first directional audio data collected by the directional sound pressure gradient sensor satisfying the first preset condition with the target direction to increase the proportion of the effective audio data in the first directional audio data collected by the directional sound pressure gradient sensor satisfying the first preset condition with the target direction.

5. The method of claim 3, wherein, The increasing of the proportion of the effective audio data in the first directional audio data collected by the directional audio collector satisfying the first preset condition with the target direction comprises: According to the target direction, determining the weight parameter of the first directional audio data collected by each directional sound pressure gradient sensor; By adjusting the beam spatial response, the directivity factor, and the white noise gain, the weight parameter of the audio data collected by the directional sound pressure gradient sensor satisfying the first preset condition with the target direction is increased, so that the proportion of the effective audio data in the first directional audio data collected by the directional sound pressure gradient sensor satisfying the first preset condition with the target direction is increased.

6. The method according to any one of claims 3 to 5, characterized in that, The increasing the proportion of the effective audio data in the first omnidirectional audio data collected by the omnidirectional audio collector satisfying the second preset condition with the target direction specifically comprises: According to the target direction, determining a weight parameter of the first omnidirectional audio data collected by the omnidirectional pressure sensor; Increasing the weight parameter of the first omnidirectional audio data collected by the omnidirectional pressure sensor satisfying the second preset condition with the target direction, so as to increase the proportion of the effective data in the first omnidirectional audio data collected by the omnidirectional pressure sensor satisfying the second preset condition with the target direction.

7. The method according to any one of claims 3 to 5, characterized in that, The increasing the proportion of the effective audio data in the first omnidirectional audio data collected by the omnidirectional audio collector satisfying the second preset condition with the target direction specifically comprises: According to the target direction, determining a weight parameter of the first omnidirectional audio data collected by the omnidirectional pressure sensor; By adjusting the beam spatial response, the directivity factor and the white noise gain, the weight parameter of the audio data collected by the omnidirectional pressure sensor satisfying the second preset condition with the target direction is increased, so as to increase the proportion of the effective audio data in the first omnidirectional audio data collected by the omnidirectional pressure sensor satisfying the second preset condition with the target direction.

8. The method of claim 2 or 3, wherein, The electronic device comprises a plurality of acoustic vector sensors, and the target direction can be determined by the following manner: Using the time difference of the first directional audio data and / or the omnidirectional audio data collected by at least two acoustic vector sensors, determining the phase difference of the first directional audio data and / or the omnidirectional audio data collected by the at least two acoustic vector sensors; Based on the phase difference, determining the target direction.

9. The method of claim 1, wherein, The case that the target direction satisfies the first preset condition comprises: The direction difference between the collection direction of the directional audio collector and the target direction is less than a first preset value.

10. The method of claim 1, wherein, The case that the collection direction and the target direction satisfy the second preset condition comprises: The direction difference between the collection direction and the target direction is less than a second preset value.

11. An electronic device, comprising: The electronic device comprises a processor, a memory, at least two directional audio collectors and at least one omnidirectional audio collector; and The memory is used for storing instructions executed by one or more processors of the electronic device; And the processor is one of the processors of the electronic device, and is used for running the instructions to enable the electronic device to implement the following operations: For the first directional audio data collected by the at least two directional audio collectors, increasing the proportion of the effective audio data in the first directional audio data collected by the directional audio collector satisfying the first preset condition with the target direction to determine second directional audio data; For the first omnidirectional audio data collected by the at least one omnidirectional audio collector, increasing the proportion of the effective audio data in the first omnidirectional audio data satisfying the second preset condition with the target direction to determine second omnidirectional audio data; Determining the audio collection data of the electronic device in the target direction, wherein the audio collection data comprises the second directional audio data and the second omnidirectional audio data superimposed; and, The parameters of the super-directive beamformer are determined, the super-directive beamformer is constructed, and the audio collection data is filtered based on the super-directive beamformer to obtain enhanced voice signals, wherein the parameters of the super-directive beamformer include at least one of a beam spatial response, a white noise gain, and a directivity factor.

12. The electronic device of claim 11, wherein, The electronic device includes a plurality of acoustic vector sensors, an omnidirectional pressure sensor in the acoustic vector sensors serving as the omnidirectional audio collector, and a directional sound pressure gradient sensor in the acoustic vector sensors serving as the directional audio collector.

13. The electronic device of claim 12, wherein, The plurality of acoustic vector sensors are arranged according to a preset rule, and Each of the acoustic vector sensors includes an omnidirectional pressure sensor and at least three directional sound pressure gradient sensors.

14. The electronic device of claim 13, wherein, The ratio of the effective audio data in the first directional audio data collected by the directional audio collector that meets the first preset condition with the target direction is increased, specifically including: According to the target direction, a weight parameter of the first directional audio data collected by each of the directional sound pressure gradient sensors is determined; The weight parameter of the first directional audio data collected by the directional sound pressure gradient sensor that meets the first preset condition with the target direction is increased to increase the ratio of the effective data in the first directional audio data collected by the directional sound pressure gradient sensor that meets the first preset condition with the target direction.

15. The electronic device of claim 13, wherein, The ratio of the effective audio data in the first directional audio data collected by the directional audio collector that meets the first preset condition with the target direction is increased, specifically including: According to the target direction, a weight parameter of the first directional audio data collected by each of the directional sound pressure gradient sensors is determined; The weight parameter of the audio data collected by the directional sound pressure gradient sensor that meets the first preset condition with the target direction is increased by adjusting the beam spatial response, the directivity factor, and the white noise gain, thereby increasing the ratio of the effective audio data in the first directional audio data collected by the directional sound pressure gradient sensor that meets the first preset condition with the target direction.

16. The electronic device of claim 13, wherein, The ratio of the effective audio data in the first omnidirectional audio data collected by the omnidirectional audio collector that meets the second preset condition with the target direction is increased, specifically including: According to the target direction, a weight parameter of the first omnidirectional audio data collected by the omnidirectional pressure sensor is determined; The weight parameter of the first omnidirectional audio data collected by the omnidirectional pressure sensor that meets the second preset condition with the target direction is increased to increase the ratio of the effective data in the first omnidirectional audio data collected by the omnidirectional pressure sensor that meets the second preset condition with the target direction.

17. The electronic device of claim 13, wherein, The ratio of the effective audio data in the first omnidirectional audio data collected by the omnidirectional audio collector that meets the second preset condition with the target direction is increased, specifically including: According to the target direction, a weight parameter of the first omnidirectional audio data collected by the omnidirectional pressure sensor is determined; The weight parameter of the first omnidirectional audio data collected by the omnidirectional pressure sensor that meets the second preset condition with the target direction is increased to increase the ratio of the effective data in the first omnidirectional audio data collected by the omnidirectional pressure sensor that meets the second preset condition with the target direction. The weight parameter of the audio data collected by the omnidirectional pressure sensor satisfying the second preset condition with the target direction is increased by adjusting the beam space response, the directivity factor and the white noise gain, so that the proportion of the effective audio data in the first omnidirectional audio data collected by the omnidirectional pressure sensor satisfying the second preset condition with the target direction is increased.

18. The electronic device of claim 11, wherein, The electronic device comprises a plurality of acoustic vector sensors, and the target direction can be determined by the following method: The phase difference of the first directional audio data and / or the omnidirectional audio data collected by at least two acoustic vector sensors is determined by using the time difference of the first directional audio data and / or the omnidirectional audio data collected by the at least two acoustic vector sensors; The target direction is determined based on the phase difference.

19. The electronic device of claim 11, wherein, The case that the target direction satisfies the first preset condition includes: The direction difference between the collection direction of the directional audio collector and the target direction is less than a first preset value.

20. The electronic device of claim 11, wherein, The case that the collection direction satisfies the second preset condition with the target direction includes: The direction difference between the collection direction and the target direction is less than a second preset value.

21. A computer readable medium characterized by The computer readable medium stores instructions, which, when executed on the electronic device, cause the electronic device to perform the audio collection method of any one of claims 1 to 10.

Citation Information

Patent Citations

  • Sound pickup device, sound pickup and sound source separating device and method for picking up sound, method for picking up sound and separating sound source and recording medium for recording sound pickup program, sound pickup and sound source separating program

    JP2002084590A