Audio acquisition method and apparatus

The audio acquisition method using an omnidirectional microphone array with DOA-based weight calculation addresses the inflexibility of 4-in-1 microphones, allowing customizable audio processing and improved signal capture.

WO2025222433A1PCT designated stage Publication Date: 2025-10-30HARMAN INT IND INC +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/089792
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing 4-in-1 microphones require a strict structural design to maintain good frequency response and directional performance, and their audio acquisition patterns are fixed and cannot be flexibly adjusted according to user requirements, leading to poor performance and low signal-to-noise ratio.

Method used

An audio acquisition method using a microphone array of omnidirectional microphones that allows for free switching of audio acquisition modes through weight calculation based on direction of arrival (DOA) and correlation parameters, enabling customizable audio signal processing.

Benefits of technology

Enables flexible adjustment of audio acquisition modes to enhance signals from specified directions and accurately capture moving targets, achieving higher signal-to-noise ratio without complex structural designs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024089792_30102025_PF_FP_ABST
    Figure CN2024089792_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio acquisition method, an audio acquisition apparatus and a device. The audio acquisition method includes: acquiring a set of audio signals using a microphone array comprising a plurality of omnidirectional microphones arranged in a predetermined rule, the set of audio signals including an audio signal acquired by each omnidirectional microphone of the microphone array; receiving a selection of an audio acquisition mode of a plurality of audio acquisition modes; determining a set of weights corresponding to the selected audio acquisition mode, where each weight of the set of weights corresponds to each microphone of the microphone array and is computed based on a weight coefficient associated with a particular direction of arrival (DOA); processing audio signals of the set of audio signals using the set of weights to generate a processed audio signal; and outputting the processed audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

AUDIO ACQUISITION METHOD AND APPARATUSTECHNICAL FIELD

[0001] The present disclosure relates to a field of audio processing, and in particular, to an audio acquisition method, an audio acquisition apparatus and a device.BACKGROUND

[0002] A microphone, also known as a transducer, is an energy conversion device that converts a sound signal into an electrical signal. To accommodate different scenarios and applications, microphones have developed a variety of audio recording patterns, such as a cardioid pattern, an omnidirectional pattern, a stereo pattern, a bidirectional pattern, and the like.

[0003] At an early stage, different microphones need to be used to realize the different audio recording patterns described above and usually require complicated setup and equipment. Then the technology was developed to combine these four patterns into one microphone product. The state-of-art 4-in-1 technique typically uses four unidirectional microphones, e.g., a first microphone facing front, a second microphone facing left, a third microphone facing right, and a fourth microphone facing rear, and then signals of the different microphones are combined by hardware to achieve a final audio signal output. However, such a 4-in-1 microphone in the prior art requires a strict structural design to maintain a good frequency response and directional performance, and its audio acquisition patterns are fixed and cannot be flexibly adjusted according to user requirements.

[0004] SUMMARY OF THE DISCLOSURE

[0005] The present disclosure proposes an audio acquisition method, an audio acquisition apparatus and device, which enables free switching of a plurality of audio acquisition modes with a simple structural design.

[0006] According to an aspect of the present disclosure, there is provided an audio acquisition method, comprising: acquiring a set of audio signals using a microphone array comprising a plurality of omnidirectional microphones arranged in a predetermined rule, the set of audio signals comprising an audio signal acquired by each omnidirectional  microphone of the microphone array; receiving a selection of an audio acquisition mode of a plurality of audio acquisition modes; determining a set of weights corresponding to the selected audio acquisition mode, wherein each weight of the set of weights respectively corresponds to each microphone of the microphone array and is computed based on a weight coefficient associated with a particular direction of arrival (DOA) ; processing audio signals of the set of audio signals using the set of weights to generate a processed audio signal; and outputting the processed audio signal.

[0007] In one or more embodiments of the present disclosure, wherein determining the set of weights corresponding to the selected audio acquisition mode comprises: determining a correlation parameter representing correlation between respective audio signals of the set of audio signals; and determining the set of weights corresponding to the selected audio acquisition mode based at least on the correlation parameter.

[0008] In one or more embodiments of the present disclosure, wherein the selected audio acquisition mode comprises a customizable acquisition mode configured to enhance an audio signal from a specified direction of arrival in the processed audio signal, and wherein determining the set of weights corresponding to the selected audio acquisition mode based at least on the correlation parameter comprises: receiving the specified direction of arrival; and determining the set of weights corresponding to the customizable acquisition mode based on the specified direction of arrival and the correlation parameter.

[0009] In one or more embodiments of the present disclosure, wherein the selected audio acquisition mode comprises a DOA estimation mode configured to enhance an audio signal from a target sound source in the processed audio signal, and wherein determining the set of weights corresponding to the selected audio acquisition mode based at least on the correlation parameter comprises: detecting or tracking a current DOA of the audio signal from the target sound source based on audio signals acquired by the plurality of omnidirectional microphones of the microphone array from the target sound source; and determining the set of weights corresponding to the DOA estimation mode based on the current DOA and the correlation parameter.

[0010] In one or more embodiments of the present disclosure, wherein the selected audio acquisition mode comprises a cardioid acquisition mode configured to enhance an audio signal from a front of the microphone array and suppress an audio signal from a  rear of the microphone array in the processed audio signal, and wherein determining the set of weights corresponding to the selected audio acquisition mode comprises: determining a first weight coefficient associated with an audio signal in the set of audio signals from the front of the microphone array as a first value and a second weight coefficient associated with an audio signal in the set of audio signals from the rear of the microphone array as a second value, wherein the first value is greater than the second value; and determining the set of weights corresponding to the cardioid acquisition mode based at least on the first weight coefficient and the second weight coefficient.

[0011] In one or more embodiments of the present disclosure, wherein the selected audio acquisition mode comprises an omnidirectional acquisition mode configured to output audio signals from all directions of arrival equally in the processed audio signal, and wherein determining the set of weights corresponding to the selected audio acquisition mode comprises: determining, in the set of weights, weights corresponding to the respective omnidirectional microphones of the microphone array as an equal value.

[0012] In one or more embodiments of the present disclosure, wherein the selected audio acquisition mode comprises a stereo acquisition mode configured to enhance audio signals from front left and front right of the microphone array in the processed audio signal for provision into a first channel and a second channel, respectively, and wherein determining the set of weights corresponding to the selected audio acquisition mode comprises: determining a first set of weights for the first channel and a second set of weights for the second channel, respectively.

[0013] In one or more embodiments of the present disclosure, wherein determining the first set of weights for the first channel and the second set of weights for the second channel, respectively, comprises: determining a first weight coefficient associated with an audio signal in the set of audio signals from the front left of the microphone array as a first value and a second weight coefficient associated with an audio signal in the set of audio signals from a rear right of the microphone array as a second value, the first value being greater than the second value, and determining the first set of weights based at least on the first weight coefficient and the second weight coefficient; and determining a third weight coefficient associated with an audio signal in the set of audio signals from a front right of the microphone array as a third value and a fourth weight coefficient associated  with an audio signal in the set of audio signals from a rear left of the microphone array as a fourth value, the third value being greater than the fourth value, and determining the second set of weights based at least on the third weight coefficient and the fourth weight coefficient.

[0014] In one or more embodiments of the present disclosure, wherein the selected audio acquisition mode comprises a bidirectional acquisition mode configured to enhance audio signals from front and rear of the microphone array and suppress audio signals from left and right of the microphone array in the processed audio signal, and wherein determining the set of weights corresponding to the selected audio acquisition mode comprises: determining a first weight coefficient and a second weight coefficient respectively associated with audio signals in the set of audio signals from the front and rear of the microphone array as a first value, and determining a third weight coefficient and a fourth weight coefficient respectively associated with audio signals in the set of audio signals from the left and right of the microphone array as a second value, the first value being greater than the second value, and determining the set of weights corresponding to the bidirectional acquisition mode based on the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient.

[0015] In one or more embodiments of the present disclosure, wherein processing audio signals of the set of audio signals using the set of weights to generate the processed audio signal comprises: performing phase delay compensation on the audio signals of the set of audio signals; and performing a weighted summation on the compensated audio signals using the set of weights to generate the processed audio signal.

[0016] In one or more embodiments of the present disclosure, wherein the predetermined rule comprises any one or any combination of a cross shape, a plane shape, a square shape, a circular shape, and a spiral shape.

[0017] According to another aspect of the present disclosure, there is provided an audio acquisition apparatus, comprising: an acquisition unit configured to acquire a set of audio signals using a microphone array comprising a plurality of omnidirectional microphones arranged in a predetermined rule, the set of audio signals comprising an audio signal acquired by each omnidirectional microphone of the microphone array; an  input unit configured to receive a selection of an audio acquisition mode of a plurality of audio acquisition modes; a processing unit configured to determine a set of weights corresponding to the selected audio acquisition mode, wherein each weight of the set of weights respectively corresponds to each microphone of the microphone array and is computed based on a weight coefficient associated with a particular DOA, and process each audio signal of the set of audio signals using the set of weights to generate a processed audio signal; and an output unit configured to output the processed audio signal.

[0018] According to another aspect of the present disclosure, there is provided an audio acquisition device, comprising: one or more processors; and one or more memories, wherein the one or more memories have stored therein computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to perform the method described above.

[0019] With the audio acquisition method, audio acquisition apparatus and device in the above aspects of the disclosure, a microphone array composed of omnidirectional microphones may be used to realize free switching of various audio acquisition modes, and a higher signal-to-noise ratio may be obtained without complicated structural design. In particular, an audio acquisition mode may be customized according to user requirements to enhance audio signals in a specified DOA, and a DOA estimation mode may be realized to accurately capture sounds of a moving target.BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above and other objects, features and advantages of embodiments of the present disclosure will become obvious from the following detailed description of embodiments of the present disclosure taken in conjunction with accompanying drawings. The accompanying drawings are used to provide further understanding of the embodiments of the present disclosure, constitute a part of the specification, explain the present disclosure together with the embodiments of the present disclosure, and do not constitute a limitation of the present disclosure. In the drawings, like reference numerals generally represent like components or steps.

[0021] FIG. 1 illustrates a flow diagram of an audio acquisition method in accordance with one or more embodiments of the present disclosure;

[0022] FIG. 2 illustrates an example arrangement of a 4-element uniform circular microphone array (UCA) in accordance with one or more embodiments of the present disclosure;

[0023] FIG. 3 illustrates a flow diagram of a method for determining a set of weights corresponding to an audio acquisition mode in accordance with one or more embodiments of the present disclosure;

[0024] FIG. 4 illustrates an example beam pattern of a customizable acquisition mode in accordance with one or more embodiments of the present disclosure;

[0025] FIG. 5 illustrates an example beam pattern of a DOA estimation mode in accordance with one or more embodiments of the present disclosure;

[0026] FIG. 6 illustrates an example beam pattern of cardioid acquisition mode in accordance with one or more embodiments of the present disclosure;

[0027] FIG. 7 illustrates an example beam pattern of an omnidirectional acquisition mode in accordance with one or more embodiments of the present disclosure;

[0028] FIG. 8 illustrates an example beam pattern of a stereo acquisition mode in accordance with one or more embodiments of the disclosure;

[0029] FIG. 9 illustrates an example beam pattern of a bidirectional acquisition mode in accordance with one or more embodiments of the present disclosure;

[0030] FIG. 10 illustrates a schematic structural diagram of an audio acquisition apparatus in accordance with one or more embodiments of the disclosure.

[0031] DESCRIPTION OF THE EMBODIMENTS

[0032] In order to make objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and thoroughly with reference to the accompanying drawings. Obviously, these described embodiments are only a part of the present disclosure, not all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without paying creative efforts fall into the protection scope of the present disclosure.

[0033] As used herein and in the claims, the words “a, ” “an, ” “an, ” and / or “the” do not refer to the singular, but may include the plural unless the context clearly dictates otherwise. In  general, the terms “comprise” and “comprising” only imply the inclusion of steps and elements specifically identified, these steps and elements do not constitute an exclusive list and a method or apparatus may also contain other steps or elements.

[0034] Flowcharts are used herein to illustrate steps of a method according to one or more embodiments of the present disclosure. It should be understood that preceding or subsequent steps do not have to be performed exactly in order. Rather, various steps may be processed in reverse order or simultaneously, as desired. Meanwhile, other steps may also be added to the method, or certain step or steps may be removed from the method.

[0035] As used in one or more embodiments of the present disclosure, an audio acquisition mode refers to a mode in which a microphone product captures and outputs audio signals, which may also be referred to, for example, as a beam pattern, a recording pattern, and the like. Typically, audio acquisition modes may include a cardioid acquisition mode, an omnidirectional acquisition mode, a stereo acquisition mode, and a bidirectional acquisition mode. Current 4-in-1 microphones may achieve the four audio acquisition modes described above in one product by combining signals from four unidirectional microphones. However, unidirectional microphones usually have a lower signal-to-noise ratio (SNR) , and as noted above, such 4-in-1 microphones require a rigorous structural design to maintain good frequency response and directional performance, and tend to perform poorly at high frequencies. In addition, the audio acquisition modes of the existing 4-in-1 microphone are fixed and cannot be flexibly adjusted according to user requirements.

[0036] To solve the above problems, the present disclosure proposes an audio acquisition method, an audio acquisition apparatus and device, which enables free switching of multiple audio capturing modes with a simple structural design. In particular, an audio acquisition mode may be customized according to user demands to enhance audio signals from a specified direction of arrival (DOA) , and a DOA estimation mode may be realized to precisely capture sounds of a moving target. An audio acquisition method according to one or more embodiments of the present disclosure will be described below with reference to FIG. 1. The audio acquisition method according to one or more embodiments of the present disclosure may be implemented, for example, by an audio acquisition apparatus or device, or by any other terminal or device (e.g., a computer, a cell phone, a smart wearable device, a television set, etc. ) including an audio acquisition component, which is not limited in the embodiments of the  present disclosure.

[0037] FIG. 1 shows a flow diagram of an audio acquisition method 100 according to one or more embodiments of the present disclosure. As shown in FIG. 1, in step S102, a set of audio signals is acquired using a microphone array comprising a plurality of omnidirectional microphones arranged in a predetermined rule. As mentioned earlier, an omnidirectional microphone has the same sensitivity to sound in all directions, i.e. audio signals from all directions may be captured indifferently. The microphone array according to one or more embodiments of the present disclosure may include a plurality of omnidirectional microphones, for example, 4 omnidirectional microphones, or a greater or fewer number of microphones, which is not particularly limited by the embodiments of the present disclosure. The respective omnidirectional microphones in the microphone array may be identical, or perfectly calibrated with respect to each other. The plurality of omnidirectional microphones in the microphone array may be arranged, for example, in a predetermined rule such as a cross shape, a plane shape, a square shape, a circular shape, a spiral shape, and the like, which is not particularly limited by the embodiments of the present disclosure. FIG. 2 shows an example arrangement of a 4-element circular microphone array (UCA) in which 4 omnidirectional microphones are uniformly arranged on a circle centered on the coordinate origin and the radius of the circle is about 0.03 m, in accordance with one or more embodiments of the present disclosure. It should be noted that the number, arrangement, relative position, etc. of the microphones of the microphone array shown in FIG. 2 are merely by way of example and are not to be construed as limiting in any sense.

[0038] Furthermore, in one or more embodiments of the present disclosure, the microphone array herein may include one or more of omnidirectional microphones, unidirectional microphones, bidirectional microphones, or any combination thereof. Hereinafter, the present disclosure will be described in the case where the microphone array includes a plurality of omnidirectional microphones. However, those skilled in the art may understand that the description in the various embodiments of the present disclosure may be easily extended to the case where the microphone array includes one or more of omnidirectional microphones, unidirectional microphones, bidirectional microphones, or any combination thereof.

[0039] The set of audio signals acquired by the microphone array comprises an audio  signal acquired by each omnidirectional microphone of the array. Due to omnidirectional acquisition characteristics of the omnidirectional microphones, the audio signals may come from arbitrary Directions of Arrival (DOAs) . In some scenarios or depending on the specific needs of a user, it may be desirable to capture audio signals only from one or more directions, while suppressing audio signals from other one or more directions, i.e. to realize a specific audio acquisition mode. The audio acquisition method according to one or more embodiments of the present disclosure may further utilize a beamforming process to process the set of audio signals acquired by the microphone array to realize a plurality of different audio acquisition modes, as will be described in further detail below. Beamforming is a signal processing technique for achieving directional signal transmission or reception, which enables coherent superposition of signals from one direction while mutually canceling signals from other directions, thereby enhancing the directivity of signal reception.

[0040] In step S104, a selection of an audio acquisition mode of a plurality of audio acquisition modes may be received. For example, in examples where the audio acquisition method according to one or more embodiments of the present disclosure is implemented by an audio acquisition device, a button, a knob, a user interface, or other component for implementing a selection function may be included on the audio acquisition device for selecting an audio acquisition mode, and the user may select a particular audio acquisition mode by pressing the button, rotating the knob, triggering the user interface, and the like.

[0041] In step S106, a set of weights corresponding to the selected audio acquisition mode is determined. Each weight of the set of weights may correspond to each microphone of the microphone array, and is computed based on a weight coefficient associated with a particular DOA. Herein, the specific DOA may be, for example, a DOA from which it is desired to enhance audio signals received, or a DOA from which it is desired to suppress audio signals received, or the like. Thus, the set of weights may be utilized to adjust the set of audio signals captured by the microphone array to achieve the selected audio acquisition mode.

[0042] In particular, in step S108, audio signals of the set of audio signals may be processed using the set of weights to generate a processed audio signal. For example, the processed audio signal may be generated by separately multiplying the respective weights of the set of weights with audio signals captured by the respective omnidirectional microphones of the microphone array, and superimposing the weighted audio signals. In particular, since  azimuth angles of the microphones of the microphone array are different, their reception times for an audio signal are also different, resulting in phase delays between the audio signals received by the respective microphones. Thus, when beamforming the set of audio signals acquired by the microphone array, the phase delays between the different microphone signals may first be compensated, after which the compensated microphone signals are subjected to a weighted summation process, for example by multiplying and summing the weights of the computed set of weights with the corresponding microphone signals, respectively, to obtain the processed audio signal. Finally, the processed audio signal is output in step S110, e.g., into a speaker, headphones, other audio delivery mechanisms or storage mechanisms, and the like.

[0043] The method of determining the set of weights corresponding to the selected audio acquisition mode is described in further detail below with reference to FIG. 3. FIG. 3 illustrates a flow diagram of a method 300 for determining a set of weights corresponding to an audio acquisition mode, in accordance with one or more embodiments of the present disclosure. As shown in FIG. 3, a correlation parameter representing the correlation between respective audio signals of the set of audio signals may first be determined in step S302, after which the set of weights corresponding to the audio acquisition mode is determined in step S304 based at least on the correlation parameter. In one or more embodiments of the present disclosure, the correlation parameter may refer to, for example, a covariance matrix, but may also be other parameters capable of representing the correlation between the respective audio signals, which is not particularly limited in the embodiments of the present disclosure. For example, the covariance matrix may be obtained by cross-multiplying the audio signals captured by the microphones of the microphone array, and then the set of weights corresponding to the audio acquisition mode may be determined at least based on the covariance matrix.

[0044] In one or more embodiments of the present disclosure, the selected audio acquisition mode may be a customizable acquisition mode configured to enhance audio signals from a specified DOA in the processed audio signal. For example, the user may specify a desired DOA, θ from which audio signals are expected to be enhanced. For the customizable acquisition mode, the set of weights may be calculated, for example, by the following Equation (1) :

[0045] where wC represents the set of weights for the customizable acquisition mode, and is the transposition of wC; R represents a covariance matrix of the audio signals in the set of audio signals received by the microphone array; R-1 is the inverse matrix of R; v (θ) represents a steering vector of the microphone array towards to the angle, θ, vH (θ) is the transposition of v (θ) , and the steering vector represents a transfer function between the DOA, θ, and each microphone of the microphone array. In this customizable acquisition mode, vH (θ) and v (θ) together may be referred to as a weight coefficient associated with the specified DOA, or v (θ) itself may be referred to as the weight coefficient associated with the specified DOA.

[0046] Taking the 4-element UCA shown in FIG. 2 as an example, for a given DOA, θ, the steering vector of the microphone array may be expressed as follows:

[0047] where r represents the radius of the microphone array; λ represents a wavelength of an audio signal; and φ1 to φ4 represent azimuth angles of various microphones of the microphone array, respectively.

[0048] Furthermore, in Equation (2) above, for ease of description, it is assumed that the audio signal and the microphone array are in the same elevation, but it may be understood by those skilled in the art that Equation (2) may be easily extended to the case where the audio signal and the microphone array are in different elevations. Further, in Equation (2) and the following description of the disclosure, audio signals are assumed to have a single frequency for convenience of description, but those skilled in the art may understand that the description in the various embodiments of the present disclosure may be easily extended to the case of wideband signals having different frequencies.

[0049] By combining Equations (1) and (2) above, the set of weights may be adaptively calculated according to the specified DOA from the user and then used to weight the audio signals captured by the various omnidirectional microphones of the microphone array in the subsequent beamforming process.

[0050] In general, the ability to enhance or suppress signals in different directions may be described using a beam pattern. The beam pattern may be expressed, for example, as follows: BP (θ) = wH (θ) v (θ) vH (θ) w (θ)   (3)

[0051] where w (θ) represents a set of weights for a given DOA, θ; wH (θ) is the transposition of w (θ) ; v (θ) represents the steering vector of the microphone array towards the angle, θ.

[0052] In this embodiment of the present disclosure, by combining the above Equations (1) to (3) , the beam pattern for the customizable acquisition mode may be obtained. For example, for the 4-element UCA as shown in FIG. 2, when the specified DOA is 20°, a beam pattern diagram as shown in FIG. 4 may be obtained. As can be seen in FIG. 4, signals from a direction of 20° are enhanced, while signals from an opposite direction, i.e. 200°, are suppressed. In the customizable audio acquisition mode, the user may choose to enhance audio signals for any DOA, thereby greatly enhancing the user experience.

[0053] In one or more embodiments of the present disclosure, the selected audio acquisition mode may be a DOA estimation mode configured to enhance audio signals from a target sound source in the processed audio signal. In a practical application scenario, the sound source may be moving, resulting in blur or distortion in captured audio signals due to such motion. The audio acquisition method according to one or more embodiments of the present disclosure may track and determine a current DOA of a target sound source and determine a set of weights according to the determined current DOA and the correlation parameter, thereby improving the audio acquisition effect for the moving sound source. Herein, the target sound source may be, for example, a sound source of interest, a sound source having the strongest audio signals, or a specific person or vocalized object designated by the user. In particular, the current DOA of an audio signal from the target sound source may be tracked based on corresponding audio signals captured by various omnidirectional microphones of the microphone array from the target sound source. For example, the current DOA may be determined based on the audio signals captured by the microphones based on sound source localization techniques. For the DOA estimation mode, the set of weights in this mode may be determined, for example, by the following Equation (4) :

[0054] where θ represents an initial DOA and Δθ represents a change in the DOA due to the movement of the target sound source; wM represents the set of weights for the DOA estimation mode;  is the transposition of wM ; R represents a covariance matrix of the audio signals in the set of audio signals received by the microphone array; R-1 is the inverse matrix of R; v (θ+Δθ) represents the steering vector of the microphone array towards to the  angle, θ+Δθ, and vH (θ+Δθ) is the transposition of v (θ+Δθ) . In this DOA estimation mode, vH (θ+Δθ) and v (θ+Δθ) together may be referred to as a weight coefficient associated with the DOA, or v (θ+Δθ) itself may be referred to as the weight coefficient associated with the DOA.

[0055] In this DOA estimation mode, by detecting or tracking the current DOA of the target sound source in real-time and updating the set of weights accordingly, the sound of the moving target can be acquired efficiently. For a 4-element UCA as shown in FIG. 2, by combining the above Equations (2) to (4) , a beam pattern for the DOA estimation mode may be obtained as shown in FIG. 5. In FIG. 5, assuming that the initial DOA of the moving target sound source is 20°, by detecting or tracking the moving target sound source and updating the set of weights in real-time, it can be ensured that audio signals from the target sound source can be consistently enhanced.

[0056] In one or more embodiments of the present disclosure, the selected audio acquisition mode may be a cardioid acquisition mode configured to enhance audio signals from the front of the microphone array and suppress audio signals from the rear of the microphone array in the processed audio signal. Herein, the front, the rear, the right side, and the left side of the microphone array may be determined with respect to a predefined reference point or reference frame and may, for example, correspond to directions of arrival of 0°, 180°, 90°, and 270°, respectively, which will be explained below as an example without limiting the present disclosure in any sense. For example, in the example 4-element UCA shown in FIG. 2, the +y, -y, +x, and -x directions may correspond to the front, the rear, the right, and the left sides of the microphone array, respectively.

[0057] In determining the set of weights corresponding to the cardioid acquisition mode, a first weight coefficient associated with an audio signal in the set of audio signals from the front of the microphone array may be determined as a first value and a second weight coefficient associated with an audio signal in the set of audio signals from the rear of the microphone array may be determined as a second value, where the first value is greater than the second value. Then, the set of weights corresponding to the cardioid acquisition mode may be determined based at least on the first weight coefficient and the second weight coefficient. Specifically, in one or more examples, the set of weights corresponding to the cardioid acquisition mode may be determined by the following Equation  (5) :

[0058] where wF represents the set of weights for the cardioid acquisition mode;  is the transposition of wF; and “+” represents a pseudo-inverse of the matrix.

[0059] In this example, the first weight coefficient associated with the audio signal from the front of the microphone array (i.e. the 0° direction) is determined to be 1 and the second weight coefficient associated with the audio signal from the rear of the microphone array (i.e. the 180° direction) is determined to be 0. In addition, in the above Equation, θ1 and θ2 are directions of arrival for adjusting the beam pattern, which may be 90° and 270°, respectively, for example, and α1 and α2 are corresponding weight coefficients. θ1, θ2, α1, and α2 may be set as necessary to adjust the shape of the beam pattern, which is not particularly limited by the embodiments of the present disclosure. For the 4-element UCA as shown in FIG. 2, FIG. 6 shows an example beam pattern of the cardioid acquisition pattern in accordance with one or more embodiments of the present disclosure. As shown in FIG. 6, audio signals from the front of the microphone array (i.e., the 0° direction) are enhanced while audio signals from the rear of the microphone array (i.e., the 180° direction) are suppressed, such that the overall shape of the beam pattern resembles a cardioid. An application example of the cardioid acquisition mode may be speech recording, where it is ensured that sound from the speaker can be captured clearly while noise from the audience is suppressed.

[0060] In one or more embodiments of the present disclosure, the selected audio acquisition mode may be an omnidirectional acquisition mode configured to output audio signals from all directions of arrival equally in the processed audio signal. In this mode, weights in the set of weights corresponding to the respective omnidirectional microphones of the microphone array may be determined to be an equal value, for example, 1. For the 4-element UCA as shown in FIG. 2, FIG. 7 shows an example beam pattern of the omnidirectional acquisition pattern in accordance with one or more embodiments of the present disclosure. As shown in FIG. 7, audio signals from different directions of arrival are equalized.

[0061] In one or more embodiments of the present disclosure, the selected audio acquisition mode may be a stereo acquisition mode configured to enhance audio signals from the front left and front right of the microphone array in the processed audio signal for provision into a first channel and a second channel, respectively. Generally, in order to achieve a stereo  sound effect, audio signals different from each other need to be output into two channels, respectively, to be provided, for example, to the left ear and the right ear of the user, respectively. Therefore, for the stereo acquisition mode, two sets of weights respectively corresponding to the first channel and the second channel need to be determined. In one or more embodiments, for the first weight set, a first weight coefficient associated with an audio signal in the set of audio signals from the front left of the microphone array may be determined as a first value and a second weight coefficient associated with an audio signal in the set of audio signals from the rear right of the microphone array may be determined as a second value, the first value being greater than the second value, and then the first set of weights may be determined based at least on the first weight coefficient and the second weight coefficient. In one or more embodiments, for the second weight set, a third weight coefficient associated with an audio signal in the set of audio signals from the front right of the microphone array may be determined as a third value and a fourth weight coefficient associated with an audio signal in the set of audio signals from the rear left of the microphone array may be determined as a fourth value, the third value being greater than the fourth value, and then the second set of weights may be determined based at least on the third weight coefficient and the fourth weight coefficient.

[0062] Specifically, in one or more examples, the first set of weights, wLF, and the second set of weights, wRF, for the stereo acquisition mode may be determined by the following Equations (6) and (7) , respectively:

[0063] where is the transposition of the first set of weights, wLF, and is the transposition of the second set of weights, wRF.

[0064] In this example, for the first set of weights, the first weight coefficient associated with an audio signal from the front left of the microphone array (e.g., the -45° direction) is determined to be 1 and the second weight coefficient associated with an audio signal from the rear right of the microphone array (e.g., the 135° direction) is determined to be 0; for the second set of weights, the third weight coefficient associated with an audio signal from the front right of the microphone array (e.g. 45° direction) is determined to be 1 and the fourth weight  coefficient associated with an audio signal from the rear left of the microphone array (e.g. -135°direction) is determined to be 0. In addition, in the above Equations, θ1 and θ2 are directions of arrival for adjusting the beam pattern, and α1and α2 are corresponding weight coefficients. θ1, θ2, α1, and α2 may be set as necessary to adjust the shape of the beam pattern, which is not particularly limited by the embodiments of the present disclosure.

[0065] For the 4-element UCA as shown in FIG. 2, FIG. 8 shows an example beam pattern of the stereo acquisition mode in accordance with one or more embodiments of the present disclosure. As shown in FIG. 8, in the beam pattern represented by the solid line, audio signals from the front left of the microphone array (-45° direction) are enhanced while audio signals from the rear right of the microphone array (135° direction) are suppressed; in the beam pattern represented by the dashed line, audio signals from the front right of the microphone array (45° direction) are enhanced while audio signals from the rear left of the microphone array (-135° direction) are suppressed.

[0066] In one or more embodiments of the present disclosure, the selected audio acquisition mode may be a bidirectional acquisition mode configured to enhance audio signals from the front and rear of the microphone array and suppress audio signals from the left and right of the microphone array in the processed audio signal. In determining the set of weights corresponding to the bidirectional acquisition mode, a first weight coefficient and a second weight coefficient respectively associated with audio signals in the set of audio signals from the front and rear of the microphone array may be determined as a first value, a third weight coefficient and a fourth weight coefficient respectively associated with audio signals in the set of audio signals from the left and right of the microphone array may be determined as a second value, the first value being greater than the second value, and then the set of weights corresponding to the bidirectional acquisition mode may be determined based on the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient. Specifically, in one or more examples, the set of weights corresponding to the omnidirectional acquisition pattern may be determined by the following Equation (8) :

[0067] where wBi represents the set of weights for the omnidirectional acquisition mode;  is the transposition of wBi.

[0068] In this example, the first weight coefficient and the second weight coefficient associated with the audio signals from the front (i.e. the 0° direction) and the rear (i.e. the 180° direction) of the microphone array are determined to be 1, and the third weight coefficient and the fourth weight coefficient associated with the audio signals from the left (i.e. the -90° direction) and the right (i.e. the 90° direction) of the microphone array are determined to be 0. For the 4-element UCA as shown in FIG. 2, FIG. 9 illustrates an example beam pattern of the bidirectional acquisition mode in accordance with one or more embodiments of the present disclosure. As shown in FIG. 9, audio signals from the front (i.e., 0° direction) and rear (i.e., 180° direction) of the microphone array are enhanced, while audio signals from the left (i.e., -90° direction) and right (i.e., 90° direction) of the microphone array are suppressed, such that the overall shape of the beam pattern resembles an “8” type, and thus the bidirectional acquisition mode may also be referred to as an “8” acquisition mode.

[0069] The audio acquisition method according to one or more embodiments of the present disclosure is described above, which enables free switching of a plurality of audio acquisition modes using a microphone array composed of omnidirectional microphones, and may achieve a higher signal-to-noise ratio without complicated structural design. In particular, an audio acquisition mode may be customized according to user requirements to enhance audio signals from a specified DOA, and a DOA estimation mode may be realized to accurately capture the sounds of a moving target.

[0070] An audio acquisition apparatus according to one or more embodiments of the present disclosure will be described below with reference to FIG. 10. FIG. 10 illustrates a schematic structural diagram of an audio acquisition apparatus 1000, in accordance with one or more embodiments of the disclosure. As shown in FIG. 10, the audio acquisition apparatus 1000 may include an acquisition unit 1002, an input unit 1004, a processing unit 1006, and an output unit 1008. In addition to these four units, the audio acquisition apparatus 1000 may further include other related components, but since these components are not relevant to the present disclosure, a detailed description thereof is omitted herein. In addition, since details of part of the functions of the audio acquisition apparatus 1000 are similar to details of the steps of the audio acquisition method 100 as described with reference to FIG. 1, repeated descriptions of some content are omitted herein for brevity. The audio acquisition apparatus 1000 may be a stand-alone apparatus or may be included in other terminals or devices (e.g., a computer, a cell  phone, a smart wearable device, a television set, etc. ) having an audio acquisition function, which is not particularly limited by the embodiments of the present disclosure.

[0071] The acquisition unit 1002 is configured to acquire a set of audio signals using a microphone array comprising a plurality of omnidirectional microphones arranged in a predetermined rule. The microphone array according to one or more embodiments of the present disclosure may include a plurality of omnidirectional microphones, for example, 4 omnidirectional microphones, or a greater or fewer number of microphones, which is not particularly limited by the embodiments of the present disclosure. The respective omnidirectional microphones in the microphone array may be identical, or perfectly calibrated with respect to each other. The plurality of omnidirectional microphones in the microphone array may be arranged, for example, in a predetermined rule such as a cross shape, a plane shape, a square shape, a circular shape, a spiral shape, and the like, which is not particularly limited by the embodiments of the present disclosure.

[0072] The input unit 1004 may be configured to receive a selection of an audio acquisition mode of a plurality of audio acquisition modes. For example, a button, a knob, a user interface, or other component for implementing a selection function may be included on the audio acquisition apparatus for selecting an audio acquisition mode, and the user may select a particular audio acquisition mode by pressing the button, rotating the knob, triggering the user interface, and the like.

[0073] The processing unit 1006 may be configured to determine a set of weights corresponding to the selected audio acquisition mode. Each weight of the set of weights may correspond to each microphone of the microphone array, and is computed based on a weight coefficient associated with a particular DOA. Herein, the specific DOA may be, for example, a DOA from which it is desired to enhance audio signals received, or a DOA from which it is desired to suppress audio signals received, or the like. Thus, the set of weights may be utilized to adjust the set of audio signals captured by the microphone array to achieve the selected audio acquisition mode. In particular, the processing unit 1006 may process audio signals of the set of audio signals using the set of weights to generate a processed audio signal. For example, the processed audio signal may be generated by separately multiplying the respective weights of the set of weights with audio signals captured by the respective omnidirectional microphones of the microphone array, and superimposing the weighted audio signals. Finally, the processed  audio signal is output by the output unit 1008 into a speaker, headphones, other audio delivery mechanisms or storage mechanisms, and the like.

[0074] In one or more embodiments of the present disclosure, the processing unit 1006 may be configured to determine a correlation parameter representing the correlation between respective audio signals of the set of audio signals, and then determine the set of weights corresponding to the audio acquisition mode based at least on the correlation parameter. In one or more embodiments of the embodiments of the present disclosure, the correlation parameter may refer to, for example, a covariance matrix, but may also be other parameters capable of representing the correlation between the respective audio signals, which is not particularly limited in the embodiments of the present disclosure. For example, the covariance matrix may be obtained by cross-multiplying the audio signals captured by the microphones of the microphone array, and then the set of weights corresponding to the audio acquisition mode may be determined at least based on the covariance matrix.

[0075] The audio acquisition apparatus 1000 according to one or more embodiments of the present disclosure may implement a plurality of audio acquisition modes, such as a customizable acquisition mode, a DOA estimation mode, a cardioid acquisition mode, an omnidirectional acquisition mode, a stereo acquisition mode, a bidirectional acquisition mode, and the like, as described in detail above. The user is free to switch between these audio acquisition modes to obtain the desired audio effect. In particular, the audio acquisition apparatus 1000 according to one or more embodiments of the present disclosure may customize an audio acquisition mode to enhance audio signals from a specified DOA according to user demand, and may implement a DOA estimation mode to precisely capture the sounds of a moving target.

[0076] In one or more embodiments of the present disclosure, there is further provided an audio acquisition device comprising one or more processors and one or more memories, where the one or more memories have stored therein computer-readable instructions which, when executed by the one or more processors, cause the one or more processors to execute the audio acquisition method as described above.

[0077] Program portions of the technology may be considered to be “product” or “article” that exists in the form of executable codes and / or related data, which are embodied or implemented by a computer-readable medium. A tangible, permanent storage medium may  include an internal memory, or a storage used by computers, processors, or similar devices or associated modules. For example, various semiconductor memories, tape drivers, disk drivers, or any similar devices capable of providing storage functionality for software.

[0078] All software or parts of it may sometimes communicate over a network, such as the Internet or other communication networks. Such communication can load software from one computer device or processor to another. For example, loading from one server or host computer to a hardware environment of one computer environment, or other computer environment implementing the system, or a system having a similar function associated with providing information needed for the communication method. Therefore, another medium capable of transmitting software elements can also be used as a physical connection between local devices, such as light waves, electric waves, electromagnetic waves, etc., to be propagated through cables, optical cables, or air. A physical medium used for carrying the waves such as cables, wireless connections, or fiber optic cables may also be considered as a medium for carrying the software. In usage herein, unless a tangible “storage” medium is defined, other terms referring to a computer or machine “readable medium” mean a medium that participates in execution of any instruction by the processor.

[0079] The present application uses specific words to describe embodiments of the present disclosure. Reference to “an embodiment, ” “one or more embodiments, ” and / or “some embodiments” means a feature, structure, or characteristic in connection with at least one embodiment of the present disclosure. Therefore, it should be emphasized and noted that two or more references to “an embodiment, ” “one embodiment, ” or “an alternative embodiment” in various places throughout this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics may be combined as suitable in one or more embodiments of the application.

[0080] Moreover, one skilled in the art will appreciate that aspects of the present disclosure may be illustrated and described in terms of a number of patentable categories or instances, including any new and useful process, machine, manufacture, or combination of matter, or any new and useful improvement thereof. Accordingly, aspects of the present disclosure may be performed entirely by hardware, entirely by software (including firmware, resident software, micro-code, etc. ) , or by a combination of hardware and software. The above  hardware or software may each be referred to as a “data block, ” “module, ” “engine, ” “unit, ” “component, ” or “system. ” Furthermore, aspects of the present disclosure may be embodied as a computer product embodied in one or more computer-readable media including computer-readable program code.

[0081] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or extremely formal sense unless expressly so defined herein.

[0082] While various embodiments of the disclosure have been described, it will be apparent to those of ordinary skill in the art that many more embodiments and implementations are possible that are within the scope of the disclosure. Accordingly, the disclosure is not to be restricted except in light of the attached claims and their equivalents.

Claims

1.An audio acquisition method, comprising:acquiring a set of audio signals using a microphone array comprising a plurality of omnidirectional microphones arranged in a predetermined rule, the set of audio signals comprising an audio signal acquired by each omnidirectional microphone of the microphone array;receiving a selection of an audio acquisition mode of a plurality of audio acquisition modes;determining a set of weights corresponding to the selected audio acquisition mode, wherein each weight of the set of weights corresponds to each microphone of the microphone array and is computed based on a weight coefficient associated with a particular direction of arrival (DOA) ;processing audio signals of the set of audio signals using the set of weights to generate a processed audio signal; andoutputting the processed audio signal.2.The method of claim 1, wherein determining the set of weights corresponding to the selected audio acquisition mode comprises:determining a correlation parameter representing correlation between respective audio signals of the set of audio signals; anddetermining the set of weights corresponding to the selected audio acquisition mode based at least on the correlation parameter.3.The method of claim 2, wherein the selected audio acquisition mode comprises a customizable acquisition mode configured to enhance an audio signal from a specified DOA in the processed audio signal, and wherein determining the set of weights corresponding to the selected audio acquisition mode based at least on the correlation parameter comprises:receiving the specified DOA; anddetermining the set of weights corresponding to the customizable acquisition mode based on the specified DOA and the correlation parameter.4.The method of claim 2, wherein the selected audio acquisition mode comprises a DOA estimation mode configured to enhance an audio signal from a target sound source in the processed audio signal, and wherein determining the set of weights corresponding to the selected audio acquisition mode based at least on the correlation parameter comprises:detecting or tracking a current DOA of the audio signal from the target sound source based on audio signals acquired by the plurality of omnidirectional microphones of the microphone array from the target sound source; anddetermining the set of weights corresponding to the DOA estimation mode based on the current DOA and the correlation parameter.5.The method of claim 1, wherein the selected audio acquisition mode comprises a cardioid acquisition mode configured to enhance an audio signal from a front of the microphone array and suppress an audio signal from a rear of the microphone array in the processed audio signal, and wherein determining the set of weights corresponding to the selected audio acquisition mode comprises:determining a first weight coefficient associated with an audio signal in the set of audio signals from the front of the microphone array as a first value and a second weight coefficient associated with an audio signal in the set of audio signals from the rear of the microphone array as a second value, wherein the first value is greater than the second value; anddetermining the set of weights corresponding to the cardioid acquisition mode based at least on the first weight coefficient and the second weight coefficient.6.The method of claim 1, wherein the selected audio acquisition mode comprises an omnidirectional acquisition mode configured to output audio signals from all directions of arrival equally in the processed audio signal, and wherein determining the set of weights corresponding to the selected audio acquisition mode comprises:determining, in the set of weights, weights corresponding to the respective omnidirectional microphones of the microphone array as an equal value.7.The method of claim 1, wherein the selected audio acquisition mode comprises a stereo acquisition mode configured to enhance audio signals from front left and front right of the microphone array in the processed audio signal for provision into a first channel and a second channel, respectively, and wherein determining the set of weights corresponding to the selected audio acquisition mode comprises:determining a first set of weights for the first channel and a second set of weights for the second channel, respectively.8.The method of claim 7, wherein determining the first set of weights for the first channel and the second set of weights for the second channel, respectively, comprises:determining a first weight coefficient associated with an audio signal in the set of audio signals from the front left of the microphone array as a first value and a second weight coefficient associated with an audio signal in the set of audio signals from a rear right of the microphone array as a second value, the first value being greater than the second value, and determining the first set of weights based at least on the first weight coefficient and the second weight coefficient; anddetermining a third weight coefficient associated with an audio signal in the set of audio signals from a front right of the microphone array as a third value and a fourth weight coefficient associated with an audio signal in the set of audio signals from a rear left of the microphone array as a fourth value, the third value being greater than the fourth value, and determining the second set of weights based at least on the third weight coefficient and the fourth weight coefficient.9.The method of claim 1, wherein the selected audio acquisition mode comprises a bidirectional acquisition mode configured to enhance audio signals from front and rear of the microphone array and suppress audio signals from left and right of the microphone array in the processed audio signal, and wherein determining the set of weights corresponding to the selected audio acquisition mode comprises:determining a first weight coefficient and a second weight coefficient respectively associated with audio signals in the set of audio signals from the front and rear of the microphone array as a first value, and determining a third weight coefficient and a fourth weight coefficient respectively associated with audio signals in the set of audio signals from the left and right of the microphone array as a second value, the first value being greater than the second value; anddetermining the set of weights corresponding to the bidirectional acquisition mode based on the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient.10.The method of claim 1, wherein processing audio signals of the set of audio signals using the set of weights to generate the processed audio signal comprises:performing phase delay compensation on the audio signals of the set of audio signals; andperforming a weighted summation on the compensated audio signals using the set of weights to generate the processed audio signal.11.The method of claim 1, wherein the predetermined rule comprises any one or any combination of a cross shape, a plane shape, a square shape, a circular shape, and a spiral shape.12.An audio acquisition apparatus, comprising:an acquisition unit configured to acquire a set of audio signals using a microphone array comprising a plurality of omnidirectional microphones arranged in a predetermined rule, the set of audio signals comprising an audio signal acquired by each omnidirectional microphone of the microphone array;an input unit configured to receive a selection of an audio acquisition mode of a plurality of audio acquisition modes;a processing unit configured to determine a set of weights corresponding to the selected audio acquisition mode, wherein each weight of the set of weights corresponds to each microphone of the microphone array and is computed based on a weight coefficient associated with a particular direction of arrival (DOA) , and process each audio signal of the set of audio signals using the set of weights to generate a processed audio signal; andan output unit configured to output the processed audio signal.13.An audio acquisition device, comprising:one or more processors; andone or more memories, wherein the one or more memories have stored therein computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to perform the method of any of claims 1-11.

Citation Information

Patent Citations

  • Planar microphone sensor array and sound source positioning method thereof

    CN115774240A