Audio generation method and apparatus, and storage medium

By using the time difference between microphones on the terminal to construct the target intensity difference processing audio signal, the problem of stereo generation on small devices is solved, and efficient and low-cost stereo acquisition is achieved.

WO2025175435A1PCT designated stage Publication Date: 2025-08-28BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/077627
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently create audio audio on small devices, and it usually requires a microphone with complex structure or high computing power, resulting in high cost and low efficiency.

Method used

By collecting audio signals using at least two non-directional microphones of the terminal, a target intensity difference is constructed based on the time difference between the microphones, and a target intensity difference is processed to generate bulk audio.

Benefits of technology

It realizes efficient production of stereo audio on small devices, reduces dependence on complex structure microphones and high computing power, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024077627_28082025_PF_FP_ABST
    Figure CN2024077627_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio generation method and apparatus, and a storage medium. The audio generation method comprises: using at least two microphones of a terminal to acquire first audio signals; constructing a required target level difference on the basis of a time difference between the first audio signals acquired by the at least two microphones, wherein the time difference is negatively correlated with the target level difference; using the target level difference to process the first audio signals to obtain second audio signals, so that a level difference between the second audio signals corresponding to the at least two microphones reaches the target level difference; and generating a stereo audio on the basis of the second audio signals. According to the present disclosure, the obtaining of the stereo audio by means of the described solution does not depend on microphones having complex structures, uses a simple algorithm, and does not depend on high computing power. The present disclosure can be applied to small devices to achieve the acquisition of stereo audios, thereby improving efficiency and reducing cost.
Need to check novelty before this filing date? Find Prior Art

Description

Audio generation method, device and storage medium Technical Field

[0001] The present disclosure relates to the field of audio technology, and in particular to an audio generation method, device, and storage medium. Background Art

[0002] With the development and advancement of science and technology, the technology of capturing spatial audio is being widely researched. Spatial audio, also known as audio in stereo format, provides users with an immersive audio experience.

[0003] Summary of the Invention

[0004] However, the current generation of stereo format audio relies on complex microphone structures or high computing power, which is difficult to implement.

[0005] The embodiments of the present disclosure provide an audio generation method, an audio generation device, and a storage medium.

[0006] According to a first aspect of an embodiment of the present disclosure, an audio generation method is proposed, comprising: collecting a first audio signal using at least two microphones of a terminal; constructing a required target intensity difference based on a time difference between the first audio signals collected by the at least two microphones; wherein the time difference and the target intensity difference are negatively correlated; processing the first audio signal using the target intensity difference to obtain a second audio signal, such that the intensity difference between the second audio signals corresponding to the at least two microphones satisfies the target intensity difference; and generating stereo audio based on the second audio signal.

[0007] According to a second aspect of an embodiment of the present disclosure, an audio generation device is proposed, comprising: an acquisition module for acquiring a first audio signal using at least two microphones of a terminal; a processing module for constructing a required target intensity difference based on a time difference between the first audio signals acquired by the at least two microphones; wherein the time difference and the target intensity difference are negatively correlated; the first audio signal is processed using the target intensity difference to obtain a second audio signal, so that the intensity difference between the second audio signals corresponding to the at least two microphones meets the target intensity difference; and stereo audio is generated based on the second audio signal.

[0008] According to a third aspect of an embodiment of the present disclosure, an electronic device is proposed, comprising: a memory for storing instructions; and a processor for calling the instructions stored in the memory to execute the audio generation method as described in the first aspect and any one of the first aspects.

[0009] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is proposed, wherein the storage medium stores instructions. When the instructions are executed by a processor, the audio generation method as described in the first aspect and any one of the first aspects is executed.

[0010] The present disclosure utilizes the time difference between audio signals collected by different microphones on a terminal to construct a target intensity difference, which is then used to process the collected audio signals to obtain an audio signal that meets the target intensity difference. Stereo audio is then generated based on the audio signal that meets the target intensity difference. The present disclosure utilizes the above-mentioned scheme to obtain a stereo audio format that does not rely on microphones with complex structures, and the algorithm is simple and does not rely on high computing power. This approach can be applied to small devices to achieve stereo audio acquisition, improving efficiency and reducing costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following drawings required for describing the embodiments are introduced. The following drawings are merely some embodiments of the present disclosure and do not impose specific limitations on the protection scope of the present disclosure.

[0012] FIG1 is a flowchart showing an audio generation method according to an embodiment of the present disclosure.

[0013] Fig. 2 is a flow chart of a method for processing an audio signal according to an exemplary embodiment.

[0014] Fig. 3 is a schematic diagram showing a first-order differential beamforming according to an exemplary embodiment.

[0015] Fig. 4 is a block diagram of an audio generating device 100 according to an exemplary embodiment.

[0016] Fig. 5 is a block diagram showing an audio generating device 200 according to an exemplary embodiment. DETAILED DESCRIPTION

[0017] The embodiments of the present disclosure provide an audio generation method, an audio generation device, and a storage medium.

[0018] In a first aspect, an embodiment of the present disclosure proposes an audio generation method, comprising: collecting a first audio signal using at least two microphones of a terminal; constructing a required target intensity difference based on a time difference between the first audio signals collected by the at least two microphones; wherein the time difference and the target intensity difference are negatively correlated; processing the first audio signal using the target intensity difference to obtain a second audio signal, so that the intensity difference between the second audio signals corresponding to the at least two microphones satisfies the target intensity difference; and generating stereo audio based on the second audio signal.

[0019] In the above embodiment, the time difference between the audio signals collected by different microphones in the terminal is used to construct a target intensity difference. This target intensity difference is then used to process the collected audio signals to obtain an audio signal that meets the target intensity difference. Stereo audio is then generated based on the audio signal that meets the target intensity difference. The above scheme achieves a stereo audio format that does not rely on complex microphone structures, and the algorithm is simple and does not require high computing power. It can be applied to small devices, enabling stereo audio capture, improving efficiency, and reducing costs.

[0020] In some optional embodiments of the first aspect, the distance between the microphones determines the time difference between the first audio signals collected by the microphones, and constructing the required target intensity difference based on the time difference between the microphones includes: determining a matrix based on the distance between the at least two microphones, the matrix being used to characterize the target intensity difference, and a preconfigured matrix function being satisfied between the distance and the matrix; and processing the first audio signal using the target intensity difference includes: processing the first audio signal using the matrix; wherein the matrix is ​​used to increase or decrease the intensity difference of the first audio signal.

[0021] In the above embodiment, since the distance between the microphones determines the time difference between the microphones, a matrix can be determined based on the distance, and the matrix can be used to process the first audio signal to obtain stereo audio through simple algorithm processing, thereby improving efficiency.

[0022] In some optional embodiments of the first aspect, the smaller the distance between the at least two microphones, the larger the value of the matrix; and the larger the distance between the at least two microphones, the smaller the value of the matrix.

[0023] In the above embodiment, the smaller the distance between the microphones, the larger the value of the matrix, and the smaller the distance between the microphones, the larger the value of the matrix, that is, there is a negative correlation between the microphones and the matrix, so that the matrix can be determined according to the distance to achieve stereo format acquisition.

[0024] In some optional embodiments of the first aspect, the distance between the at least two microphones is determined in the following manner: according to the distance between the microphones pre-stored in the terminal.

[0025] In the above embodiment, the distance between the microphones may be pre-stored in the terminal, and the terminal may directly read the pre-stored distance value in the terminal when generating the audio signal in a stereo format, so as to improve efficiency.

[0026] In some optional embodiments of the first aspect, the distance between the at least two microphones is determined in the following manner: the distance between the microphones is acquired in real time according to a folding state of the terminal.

[0027] In the above embodiment, the distance between the microphones can be determined according to the folding state of the terminal, so that for a terminal with a folding function, the distance between the microphones can be flexibly determined to ensure that stereo audio can be accurately captured.

[0028] In some optional embodiments of the first aspect, calculating the distance between the at least two microphones according to the folding state of the terminal includes: calculating the distance between the at least two microphones according to a folding angle of the terminal.

[0029] In the above embodiment, the distance between the microphones can be calculated according to the folding angle of the terminal, so that the distance between the microphones can be accurately calculated, and the distance between the microphones can be flexibly obtained for different angles to ensure that the audio in stereo format can be accurately captured.

[0030] In some optional embodiments of the first aspect, the terminal includes multiple sub-terminals, and the distance between the at least two microphones is determined in the following manner: the distance between the microphones is determined according to the distances between different sub-terminals.

[0031] In the above embodiment, if the terminal includes multiple sub-terminals, the distance between the microphones can be determined according to the distance between the sub-terminals, so as to quickly obtain audio in a stereo format.

[0032] In a second aspect, an audio generation device is provided, comprising: an acquisition module for acquiring a first audio signal using at least two microphones of a terminal; a processing module for constructing a required target intensity difference based on a time difference between the first audio signals acquired by the at least two microphones; wherein the time difference and the target intensity difference are negatively correlated; using the target intensity difference, the first audio signal is processed to obtain a second audio signal, so that the intensity difference between the second audio signals corresponding to the at least two microphones meets the target intensity difference; and based on the second audio signal, stereo audio is generated.

[0033] In some optional embodiments of the second aspect, the distance between the microphones determines the time difference between the first audio signals collected by the microphones, and the processing module constructs the required target intensity difference based on the time difference between the first audio signals collected by the at least two microphones in the following manner: a matrix is ​​determined based on the distance between the microphones, the matrix is ​​used to characterize the target intensity difference, and a preconfigured matrix function is satisfied between the distance and the matrix; the processing module uses the target intensity difference to process the first audio signal in the following manner: the first audio signal is processed using the matrix; wherein the matrix is ​​used to increase or decrease the intensity difference of the first audio signal.

[0034] In some optional embodiments of the second aspect, the smaller the distance between the at least two microphones, the larger the value of the matrix; and the larger the distance between the at least two microphones, the smaller the value of the matrix.

[0035] In some optional embodiments of the second aspect, the processing module determines the distance between the microphones in the following manner: determining according to the distance between the microphones pre-stored in the terminal.

[0036] In some optional embodiments of the second aspect, the processing module determines the distance between the microphones in the following manner: acquiring the distance between the at least two microphones in real time according to a folding state of the terminal.

[0037] In some optional embodiments of the second aspect, the processing module obtains the distance between the at least two microphones in real time according to the folding state of the terminal in the following manner: calculating the distance between the at least two microphones according to the folding angle of the terminal.

[0038] In some optional embodiments of the second aspect, the terminal includes multiple sub-terminals, and the processing module determines the distance between the microphones in the following manner: determining the distance between the at least two microphones based on the distances between different sub-terminals.

[0039] In a third aspect, an electronic device is proposed, comprising: a memory for storing instructions; and a processor for calling the instructions stored in the memory to execute the audio generation method as described in the first aspect and any one of the first aspects.

[0040] In a fourth aspect, a storage medium is proposed, wherein the storage medium stores instructions. When the instructions are executed by a processor, the audio generation method as described in the first aspect and any one of the first aspects is executed.

[0041] In a fifth aspect, an embodiment of the present disclosure proposes a program product. When the program product is executed by a communication device, the communication device executes the method described in the optional implementation manner of the first aspect or the second aspect.

[0042] In a sixth aspect, an embodiment of the present disclosure proposes a computer program, which, when executed on a computer, enables the computer to execute the method described in the optional implementation of the first aspect or the second aspect.

[0043] In a seventh aspect, an embodiment of the present disclosure provides a chip or a chip system, which includes a processing circuit configured to execute the method described in the optional implementation of the first or second aspect.

[0044] It is understood that the audio generating device, electronic device, storage medium, program product, computer program, chip, or chip system involved in each embodiment of the present disclosure is used to perform the method proposed in the embodiment of the present disclosure. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding method and will not be repeated here.

[0045] The present disclosure provides an audio generation method, device, and storage medium. In some embodiments, the terms audio generation method and data processing method are interchangeable, and the terms audio generation device and data processing device are interchangeable.

[0046] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.

[0047] In each embodiment of the present disclosure, unless otherwise specified or provided for, the terms and / or descriptions between the embodiments are consistent and may be referenced by each other. The technical environments in different embodiments may be combined to form new embodiments based on their inherent logical relationships.

[0048] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.

[0049] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.

[0050] In the embodiments of the present disclosure, “plurality” refers to two or more.

[0051] In some embodiments, the terms "at least one," "one or more," "a plurality of," "multiple," etc. may be used interchangeably.

[0052] In some embodiments, descriptions such as "at least one of A and B," "A and / or B," "A in one case, B in another case," or "in response to one case A, in response to another case B" may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); and in some embodiments, A and B (both A and B are executed). The above is also applicable when there are more branches such as A, B, and C.

[0053] In some embodiments, "A or B" and other descriptions may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The above is also applicable when there are more branches such as A, B, C, etc.

[0054] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for example, if the description object is "information", then the "first information" and "the performance of each AI model" can be the same information or different information, and their contents can be the same or different.

[0055] In some embodiments, “including A,” “comprising A,” “used to indicate A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.

[0056] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "at the time of...", "when...", "if...", "if...", etc. can be used interchangeably.

[0057] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not less than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.

[0058] In some embodiments, devices and equipment can be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they can also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", etc.

[0059] In some embodiments, "network" can be interpreted as devices included in the network, such as access network equipment, core network equipment, etc.

[0060] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)", "base station (BS)", "radio base station", "fixed station", and in some embodiments may also be understood as "node", "access point", "transmission point (TP)", "reception point (RP)", "transmission and / or reception point (TRP)" "panel", "antenna panel", "antenna array", "cell", "macro cell", "small cell", "femto cell", "pico cell", "sector", "cell group", "serving cell", "carrier", "component carrier", "bandwidth part (BWP)", etc.

[0061] In some embodiments, "terminal" or "terminal device" may be referred to as "user equipment (UE)", "user terminal" "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc.

[0062] In some embodiments, obtaining data, information, etc. may comply with the laws and regulations of the country where the data is obtained.

[0063] In some embodiments, data, information, etc. may be obtained with the user's consent.

[0064] In addition, each element, each row, or each column in the table of the embodiment of the present disclosure can be implemented as an independent embodiment, and the combination of any elements, any rows, and any columns can also be implemented as an independent embodiment.

[0065] Currently, stereo audio capture is often achieved through two methods. First, spatial audio is captured using a directional microphone pair. A directional microphone pair is a microphone composed of two diaphragms. Its structure is more complex than an omnidirectional microphone, and its low-frequency response is poor. Its disadvantages are its large size, making it difficult to install on small devices such as mobile phones and AR glasses. Its components are fragile and easily affected by the environment, and its cost is high. Second, stereo audio is generated through microphone array processing. However, this processing requires high computing power, making it difficult to implement on small devices due to computing power limitations. In other words, due to its large size and high computing power requirements, terminals cannot capture stereo audio, or the efficiency of stereo audio capture is low, resulting in high costs.

[0066] Therefore, the present disclosure provides an audio generation method that uses at least two microphones of a terminal to collect original audio signals and constructs a desired target intensity difference based on the time difference between the original audio signals collected by the at least two microphones. The original audio signals are processed using the target intensity difference so that the intensity difference between the processed audio signals meets the target intensity difference, thereby generating stereo audio based on the processed audio signals. Through the above-mentioned scheme for generating stereo audio, the present disclosure can realize stereo audio collection on small devices such as terminals.

[0067] In some embodiments, the terminal includes, for example, a mobile phone, a wearable device, an Internet of Things device, a car with communication function, a smart car, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical surgery, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, and at least one of a wireless terminal device in a smart home, but is not limited thereto.

[0068] Figure 1 is a flow chart of an audio generation method according to an embodiment of the present disclosure. As shown in Figure 1 , the audio generation method includes the following steps.

[0069] Step S11: Collect a first audio signal using at least two microphones of the terminal.

[0070] In some embodiments, the terminal may be configured with at least two microphones. The microphones are non-directional, simple in structure, and small in size, and can be configured on the terminal. The terminal may use the at least two microphones to collect a first audio signal, which can be understood as a raw audio signal, i.e., an unprocessed audio signal.

[0071] It is understandable that the terminal has at least two microphones for collecting the first audio signal, but is not limited to this. It can be set according to actual conditions. For example, while ensuring efficiency and cost, more microphones can be flexibly used to collect the first audio signal so that the subsequent generated stereo audio effect is better.

[0072] Step S12: constructing a required target intensity difference according to the time difference between the first audio signals collected by at least two microphones.

[0073] In some embodiments, due to the different positions of different microphones, their distances from the sound source also vary. Therefore, the first audio signals captured by different microphones have an interchannel time difference (ICTD) and an interchannel level difference (ICLD). For example, a microphone closer to the sound source captures the first audio signal faster, while a microphone farther from the sound source captures the first audio channel slower, resulting in a time difference between the first audio signals captured by the two microphones. For another example, a microphone closer to the sound source captures a higher intensity first audio signal, while a microphone farther from the sound source captures a lower intensity first audio signal. For a directional microphone, there may be a certain angle between the direction in which it captures the first audio signal and the orientation of the sound source relative to the microphone. The larger the angle, the lower the intensity of the captured first audio signal. Conversely, the smaller the angle, the higher the intensity of the captured first audio signal. To achieve an effect similar to that of a directional microphone, the intensity difference of the first audio signal can be adjusted so that the adjusted first audio signal has a similar effect to the original audio signal captured by the directional microphone, thereby achieving a stereo effect. The intensity of the first audio signal is also affected by the time difference. Therefore, in this embodiment, a target intensity difference is established based on the time difference between the first audio signals captured by different microphones. The target intensity difference represents the intensity difference between the second audio signals obtained after processing the first audio signals. The time difference and the target intensity difference are negatively correlated. That is, the greater the time difference, the greater the target intensity difference, i.e., the greater the intensity difference between the second audio signals obtained after processing. The smaller the time difference, the smaller the target intensity difference, i.e., the smaller the intensity difference between the second audio signals obtained after processing.

[0074] Step S13: Process the first audio signal using the target intensity difference to obtain a second audio signal, so that the intensity difference between the second audio signals corresponding to at least two microphones satisfies the target intensity difference.

[0075] In some embodiments, the first audio signal may be processed to obtain a second audio signal such that the intensity difference of the second audio signal satisfies a target intensity difference. The processed second audio signal may be understood as an audio signal collected by a virtual microphone, which has directionality.

[0076] Step S14: Generate stereo audio based on the second audio signal.

[0077] In some embodiments, the microphone may be configured to output the second audio signal as stereo audio, which may also be referred to as spatial audio.

[0078] The present disclosure utilizes at least two microphones of a terminal to collect a first audio signal, and determines a target intensity difference based on the time difference of the first audio signals collected by the at least two microphones. The first audio signal is processed using the target intensity difference to obtain a second audio signal. The intensity difference between the second audio signals is ensured to meet the target intensity difference, thereby generating stereo audio based on the second audio signals. The present disclosure, through the above-described scheme for generating stereo audio, enables stereo audio to be collected on small devices such as terminals.

[0079] FIG2 is a flow chart of an audio signal processing method according to an exemplary embodiment. As shown in FIG2 , the present disclosure provides an audio signal processing method, which includes the following steps.

[0080] Step S21 : determining a matrix based on the distance between at least two microphones, where the matrix is ​​used to represent the target intensity difference, and a pre-configured matrix function is satisfied between the distance and the matrix.

[0081] Step S22: Process the first audio signal using a matrix.

[0082] In some embodiments, the distance between different microphones can determine the time difference between the first audio signals collected by different microphones. Therefore, a matrix can be determined based on the distance between the microphones. The matrix can be used to represent a target intensity difference. That is, the matrix can be used to enhance or reduce the intensity difference of the first audio signal so that the intensity difference between the obtained second audio signals reaches the target intensity difference. In other words, the distance between the microphones can be used to determine the matrix, and the matrix can be used to process the first audio signal to obtain the second audio signal.

[0083] In some embodiments, the terminal may pre-configure a matrix function. A matrix function is a function between a matrix and a distance, that is, the distance can be understood as the independent variable of the matrix function, and the matrix can be understood as the dependent variable of the matrix function. The matrix function refers to Formula 1. H = [1, -e jwr0a1,1 ] Formula 1

[0084] In formula 1, H represents the matrix, e jwr0a1,1 represents a constant e as the base and jwr0a1,1 as the exponent. Wherein, the magnitude of e is approximately equal to 2.72. j represents a complex number and w represents the frequency of the audio signal. r0 represents the distance between the microphones. a1,1 represents the parameter that controls the beam performance. That is, the matrix can be determined based on the distance r0 between the microphones and the frequency w of the audio signal. For example, when w = A and r0 = B, the matrix is ​​assumed to be H1, which is equal to [1, -e jABa1,1 ]. When determining the matrix, the frequency w can be obtained according to the actual situation and substituted into Formula 1, or it can be set to a fixed value, which is not limited in this disclosure. When determining the matrix formula, the actual distance between the microphones is substituted into Formula 1 to obtain the matrix.

[0085] In some embodiments, the first audio signal can be processed using a determined matrix. This disclosure will be described using the example of two left and right microphones collecting audio signals, but it is understood that the embodiments of this disclosure are not limited to two left and right microphones. That is, the first audio signal includes a first audio signal of a left channel and a first audio signal of a right channel. The following formulas 2 and 3 can be used to process the first audio signal using a matrix. CH L =1*mic L +e jwr0a1,1 *mic R Formula 2 CH R =1*mic R +e jwr0a1,1 *mic L Formula 3

[0086] In Formula 2 and Formula 3, CH L Represents the second audio signal of the left channel, mic L Represents the frequency domain expression of the first audio signal of the left channel, mic R represents the frequency domain expression of the first audio signal of the right channel. 1 and e jwr0a1,1 Represents two elements in a determined matrix. That is, the first audio signal of the left channel can be multiplied by 1, and the first audio signal of the right channel can be multiplied by e jwr0a1,1 , and add them together to get the second audio signal of the left channel. The first audio signal of the left channel can be multiplied by e jwr0a1,1, multiply the first audio signal of the right channel by 1, and add them to obtain the second audio signal of the right channel.

[0087] In some embodiments, FIG3 is a schematic diagram of a first-order differential beam according to an exemplary embodiment. As shown in FIG3 , two microphones can form a first-order differential beam, and two microphones can be used to form two mutually symmetrical beams, each pointing to two end-fire directions opposite to the microphone array. The left channel is a beam with a zero point set at 0° and an end-fire direction of 180°, thereby suppressing the sound on the right. The right channel is a beam with a zero point set at 180° and an end-fire direction of 0°, thereby suppressing the sound on the left. As a result, the left and right channels construct different intensity differences (interchannel level difference, ICLD) for sounds incident at different angles. The matrix H can be called a filter for the differential beam. For the left channel, the filter e of the differential beam jwr0a1,1 You can add a delay to the micR signal. For the right channel, the filter of the differential beam is e jwr0a1,1 By adding a delay to the micL signal and adding the mic spacing d, the left and right sounds arrive at different time intervals between the left and right channels, with the left sound arriving in the left channel before the right, and the right sound arriving in the right channel before the left. By adjusting the beam filter, you can control the ICLD between the left and right channels. ICTD is affected by the mic placement. Adjusting the ICLD and ICTD between the left and right channels based on the desired sound field creates a suitable stereo signal.

[0088] In the present disclosure, since the distance between microphones determines the time difference between the microphones, a matrix can be determined based on the distance, and the matrix can be used to process the first audio signal to obtain stereo audio through simple algorithm processing, thereby improving efficiency.

[0089] The audio signal processing method provided by the present disclosure includes that the smaller the distance between microphones, the larger the value of the matrix; and the larger the distance between microphones, the smaller the value of the matrix.

[0090] In some embodiments, a smaller distance between microphones indicates a smaller time difference between the microphones, requiring a larger target intensity difference, i.e., a larger matrix value is required. A larger distance between microphones indicates a larger time difference between the microphones, requiring a smaller target intensity difference, i.e., a smaller matrix value is required.

[0091] In some embodiments, stereo audio is captured by adjusting the distance between the microphones and applying a processing algorithm to achieve the desired directivity. When the microphones are spaced farther apart, i.e., the ICTD is larger, the reliance on directivity is reduced, i.e., the ICLD is smaller. Conversely, when the microphones are spaced closer, i.e., the ICTD is smaller, more directivity is required, i.e., the ICLD is larger. By adjusting the ICLD and ICTD values, a desired sound field width can be created.

[0092] In the present disclosure, the smaller the distance between the microphones, the larger the value of the matrix, and the smaller the distance between the microphones, the larger the value of the matrix, that is, there is a negative correlation between the microphones and the matrix, so that the matrix can be determined according to the distance to achieve stereo format acquisition.

[0093] The present disclosure provides a method for determining the distance between microphones, including: determining the distance between microphones according to pre-stored data in a terminal.

[0094] In some embodiments, the distance between the microphones can be pre-stored in the terminal. During the audio processing process, the pre-stored distance between the microphones can be read from the terminal to enable fast and efficient audio generation.

[0095] In the present disclosure, the distance between the microphones may be pre-stored in the terminal, and the terminal may directly read the pre-stored distance value in the terminal when generating an audio signal in a stereo format, so as to improve efficiency.

[0096] The present disclosure provides a method for determining the distance between microphones, including: acquiring the distance between microphones in real time according to a folding state of a terminal.

[0097] In some embodiments, the distance between the microphones can be obtained in real time based on the folding state of the terminal. For example, for a terminal with a folding function, the distance between the microphones can be obtained in real time based on whether the terminal is folded and the degree of folding. For example, the terminal can pre-store the distance between the microphones at different folding degrees. For another example, the terminal can pre-store the distance between the microphones at different folding methods, which is not limited in this disclosure.

[0098] In the present disclosure, the distance between the microphones can be determined according to the folding state of the terminal, so that for a terminal with a folding function, the distance between the microphones can be flexibly determined to ensure that the audio in stereo format can be accurately captured.

[0099] In some embodiments, the present disclosure provides a method for determining a distance between microphones, comprising: calculating the distance between the microphones according to a folding angle of a terminal.

[0100] In some embodiments, the distance between the microphones can be calculated based on the angle at which the terminal is folded. For example, for different terminals, there is a corresponding relationship between the angle at which the terminal is folded and the distance between the microphones. Therefore, the distance between the microphones can be calculated based on the angle at which the terminal is folded.

[0101] In the present disclosure, the distance between the microphones can be calculated according to the folding angle of the terminal, so that the distance between the microphones can be accurately calculated, and the distance between the microphones can be flexibly obtained for different angles, ensuring that the audio in stereo format can be accurately captured.

[0102] In some embodiments, the terminal may be a combined terminal including multiple terminals, or a distributed terminal. For these terminals, in order to more quickly obtain the distance between microphones, the present disclosure provides a method for determining the distance between microphones, including: determining the distance between microphones based on the distances between different sub-terminals.

[0103] In some embodiments, a terminal may include multiple sub-terminals, and the terminal's microphones may be distributed across multiple sub-terminals. A sub-terminal can be understood as an independent entity that makes up the terminal. For example, when the terminal is a combination terminal, each terminal in the combination terminal can be referred to as a sub-terminal of the combination terminal. For example, when a mobile phone and a tablet are connected to the same network or establish another wireless connection, the mobile phone and tablet can be referred to as a combination terminal, the mobile phone can be referred to as a sub-terminal of the combination terminal, and the tablet can be referred to as another sub-terminal of the combination terminal. For another example, when the terminal is a distributed terminal, each node in the distributed terminal can be referred to as a sub-terminal of the distributed terminal. A distributed terminal is an intelligent device based on distributed technology, consisting of multiple nodes, each capable of data processing and communication. For example, a distributed terminal can be a pair of Bluetooth headsets, each of which can be referred to as a sub-terminal of the pair. Therefore, the distance between at least two microphones can be determined based on the distance between the different sub-terminals. For example, if microphone A is located on sub-terminal X and microphone B is located on sub-terminal Y, the distance between microphones A and B can be determined based on the distance between sub-terminal X and sub-terminal Y.

[0104] In the present disclosure, if the terminal includes multiple sub-terminals, the distance between the microphones can be determined according to the distance between the sub-terminals, so as to quickly obtain audio in a stereo format.

[0105] The present disclosure also provides an audio generation method as follows:

[0106] In some embodiments, two microphones are spaced apart by a distance d to form a microphone array, and two virtual microphones with different orientations are constructed at different positions to form the left and right channels required for stereo sound.

[0107] In some embodiments, the position of virtual mic1 is similar to that of mic1, and the position of virtual mic2 is similar to that of mic2. Therefore, incident sound from different angles experiences ICTD when reaching virtual mic1 and virtual mic2, due to the distance d between the two virtual mics. Signal processing is also used to control the directivity of the virtual mics based on the models of mic1 and mic2, including but not limited to beamforming (e.g., differential beamforming, summation beamforming, and noise reduction). This ensures that incident sound from different angles receives different signal strengths when reaching virtual mic1 and virtual mic2 due to the difference in directivity, resulting in ICTD for the same incident sound on both virtual mics. To achieve a stereo effect, appropriate ICTD and ICTD must be established between the left and right channels. Stereo audio is captured by adjusting the distance between the microphones and applying processing algorithms to achieve the desired directivity. When the microphones are spaced farther apart, ICTD increases, reducing reliance on directivity (i.e., ICLD decreases). Conversely, when the microphones are spaced closer together, ICTD decreases, requiring greater directivity (i.e., ICLD increases). By adjusting the ICLD and ICTD values, a desired sound field width can be created.

[0108] In some embodiments, the terminal is provided with two microphones, which form two symmetrical differential beams that are output to the left channel and the right channel as stereo audio respectively. Please refer to the above formulas 2 and 3. As shown in Figure 3, the two microphones can form a first-order differential beam, and two microphones can be used to form two mutually symmetrical beams, each pointing to the two end-fire directions opposite to the microphone array. The left channel is a beam with a zero point set at 0° and an end-fire direction of 180°, thereby suppressing the sound on the right. The right channel is a beam with a zero point set at 180° and an end-fire direction of 0°, thereby suppressing the sound on the left. As a result, the left and right channels construct different ICLDs for sounds incident at different angles. The matrix H can be called a filter for the differential beam. For the left channel, the filter e of the differential beam is jwr0a1,1 You can add a delay to the micR signal. For the right channel, the filter of the differential beam is e jwr0a1,1 By adding a delay to the micL signal and adding the mic spacing d, the left and right sounds arrive at different time intervals between the left and right channels, with the left sound arriving in the left channel before the right, and the right sound arriving in the right channel before the left. By adjusting the beam filter, you can control the ICLD between the left and right channels. ICTD is affected by the mic placement. Adjusting the ICLD and ICTD between the left and right channels based on the desired sound field creates a suitable stereo signal.

[0109] Based on the same concept, an embodiment of the present disclosure also provides an audio generating device.

[0110] It is understandable that the audio generation device provided by the embodiment of the present disclosure includes hardware structures and / or software modules corresponding to the execution of each function in order to realize the above functions. In combination with the units and algorithm steps of the various examples disclosed in the embodiment of the present disclosure, the embodiment of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiment of the present disclosure.

[0111] It should be noted that those skilled in the art will appreciate that the various implementation methods / embodiments involved in the embodiments of the present disclosure can be used in conjunction with the aforementioned embodiments or can be used independently. Whether used alone or in conjunction with the aforementioned embodiments, the implementation principles are similar. In the implementation of the present disclosure, some embodiments are described in terms of implementation methods used together. Of course, those skilled in the art will appreciate that such examples are not limitations on the embodiments of the present disclosure.

[0112] Fig. 4 is a block diagram of an audio generation device 100 according to an exemplary embodiment. Referring to Fig. 4 , the device 100 includes: a collection module 101 and a processing module 102 .

[0113] The acquisition module 101 is configured to acquire a first audio signal using at least two microphones of the terminal. The processing module 102 is configured to establish a desired target intensity difference based on the time difference between the first audio signals acquired by the at least two microphones. The time difference and the target intensity difference are negatively correlated. The first audio signal is processed using the target intensity difference to obtain a second audio signal, such that the intensity difference between the second audio signals corresponding to the at least two microphones meets the target intensity difference. Stereo audio is generated based on the second audio signal.

[0114] In some embodiments, the distance between the microphones determines the time difference between the first audio signals captured by the microphones. Processing module 102 constructs a desired target intensity difference based on the time difference between at least two microphones by determining a matrix based on the distance between the at least two microphones. The matrix is ​​used to represent the target intensity difference, and the distance and the matrix satisfy a preconfigured matrix function. Processing module 102 processes the first audio signal using the target intensity difference by processing the first audio signal using the matrix. The matrix is ​​used to increase or decrease the intensity difference of the first audio signal.

[0115] In some embodiments, the smaller the distance between the at least two microphones, the larger the value of the matrix. The larger the distance between the at least two microphones, the smaller the value of the matrix.

[0116] In some embodiments, the processing module 102 determines the distance between the microphones in the following manner: determining the distance between the microphones according to pre-stored distances between the microphones in the terminal.

[0117] In some embodiments, the processing module 102 determines the distance between the microphones in the following manner: obtaining the distance between at least two microphones in real time according to the folding state of the terminal.

[0118] In some embodiments, the processing module 102 acquires the distance between at least two microphones in real time according to the folding state of the terminal in the following manner: calculating the distance between the microphones according to the folding angle of the terminal.

[0119] In some embodiments, the terminal includes multiple sub-terminals, and the processing module 102 determines the distance between at least two microphones in the following manner: determining the distance between at least two microphones according to the distances between different sub-terminals.

[0120] Fig. 5 is a block diagram showing an audio generating device 200 according to an exemplary embodiment.

[0121] As shown in FIG. 5 , apparatus 200 may include one or more of the following components: a processing component 202 , a memory 204 , a power component 206 , a multimedia component 208 , an audio component 210 , an input / output (I / O) interface 212 , a sensor component 214 , and a communication component 216 .

[0122] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 202 may include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate interaction between the multimedia component 208 and the processing component 202.

[0123] The memory 204 is configured to store various types of data to support operations on the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0124] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 200.

[0125] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0126] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.

[0127] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0128] The sensor assembly 214 includes one or more sensors for providing various aspects of the status assessment of the device 200. For example, the sensor assembly 214 can detect the open / closed state of the device 200, the relative positioning of components, such as the display and keypad of the device 200. The sensor assembly 214 can also detect changes in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and temperature changes of the device 200. The sensor assembly 214 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0129] The communication component 216 is configured to facilitate wired or wireless communication between the device 200 and other devices. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0130] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0131] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 204 including instructions, which can be executed by the processor 220 of the apparatus 200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0132] The present disclosure utilizes the time difference between audio signals collected by different microphones on a terminal to construct a target intensity difference, which is then used to process the collected audio signals to obtain an audio signal that meets the target intensity difference. Stereo audio is then generated based on the audio signal that meets the target intensity difference. The present disclosure utilizes the above-mentioned scheme to obtain a stereo audio format that does not rely on microphones with complex structures, and the algorithm is simple and does not rely on high computing power. This approach can be applied to small devices to achieve stereo audio acquisition, improving efficiency and reducing costs.

[0133] It is understood that in this disclosure, "plurality" refers to two or more than two, and other quantifiers are similar. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects before and after are in an "or" relationship. The singular forms "a", "an", and "the" are also intended to include plural forms, unless the context clearly indicates otherwise.

[0134] It will be further understood that the terms "first," "second," and the like are used to describe various types of information, but such information should not be limited to these terms. These terms are used solely to distinguish information of the same type from one another and do not indicate a particular order or level of importance. In fact, the terms "first," "second," and the like are fully interchangeable. For example, first information could be referred to as second information, and similarly, second information could be referred to as first information without departing from the scope of this disclosure.

[0135] It is further understood that although operations are described in a particular order in the drawings in the embodiments of the present disclosure, this should not be construed as requiring that the operations be performed in the particular order shown or in a serial order, or that all of the operations shown be performed to obtain the desired results. In certain circumstances, multitasking and parallel processing may be advantageous.

[0136] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein.

[0137] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

Claims

1. An audio generation method, characterized in that: The method comprises: collecting a first audio signal using at least two microphones of the terminal; constructing a required target intensity difference according to a time difference between the first audio signals collected by the at least two microphones; wherein the time difference and the target intensity difference are negatively correlated; Processing the first audio signal using the target intensity difference to obtain a second audio signal, such that an intensity difference between the second audio signals corresponding to the at least two microphones satisfies the target intensity difference; Based on the second audio signal, stereo audio is generated.

2. The method according to claim 1, characterized in that The distance between the microphones determines the time difference between the first audio signals collected by the microphones. The step of constructing the required target intensity difference based on the time difference between the first audio signals collected by the at least two microphones includes: determining a matrix based on the distance between the at least two microphones, the matrix being used to represent the target intensity difference, wherein a preconfigured matrix function is satisfied between the distance and the matrix; The processing of the first audio signal by using the target intensity difference includes: Processing the first audio signal using the matrix; The matrix is ​​used to increase or decrease the intensity difference of the first audio signal.

3. The method according to claim 2, characterized in that The smaller the distance between the at least two microphones, the larger the value of the matrix; The greater the distance between the at least two microphones, the smaller the value of the matrix.

4. The method according to claim 2 or 3, characterized in that The distance between the at least two microphones is determined in the following manner: The distance between microphones is determined based on the pre-stored distance between the terminals.

5. The method according to claim 2 or 3, characterized in that The distance between the at least two microphones is determined in the following manner: The distance between the at least two microphones is acquired in real time according to the folding state of the terminal.

6. The method according to claim 5, characterized in that The acquiring the distance between the at least two microphones in real time according to the folding state of the terminal includes: The distance between the at least two microphones is calculated according to the folding angle of the terminal.

7. The method according to claim 2 or 3, characterized in that The terminal includes multiple sub-terminals, and the distance between the at least two microphones is determined in the following manner: The distance between the at least two microphones is determined according to the distances between different sub-terminals.

8. An audio generating device, characterized in that: The device comprises: an acquisition module, configured to acquire a first audio signal using at least two microphones of the terminal; A processing module is configured to construct a desired target intensity difference based on a time difference between first audio signals collected by the at least two microphones; wherein the time difference and the target intensity difference are negatively correlated; process the first audio signal using the target intensity difference to obtain a second audio signal such that an intensity difference between the second audio signals corresponding to the at least two microphones satisfies the target intensity difference; and generate stereo audio based on the second audio signal.

9. The device according to claim 8, characterized in that The distance between the microphones determines the time difference between the first audio signals collected by the microphones. The processing module constructs the required target intensity difference based on the time difference between the first audio signals collected by the at least two microphones in the following manner: determining a matrix based on the distance between the at least two microphones, the matrix being used to represent the target intensity difference, wherein a preconfigured matrix function is satisfied between the distance and the matrix; The processing module processes the first audio signal using the target intensity difference in the following manner: Processing the first audio signal using the matrix; The matrix is ​​used to increase or decrease the intensity difference of the first audio signal.

10. The device according to claim 9, characterized in that The smaller the distance between the at least two microphones, the larger the value of the matrix; The greater the distance between the at least two microphones, the smaller the value of the matrix.

11. The device according to claim 9 or 10, characterized in that The processing module determines the distance between the at least two microphones in the following manner: The distance between microphones is determined based on the pre-stored distance between the terminals.

12. The device according to claim 9 or 10, characterized in that The processing module determines the distance between the at least two microphones in the following manner: The distance between the microphones is obtained in real time according to the folding state of the terminal.

13. The device according to claim 12, characterized in that The processing module acquires the distance between the at least two microphones in real time according to the folding state of the terminal in the following manner: Calculate the distance between the microphones based on the folding angle of the terminal.

14. The device according to claim 9 or 10, characterized in that The terminal includes a plurality of sub-terminals, and the processing module determines the distance between the at least two microphones in the following manner: The distance between the at least two microphones is determined according to the distances between different sub-terminals.

15. An electronic device, characterized in that: include: a memory for storing instructions; as well as A processor, configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

16. A storage medium, characterized in that The storage medium stores instructions, and when the instructions are executed by the processor, the method according to any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Determining the inter-channel time difference of a multi-channel audio signal

    CN103403800A

  • Method for obtaining the same sound source through two microphones, and acquisition equipment

    CN107040843A

  • Multichannel directional pickup audio output method and system

    CN112261528A

  • Method and apparatus for voice or sound activity detection for spatial audio

    US20200314580A1