Multi-microphone array signal processing method and related product
By establishing a coordinate system and distance judgment, and combining fixed beam energy calculation and weighted mixing, the problem of near-field and far-field sound source identification in multi-microphone array signal processing is solved, achieving smooth audio processing effects on desktop devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AISPEECH CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-04-28
AI Technical Summary
Existing multi-microphone array signal processing technology is difficult to use to distinguish near-field and far-field sound sources on desktop devices, and its hardware design is complex and costly, making it unsuitable for effective application on small devices.
By establishing a coordinate system, calculating the distance from the sound source to the origin, determining whether gain is needed based on the distance, and applying linear gain or attenuation to the microphone array signal, combined with fixed beam energy calculation and weighted mixing, near-field pickup and far-field suppression are achieved.
This technology enables smooth audio transitions between channels as the sound source moves, avoiding abruptness, reducing far-field sound loudness, preventing near-field speech quality from being affected by distance estimation errors, and achieving a smooth transition between near-field pickup and far-field suppression.
Smart Images

Figure CN121940690A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech enhancement, and in particular relates to a multi-microphone array signal processing method and related products. Background Technology
[0002] Among related technologies, multi-microphone array signal processing techniques include the following: The first is a hardware method, using directional microphones to achieve directional sound pickup to a certain extent. The second is to use beamforming technology to locate the sound source, and then use a fixed beam to output only the speech from the specified direction. The third is to use a neural network to train a system that only outputs speech from a fixed direction.
[0003] These desktop microphone technologies are all used for remote sound pickup and can only identify direction, not distance. Currently, there is no practical near-field sound pickup technology for desktop sound reinforcement. To distinguish between near and far-field sound sources, the microphone array needs a sufficiently large microphone spacing, which exceeds the size limits of commonly used desktop devices. Some multi-microphone arrays achieve distance detection through spatial layout design and corresponding algorithms. However, they are expensive, have complex hardware designs (considering uniform microphone distribution), impose stringent requirements on product appearance and structural design, and are greatly affected by environmental noise and reverberation.
[0004] Human voices exist in a very wide frequency range, and the wavelength of the sound signal at each frequency is different. Therefore, to detect a sufficient number of frequencies, the microphone array needs to be able to detect a sufficiently large wavelength. This requires the microphone array to be arranged with a sufficiently large spacing, which is a significant limitation on hardware and appearance design. Consumer electronics emphasize small and exquisite products, and products with a length of one to two meters are not acceptable to the market. Summary of the Invention
[0005] This invention provides a multi-microphone array signal processing method, electronic device, and storage medium to at least solve one of the above-mentioned technical problems.
[0006] In a first aspect, embodiments of the present invention provide a multi-microphone array signal processing method, comprising: determining a coordinate system and the origin of the coordinate system based on the positions of multiple microphone arrays; acquiring multiple angles from the multiple microphone arrays to a sound source, and calculating the distance from the sound source to the origin based on the multiple angles; determining whether gain is required based on the distance from the sound source to the origin; if no gain is required, no processing is performed; if gain is required, calculating the multiple distances from the multiple microphone arrays to the sound source; and performing linear gain or linear attenuation on the signals acquired by the multiple microphone arrays based on the multiple distances to obtain multiple processed signals.
[0007] In a second aspect, embodiments of the present invention also provide an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the method described in the first aspect.
[0008] Thirdly, embodiments of the present invention also provide a storage medium storing a computer program thereon, characterized in that the computer program, when executed by a processor, implements the steps of the method described in the first aspect.
[0009] Fourthly, embodiments of this application also provide a computer program product, the computer program product including a computer program stored on a storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform any of the above-described multi-microphone array signal processing methods.
[0010] In the method of this application embodiment, after obtaining the positions of multiple microphone arrays, a coordinate system is established based on the positions of the microphone arrays and the origin of the coordinate system is determined. The angle of the sound source relative to its own coordinate system is obtained by combining fixed beam and energy calculation. Then, the distance of the sound source from the origin is calculated based on the obtained angle of each microphone array to the sound source. After obtaining the distance of the sound source to the origin of the coordinate system, it is detected whether the distance of the sound source to the origin of the coordinate system is close or far. If it is far, no processing is done. If it is close, the distance of the microphone array to the sound source is calculated separately. The signal obtained by the entire array of microphones is linearly amplified or linearly attenuated according to the distance. The signal obtained by each microphone is processed, so that when the sound source moves, the audio will smoothly "transfer" between channels without any abruptness. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart illustrating a multi-microphone array signal processing method according to an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the sound source calculation principle of a multi-microphone array signal processing method according to an embodiment of the present invention. Figure 3 This is a logic diagram for calculating the sound source coordinates of a multi-microphone array signal processing method according to an embodiment of the present invention; Figure 4 A signal mixing logic diagram of a multi-microphone array signal processing method provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.
[0015] This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, elements, data structures, etc., that perform a specific task or implement a specific abstract data type. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0016] In this invention, terms such as "module," "device," and "system" refer to relevant entities applied to a computer, such as hardware, combinations of hardware and software, software, or software in execution. More specifically, for example, an element can be, but is not limited to, a process running on a processor, a processor, an object, an executable element, an execution thread, a program, and / or a computer. Furthermore, an application program or script running on a server, and the server itself, can also be an element. One or more elements may be in an execution process and / or thread, and elements may be localized on a single computer and / or distributed across two or more computers, and may be run on various computer-readable media. Elements can also communicate via local and / or remote processes based on signals having one or more data packets, for example, signals from data interacting with another element in a local system, a distributed system, and / or interacting with other systems via signals over a network on the Internet.
[0017] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising" or "including" include not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0018] Currently, far-field local sound reinforcement products are in their initial stage, with related products only developing in the last two years. Many functions are still in the conceptual stage, and current far-field sound reinforcement products lack desktop sound reinforcement capabilities. Based on existing technology and design flaws, those skilled in the art typically employ the following approach: calculating the true sound source location by combining the directional sound source localization of a single device with the relative positions of two devices, thereby selecting the appropriate beam for controlling sound reinforcement at both near and far distances.
[0019] For multi-microphone array signal processing, please refer to [link / reference]. Figure 1 This illustrates a multi-microphone array signal processing method provided by an embodiment of the present invention.
[0020] As shown in 1, in step 101, the coordinate system and the origin of the coordinate system are determined according to the positions of the multiple microphone arrays; In correction 102, multiple angles from the multiple microphone arrays to the sound source are obtained respectively, and the distance from the sound source to the origin is calculated based on the multiple angles; In correction 103, it is determined whether gain needs to be applied based on the distance from the sound source to the origin. If no gain is needed, no processing is performed. If gain is needed, multiple distances from the multiple microphone arrays to the sound source are calculated respectively. In step 104, the signals acquired by the multiple microphone arrays are linearly gained or linearly attenuated according to the multiple distances to obtain multiple processed signals.
[0021] In this embodiment, for step 101, a coordinate system and the origin of the coordinate system are determined based on the positions of the multiple microphone arrays. For example, after obtaining the positions of the multiple microphone arrays, a coordinate system is established based on the positions of the microphone arrays and the origin of the coordinate system is determined.
[0022] For step 102, multiple angles from the multiple microphone arrays to the sound source are obtained respectively, and the distance from the sound source to the origin is calculated based on the multiple angles. For example, the angles of the sound source relative to its own coordinate system are obtained by using fixed beams combined with energy calculation, and then the distance from the sound source to the origin is calculated based on the obtained angles from each microphone array to the sound source. The fixed beams combined with energy calculation includes: pre-forming a set of "virtual microphones" (beams) pointing in different directions, and then determining the direction of the sound source by comparing which beam receives the strongest signal energy.
[0023] For step 103, it is determined whether gain needs to be applied based on the distance from the sound source to the origin. If no gain is needed, no processing is performed. If gain is needed, multiple distances from the multiple microphone arrays to the sound source are calculated. For example, after obtaining the distance from the sound source to the origin of the coordinate system, it is detected whether the distance from the sound source to the origin of the coordinate system is greater than 3m to determine whether it is a close distance or a long distance. If it is greater than 3m, it is a long distance and no processing is performed. If it is less than 3m, it is a close distance and the distance from the microphone arrays to the sound source is calculated. Here, 3m is just a specific example, and other distances are also possible. This application does not limit this.
[0024] For step 104, the signals acquired by the multiple microphone arrays are linearly amplified or linearly attenuated according to the multiple distances to obtain multiple processed signals. For example, after obtaining the distance between each microphone array and the sound source, the signals acquired by the entire array of microphones are linearly amplified or linearly attenuated according to the distance, thereby processing the signals acquired by each microphone.
[0025] The method in this embodiment, after obtaining the positions of multiple microphone arrays, establishes a coordinate system based on the positions of the microphone arrays and determines the origin of the coordinate system. It then obtains the angle of the sound source relative to its own coordinate system using a fixed beam and energy calculation. Next, it calculates the distance from the sound source to the origin based on the angle of each microphone array to the sound source. After obtaining the distance from the sound source to the origin of the coordinate system, it checks whether the distance is greater than 3m to determine whether it is a near or far distance. If it is greater than 3m, it is considered a far distance and no processing is performed; if it is less than 3m, it is considered a near distance and the distance from each microphone array to the sound source is calculated. Here, 3m is just a specific example; other distances can also be used, and this application is not limited in this regard. Based on the distance, the signals obtained by the entire microphone array are linearly amplified or linearly attenuated. The signals obtained by each microphone are processed to achieve a smooth "transition" of audio between channels when the sound source moves, without any abruptness.
[0026] In some optional embodiments, after linearly gaining or linearly attenuating the signals acquired by the multiple microphone arrays according to the multiple distances, the method further includes: weighted mixing of the multiple processed signals and outputting a mixed signal. For example, after acquiring multiple signals after linear gain or linear attenuation, the acquired signals are mixed and optimized by weighted mixing, and the optimized audio is output, thereby achieving the optimal output of the multi-microphone array and achieving the effect of near-field sound pickup and far-field suppression.
[0027] In some optional embodiments, after calculating the multiple distances from the multiple microphone arrays to the sound source, the method further includes: detecting multiple short-time signal-to-noise ratios (SNRs) of the signals acquired by the multiple microphone arrays in real time, and determining whether the multiple distances are accurate based on the multiple SNRs; for each distance, if it is inaccurate, recalculating it, and applying linear gain or linear attenuation based on the recalculated distance. For example, after obtaining the distances from the multiple microphone arrays to the sound source, the short-time SNR of the signals acquired by each microphone array is detected in real time, and the accuracy of the distance from the microphone array to the sound source is determined based on the short-time SNR. If the short-time SNR is very high, a misjudgment may have occurred (e.g., the sound source is actually very close, but the localization algorithm is temporarily erroneous). In this case, the distance from the microphone array to the sound source is recalculated, and the newly calculated distance is used to overwrite the previous distance. Linear gain or linear attenuation is applied based on the new distance from the microphone array to the sound source, thereby preventing the near-field speech quality from being suppressed due to temporary localization errors when the near-field speech quality is extremely high.
[0028] In some optional embodiments, the step of linearly gaining or linearly attenuating the signals acquired by the multiple microphone arrays according to the multiple distances to obtain multiple processed signals includes: acquiring multiple distances respectively, and detecting whether the multiple distances exceed a first threshold; if they do not exceed the first threshold, they are completely preserved and / or gained; if they exceed the first threshold, they detect whether the multiple distances exceed a second threshold; if they exceed the second threshold, they are strongly suppressed, wherein the first threshold is less than the second threshold; if they do not exceed the second threshold, they linearly attenuate the signals according to the distances exceeding the first threshold. For example, a first threshold and a second threshold are set, wherein the second threshold is greater than the first threshold; detecting whether the distances of the multiple microphone arrays from the sound source are greater than the first threshold; if they are less than the first threshold, they linearly gain the signals acquired by the corresponding microphone arrays; if they are greater than the first threshold, they detect whether the distances of the microphone arrays from the sound source are greater than the second threshold; if they are greater than the second threshold, they strongly suppress the signals acquired by the corresponding microphone arrays; if they are greater than the first threshold but less than the second threshold, they detect how much the distances of the microphone arrays from the sound source are greater than the first threshold, and linearly attenuate the signals according to the distances, thereby ensuring that the changes in gain or attenuation of the acquired signals are smooth and do not exhibit strong fluctuations. The number of thresholds can be one or more, and this application does not limit this. In a specific example, when the sound source is close to one microphone array and far from another, the signal acquired by the microphone array closer to the sound source is linearly amplified, while the signal acquired by the microphone array farther from the sound source is linearly attenuated. Furthermore, the processing method is adjusted according to the real-time changes in the sound source's position as the sound source moves. If the sound source gradually moves closer to the microphone array that was previously far away, the processing of the signal acquired by that microphone array will gradually change from linear attenuation to linear gain according to the distance. This achieves a corresponding change in signal processing method based on the distance the sound source moves, such as changing from linear attenuation to linear gain or vice versa.
[0029] In some optional embodiments, the plurality of microphone arrays are two microphone arrays. Determining the coordinate system and the origin of the coordinate system based on the positions of the plurality of microphone arrays includes: obtaining the positions of the plurality of microphone arrays; connecting the positions of the two microphone arrays with line segments, and establishing a coordinate system with the line segment as the coordinate axis, setting the center point of the line segment as the origin. For example, when there are two microphone arrays, the positions of the two microphone arrays are first obtained, then the x-axis is determined based on the line segment containing the two microphone arrays, and the center point of the line segment containing the two microphone arrays is set as the origin. Then the y-axis is set to establish the coordinate system, thereby ensuring that the coordinate system can be established based on the positions of the plurality of microphone arrays.
[0030] In some optional embodiments, obtaining the multiple angles from the multiple microphone arrays to the sound source includes: calculating the multiple angles from the multiple microphone arrays to the sound source by using fixed beamforming energy. For example, each microphone array can estimate the angle of the sound source relative to its own coordinate system by using fixed beamforming energy, thereby obtaining the multiple angles from the multiple microphone arrays to the sound source.
[0031] In some optional embodiments, calculating the distance from the sound source to the origin based on the plurality of angles includes: calculating the coordinates of the sound source in the coordinate system based on the plurality of angles; calculating the distance from the sound source to the origin of the coordinate system based on the coordinates of the sound source in the coordinate system. For example, after obtaining the angles from the multiple microphone arrays to the sound source, the position of the sound source in the coordinate system can be found through the multiple angles, and then the distance from the sound source to the origin can be calculated based on the position of the sound source in the coordinate system, thereby realizing the derivation of the actual coordinates of the sound source through planar positioning.
[0032] For details, please refer to [link / reference]. Figure 2 and Figure 3 It shows a schematic diagram of sound source calculation principle and a logic diagram of sound source coordinate calculation according to the present invention.
[0033] like Figure 2 and Figure 3 As shown, assume that two conventional microphone arrays (microphone boards) are placed on the same coordinate axis, with the origin in the middle and the two arrays on either side of the origin.
[0034] Usually, two microphone arrays are enough to define the coordinate system. For more microphone arrays, the coordinates are fixed within this coordinate system, so selecting two microphone arrays is sufficient to define the coordinate system.
[0035] Let array 1 be located at (-d, 0) and array 2 be located at (d, 0), where 2d is the distance between the two arrays (baseline length).
[0036] Each array has its own geometry (e.g., a linear array with known microphone spacing), allowing for independent estimation of the angle of the sound source relative to the array itself.
[0037] First, each array estimates the angle of the sound source relative to its own coordinate system by combining fixed beams with energy calculations.
[0038] Taking array 1 as an example: Assuming that the microphones of array 1 are arranged along the x-axis (or have a known geometric structure), then the estimated angle θ1 is the angle between the direction of the sound source and the x-axis of array 1's own coordinate system.
[0039] Secondly, in the global coordinate system: the position of array 1 is (-d, 0), and the position of array 2 is (d, 0). The position of the sound source in the global coordinate system is set to (x, y). From the perspective of array 1, the direction angle θ1 of the sound source satisfies: tan(θ1) = y / (x - (-d)) = y / (x+d) Similarly, from array 2, the direction angle θ2 of the sound source satisfies: tan(θ²) = y / (x - d) From the two equations above: y = tan(θ1) * (x + d) ... (1) y = tan(θ2) * (x - d) ... (2) Combine (1) and (2): tan(θ1)*(x+d) = tan(θ2)*(xd) => x * tan(θ1) + d * tan(θ1) = x * tan(θ2) - d * tan(θ2) => x * (tan(θ1) - tan(θ2)) = -d * tan(θ2) - d * tan(θ1) => x = [ -d * (tan(θ1) + tan(θ2)) ] / (tan(θ1) - tan(θ2)) = [ d * (tan(θ1) + tan(θ2)) ] / (tan(θ2) - tan(θ1)) Then substitute (1) or (2) to find y.
[0040] Calculate the distance from the sound source to the origin (0,0): r = sqrt(x^2 + y^2) Final judgment: If r <= 3 meters, it is a short distance; otherwise, it is a long distance.
[0041] Then, by matching the corresponding sound source beam with the sound source coordinates, a gentle and robust attenuation strategy is implemented, rather than a hard switch that is either on or off. The core goal is to significantly reduce the loudness of far-field sounds while avoiding damage to near-field speech due to distance estimation errors. This is an approach that is closer to "intelligent gain control" or "adaptive noise suppression".
[0042] In this context, a fixed beam involves determining the frequency, wavelength, array structure, number of microphones, and target direction, calculating the steering vector, and then calculating the phase difference generated by the wavefront from the target direction at each antenna. Finally, a weight vector is obtained to form a beam, which is then applied to receive data. The fixed beam at each point is a set of filters with defined parameters.
[0043] in, Figure 3 The dot product serves the following purpose: it determines whether the directions from the sound source to the two microphone arrays are parallel, thus providing a preliminary estimate of the sound source distance. The dot product value directly reflects the cosine of the angle (θ1-θ2) between the two direction vectors. When the two directions are perfectly aligned (parallel and in the same direction), the angle is 0°, and cos(0°) = 1. Therefore, the dot product ≈ 1. This allows for the exclusion of the far field (long distance) by ensuring the dot product is greater than cos5°, allowing subsequent gain adjustments only for the near field (short distance).
[0044] Please continue to refer to this. Figure 4 This illustrates a signal mixing logic diagram of the present invention.
[0045] like Figure 4 As shown, this application achieves smooth and continuous attenuation by mapping the sound source distance to a gain coefficient.
[0046] The linear decay method is as follows: Set two distance thresholds: R_near (e.g., 2.5m) and R_far (e.g., 3.5m). Define a gain factor G, ranging from G_min (e.g., 0.1, representing -20dB attenuation) to 1.0 (0dB attenuation, no attenuation). If R <= R_near: G = 1.0 (fully preserved, or even slightly gainable); if R >= R_far: G = G_min (strong suppression); if R_near < R < R_far: G = 1 - (1 - G_min) * (R - R_near) / (R_far - R_near). This avoids the "click" sound and voice interruption of hard switching. Near the 3m threshold, even with slight fluctuations in distance estimation, the gain change is gradual and does not cause drastic jumps in the output audio level. At the same time, the attenuation increases with distance.
[0047] To further avoid misjudgments, a protection mechanism is added. The short-time signal-to-noise ratio (SNR) of each array output signal is estimated in real time. If a channel is judged to be far-field (gain G is lowered) but its SNR is very high, a misjudgment may have occurred (e.g., the sound source is actually very close, but the localization algorithm temporarily malfunctions). In this case, the distance judgment can be overridden, limiting the attenuation of gain G, or slowly restoring it to a safe value (e.g., G > 0.5), instead of directly attenuating it to G_min. This prevents near-field speech quality from being suppressed due to temporary localization errors. Even if a brief, loud noise suddenly appears at a distance, its low SNR will not trigger protection and it will still be suppressed.
[0048] The innovation lies in how to mix the audio from two arrays to maximize their advantages. The aforementioned distance-based gain control processing is applied separately to the audio from both arrays. Instead of simply choosing one, a weighted mix is performed.
[0049] Output = (G1 * Signal1) + (G2 * Signal2), In this system, G1 and G2 are the gain coefficients calculated by arrays 1 and 2 based on their distance from the sound source, respectively. For a near-field sound source, the G value of the array closer to it is close to 1, while the G value of the array farther away is very small. The final output naturally prioritizes the signal from the near-field arrays, while the signal from the far-field arrays is attenuated to almost inaudible, achieving "relative suppression." Even if the localization or pickup of one array temporarily fails, the system can still rely on the other array. When the sound source moves, the audio smoothly transitions between the two channels without any abruptness.
[0050] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of combined actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps can be performed in other orders or simultaneously according to the present invention. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0051] In some embodiments, the present invention provides a non-volatile computer-readable storage medium storing one or more programs including execution instructions, which can be read and executed by an electronic device (including but not limited to a computer, server, or network device, etc.) to perform any of the multi-microphone array signal processing methods described above.
[0052] In some embodiments, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform any of the above-described multi-microphone array signal processing methods.
[0053] In some embodiments, the present invention also provides an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a multi-microphone array signal processing method.
[0054] Figure 5 This is a schematic diagram of the hardware structure of an electronic device for performing a multi-microphone array signal processing method according to another embodiment of this application, as shown below. Figure 5 As shown, the device includes: One or more processors 510 and memory 520, Figure 5 Take the 510 processor as an example.
[0055] The apparatus for performing the multi-microphone array signal processing method may further include an input device 530 and an output device 540.
[0056] The processor 510, memory 520, input device 530, and output device 540 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.
[0057] The memory 520, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the multi-microphone array signal processing method in the embodiments of this application. The processor 510 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 520, thereby implementing the multi-microphone array signal processing method in the above-described method embodiments.
[0058] Memory 520 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the multi-microphone array signal processing device, etc. Furthermore, memory 520 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 520 may optionally include memory remotely located relative to processor 510, and this remote memory may be connected to the multi-microphone array signal processing device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0059] Input device 530 can receive input digital or character information, and generate signals related to user settings and function control of the multi-microphone array signal processing device. Output device 540 may include display devices such as a display screen.
[0060] The one or more modules are stored in the memory 520, and when executed by the one or more processors 510, they perform the multi-microphone array signal processing method in any of the above method embodiments.
[0061] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.
[0062] The electronic devices in this application embodiments exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.
[0063] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include: PDAs, MIDs, and UMPCs, etc.
[0064] (3) Portable entertainment devices: These devices can display and play multimedia content. This category includes audio and video players, handheld game consoles, e-book readers, as well as smart toys and portable car navigation devices.
[0065] (4) Other airborne electronic devices with data interaction capabilities, such as vehicle-mounted systems installed on vehicles.
[0066] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0067] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A multi-microphone array signal processing method, comprising: The coordinate system and its origin are determined based on the positions of multiple microphone arrays; Multiple angles from the multiple microphone arrays to the sound source are obtained respectively, and the distance from the sound source to the origin is calculated based on the multiple angles; Whether gain needs to be applied is determined based on the distance from the sound source to the origin. If no gain is needed, no processing is performed. If gain is needed, the distances from the multiple microphone arrays to the sound source are calculated separately. Based on the multiple distances, the signals acquired by the multiple microphone arrays are linearly gained or linearly attenuated to obtain multiple processed signals.
2. The method according to claim 1, characterized in that, After applying linear gain or linear attenuation to the signals acquired by the plurality of microphone arrays according to the plurality of distances, the method further includes: The multiple processed signals are weighted and mixed to output the mixed signal.
3. The method according to claim 1, characterized in that, After calculating the plurality of distances from the plurality of microphone arrays to the sound source, the method further includes: The short-time signal-to-noise ratios of the signals acquired by the multiple microphone arrays are detected in real time, and the accuracy of the multiple distances is determined based on the multiple short-time signal-to-noise ratios. For each distance, if it is inaccurate, it is recalculated, and linear gain or linear decay is applied based on the recalculated distance.
4. The method according to claim 1, characterized in that, The process of linearly gaining or linearly attenuating the signals acquired by the multiple microphone arrays based on the multiple distances to obtain multiple processed signals includes: Multiple distances are acquired, and it is detected whether the multiple distances exceed a first threshold. If they do not exceed the first threshold, they are completely preserved and / or gain is applied. If the first threshold is exceeded, then it is detected whether the plurality of distances exceed the second threshold. If the second threshold is exceeded, then it is strongly suppressed, wherein the first threshold is less than the second threshold. If the second threshold is not exceeded, then linear decay is applied based on the distance exceeding the first threshold.
5. The method according to claim 1, characterized in that, The plurality of microphone arrays comprises two microphone arrays, and the determination of the coordinate system and the origin of the coordinate system based on the positions of the plurality of microphone arrays includes: Obtain the positions of multiple microphone arrays; Connect the positions of the two microphone arrays with a line segment, and establish the coordinate system with the line segment as the coordinate axis, setting the center point of the line segment as the origin.
6. The method according to claim 1, characterized in that, The step of obtaining the multiple angles from the multiple microphone arrays to the sound source includes: calculating the multiple angles from the multiple microphone arrays to the sound source by combining energy with a fixed beam.
7. The method according to claim 1, characterized in that, The step of calculating the distance from the sound source to the origin based on the multiple angles includes: The coordinates of the sound source in the coordinate system are calculated based on the multiple angles. Calculate the distance from the sound source to the origin of the coordinate system based on the coordinates of the sound source in the coordinate system.
8. An electronic device comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1 to 7.
9. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 7.