A method, apparatus and medium for directional sound pickup of a microphone device
By constructing a set of directional angles and calculating cross-correlation energy to determine the direction of the main sound source, and combining it with frequency dimension speech enhancement weighting, the problems of unstable main sound source recognition and insufficient frequency dimension processing in microphone devices under complex environments are solved, achieving higher speech clarity and anti-interference capability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING YITONG TECH
- Filing Date
- 2025-10-24
- Publication Date
- 2026-04-28
AI Technical Summary
Existing microphone equipment suffers from unstable identification of the main sound source direction in complex environments and has limited ability to refine the frequency dimension speech enhancement weights, resulting in insufficient sound pickup accuracy and clarity.
By constructing a set of directional angles, calculating the cross-correlation energy to determine the direction of the main sound source, constructing enhancement weights based on the direction of the main sound source, identifying the suppression weights of the interference direction, generating the fused spectrum of the main direction and the interference direction, and performing weighted processing in combination with the speech enhancement weights of the frequency dimension, the enhanced speech signal is finally generated.
It achieves accurate identification of the direction of the target sound source in scenarios with multiple interfering sound sources, improves speech clarity and anti-interference ability, and enhances the enhancement and suppression effect of speech signals.
Smart Images

Figure CN121310040B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech signal processing technology, and in particular to a directional pickup method, device and medium for a microphone. Background Technology
[0002] In the field of speech acquisition and processing, the directional pickup method of microphone devices has always been an important means to improve speech intelligibility and anti-interference performance. Conventional methods mostly rely on microphone arrays to acquire sound source signals, and perform energy analysis on signals from different directions through spatial filtering, beamforming, and cross-correlation operations. They also combine weighted superposition strategies to enhance the target sound source and suppress non-target sound sources. This type of method is widely used in scenarios such as conference systems, intelligent voice interaction, and remote communication. By processing signals in the time and frequency domain, the intelligibility and robustness of the target speech can be improved to a certain extent.
[0003] However, conventional methods still have certain limitations in complex environments: on the one hand, conventional methods mainly rely on the comparison of directional angle energy to determine the direction of the main sound source. When multiple interference sources exist at the same time and their energies are similar, the identification of the main direction may not be stable enough, thus affecting the accuracy of sound pickup; on the other hand, conventional methods often use a uniform weighting or simple energy ratio method to perform speech enhancement weighting in the frequency dimension, which makes it difficult to perform fine processing on the differentiated noise characteristics of different frequency bands. Therefore, it is easy to have insufficient enhancement or residual interference in the spectrum recovery and temporal reconstruction stages. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a directional sound pickup method for microphone devices to solve the problems of insufficient stability in main sound source direction recognition and limited fine-grained processing capability of frequency dimension speech enhancement weights in existing technologies.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides a directional sound pickup method for a microphone device, comprising: acquiring audio signals from each microphone and performing frequency domain transformation on the audio signals to construct a set of directional angles;
[0008] The cross-correlation energy of each direction is calculated based on the set of directional angles to form a directional energy vector, and the direction with the maximum energy in the directional energy vector is taken as the direction of the main sound source at the current time.
[0009] Centered on the direction of the main sound source, a main direction enhancement weight is constructed, and interference directions with similar energy to the main direction are identified to generate a suppression weight for each interference direction.
[0010] The directional gain at the current time is synthesized based on the main direction enhancement weight, the interference direction suppression weight, and the neutral direction weight.
[0011] The signals of each channel are weighted and fused according to the directional gain to generate the main directional fused spectrum and the interference directional fused spectrum, respectively.
[0012] Based on the energy comparison between the main direction and the interference direction, the speech enhancement weight in the frequency dimension is calculated. The speech enhancement weight is used to weight the fused spectrum in the main direction to obtain the enhanced speech spectrum. Finally, the pickup signal is generated by transforming from the frequency domain to the time domain.
[0013] In a preferred embodiment of the directional sound pickup method for the microphone device described in this invention, the steps of acquiring the audio signals from each microphone and performing frequency domain transformation on the audio signals to construct a set of directional angles are as follows:
[0014] The audio signals from each microphone are collected to generate a multi-channel time-domain audio signal set;
[0015] A short-time Fourier transform is performed on a collection of multi-channel time-domain audio signals to obtain a complex spectral tensor.
[0016] Based on the spatial orientation of the discrete microphone array with a preset angle step, construct a set of directional angles.
[0017] In a preferred embodiment of the directional sound pickup method for the microphone device described in this invention, the following steps are taken: Calculating the cross-correlation energy of each direction based on the set of directional angles to form a directional energy vector, and using the direction with the maximum energy in the directional energy vector as the direction of the main sound source at the current time.
[0018] Select the complex spectral tensor and the set of directional angles at the current time, and construct a microphone propagation time delay table for each directional angle;
[0019] Based on the complex spectral tensor and propagation time delay table, calculate the cross-correlation energy of each direction angle at the current time, and generate the cross-correlation energy of the direction angle;
[0020] Combine all cross-correlation values of direction angles in order of direction angle index to form the direction energy vector at the current time;
[0021] The direction angle with the largest cross-correlation energy under the direction energy vector is selected as the direction of the main sound source at the current time.
[0022] In a preferred embodiment of the directional pickup method for the microphone device described in this invention, the steps include: constructing a main direction enhancement weight centered on the main sound source direction, identifying interference directions with energy similar to the main direction, and generating a suppression weight for each interference direction.
[0023] Use the current direction of the main sound source as the central direction angle for constructing the main direction enhancement weights;
[0024] The Gaussian function exponent is calculated based on the angle difference between the main sound source direction and each direction angle in the set of direction angles to obtain the main direction enhancement weight;
[0025] The current time's main sound source direction and direction energy vector are used as the basis for identifying the direction of interference;
[0026] The cross-correlation energy of all non-main sound source directions in the directional energy vector is compared with the cross-correlation energy of the main sound source direction. The directional angles that meet the relative energy determination conditions are selected to form the interference directions.
[0027] The suppression weight for each interference direction is calculated based on the interference direction and the direction energy vector.
[0028] As a preferred embodiment of the directional pickup method for the microphone device described in this invention, the specific steps for synthesizing the directional gain at the current time based on the main direction enhancement weight, the interference direction suppression weight, and the neutral direction weight are as follows:
[0029] Set neutral direction weights, and divide the neutral direction set according to the main direction enhancement weight, the interference direction suppression weight, and the direction angle set;
[0030] According to the assignment rules for the main direction enhancement weight, the interference direction suppression weight, and the neutral direction weight, a direction gain is assigned to each direction angle in the direction angle set.
[0031] In a preferred embodiment of the directional pickup method for the microphone device described in this invention, the step of weighted fusing of the signals from each channel based on directional gain to generate a main directional fused spectrum and an interference directional fused spectrum comprises the following steps:
[0032] Based on the directional gain, complex spectral tensor, and directional angle set, construct the relationship between the directional angle and channel index for each microphone;
[0033] The current time is divided into the main direction channel and the interference direction channel based on the set of directional angles, the direction of the main sound source, and the direction of interference.
[0034] The spectrum signals of each channel in the main direction channel are weighted and summed according to the weight of the direction angle in the direction gain to generate the main direction fused spectrum;
[0035] The spectrum signals of each channel in the interference direction channel are weighted and summed according to the weight of the direction angle in the direction gain to generate the interference direction fused spectrum.
[0036] In a preferred embodiment of the directional pickup method for the microphone device described in this invention, the steps include: calculating the speech enhancement weights in the frequency dimension based on the energy comparison between the main direction and the interference direction, and then using these speech enhancement weights to weight the fused spectrum in the main direction to obtain the enhanced speech spectrum.
[0037] The energy value of the main direction fused spectrum at each frequency point is calculated based on the main direction fused spectrum, and the main direction energy spectrum sequence is generated.
[0038] The energy value of the interference direction fused spectrum at each frequency point is calculated based on the interference direction fused spectrum, and an interference direction energy spectrum sequence is generated;
[0039] Speech enhancement weights are constructed based on the energy ratio of the main direction energy spectrum sequence and the interference direction fused spectrum sequence.
[0040] The enhanced speech spectrum is obtained by weighting the fusion spectrum in the main direction based on the speech enhancement weights.
[0041] In a preferred embodiment of the directional pickup method for the microphone device described in this invention, the specific steps for generating the final pickup signal through frequency-domain to time-domain transformation are as follows:
[0042] Phase recovery is performed on the enhanced speech spectrum to generate the enhanced speech complex spectrum;
[0043] Perform an inverse Fourier transform on the complex spectrum of the enhanced speech to generate an enhanced speech frame sequence;
[0044] The enhanced speech frame sequence is reconstructed by overlapping addition to generate the final pickup signal.
[0045] In a second aspect, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the directional pickup method of the microphone device as described in the first aspect of the present invention.
[0046] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the directional pickup method of the microphone device as described in the first aspect of the present invention.
[0047] The beneficial effects of this invention are as follows: By calculating the cross-correlation energy of each direction based on the set of direction angles and determining the direction of the main sound source, accurate identification of the target sound source direction is achieved, improving the reliability of direction selection in scenarios with multiple interfering sound sources; by generating the main direction fusion spectrum and the interference direction fusion spectrum separately based on the weighted directional gain, the separation processing of the target spectrum and the interference spectrum is achieved, enhancing the fine control capability of the speech enhancement weighting link; by constructing the main direction enhancement weight with the main sound source direction as the center and generating suppression weights for interference directions with similar energy, the enhancement of the target speech signal and the effective suppression of the interference speech signal are achieved, thereby improving speech clarity. Attached Figure Description
[0048] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 This is a flowchart of a directional pickup method for microphone devices.
[0050] Figure 2 A flowchart for determining the direction of the main sound source.
[0051] Figure 3 This is a flowchart for synthesizing directional gain.
[0052] Figure 4 The flowchart for generating the final pickup signal. Detailed Implementation
[0053] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0054] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0055] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0056] Reference Figures 1-4As one embodiment of the present invention, this embodiment provides a directional sound pickup method for a microphone device, comprising the following steps:
[0057] S1. Collect audio signals from each microphone and perform frequency domain transformation on the audio signals to construct a set of directional angles.
[0058] The audio signals from each microphone are collected to generate a multi-channel time-domain audio signal set.
[0059] Furthermore, a microphone array is deployed in the scene to be picked up, and the raw audio signal of each microphone is collected through synchronous sampling.
[0060] Assume the sampling frequency of each microphone is The duration of a single frame is Then the length of each sampling frame is .
[0061] After all channels are acquired, a multi-channel time-domain signal set is formed, represented as:
[0062] ;
[0063] in, Represents a collection of multi-channel time-domain signals. Indicates the first The microphone in time The sampled values.
[0064] A short-time Fourier transform is performed on a collection of multi-channel time-domain audio signals to obtain a complex spectral tensor.
[0065] Furthermore, a short-time Fourier transform is performed on each channel signal in the multi-channel time-domain audio signal set to convert the time-domain signal into a complex spectral tensor in the time-frequency domain.
[0066] Specifically, the short-time Fourier transform uses a frame length of... The sliding window function is used for frame division and sliding window function weighting; based on the frame division and sliding window function weighting, the spectrum is calculated to obtain the complex spectrum tensor.
[0067] Based on the spatial orientation of the discrete microphone array with a preset angle step, construct a set of directional angles.
[0068] Furthermore, based on the requirements of the microphone array spatial angle in the sound pickup scenario, the directional angle step size is set, which is the minimum interval between two adjacent discrete directional angles.
[0069] The principles that the direction angle step size must satisfy are as follows: A1-A2:
[0070] A1. The directional angle step size is less than or equal to the minimum resolvable angle allowed by the microphone array, wherein the minimum resolvable angle is determined by the array aperture and wavelength.
[0071] A2, the set of orientation angles covers the entire 360° range.
[0072] Given a fixed step size, the set of direction angles is generated discretely using a linearly increasing method, as follows:
[0073] ;
[0074] in, Represents the set of direction angles. Indicates the direction angle. Indicates the direction angle index. Indicates the direction angle step size. Represents a set of integers.
[0075] S2. Calculate the cross-correlation energy of each direction based on the direction angle set, form a direction energy vector, and take the direction with the maximum energy in the direction energy vector as the main sound source direction at the current time.
[0076] Select the complex spectral tensor and the set of directional angles at the current time, and construct a microphone propagation time delay table for each directional angle.
[0077] Furthermore, the microphone propagation time delay for each directional angle is expressed as:
[0078] ;
[0079] in, Indicates the first The microphone is at the azimuth angle The propagation time delay below Indicates the microphone spacing. This indicates the speed of sound in the air.
[0080] Based on the calculated relationship between microphone propagation time delay at each directional angle, a microphone propagation time delay table is constructed for each directional angle.
[0081] Based on the complex spectral tensor and propagation time delay table, calculate the cross-correlation energy of each direction angle at the current time, and generate the cross-correlation energy of the direction angle.
[0082] Furthermore, the cross-correlation energy is calculated based on the propagation time delay between the first microphone and the remaining microphones, and is expressed as:
[0083] ;
[0084] in, Indicates time Downward Angle The cross-correlation energy, Indicates the number of microphones. Indicates the time of the first microphone. frequency The complex spectral tensor under the given conditions, Indicates the first The microphone in time frequency The conjugate of the complex spectrum, Represents the imaginary unit. This represents the function for extracting the real part of a complex number.
[0085] The cross-correlation values of all direction angles are combined in the order of their direction angle indices to form the direction energy vector for the current time.
[0086] The direction angle with the highest cross-correlation energy under the direction energy vector is selected as the direction of the main sound source at the current time.
[0087] S3. Using the direction of the main sound source as the center, construct the main direction enhancement weight, identify the interference direction with similar energy to the main direction, and generate the suppression weight for each interference direction.
[0088] The direction of the main sound source at the current time is used as the central direction angle for constructing the main direction enhancement weight.
[0089] The Gaussian function exponent is calculated based on the angle difference between the main sound source direction and each direction angle in the set of direction angles to obtain the main direction enhancement weight.
[0090] Furthermore, the principal direction enhancement weight function is expressed as:
[0091] ;
[0092] in, Indicates the first Each direction angle in time The main direction below increases the weight. Represents the first in the set of direction angles One direction angle, Indicates time The direction angle of the main sound source below, This represents the standard deviation parameter of the angle.
[0093] It should be noted that the angle standard deviation parameter is used to control the expansion range of the main direction enhancement weight function. The angle standard deviation parameter is set based on the angle step size, the acceptable offset range of the main sound source, and the directional resolution capability of the microphone array. Specifically, the minimum resolvable angle interval is determined according to the angle step size of the directional angle set. For example, when the angle step size of the directional angle set is five degrees, the interval between each two adjacent directional angles is five degrees. Then, the coverage width of the main direction enhancement weight is determined according to the possible angular offset range of the main sound source in the sound pickup scenario. For example, the enhancement range of the main sound source is set to ±15 degrees. Finally, the angle step size is correlated with the angle coverage width to calculate the number of directional angles that can be covered, thereby setting the angle standard deviation parameter to the angle value corresponding to the enhancement range. For example, the angle standard deviation parameter is set to 15 degrees.
[0094] The current time's main sound source direction and direction energy vector are used as the basis for identifying the direction of interference.
[0095] The cross-correlation energy of all non-main sound source directions in the directional energy vector is compared with the cross-correlation energy of the main sound source direction. The directional angles that meet the relative energy determination conditions are selected to form the interference directions.
[0096] Furthermore, based on historical interference identification data, an interference identification threshold is set.
[0097] The relative energy criterion is as follows:
[0098] If the direction angle The direction angle is determined if the following conditions are met. Interference direction angle:
[0099] ;
[0100] in, Indicates time The first non-main sound source direction Cross-correlation energy of each direction angle Indicates time The cross-correlation energy of the directional angle of the main sound source direction. This represents the interference identification threshold. The interference direction is formed by combining all identified interference direction angles.
[0101] The suppression weight for each interference direction is calculated based on the interference direction and the direction energy vector.
[0102] Furthermore, the formula for calculating the suppression weight is expressed as follows:
[0103] ;
[0104] in, Indicates time The first in the direction of lower interference Suppression weights for each directional angle, Represents the first in the set of direction angles One direction angle, Represents the set of direction angles. Indicates time The first in the set of downward direction angles Cross-correlation energy of each direction angle.
[0105] S4. Based on the main direction enhancement weight, the interference direction suppression weight, and the neutral direction weight, synthesize the direction gain for the current time.
[0106] Set neutral direction weights and divide the neutral direction set according to the main direction enhancement weight, the interference direction suppression weight, and the direction angle set.
[0107] Furthermore, the median of the weights for enhancing the main direction and suppressing the interference direction is taken as the weight for the neutral direction.
[0108] In time Using the directional angle set as the basic index range, for directional angles belonging to the main sound source direction, directional angles belonging to the interference direction, and directional angles belonging to neither the main sound source direction nor the interference direction, the weight values corresponding to the main direction enhancement weight, the suppression weight values corresponding to the interference direction suppression weight, and the neutral direction weight are read.
[0109] According to the assignment rules for the main direction enhancement weight, the interference direction suppression weight, and the neutral direction weight, a direction gain is assigned to each direction angle in the direction angle set.
[0110] Furthermore, the assignment rule for directional gain is expressed as follows:
[0111] ;
[0112] in, Indicates direction angle In time Directional gain below, Indicates the weight in the neutral direction. Indicates the directional angle of the main sound source. The azimuth angle of the standard interference direction. The direction angle indicates the neutral direction.
[0113] S5. Weighted fusion of signals from each channel is performed based on directional gain to generate the main directional fusion spectrum and the interference directional fusion spectrum, respectively.
[0114] Based on the directional gain, complex spectral tensor, and directional angle set, construct the relationship between the directional angle and channel index for each microphone.
[0115] Furthermore, the microphone array is configured as a uniform linear distribution, containing... There are 10 microphones, which are evenly arranged horizontally with a constant spacing between them. ; Select the center position of the microphone array as the origin of the relative coordinate system, and construct a centrally symmetric spatial coordinate system with the x-axis in the two-dimensional plane as the arrangement direction.
[0116] Based on the relative spatial coordinates in the constructed microphone array and combined with the orientation of the microphone array, the main orientation of each microphone is derived. The main orientation angle of each microphone is mapped to the corresponding orientation angle in the orientation angle set by rounding down. According to the matching results between the channel index of each microphone and the main orientation angle, an association set between each orientation angle and multiple channel indices in the orientation angle set is established to form a relationship table between orientation angles and channel indices.
[0117] The current time is divided into the main direction channel and the interference direction channel based on the set of directional angles, the direction of the main sound source, and the direction of interference.
[0118] Furthermore, in the direction angle and through index relationship table, all channel indices associated with direction angles that are completely consistent with the direction of the main sound source are filtered to form the main direction channel at the current time.
[0119] For each direction angle in the interference direction, find the corresponding channel index in the direction angle and channel index relationship table, and merge all channel indices into the interference direction channel.
[0120] If there are intersecting channels in the main direction channel and the interference direction channel, a priority processing strategy is adopted, and the intersecting channel is given priority to be assigned to the main direction channel and removed from the interference direction channel.
[0121] The spectral signals of each channel in the main direction channel are weighted and summed according to the weight of the directional angle in the directional gain to generate the main direction fused spectrum.
[0122] Furthermore, the formula for calculating the fused spectrum in the main direction is expressed as:
[0123] ;
[0124] in, This represents the main direction fused spectrum at time and frequency. This indicates the main direction channel at the current frame time. Indicates direction angle In time Directional gain below, Indicates the first, The microphone in time frequency The complex spectrum below.
[0125] The spectrum signals of each channel in the interference direction channel are weighted and summed according to the weight of the direction angle in the direction gain to generate the interference direction fused spectrum.
[0126] Furthermore, the formula for calculating the interference direction fusion spectrum is expressed as:
[0127] ;
[0128] in, This represents the interference direction fusion spectrum at time and frequency. This indicates the interference direction channel at the current frame time.
[0129] S6. Based on the energy comparison between the main direction and the interference direction, calculate the speech enhancement weight in the frequency dimension, use the speech enhancement weight to weight the fused spectrum in the main direction to obtain the enhanced speech spectrum, and generate the final pickup signal through frequency domain to time domain transformation.
[0130] The energy value of the main direction fused spectrum at each frequency point is calculated based on the main direction fused spectrum, and the main direction energy spectrum sequence is generated.
[0131] The energy value of the interference direction fused spectrum at each frequency point is calculated based on the interference direction fused spectrum, and an interference direction energy spectrum sequence is generated.
[0132] Furthermore, assuming the fused spectrum in the main direction or the fused spectrum in the interference direction is a complex spectral sequence, the energy value calculation formula is expressed as:
[0133] ;
[0134] in, Indicates frequency The energy value at that location, Indicates frequency The amplitude of the complex spectrum at that point. Indicates frequency The real part of the place, Indicates frequency The virtual part of the location.
[0135] Based on the energy value calculation formula, the main direction energy spectrum sequence is calculated and generated. and interference direction energy spectrum sequence .
[0136] Speech enhancement weights are constructed based on the energy ratio of the main direction energy spectrum sequence and the interference direction fused spectrum sequence.
[0137] Furthermore, the speech enhancement weights are expressed in the form of proportional normalization for the fused spectral sequence of the main direction energy spectrum and the interference direction energy spectrum.
[0138] The enhanced speech spectrum is obtained by weighting the fusion spectrum in the main direction based on the speech enhancement weights.
[0139] Furthermore, the speech enhancement weights are applied to the main direction fusion spectrum, and weighted point by point according to the frequency points to obtain the enhanced speech spectrum.
[0140] Phase recovery is performed on the enhanced speech spectrum to generate the enhanced speech complex spectrum.
[0141] Furthermore, at each time and frequency point in the enhanced speech spectrum, the amplitude value of the enhanced speech spectrum is used as the complex modulus, and the phase value of the main direction fused spectrum is used as the complex argument to reconstruct the complex spectrum of the enhanced speech.
[0142] Perform an inverse Fourier transform on the complex spectrum of the enhanced speech to generate an enhanced speech frame sequence.
[0143] Furthermore, an inverse fast Fourier transform is performed on each time step of the enhanced speech complex spectrum to obtain a sequence of equal-length sampling points in the time domain, thereby constructing an enhanced speech frame sequence.
[0144] The enhanced speech frame sequence is reconstructed by overlapping addition to generate the final pickup signal.
[0145] Furthermore, temporal reconstruction is performed on all enhanced speech frame sequences using a frame-shift overlap method. Specifically, by setting the frame length and frame shift, adjacent frames are aligned according to the frame shift interval, and the sampling points in the overlapping area are summed with amplitude weighting to recover the continuous temporal speech signal and finally generate the pickup signal.
[0146] This embodiment also provides a computer device applicable to the directional pickup method of a microphone device, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the directional pickup method of the microphone device as proposed in the above embodiment.
[0147] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0148] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the directional pickup method for a microphone device as described in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0149] In summary, this invention achieves accurate identification of the target sound source direction by calculating the cross-correlation energy of each direction based on the set of direction angles and determining the direction of the main sound source, thus improving the reliability of direction selection in scenarios with multiple interfering sound sources. By generating the main direction fusion spectrum and the interference direction fusion spectrum separately based on the weighted directional gain, the target spectrum and the interference spectrum are separated and processed, enhancing the fine control capability of the speech enhancement weighting process. By constructing the main direction enhancement weight with the main sound source direction as the center and generating suppression weights for interference directions with similar energy, the enhancement of the target speech signal and the effective suppression of the interference speech signal are achieved, thereby improving speech clarity.
[0150] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A directional sound pickup method for a microphone device, characterized in that: include, Collect audio signals from each microphone and perform frequency domain transformation on the audio signals to construct a set of directional angles; The cross-correlation energy of each direction is calculated based on the set of directional angles to form a directional energy vector, and the direction with the maximum energy in the directional energy vector is taken as the direction of the main sound source at the current time. Centered on the direction of the main sound source, a main direction enhancement weight is constructed, and interference directions with similar energy to the main direction are identified to generate a suppression weight for each interference direction. The directional gain at the current time is synthesized based on the main direction enhancement weight, the interference direction suppression weight, and the neutral direction weight. The signals of each channel are weighted and fused according to the directional gain to generate the main directional fused spectrum and the interference directional fused spectrum, respectively. Based on the energy comparison between the main direction and the interference direction, the speech enhancement weight in the frequency dimension is calculated. The speech enhancement weight is used to weight the fused spectrum in the main direction to obtain the enhanced speech spectrum. Finally, the pickup signal is generated by transforming from the frequency domain to the time domain. The generation of suppression weights for each interference direction includes using the current time's main sound source direction as the center direction angle for constructing the main direction enhancement weights; The Gaussian function exponent is calculated based on the angle difference between the main sound source direction and each direction angle in the set of direction angles to obtain the main direction enhancement weight; The current time's main sound source direction and direction energy vector are used as the basis for identifying the direction of interference; The cross-correlation energy of all non-main sound source directions in the directional energy vector is compared with the cross-correlation energy of the main sound source direction. The directional angles that meet the relative energy determination conditions are selected to form the interference directions. Calculate the suppression weight for each interference direction based on the interference direction and the direction energy vector; The synthesized current-time directional gain includes setting neutral direction weights and dividing the neutral direction set according to the main direction enhancement weight, the interference direction suppression weight, and the direction angle set; The neutral direction weight includes taking the median value of the main direction enhancement weight and the interference direction suppression weight as the neutral direction weight; According to the assignment rules of the main direction enhancement weight, the interference direction suppression weight and the neutral direction weight, each direction angle in the direction angle set is assigned a direction gain. The process of generating the main direction fusion spectrum and the interference direction fusion spectrum respectively includes constructing the relationship between the directional angle and the channel index of each microphone based on the directional gain, the complex spectral tensor, and the directional angle set. The current time is divided into the main direction channel and the interference direction channel based on the set of directional angles, the direction of the main sound source, and the direction of interference. The spectrum signals of each channel in the main direction channel are weighted and summed according to the weight of the direction angle in the direction gain to generate the main direction fused spectrum; The spectrum signals of each channel in the interference direction channel are weighted and summed according to the weight of the direction angle in the direction gain to generate the interference direction fused spectrum; The acquisition of the enhanced speech spectrum includes calculating the energy value of the main direction fusion spectrum at each frequency point based on the main direction fusion spectrum, and generating a main direction energy spectrum sequence. The energy value of the interference direction fused spectrum at each frequency point is calculated based on the interference direction fused spectrum, and an interference direction energy spectrum sequence is generated; Speech enhancement weights are constructed based on the energy ratio of the main direction energy spectrum sequence and the interference direction fused spectrum sequence. The enhanced speech spectrum is obtained by weighting the fusion spectrum in the main direction based on the speech enhancement weights.
2. The directional pickup method of the microphone device as described in claim 1, characterized in that: The specific steps for acquiring audio signals from each microphone and performing frequency domain transformation on the audio signals to construct a set of directional angles are as follows: The audio signals from each microphone are collected to generate a multi-channel time-domain audio signal set; A short-time Fourier transform is performed on a collection of multi-channel time-domain audio signals to obtain a complex spectral tensor. Based on the spatial orientation of the discrete microphone array with a preset angle step, construct a set of directional angles.
3. The directional pickup method for a microphone device as described in claim 2, characterized in that: The specific steps are as follows: Calculating the cross-correlation energy of each direction based on the set of direction angles to form a direction energy vector, and taking the direction with the maximum energy in the direction energy vector as the direction of the main sound source at the current time. Select the complex spectral tensor and the set of directional angles at the current time, and construct a microphone propagation time delay table for each directional angle; Based on the complex spectral tensor and propagation time delay table, calculate the cross-correlation energy of each direction angle at the current time, and generate the cross-correlation energy of the direction angle; Combine all cross-correlation values of direction angles in order of direction angle index to form the direction energy vector at the current time; The direction angle with the largest cross-correlation energy under the direction energy vector is selected as the direction of the main sound source at the current time.
4. The directional pickup method of the microphone device as described in claim 3, characterized in that: The process of generating the final pickup signal through frequency-domain to time-domain transformation involves the following steps: Phase recovery is performed on the enhanced speech spectrum to generate the enhanced speech complex spectrum; Perform an inverse Fourier transform on the complex spectrum of the enhanced speech to generate an enhanced speech frame sequence; The enhanced speech frame sequence is reconstructed by overlapping addition to generate the final pickup signal.
5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the directional pickup method of the microphone device according to any one of claims 1 to 4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the directional pickup method of the microphone device according to any one of claims 1 to 4.
Citation Information
Patent Citations
Sound source direction positioning method and device, voice equipment and voice system
CN111025233A
FPGA-based microphone array directional pickup method
CN115665616A