Flexible microphone array speech enhancement method and device, electronic equipment, and medium
Patent Information
- Application Number
- CN202310349782.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-28
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2043-03-28
AI Technical Summary
现有的麦克风阵列技术存在一些限制,例如,阵列中麦克风的个数受限于设备尺寸以及功耗而无法大幅增加,阵列距离声源的距离太远使录制的音频信噪比较低
[0053]与现有技术相比,本发明的有益效果是:本发明提供了一种柔性麦克风阵列语音增强方法,通过对麦克风阵列接收到的信号的协方差矩阵求谱函数,进行谱峰搜索,找到谱函数的极大值,将极大值对应的角度作为声源方向角;对声源方向的信号进行波束响应优化以增强声源方向语音信号,再经维纳滤波处理后输出增强的语音信号,本发明方法实现复杂环境下的多人语音分离与增强。同时,本发明还设计了一种柔性麦克风阵列语音增强装置,通过按键进行切换,收听每一处声源方向的增强语音,解决了在复杂场景下针对多人混合语音难以区分的问题,便捷地进行实时语音处理与增强,尤其适用于户外嘈杂环境下多人会话场景。
Smart Images

Figure CN116343808B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to flexible circuits, speech separation and enhancement, and more particularly to a method and apparatus for speech enhancement using a flexible microphone array, electronic equipment, and a medium. Background Technology
[0002] Multi-person speech recognition and separation in complex environments is an extremely important and practical task. Many scenarios in life, such as indoor meetings with multiple participants or outdoor team activities, take place in noisy environments. Traditional sensor systems record signals that simultaneously contain background noise and multiple people's speech signals, making it difficult to effectively distinguish the location and content of each person's voice. Therefore, traditional audio transceiver systems cannot achieve the enhancement and transmission of signals from desired sound sources in a specific direction.
[0003] A microphone array is a group of acoustic sensors (called microphones) arranged in a specific order. Through the interaction of minute time differences between each microphone in the array, sound waves arrive at the array. Microphone arrays can achieve better spatial directivity than a single microphone. Microphone arrays are generally used for sound source localization, background noise suppression, signal extraction, and separation. Microphone arrays do not restrict the speaker's movement and can locate sound sources at any position in space, making them an important fundamental device in human-computer interaction and speech direction picking. The speech separation problem originated from the "cocktail party problem," aiming to separate the desired speaker's voice from a noisy environment (interference from other voices or background noise) to make the desired sound clearer. Existing microphone array technology has some limitations. For example, the number of microphones in the array is limited by device size and power consumption, and a large distance between the array and the sound source results in a low signal-to-noise ratio in the recorded audio.
[0004] Wearable devices are portable devices that fit closely to the user, and can be used in fields such as health monitoring and virtual display. Most existing wearable devices are primarily watches, headphones, and glasses, with relatively fixed designs. Flexible wearable devices, however, possess high mechanical flexibility, allowing for better skin contact and enabling a more seamless integration between the user and sensors. Flexible MEMS microphone arrays are small in size, low in power consumption, and can conform well to the human skin surface, making them easy to wear. During actual movement, they achieve real-time integrated acquisition, storage, and processing of voice signals, and provide real-time feedback of the desired voice signal to the designated person.
[0005] This invention proposes a voice enhancement method based on wearable devices to achieve multi-person voice separation and enhancement in complex environments. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a method and apparatus for enhancing voice using a flexible microphone array, as well as an electronic device and a medium.
[0007] According to a first aspect of the present invention, a method for enhancing speech using a flexible microphone array is provided, characterized in that the method comprises:
[0008] Get the voice threshold value;
[0009] Voice presence detection based on voice threshold;
[0010] The covariance matrix of the signal received by the microphone array is used to calculate the spectral function, and a spectral peak search is performed to find the maximum value of the spectral function. The angle corresponding to the maximum value is the sound source direction angle.
[0011] Beam response optimization is performed on the signal in the direction of the sound source to enhance the speech signal in that direction, and then the enhanced speech signal is output after Wiener filtering.
[0012] Furthermore, obtaining the voice threshold includes:
[0013] Voice threshold L τ The expression is as follows:
[0014] L τ = (1-β)L0+βL1
[0015] In the formula, L0 is the general environmental noise energy, L1 is the on-site environmental energy collected within the preset time period, and β is the weighting coefficient.
[0016] Furthermore, speech presence detection based on speech thresholds includes:
[0017] The signal received by the microphone array is divided into several sub-bands;
[0018] Calculate the bivariate Gaussian log-likelihood ratio for each subband;
[0019] The weighted summation of the bivariate Gaussian log-likelihood ratios for all subbands is performed.
[0020] If the sum of the binary Gaussian log-likelihood ratios of all subbands is greater than the speech threshold, then it is determined that there is speech in the signal received by the microphone array.
[0021] Furthermore, the spectral function of the covariance matrix of the signal received by the microphone array is calculated, and a spectral peak search is performed to find the maxima of the spectral function. The angles corresponding to the maxima, i.e., the sound source direction angles, include:
[0022] Construct the covariance matrix of the received signals of the microphone array;
[0023] Eigenvalue decomposition of the covariance matrix yields λ1, λ2…λ M Eigenvalues;
[0024] Let λ1, λ2…λM The eigenvalues are arranged in descending order, i.e., λ1≥…≥λ j >λ j+1 =…=λ M =σ 2 , where σ 2 It is noise power;
[0025] λ j+1 …λ M The eigenvectors corresponding to the eigenvalues span a noisy subspace;
[0026] A spectral function is constructed based on the steering vector and noise subspace of a flexible ring microphone array. A spectral peak search is performed on the spectral function to find all maxima within 0 to 360 degrees. The angle corresponding to the maxima is the angle estimate of the direction of the sound source.
[0027] Furthermore, the spectral function constructed based on the steering vector and noise subspace of the flexible ring microphone array includes:
[0028] The expression for the spectral function is as follows:
[0029]
[0030] In the formula, A(θ)=[α(θ1),α(θ2)…α(θ n [)] is the guide vector for the flexible ring microphone array. θ n V is the angle of incidence of the signal relative to a single microphone in the array, n = 1, 2…N; c is the speed of sound in air; d is the distance between the microphones in the circular array; Noise Represents the noise subspace.
[0031] Furthermore, beam response optimization of the signal in the direction of the sound source to enhance the speech signal in that direction includes:
[0032] Obtain the beam response S(θ) of a single sound source at angle θ. i ) = W H A(θ); i = 1, 2...D, where D is the number of sound sources determined by spectral peak search, and W is the weight vector;
[0033] By minimizing the beam response S(θ) i This is used to enhance the speech signal in the direction of the sound source.
[0034] Furthermore, by minimizing the beam response S(θ) i To enhance the speech signal in the direction of the sound source, the following methods are used:
[0035] Minimize beam response S(θ) i The expression for the minimization problem of ) is:
[0036] min W H A(θ)
[0037] stW H A(θ0)=1
[0038] Where A(θ0) is the steering vector of the enhanced sound source direction;
[0039] Add a penalty term to the minimization problem, updating the minimization problem to:
[0040]
[0041] in, This is called the penalty parameter;
[0042] Introduce auxiliary variable b for W H Find the following minimization problem:
[0043]
[0044] In the formula, q represents the number of iterations;
[0045] For the variable W in equation (1) H Taking the derivative and setting the result to 0, we obtain the following about W. H The expression:
[0046]
[0047] Expanding the Laplace operator in equation (2) using the central difference method, we obtain the iterative calculation W. H Numerical format:
[0048]
[0049] Set the stopping condition for iterative optimization as follows: ε is a real number greater than 0.
[0050] According to a second aspect of the present invention, a flexible microphone array voice enhancement device is provided for implementing the above-described flexible microphone array voice enhancement method. The device includes: a circular flexible microphone ring array; a plurality of acoustic sensors are arranged at equal intervals in the circular flexible microphone ring array, and a power switch button, a sound source enhancement switching button, and headphones are also installed on the circular flexible microphone ring array.
[0051] According to a third aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-described flexible microphone array voice enhancement method.
[0052] According to a fourth aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described flexible microphone array voice enhancement method.
[0053] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a flexible microphone array speech enhancement method. By calculating the spectral function of the covariance matrix of the signal received by the microphone array, spectral peak search is performed to find the maximum value of the spectral function, and the angle corresponding to the maximum value is taken as the sound source direction angle. Beam response optimization is performed on the signal in the sound source direction to enhance the speech signal in the sound source direction. After Wiener filtering, the enhanced speech signal is output. This invention achieves multi-person speech separation and enhancement in complex environments. Simultaneously, this invention also designs a flexible microphone array speech enhancement device, which allows switching via buttons to listen to the enhanced speech from each sound source direction. This solves the problem of difficulty in distinguishing mixed multi-person speech in complex scenarios, enabling convenient real-time speech processing and enhancement, and is particularly suitable for multi-person conversation scenarios in noisy outdoor environments. Attached Figure Description
[0054] Figure 1 This is a schematic diagram of the flexible microphone array device of the present invention;
[0055] Figure 2 This is a flowchart of the speech separation and enhancement process of the present invention;
[0056] Figure 3 This is a schematic diagram of the sum of the likelihood ratios for speech presence detection in this invention.
[0057] Figure 4 This is a schematic diagram of the speech presence detection results of the present invention;
[0058] Figure 5 This is a schematic diagram of the sound source localization results of the present invention;
[0059] Figure 6 This is a waveform diagram of the input signal of a single sensor of the microphone array in an embodiment of the present invention;
[0060] Figure 7 This is a waveform diagram of a speech signal from a sound source direction after signal enhancement by a single sensor in an embodiment of the present invention;
[0061] Figure 8This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0062] The present invention will be further described below with reference to specific embodiments and accompanying drawings. The description of the embodiments below is only for the purpose of helping to understand the present invention. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principle of the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
[0063] To address the challenge of distinguishing mixed speech in noisy multi-person conversation scenarios, this invention provides a flexible microphone array speech enhancement method. For example... Figure 1 As shown, the method is based on a flexible microphone array voice enhancement device. The device includes a circular flexible microphone array with several acoustic sensors evenly spaced within it. A power switch button B1 and an enhanced sound source switching button B2 are also installed on the circular flexible microphone array. In use, pressing the power switch button B1 turns the device on, and pressing and holding B1 for 3 seconds turns it off. Each press of the enhanced sound source switching button B2 switches the desired sound source signal separated by the device, satisfying multi-source separation scenarios. A small earphone is also attached to the circular flexible microphone array for the user to listen to both the array's received signal and the enhanced voice signal. The circular flexible microphone array fits within a ring-shaped area formed by the forehead, above the left ear, the back of the head, and above the right ear, and can be worn on the head like a headband.
[0064] Figure 2 The flowchart shown is a speech enhancement method for a flexible microphone array provided by an embodiment of the present invention. The method specifically includes the following steps:
[0065] Step S1: Obtain the voice threshold value.
[0066] It should be noted that in step S1, the flexible microphone array voice enhancement device is turned on each time it is used, and the sound sources present in the environment are collected for 10 seconds. The flexible microphone array voice enhancement device is automatically calibrated, and the threshold value L for judging whether there is voice in the environment is calculated. τ .
[0067] Among them, L τ It is an automatically updated threshold, and its expression is:
[0068] L τ = (1-β)L0+βL1
[0069] In the formula, L0 is the general environmental noise energy collected a priori, and L1 is the ambient environmental energy collected in the first 10 seconds after the flexible microphone array voice enhancement device is powered on each time; β is the weighting coefficient, which is taken as 0.95 in this example based on experience; the ambient environmental energy L1 is calculated by summing the squares of the signals from the flexible microphone array, and its expression is: L1=∑X 2 (t), where t represents time. The speech signal acquired by the flexible microphone array is denoted as X(t) = (x1(t), ..., x2(t)). n (t)), n = 1, 2, ..., N, where x n (t) is the voice signal collected by a single microphone, and N is the number of microphones in the array. In this embodiment, N = 32.
[0070] Step S2: Perform speech presence detection based on the speech threshold value obtained in step S1.
[0071] In the absence of speech, speech separation is unnecessary. Therefore, the flexible microphone array speech enhancement device must first perform speech presence detection on the acquired signal.
[0072] This invention determines whether a signal received by a flexible microphone array contains speech by calculating the sub-band log-likelihood ratio of the array signal. Specifically, in this example, the signal received by the flexible microphone array is divided into several sub-bands, the sum of the binary Gaussian log-likelihood ratios of all sub-bands is calculated, and speech presence detection is performed based on the speech threshold value obtained in step S1. The signal received by the nth microphone is represented as:
[0073] x n (t)=g n,d (t)*s d (t)+g n,i (t)*s i (t)+v n (t).
[0074] Among them, s d (t) represents the desired speech signal, s i (t) represents the interference acoustic signal, v n (t) represents other environmental noise. g n,d (t) represents the acoustic impulse response function of the nth microphone and the desired speech signal, g n,i (t) represents the acoustic impulse response function of the nth microphone and the interfering speech signal.
[0075] The speech and interference signals are considered as two independent and uncorrelated variables, where d is the characteristic value of the corresponding speech signal. In this example, the received signal frequency band is divided into four sub-bands: 100–400 Hz, 400–1000 Hz, 1000–2000 Hz, and 2000–3500 Hz.
[0076] The bivariate Gaussian log-likelihood ratio for each sub-band is calculated using the following expression:
[0077]
[0078] Where, μμ ds It is the power mean of the corresponding speech signal within a sub-band, μ is σ is the average power of the corresponding interference acoustic signal within a sub-band. ds σ is the power variance of the corresponding speech signal within a sub-band. is It is the power variance of the corresponding interference acoustic signal within a sub-band. μ dE It is the power mean of the noise associated with the speech signal within a subband, μ. iE σ is the power mean of the noise associated with the interfering acoustic signal and the ambient noise within a sub-band. dE σ is the power variance of speech-related noise in a subband. iE It is the power variance of noise and ambient noise associated with the interfering acoustic signal within a sub-band. k represents the number of sub-bands.
[0079] The weighted sum of the binary Gaussian log-likelihood ratios of all sub-bands is taken. If the sum of the binary Gaussian log-likelihood ratios of all sub-bands is greater than the speech threshold, then speech is determined to be present in the signal received by the microphone array. The expression is as follows:
[0080]
[0081] Where, α k These are the weights of each sub-band. If L ≥ L τ If so, it is assumed that the received signal contains speech.
[0082] Figure 3 The results of calculating the sum of the log-likelihood ratios of the four sub-bands over a period of time are shown in this embodiment. Figure 4 The signal in the box represents the detected speech signal range, corresponding to Figure 3 The middle likelihood ratio and the sum exceeding the threshold L τ The area. In this embodiment, L is tested in a semi-anechoic chamber environment. τ =3.1, α k =0.25.
[0083] Step S3: Calculate the spectral function of the covariance matrix of the signal received by the microphone array, perform spectral peak search, find the maximum value of the spectral function, and the angle corresponding to the maximum value is the sound source direction angle.
[0084] The covariance matrix of the received signal from the microphone array is expressed as:
[0085]
[0086] In the formula, H is the conjugate transpose operator, and n is the number of a single microphone in the microphone array.
[0087] Eigenvalue decomposition is performed on the covariance matrix R, and eigenvectors are used to construct a speech signal subspace and a noise subspace. Since the speech signal and noise are independent, these two subspaces are orthogonal. The following results are obtained:
[0088] R = V∑V H
[0089] Where, V = [V Speech V Noise ],∑=diag(λ1,λ2…λ M ), λ M Let be all the eigenvalues of the covariance matrix R, and arrange these eigenvalues in descending order, i.e., λ1≥…≥λ j >λ j+1 =…=λ M =σ 2 , where σ 2 This is the noise power; the first j eigenvalues are related to the speech signal and their values are greater than σ. 2 The eigenvectors corresponding to these j eigenvalues span the speech signal subspace V. Speech From λ j+1 to λ M The eigenvectors corresponding to these eigenvalues span the noise subspace V. Noise .
[0090] Therefore, the source direction angle θ of the speech signal can be obtained by searching for spectral peaks in the spatial spectral function, the expression of which is as follows:
[0091]
[0092] In the formula, A(θ)=[α(θ1),α(θ2)…α(θ n [)] is the guide vector for the flexible ring microphone array. θ n Let n be the angle of incidence of the signal relative to a single microphone in the array, where n = 1, 2, ..., N. Let c be the speed of sound in air, and d be the distance between the microphones in the circular array.
[0093] A spectral peak search is performed on the spectral function P(θ) to find all maxima of P(θ) within 0 to 360 degrees. The θ corresponding to the maxima is the angle estimate of the direction of the sound source.
[0094] In the case of multiple sound sources, the output sound source localization direction vector θ = (θ1, θ2, ... θ) D D represents the number of sound sources determined through spectral peak search. For example... Figure 5 As shown, in this embodiment, three sound sources were found, with directional angles of 40 degrees, 128 degrees, and 220 degrees, respectively.
[0095] Step S4: Beam response optimization is performed on the signal in the direction of the sound source to enhance the speech signal in the direction of the sound source, and then the enhanced speech signal is output after Wiener filtering.
[0096] The beam response S(θ) of a single sound source at angle θ i ) = W H A(θ), i = 1, 2...D, speech enhancement in the direction of the sound source is achieved by minimizing the beam response S(θ). i And obtain a weight vector. This optimization problem can be expressed as:
[0097] min W H A(θ)
[0098] stW H A(θ0)=1
[0099] Where A(θ0) is the steering vector of the enhanced sound source direction.
[0100] Adding a penalty term to this constrained minimization problem transforms the optimization problem into:
[0101]
[0102] in, This is called the penalty parameter.
[0103] Introduce auxiliary variable b for W H Find the following minimization problem:
[0104]
[0105] In the formula, q represents the number of iterations.
[0106] For the variable W in equation (1) H Taking the derivative and setting the result to 0, we obtain the following about W. H The expression:
[0107]
[0108] Expanding the Laplace operator in equation (2) using the central difference method, we obtain the iterative calculation W. H Numerical format:
[0109]
[0110] The stopping condition for iterative calculation is ε is a very small number greater than 0. The output speech signal from the desired sound source is further processed by Wiener filtering to suppress noise and enhance the desired speech enhancement effect in the direction of the sound source, resulting in the final enhanced output speech signal.
[0111] Furthermore, users can switch between enhanced voice signals from different sound source directions using the enhanced sound source switching button B2.
[0112] In this embodiment, the mixed speech signal collected by the fourth microphone of the flexible microphone array is as follows: Figure 6 As shown, the waveform of the enhanced speech signal output from the 40-degree sound source direction after separation is as follows. Figure 7 As shown. Figure 6 and Figure 7 Compared to the waveform of the mixed speech signal, by Figure 7 As can be seen, the enhanced speech signal waveform exhibits a significant suppression effect on the interference signal, with only the amplitude of the desired speech signal being relatively prominent.
[0113] like Figure 8 As shown, this application provides an electronic device including a memory 101 for storing one or more programs and a processor 102. When the one or more programs are executed by the processor 102, they implement the method as described in any of the first aspects above.
[0114] The system also includes a communication interface 103. The memory 101, processor 102, and communication interface 103 are electrically connected directly or indirectly to each other to enable data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines. The memory 101 can be used to store software programs and modules, and the processor 102 executes various functional applications and data processing by executing the software programs and modules stored in the memory 101. The communication interface 103 can be used for signaling or data communication with other node devices.
[0115] The memory 101 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0116] The processor 102 can be an integrated circuit chip with signal processing capabilities. The processor 102 can be a general-purpose processor 102, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0117] In the embodiments provided in this application, it should be understood that the disclosed methods and systems can also be implemented in other ways. The method and system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0118] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0119] On the other hand, embodiments of this application provide a computer-readable storage medium storing a computer program thereon. When executed by processor 102, the computer program implements the methods described in any of the first aspects above. If the functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory 101 (ROM), random access memory 101 (RAM), magnetic disks, or optical disks.
[0120] The above embodiments are only used to illustrate the design concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made based on the principles and design ideas disclosed in the present invention are within the protection scope of the present invention.
Claims
1. A method for enhancing speech using a flexible microphone array, characterized in that, The method includes: Obtain the voice threshold value; the voice threshold value The expression is: In the formula, This is general environmental noise energy. It is the on-site environmental energy collected within a pre-set time period; These are the weighting coefficients; Voice presence detection based on voice threshold; The covariance matrix of the signal received by the microphone array is used to calculate the spectral function, and a spectral peak search is performed to find the maximum value of the spectral function. The angle corresponding to the maximum value is the sound source direction angle. Beam response optimization is performed on the signal from the direction of the sound source to enhance the speech signal from that direction, and then the enhanced speech signal is output after Wiener filtering; this includes: obtaining the angle. Single sound source directional beam response ; , This represents the number of sound sources identified through spectral peak search. The weight vector is used to minimize the beam response. To enhance the speech signal in the direction of the sound source; Specifically, the covariance matrix of the signal received by the microphone array is used to calculate the spectral function, and a spectral peak search is performed to find the maxima of the spectral function. The angles corresponding to the maxima, i.e., the sound source direction angles, include: Construct the covariance matrix of the received signals of the microphone array; Perform eigenvalue decomposition on the covariance matrix to obtain Eigenvalues; Will The eigenvalues are arranged in descending order, i.e. ,in It is noise power; Will The eigenvectors corresponding to the eigenvalues span a noisy subspace; A spectral function is constructed based on the steering vector and noise subspace of a flexible ring microphone array. A spectral peak search is performed on the spectral function to find all maxima within 0 to 360 degrees. The angle corresponding to the maxima is the angle estimate of the direction of the sound source.
2. The flexible microphone array speech enhancement method according to claim 1, characterized in that, Speech presence detection based on speech thresholds includes: The signal received by the microphone array is divided into several sub-bands; Calculate the bivariate Gaussian log-likelihood ratio for each subband; The weighted summation of the bivariate Gaussian log-likelihood ratios for all subbands is performed. If the sum of the binary Gaussian log-likelihood ratios of all subbands is greater than the speech threshold, then it is determined that there is speech in the signal received by the microphone array.
3. The flexible microphone array speech enhancement method according to claim 1, characterized in that, The spectral function is constructed based on the steering vector and noise subspace of the flexible ring microphone array, including: The expression for the spectral function is as follows: ; In the formula, This is the guide vector for the flexible ring microphone array. , The angle of incidence of the signal relative to a single microphone in the array. The speed of sound in air. The spacing between microphones in a circular array; Represents the noise subspace.
4. The flexible microphone array speech enhancement method according to claim 1, characterized in that, By minimizing beam response To enhance the speech signal in the direction of the sound source, the following methods are included: Minimize beam response The expression for the minimization problem is: ; ; in, It is the guiding vector that enhances the direction of the sound source; Add a penalty term to the minimization problem, updating the minimization problem to: ; in, This is called the penalty parameter; Introducing auxiliary variables ,right Find the following minimization problem: ; In the formula, This represents the number of iterations. For the variables in equation (1) Taking the derivative and setting the result to 0, we obtain the following about The expression: ; Expanding the Laplace operator in equation (2) using the central difference method, we obtain the iterative calculation. Numerical format: ; Set the stopping condition for iterative optimization as follows: , It is a real number greater than 0.
5. A flexible microphone array speech enhancement device, used to implement the flexible microphone array speech enhancement method according to any one of claims 1-4, characterized in that, The device includes: a circular flexible microphone array; several acoustic sensors are arranged at equal intervals in the circular flexible microphone array, and a power switch button, a sound source enhancement switching button, and headphones are also installed on the circular flexible microphone array.
6. An electronic device comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the flexible microphone array voice enhancement method according to any one of claims 1-4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the flexible microphone array speech enhancement method as described in any one of claims 1-4.
Citation Information
Patent Citations
Interactive intelligent voice home control device and control method based on open source hardware
CN110265012A
Microphone array speech enhancement method and device
CN110517701A