Speech presence probability calculation method, system, speech enhancement method and earphone

The microphone signal is iteratively optimized through maximum likelihood estimation and expectation maximization algorithms, which solves the problem of inaccurate speech probability estimation on devices with a small number of microphones and small spacing, improves the estimation accuracy of the noise covariance matrix, and enhances the speech enhancement effect of the MVDR algorithm.

CN115966215BActive Publication Date: 2025-10-03SHENZHEN SHOKZ CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111182432.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-11
Publication Date
2025-10-03
Estimated Expiration
2041-10-11

AI Technical Summary

Technical Problem

In the existing technology, the speech enhancement method based on the MVDR algorithm is less effective on devices with a small number of microphones and a small distance between them, such as headphones. This is mainly due to the low accuracy of the estimation of the probability of speech existence, which leads to low accuracy in the estimation of the noise covariance matrix.

Method used

The maximum likelihood estimation and expectation maximization algorithms are used to iteratively optimize the microphone signal. The speech presence model is determined by comparing the entropy of the speech presence probability and the speech absence probability, thereby improving the accuracy of the speech presence probability and calculating a more accurate noise covariance matrix.

Benefits of technology

The calculation accuracy of the speech existence probability and the estimation accuracy of the noise covariance matrix are improved, which enhances the speech enhancement effect of the MVDR algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115966215B_ABST
    Figure CN115966215B_ABST
Patent Text Reader

Abstract

The speech presence probability calculation method, system, speech enhancement method and headphones provided in this specification correct the speech presence probability and speech non-existence probability in the iterative process by comparing the entropy of the speech presence probability and the entropy of the speech non-existence probability to obtain faster convergence speed and better convergence results, thereby making the speech presence probability and noise covariance matrix estimation more accurate, thereby improving the speech enhancement effect of MVDR.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of speech signal processing technology, and in particular to a method and system for calculating speech presence probability, a speech enhancement method, and a headset. Background Art

[0002] In speech enhancement technologies based on beamforming algorithms, especially in adaptive beamforming algorithms based on minimum variance distortionless response (MVDR), solving for the noise covariance matrix—a parameter that describes the relationship between the statistical characteristics of noise between different microphones—is crucial. The main method in the prior art is to calculate the noise covariance matrix based on the probability of speech presence. For example, the voice activity detection (VAD) method is used to estimate the probability of speech presence and then calculate the noise covariance matrix. However, the accuracy of the speech presence probability estimation in the prior art is insufficient, resulting in low precision in the noise covariance matrix estimation, which in turn leads to poor speech enhancement effect of the MVDR algorithm. This is especially true when the number of microphones is small, such as less than 5, and the effect drops sharply. Therefore, the MVDR algorithm in the prior art is mostly used in microphone array devices with a large number of microphones and large spacing, such as mobile phones and smart speakers. However, the speech enhancement effect is poor for devices with a small number of microphones and small spacing, such as headphones.

[0003] Therefore, it is necessary to provide a method, system, speech enhancement method and headset with higher accuracy for calculating the probability of speech presence. Summary of the Invention

[0004] This specification provides a more accurate method and system for calculating the probability of speech presence, a speech enhancement method, and a headset.

[0005] In a first aspect, the present specification provides a method for calculating the probability of speech presence, for M microphones distributed in a preset array shape, where M is an integer greater than 1, comprising: obtaining microphone signals output by the M microphones, the microphone signals satisfying a first model or a second model of a Gaussian distribution, one of the first model and the second model being a speech presence model, and the other being a speech absence model; iteratively optimizing the first model and the second model based on maximum likelihood estimation and expectation maximization algorithms until convergence, and during the iterative process, determining whether the speech presence model is the first model or the second model based on the entropy of a first probability when the microphone signal is the first model and the entropy of a second probability when the microphone signal is the second model, the first probability being complementary to the second probability; and when the maximum likelihood estimation and expectation maximization algorithms converge, taking the probability that the microphone signal is the speech presence model as the speech presence probability of the microphone signal and outputting it.

[0006] In some embodiments, the first variance of the Gaussian distribution corresponding to the first model includes the product of the first parameter and the first spatial covariance matrix; and the second variance of the Gaussian distribution corresponding to the second model includes the product of the second parameter and the second spatial covariance matrix; the iterative optimization of the first model and the second model based on the maximum likelihood estimation and expectation maximization algorithm respectively includes: constructing an objective function based on the maximum likelihood estimation and expectation maximization algorithm; determining optimization parameters, the optimization parameters including the first spatial covariance matrix and the second spatial covariance matrix; determining the initial values ​​of the optimization parameters; based on the objective function and the initial values ​​of the optimization parameters, iterating the optimization parameters multiple times until the objective function converges, including: determining whether the probability of speech existence is the first model or the second model based on the entropy of the first probability and the entropy of the second probability in the multiple iterations; and outputting the converged value of the optimization parameter and its corresponding first probability and second probability.

[0007] In some embodiments, determining whether the probability of speech existence is the first model or the second model based on the entropy of the first probability and the entropy of the second probability in the multiple iterations includes: in any one iteration of the multiple iterations, calculating the entropy of the first probability and the entropy of the second probability, and determining whether the probability of speech existence is the first model or the second model, including: determining that the entropy of the first probability is greater than the entropy of the second probability, determining that the speech existence model is the second model; or determining that the entropy of the first probability is less than the entropy of the second probability, determining that the speech existence model is the first model.

[0008] In some embodiments, determining whether the probability of speech existence is the first model or the second model based on the entropy of the first probability and the entropy of the second probability in the multiple iterations includes: in the first iteration of the multiple iterations, calculating the entropy of the first probability and the entropy of the second probability, and determining whether the probability of speech existence is the first model or the second model, including: determining that the entropy of the first probability is greater than the entropy of the second probability, determining that the speech existence model is the second model; or determining that the entropy of the first probability is less than the entropy of the second probability, determining that the speech existence model is the first model.

[0009] In some embodiments, the multiple iterations of the optimization parameters also include: in each iteration of the multiple iterations: modifying the first probability and the second probability based on the entropy of the first probability and the entropy of the second probability, including: determining that the first model is the speech existence model, and the entropy of the first probability is greater than the entropy of the second probability, and exchanging the value corresponding to the first probability with the value corresponding to the second probability; or determining that the second model is the speech existence model, and the entropy of the second probability is greater than the entropy of the first probability, and exchanging the value corresponding to the first probability with the value corresponding to the second probability; and updating the optimization parameters based on the modified first probability and second probability.

[0010] In some embodiments, the multiple iterations of the optimization parameters also include: in each iteration of the multiple iterations: performing a reversible correction on the optimization parameters, including: determining that the optimization parameters are irreversible, and correcting the optimization parameters through a deviation matrix, wherein the deviation matrix includes a unit matrix, a random matrix that obeys a normal distribution or a uniform distribution.

[0011] In a second aspect, the present specification also provides a system for calculating the probability of speech presence, comprising at least one storage medium and at least one processor, wherein the at least one storage medium stores at least one instruction set for calculating the probability of speech presence; the at least one processor is communicatively connected to the at least one storage medium, wherein when the system for calculating the probability of speech presence is running, the at least one processor reads the at least one instruction set and implements the method for calculating the probability of speech presence described in the first aspect of this specification.

[0012] In a third aspect, the present specification also provides a speech enhancement method for M microphones distributed in a preset array shape, where M is an integer greater than 1, comprising: obtaining microphone signals output by the M microphones; determining the speech existence probability of the microphone signal based on the speech existence probability calculation method described in any one of claims 1 to 7; determining the noise covariance matrix of the microphone signal based on the speech existence probability; determining the filter coefficient corresponding to the microphone signal based on the MVDR method and the noise space covariance matrix; and merging the microphone signals based on the filter coefficient to output a target audio signal.

[0013] In a fourth aspect, this specification also provides a headset, comprising a microphone array and a computing device, wherein the microphone array comprises M microphones distributed in a preset array shape, where M is an integer greater than 1; when the computing device is running, it communicates with the microphone array and executes the speech enhancement method described in the third aspect of this specification.

[0014] In some embodiments, the M microphones are linearly distributed, and M is not greater than 5, and the spacing between adjacent microphones in the M microphones is between 20 mm and 40 mm. The headset also includes a first shell and a second shell, and the microphone array is installed on the first shell. The first shell includes a first interface and a contact, and the first interface is provided with a first magnetic device. The contact is provided at the first interface and is communicatively connected to the microphone array; the computing device is installed on the second shell, and the second shell includes a second interface and a guide rail. The second interface is provided with a second magnetic device, and the guide rail is provided at the second interface and is communicatively connected to the computing device. The adsorption force between the first magnetic device and the second magnetic device makes the first shell and the second shell detachable. When the first shell and the second shell are connected, the contact contacts the guide rail, so that the microphone array is communicatively connected to the computing device.

[0015] As can be seen from the above technical solutions, the method, system, speech enhancement method and headset provided in this specification are used for a microphone array composed of multiple microphones. Each microphone in the microphone array can collect audio from multiple sound sources in the space and output corresponding microphone signals. The audio signal of each sound source satisfies the Gaussian distribution. The multiple microphone signals output by the multiple microphone arrays satisfy the joint Gaussian distribution. In order to obtain the speech presence probability in the multiple microphone signals, the method, system, speech enhancement method and headset for calculating the speech presence probability can respectively obtain the speech presence model when speech is present and the speech absence model when speech is absent in the multiple microphone signals, and optimize through multiple iterations based on the maximum likelihood estimation and expectation maximization algorithm, and in the iterative process, according to the entropy of the speech presence probability and the entropy of the speech absence probability, correct the speech presence probability and the speech absence probability, thereby calculating and determining the model parameters of the speech presence model and the model parameters when speech is absent, and when the maximum likelihood estimation and expectation maximization algorithm converge, obtain the speech presence probability corresponding to the speech presence model. The speech presence probability calculation method, system, speech enhancement method and headphones correct the speech presence probability and speech non-existence probability in the iterative process by comparing the entropy of the speech presence probability and the entropy of the speech non-existence probability, so as to obtain a faster convergence speed and better convergence results, thereby making the speech presence probability and the noise covariance matrix estimation more accurate, thereby improving the speech enhancement effect of MVDR.

[0016] The speech presence probability calculation method, system, speech enhancement method, and other features of the headset provided in this specification are partially outlined in the following description. The following figures and examples will be readily apparent to those skilled in the art based on the description. The inventive aspects of the speech presence probability calculation method, system, speech enhancement method, and headset provided in this specification can be fully explained through practice or use of the methods, devices, and combinations described in the following detailed examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 FIG2 shows a hardware schematic diagram of a system for calculating the probability of speech presence provided in accordance with an embodiment of the present specification;

[0019] Figure 2AA schematic diagram of an exploded structure of an electronic device provided according to an embodiment of this specification is shown;

[0020] Figure 2B shows a front view of a first housing provided according to an embodiment of this specification;

[0021] Figure 2C shows a top view of a first housing provided according to an embodiment of this specification;

[0022] Figure 2D shows a front view of a second housing provided according to an embodiment of this specification;

[0023] Figure 2E shows a bottom view of a second housing provided according to an embodiment of this specification;

[0024] Figure 3 A flowchart of a method for calculating the probability of speech presence is shown according to an embodiment of this specification;

[0025] Figure 4 A flowchart of an iterative optimization method according to an embodiment of the present invention is shown;

[0026] Figure 5 A flowchart of multiple iterations provided according to an embodiment of the present specification is shown;

[0027] Figure 6 A flowchart showing another multiple iterations provided according to an embodiment of the present specification; and

[0028] Figure 7 The flowchart of a speech enhancement method provided according to an embodiment of this specification is shown. DETAILED DESCRIPTION

[0029] The following description provides specific application scenarios and requirements for this specification, with the goal of enabling those skilled in the art to make and use the contents of this specification. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but is intended to be accorded the broadest scope consistent with the claims.

[0030] The terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. For example, as used herein, the singular forms "a," "an," and "the" may also include the plural forms unless the context clearly indicates otherwise. When used in this specification, the terms "comprise," "include," and / or "contain" are intended to refer to the presence of the associated integers, steps, operations, elements, and / or components, but do not preclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups or the addition of other features, integers, steps, operations, elements, components, and / or groups in the system / method.

[0031] These and other features of this specification, as well as the operation and function of the associated elements of the structure, and the economical assembly and manufacture of the components, can be significantly improved with consideration of the following description. Reference is made to the accompanying drawings, all of which form a part of this specification. However, it should be expressly understood that the drawings are for illustration and description purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0032] The flowcharts used in this specification illustrate operations implemented by systems according to some embodiments of the present specification. It should be clearly understood that the operations of the flowcharts may not be implemented in sequence. Rather, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.

[0033] For the convenience of description, the following terms will be explained in the specification:

[0034] Minimum Variance Distortionless Response (MVDR): An adaptive beamforming algorithm based on the maximum signal-to-interference-plus-noise ratio (SINR) criterion, the MVDR algorithm adaptively minimizes the array output power in the desired direction while maximizing the SINR. Its goal is to minimize the variance of the recorded signal. If the noise signal and the desired signal are uncorrelated, the variance of the recorded signal is the sum of the variances of the desired signal and the noise signal. Therefore, the MVDR solution seeks to minimize this sum, thereby mitigating the impact of the noise signal. The principle is to select appropriate filter coefficients to minimize the average power of the array output, under the constraint that the desired signal is distortion-free.

[0035] Speech presence probability: the probability that the target speech signal exists in the current audio signal.

[0036] Gaussian distribution: Normal distribution, also known as "normal distribution", also known as Gaussian distribution, the normal curve is bell-shaped, low at both ends, high in the middle, and symmetrical on both sides. Because of its bell-shaped curve, people often call it the bell curve. If the random variable X obeys a mathematical expectation of μ and variance of σ 2 Normal distribution, denoted as N(μ, σ 2 The probability density function is the normal distribution. The expected value μ determines its location, and the standard deviation σ determines the amplitude of the distribution. When μ = 0 and σ = 1, the normal distribution is the standard normal distribution.

[0037] Figure 1 FIG2 shows a hardware schematic diagram of a system for calculating the probability of speech presence provided according to an embodiment of the present disclosure. The system for calculating the probability of speech presence can be applied to an electronic device 200 .

[0038] In some embodiments, the electronic device 200 may be a wireless headset, a wired headset, a smart wearable device, such as smart glasses, a smart helmet, or a smart watch, etc., which have audio processing capabilities. The electronic device 200 may also be a mobile device, a tablet computer, a laptop computer, a built-in device in a motor vehicle, or the like, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, or the like, or any combination thereof. For example, the smart mobile device may include a mobile phone, a personal digital assistant, a gaming device, a navigation device, an ultra-mobile personal computer (UMPC), etc., or any combination thereof. In some embodiments, the smart home device may include a smart TV, a desktop computer, etc., or any combination thereof. In some embodiments, the built-in device in a motor vehicle may include an onboard computer, an onboard TV, etc.

[0039] In this specification, we take the electronic device 200 as a headset as an example for description. The headset can be a wireless headset or a wired headset. Figure 1 As shown, the electronic device 200 may include a microphone array 220 and a computing device 240 .

[0040] Microphone array 220 may be the audio capture device of electronic device 200. Microphone array 220 may be configured to capture local audio and output microphone signals, that is, electronic signals carrying audio information. Microphone array 220 may include M microphones 222 distributed in a predetermined array shape. M is an integer greater than 1. The M microphones 222 may be evenly or unevenly distributed. The M microphones 222 may output microphone signals. The M microphones 222 may output M microphone signals. Each microphone 222 corresponds to a microphone signal. The M microphone signals are collectively referred to as the microphone signals. In some embodiments, the M microphones 222 may be distributed linearly. In some embodiments, the M microphones 222 may also be distributed in arrays of other shapes, such as a circular array, a rectangular array, etc. For ease of description, the following description will use the linear distribution of M microphones 222 as an example. In some embodiments, M may be any integer greater than 1, such as 2, 3, 4, 5, or even more. In some embodiments, due to space limitations, M may be an integer greater than 1 but not greater than 5, such as in products such as headphones. When the electronic device 200 is a headset, the distance between adjacent microphones 222 in the M microphones 222 may be between 20 mm and 40 mm. In some embodiments, the distance between adjacent microphones 222 may be smaller, such as between 10 mm and 20 mm.

[0041] In some embodiments, microphone 222 may be a bone conduction microphone that directly collects human body vibration signals. A bone conduction microphone may include a vibration sensor, such as an optical vibration sensor, an accelerometer, or the like. The vibration sensor may collect mechanical vibration signals (e.g., signals generated by the vibration of the skin or bones when a user speaks) and convert the mechanical vibration signals into electrical signals. The mechanical vibration signals referred to herein primarily refer to vibrations transmitted through solids. The bone conduction microphone contacts the user's skin or bones through the vibration sensor or a vibrating component connected to the vibration sensor, thereby collecting vibration signals generated by the bones or skin when the user makes a sound and converting the vibration signals into electrical signals. In some embodiments, the vibration sensor may be a device that is sensitive to mechanical vibrations but insensitive to air vibrations (i.e., the vibration sensor's response to mechanical vibrations exceeds its response to air vibrations). Because the bone conduction microphone can directly pick up vibration signals at the site of vocalization, it can reduce the impact of ambient noise.

[0042] In some embodiments, the microphone 222 may also be an air conduction microphone that directly collects air vibration signals. The air conduction microphone collects air vibration signals caused by the user making sounds and converts the air vibration signals into electrical signals.

[0043] In some embodiments, the M microphones 220 may be M bone conduction microphones. In some embodiments, the M microphones 220 may also be M air conduction microphones. In some embodiments, the M microphones 220 may include both bone conduction microphones and air conduction microphones. Of course, the microphones 222 may also be other types of microphones, such as optical microphones, microphones that receive electromyographic signals, and so on.

[0044] The computing device 240 can be communicatively connected to the microphone array 220. A communicatively connected connection refers to any form of connection capable of directly or indirectly receiving information. In some embodiments, the computing device 240 and the microphone array 220 can communicate data via wireless communication. In some embodiments, the computing device 240 and the microphone array 220 can also communicate data via a direct wired connection. In some embodiments, the computing device 240 and the microphone array 220 can also be indirectly connected to the microphone array 220 via a direct wired connection to other circuits to achieve data transmission. This description uses the example of a direct wired connection between the computing device 240 and the microphone array 220.

[0045] Computing device 240 can be a hardware device with data processing capabilities. In some embodiments, the speech presence probability calculation system can include computing device 240. In some embodiments, the speech presence probability calculation system can be applied to computing device 240. That is, the speech presence probability calculation system can run on computing device 240. The speech presence probability calculation system can include a hardware device with data processing capabilities and the necessary programs to operate the hardware device. Of course, the speech presence probability calculation system can also be simply a hardware device with data processing capabilities, or simply a program running on the hardware device.

[0046] The voice presence probability calculation system may store data or instructions for executing the voice presence probability calculation method described in this specification, and may execute the data and / or instructions. When the voice presence probability calculation system is running on the computing device 240, the voice presence probability calculation system may obtain the microphone signal from the microphone array 220 based on the communication connection, and execute the data or instructions of the voice presence probability calculation method described in this specification to calculate the voice presence probability in the microphone signal. The voice presence probability calculation method is introduced in other parts of this specification. For example, in Figures 3 to 6 The method for calculating the probability of speech existence is introduced in the description.

[0047] like Figure 1As shown, the computing device 240 may include at least one storage medium 243 and at least one processor 242. In some embodiments, the electronic device 200 may further include a communication port 245 and an internal communication bus 241.

[0048] The internal communication bus 241 can connect various system components, including the storage medium 243 , the processor 242 , and the communication port 245 .

[0049] The communication port 245 can be used for data communication between the computing device 240 and the outside world. For example, the computing device 240 can obtain the microphone signal from the microphone array 220 through the communication port 245 .

[0050] At least one storage medium 243 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk, a read-only storage medium (ROM), or a random access storage medium (RAM). When the speech presence probability calculation system can be run on the computing device 240, the storage medium 243 may also include at least one instruction set stored in the data storage device for performing speech presence probability calculation on the microphone signal. The instruction is a computer program code, and the computer program code may include a program, routine, object, component, data structure, process, module, etc. for executing the speech presence probability calculation method provided in this specification.

[0051] At least one processor 242 can be communicatively connected to at least one storage medium 243 via an internal communication bus 241. The communication connection refers to any form of connection capable of directly or indirectly receiving information. The at least one processor 242 is configured to execute the at least one instruction set described above. When the speech presence probability calculation system is executed on the computing device 240, the at least one processor 242 reads the at least one instruction set and, in accordance with the instructions of the at least one instruction set, executes the speech presence probability calculation method provided herein. The processor 242 can execute all steps included in the speech presence probability calculation method. The processor 242 can be in the form of one or more processors. In some embodiments, the processor 242 can include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physical processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field-programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, or any combination thereof. For illustrative purposes only, only one processor 242 is described in this specification as being part of the computing device 240. However, it should be noted that the computing device 240 described herein may also include multiple processors 242. Therefore, the operations and / or method steps disclosed herein may be performed by a single processor as described herein, or may be performed jointly by multiple processors. For example, if the processor 242 of the computing device 240 is described herein as performing steps A and B, it should be understood that steps A and B may also be performed jointly or separately by two different processors 242 (e.g., a first processor performing step A and a second processor performing step B, or a first and second processor performing steps A and B together).

[0052] Figure 2A FIG1 shows an exploded structural diagram of an electronic device 200 provided according to an embodiment of this specification. Figure 2A As shown, the electronic device 200 may include a microphone array 220 , a computing device 240 , a first housing 260 , and a second housing 280 .

[0053] The first shell 260 can be a mounting base for the microphone array 220. The microphone array 220 can be installed inside the first shell 260. The shape of the first shell 260 can be adaptively designed according to the distribution shape of the microphone array 220, and this specification does not impose too many restrictions on this. The second shell 280 can be a mounting base for the computing device 240. The computing device 240 can be installed inside the second shell 280. The shape of the second shell 280 can be adaptively designed according to the shape of the computing device 240, and this specification does not impose too many restrictions on this. When the electronic device 200 is a headset, the second shell 280 can be connected to the wearing part. The second shell 280 can be connected to the first shell 260. As mentioned above, the microphone array 220 can be electrically connected to the computing device 240. Specifically, the microphone array 220 can be electrically connected to the computing device 240 through the connection between the first shell 260 and the second shell 280.

[0054] In some embodiments, the first housing 260 can be fixedly connected to the second housing 280, for example, by integral molding, welding, riveting, adhesive bonding, etc. In some embodiments, the first housing 260 can be detachably connected to the second housing 280. The computing device 240 can be communicatively connected to different microphone arrays 220. Specifically, different microphone arrays 220 can have different numbers of microphones 222, different array shapes, different spacing between microphones 222, different installation angles of the microphone arrays 220 within the first housing 260, different installation positions of the microphone arrays 220 within the first housing 260, etc. Users can replace corresponding microphone arrays 220 according to different application scenarios to make the electronic device 200 suitable for a wider range of scenarios. For example, when the user is close to the electronic device 200 in an application scenario, the user can replace the microphone array 220 with a closer spacing. For another example, when the user is close to the electronic device 200 in an application scenario, the user can replace the microphone array 220 with a larger spacing and a larger number of microphones, etc.

[0055] The detachable connection can be any form of physical connection, such as a threaded connection, a snap connection, a magnetic connection, etc. In some embodiments, the first housing 260 and the second housing 280 can be connected magnetically. That is, the first housing 260 and the second housing 280 are detachably connected by the adsorption force of a magnetic device.

[0056] Figure 2B 1 shows a front view of a first housing 260 provided according to an embodiment of this specification; Figure 2C FIG. 2 shows a top view of a first housing 260 provided according to an embodiment of the present specification. Figure 2B and Figure 2CAs shown, the first housing 260 may include a first interface 262. In some embodiments, the first housing 260 may further include a contact 266. In some embodiments, the first housing 260 may further include an angle sensor ( Figure 2B and Figure 2C not shown).

[0057] The first interface 262 may be a mounting interface between the first housing 260 and the second housing 280. In some embodiments, the first interface 262 may be circular. The first interface 262 may be rotatably connected to the second housing 280. When the first housing 260 is mounted on the second housing 280, the first housing 260 may rotate relative to the second housing 280, adjusting the angle of the first housing 260 relative to the second housing 280 and thereby adjusting the angle of the microphone array 220.

[0058] A first magnetic device 263 may be provided on the first interface 262. The first magnetic device 263 may be provided at a position of the first interface 262 close to the second shell 280. The first magnetic device 263 may generate a magnetic adsorption force, thereby achieving a detachable connection with the second shell 280. When the first shell 260 is close to the second shell 260, the first shell 260 and the second shell 280 are quickly connected by the adsorption force. In some embodiments, after the first shell 260 and the second shell 280 are connected, the first shell 260 can also be rotated relative to the second shell 280 to adjust the angle of the microphone array 220. Under the action of the adsorption force, when the first shell 260 rotates relative to the second shell 280, the connection between the first shell 260 and the second shell 280 can still be maintained.

[0059] In some embodiments, the first interface 262 may also be provided with a first positioning device ( Figure 2B and Figure 2C The first positioning device may be a positioning step protruding outward or a positioning hole extending inward. The first positioning device may cooperate with the second housing 280 to achieve quick installation of the first housing 260 and the second housing 280.

[0060] like Figure 2B and Figure 2CAs shown, in some embodiments, the first shell 260 may further include contacts 266. The contacts 266 may be installed at the first interface 262. The contacts 266 may protrude outward from the first interface 262. The contacts 266 may be elastically connected to the first interface 262. The contacts 266 may be communicatively connected to the M microphones 222 in the microphone array 220. The contacts 266 may be made of elastic metal to enable data transmission. When the first shell 260 is connected to the second shell 280, the microphone array 220 may be communicatively connected to the computing device 240 through the contacts 266. In some embodiments, the contacts 266 may be distributed in a circular shape. After the first shell 260 is connected to the second shell 280, when the first shell 260 rotates relative to the second shell 280, the contacts 266 may also rotate relative to the second shell 280 and maintain a communicatively connected connection with the computing device 240.

[0061] In some embodiments, the first housing 260 may also be provided with an angle sensor ( Figure 2B and Figure 2C (not shown). The angle sensor can be in communication with the contact 266, thereby achieving communication with the computing device 240. The angle sensor can collect angle data of the first housing 260 to determine the angle of the microphone array 220, providing reference data for subsequent calculation of the probability of speech presence.

[0062] Figure 2D 1 shows a front view of a second housing 280 provided according to an embodiment of this specification; Figure 2E FIG2 shows a bottom view of a second housing 280 provided according to an embodiment of the present specification. Figure 2D and Figure 2E As shown, the second housing 280 may include a second interface 282. In some embodiments, the second housing 280 may further include a guide rail 286.

[0063] The second interface 282 may be a mounting interface between the second housing 280 and the first housing 260. In some embodiments, the second interface 282 may be circular. The second interface 282 may be rotatably connected to the first interface 262 of the first housing 260. When the first housing 260 is mounted on the second housing 280, the first housing 260 may rotate relative to the second housing 280, adjusting the angle of the first housing 260 relative to the second housing 280, thereby adjusting the angle of the microphone array 220.

[0064] The second interface 282 may be provided with a second magnetic device 283. The second magnetic device 283 may be provided near the first housing 260 near the second interface 282. The second magnetic device 283 can generate a magnetic attraction force, thereby achieving a removable connection with the first interface 262. The second magnetic device 283 can be used in conjunction with the first magnetic device 263. When the first housing 260 is close to the second housing 260, the attraction force between the second magnetic device 283 and the first magnetic device 263 allows the first housing 260 to be quickly attached to the second housing 280. When the first housing 260 is attached to the second housing 260, the second magnetic device 283 and the first magnetic device 263 are positioned opposite each other. In some embodiments, after the first housing 260 and the second housing 280 are connected, the first housing 260 can be rotated relative to the second housing 280 to adjust the angle of the microphone array 220. Due to the attraction force, the connection between the first housing 260 and the second housing 280 can be maintained even when the first housing 260 rotates relative to the second housing 280.

[0065] In some embodiments, the second interface 282 may also be provided with a second positioning device ( Figure 2D and Figure 2E (not shown). The second positioning device can be an outwardly protruding positioning step or an inwardly extending positioning hole. The second positioning device can cooperate with the first positioning device of the first housing 260 to achieve quick installation of the first housing 260 and the second housing 280. When the first positioning device is the positioning step, the second positioning device can be the positioning hole. When the first positioning device is the positioning hole, the second positioning device can be the positioning step.

[0066] like Figure 2D and Figure 2EAs shown, in some embodiments, the second shell 280 may further include a guide rail 286. The guide rail 286 may be installed at the second interface 282. The guide rail 286 may be communicatively connected to the computing device 240. The guide rail 286 may be made of a metal material to enable data transmission. When the first shell 260 is connected to the second shell 280, the contact 266 may contact the guide rail 286 to form a communication connection, thereby enabling a communication connection between the microphone array 220 and the computing device 240 to enable data transmission. As previously described, the contact 266 may be elastically connected to the first interface 262. Therefore, after the first shell 260 is connected to the second shell 280, the contact 266 may be fully contacted with the guide rail 286 under the action of the elastic force of the elastic connection to achieve a reliable communication connection. In some embodiments, the guide rail 286 may be distributed in a circular shape. After the first housing 260 is connected to the second housing 280 , when the first housing 260 rotates relative to the second housing 280 , the contact 266 can also rotate relative to the guide rail 286 and maintain a communication connection with the guide rail 286 .

[0067] Figure 3 FIG2 shows a flow chart of a method P100 for calculating the probability of speech presence according to an embodiment of the present specification. The method P100 can calculate the probability of speech presence of the microphone signal. Specifically, the processor 242 can execute the method P100. Figure 3 As shown, the method P100 may include:

[0068] S120 : Acquire microphone signals output by the M microphones 222 .

[0069] As previously described, each microphone 222 can output a corresponding microphone signal. M microphones 222 correspond to M microphone signals. When calculating the probability of speech presence, method P100 can perform the calculation based on all of the M microphone signals, or based on a portion of the microphone signals. Therefore, the microphone signal may include M microphone signals corresponding to M microphones 222 or a portion of the microphone signals. The following description of this specification will use the example of the microphone signal including M microphone signals corresponding to M microphones 222 as an example.

[0070] As mentioned above, the microphone 222 can collect noise in the surrounding environment and can also collect the target voice of the target user. Assume that there are N signal sources around the microphone 222, namely s1(t), ..., s N (t). For the convenience of description, we define N signal sources as s v (t).s v (t) is composed of N signal sources s1(t), ..., s N(t) is a signal source vector. Where v = n or s + n. When v = n, it means N signal sources s v (t) is all noise signal. When v=s+n, it means N signal sources s v (t) consists of a noise signal and a target speech signal. N signal sources s v The sound field mode of (t) is the far field mode. N signal sources s v (t) can be considered as a plane wave. For ease of description, we mark the microphone signal at time t as x(t). The microphone signal x(t) can be a signal vector composed of M microphone signals. In this case, the microphone signal x(, t) can be expressed as the following formula:

[0071]

[0072] Among them, a v (θ) is N signal sources s v (t) is the steering vector. θ1, ..., θ N There are N signal sources s1(t), ..., s N (t) The incident angle between the microphone 222. v (θ) can be the same as θ1, ..., θ N and the distances d1, ..., d between adjacent microphones 222 M-1 The computing device 240 pre-stores the relative position relationship of the M microphones 222, such as relative distance or relative coordinates. That is, the computing device 240 pre-stores d1, ..., d M-1 .

[0073] The microphone signal x(t) is a time domain signal. In some embodiments, in step S120, the computing device 240 may further perform spectrum analysis on the microphone signal x(t). Specifically, the computing device 240 may perform Fourier transform based on the time domain signal x(t) of the microphone signal to obtain the frequency domain signal x(t) of the microphone signal. f,t In the following description, the microphone signal x in the frequency domain will be used. f,t Provide a description.

[0074] At this time, the microphone signal x f,t It can be expressed as the following formula:

[0075]

[0076] in, is the steering vector in the frequency domain. is the complex amplitude of the signal corresponding to the N signal sources in the frequency domain. In some embodiments, the N signal sources It can satisfy the Gaussian distribution. It can be expressed as the following formula:

[0077]

[0078] In some embodiments, the Gaussian distribution It can be a complex Gaussian distribution. for When v = n, There is no model for speech that satisfies Gaussian distribution. When v=s+n, The variance of the speech absence model when v = n is the speech presence model that satisfies the Gaussian distribution. Different from the variance of the speech presence model when v = s + n

[0079] According to formula (2) and formula (3), the microphone signal x f,t Also satisfies the Gaussian distribution. Specifically, the microphone signal x f,t It can be a speech presence model or a speech absence model that satisfies Gaussian distribution. f,t It can be expressed as the following formula:

[0080]

[0081] in, is x f,t The variance of . For the convenience of description, we will Defined as the spatial covariance matrix. When v = n, x f,t There is no model for speech that satisfies Gaussian distribution. When v = s + n, x f,t There is a model for speech that satisfies the Gaussian distribution.

[0082] The microphone signal x f,t The corresponding speech presence probability can be the microphone signal x f,t The probability of belonging to the speech presence model. For the convenience of description, we will f,t The corresponding speech existence probability is defined as The microphone signal x f,t The corresponding speech absence probability is defined as We will model the presence of speech, the microphone signal x f,t The corresponding speech existence distribution probability is defined as We will model the absence of speech, the microphone signal x f,t The corresponding speech existence distribution probability is defined as The microphone signal x f,t The corresponding speech existence probability It can be expressed as the following formula:

[0083]

[0084] To calculate The computing device 240 needs to determine the speech presence variance corresponding to the speech presence model And the speech absence variance corresponding to the speech absence model Assume the microphone signal x f,t It can be the first model or the second model that satisfies Gaussian distribution. One of the first model and the second model is a speech presence model, and the other is a speech absence model.

[0085] For the convenience of description, we define the first model as the following formula:

[0086]

[0087] in, is the first variance of the Gaussian distribution corresponding to the first model. The first parameter and the first spatial covariance matrix The product of .

[0088] We define the second model as the following formula:

[0089]

[0090] in, is the second variance of the Gaussian distribution corresponding to the second model. The second parameter and the second spatial covariance matrix The product of .

[0091] To calculate The computing device 240 needs to determine which of the first model and the second model is the speech presence model and which is the speech absence model.

[0092] S140: Iteratively optimizing the first model and the second model based on maximum likelihood estimation and expectation maximization algorithms respectively until convergence.

[0093] The computing device 240 may use an iterative optimization method to iteratively optimize the first model and the second model respectively to obtain the first variance of the first model. and the second variance of the second model In the iterative process, the computing device 240 may calculate the value of the microphone signal x based on the f,t The first probability when it is the first model Entropy and the microphone signal x f,t The second probability when it is the second model Entropy It is determined whether the speech presence model is the first model or the second model.

[0094] First probability It can be that in the first model and the second model, the microphone signal x f,t The probability of belonging to the first model. The second probability It can be that in the first model and the second model, the microphone signal x f,t The probability of belonging to the second model. Among them, the first probability With the second probability Complementarity, that is We will first model that the microphone signal x f,t The corresponding first distribution probability is defined as In the second model, the microphone signal x f,t The corresponding second distribution probability is defined as The microphone signal x f,t The corresponding first probability It can be expressed as the following formula:

[0095]

[0096] The microphone signal x f,t The corresponding second probability It can be expressed as the following formula:

[0097]

[0098] Figure 4 A flowchart of an iterative optimization provided according to an embodiment of this specification is shown. Figure 4 The step shown is step S140. Figure 4 As shown, step S140 may include:

[0099] S142: Construct an objective function based on maximum likelihood estimation and expectation maximization algorithm.

[0100] As mentioned before, the unknown parameters include the first variance of the first model and the second variance of the second model Among them, the hidden variable is the microphone signal x f,t The first probability of belonging to the first model and the microphone signal x f,t The first probability belongs to the second model Therefore, the maximum likelihood estimation and expectation maximization algorithm are used to estimate the first variance. and the second variance of the second model Perform iterative optimization. The objective function is the maximum likelihood estimation function. The maximum likelihood estimation function can be expressed as the following formula:

[0101]

[0102] S144: Determine optimization parameters.

[0103] First parameter and the first spatial covariance matrix The relationship between can be expressed as the following formula:

[0104]

[0105] Second parameter and the second spatial covariance matrix The relationship between can be expressed as the following formula:

[0106]

[0107] Therefore, the optimization parameters may include the first spatial covariance matrix and the second spatial covariance matrix

[0108] S145: Determine the initial value of the optimization parameter.

[0109] For the convenience of description, we will first space covariance matrix The initial value is defined as The second spatial covariance matrix The initial value is defined as The first spatial covariance matrix Initial value of and the second spatial covariance matrix Initial value In some embodiments, the first spatial covariance matrix Initial value of and / or the second spatial covariance matrix Initial value It can be the identity matrix IN. In some embodiments, the first spatial covariance matrix Initial value of and / or the second spatial covariance matrix Initial value It can be directly calculated based on the microphone signals of several adjacent frames. and / or It can be expressed as the following formula:

[0110]

[0111] S146: Based on the objective function and the initial values ​​of the optimization parameters, perform multiple iterations on the optimization parameters until the objective function converges.

[0112] As mentioned above, the computing device 240 may calculate the probability of Entropy and the second probability Entropy It is determined whether the speech presence probability is the first model or the second model.

[0113] In some embodiments, the computing device 240 may calculate the probability of the first Entropy and the second probability Entropy Determine whether the speech existence probability is the first model or the second model, such as Figure 5 shown. Figure 5 FIG. 4 shows a flowchart of multiple iterations provided according to an embodiment of this specification, corresponding to step S146. Figure 5 As shown, step S146 may be included in each iteration:

[0114] S146-2: Perform reversible correction on the optimization parameters.

[0115] Specifically, step S146-2 may be, when it is determined that the optimization parameter is irreversible, modifying the optimization parameter by using a deviation matrix. The deviation matrix may include a unit matrix, a random matrix that obeys a normal distribution or a uniform distribution. As mentioned above, the optimization parameter includes a first spatial covariance matrix and the second spatial covariance matrix According to formula (11) and formula (12), to obtain the first parameter and the second parameter The first spatial covariance matrix and the second spatial covariance matrix Need to be reversible. The larger the matrix condition number, the closer the matrix is ​​to a singular matrix (irreversible matrix). When the first space covariance matrix or the second spatial covariance matrix When it is not invertible (i.e. the matrix condition number is greater than a certain threshold η), the first spatial covariance matrix is ​​given or the second spatial covariance matrix Add a slight perturbation to make the correction and ensure its reversibility.

[0116] Specifically, the computing device 240 may calculate the first spatial covariance matrix and the second spatial covariance matrix Make a reversible judgment. If or It represents the first spatial covariance matrix Or the second spatial covariance matrix Irreversible, requiring reversibility correction. Wherein, η is the condition number threshold. In some embodiments, η = 10000. In some embodiments, η can be larger or smaller.

[0117] When the first spatial covariance matrix Or the second spatial covariance matrix When it is not reversible, the first space covariance matrix can be obtained by the deviation matrix Q Or the second spatial covariance matrix Make corrections. At this time, the first spatial covariance matrix Or the second spatial covariance matrix It can be expressed as the following formula:

[0118]

[0119]

[0120] Where Q is the deviation matrix, μ is the deviation coefficient, and in some embodiments, μ=0.001.

[0121] When the first spatial covariance matrix and the second spatial covariance matrix When both are reversible, no correction is required.

[0122] S146-3: Determine the first parameter based on formula (11) and formula (12) and the second parameter

[0123] S146-4: Determine the first probability based on formula (8) and formula (9) And the second probability

[0124] S146-5: Based on the first probability And the second probability Update the first spatial covariance matrix of the optimization parameters and the second spatial covariance matrix

[0125] The first spatial covariance matrix and the second spatial covariance matrix It can be expressed as the following formula:

[0126]

[0127]

[0128] S146-6: Based on the objective function, determine whether to stop the iteration.

[0129] Step S146-6 may include:

[0130] S146-7: Determine the end of the iteration and output the convergence value of the optimization parameter. Or

[0131] S146-8: Determine that the iteration has not stopped and proceed to the next iteration.

[0132] like Figure 5 As shown, step S146 may further include:

[0133] S146-9: In any one of the multiple iterations, based on the first probability Entropy and the second probability Entropy Determine whether the speech presence probability is the first model or the second model.

[0134] Step S146-9 can be performed during the iteration process or after the iteration is completed, with the first probability in any of the multiple iterations. and the second probability To calculate the parameters, calculate the first probability Entropy and the second probability Entropy This determines whether the speech presence probability corresponds to the first model or the second model. Entropy represents the degree of chaos, or disorder, in a system. The more disordered a system is, the greater its entropy value; the more ordered a system is, the smaller its entropy value. N signal sources containing only noise are more disordered than N signal sources containing speech signals. Therefore, the entropy of the speech-absent model is greater than the entropy of the speech-present model.

[0135] Specifically, in step S146-9, the computing device 240 may obtain the first probability in any iteration and the second probability And calculate the first probability Entropy and the second probability Entropy When the first probability Entropy Greater than the second probability Entropy When the first probability Entropy Less than the second probability Entropy When , the computing device 240 may determine that the speech presence model is the first model and the second model is the speech absence model.

[0136] In some embodiments, the computing device 240 may, in a first iteration of the plurality of iterations, calculate the probability of Entropy and the second probability Entropy Determine whether the speech existence probability is the first model or the second model, and in each iteration of the subsequent multiple iterations, the first probability and the second probability Correction is made to correct the probability misjudgment of speech, such as Figure 6 shown. Figure 6 FIG. 4 shows another multi-iteration flow chart according to an embodiment of the present specification, corresponding to step S146. Figure 6 As shown, step S146 may include:

[0137] S146-10: In the first iteration of the multiple iterations, calculate the first probability Entropy and the second probability Entropy Determine whether the speech presence probability is the first model or the second model.

[0138] Specifically, in step S146-1, the computing device 240 may determine the first parameter based on formula (11) and formula (12) in the first iteration. and the second parameter Then, the first probability is determined based on formula (8) and formula (9): And the second probability Then calculate the first probability Entropy and the second probability Entropy And compare. When the first probability Entropy Greater than the second probability Entropy When the first probability Entropy Less than the second probability Entropy When , the computing device 240 may determine that the speech presence model is the first model and the second model is the speech absence model.

[0139] like Figure 6 As shown, step S146 may also be included in each iteration after the first iteration:

[0140] S146-11: Perform reversible correction on the optimization parameters. As described above, step S146-2 will not be repeated here.

[0141] S146-12: Determine the first parameter based on formula (11) and formula (12) and the second parameter

[0142] S146-13: Determine the first probability based on formula (8) and formula (9) And the second probability

[0143] S146-14: Based on the first probability Entropy and the second probability Entropy For the first probability and the second probability Make corrections.

[0144] Specifically, step S146-14 may be that the computing device 240 calculates the first probability Entropy and the second probability Entropy And compare. When the voice existence model is the first model, if the first probability Entropy Greater than the second probability Entropy Then the first probability The corresponding value and the second probability The corresponding values ​​are swapped. The corresponding value is updated to the second probability The corresponding value is the second probability The corresponding value is updated to the first probability Corresponding value. When the voice presence model is the first model, if the first probability Entropy Less than the second probability Entropy Then the first probability is not correct and the second probability When the speech existence model is the second model, if the first probability Entropy Less than the second probability Entropy Then the first probability The corresponding value and the second probability The corresponding values ​​are swapped. The corresponding value is updated to the second probability The corresponding value is the second probability The corresponding value is updated to the first probability When the voice presence model is the second model, if the first probability Entropy Greater than the second probability Entropy Then the first probability is not correct and the second probability Make corrections.

[0145] S146-15: Based on the revised first probability and the second probability Update the optimization parameters, the first spatial covariance matrix And the second space anti-variance matrix

[0146] In steps S146-14 and S146-15, the entropy of the speech presence model can be made smaller than the entropy of the speech absence model during each iteration to ensure that each iteration converges toward the target direction, thereby accelerating the convergence speed.

[0147] S146-16: Based on the objective function, determine whether to stop the iteration.

[0148] Step S146-16 may include:

[0149] S146-17: Determine the end of the iteration and output the convergence value of the optimization parameter. Or

[0150] S146-18: Determine that the iteration has not stopped and proceed to the next iteration.

[0151] like Figure 4 As shown, step S140 may further include:

[0152] S148: Output the convergence value of the optimization parameter and its corresponding first probability and the second probability

[0153] As mentioned above, when the objective function converges, the computing device 240 can output the value of the optimization parameter corresponding to the convergence of the objective function as the convergence value of the optimization parameter. At the same time, the computing device 240 can output the first probability corresponding to the convergence value of the optimization parameter and the second probability Output. As shown in formula (16) and formula (17), the first spatial covariance matrix of the optimization parameter is and the second spatial covariance matrix Based on the first probability and the second probability When the objective function converges, the computing device 240 can calculate the first spatial covariance matrix and the second spatial covariance matrix The corresponding first probability and the second probability Output.

[0154] like Figure 3 As shown, the method P100 may further include:

[0155] S160: When the maximum likelihood estimation and expectation maximization algorithms converge, the microphone signal x f,t is the probability of speech presence model as the microphone signal x f,t The probability of speech existence And output.

[0156] As mentioned above, in step S140, the computing device 240 may calculate the probability of Entropy and the second probability Entropy Determine whether the speech presence model is the first model or the second model. When the speech presence model is the first model, the microphone signal x f,t The probability of the presence of speech model can be the microphone signal x f,t is the first probability of the first model At this time, the microphone signal x f,t The probability of speech existence It can be the first spatial covariance matrix when the objective function converges The first probability corresponding to the convergence value of When the speech presence model is the first model, the microphone signal x f,t The probability of the presence of speech model can be the microphone signal x f,t is the second probability of the second model At this time, the microphone signal x f,t The probability of speech existence It can be the second space covariance matrix when the objective function converges The second probability corresponding to the convergence value of

[0157] The computing device 240 can calculate the probability of speech existence Output to other computing modules, such as speech enhancement module, etc.

[0158] In summary, in the system and method P100 for calculating the probability of speech existence provided in this specification, the calculation device 240 can calculate the probability of speech existence according to the first probability corresponding to the first model. Entropy and the second probability corresponding to the second model Entropy To determine which of the first model and the second model is the speech presence model and which is the speech absence model, thereby obtaining the microphone signal x f,t The probability of speech existence To correct the misjudgment of speech probability in the iterative process and improve the probability of speech existence At the same time, the calculation device 240 can calculate the accuracy of the calculation according to the first probability in the iterative process. Entropy and the second probability Entropy For the first probability and the second probability Correction is made to iterate the optimization parameters in the target direction, thereby accelerating the convergence speed and further improving the probability of speech existence. The calculation accuracy of .

[0159] This specification also provides a speech enhancement system. The speech enhancement system can also be applied to electronic device 200. In some embodiments, the speech enhancement system can include a computing device 240. In some embodiments, the speech enhancement system can be applied to computing device 240. That is, the speech enhancement system can run on computing device 240. The speech enhancement system can include a hardware device with data information processing capabilities and the necessary programs required to drive the hardware device. Of course, the speech enhancement system can also be simply a hardware device with data processing capabilities, or simply a program running on the hardware device.

[0160] The speech enhancement system may store data or instructions for executing the speech enhancement method described in this specification and may execute the data and / or instructions. When the speech enhancement system is running on the computing device 240, the speech enhancement system may obtain the microphone signal from the microphone array 220 based on the communication connection and execute the data or instructions of the speech enhancement method described in this specification. The speech enhancement method is described in other parts of this specification. For example, in Figure 7 The speech enhancement method is introduced in the description of .

[0161] When the speech enhancement system is running on the computing device 240, the speech enhancement system is in communication with the microphone array 220. The storage medium 243 may also include at least one instruction set stored in the data storage device for performing MVDR-based speech enhancement calculations on the microphone signals. The instructions are computer program codes, which may include programs, routines, objects, components, data structures, processes, modules, etc. for executing the speech enhancement method provided in this specification. The processor 242 can read the at least one instruction set and execute the speech enhancement method provided in this specification according to the instructions of the at least one instruction set. The processor 242 can execute all steps included in the speech enhancement method.

[0162] Figure 7 FIG2 shows a flow chart of a speech enhancement method P200 provided according to an embodiment of the present specification. The method P200 can perform speech enhancement on the microphone signal based on the MVDR method. Specifically, the processor 242 can execute the method P200. Figure 7 As shown, the method P200 may include:

[0163] S220: Obtain microphone signals x output by the M microphones f,t .

[0164] As described in step S120, details will not be repeated here.

[0165] S240: Determine the microphone signal x based on the voice presence probability calculation method P100. f,t The probability of speech existence

[0166] S260: Based on the probability of speech presence Determine the microphone signal x f,t The noise covariance matrix

[0167] Noise covariance matrix It can be expressed as the following formula:

[0168]

[0169] S280: Based on the MVDR method and the noise space covariance matrix Determine the microphone signal x f,t The corresponding filter coefficient ω f,t .

[0170] Filter coefficient ω f,t It can be expressed as the following formula:

[0171]

[0172] in, is the steering vector corresponding to the target direction where the target voice is located. θs is the signal incident angle corresponding to the target direction. In some embodiments, θs is known. In some embodiments, θ s is unknown, the computing device 240 may calculate the noise covariance matrix based on Perform subspace decomposition and calculate

[0173] In some embodiments, the filter coefficient ω f , t can also be expressed as the following formula:

[0174]

[0175] in, is the convergence value corresponding to the speech non-existence model. When the first model is the speech non-existence model, for The corresponding convergence value. When the second model is a speech non-existent model, for The corresponding convergence value.

[0176] S290: Based on the filter coefficient ω f,t For the microphone signal x f,t Merge and output the target audio signal y f,t .

[0177] Target audio signal y f,t It can be expressed as the following formula:

[0178] y f,t =ω f,t H x f,t Formula (21)

[0179] The computing device 240 can convert the target audio signal y f,t Output to other electronic devices, such as remote communication devices.

[0180] In summary, the speech presence probability calculation system and method P100, speech enhancement system and method P200, and electronic device 200 provided in this specification are used for a microphone 220 array composed of multiple microphones 222. The speech presence probability calculation system and method P100, speech enhancement system and method P200, and electronic device 200 can respectively obtain a speech presence model when speech is present and a speech absence model when speech is not present in multiple microphone signals, and optimize through multiple iterations based on maximum likelihood estimation and expectation maximization algorithms, and during the iteration process, according to the entropy of the speech presence probability and the entropy of the speech absence probability, correct the speech presence probability and the speech absence probability, thereby calculating and determining the model parameters of the speech presence model and the model parameters when speech is absent, and when the maximum likelihood estimation and expectation maximization algorithms converge, obtain the speech presence probability corresponding to the speech presence model. The speech presence probability calculation system and method P100, speech enhancement system and method P200, and electronic device 200 correct the speech presence probability and speech non-existence probability in the iterative process by comparing the entropy of the speech presence probability and the entropy of the speech non-existence probability to obtain faster convergence speed and better convergence results, thereby making the speech presence probability and noise covariance matrix estimation more accurate, thereby improving the speech enhancement effect of MVDR.

[0181] Another aspect of this specification provides a non-transitory storage medium storing at least one set of executable instructions for calculating the probability of speech presence. When the executable instructions are executed by a processor, the executable instructions direct the processor to implement the steps of the method for calculating the probability of speech presence P100 described in this specification. In some possible implementations, various aspects of this specification may also be implemented in the form of a program product comprising program code. When the program product is executed on a computing device (such as computing device 240), the program code is used to cause the computing device to perform the steps for calculating the probability of speech presence described in this specification. The program product for implementing the above method may include the program code in a portable compact disk read-only memory (CD-ROM) and may be executed on the computing device. However, the program product of this specification is not limited to this. In this specification, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system (such as processor 242). The program product may utilize any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples of computer-readable storage media include: an electrical connection having one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable storage medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing. Program code for performing the operations described herein may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may execute entirely on the computing device, partially on the computing device, as a stand-alone software package, partially on the computing device and partially on a remote computing device, or entirely on the remote computing device.

[0182] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0183] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and may not be limiting. Although not expressly stated herein, those skilled in the art will understand that this specification encompasses various reasonable changes, improvements, and modifications to the embodiments. Such changes, improvements, and modifications are intended to be suggested by this specification and are within the spirit and scope of the exemplary embodiments of this specification.

[0184] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, “one embodiment,” “an embodiment,” and / or “some embodiments” mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is emphasized and should be understood that two or more references to “an embodiment,” “one embodiment,” or “an alternative embodiment” in various parts of this specification do not necessarily refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.

[0185] It should be understood that in the foregoing descriptions of the embodiments of this specification, to facilitate understanding of a feature and to simplify this specification, various features are combined in a single embodiment, figure, or description thereof. However, this does not necessarily mean that these features are combined. When reading this specification, those skilled in the art may extract some of the features and understand them as separate embodiments. In other words, the embodiments of this specification can also be understood as the integration of multiple sub-embodiments. This also applies when each sub-embodiment contains fewer than all the features of a single previously disclosed embodiment.

[0186] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, documents, articles, etc., cited herein is hereby incorporated by reference in its entirety for all purposes, except for any prosecution document history related thereto, any equivalent that may be inconsistent or conflicting with this document, or any equivalent prosecution document history that may have a limiting effect on the broadest scope of the claims now or hereafter associated with this document. For example, if there is any inconsistency or conflict between the description, definition, and / or use of terms associated with any incorporated material and the terminology, description, definition, and / or use associated with this document, the terminology in this document shall control.

[0187] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.

Claims

1. A method for calculating the probability of speech presence, characterized in that: M microphones distributed in a preset array shape, where M is an integer greater than 1, including: Acquire microphone signals output by the M microphones, where the microphone signals satisfy a first model or a second model of a Gaussian distribution, where one of the first model and the second model is a speech presence model, and the other is a speech absence model; Iteratively optimizing the first model and the second model based on maximum likelihood estimation and expectation maximization algorithms respectively until convergence, and during the iteration process, determining whether the speech presence model is the first model or the second model based on an entropy of a first probability when the microphone signal is the first model and an entropy of a second probability when the microphone signal is the second model, the first probability and the second probability being complementary; and When the maximum likelihood estimation and expectation maximization algorithms converge, the probability that the microphone signal is the speech presence model is output as the speech presence probability of the microphone signal.

2. The method for calculating the probability of speech existence according to claim 1, wherein: A first variance of the Gaussian distribution corresponding to the first model includes a product of a first parameter and a first spatial covariance matrix; and The second variance of the Gaussian distribution corresponding to the second model includes the product of the second parameter and the second spatial covariance matrix; The iterative optimization of the first model and the second model based on maximum likelihood estimation and expectation maximization algorithm respectively includes: Construct the objective function based on maximum likelihood estimation and expectation maximization algorithm; Determining optimization parameters, where the optimization parameters include the first spatial covariance matrix and the second spatial covariance matrix; Determining initial values ​​of the optimization parameters; Based on the objective function and the initial values ​​of the optimization parameters, the optimization parameters are iterated multiple times until the objective function converges, including: determining, in the plurality of iterations, whether the speech presence probability is the first model or the second model based on the entropy of the first probability and the entropy of the second probability; and Output the convergence value of the optimization parameter and the first probability and the second probability corresponding to the value.

3. The method for calculating the probability of speech existence according to claim 2, wherein: The determining, in the plurality of iterations, whether the speech presence probability is the first model or the second model based on the entropy of the first probability and the entropy of the second probability comprises: In any one of the multiple iterations, calculating the entropy of the first probability and the entropy of the second probability, and determining whether the speech presence probability is the first model or the second model, includes: determining that the entropy of the first probability is greater than the entropy of the second probability, and determining that the speech presence model is the second model; or It is determined that the entropy of the first probability is less than the entropy of the second probability, and the speech presence model is determined to be the first model.

4. The method for calculating the probability of speech existence according to claim 2, wherein: The determining, in the plurality of iterations, whether the speech presence probability is the first model or the second model based on the entropy of the first probability and the entropy of the second probability comprises: In a first iteration of the multiple iterations, calculating the entropy of the first probability and the entropy of the second probability, and determining whether the speech presence probability is the first model or the second model, includes: determining that the entropy of the first probability is greater than the entropy of the second probability, and determining that the speech presence model is the second model; or Determining that the entropy of the first probability is less than the entropy of the second probability, and determining that the speech presence model is the first model; 5. The method for calculating the probability of speech existence according to claim 4, wherein: The iterating the optimization parameters multiple times further includes, in each of the multiple iterations: Modifying the first probability and the second probability based on the entropy of the first probability and the entropy of the second probability includes: Determining that the first model is the speech presence model, and the entropy of the first probability is greater than the entropy of the second probability, and exchanging the value corresponding to the first probability with the value corresponding to the second probability; or Determining that the second model is the speech presence model, and that the entropy of the second probability is greater than the entropy of the first probability, and exchanging a value corresponding to the first probability with a value corresponding to the second probability; and The optimization parameter is updated based on the revised first probability and the second probability.

6. The method for calculating the probability of speech existence according to claim 2, wherein: The iterating the optimization parameters multiple times further includes, in each of the multiple iterations: Performing a reversible correction on the optimization parameters includes: It is determined that the optimization parameters are irreversible, and the optimization parameters are corrected using a deviation matrix, where the deviation matrix includes one of a unit matrix and a random matrix that obeys a normal distribution or a uniform distribution.

7. A system for calculating the probability of speech existence, characterized in that: include: at least one storage medium storing at least one instruction set for calculating the probability of speech presence; as well as at least one processor, in communication with the at least one storage medium; When the speech existence probability calculation system is running, the at least one processor reads the at least one instruction set and implements the speech existence probability calculation method according to any one of claims 1 to 6.

8. A speech enhancement method, characterized in that: M microphones distributed in a preset array shape, where M is an integer greater than 1, including: Obtaining microphone signals output by the M microphones; Determining the speech presence probability of the microphone signal based on the speech presence probability calculation method according to any one of claims 1 to 6; Determining a noise covariance matrix of the microphone signal based on the speech presence probability; Determining a filter coefficient corresponding to the microphone signal based on the MVDR method and the noise covariance matrix; and The microphone signals are combined based on the filter coefficients to output a target audio signal.

9. A headset, characterized in that: include: A microphone array, comprising M microphones distributed in a preset array shape, where M is an integer greater than 1; as well as A computing device is communicatively connected to the microphone array during operation and executes the speech enhancement method according to claim 8.

10. The earphone according to claim 9, wherein The M microphones are linearly distributed, and M is no greater than 5. The spacing between adjacent microphones in the M microphones is between 20 mm and 40 mm. The headset further includes: A first housing, on which the microphone array is mounted, comprising: A first interface is provided with a first magnetic device; and a contact, provided at the first interface, and communicatively connected to the microphone array; and a second housing on which the computing device is mounted, comprising: The second interface is provided with a second magnetic device; and a guide rail, provided at the second interface, in communication with the computing device, The adsorption force between the first magnetic device and the second magnetic device enables the first shell and the second shell to be detachably connected. When the first shell and the second shell are connected, the contact contacts the guide rail, so that the microphone array is communicatively connected with the computing device.

Citation Information

Patent Citations

  • MMSE-LSA speech enhancement method based on improved noise estimation

    CN112201269A

  • Multi-channel speech enhancement method and device

    CN113030862A