Speech Activity Detection Method, System, Speech Enhancement Method, and System

By adopting a joint optimization method of likelihood function maximization and noise covariance matrix rank minimization in microphone signals, combined with statistical hypothesis testing, the problem of poor speech enhancement effect on devices with small number of microphones and small spacing is solved, and higher precision speech activity detection and enhancement is achieved.

CN116110421BActive Publication Date: 2025-07-04SHENZHEN SHOKZ CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111331926.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-11
Publication Date
2025-07-04
Estimated Expiration
2041-11-11

AI Technical Summary

Technical Problem

In the prior art, the speech enhancement method based on the beamforming algorithm is not effective on devices with small microphones and small spacing, especially for devices such as headphones. This is mainly due to the insufficient estimation accuracy of the noise covariance matrix, resulting in poor speech enhancement effect of the MVDR algorithm.

Method used

The likelihood function maximization and noise covariance matrix rank minimization are used as joint optimization goals. The statistical hypothesis test is used to determine whether the target speech signal exists in the microphone signal, and the estimated value of the noise covariance matrix is ​​optimized to improve the detection and enhancement effect of speech activity.

Benefits of technology

The estimation accuracy of the noise covariance matrix is ​​improved, and the accuracy of speech activity detection and speech enhancement effect are enhanced, especially on devices with small microphone counts and small spacing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116110421B_ABST
    Figure CN116110421B_ABST
Patent Text Reader

Abstract

In the voice activity detection method, system, voice enhancement method, and system provided in this specification, the microphone signals output by the microphone array satisfy a first model corresponding to a noise signal or a second model corresponding to a mixture of a target voice signal and the noise signal. The method and system can use maximizing the likelihood function and minimizing the rank of the noise covariance matrix as a joint optimization objective to optimize the first model and the second model respectively, determine a first estimate of the noise covariance matrix of the first model and a second estimate of the noise covariance matrix of the second model, and determine whether the microphone signal satisfies the first model or the second model through a statistical hypothesis testing method, so as to determine whether there is a target voice signal in the microphone signal, determine the noise covariance matrix of the microphone signal, and further perform voice enhancement on the microphone signal. The method and system can improve the accuracy of noise covariance estimation, and thus improve the voice enhancement effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the technical field of target voice signal processing, and in particular, to a voice activity detection method, system, voice enhancement method, and system. Background Art

[0002] In the voice enhancement technology based on the beamforming algorithm, especially in the adaptive beamforming algorithm of Minimum Variance Distortionless Response (MVDR for short), it is crucial to solve the parameter - the noise covariance matrix that describes the relationship of noise statistical characteristics between different microphones. The main method in the prior art is to calculate the noise covariance matrix based on the voice presence probability method. For example, the voice presence probability is estimated through the Voice Activity Detection (VAD for short) method, and then the noise covariance matrix is calculated. However, the accuracy of the voice presence probability estimation in the prior art is not high enough, resulting in a low estimation accuracy of the noise covariance matrix, and further resulting in a poor voice enhancement effect of the MVDR algorithm. Especially when the number of microphones is small, such as less than 5, the effect drops sharply. Therefore, the MVDR algorithm in the prior art is mostly used in microphone array devices with a large number of microphones and large spacings, such as mobile phones and smart speakers, while the voice enhancement effect for devices with a small number of microphones and small spacings, such as headphones, is poor.

[0003] Therefore, there is a need to provide a voice activity detection method, system, voice enhancement method, and system with higher accuracy. Summary of the Invention

[0004] This specification provides a voice activity detection method, system, voice enhancement method, and system with higher accuracy.

[0005] In a first aspect, this specification provides a voice activity detection method for M microphones distributed in a preset array shape, where M is an integer greater than 1, including: obtaining microphone signals output by the M microphones, where the microphone signals satisfy a first model corresponding to the absence of a target voice signal or a second model corresponding to the presence of a target voice signal; taking the maximization of the likelihood function and the minimization of the rank of the noise covariance matrix as a joint optimization objective, respectively optimizing the first model and the second model, and determining a first estimate value of the noise covariance matrix of the first model and a second estimate value of the noise covariance matrix of the second model; and based on statistical hypothesis testing, determining the target model corresponding to the microphone signals and the noise covariance matrix, where the target model includes one of the first model and the second model, and the noise covariance matrix of the microphone signals is the noise covariance matrix of the target model.

[0006] In some embodiments, the microphone signal includes K consecutive frames of audio signals, where K is a positive integer greater than 1, and the microphone signal includes an M×K data matrix.

[0007] In some embodiments, the microphone signal is a complete observation signal or an incomplete observation signal. All data in the M×K data matrix in the complete observation signal is complete, and some data in the M×K data matrix in the incomplete observation signal is missing. When the microphone signal is the incomplete observation signal, obtaining the microphone signals output by the M microphones includes: obtaining the incomplete observation signal; performing row and column permutation on the microphone signal based on the data missing positions in each column of the M×K data matrix, and dividing the microphone signal into at least one sub-microphone signal, and the microphone signal includes the at least one sub-microphone signal.

[0008] In some embodiments, taking the maximization of the likelihood function and the minimization of the rank of the noise covariance matrix as a joint optimization objective, optimizing the first model and the second model respectively, includes: using the microphone signal as sample data, establishing a first likelihood function corresponding to the first model, and the likelihood function includes the first likelihood function; taking the maximization of the first likelihood function and the minimization of the rank of the noise covariance matrix of the first model as the optimization objective, optimizing the first model to determine the first estimated value; using the microphone signal as sample data, establishing a second likelihood function of the second model, and the likelihood function includes the second likelihood function; and taking the maximization of the second likelihood function and the minimization of the rank of the noise covariance matrix of the second model as the optimization objective, optimizing the second model to determine the second estimated value and the amplitude estimated value of the target speech signal.

[0009] In some embodiments, the microphone signal includes a noise signal, the noise signal follows a Gaussian distribution, and the noise signal at least includes: a colored noise signal, which follows a Gaussian distribution with zero mean, and its corresponding noise covariance matrix is a low-rank positive semi-definite matrix.

[0010] In some embodiments, determining the target model and the noise covariance matrix corresponding to the microphone signal based on the statistical hypothesis test includes: establishing a binary hypothesis test model based on the microphone signal, where the null hypothesis of the binary hypothesis test model includes that the microphone signal satisfies the first model, and the alternative hypothesis of the binary hypothesis test model includes that the microphone signal satisfies the second model; substituting the first estimate value, the second estimate value, and the amplitude estimate value into the decision criterion of the detector of the binary hypothesis test model to obtain a test statistic; and determining the target model of the microphone signal based on the test statistic.

[0011] In some embodiments, determining the target model of the microphone signal based on the test statistic includes: determining that the test statistic is greater than the preset decision threshold, determining that the target voice signal exists in the microphone signal, determining that the target model is the second model, and the noise covariance matrix of the microphone signal is the second estimate value; or determining that the test statistic is less than the preset decision threshold, determining that the target voice signal does not exist in the microphone signal, determining that the target model is the first model, and the noise covariance matrix of the microphone signal is the first estimate value.

[0012] In a second aspect, this specification also provides a voice activity detection system, including at least one storage medium and at least one processor. The at least one storage medium stores at least one instruction set for voice activity detection; the at least one processor is communicatively connected to the at least one storage medium. When the voice activity detection system runs, the at least one processor reads the at least one instruction set and implements the voice activity detection method described in the first aspect of this specification.

[0013] In a third aspect, this specification also provides a voice enhancement method for M microphones distributed in a preset array shape, where M is an integer greater than 1, including: obtaining the microphone signals output by the M microphones; determining the target model of the microphone signals and the noise covariance matrix of the microphone signals based on the voice activity detection method according to any one of claims 1-8, and the noise covariance matrix of the microphone signals is the noise covariance matrix of the target model; determining the filtering coefficients corresponding to the microphone signals based on the MVDR method and the noise covariance matrix of the microphone signals; and combining the microphone signals based on the filtering coefficients to output a target audio signal.

[0014] Fourthly, this specification also provides a voice enhancement system, including at least one storage medium and at least one processor. The at least one storage medium stores at least one instruction set for voice enhancement. The at least one processor is communicatively connected to the at least one storage medium. When the voice enhancement system runs, the at least one processor reads the at least one instruction set and implements the voice enhancement method described in the third aspect of this specification.

[0015] As can be seen from the above technical solutions, the voice activity detection methods, systems, voice enhancement methods, and systems provided in this specification are used for a microphone array composed of multiple microphones. Among them, the microphone signals output by the microphone array satisfy the first model corresponding to the noise signal or the second model corresponding to the mixture of the target voice signal and the noise signal. In order to obtain whether there is a target voice signal in the microphone signals, the methods and systems can use the maximization of the likelihood function and the minimization of the rank of the noise covariance matrix as the joint optimization objectives, optimize the first model and the second model respectively, determine the first estimate of the noise covariance matrix of the first model and the second estimate of the noise covariance matrix of the second model, and judge whether the microphone signals satisfy the first model or the second model through statistical hypothesis testing, so as to determine whether there is a target voice signal in the microphone signals, and determine the noise covariance matrix of the microphone signals, and then perform voice enhancement on the microphone signals based on the MVDR method. The methods and systems can improve the accuracy of noise covariance estimation, and thus improve the voice enhancement effect.

[0016] Some of the other functions of the voice activity detection methods, systems, voice enhancement methods, and systems provided in this specification will be listed below. According to the description, the content introduced by the following numbers and examples will be obvious to those of ordinary skill in the art. The creative aspects of the voice activity detection methods, systems, voice enhancement methods, and systems provided in this specification can be fully explained by practicing or using the methods, devices, and combinations described in the detailed examples below. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 Shows a hardware schematic diagram of a voice activity detection system provided according to an embodiment of this specification;

[0019] Figure 2A Shows an exploded structural schematic diagram of an electronic device provided according to an embodiment of the present specification;

[0020] Figure 2B Shows a front view of a first housing provided according to an embodiment of the present specification;

[0021] Figure 2C Shows a top view of a first housing provided according to an embodiment of the present specification;

[0022] Figure 2D Shows a front view of a second housing provided according to an embodiment of the present specification;

[0023] Figure 2E Shows a bottom view of a second housing provided according to an embodiment of the present specification;

[0024] Figure 3 Shows a flowchart of a voice activity detection method provided according to an embodiment of the present specification;

[0025] Figure 4 Shows a schematic diagram of a complete observation signal provided according to an embodiment of the present specification;

[0026] Figure 5A Shows a schematic diagram of an incomplete observation signal provided according to an embodiment of the present specification;

[0027] Figure 5B Shows a schematic diagram of the rearrangement of an incomplete observation signal provided according to an embodiment of the present specification;

[0028] Figure 5C Shows a schematic diagram of the rearrangement of an incomplete observation signal provided according to an embodiment of the present specification;

[0029] Figure 6 Shows a flowchart of an iterative optimization provided according to an embodiment of the present specification;

[0030] Figure 7 Shows a flowchart of determining a target model provided according to an embodiment of the present specification; and

[0031] Figure 8 Shows a flowchart of a voice enhancement method provided according to an embodiment of the present specification. Detailed implementation manners

[0032] The following description provides specific application scenarios and requirements of this specification, aiming to enable those skilled in the art to manufacture and use the content in this specification. For those skilled in the art, various partial modifications to the disclosed embodiments are obvious, and without departing from the spirit and scope of this specification, the general principles defined here can be applied to other embodiments and applications. Therefore, this specification is not limited to the illustrated embodiments, but rather to the broadest scope consistent with the claims.

[0033] The terms used herein are for the purpose of describing specific example embodiments only and are not restrictive. For example, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. When used in this specification, the terms "comprising", "including", and / or "containing" mean that the associated integers, steps, operations, elements, and / or components exist, but do not exclude the existence of one or more other features, integers, steps, operations, elements, components, and / or groups, or the addition of other features, integers, steps, operations, elements, components, and / or groups in the system / method.

[0034] Considering the following description, these features and other features of this specification, as well as the operations and functions of the related elements of the structure, and the economy of the combination and manufacture of the components can be significantly improved. Referring to the accompanying drawings, all of these form a part of this specification. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0035] The flowcharts used in this specification illustrate the operations implemented by the system according to some embodiments in this specification. It should be clearly understood that the operations in the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.

[0036] For the convenience of description, the terms that will appear in this specification are first explained as follows:

[0037] Statistical hypothesis testing: It is a method in mathematical statistics to infer the population from samples according to certain assumptions. The specific approach is as follows: Make a certain assumption about the population under study according to the needs of the problem, denoted as the original hypothesis H0; Select an appropriate statistic, and the selection of this statistic should make its distribution known when the original hypothesis H0 holds; Calculate the value of the statistic from the measured samples, and conduct a test according to the pre-given significance level to make a judgment of rejecting or accepting the original hypothesis H0. Common statistical hypothesis testing methods include the u-test method, t-test method, χ2-test method (chi-square test), F-test method, rank sum test, etc.

[0038] Minimum Variance Distortionless Response (MVDR): It is an adaptive beamforming algorithm based on the maximum signal-to-interference-plus-noise ratio (SINR) criterion. The MVDR algorithm can adaptively minimize the power of the array output in the desired direction while maximizing the signal-to-interference-plus-noise ratio. Its goal is to minimize the variance of the recorded signal. If the noise signal and the desired signal are uncorrelated, then the variance of the recorded signal is the sum of the variances of the desired signal and the noise signal. Therefore, the MVDR solution seeks to minimize this sum, thereby reducing the impact of the noise signal. Its principle is to select appropriate filter coefficients under the constraint that the desired signal is undistorted, so as to minimize the average power of the array output.

[0039] Voice activity detection: The process of segmenting the speaking speech period and the non-speaking period in the target speech signal.

[0040] Gaussian distribution: Normal distribution, also known as "normal distribution", also known as Gaussian distribution (Gaussian distribution). The normal curve is bell-shaped, low at both ends, high in the middle, and symmetrical about the y-axis. Because its curve is bell-shaped, it is often called the bell curve. If the random variable X follows a normal distribution with a mathematical expectation of μ and a variance of σ 2 , it is denoted as N(μ, σ 2 ). The expected value μ of the probability density function of the normal distribution determines its position, and its standard deviation σ determines the amplitude of the distribution. When μ = 0 and σ = 1, the normal distribution is the standard normal distribution.

[0041] Figure 1 FIG. shows a hardware schematic diagram of a voice activity detection system provided according to an embodiment of the present specification. The voice activity detection system can be applied to the electronic device 200.

[0042] In some embodiments, the electronic device 200 may be a wireless headset, a wired headset, a smart wearable device, such as a smart glasses, a smart helmet, or a smart watch, etc., which are devices with audio processing functions. The electronic device 200 may also be a mobile device, a tablet computer, a laptop computer, an in-vehicle device of a motor vehicle, or the like, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, or the like, or any combination thereof. For example, the smart mobile device may include a mobile phone, a personal digital assistant, a gaming device, a navigation device, an ultra-mobile personal computer (UMPC), etc., or any combination thereof. In some embodiments, the smart home device may include a smart TV, a desktop computer, etc., or any combination. In some embodiments, the in-vehicle device in a motor vehicle may include an in-vehicle computer, an in-vehicle TV, etc.

[0043] In this specification, we will describe the electronic device 200 as an earphone as an example. The earphone may be a wireless earphone or a wired earphone. As Figure 1 shown, the electronic device 200 may include a microphone array 220 and a computing device 240.

[0044] The microphone array 220 may be an audio acquisition device of the electronic device 200. The microphone array 220 may be configured to acquire local audio and output a microphone signal, that is, an electronic signal carrying audio information. The microphone array 220 may include M microphones 222 distributed in a preset array shape. Wherein, the M is an integer greater than 1. The M microphones 222 may be evenly distributed or unevenly distributed. The M microphones 222 may output microphone signals. The M microphones 222 may output M microphone signals. Each microphone 222 corresponds to a microphone signal. The M microphone signals are collectively referred to as the microphone signal. In some embodiments, the M microphones 222 may be linearly distributed. In some embodiments, the M microphones 222 may also be distributed in an array of other shapes, such as a circular array, a rectangular array, etc. For the convenience of description, in the following description, we will describe the case where the M microphones 222 are linearly distributed as an example. In some embodiments, M may be any integer greater than 1, such as 2, 3, 4, 5, or even more, etc. In some embodiments, due to space limitations, M may be an integer greater than 1 and not greater than 5, such as in products such as earphones. When the electronic device 200 is an earphone, the distance between adjacent microphones 222 among the M microphones 222 may be between 20 mm and 40 mm. In some embodiments, the distance between adjacent microphones 222 may be smaller, such as between 10 mm and 20 mm.

[0045] In some embodiments, the microphone 222 can be a bone conduction microphone that directly collects human body vibration signals. The bone conduction microphone can include vibration sensors, such as optical vibration sensors, acceleration sensors, etc. The vibration sensor can collect mechanical vibration signals (such as signals generated by vibrations of the skin or bones when the user speaks), and convert the mechanical vibration signals into electrical signals. The mechanical vibration signals mentioned here mainly refer to vibrations transmitted through solids. The bone conduction microphone contacts the user's skin or bones through the vibration sensor or a vibration component connected to the vibration sensor, so as to collect the vibration signals generated by the bones or skin when the user makes a sound, and convert the vibration signals into electrical signals. In some embodiments, the vibration sensor can be a device that is sensitive to mechanical vibrations and insensitive to air vibrations (that is, the response ability of the vibration sensor to mechanical vibrations exceeds the response ability of the vibration sensor to air vibrations). Since the bone conduction microphone can directly pick up the vibration signals of the sound generation part, the bone conduction microphone can reduce the influence of environmental noise.

[0046] In some embodiments, the microphone 222 can also be an air conduction microphone that directly collects air vibration signals. The air conduction microphone collects the air vibration signals generated when the user makes a sound, and converts the air vibration signals into electrical signals.

[0047] In some embodiments, the M microphones 220 can be M bone conduction microphones. In some embodiments, the M microphones 220 can also be M air conduction microphones. In some embodiments, the M microphones 220 can include both bone conduction microphones and air conduction microphones. Of course, the microphone 222 can also be other types of microphones. Such as optical microphones, microphones that receive electromyographic signals, and so on.

[0048] The computing device 240 can be communicatively connected to the microphone array 220. The communicative connection refers to any form of connection that can directly or indirectly receive information. In some embodiments, the computing device 240 can communicate with the microphone array 220 to transfer data to each other through a wireless communication connection; in some embodiments, the computing device 240 can also communicate with the microphone array 220 to transfer data to each other through a direct wire connection; in some embodiments, the computing device 240 can also establish an indirect connection with the microphone array 220 by directly connecting to other circuits through wires, so as to achieve data transfer to each other. In this specification, the example of the computing device 240 being directly connected to the microphone array 220 by wire will be described.

[0049] The computing device 240 can be a hardware device with data information processing capabilities. In some embodiments, the voice activity detection system can include the computing device 240. In some embodiments, the voice activity detection system can be applied to the computing device 240. That is, the voice activity detection system can run on the computing device 240. The voice activity detection system can include a hardware device with data information processing capabilities and the necessary programs required to drive the operation of the hardware device. Of course, the voice activity detection system can also be only a hardware device with data processing capabilities, or only a program running on the hardware device.

[0050] The voice activity detection system can store data or instructions for performing the voice activity detection method described in this specification and can execute the data and / or instructions. When the voice activity detection system runs on the computing device 240, the voice activity detection system can obtain the microphone signals from the microphone array 220 based on the communication connection and execute the data or instructions of the voice activity detection method described in this specification to calculate whether there is a target voice signal in the microphone signals. The voice activity detection method is introduced in other parts of this specification. For example, in Figures 3 to 8 the description of

[0051] As Figure 1 shown, the computing device 240 can include at least one storage medium 243 and at least one processor 242. In some embodiments, the electronic device 200 can also include a communication port 245 and an internal communication bus 241.

[0052] The internal communication bus 241 can connect different system components, including the storage medium 243, the processor 242, and the communication port 245.

[0053] The communication port 245 can be used for data communication between the computing device 240 and the outside world. For example, the computing device 240 can obtain the microphone signals from the microphone array 220 through the communication port 245.

[0054] At least one storage medium 243 may include a data storage device. The data storage device may be a non-transitory storage medium or a transitory storage medium. For example, the data storage device may include one or more of a magnetic disk, a read-only storage medium (ROM), or a random access storage medium (RAM). When the voice activity detection system can run on the computing device 240, the storage medium 243 may further include at least one instruction set stored in the data storage device for performing voice activity detection on the microphone signal. The instructions are computer program code, and the computer program code may include programs, routines, objects, components, data structures, procedures, modules, etc. for performing the voice activity detection method provided in this specification.

[0055] At least one processor 242 may be communicatively connected to at least one storage medium 243 via an internal communication bus 241. The communicative connection refers to any form of connection capable of directly or indirectly receiving information. At least one processor 242 is configured to execute the above at least one instruction set. When the voice activity detection system can run on the computing device 240, at least one processor 242 reads the at least one instruction set and performs the voice activity detection method provided in this specification according to the instructions of the at least one instruction set. The processor 242 may execute all steps included in the voice activity detection method. The processor 242 may be in the form of one or more processors. In some embodiments, the processor 242 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, etc., or any combination thereof. For illustrative purposes only, only one processor 242 is described in the computing device 240 in this specification. However, it should be noted that the computing device 240 in this specification may further include multiple processors 242. Therefore, the operations and / or method steps disclosed in this specification may be executed by one processor as described in this specification or jointly executed by multiple processors. For example, if the processor 242 of the computing device 240 executes steps A and B in this specification, it should be understood that steps A and B may also be jointly or separately executed by two different processors 242 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).

[0056] Figure 2A Shows an exploded schematic view of an electronic device 200 provided according to an embodiment of the present specification. As Figure 2A shown, the electronic device 200 may include a microphone array 220, a computing device 240, a first housing 260, and a second housing 280.

[0057] The first housing 260 may be a mounting base for the microphone array 220. The microphone array 220 may be mounted inside the first housing 260. The shape of the first housing 260 may be adaptively designed according to the distribution shape of the microphone array 220, and the present specification does not make excessive limitations on this. The second housing 280 may be a mounting base for the computing device 240. The computing device 240 may be mounted inside the second housing 280. The shape of the second housing 280 may be adaptively designed according to the shape of the computing device 240, and the present specification does not make excessive limitations on this. When the electronic device 200 is an earphone, the second housing 280 may be connected to the wearing part. The second housing 280 may be connected to the first housing 260. As described above, the microphone array 220 may be electrically connected to the computing device 240. Specifically, the microphone array 220 may be electrically connected to the computing device 240 through the connection between the first housing 260 and the second housing 280.

[0058] In some embodiments, the first housing 260 may be fixedly connected to the second housing 280, for example, integrally formed, welded, riveted, bonded, and so on. In some embodiments, the first housing 260 may be detachably connected to the second housing 280. The computing device 240 may be communicatively connected to different microphone arrays 220. Specifically, the different microphone arrays 220 may be different in the number of microphones 222 in the microphone array 220, the array shape, the spacing between the microphones 222, the mounting angle of the microphone array 220 in the first housing 260, the mounting position of the microphone array 220 in the first housing 260, and so on. The user may replace the corresponding microphone array 220 according to different application scenarios so that the electronic device 200 is applicable to a wider range of scenarios. For example, when the distance between the user and the electronic device 200 is relatively close in the application scenario, the user may replace it with a microphone array 220 with a smaller spacing. For another example, when the distance between the user and the electronic device 200 is relatively close in the application scenario, the user may replace it with a microphone array 220 with a larger spacing and a larger number, and so on.

[0059] The detachable connection may be any form of physical connection, for example, screw connection, snap connection, magnetic attraction connection, and so on. In some embodiments, the first housing 260 and the second housing 280 may be magnetically connected. That is, the first housing 260 and the second housing 280 are detachably connected through the adsorption force of the magnetic device.

[0060] Figure 2B shows a front view of a first housing 260 provided according to an embodiment of the present specification; Figure 2C shows a top view of a first housing 260 provided according to an embodiment of the present specification. As Figure 2B and Figure 2C shown, the first housing 260 may include a first interface 262. In some embodiments, the first housing 260 may further include contacts 266. In some embodiments, the first housing 260 may further include an angle sensor ( Figure 2B and Figure 2C not shown in the figures).

[0061] The first interface 262 may be an installation interface between the first housing 260 and the second housing 280. In some embodiments, the first interface 262 may be circular. The first interface 262 may be rotatably connected to the second housing 280. When the first housing 260 is installed on the second housing 280, the first housing 260 may rotate relative to the second housing 280 to adjust the angle of the first housing 260 relative to the second housing 280, thereby adjusting the angle of the microphone array 220.

[0062] A first magnetic device 263 may be provided on the first interface 262. The first magnetic device 263 may be disposed at a position on the first interface 262 close to the second housing 280. The first magnetic device 263 may generate a magnetic adsorption force to achieve a detachable connection with the second housing 280. When the first housing 260 approaches the second housing 260, the first housing 260 and the second housing 280 are quickly connected through the adsorption force. In some embodiments, after the first housing 260 is connected to the second housing 280, the first housing 260 may still rotate relative to the second housing 280 to adjust the angle of the microphone array 220. Under the action of the adsorption force, when the first housing 260 rotates relative to the second housing 280, the connection between the first housing 260 and the second housing 280 can still be maintained.

[0063] In some embodiments, a first positioning device ( Figure 2B and Figure 2C not shown in the figures) may also be provided on the first interface 262. The first positioning device may be a positioning step protruding outward or a positioning hole extending inward. The first positioning device may cooperate with the second housing 280 to achieve quick installation of the first housing 260 and the second housing 280.

[0064] As Figure 2B and Figure 2CAs shown, in some embodiments, the first housing 260 may further include a contact 266. The contact 266 may be installed at the first interface 262. The contact 266 may protrude outward from the first interface 262. The contact 266 may be elastically connected to the first interface 262. The contact 266 may be communicatively connected to M microphones 222 in the microphone array 220. The contact 266 may be made of elastic metal to enable data transmission. When the first housing 260 is connected to the second housing 280, the microphone array 220 may be communicatively connected to the computing device 240 through the contact 266. In some embodiments, the contacts 266 may be circularly distributed. After the first housing 260 is connected to the second housing 280, when the first housing 260 rotates relative to the second housing 280, the contacts 266 may also rotate relative to the second housing 280 and maintain communication connection with the computing device 240.

[0065] In some embodiments, an angle sensor ( Figure 2B and Figure 2C not shown in ) may also be provided on the first housing 260. The angle sensor may be communicatively connected to the contact 266, thereby achieving communication connection with the computing device 240. The angle sensor may collect the angle data of the first housing 260, thereby determining the angle at which the microphone array 220 is located, and providing reference data for the subsequent calculation of the voice presence probability.

[0066] Figure 2D FIG. shows a front view of a second housing 280 provided according to an embodiment of the present specification; Figure 2E FIG. shows a bottom view of a second housing 280 provided according to an embodiment of the present specification. As Figure 2D and Figure 2E shown, the second housing 280 may include a second interface 282. In some embodiments, the second housing 280 may further include a guide rail 286.

[0067] The second interface 282 may be an installation interface between the second housing 280 and the first housing 260. In some embodiments, the second interface 282 may be circular. The second interface 282 may be rotatably connected to the first interface 262 of the first housing 260. When the first housing 260 is installed on the second housing 280, the first housing 260 may rotate relative to the second housing 280 to adjust the angle of the first housing 260 relative to the second housing 280, thereby adjusting the angle of the microphone array 220.

[0068] A second magnetic device 283 may be provided on the second interface 282. The second magnetic device 283 may be disposed at a position of the second interface 282 close to the first housing 260. The second magnetic device 283 may generate a magnetic adsorption force to achieve a detachable connection with the first interface 262. The second magnetic device 283 may be used in cooperation with the first magnetic device 263. When the first housing 260 approaches the second housing 280, the first housing 260 can be quickly mounted on the second housing 280 by the adsorption force between the second magnetic device 283 and the first magnetic device 263. When the first housing 260 is mounted on the second housing 260, the positions of the second magnetic device 283 and the first magnetic device 263 are opposite. In some embodiments, after the first housing 260 is connected to the second housing 280, the first housing 260 may still rotate relative to the second housing 280 to adjust the angle of the microphone array 220. Under the action of the adsorption force, when the first housing 260 rotates relative to the second housing 280, the connection between the first housing 260 and the second housing 280 can still be maintained.

[0069] In some embodiments, a second positioning device ( Figure 2D and Figure 2E not shown in

[0070] such as Figure 2D and Figure 2EAs shown, in some embodiments, the second housing 280 may further include a guide rail 286. The guide rail 286 may be installed at the second interface 282. The guide rail 286 may be communicatively connected to the computing device 240. The guide rail 286 may be made of a metal material to enable data transmission. When the first housing 260 is connected to the second housing 280, the contact 266 may contact the guide rail 286 to form a communication connection, thereby enabling communication between the microphone array 220 and the computing device 240 to achieve data transmission. As described above, the contact 266 may be elastically connected to the first interface 262. Therefore, after the first housing 260 is connected to the second housing 280, under the elastic force of the elastic connection, the contact 266 may be brought into full contact with the guide rail 286 to achieve a reliable communication connection. In some embodiments, the guide rail 286 may be circularly distributed. After the first housing 260 is connected to the second housing 280, when the first housing 260 rotates relative to the second housing 280, the contact 266 may also rotate relative to the guide rail 286 and maintain a communication connection with the guide rail 286.

[0071] Figure 3 The flowchart of a voice activity detection method P100 provided according to an embodiment of the present specification is shown. The method P100 may calculate whether there is a target voice signal in the microphone signal. Specifically, the processor 242 may execute the method P100. As Figure 3 shown, the method P100 may include:

[0072] S120: Obtain microphone signals output by M microphones 222.

[0073] As described above, each microphone 222 may output a corresponding microphone signal. The M microphones 222 correspond to M microphone signals. When calculating whether there is a target voice signal in the microphone signal, the method P100 may calculate based on all the microphone signals among the M microphone signals or based on some of the microphone signals. Therefore, the microphone signal may include the M microphone signals corresponding to the M microphones 222 or some of the microphone signals. In the following description of this specification, the case where the microphone signal may include the M microphone signals corresponding to the M microphones 222 will be used as an example for description.

[0074] In some embodiments, the microphone signal may be a time-domain signal. In some embodiments, in step S120, the computing device 240 may perform frame segmentation and windowing on the microphone signal to divide the microphone signal into a plurality of consecutive audio signals. In some embodiments, in step S120, the computing device 240 may also perform a time-frequency transform on the microphone signal to obtain the frequency-domain signal of the microphone signal. For convenience of description, we label the microphone signal at any frequency point as X. In some embodiments, the microphone signal X may include K consecutive frames of audio signals. K is any positive integer greater than 1. For convenience of description, we label the k-th frame of the microphone signal as x k . The k-th frame of the microphone signal x k can be expressed by the following formula:

[0075] x k = [x 1,k , x 2,k , …, x M,k T Formula (1)

[0076] The k-th frame of the microphone signal x k can be an M-dimensional signal vector composed of M microphone signals. The microphone signal X can be represented as an M×K data matrix. The microphone signal X can be expressed by the following formula:

[0077]

[0078] where the microphone signal X is an M×K data matrix. The m-th row in the data matrix represents the microphone signal received by the m-th microphone, and the k-th column represents the microphone signal of the k-th frame.

[0079] As mentioned above, the microphone 222 can collect the noise in the surrounding environment and output a noise signal, and can also collect the speech of the target user and output the target speech signal. When the target user does not speak, the microphone signal only contains the noise signal. When the target user speaks, the microphone signal contains the target speech signal and the noise signal. The k-th frame of the microphone signal x k can be expressed by the following formula:

[0080] x k = Ps k + d k Formula (3)

[0081] where k = 1, 2, …, K. d k is the noise signal in the k-th frame of the microphone signal x k . s k ​is the amplitude of the target voice signal. P is the target steering vector of the target voice signal.

[0082] The microphone signal X can be expressed by the following formula:

[0083] X = [x1, x2, …, x K = PS + D Formula (4)

[0084] where S is the amplitude of the target voice signal. S = [s1, s2, …, s K . D is the noise signal. D = [d1, d2, …, d K .

[0085] The noise signal d k can be expressed by the following formula:

[0086] d k = [d 1,k , d 2,k , …, d M,k T Formula (5)

[0087] The noise signal d k in the k-th frame microphone signal x k can be an M-dimensional signal vector composed of M microphone signals.

[0088] In some embodiments, the noise signal d k can at least include a colored noise signal c k . In some embodiments, the noise signal d k can also include a white noise signal n k . The noise signal d k can be expressed by the following formula:

[0089] d k = c k + n k Formula (6)

[0090] Then the noise signal D = C + N. Where C is the colored noise signal, C = [c1, c2, …, c K . N is the white noise signal, N = [n1, n2, …, n K .

[0091] The computing device 240 can utilize the unified mapping relationship between the clustering (Cluster) feature of the sound source spatial distribution of the noise signal d k and the parameters of the microphone array 220 to establish a parametric clustering model, and cluster the sound source of the noise signal d k so as to cluster the noise signal dk Divided into a colored noise signal c k and a white noise signal n k .

[0092] In some embodiments, the noise signal D follows a Gaussian distribution. The noise signal d k ~CN(0,M). M is the noise covariance matrix of the noise signal d k . Among them, the colored noise signal c k follows a Gaussian distribution with zero mean. That is, c k ~CN(0,M c ). The noise covariance matrix M corresponding to the colored noise signal c k has a low-rank property and is a low-rank positive semi-definite matrix. The white noise signal n c also follows a Gaussian distribution with zero mean. That is, n k ~CN(0,M k ). The power of the white noise signal n n is k That is That is The noise covariance matrix M of the noise signal d k can be expressed by the following formula:

[0093]

[0094] The noise covariance matrix M of the noise signal d k can be decomposed into the sum of the identity matrix I n and the low-rank positive semi-definite matrix M c .

[0095] In some embodiments, the power of the white noise signal n k can be pre-stored in the computing device 240 In some embodiments, the power of the white noise signal n k can be pre-estimated in the computing device 240 For example, the computing device 240 can estimate the power of the white noise signal n k based on methods such as minimum value tracking and histogram In some embodiments, the computing device 240 can estimate the power of the white noise signal n k based on the method P100

[0096] s k is the complex amplitude of the target speech signal. In some embodiments, there is a target speech signal source around the microphone 222. In some embodiments, there are L target speech signal sources around the microphone 222. At this time, s k can be an L×1 dimensional vector.

[0097] The target guiding vector P is an M×L-dimensional matrix. The target guiding vector P can be expressed by the following formula:

[0098]

[0099] Where f0 is the carrier frequency. d is the distance between adjacent microphones 222. c is the speed of sound. θ1, ……, θ N are respectively the incident angles between the L target voice signal sources and the microphone 222. In some embodiments, the angles of the target voice signal sources s k are usually distributed within a specific set of angular ranges. Therefore, θ1, ……, θ N are known. The relative position relationships of the M microphones 222, such as relative distances or relative coordinates, are pre-stored in the computing device 240. That is, the distance d between adjacent microphones 222 is pre-stored in the computing device 240.

[0100] Figure 4 FIG. shows a schematic diagram of a complete observation signal provided according to an embodiment of the present specification. In some embodiments, the microphone signal X is a complete observation signal, as Figure 4 shown. All the data in the M×K data matrix in the complete observation signal is complete. As Figure 4 shown, the horizontal axis is the frame number k of the microphone signal X, and the vertical axis is the microphone signal number m in the microphone array 220. The m-th row represents the microphone signal received by the m-th microphone 222, and the k-th column represents the microphone signal of the k-th frame.

[0101] Figure 5A FIG. shows a schematic diagram of an incomplete observation signal provided according to an embodiment of the present specification. In some embodiments, the microphone signal X is an incomplete observation signal, as Figure 5A shown. Some of the data in the M×K data matrix in the incomplete observation signal is missing. The computing device 240 can rearrange the incomplete observation signal. As Figure 5A shown, the horizontal axis is the frame number k of the microphone signal X, and the vertical axis is the microphone signal channel number m. The m-th row represents the microphone signal received by the m-th microphone 222, and the k-th column represents the microphone signal of the k-th frame.

[0102] When the microphone signal X is the incomplete observation signal, step S120 may further include rearranging the incomplete observation signal. Figure 5B FIG. shows a schematic diagram of the rearrangement of an incomplete observation signal provided according to an embodiment of the present specification; Figure 5CA schematic diagram of rearranging an incomplete observation signal provided according to an embodiment of this specification is shown. When the computing device 240 rearranges the incomplete observation signal, it can be as follows: The computing device 240 obtains the incomplete observation signal; based on the data missing positions in each column of the M×K data matrix, the computing device 240 performs row and column permutations on the microphone signal X and divides the microphone signal X into at least one sub-microphone signal. The microphone signal X includes the at least one sub-microphone signal.

[0103] In the incomplete observation signal, since the data missing positions in the microphone signals x of different frame numbers k may be the same, in order to reduce the algorithm operation amount and operation time, the computing device 240 can classify the K-frame microphone signals X according to the data missing positions in the microphone signals x k of different frame numbers, divide the microphone signals x k with the same data missing positions into the same sub-microphone signal, and perform a permutation on the row positions in the data matrix of the microphone signal X to make the microphone signal positions in the same sub-microphone signal adjacent, as Figure 5B shown. We divide the K-frame microphone signal X into at least one sub-microphone signal. For convenience of description, we define the number of at least one sub-microphone signal as G. Among them, G is a positive integer not less than 1. We define the g-th sub-microphone signal as X g . Among them, g = 1, 2,..., g.

[0104] The computing device 240 can also perform a row permutation on the microphone signal X according to the data missing positions in each sub-microphone signal X g to make the data missing positions in all sub-microphone signals adjacent, as Figure 5C shown.

[0105] In summary, in the incomplete observation signal, the sub-microphone signal X g can be expressed by the following formula:

[0106] X g = P g S g + D g Formula (9)

[0107] Among them, P g = Q g P, S g = B g S. The matrices Q g , B g are matrices composed of 0 and 1 elements determined by the data missing positions.

[0108] The microphone signal X can be expressed by the following formula:

[0109] X = [X1, X2, …, X G Formula (10)

[0110] For the sake of convenience in description, in the following description, we will describe the microphone signal X as an incomplete observation signal.

[0111] As described above, the microphone 222 can collect both the noise signal D and the target voice signal. When there is no target voice signal in the microphone signal X, the microphone signal X satisfies the first model corresponding to the noise signal D. When there is a target voice signal in the microphone signal X, the microphone signal satisfies the second model corresponding to the mixture of the target voice signal and the noise signal D.

[0112] For the sake of convenience in description, we define the first model as the following formula:

[0113] X = D Formula (11)

[0114] When the microphone signal X is a complete observation signal, the first model can be expressed by the following formula:

[0115] x k = d k Formula (12)

[0116] When the microphone signal X is an incomplete observation signal, the first model can be expressed by the following formula:

[0117] X g = D g Formula (13)

[0118] We define the second model as the following formula:

[0119] X = PS + D Formula (14)

[0120] When the microphone signal X is a complete observation signal, the second model can be expressed by the following formula:

[0121] x k = Ps k + d k Formula (15)

[0122] When the microphone signal X is an incomplete observation signal, the second model can be expressed by the following formula:

[0123] X g = P g S g + D gFormula (16)

[0124] For the convenience of presentation, in the following description, we will take the microphone signal X as an example of an incomplete observation signal for description.

[0125] As Figure 3 shown, the method P100 may further include:

[0126] S140: Taking the maximization of the likelihood function and the minimization of the rank of the noise covariance matrix as the joint optimization objective, optimize the first model and the second model respectively, and determine the first estimate of the noise covariance matrix M1 of the first model and the second estimate of the noise covariance matrix M2 of the second model

[0127] There is a noise covariance matrix M of the unknown parameter noise signal D in the first model. For the convenience of description, we define the noise covariance matrix M of the unknown parameter noise signal D in the first model as M1. There is a noise covariance matrix M of the unknown parameter noise signal D and the amplitude S of the target speech signal in the second model. For the convenience of description, we define the noise covariance matrix M of the unknown parameter noise signal D in the second model as M2. The computing device 240 can optimize the first model and the second model respectively based on the optimization method, and determine the first estimate of the unknown parameter M1 the second estimate of M2 and the estimate of the amplitude S of the target speech signal

[0128] On the one hand, the computing device 240 can be triggered from the perspective of the likelihood function, and take the maximization of the likelihood function as the optimization objective to optimize and design the first model and the second model respectively. On the other hand, as mentioned above, the colored noise signal c k the corresponding noise covariance matrix M c has the low-rank property and is a low-rank positive semi-definite matrix. Therefore, the noise covariance matrix M of the noise signal d k also has the low-rank property. Especially for the incomplete observation signal, the low-rank property of the noise covariance matrix M of the noise signal d k still needs to be maintained during the rearrangement process of the incomplete observation signal. Therefore, the computing device 240 can be based on the noise signal d kBased on the low-rank property of the noise covariance matrix M, with the minimization of the rank of the noise covariance matrix M as the optimization objective, the first model and the second model are respectively optimized. Therefore, the computing device 240 can optimize the first model and the second model respectively with the maximization of the likelihood function and the minimization of the rank of the noise covariance matrix as the joint optimization objective to determine the first estimate of the unknown parameter M1 The second estimate of M2 And the estimated value of the amplitude S of the target speech signal

[0129] Figure 6 Shows a flowchart of an iterative optimization provided according to an embodiment of the present specification Figure 6 What is shown is step S140. As Figure 6 Shown, step S140 may include:

[0130] S142: Using the microphone signal X as sample data, establish the first likelihood function L1(M1) corresponding to the first model

[0131] The likelihood function includes the first likelihood function L1(M1). According to formulas (11)-(13), the first likelihood function L1(M1) can be expressed as the following formula:

[0132]

[0133] Among them, formula (17) respectively represents the first likelihood function L1(M1) under the complete observation signal and the incomplete observation signal Represents the maximum likelihood estimate of the parameter M1 And Represents that under the first model, given the parameter After that, the probability that the microphone signal X appears

[0134] S144: With the maximization of the first likelihood function L1(M1) and the minimization of the rank Rank(M1) of the noise covariance matrix M1 of the first model as the optimization objective, optimize the first model to determine the first estimate of M1

[0135] The maximization of the first likelihood function L1(M1) can be expressed as min(-log(L1(M1))). The minimization of the rank Rank(M1) of the noise covariance matrix M1 of the first model can be expressed as min(Rank(M1)). As mentioned before, we use the white noise signal n k Of the noise covariance matrix For description by taking a known example, according to formula (7), it can be known that minimizing the rank Rank(M1) of the noise covariance matrix M1 of the first model can be expressed as the noise covariance matrix M of the colored noise signal C c Minimize min(Rank(M c ))). Therefore, the objective function of the optimization objective can be expressed as the following formula:

[0136] min(-log(L1(M1)) + γRank(M c )) Formula (18)

[0137] where γ is the regularization coefficient. Since minimizing the matrix rank can be relaxed to a problem of minimizing the nuclear norm. Therefore, formula (18) can be expressed as the following formula:

[0138] min(-log(L1(M1)) + γ‖M c ‖ * ) Formula (19)

[0139] The iterative constraint condition of the first model can be expressed as the following formula:

[0140]

[0141] where M c ≥0 is the positive definiteness constraint of the noise covariance matrix M c of the colored noise signal C. The optimization problem of the first model can be expressed as the following formula:

[0142]

[0143] After determining the objective function and the constraint conditions, the computing device 240 can take the objective function as the optimization objective and iteratively optimize the unknown parameter M1 of the first model, so as to determine the first estimated value of the noise covariance matrix M1 of the first model

[0144] Formula (21) is a semi - definite programming problem, and the computing device 240 can solve it through various algorithms. For example, the gradient projection algorithm can be used. Specifically, in each step of the gradient projection algorithm iteration, we first solve formula (19) by the gradient method without any constraints, and then project the obtained solution onto the semi - definite cone to make it satisfy the matrix semi - positive definiteness constraint condition formula (20).

[0145] As Figure 6 shown, step S140 may further include:

[0146] S146: Taking the microphone signal X as sample data, establish the second likelihood function L2(S, M2) of the second model.

[0147] The likelihood function includes a second likelihood function L2(S, M2). According to formulas (14) to (16), the second likelihood function L2(S, M2) can be expressed as the following formula:

[0148]

[0149] Among them, formula (22) represents the second likelihood function under the complete observation signal and the incomplete observation signal respectively. Represents the maximum likelihood estimates of the parameters S and M2. And Respectively represent the probabilities of the microphone signal X appearing after the parameters S and M2 are given under the second model.

[0150] S148: Optimize the second model with the maximization of the second likelihood function L2(S, M2) and the minimization of the rank Rank(M2) of the noise covariance matrix M2 of the second model as the optimization objectives, and determine the second estimated value of M2 And the estimated value of the amplitude S of the target speech signal

[0151] The maximization of the second likelihood function L2(S, M2) can be expressed as min(-log(L2(S, M2))). The minimization of the rank Rank(M2) of the noise covariance matrix M2 of the second model can be expressed as min(Rank(M2)). As mentioned above, we use the white noise signal n k The noise covariance matrix of Known as an example for description. According to formula (7), it can be known that the minimization of the rank Rank(M2) of the noise covariance matrix M2 of the second model can be expressed as the minimization of the noise covariance matrix M of the colored noise signal C c Minimize min(Rank(M c ))). Therefore, the objective function of the optimization objective can be expressed as the following formula:

[0152] min(-log(L2(S, M2)) + γRank(M c )) Formula (23)

[0153] Among them, γ is the regularization coefficient. Since the minimization of the matrix rank can be relaxed to the minimization problem of the nuclear norm. Therefore, formula (23) can be expressed as the following formula:

[0154] min(-log(L2(S, M2)) + γ‖M c ‖ * ) Formula (24)

[0155] The iterative constraint conditions of the second model can be expressed as the following formula:

[0156]

[0157] where M c ≥ 0 is the positive definiteness constraint of the noise covariance matrix M c of the colored noise signal C. The optimization problem of the second model can be expressed as the following formula:

[0158]

[0159] After determining the objective function and the constraint conditions, the computing device 240 can iteratively optimize the unknown parameters M2 and S of the second model with the objective function as the optimization target, so as to determine the second estimated value of the noise covariance matrix M2 of the second model and the estimated value

[0160] of the amplitude S of the target speech signal. Formula (26) is a semi-definite programming problem, and the computing device 240 can solve it through various algorithms. For example, the gradient projection algorithm can be used. Specifically, in each iteration of the gradient projection algorithm, we first solve formula (24) by the gradient method without any constraints, and then project the obtained solution onto the semi-definite cone to make it satisfy the matrix semi-definiteness constraint condition formula (25).

[0161] In summary, the method P100 can jointly optimize the first model and the second model with the maximization of the likelihood function and the minimization of the rank of the noise covariance matrix as the joint optimization target, so as to determine the first estimated value of the unknown parameter M1 and the second estimated value of M2, so that the estimation accuracy of M1 and M2 is higher, providing a higher-precision data model for subsequent statistical hypothesis testing, thereby improving the accuracy of voice activity detection and the voice enhancement effect.

[0162] As Figure 3 shown, the method P100 may further include:

[0163] S160: Based on statistical hypothesis testing, determine the target model corresponding to the microphone signal X and the noise covariance matrix M.

[0164] The target model includes one of the first model and the second model. The noise covariance matrix M of the microphone signal X is the noise covariance matrix of the target model. When the target model of the microphone signal X is the first model, the noise covariance matrix of the microphone signal X. When the target model of the microphone signal X is the second model, the noise covariance matrix

[0165] The computing device 240 can determine whether the microphone signal X satisfies the first model or the second model based on the method of statistical hypothesis testing, so as to determine whether there is a target voice signal in the microphone signal X.

[0166] Figure 7 A flowchart of determining a target model provided according to an embodiment of the present specification is shown. Figure 7 The shown flowchart is step S160. As Figure 7 shown, step S160 may include:

[0167] S162: Based on the microphone signal X, establish a binary hypothesis testing model.

[0168] Among them, the null hypothesis H0 of the binary hypothesis testing model may be that there is no target voice signal in the microphone signal X, that is, the microphone signal X satisfies the first model. The alternative hypothesis H1 of the binary hypothesis testing model may be that there is a target voice signal in the microphone signal X, that is, the microphone signal X satisfies the second model. The binary hypothesis testing model can be expressed by the following formula:

[0169]

[0170]

[0171] Among them, the microphone signal X in formula (27) is a complete observation signal. The microphone signal X in formula (28) is an incomplete observation signal.

[0172] S164: Substitute the first estimated value the second estimated value and the estimated value of the amplitude S into the decision criterion of the detector of the binary hypothesis testing model to obtain a test statistic ψ.

[0173] The detector can be any one or more detectors. In some embodiments, the detector may be one or more of a GLRT detector, a Rao detector, and a Wald detector. In some embodiments, the detector may also be a u-detector, a t-detector, a χ2 detector (chi-square test), an F-detector, a rank sum detector, and so on. The test statistic ψ of different detectors is different.

[0174] Taking the GLRT detector (Generalized Likelihood Ratio Test) as an example for illustration. When the microphone signal X is a complete observation signal, in the GLRT detector, the test statistic ψ can be expressed as the following formula:

[0175]

[0176] where, and are the likelihood functions under the null hypothesis H0 and the alternative hypothesis H1 respectively.

[0177] When the microphone signal X is an incomplete observation signal, in the GLRT detector, the test statistic ψ can be expressed as the following formula:

[0178]

[0179] where, and are the likelihood functions under the null hypothesis H0 and the alternative hypothesis H1 respectively.

[0180] In the GLRT detector, it is necessary to estimate the unknown parameters under both the null hypothesis H0 and the alternative hypothesis H1, and there are many parameters to be estimated. While the Rao detector only needs to estimate the unknown parameter under the null hypothesis H0. When the number of frames K, the Rao test has the same detection performance as the GLRT detector. When the number of frames K is limited, although the Rao detector cannot achieve the same detection performance as the GLRT detector, it has the advantages of simpler calculation and being more suitable for the case where it is difficult to solve the unknown parameters under the alternative hypothesis H1.

[0181] Therefore, in view of the balanced requirements of the actual system for detection performance and computational complexity, the computing device 240 proposes a Rao detector based on the aforementioned GLRT detector. Taking the incomplete observation signal as an example, the test statistic ψ of the Rao detector can be expressed as the following formula:

[0182]

[0183] where, f(X1,X2,…,X G |θ,M) represents the probability density function under the alternative hypothesis H1. M = M2. θ r =[PS R,1 ,PS R,2 ,…,PS R,M, PS L,1 , PS L,2 , …, PS L,M ][ T 。Among them, PS R,m is the real part of the amplitude of the target speech signal in the audio signal of the m-th microphone 222. PS L,m is the imaginary part of the amplitude of the target speech signal in the audio signal of the m-th microphone 222. m = 1, 2, …, M.. θ r is a 2M-dimensional vector. Among them, θ s is a real vector containing redundant parameters. It includes the real and imaginary parts of the elements on the non-diagonal of M and the elements on the diagonal. Formula (31) can be simplified to the following formula:

[0184]

[0185] Among them,

[0186] In formula (32), as long as the estimator of the unknown parameter under the null hypothesis H0 can be obtained then the test statistic ψ of the Rao test can be obtained.

[0187] S166: Determine the target model of the microphone signal X based on the test statistic ψ.

[0188] Specifically, step S166 may include:

[0189] S166-2: Determine that the test statistic ψ is greater than the preset decision threshold η, determine that there is a target speech signal in the microphone signal X, determine the target model as the second model, and the noise covariance matrix of the microphone signal is the second estimated value Or

[0190] S166-4: Determine that the test statistic ψ is less than the preset decision threshold η, determine that there is no target speech signal in the microphone signal X, determine the target model as the first model, and the noise covariance matrix of the microphone signal is the first estimated value

[0191] Step S166 can be expressed as the following formula:

[0192]

[0193] The decision threshold η is a parameter related to the false alarm probability. The false alarm probability can be obtained through experiments, or through machine learning, or through experience.

[0194] For exampleFigure 3 As shown, the method P100 may further include:

[0195] S180: Output the target mode of the microphone signal X and the noise covariance matrix M.

[0196] The computing device 240 may output the target mode of the microphone signal X and the noise covariance matrix M to other computing modules, such as the voice enhancement module, etc.

[0197] In summary, in the voice activity detection system and method P100 provided in this specification, the computing device 240 may use maximizing the likelihood function and minimizing the rank of the noise covariance matrix as the joint optimization objectives to optimize the first model and the second model respectively, so as to determine the first estimate of the unknown parameter M1 and the second estimate of M2 so that the estimation accuracies of M1 and M2 are higher, providing a data model with higher accuracy for subsequent statistical hypothesis testing, thereby improving the accuracy of voice activity detection and the voice enhancement effect.

[0198] This specification also provides a voice enhancement system. The voice enhancement system may also be applied to the electronic device 200. In some embodiments, the voice enhancement system may include a computing device 240. In some embodiments, the voice enhancement system may be applied to the computing device 240. That is, the voice enhancement system may run on the computing device 240. The voice enhancement system may include a hardware device with data information processing capabilities and the necessary programs required to drive the hardware device to work. Of course, the voice enhancement system may also be only a hardware device with data processing capabilities, or only a program running in the hardware device.

[0199] The voice enhancement system may store data or instructions for executing the voice enhancement method described in this specification and may execute the data and / or instructions. When the voice enhancement system runs on the computing device 240, the voice enhancement system may obtain the microphone signal from the microphone array 220 based on the communication connection and execute the data or instructions of the voice enhancement method described in this specification. The voice enhancement method is introduced in other parts of this specification. For example, in Figure 8 the description introduces the voice enhancement method.

[0200] When the voice enhancement system runs on the computing device 240, the voice enhancement system is communicatively connected to the microphone array 220. The storage medium 243 may further include at least one instruction set stored in the data storage device for performing voice enhancement calculations on the microphone signals. The instructions are computer program codes, and the computer program codes may include programs, routines, objects, components, data structures, procedures, modules, etc. for executing the voice enhancement method provided in this specification. The processor 242 may read the at least one instruction set and execute the voice enhancement method provided in this specification according to the instructions of the at least one instruction set. The processor 242 may execute all steps included in the voice enhancement method.

[0201] Figure 8 The flowchart of the voice enhancement method P200 provided according to an embodiment of this specification is shown. The method P200 may perform voice enhancement on the microphone signals. Specifically, the processor 242 may execute the method P200. As Figure 8 shown, the method P200 may include:

[0202] S220: Obtain the microphone signal X output by the M microphones.

[0203] As described in step S120, it will not be elaborated here.

[0204] S240: Based on the voice activity detection method P100, determine the target model of the microphone signal X and the noise covariance matrix M of the microphone signal X.

[0205] The noise covariance matrix M of the microphone signal X is the noise covariance matrix of the target model. When the target model of the microphone signal X is the first model, the noise covariance matrix of the microphone signal X When the target model of the microphone signal X is the second model, the noise covariance matrix of the microphone signal X

[0206] S260: Based on the MVDR method and the noise covariance matrix M of the microphone signal X, determine the filtering coefficient ω corresponding to the microphone signal.

[0207] The filtering coefficient ω may be a vector of M×1 dimension. The filtering coefficient ω can be expressed by the following formula:

[0208] ω = [ω1, ω2, …, ω M H Formula (34)

[0209] where, the filtering coefficient corresponding to the m-th microphone 222 is ω m . m = 1, 2, …, M.​

[0210] The filtering coefficient ω can be expressed by the following formula:

[0211]

[0212] As described above, P is the target steering vector of the target voice signal. In some embodiments, P is known.

[0213] S280: Combine the microphone signal X based on the filtering coefficient, and output the target audio signal y k .

[0214] The target audio signal Y can be expressed by the following formula:

[0215] Y = ω H X Formula (36)

[0216] The computing device 240 can output the target audio signal Y to other electronic devices, such as a remote call device.

[0217] In summary, the voice activity detection system and method P100 and the voice enhancement system and method P200 provided in this specification are used for the microphone 220 array composed of multiple microphones 222. The voice activity detection system and method P100 and the voice enhancement system and method P200 can obtain the microphone signal X collected by the microphone array 220. The microphone signal X can be the first model corresponding to the noise signal or the second model corresponding to the mixture of the target voice signal and the noise signal. The voice activity detection system and method P100 and the voice enhancement system and method P200 can use the microphone signal X as a sample, and use the maximization of the likelihood function and the minimization of the rank of the noise covariance matrix M of the microphone signal X as the joint optimization objective, and optimize the first model and the second model respectively to determine the first estimate value of the noise covariance matrix M1 of the first model and the second estimate value of the noise covariance matrix M2 of the second model and determine whether the microphone signal X satisfies the first model or the second model through the method of statistical hypothesis testing, so as to determine whether there is a target voice signal in the microphone signal X, and determine the noise covariance matrix M of the microphone signal X, and then perform voice enhancement on the microphone signal X based on the MVDR method. The voice activity detection system and method P100 and the voice enhancement system and method P200 can make the estimation accuracy of the noise covariance matrix M and the accuracy of voice activity detection higher, and then improve the voice enhancement effect.

[0218] On the other hand, this specification provides a non-transitory storage medium storing at least one set of executable instructions for voice activity detection. When the executable instructions are executed by a processor, the executable instructions direct the processor to perform the steps of the voice activity detection method P100 described in this specification. In some possible implementation manners, various aspects of this specification can also be implemented in the form of a program product, which includes program code. When the program product runs on a computing device (such as computing device 240), the program code is used to cause the computing device to execute the voice activity detection steps described in this specification. The program product for implementing the above method can use a portable compact disc read-only memory (CD-ROM) to include the program code and can run on a computing device. However, the program product of this specification is not limited to this. In this specification, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system (such as processor 242). The program product can use any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, be but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code included on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification can be written in any combination of one or more programming languages, including object-oriented programming languages - such as Java, C++, etc., and also including conventional procedural programming languages - such as the "C" language or similar programming languages.The program code can be executed entirely on a computing device, partially on a computing device, executed as a stand-alone software package, partially on a computing device and partially on a remote computing device, or entirely on a remote computing device.

[0219] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0220] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and is not limiting. Although not explicitly stated herein, those skilled in the art will understand that this specification is intended to encompass various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.

[0221] In addition, certain terms in this specification have been used to describe embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.

[0222] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, figure, or its description. However, this does not mean that the combination of these features is necessary, and those skilled in the art may well extract some of these features as separate embodiments for understanding when reading this specification. That is to say, the embodiments in this specification can also be understood as an integration of multiple sub-embodiments. And it also holds when the content of each sub-embodiment contains fewer features than all the features of a single foregoing disclosed embodiment.

[0223] Each patent, patent application, published patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, items, etc., may be incorporated herein by reference. All of the content for all purposes, except any prosecution file history associated therewith, any same that may be inconsistent or in conflict with this document, or any same prosecution file history that may have a limiting effect on the broadest scope of the claims. Now or hereafter associated with this document. By way of example, if there is any inconsistency or conflict between the description, definition, and / or use of terms associated with any of the materials incorporated herein and the terms, descriptions, definitions, and / or used in this document, the terms in this document shall prevail.

[0224] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.

Claims

1. A voice activity detection method, characterized in that, For M microphones distributed in a preset array shape, where M is an integer greater than 1, including: Obtain microphone signals output by the M microphones, where the microphone signals satisfy the first model corresponding to no target voice signal or the second model corresponding to a target voice signal; Taking the maximization of the likelihood function and the minimization of the rank of the noise covariance matrix as a joint optimization objective, optimize the first model and the second model respectively, and determine the first estimate of the noise covariance matrix of the first model and the second estimate of the noise covariance matrix of the second model; and Based on statistical hypothesis testing, determine the target model and the noise covariance matrix corresponding to the microphone signals, where the target model includes one of the first model and the second model, and the noise covariance matrix of the microphone signals is the noise covariance matrix of the target model.

2. The voice activity detection method according to claim 1, characterized in that, The microphone signals include K consecutive frames of audio signals, where K is a positive integer greater than 1, and the microphone signals include an M×K data matrix.

3. The voice activity detection method according to claim 2, characterized in that The microphone signals are complete observation signals or incomplete observation signals. In the complete observation signals, all data in the M×K data matrix are complete. In the incomplete observation signals, some data in the M×K data matrix are missing. When the microphone signals are incomplete observation signals, obtaining the microphone signals output by the M microphones includes: Obtain the incomplete observation signals; Based on the data missing positions in each column of the M×K data matrix, perform row and column permutations on the microphone signals, and divide the microphone signals into at least one sub-microphone signal. The microphone signals include the at least one sub-microphone signal.

4. The voice activity detection method according to claim 1, wherein, Taking the maximization of the likelihood function and the minimization of the rank of the noise covariance matrix as a joint optimization objective, optimizing the first model and the second model respectively includes: Taking the microphone signals as sample data, establish a first likelihood function corresponding to the first model, where the likelihood function includes the first likelihood function; Taking the maximization of the first likelihood function and the minimization of the rank of the noise covariance matrix of the first model as the optimization objective, optimize the first model, and determine the first estimate; Taking the microphone signals as sample data, establish a second likelihood function of the second model, where the likelihood function includes the second likelihood function; and Taking the maximization of the second likelihood function and the minimization of the rank of the noise covariance matrix of the second model as the optimization objective, optimize the second model, and determine the second estimate and the amplitude estimate of the target voice signal.

5. The voice activity detection method according to claim 4, characterized in that The microphone signals include noise signals, and the noise signals follow a Gaussian distribution. The noise signals at least include: Colored noise signals, following a Gaussian distribution with zero mean, and the corresponding noise covariance matrix is a low-rank positive semi-definite matrix.

6. The voice activity detection method according to claim 1, characterized in that, The determining the target model and the noise covariance matrix corresponding to the microphone signals based on statistical hypothesis testing includes: Based on the microphone signal, a binary hypothesis testing model is established, wherein the null hypothesis of the binary hypothesis testing model includes that the microphone signal satisfies the first model, and the alternative hypothesis of the binary hypothesis testing model includes that the microphone signal satisfies the second model; Substitute the first estimated value, the second estimated value, and the amplitude estimated value into the decision criterion of the detector of the binary hypothesis testing model to obtain a test statistic; and Based on the test statistic, determine the target model of the microphone signal.

7. The voice activity detection method according to claim 6, characterized in that The determining the target model of the microphone signal based on the test statistic includes: Determine that the test statistic is greater than the preset decision threshold, determine that the target voice signal exists in the microphone signal, determine that the target model is the second model, and the noise covariance matrix of the microphone signal is the second estimated value; or Determine that the test statistic is less than the preset decision threshold, determine that the target voice signal does not exist in the microphone signal, determine that the target model is the first model, and the noise covariance matrix of the microphone signal is the first estimated value.

8. A voice activity detection system, characterized in that, including: At least one storage medium storing at least one instruction set for voice activity detection; and At least one processor communicatively connected to the at least one storage medium, wherein when the voice activity detection system runs, the at least one processor reads the at least one instruction set and implements the voice activity detection method according to any one of claims 1-7.

9. A voice enhancement method, characterized in that, For M microphones distributed in a preset array shape, where M is an integer greater than 1, including: Obtain the microphone signals output by the M microphones; Based on the voice activity detection method according to any one of claims 1-7, determine the target model of the microphone signal and the noise covariance matrix of the microphone signal, and the noise covariance matrix of the microphone signal is the noise covariance matrix of the target model; Based on the MVDR method and the noise covariance matrix of the microphone signal, determine the filter coefficients corresponding to the microphone signals; and Based on the filter coefficients, merge the microphone signals and output a target audio signal.

10. A voice enhancement system, characterized in that, including: At least one storage medium storing at least one instruction set for voice enhancement; and At least one processor communicatively connected to the at least one storage medium, wherein when the voice enhancement system runs, the at least one processor reads the at least one instruction set and implements the voice enhancement method according to claim 9.

Citation Information

Patent Citations

  • Speech enhancement method

    CN109087664A

  • Noise estimation for use with noise reduction and echo cancellation in personal communication

    US20140056435A1