Voice activity detection method and system, voice enhancement method and system
By optimizing the model in the micro microphone array to maximize the likelihood function and minimize the rank of the noise covariance matrix, the problem of inaccurate estimation of the noise covariance matrix in the prior art is solved, and the speech enhancement effect is improved.
Patent Information
- Application Number
- JP2023555858
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-11
AI Technical Summary
The prior art is difficult to accurately estimate the noise covariance matrix when using the Minimum Variance Distortionless Response (MVDR) algorithm for speech enhancement, especially in micro microphone arrays, resulting in low speech enhancement effect.
The target model and the noise covariance matrix are determined by optimizing the first and second models in M microphones distributed in the preset matrix to maximize the likelihood function and minimize the rank of the noise covariance matrix as a joint optimization goal.
The accuracy of speech activity detection and the estimation accuracy of noise covariance matrix are improved, thereby improving the speech enhancement effect of the MVDR algorithm, especially in micro microphone arrays.
Smart Images

Figure 0007675837000085 
Figure 0007675837000086 
Figure 0007675837000087
Abstract
Description
[Technical field]
[0001] The present specification relates to the technical field of target voice signal processing, in particular to a voice activity detection method and system, and a voice enhancement method and system. [Background technology]
[0002] In a voice enhancement technique based on a beamforming algorithm, especially in an adaptive beamforming algorithm of Minimum Variance Distortionless Response (abbreviated as MVDR), how to solve the parameter that describes the relationship of the statistical characteristics of noise between different microphones, the noise covariance matrix, is extremely important. The main method in the prior art is to calculate the noise covariance matrix based on the method of voice presence probability, for example, estimating the voice presence probability by a voice activity detection method (abbreviated as Voice Activity Detection (VAD)), and then calculating the noise covariance matrix. However, the estimation accuracy rate of the voice presence probability in the prior art is not sufficient, so the estimation accuracy of the noise covariance matrix is low, and the voice enhancement effect of the MVDR algorithm is low. In particular, when the number of microphones is small, for example, less than five, the effect drops sharply. Therefore, the MVDR algorithm in the prior art is often used in microphone array devices with a large number of microphones and large spacing, such as mobile phones and smart speakers, but the voice enhancement effect is low in devices with a small number of microphones and small spacing, such as earphones.
[0003] Therefore, there is a need to provide more accurate voice activity detection methods and systems, and speech enhancement methods and systems. Summary of the Invention
[0004] The present specification provides more accurate voice activity detection methods and systems, and speech enhancement methods and systems.
[0005] According to a first aspect, the present specification provides a voice activity detection method, for use with M microphones distributed in a pre-defined array shape, where M is an integer greater than 1, the method comprising: obtaining microphone signals output by the M microphones, where either there is no first model corresponding to a target speech signal or there is a second model corresponding to a target speech signal; optimizing the first model and the second model respectively using a joint optimization objective of maximizing a likelihood function and minimizing a rank of a noise covariance matrix, to determine a first estimate of a noise covariance matrix of the first model and a second estimate of a noise covariance matrix of the second model; and determining a target model and a noise covariance matrix corresponding to the microphone signals based on a statistical hypothesis test, where the target model comprises one of the first model and the second model, and the noise covariance matrix of the microphone signals is the noise covariance matrix of the target model.
[0006] In some embodiments, the microphone signal comprises K frames of consecutive audio signals, where K is a positive integer greater than 1, and the microphone signal comprises an M×K data matrix.
[0007] In some embodiments, the microphone signal is a full observation signal or a non-full observation signal, in which all data in the M×K data matrix is complete in the full observation signal and some data in the M×K data matrix is missing in the non-full observation signal, and when the microphone signal is the non-full observation signal, acquiring the microphone signals output by the M microphones includes acquiring the non-full observation signal, and performing row permutation and column permutation on the microphone signal based on the data missing positions in each column of the M×K data matrix, and splitting the microphone signal into at least one sub-microphone signal, and the microphone signal includes the at least one sub-microphone signal.
[0008] In some embodiments, optimizing the first model and the second model, respectively, with a joint optimization objective of maximizing a likelihood function and minimizing the rank of a noise covariance matrix includes: establishing a first likelihood function included in the likelihood function corresponding to the first model, using the microphone signal as sample data; optimizing the first model with an optimization objective of maximizing the first likelihood function and minimizing the rank of the noise covariance matrix of the first model, and determining the first estimate; determining a second likelihood function included in the likelihood function of the second model, using the microphone signal as sample data; and optimizing the second model with an optimization objective of maximizing the second likelihood function and minimizing the rank of the noise covariance matrix of the second model, and determining the second estimate and an amplitude estimate of the target speech signal.
[0009] In some embodiments, the microphone signals include noise signals following a Gaussian distribution, the noise signals including at least a colored noise signal following a zero-mean Gaussian distribution and a corresponding noise covariance matrix which is a low-rank positive semidefinite matrix.
[0010] In some embodiments, determining a target model and a noise covariance matrix corresponding to the microphone signal based on the statistical hypothesis testing includes establishing a binary hypothesis testing model based on the microphone signal, where a null hypothesis of the binary hypothesis testing model includes that the microphone signal satisfies the first model and an alternative hypothesis of the binary hypothesis testing model includes that the microphone signal satisfies the second model; substituting the first estimate, the second estimate and the amplitude estimate into a detector criterion of the binary hypothesis testing model to obtain a test statistic; and determining the target model of the microphone signal based on the test statistic.
[0011] In some embodiments, determining the target model of the microphone signal based on the test statistic includes determining that the test statistic is greater than the predetermined decision threshold, determining that the target speech signal is present in the microphone signal, and determining that the target model is the second model and that a noise covariance matrix of the microphone signal is the second estimate, or determining that the test statistic is less than the predetermined decision threshold, determining that the target speech signal is not present in the microphone signal, and determining that the target model is the first model and that a noise covariance matrix of the microphone signal is the first estimate.
[0012] In some embodiments, the detector includes at least one of a GLRT detector, a Rao checker, and a Wald checker.
[0013] According to a second aspect, the present specification further provides a voice activity detection system, the system comprising at least one storage medium and at least one processor, the at least one storage medium having stored therein at least one instruction set for voice activity detection, the at least one processor being communicatively connected to the at least one storage medium, wherein when the voice activity detection system is operational, the at least one processor reads the at least one instruction set and performs the voice activity detection method as described in the first aspect of the present specification.
[0014] According to a third aspect, the present specification further provides a speech enhancement method for M microphones distributed in a pre-defined array shape, where M is an integer greater than 1, the method comprising: obtaining microphone signals output by the M microphones; determining a target model of the microphone signals and a noise covariance matrix of the microphone signals, which is a noise covariance matrix of the target model, based on a voice activity detection method according to any one of claims 1 to 8; determining filtering coefficients corresponding to the microphone signals based on an MVDR method and the noise covariance matrix of the microphone signals; and integrating the microphone signals based on the filtering coefficients to output a target audio signal.
[0015] According to a fourth aspect, the present specification further provides a voice enhancement system, the system including at least one storage medium and at least one processor, the at least one storage medium storing at least one instruction set for performing voice enhancement, the at least one processor being communicatively connected to the at least one storage medium, wherein when the voice enhancement system is operating, the at least one processor reads the at least one instruction set and performs the voice enhancement method described in the third aspect of the present specification.
[0016] As can be seen from the above technical solutions, the voice activity detection method and system, the speech enhancement method and system according to the present specification are used in a microphone array consisting of multiple microphones, where the microphone signal output by the microphone array satisfies a first model corresponding to a noise signal or a second model corresponding to a combination of a target speech signal and the noise signal. In order to obtain whether a target speech signal exists in the microphone signal, the method and system respectively optimize the first model and the second model using the joint optimization objectives of maximizing the likelihood function and minimizing the rank of the noise covariance matrix, determine a first estimate of the noise covariance matrix of the first model and a second estimate of the noise covariance matrix of the second model, and determine whether the target speech signal exists in the microphone signal by using a statistical hypothesis testing method to determine whether the microphone signal satisfies the first model or the second model, determine the noise covariance matrix of the microphone signal, and perform speech enhancement on the microphone signal based on the MVDR method. The method and system can improve the estimation accuracy of the noise covariance and further improve the speech enhancement effect.
[0017] Other features of the voice activity detection method, system, speech enhancement method and system according to the present specification are described in part in the following description. According to the description, the contents shown in the following figures and examples will be self-evident to those skilled in the art. The creative aspects of the voice activity detection method, system, speech enhancement method and system according to the present specification can be fully understood by practice or use of the methods, apparatus and combinations described in the following detailed examples. [Brief description of the drawings]
[0018] In order to more clearly describe the technical solutions in the embodiments of the present specification, the following provides a brief description of the drawings that need to be used in the description of the embodiments. It should be apparent that the drawings in the following description are only some embodiments of the present specification, and those skilled in the art can obtain other drawings based on these drawings without any creative efforts. [Figure 1] 1 is a hardware schematic diagram of a voice activity detection system according to an embodiment of the present disclosure. [Figure 2A] 1 is a schematic exploded view of an electronic device according to an embodiment of the present specification; [Figure 2B] FIG. 2 is a front view of a first case according to an embodiment of the present disclosure. [Figure 2C] FIG. 2 is a plan view of a first case according to an embodiment of the present disclosure. [Figure 2D] FIG. 13 is a front view of a second case according to an embodiment of the present disclosure. [Figure 2E] FIG. 13 is a bottom view of a second case according to an embodiment of the present disclosure. [Diagram 3] 1 is a flow chart of a voice activity detection method according to an embodiment of the present disclosure; [Figure 4] FIG. 2 is a schematic diagram of a full observation signal according to an embodiment of the present disclosure. [Figure 5A] FIG. 2 is a schematic diagram of a non-full observation signal according to an embodiment of the present disclosure. [Figure 5B] FIG. 2 is a schematic diagram of rearrangement of non-full observed signals according to an embodiment of the present disclosure; [Figure 5C] FIG. 2 is a schematic diagram of rearrangement of non-full observed signals according to an embodiment of the present disclosure; [Figure 6] 1 is a flow chart of an iterative optimization according to an embodiment of the present disclosure. [Figure 7] 1 is a flow chart of determining a target model according to an embodiment of the present disclosure. [Figure 8] 1 is a flowchart of a speech enhancement method according to an embodiment of the present disclosure; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0019] The following description provides specific application scenarios and requirements of the present specification to enable those skilled in the art to make and use the contents of the present specification. Various partial modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the present specification. Therefore, the present specification is not intended to be limited to the embodiments shown, but is accorded the widest scope consistent with the claims.
[0020] The terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting. For example, the singular forms "a", "an" and "the" as used herein can also include the plural unless the context clearly dictates otherwise. As used herein, the terms "comprise", "include", and / or "containing" refer to the presence of associated integers, steps, operations, elements and / or components, but do not exclude the presence of one or more other features, integers, steps, operations, elements, components and / or groups, or that other features, integers, steps, operations, elements, components and / or groups may be added to the system / method.
[0021] These and other features of the present specification, as well as the operation and function of the associated elements of structure, and the combination of parts and economies of manufacture, can be clearly improved upon in consideration of the following description. Reference is made to the drawings, all of which form a part of this specification. It is to be expressly understood, however, that the drawings are for the purpose of illustration and description only and are not intended to limit the scope of this specification. It is also to be understood that the drawings are not drawn to scale.
[0022] As used herein, flowcharts illustrate operations of a system implementation according to some embodiments of the present invention. It should be clearly understood that the operations of the flowcharts may be implemented out of order. Conversely, the operations may be implemented in reverse order or simultaneously. It should be noted that one or more other operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.
[0023] To facilitate the explanation, the terms appearing in this specification are first explained as follows.
[0024] <Statistical hypothesis testing> It is a mathematical statistical method to estimate a population from a sample based on a certain hypothetical condition. The specific method is as follows: According to the needs of the problem, some hypothesis is made about the population to be studied, the null hypothesis is written as H_0, and when the null hypothesis H_0 is established, an appropriate statistic is selected so that its distribution is known, the value of the statistic is calculated from an actual sample, and it is tested based on a pre-given significance level to determine whether to reject or accept the null hypothesis H_0. Common statistical hypothesis testing methods include u-test, t-test, chi-square test, F-test, rank sum test, etc.
[0025] <Minimum Variance Distortionless Response (abbreviated as MVDR)> An adaptive beamforming algorithm based on the maximum signal-to-interference-and-noise ratio (SINR) criterion, the MVDR algorithm can adaptively minimize the power of the array output in the desired direction while maximizing the signal-to-interference-and-noise ratio. The goal is to minimize the variance of the recorded signal. If the noise signal and the desired signal are uncorrelated, the variance of the recorded signal is the sum of the variances of the desired and noise signals. Therefore, the MVDR solution seeks to reduce the effect of the noise signal by minimizing this sum. The principle is to select appropriate filter coefficients to minimize the average power of the array output under the constraint that the desired signal is undistorted.
[0026] <Voice Activity Detection> This is a processing procedure for dividing a target speech signal into speech segments and non-speech segments.
[0027] <Gaussian distribution> Normal distribution, also known as "stationary distribution" and also known as Gaussian distribution, is a bell-shaped curve, low at both ends, high in the middle, symmetrical, and often called a bell curve. A random variable X has an expectation of μ and a variance of σ. 2 If the normal distribution is 2 ) The desired value μ for a probability density function of a normal distribution determines its location, and its standard deviation σ determines the amplitude of the distribution. When μ=0 and σ=1, the normal distribution is the standard normal distribution.
[0028] 1 shows a hardware schematic diagram of a voice activity detection system according to an embodiment of the present disclosure. The voice activity detection system may be used in an electronic device 200.
[0029] In some embodiments, the electronic device 200 may be a device with audio processing capabilities, such as a wireless earphone, a wired earphone, a smart wearable device, such as a smart glass, a smart helmet, or a smart watch. The electronic device 200 may also be a mobile device, a tablet computer, a laptop, an in-car device, or the like, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, or the like, or any combination thereof. For example, the smart mobile device may include a mobile phone, a personal digital assistant, a gaming device, a navigation device, an ultra-mobile personal computer (UMPC), or the like, or any combination thereof. In some embodiments, the smart home device may include a smart television, a desktop computer, or the like, or any combination thereof. In some embodiments, the in-car device may include an in-car computer, an in-car television, or the like.
[0030] In this specification, the inventors take the electronic device 200 as an example of an earphone. The earphone may be a wireless earphone or a wired earphone. As shown in FIG. 1, the electronic device 200 may include a microphone array 220 and a computing device 240.
[0031] The microphone array 220 may be an audio collecting device of the electronic device 200. The microphone array 220 may be configured to acquire local audio and output a microphone signal, i.e., an electronic signal with audio information. The microphone array 220 may include M microphones 222 distributed in a predefined array shape, where M is an integer greater than 1. The M microphones 222 may be uniformly or non-uniformly distributed. The M microphones 222 may output microphone signals. The M microphones 222 may output M microphone signals. Each microphone 222 corresponds to one microphone signal. The M microphone signals are collectively referred to as the microphone signals. In some embodiments, the M microphones 222 may be linearly distributed. In some embodiments, the M microphones 222 may be distributed as an array of other shapes, for example, a circular array, a rectangular array, etc. For ease of explanation, in the following description, the inventors will take the M microphones 222 as an example of being linearly distributed. In some embodiments, M may be any integer greater than 1, such as 2, 3, 4, 5, or more. In some embodiments, due to space constraints, M may be an integer greater than 1 and less than or equal to 5, for example, in products such as earphones. When the electronic device 200 is an earphone, the spacing between adjacent microphones 222 among the M microphones 222 may be 20 mm to 40 mm. In some embodiments, the spacing between adjacent microphones 222 may be smaller, such as 10 mm to 20 mm.
[0032] In some embodiments, the microphone 222 may be a bone conduction microphone that directly collects human body vibration signals. The bone conduction microphone may include a vibration sensor, such as an optical vibration sensor, an acceleration sensor, etc. The vibration sensor can collect mechanical vibration signals (e.g., signals due to vibrations generated by the skin or bones when the user is speaking) and convert the mechanical vibration signals into electrical signals. The mechanical vibration signals here mainly refer to vibrations that propagate through solid bodies. The bone conduction microphone collects vibration signals generated by the skin or bones when the user speaks by contacting the skin or bones of the user through the vibration sensor or a vibration component connected to the vibration sensor, and converts the vibration signals into electrical signals. In some embodiments, the vibration sensor may be a device that is sensitive to mechanical vibrations but not air vibrations (i.e., the response capability of the vibration sensor to mechanical vibrations exceeds the response capability of the vibration sensor to air vibrations). The bone conduction microphone can directly pick up vibration signals from the speaking part, thereby reducing the influence of environmental noise.
[0033] In some embodiments, the microphone 222 may be an air conduction microphone that directly collects air vibration signals generated when a user speaks and converts the air vibration signals into electrical signals.
[0034] In some embodiments, the M microphones 222 may be M bone conduction microphones. In some embodiments, the M microphones 222 may be M air conduction microphones. In some embodiments, the M microphones 222 may include bone conduction microphones or air conduction microphones. Of course, the microphones 222 may be other types of microphones, such as optical microphones, microphones that receive myoelectric signals, etc.
[0035] The computing device 240 may be communicatively connected to the microphone array 220. The communicative connection refers to any form of connection that can receive information directly or indirectly. In some embodiments, the computing device 240 can communicate data with the microphone array 220 via a wireless communication connection, in some embodiments, the computing device 240 can also communicate data with the microphone array 220 directly connected to the microphone array 220 by wires, and in some embodiments, the computing device 240 can also directly connect to other circuits by wires to establish an indirect connection with the microphone array 220 to realize data communication between them. In this specification, the computing device 240 is directly connected to the microphone array 220 by wires as an example.
[0036] The computing device 240 may be a hardware device having data processing capabilities. In some embodiments, the voice activity detection system may include the computing device 240. In some embodiments, the voice activity detection system may be used in the computing device 240, i.e., the voice activity detection system may run on the computing device 240. The voice activity detection system may include a hardware device having data processing capabilities and a program necessary to drive the operation of the hardware device. Of course, the voice activity detection system may be only a hardware device having data processing capabilities, or only a program running on the hardware device.
[0037] The voice activity detection system may store data or instructions and may execute the data and / or instructions to perform the voice activity detection methods described herein. When the voice activity detection system operates on a computing device 240, the voice activity detection system may obtain the microphone signals from the microphone array 220 based on the communication connection, execute data or instructions of the voice activity detection methods described herein, and calculate whether a target voice signal is present in the microphone signals. The voice activity detection methods are introduced elsewhere herein. For example, the voice activity detection methods are introduced in the description of Figures 3-8.
[0038] 1, the computing device 240 may include at least one storage medium 243 and at least one processor 242. In some embodiments, the electronic device 200 may further include a communication port 245 and an internal communication bus 241.
[0039] The internal communication bus 241 may connect to different system components including a storage medium 243 , a processor 242 , and a communication port 245 .
[0040] The communication port 245 can be used for data communication between the computing device 240 and the outside world. For example, the computing device 240 can obtain the microphone signals from the microphone array 220 via the communication port 245.
[0041] At least one storage medium 243 may include a data storage device. The data storage device may be a non-transitory or a transitory storage medium. For example, the data storage device may include one or more of a magnetic disk, a read-only storage medium (ROM), or a random access storage medium (RAM). When the voice activity detection system is operable on a computing device 240, the storage medium 243 may further include at least one set of instructions stored in the data storage device for performing voice activity detection on the microphone signal. The instructions may be computer program code, and the computer program code may include programs, routines, objects, components, data structures, processes, modules, etc., for performing the voice activity detection method according to the present disclosure.
[0042] The at least one processor 242 can be communicatively connected to the at least one storage medium 243 via an internal communication bus 241. The communicative connection refers to any form of connection capable of receiving information directly or indirectly. The at least one processor 242 is for executing the at least one instruction set. When the voice activity detection system is operable on the computing device 240, the at least one processor 242 reads the at least one instruction set and executes the voice activity detection method according to the present disclosure according to the instructions of the at least one instruction set. The processor 242 can execute all steps included in the voice activity detection method. The processor 242 may be in the form of one or more processors, and in some embodiments, the processor 242 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), a special-purpose integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physical processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, or any combination thereof. For the sake of illustrative purposes only, only one processor 242 is described herein for the computing device 240. However, it should be noted that the computing device 240 herein may further include multiple processors 242, such that the operations and / or method steps disclosed herein may be performed by one processor or jointly by multiple processors as described herein.For example, where in this specification a processor 242 of a computing device 240 performs steps A and B, it should be understood that steps A and B may be performed jointly or separately by two different processors 242 (e.g., a first processor performs step A and a second processor performs step B, or the first and second processors perform steps A and B jointly).
[0043] 2A illustrates an exploded structural schematic diagram of an electronic device 200 according to an embodiment of the present disclosure. As shown in FIG. 2A, the electronic device 200 may include a microphone array 220, a computing device 240, a first case 260, and a second case 280.
[0044] The first case 260 may be a mounting substrate for the microphone array 220. The microphone array 220 may be mounted inside the first case 260. The shape of the first case 260 may be adaptively designed according to the distribution shape of the microphone array 220, and this specification does not limit this too much. The second case 280 may be a mounting substrate for the computing device 240. The computing device 240 may be mounted inside the second case 280. The shape of the second case 280 may be adaptively designed according to the shape of the computing device 240, and this specification does not limit this too much. If the electronic device 200 is an earphone, the second case 280 may be connected to the wearing site. The second case 280 may be connected to the first case 260. As described above, the microphone array 220 may be electrically connected to the computing device 240. Specifically, the microphone array 220 can realize an electrical connection with the computing device 240 through the connection between the first case 260 and the second case 280.
[0045] In some embodiments, the first case 260 may be fixedly connected to the second case 280 by integral molding, welding, crimping, adhesive, or the like. In some embodiments, the first case 260 may be removably connected to the second case 280. The computing device 240 may be communicatively connected to different microphone arrays 220. Specifically, different microphone arrays 220 may differ in the number of microphones 222 in the microphone array 220, the array shape, the spacing between the microphones 222, the mounting angle of the microphone array 220 in the first case 260, the mounting position of the microphone array 220 in the first case 260, and the like. The wearer can replace the corresponding microphone array 220 according to different application scenarios to apply the electronic device 200 to a wider range of scenarios. For example, when the distance between the wearer and the electronic device 200 is short in the application scenario, the wearer can replace the microphone array 220 with a microphone array 220 with a smaller spacing. Further, for example, if the distance between the wearer and the electronic device 200 is long in an application scenario, the wearer can switch to a microphone array 220 with a larger number of microphones spaced apart.
[0046] The detachable connection may be any form of physical connection, such as a screw connection, a snap connection, a magnetic attraction connection, etc. In some embodiments, the first case 260 and the second case 280 may be connected by a magnetic attraction, i.e., the first case 260 and the second case 280 are detachably connected by the attraction force of a magnetic device.
[0047] 2B shows a front view of the first case 260 according to an embodiment of the present specification, and FIG. 2C shows a plan view of the first case 260 according to an embodiment of the present specification. As shown in FIG. 2B and FIG. 2C, the first case 260 may include a first interface 262. In some embodiments, the first case 260 may further include a touch point 266. In some embodiments, the first case 260 may further include an angle sensor (not shown in FIG. 2B and FIG. 2C).
[0048] The first interface 262 may be a mounting interface of the first case 260 and the second case 280. In some embodiments, the first interface 262 may be circular. The first interface 262 may be rotatably connected to the second case 280. When the first case 260 is mounted on the second case 280, the angle of the microphone array 220 can be adjusted by rotating the first case 260 relative to the second case 280 and adjusting the angle of the first case 260 relative to the second case 280.
[0049] A first magnetic device 263 may be installed on the first interface 262. The first magnetic device 263 may be installed at a position close to the second case 280 of the first interface 262. The first magnetic device 263 may realize a detachable connection with the second case 280 by generating a magnetic attraction force. When the first case 260 approaches the second case 280, the attraction force quickly connects the first case 260 to the second case 280. In some embodiments, after the first case 260 is connected to the second case 280, the first case 260 can still rotate relative to the second case 280, thereby adjusting the angle of the microphone array 220. Due to the action of the attraction force, the connection between the first case 260 and the second case 280 can be maintained even if the first case 260 rotates relative to the second case 280.
[0050] In some embodiments, the first interface 262 may further include a first positioning device (not shown in FIGS. 2B and 2C ). The first positioning device may be a positioning step protruding outward or a positioning hole extending inward. The first positioning device may be engaged with the second case 280 to realize quick mounting of the first case 260 and the second case 280.
[0051] As shown in FIG. 2B and FIG. 2C, in some embodiments, the first case 260 may further include a touch point 266. The touch point 266 may be mounted at the first interface 262. The touch point 266 may protrude outward from the first interface 262. The touch point 266 may be elastically connected to the first interface 262. The touch point 266 may be communicatively connected to the M microphones 222 in the microphone array 220. The touch point 266 may be made of elastic metal to realize data transmission. When the first case 260 is connected to the second case 280, the microphone array 220 may realize a communicative connection with the computing device 240 through the touch point 266. In some embodiments, the touch points 266 may be distributed in a circular shape. After the first case 260 is connected to the second case 280, when the first case 260 rotates relative to the second case 280, the touch points 266 can also rotate relative to the second case 280 and maintain a communication connection with the computing device 240.
[0052] In some embodiments, an angle sensor (not shown in FIGS. 2B and 2C ) may further be installed on the first case 260. The angle sensor may be communicatively connected to the touch point 266, thereby achieving a communicative connection with the computing device 240. The angle sensor may collect angle data of the first case 260 to determine the angle at which the microphone array 220 is located, and provide reference data for subsequent calculation of the voice presence probability.
[0053] 2D illustrates a front view of second case 280 according to an embodiment of the present disclosure, and FIG. 2E illustrates a bottom view of second case 280 according to an embodiment of the present disclosure. As shown in FIG. 2D and FIG. 2E, second case 280 may include a second interface 282. In some embodiments, second case 280 may further include a guide rail 286.
[0054] The second interface 282 may be a mounting interface of the second case 280 and the first case 260. In some embodiments, the second interface 282 may be circular. The second interface 282 may be rotatably connected to the first interface 262 of the first case 260. When the first case 260 is mounted on the second case 280, the angle of the microphone array 220 can be adjusted by rotating the first case 260 relative to the second case 280 and adjusting the angle of the first case 260 relative to the second case 280.
[0055] A second magnetic device 283 may be installed on the second interface 282. The second magnetic device 283 may be installed at a position of the second interface 282 close to the first case 260. The second magnetic device 283 can realize a detachable connection with the first interface 262 by generating a magnetic attraction force. The second magnetic device 283 can be used by engaging with the first magnetic device 263. When the first case 260 approaches the second case 280, the attraction force between the second magnetic device 283 and the first magnetic device 263 can quickly mount the first case 260 on the second case 280. When the first case 260 is mounted on the second case 280, the second magnetic device 283 faces the position of the first magnetic device 263. In some embodiments, after the first case 260 is connected to the second case 280, the first case 260 can still rotate relative to the second case 280, thereby adjusting the angle of the microphone array 220. Due to the action of the suction force, the connection between the first case 260 and the second case 280 can be maintained even if the first case 260 rotates relative to the second case 280.
[0056] In some embodiments, a second positioning device (not shown in FIGS. 2D and 2E ) may be further installed on the second interface 282. The second positioning device may be a positioning step protruding outward or a positioning hole extending inward. The second positioning device may engage with the first positioning device of the first case 260 to realize quick mounting of the first case 260 and the second case 280. When the first positioning device is the positioning step, the second positioning device may be the positioning hole. When the first positioning device is the positioning hole, the second positioning device may be the positioning step.
[0057] As shown in FIG. 2D and FIG. 2E, in some embodiments, the second case 280 may further include a guide rail 286. The guide rail 286 may be mounted at the second interface 282. The guide rail 286 may be communicatively connected to the computing device 240. The guide rail 286 may be made of a metal material to realize data transmission. When the first case 260 is connected to the second case 280, the touch point 266 may contact the guide rail 286 to form a communicative connection, thereby realizing the communicative connection between the microphone array 220 and the computing device 240 and realizing data transmission. As mentioned above, the touch point 266 may be elastically connected to the first interface 262. Therefore, after the first case 260 is connected to the second case 280, the elastic action of the elastic connection can make the touch point 266 fully contact the guide rail 286 to realize a reliable communicative connection. In some embodiments, the guide rail 286 may be distributed in a circular shape. After the first case 260 is connected to the second case 280, when the first case 260 rotates relative to the second case 280, the touch points 266 can also rotate relative to the guide rails 286 and maintain a communication connection with the guide rails 286.
[0058] 3 shows a flowchart of a voice activity detection method P100 according to an embodiment of the present disclosure. The method P100 can calculate whether a target voice signal is present in the microphone signal. Specifically, the processor 242 can execute the method P100.
[0059] As shown in FIG. 3, the method P100 may include the following steps. S120: The microphone signals output by the M microphones 222 are acquired.
[0060] As described above, each microphone 222 can output a corresponding microphone signal. The M microphones 222 correspond to the M microphone signals. When the method P100 calculates whether the target voice signal exists in the microphone signals, the calculation may be based on all of the M microphone signals or on some of the microphone signals. Therefore, the microphone signals may include M microphone signals corresponding to the M microphones 222 or some of the microphone signals. In the following description of this specification, it is taken as an example that the microphone signals may include M microphone signals corresponding to the M microphones 222.
[0061] In some embodiments, the microphone signal may be a time-domain signal. In some embodiments, in step S120, the calculation device 240 may perform frame division and windowing on the microphone signal to divide the microphone signal into a plurality of consecutive audio signals. In some embodiments, in step S120, the calculation device 240 may further perform a time-frequency transform on the microphone signal to obtain a frequency-domain signal of the microphone signal. For ease of description, we label the microphone signal of an arbitrary frequency point as X. In some embodiments, the microphone signal X may include K frames of consecutive audio signals, where K is any positive integer greater than 1. For ease of description, we label the microphone signal of the kth frame as x. k The microphone signal x of the kth frame is k may be expressed as follows:
[0062]
number
[0063] Microphone signal x at the kth frame k may be an M-dimensional signal vector consisting of M microphone signals. The microphone signal X may be represented by an M×K data matrix. The microphone signal X may be represented by the following equation:
[0064]
number
[0065] Here, the microphone signal X is a data matrix of M×K, in which the m-th row represents the microphone signal received by the m-th microphone, and the k-th column represents the microphone signal of the k-th frame.
[0066] As described above, the microphone 222 can collect noise from the surrounding environment and output a noise signal, and can also collect the voice of the target user and output the target voice signal. When the target user does not make a sound, the microphone signal includes only the noise signal. When the target user makes a sound, the microphone signal includes the target voice signal and the noise signal. The microphone signal x of the kth frame k may be expressed as follows:
[0067]
number
[0068] where k = 1, 2, . . . , K. d k is the microphone signal x of the kth frame. k is the noise signal at s k is the amplitude of the target audio signal. P is the target steering vector of the target audio signal.
[0069] The microphone signal X may be expressed as:
number
[0070] Here, S is the amplitude of the target audio signal. S=[s1, s2, . . . , s K ]. D is the noise signal. D=[d1,d2,...,d K ].
[0071] Noise signal d k may be expressed as follows:
number
[0072] Microphone signal x at the kth frame k Noise signal d kmay be an M-dimensional signal vector consisting of M microphone signals.
[0073] In some embodiments, the noise signal d k is at least the colored noise signal c k In some embodiments, the noise signal d k is the white noise signal n k The noise signal d k may be expressed as follows:
number
[0074] If so, then the noise signal D = C + N, where C is the colored noise signal and C = [c1, c2, , c K ]. N is a white noise signal, and N=[n1,n2,...,n K ].
[0075] The computing device 240 calculates the noise signal d k A parameterized clustering model is established by using the cluster feature of the sound source spatial distribution of the noise signal d and the unified mapping relationship between the parameters of the microphone array 220. k The noise signal d k The colored noise signal C k and a white noise signal n k It can be divided into:
[0076] In some embodiments, the noise signal D follows a Gaussian distribution. k ~CN(0,M), where M is the noise signal d k where the noise covariance matrix of the colored noise signal c k follows a zero-mean Gaussian distribution, i.e., c k ~CN(0,M c ). Colored noise signal c k The noise covariance matrix Mc is a low-rank positive semidefinite matrix with the low-rank property. k also follows a zero-mean Gaussian distribution. That is, n k ~CN(0,M n ). White noise signal n k The power of δ0 2 It is. n = δ0 2 I n That is, n k ~CN(0,δ0 2 ). Noise signal d k The noise covariance matrix M of may be expressed as follows:
number
[0077] Noise signal d k The noise covariance matrix M of n and a low-rank positive semidefinite matrix M c It can be decomposed into the sum of
[0078] In some embodiments, the computing device 240 includes a white noise signal n k Power of δ0 2 In some embodiments, the white noise signal n k Power of δ0 2 For example, the calculation device 240 may estimate the white noise signal n k Power of δ0 2 In some embodiments, the computing device 240 can estimate the white noise signal n based on the method P100. k Power of δ0 2 can be estimated.
[0079] s kis the complex amplitude of the target sound signal. In some embodiments, there is one target sound signal source around the microphone 222. In some embodiments, there are L target sound signal sources around the microphone 222. In this case, s k may be an L×1 dimensional vector.
[0080] The target steering vector P is a matrix with dimensions M×L. The target steering vector P may be expressed by the following equation:
number
[0081] where f0 is the carrier frequency; d is the distance between adjacent microphones 222; c is the speed of sound; θ1, . . . , θ N are the angles of incidence between the L target sound signal sources and the microphone 222, respectively. In some embodiments, the target sound signal sources s k The angles of θ1, θ2, θ3, θ4, θ5, θ6, θ7, θ8, θ9, θ10, θ11, θ12, θ13, θ14, θ15, θ16, θ17, θ18, θ19, θ111, θ112, θ193, θ194, θ195, θ196, θ N is known. The calculation device 240 stores in advance relative positional relationships, such as relative distances or relative coordinates, of the M microphones 222. That is, the calculation device 240 stores in advance distances d between adjacent microphones 222.
[0082] FIG. 4 shows a schematic diagram of a full observation signal according to an embodiment of the present specification. In some embodiments, the microphone signal X is a full observation signal as shown in FIG. 4. In the full observation signal, all data in the M×K data matrix is complete. As shown in FIG. 4, the horizontal direction is the frame number k of the microphone signal X, and the vertical direction is the microphone signal number m in the microphone array 220. The m-th row represents the microphone signal received by the m-th microphone 222, and the k-th column represents the microphone signal of the k-th frame.
[0083] FIG. 5A shows a schematic diagram of a non-full observation signal according to an embodiment of the present specification. In some embodiments, the microphone signal X is a non-full observation signal as shown in FIG. 5A. In the non-full observation signal, some data in the M×K data matrix is missing. The calculation device 240 can rearrange the non-full observation signal. As shown in FIG. 5A, the horizontal direction is the frame number k of the microphone signal X, and the vertical direction is the channel number m of the microphone signal. The m-th row represents the microphone signal received by the m-th microphone 222, and the k-th column represents the microphone signal of the k-th frame.
[0084] If the microphone signal X is the non-full observation signal, step S120 may further include rearranging the non-full observation signal. FIG. 5B shows a schematic diagram of rearranging the non-full observation signal according to an embodiment of the present specification, and FIG. 5C shows a schematic diagram of rearranging the non-full observation signal according to an embodiment of the present specification. The case where the calculation device 240 rearranges the non-full observation signal may be as follows: The calculation device 240 obtains the non-full observation signal, and the calculation device 240 performs row permutation and column permutation on the microphone signal X according to data missing positions in each column of the M×K data matrix, and divides the microphone signal X into at least one sub-microphone signal. The microphone signal X includes the at least one sub-microphone signal.
[0085] In the non-full observation signal, microphone signals x of different frame numbers k Since the data missing positions in the frames x and y may be the same, in order to reduce the computation amount and computation time of the algorithm, the computation device 240 may k According to the data missing position in , the microphone signal X of K frames is classified, and the microphone signal x with the same data missing position is classified. kThe present inventors divide K frames of microphone signals X into at least one sub-microphone signal. For ease of explanation, the present inventors define the number of at least one sub-microphone signal as G, where G is a positive integer equal to or greater than 1. The present inventors divide the gth sub-microphone signal X into at least one sub-microphone signal X, and then permute the row positions in the data matrix of the microphone signal X so that the microphone signal positions in the same sub-microphone signal are adjacent, as shown in FIG. 5B. g where g = 1, 2, , G.
[0086] The computing device 240 further calculates each sub-microphone signal X g According to the data missing positions in , row permutation can be performed on the microphone signal X to make the data missing positions in all the sub-microphone signals adjacent, as shown in FIG. 5C.
[0087] As described above, in the non-full observation signal, the sub-microphone signal X g may be expressed as follows:
number
[0088] where X g =Q g XB g T and D g =Q g DB g T and P g =Q g P and S g =B g S. Matrix Q g , B g is a matrix consisting of 0 and 1 elements determined by the location of missing data.
[0089] The microphone signal X may be expressed as:
number
[0090] For ease of explanation, in the following description we will assume that the microphone signal X is a non-full observation signal.
[0091] As described above, the microphone 222 can collect a noise signal D and can also collect a target speech signal. When the target speech signal is not present in the microphone signal X, the microphone signal X satisfies a first model corresponding to the noise signal D. When the target speech signal is present in the microphone signal X, the microphone signal satisfies a second model corresponding to the combination of the target speech signal and the noise signal D.
[0092] For ease of explanation, we define the first model as the following equation:
number
[0093] When the microphone signal X is a full observation signal, the first model may be expressed as:
number
[0094] When the microphone signal X is a non-full observation signal, the first model may be expressed as:
number
[0095] We define the second model as follows:
number
[0096] When the microphone signal X is a full observation signal, the second model may be expressed as:
number
[0097] When the microphone signal X is a non-full observation signal, the second model may be expressed as:
number
[0098] For ease of explanation, in the following description, we take as an example that the microphone signal X is a non-full observation signal.
[0099] As shown in FIG. 3, the method P100 may include the following steps. S140: Optimizing the first model and the second model with a joint optimization objective of maximizing a likelihood function and minimizing the rank of a noise covariance matrix, and obtaining a first estimate of a noise covariance matrix M1 of the first model.
number
number
[0100] In the first model, there exists a noise covariance matrix M of the noise signal D of unknown parameters. For ease of explanation, we define the noise covariance matrix M of the noise signal D of unknown parameters in the first model as M1. In the second model, there exists a noise covariance matrix M of the noise signal D of unknown parameters and an amplitude S of the target speech signal. For ease of explanation, we define the noise covariance matrix M of the noise signal D of unknown parameters in the second model as M2. The computing device 240 optimizes the first model and the second model respectively based on an optimization method, and obtains a first estimate value of the unknown parameter M_1.
number
number
number
[0101] According to a first aspect, the computing device 240 can be triggered in terms of the likelihood function and can perform optimization design for each of the first model and the second model with the optimization goal being maximization of the likelihood function. According to another aspect, as described above, the colored noise signal c k The noise covariance matrix M c Since, has the low-rank property and is a low-rank positive semidefinite matrix, the noise signal d k The noise covariance matrix M of also has the low rank property. In particular, in the case of non-full observed signals, during the rearrangement of the non-full observed signals, the noise signal d k It is necessary to maintain the low rank property of the noise covariance matrix M of the noise signal d kBased on the low-rank characteristic of the noise covariance matrix M, the first model and the second model can be optimized with the rank minimization of the noise covariance matrix M as an optimization objective. Therefore, the computing device 240 optimizes the first model and the second model with the joint optimization objectives of maximizing the likelihood function and minimizing the rank of the noise covariance matrix to obtain a first estimate of the unknown parameter M1.
number
number
number
[0102] 6 shows a flowchart of the iterative optimization according to an embodiment of the present specification. Shown in FIG. 6 is step S140. As shown in FIG. 6, step S140 may include:
[0103] S142: Using the microphone signal X as sample data, establish a first likelihood function L1(M1) corresponding to the first model.
[0104] The likelihood function includes the first likelihood function L1(M1). According to equations (11) to (13), the first likelihood function L1(M1) may be expressed by the following equation:
number
[0105] Here, equation (17) represents a first likelihood function L1(M1) for each of the fully observed signal and the non-fully observed signal.
number
number
number
number
[0106] S144: Optimizing the first model with the optimization objectives of maximizing a first likelihood function L1(M1) and minimizing a rank Rank(M1) of a noise covariance matrix M1 of the first model, and obtaining a first estimate of M1.
number
[0107] The maximization of the first likelihood function L1(M1) may be expressed as min(-log(L1(M1))). The minimization of the rank Rank(M1) of the noise covariance matrix M1 of the first model may be expressed as min(Rank(M1)). As mentioned above, we consider the white noise signal n k The noise covariance matrix δ0 2 I n As an example, it is assumed that the noise covariance matrix M1 of the first model is known. As can be seen from equation (7), the rank minimization of the noise covariance matrix M1 of the colored noise signal C is c Minimize min(Rank(M c )) Therefore, the target function for the optimization goal may be expressed as
number
[0108] where γ is a regularization coefficient. The matrix rank minimization can be relaxed to a nuclear norm minimization problem. Therefore, according to (18), it can be expressed as
number
[0109] The iteration constraint of the first model may be expressed as:
number
[0110] Here, M c ≧0 is the noise covariance matrix M of the colored noise signal C c The optimization problem of the first model may be expressed as:
number
[0111] After determining the target function and the constraint condition, the computing device 240 performs iterative optimization on the unknown parameters M1 of the first model with the target function as an optimization target to obtain a first estimate (
number
[0112] Equation (21) is a semidefinite programming problem, which the computing device 240 can solve by a number of algorithms. For example, the gradient projection algorithm may be used. Specifically, in each iteration of the gradient projection algorithm, we first solve equation (19) by a gradient method without imposing any constraints, and then project the obtained solution onto a semidefinite cone to satisfy the matrix semidefinite constraint equation (20).
[0113] As shown in FIG. 6, step S140 may further include: S146: Using the microphone signal X as sample data, establish a second likelihood function L2(S,M2) of the second model.
[0114] The likelihood function includes a second likelihood function L2(S, M2). According to equations (14) to (16), the second likelihood function L2(S, M2) may be expressed by the following equation:
number
[0115] Here, equation (22) represents a second likelihood function for each of the fully observed signal and the non-fully observed signal.
number
number
number
[0116] S148: Optimizing the second model with the optimization objectives of maximizing a second likelihood function L2(S,M2) and minimizing a rank Rank(M2) of a noise covariance matrix M2 of the second model, and obtaining a second estimate of M2
number
number
[0117] The maximization of the second likelihood function L2(S,M2) may be expressed as min(-log(L2(S,M2))). The minimization of the rank Rank(M2) of the noise covariance matrix M2 of the second model may be expressed as min(Rank(M2)). As mentioned above, we consider the white noise signal n k The noise covariance matrix δ0 2 I n As an example, it is assumed that the noise covariance matrix M2 of the second model is known. As can be seen from equation (7), minimizing the rank Rank(M2) of the noise covariance matrix M2 of the colored noise signal C is c Minimize min(Rank(M c )) Therefore, the target function for the optimization goal may be expressed as
number
[0118] where γ is a regularization coefficient. The matrix rank minimization can be relaxed to a nuclear norm minimization problem. Therefore, according to (23), it can be expressed as
number
[0119] The iteration constraint of the second model may be expressed as follows:
number
[0120] Here, M c ≧0 is the noise covariance matrix M of the colored noise signal C c The optimization problem of the second model may be expressed as:
number
[0121] After determining the target function and the constraints, the computing device 240 performs iterative optimization on the unknown parameters M and S of the second model with the target function as an optimization goal to obtain a second estimate of the noise covariance matrix M of the second model.
number
number
[0122] Equation (26) is a semidefinite programming problem, which the computing device 240 can solve by multiple algorithms. For example, the gradient projection algorithm may be used. Specifically, in each iteration of the gradient projection algorithm, we first solve equation (24) by a gradient method without imposing any constraints, and then project the obtained solution onto a semidefinite cone to satisfy the semidefinite constraint condition equation (25) of the matrix.
[0123] As described above, the method P100 optimizes the first model and the second model, respectively, with the joint optimization objectives of maximizing the likelihood function and minimizing the rank of the noise covariance matrix to obtain a first estimate of the unknown parameter M1.
number
number
[0124] As shown in FIG. 3, the method P100 may further include: S160: Determine a target model and a noise covariance matrix M corresponding to the microphone signal X based on a statistical hypothesis test.
[0125] The target model includes one of a first model and a second model. The noise covariance matrix M of the microphone signal X is the noise covariance matrix of the target model. When the target model of the microphone signal X is the first model, the noise covariance matrix M of the microphone signal X is the noise covariance matrix of the target model.
number
number
[0126] The computing device 240 can determine whether a target voice signal is present in the microphone signal X by determining whether the microphone signal X satisfies a first model or a second model based on a statistical hypothesis testing method.
[0127] 7 shows a flowchart of determining a target model according to an embodiment of the present disclosure. The flowchart shown in FIG. 7 is step S160.
[0128] As shown in FIG. 7, step S160 may include: S162: Based on the microphone signal X, a binary hypothesis testing model is established.
[0129] Here, the null hypothesis H0 of the binary hypothesis testing model may be that the target voice signal is not present in the microphone signal X, i.e., the microphone signal X satisfies a first model. The alternative hypothesis H1 of the binary hypothesis testing model may be that the target voice signal is present in the microphone signal X, i.e., the microphone signal X satisfies a second model. The binary hypothesis testing model may be expressed by the following formula:
number
number
[0130] Here, the microphone signal X in equation (27) is a full observation signal, and the microphone signal X in equation (28) is a non-full observation signal.
[0131] S164: The first estimate
number
number
number
[0132] The detector may be any one or more detectors. In some embodiments, the detector may be one or more of a GLRT detector, a Rao checker, and a Wald checker. In some embodiments, the detector may also be a u-checker, a t-checker, a chi-squared test, an F-checker, a rank sum detector, etc. Different detectors have different test statistics ψ.
[0133] Take the GLRT detector (Generalized Likelihood Ratio Test) as an example. When the microphone signal X is a full observation signal, in the GLRT detector, the test statistic ψ may be expressed by the following equation:
number
[0134] Where:
number
number
number
number
[0135] When the microphone signal X is a non-full observation signal, in the GLRT detector, the test statistic ψ may be expressed as:
number
[0136] Where:
number
number
number
number
[0137] In the GLRT detector, unknown parameters under the null hypothesis H0 and alternative hypothesis H1
number
number
[0138] Therefore, in order to meet the equalization requirements for detection performance and computational complexity of the actual system, the computing device 240 proposes a Rao detector based on the above-mentioned GLRT detector. Taking a non-full observation signal as an example, the test statistic ψ of the Rao detector may be expressed as follows:
number
[0139] Here, f(X1,X2, ,X G │θ,M) represents the probability density function for the alternative hypothesis H1. M=M2. θ r =[PS R,1’ ,PS R,2’ ,···,PS R,M’ ,PS L,1’ ,PS L,2’ ,···,PS L,M’ ] T Here, PS R,mis the real part of the amplitude of the audio signal of the mth microphone 222 of the target audio signal. L,m is the imaginary part of the amplitude of the audio signal of the m-th microphone 222 of the target audio signal, where m=1, 2, . . . , M. θ r is a 2M-dimensional vector. θ=[θ r T θ S T ] T where θ s is a real vector containing extra parameters. It contains the real and imaginary parts of the M off-diagonal elements and the diagonal elements. Equation (31) can be simplified to the following equation:
number
[0140] In equation (32), the unknown parameters in the null hypothesis H0
number
number
[0141] S166: Determine a target model of the microphone signal X based on the test statistic ψ.
[0142] Specifically, step S166 is S166-2: Determine that the test statistic ψ is greater than a preset decision threshold η, and determine that a target voice signal exists in the microphone signal X, and the target model is a second model, and the noise covariance matrix of the microphone signal is the second estimated value
number
number
[0143] Step S166 may be expressed by the following equation:
number
[0144] The decision threshold η is a parameter related to the false alarm probability, which can be obtained by experiment, machine learning, or even experience.
[0145] As shown in FIG. 3, the method P100 includes: S180: The method may further include outputting a target mode and a noise covariance matrix M of the microphone signal X.
[0146] The computation unit 240 may output the target mode and noise covariance matrix M of the microphone signal X to other computation modules, such as a speech enhancement module.
[0147] As described above, in the voice activity detection system and method P100 according to the present specification, the computing device 240 optimizes the first model and the second model, respectively, with the joint optimization objectives of maximizing the likelihood function and minimizing the rank of the noise covariance matrix to obtain a first estimate of the unknown parameter M1.
number
number
[0148] The present specification further provides a voice enhancement system. The voice enhancement system can also be used in the electronic device 200. In some embodiments, the voice enhancement system can include a computing device 240. In some embodiments, the voice enhancement system can be used in the computing device 240. That is, the voice enhancement system can run on the computing device 240. The voice enhancement system can include a hardware device having a data information processing function and a program required to drive the operation of the hardware device. Of course, the voice enhancement system can also be only a hardware device having a data processing function, or only a program running on the hardware device.
[0149] The voice enhancement system may store data or instructions and may execute the data and / or instructions to perform the voice enhancement methods described herein. When the voice enhancement system runs on the computing device 240, the voice enhancement system may obtain the microphone signals from the microphone array 220 based on the communication connection and execute data or instructions of the voice enhancement methods described herein. The voice enhancement methods are introduced elsewhere herein. For example, the voice enhancement methods are introduced in the description of FIG. 8.
[0150] When the voice enhancement system runs on computing device 240, the voice enhancement system is in communication with microphone array 220. Storage medium 243 may further include at least one set of instructions stored in the data storage device for performing voice enhancement calculations on the microphone signals. The instructions may be computer program code, which may include programs, routines, objects, components, data structures, processes, modules, etc., for performing a voice enhancement method according to the present disclosure. Processor 242 may read the at least one set of instructions and perform a voice enhancement method according to the present disclosure according to instructions of the at least one set of instructions. Processor 242 may perform all steps included in the voice enhancement method.
[0151] FIG. 8 shows a flowchart of a speech enhancement method P200 according to an embodiment of the present specification. The method P200 can perform speech enhancement on the microphone signal. Specifically, the processor 242 can execute the method P200. As shown in FIG. 9, the method P200 may include:
[0152] S220: Acquire microphone signals X output by the M microphones. This is as described in step S120, and the description will be omitted here.
[0153] S240: Based on the voice activity detection method P100, a target model of the microphone signal X and a noise covariance matrix M of the microphone signal X are determined.
[0154] The noise covariance matrix M of the microphone signal X is the noise covariance matrix of the target model. When the target model of the microphone signal X is the first model, the noise covariance matrix M of the microphone signal X is
number
number
[0155] S260: Based on the MVDR method and the noise covariance matrix M of the microphone signal X, a filtering coefficient ω corresponding to the microphone signal X is determined.
[0156] The filtering coefficient ω may be an M×1 dimensional vector. The filtering coefficient ω may be expressed by the following formula:
number
[0157] where the filtering coefficients corresponding to the mth and microphones 222 are ω m m = 1, 2, , M.
[0158] The filtering coefficient ω may be expressed as follows:
number
[0159] As previously mentioned, P is the target steering vector of the target audio signal. In some embodiments, P is known.
[0160] S280: Integrate the microphone signal X based on the filtering coefficients to generate a target audio signal y k Output.
[0161] The target audio signal Y may be expressed as follows:
number
[0162] The computing device 240 can output the target audio signal Y to other electronic devices, such as a remote communication device.
[0163] As described above, the voice activity detection system and method P100 and the speech enhancement system and method P200 according to the present specification are used for a microphone array 220 consisting of a plurality of microphones 222. The voice activity detection system and method P100 and the speech enhancement system and method P200 can acquire a microphone signal X collected by the microphone array 220. The microphone signal X may be a first model corresponding to a noise signal, or a second model corresponding to a combination of a target voice signal and the noise signal. The voice activity detection system and method P100 and the speech enhancement system and method P200 take the microphone signal X as a sample, optimize the first model and the second model respectively with the joint optimization objectives of maximizing a likelihood function and minimizing the rank of the noise covariance matrix M of the microphone signal X, and obtain a first estimate of the noise covariance matrix M1 of the first model.
number
number
[0164] Another aspect of the present specification provides a non-transitory storage medium, storing at least one set of executable instructions for voice activity detection, which when executed by a processor, directs the processor to perform steps of the voice activity detection method P100 described herein. In some possible embodiments, each aspect of the present specification may be further realized in the form of a program product including a program code. When the program product runs on a computing device (e.g., computing device 240), the program code is for causing the computing device to perform the voice activity detection steps described herein. The program product for implementing the above method may use a portable compact disc read only memory (CD-ROM), including the program code and operable on a computing device. However, the program product of the present specification is not limited thereto, and in the present specification, a readable storage medium may be any tangible medium that includes or stores a program, which may be used by or in combination with an instruction execution system (e.g., processor 242). The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. Further examples of the readable storage medium include an electrical connection having one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical memory device, a magnetic memory device, or any suitable combination thereof. The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, in which the readable program code is embedded.Such propagated data signals may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may transmit, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the readable storage medium may be transmitted in any suitable medium, including, but not limited to, wireless, wired, optical cable, RF, or the like, or any suitable combination of the above. The program code for performing the operations herein may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, or the like, general procedural programming languages such as the "C" language, or similar programming languages. The program code may be executed entirely on the computing device, partially on the computing device, as separate software packets, partially on the computing device and partially on a remote computing device, or entirely on a remote computing device.
[0165] The above describes certain embodiments of the present specification. Other embodiments are within the scope of the following claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the examples and still achieve desirable results. Also, the processes depicted in the figures do not necessarily require a particular order or sequential order to achieve desirable results. In some embodiments, multitasking and parallel processing may also be possible or advantageous.
[0166] As described above, upon reading this detailed disclosure, those skilled in the art will understand that the detailed disclosure above may be presented by way of example only and may not be limiting. Although not expressly stated herein, the present specification should cover various reasonable changes, modifications, and alterations to the embodiments, as would be understood by those skilled in the art. These changes, modifications, and alterations are intended to be presented by this specification and are within the spirit and scope of the exemplary embodiments of the present specification.
[0167] It should be noted that certain terms in this specification are used to describe embodiments of this specification. For example, "one embodiment," "embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of this specification. Therefore, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily refer to the same embodiment. It should be noted that certain features, structures, or characteristics may be suitably combined in one or more embodiments of this specification.
[0168] It should be understood that in the above description of the embodiments of this specification, in order to facilitate the understanding of one feature, this specification combines various features into a single embodiment, drawing, or description thereof for the purpose of simplifying this specification. However, it is not necessary to combine these features, and a person skilled in the art can fully extract some of the features and understand them as a single embodiment when reading this specification. In other words, the embodiments in this specification can be understood as a combination of multiple secondary embodiments. The content of each secondary embodiment can be valid even if it has less than all the features of the single embodiment disclosed above.
[0169] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, documents, articles, etc., cited herein may be incorporated herein by reference. All content for all purposes, now or hereafter, is associated with this document, except for any claim history related thereto, any identical claim history that is inconsistent or inconsistent with this document, or any identical claim history that has a limiting effect on the broadest scope of the claims. For example, in the event of any inconsistency or discrepancy between the explanation, definition, and / or use of a term associated with any of the included materials and the explanation, definition, and / or use of the term associated with this document, the term in this document shall control.
[0170] Finally, it should be understood that the embodiments of the application disclosed herein are explanations of the principles of the embodiments of the present specification. Other modified embodiments are also within the scope of the present specification. Therefore, the embodiments disclosed herein are merely examples and are not limiting. Those skilled in the art can realize the application of the present specification using alternative configurations based on the embodiments of the present specification. Therefore, the embodiments of the present specification are not limited to the embodiments exactly described in the application.
Claims
1. A voice activity detection method for M microphones distributed in a predefined array shape, where M is an integer greater than 1, comprising: Obtaining microphone signals output by the M microphones that satisfy a first model in which a target sound signal is not present or a second model in which a target sound signal is present; optimizing the first model and the second model, respectively, with a joint optimization objective of maximizing a likelihood function and minimizing the rank of a noise covariance matrix to determine a first estimate of the noise covariance matrix of the first model and a second estimate of the noise covariance matrix of the second model; determining a target model and a noise covariance matrix corresponding to the microphone signal based on a statistical hypothesis test, wherein the target model comprises one of the first model and the second model, and the noise covariance matrix of the microphone signal is the noise covariance matrix of the target model.
2. 2. The method of claim 1, wherein the microphone signal comprises K frames of consecutive audio signals, where K is a positive integer greater than 1, and the microphone signal comprises an MxK data matrix.
3. The microphone signals are full observation signals or non-full observation signals, in which all data in the M×K data matrix is complete in the full observation signal, and in which some data in the M×K data matrix is missing in the non-full observation signal, and when the microphone signals are the non-full observation signals, obtaining the microphone signals output by the M microphones includes: acquiring the non-full observed signal; 3. The method of claim 2, further comprising: performing row permutation and column permutation on the microphone signals based on data missing positions in each column of the M×K data matrix, and dividing the microphone signals into at least one sub-microphone signal, wherein the microphone signals include the at least one sub-microphone signal.
4. optimizing the first model and the second model with a joint optimization objective of maximizing a likelihood function and minimizing the rank of a noise covariance matrix, respectively, Establishing a first likelihood function included in the likelihood function, the first likelihood function corresponding to the first model, using the microphone signal as sample data; optimizing the first model with optimization objectives of maximizing the first likelihood function and minimizing the rank of a noise covariance matrix of the first model to determine the first estimate; Establishing a second likelihood function of the second model, the second likelihood function being included in the likelihood function using the microphone signal as sample data; optimizing the second model with optimization objectives of maximizing the second likelihood function and minimizing the rank of a noise covariance matrix of the second model to determine the second estimate and an amplitude estimate of the target speech signal.
5. The microphone signal includes a noise signal that follows a Gaussian distribution, the noise signal having at least 5. The voice activity detection method of claim 4, comprising a colored noise signal that follows a zero-mean Gaussian distribution and whose corresponding noise covariance matrix is a low-rank positive semidefinite matrix.
6. Determining a target model and a noise covariance matrix corresponding to the microphone signals based on the statistical hypothesis testing includes: establishing a binary hypothesis testing model based on the microphone signal, where a null hypothesis of the binary hypothesis testing model includes that the microphone signal satisfies the first model and an alternative hypothesis of the binary hypothesis testing model includes that the microphone signal satisfies the second model; substituting the first estimate, the second estimate and the amplitude estimate into a detector criterion of the binary hypothesis testing model to obtain a test statistic; and determining the target model of the microphone signal based on the test statistic.
7. Determining the target model of the microphone signal based on the test statistic includes: determining that the test statistic is greater than the predetermined decision threshold, determining that the target speech signal is present in the microphone signal, and determining that the target model is the second model and that the noise covariance matrix of the microphone signal is the second estimate; or 7. The method of claim 6, further comprising: determining that the test statistic is less than the predetermined decision threshold; determining that the target voice signal is not present in the microphone signal; and determining that the target model is the first model and that a noise covariance matrix of the microphone signal is the first estimate.
8. 7. The voice activity detection method of claim 6, wherein the detector comprises at least one of a GLRT detector, a Rao checker, and a Wald checker.
9. 1. A voice activity detection system comprising: at least one storage medium having stored thereon at least one instruction set for voice activity detection; at least one processor in communication with the at least one storage medium; 9. The voice activity detection system according to claim 1 , wherein, when the voice activity detection system is in operation, the at least one processor reads the at least one instruction set and implements the voice activity detection method according to any one of claims 1 to 8.
10. A method for speech enhancement, which is used with M microphones distributed in a predefined array shape, where M is an integer greater than 1, acquiring microphone signals output by the M microphones; - determining the target model of the microphone signal and a noise covariance matrix of the microphone signal, which is the noise covariance matrix of the target model, according to a voice activity detection method according to any one of claims 1 to 8; determining filtering coefficients corresponding to the microphone signals based on an MVDR method and a noise covariance matrix of the microphone signals; and integrating the microphone signal based on the filtering coefficients to output a target audio signal.
11. 1. A speech enhancement system, comprising: at least one storage medium having stored thereon at least one instruction set for performing speech enhancement; at least one processor in communication with the at least one storage medium; wherein, when the speech enhancement system is in operation, the at least one processor:
11. A speech enhancement system for reading said at least one instruction set and implementing the speech enhancement method of claim 10.
Citation Information
Patent Citations
Speech recognition method
JP2006154819A
Acoustic processing device, acoustic processing system and acoustic processing method
JP2018036332A
Globally optimized least-squares postfiltering for speech enhancement.
JP2019508719A