Voice activity detection method, system, speech enhancement method and system

By calculating the linear correlation between the signal subspace of the microphone signal and the target subspace of the target speech signal, the probability of speech presence is determined and filtered. This solves the problem of low accuracy in noise covariance matrix estimation on devices with a small number of microphones and small spacing, thus improving the speech enhancement effect.

CN116364100BActive Publication Date: 2026-02-06SHENZHEN SHOKZ CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111563890.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-20
Publication Date
2026-02-06
Estimated Expiration
2041-12-20

AI Technical Summary

Technical Problem

Existing speech activity detection methods suffer from low accuracy in noise covariance matrix estimation on devices with few microphones and small spacing, especially in devices such as headphones, resulting in poor speech enhancement performance of the MVDR algorithm.

Method used

By acquiring signals from multiple microphones, the linear correlation between the signal subspace and the target subspace is determined, the probability of speech presence is calculated, and the filtering coefficients are calculated based on this probability to perform speech enhancement.

Benefits of technology

The accuracy of speech presence probability calculation has been improved, thereby enhancing speech enhancement, especially on devices with a small number of microphones and close spacing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116364100B_ABST
    Figure CN116364100B_ABST
Patent Text Reader

Abstract

The voice activity detection method and system and the voice enhancement method and system provided in the specification determine the voice presence probability of the target voice signal in the microphone signal by calculating the linear correlation between the signal subspace where the microphone signal is located and the target subspace where the target voice signal is located. The voice enhancement method and system can calculate the filtering coefficient based on the voice presence probability, so as to perform voice enhancement on the microphone signal. The method and system can improve the calculation accuracy of the voice presence probability, and further improve the voice enhancement effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of target speech signal processing, and particularly relates to a voice activity detection method and system, and a speech enhancement method and system. BACKGROUND

[0002] In the speech enhancement technology based on the beamforming algorithm, especially in the adaptive beamforming algorithm of the Minimum Variance Distortionless Response (MVDR), how to solve the parameter that describes the statistical characteristic relationship of the noise between different microphones, the noise covariance matrix, is crucial. The main method in the prior art is to calculate the noise covariance matrix based on the speech presence probability, such as estimating the speech presence probability by the Voice Activity Detection (VAD), and then calculating the noise covariance matrix. However, the accuracy of the speech presence probability estimation in the prior art is not enough, which leads to low accuracy of the noise covariance matrix estimation, and further leads to poor speech enhancement effect of the MVDR algorithm. Especially when the number of microphones is small, such as less than 5, the effect decreases sharply. Therefore, the MVDR algorithm in the prior art is mostly used in microphone array devices with many microphones and large spacing, such as mobile phones and smart speakers, and the speech enhancement effect is poor for devices with few microphones and small spacing, such as earphones.

[0003] Therefore, it is necessary to provide a voice activity detection method and system with higher precision, and a speech enhancement method and system. SUMMARY

[0004] The present specification provides a voice activity detection method and system with higher precision, and a speech enhancement method and system.

[0005] In a first aspect, the present specification provides a voice activity detection method for M microphones arranged in a preset array shape, where M is an integer greater than 1, comprising: obtaining microphone signals output by the M microphones; determining a signal subspace formed by the microphone signals based on the microphone signals; determining a target subspace formed by a target speech signal; and determining a speech presence probability that the target speech signal exists in the microphone signals based on the linear correlation between the signal subspace and the target subspace, and outputting.

[0006] In some embodiments, the determining, based on the microphone signals, a signal subspace formed by the microphone signals comprises: determining, based on the microphone signals, a sample covariance matrix of the microphone signals; performing eigen decomposition on the sample covariance matrix to determine a plurality of eigenvectors of the sample covariance matrix; and taking a matrix formed by at least part of the plurality of eigenvectors as a basis matrix of the signal subspace.

[0007] In some embodiments, the determining, based on the microphone signals, a signal subspace formed by the microphone signals comprises: determining, based on the microphone signals, a signal direction vector of the microphone signals by a spatial estimation method, the spatial estimation method comprising at least one of a DOA estimation method and a spatial spectrum estimation method; and taking the signal direction vector as a basis matrix of the signal subspace.

[0008] In some embodiments, the determining a target subspace formed by the target speech signal comprises: determining a preset target direction vector corresponding to the target speech signal as a basis matrix of the target subspace.

[0009] In some embodiments, the determining, based on the linear correlation between the signal subspace and the target subspace, a speech presence probability that the target speech signal exists in the microphone signals and outputting the speech presence probability comprises: determining a volume correlation function of the signal subspace and the target subspace; determining, based on the volume correlation function, a linear correlation coefficient of the signal subspace and the target subspace, wherein the linear correlation coefficient is negatively correlated with the volume correlation function; and taking the linear correlation coefficient as the speech presence probability and outputting the speech presence probability, wherein the determining, based on the volume correlation function, the linear correlation coefficient of the signal subspace and the target subspace comprises one of the following: determining that the volume correlation function is greater than a first threshold, and determining that the linear correlation coefficient is 0; determining that the volume correlation function is less than a second threshold, and determining that the linear correlation coefficient is 1, wherein the second threshold is less than the first threshold; and determining that the volume correlation function is between the first threshold and the second threshold, determining that the linear correlation coefficient is between 0 and 1, and the linear correlation coefficient is a negative correlation function of the volume correlation function.

[0010] In a second aspect, the present specification also provides a voice activity detection system, comprising at least one storage medium storing at least one instruction set for voice activity detection; and at least one processor in communication with the at least one storage medium, wherein when the voice activity detection system is running, the at least one processor reads the at least one instruction set and implements the voice activity detection method of the first aspect of the present specification.

[0011] In a third aspect, the present specification also provides a voice enhancement method for M microphones arranged in a preset array shape, wherein M is an integer greater than 1, comprising: obtaining microphone signals output by the M microphones; determining a voice presence probability that the target voice signal exists in the microphone signals based on the voice activity detection method of the first aspect of the present specification; determining a filter coefficient vector corresponding to the microphone signals based on the voice presence probability; and merging the microphone signals based on the filter coefficient vector to obtain a target audio signal and output.

[0012] In some embodiments, the determining of the filter coefficient vector corresponding to the microphone signals based on the voice presence probability comprises: determining a noise covariance matrix of the microphone signals based on the voice presence probability; and determining the filter coefficient vector based on an MVDR method and the noise covariance matrix.

[0013] In some embodiments, the determining of the filter coefficient vector corresponding to the microphone signals based on the voice presence probability comprises: taking the voice presence probability as a filter coefficient corresponding to a target microphone signal in the microphone signals, the target microphone signal comprising a microphone signal with the highest signal-to-noise ratio in the microphone signals; and determining filter coefficients corresponding to the remaining microphone signals other than the target microphone signal in the microphone signals as 0, the filter coefficient vector comprising a vector composed of the filter coefficient corresponding to the target microphone signal and the filter coefficients corresponding to the remaining microphone signals.

[0014] In a fourth aspect, the present specification also provides a voice enhancement system, comprising at least one storage medium storing at least one instruction set for voice enhancement; and at least one processor in communication with the at least one storage medium, wherein when the voice enhancement system is running, the at least one processor reads the at least one instruction set and implements the voice enhancement method of the second aspect of the present specification.

[0015] It can be known from the technical solution that the voice activity detection method and system and the voice enhancement method and system provided in the specification are used for a microphone array composed of multiple microphones. The microphone array can collect noise signals and target voice signals and output microphone signals. The target voice signals and the noise signals belong to two non-interconnected signals. A target subspace where the target voice signals are located and a noise subspace where the noise signals are located belong to two non-interconnected subspaces. When the target voice signals do not exist in the microphone signals, only the noise signals exist in the microphone signals. At this time, a signal subspace where the microphone signals are located and the target subspace where the target voice signals are located belong to two non-interconnected subspaces, and the linear correlation between the signal subspace and the target subspace is low. When the target voice signals exist in the microphone signals, both the target voice signals and the noise signals exist in the microphone signals. At this time, the signal subspace where the microphone signals are located and the target subspace where the target voice signals are located belong to two interconnected subspaces, and the linear correlation between the signal subspace and the target subspace is high. Therefore, the voice activity detection method and system provided in the specification can determine the voice existence probability of the target voice signals existing in the microphone signals by calculating the linear correlation between the signal subspace where the microphone signals are located and the target subspace where the target voice signals are located. The voice enhancement method and system can calculate the filter coefficient based on the voice existence probability, so as to perform voice enhancement on the microphone signals. The method and system can improve the calculation accuracy of the voice existence probability, and further improve the voice enhancement effect.

[0016] Other functions of the voice activity detection method and system and the voice enhancement method and system provided in the specification will be partially listed in the following description. According to the description, the following numbers and examples will be apparent to those skilled in the art. The creative aspects of the voice activity detection method and system and the voice enhancement method and system provided in the specification can be fully explained by practicing or using the methods, devices and combinations described in the following detailed examples. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the specification, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the specification, and those skilled in the art can also obtain other drawings according to these drawings without any creative labor.

[0018] Figure 1 A hardware schematic diagram of a voice activity detection system provided according to an embodiment of the specification is shown;

[0019] Figure 2A An exploded structural schematic diagram of an electronic device provided according to an embodiment of the specification is shown.

[0020] Figure 2B A front view of a first housing is shown, according to an embodiment of the present specification;

[0021] Figure 2C A top view of a first housing is shown, according to an embodiment of the present specification;

[0022] Figure 2D A front view of a second housing is shown, according to an embodiment of the present specification;

[0023] Figure 2E A bottom view of a second housing is shown, according to an embodiment of the present specification;

[0024] Figure 3 A flowchart of a voice activity detection method is shown, according to an embodiment of the present specification;

[0025] Figure 4 A flowchart of determining a basis matrix of a signal subspace is shown, according to an embodiment of the present specification;

[0026] Figure 5 A flowchart of another determining a basis matrix of a signal subspace is shown, according to an embodiment of the present specification;

[0027] Figure 6 A flowchart of calculating a voice presence probability is shown, according to an embodiment of the present specification;

[0028] Figure 7 A spatial principal angle diagram is shown, according to an embodiment of the present specification; and

[0029] Figure 8 A flowchart of a voice enhancement method is shown, according to an embodiment of the present specification. DETAILED DESCRIPTION

[0030] The following description provides specific details of particular applications and requirements of the present specification in order to provide a thorough description for those skilled in the art. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and generic principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the present specification. Thus, the present specification is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the claims.

[0031] The terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting. For example, singular articles "a," "an," and "the" can be read to include the plural unless the context clearly indicates otherwise. The terms "comprises," "comprising," "includes," "including," and / or "contains," when used in this specification, mean that the associated integer, step, operation, element, and / or component can be present but not exclude the addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0032] These and other features, and characteristics of the present specification, as well as the methods of operation and functions of the related elements of structure and the combination of parts and economies of manufacture, will become more apparent upon consideration of the following description and the appended claims with reference to the accompanying drawings. The current description is presented to enable any person skilled in the art to make and use the applications, but is not intended to limit the scope of what the applications purport to convey to others skilled in the art. In this manner, the current description is presented in sections to enable a person skilled in the art to convey the full scope of the applications as set forth in the appended claims. The current description is presented with information that is, at present, thought to be true, complete and accurate but is subject to one or more of several corrections or improvements as more information becomes available or as the present disclosure assumes a different context. The drawings shown in this current patent document are only meant to be illustrative and not meant to be actual drawings of the subject matter shown in this current patent document.

[0033] The flow diagrams herein illustrate the operations according to some embodiments of the present specification in terms of a particular implementation. It should be apparent that the operations of the flow diagrams can be implemented in a different order than as shown or that the operations can be implemented concurrently. Additionally, one or more operations can be added or omitted from the flow diagrams. Further, the flow diagrams can include one or more additional operations not presently contemplated.

[0034] For the purposes of this specification, the following terms shall have the following meanings. These terms are not intended to be limited by their meanings in other contexts.

[0035] Minimum Variance Distortionless Response (MVDR): is an adaptive beamforming algorithm based on the maximum Signal to Interference and Noise Ratio (SINR) criterion. The MVDR algorithm can adaptively minimize the array output power in the desired direction while maximizing the Signal to Interference and Noise Ratio. The goal is to minimize the variance of the recorded signal. If the noise signal and the desired signal are uncorrelated, the variance of the recorded signal is the sum of the variances of the desired signal and the noise signal. Therefore, the MVDR solution seeks to minimize this sum, thereby mitigating the impact of the noise signal. The principle is to select appropriate filter coefficients under the constraint of distortionless desired signal, so as to minimize the average power of the array output.

[0036] Voice Activity Detection: is a process of segmenting the speech active periods and non-speech periods in a target speech signal.

[0037] Gaussian distribution: Normal distribution, also known as "normal distribution", also known as Gaussian distribution, normal curve is bell-shaped, low at both ends and high in the middle, and left-right symmetric. Because the curve is bell-shaped, it is often called bell-shaped curve. If a random variable X follows a normal distribution with a mathematical expectation of μ and a variance of σ 2 , it is denoted as N(μ, σ 2 ). The probability density function of the normal distribution is determined by the expected value μ of the normal distribution and the standard deviation σ. When μ = 0 and σ = 1, the normal distribution is a standard normal distribution.

[0038] Figure 1 A hardware schematic diagram of a voice activity detection system according to an embodiment of the present specification is shown. The voice activity detection system can be applied to an electronic device 200.

[0039] In some embodiments, the electronic device 200 can be a wireless earphone, a wired earphone, a smart wearable device, such as a smart glasses, a smart helmet, or a smart watch, and the like, which has an audio processing function. The electronic device 200 can also be a mobile device, a tablet computer, a notebook computer, a built-in device of a motor vehicle, or the like, or any combination thereof. In some embodiments, the mobile device can include a smart home device, a smart mobile device, or the like, or any combination thereof. For example, the smart mobile device can include a mobile phone, a personal digital assistant, a game device, a navigation device, an ultra-mobile personal computer (UMPC), or the like, or any combination thereof. In some embodiments, the smart home device can include a smart television, a desktop computer, or the like, or any combination thereof. In some embodiments, the built-in device in the motor vehicle can include a vehicle-mounted computer, a vehicle-mounted television, or the like.

[0040] In the present specification, we take the electronic device 200 as an example of an earphone for description. The earphone can be a wireless earphone or a wired earphone. As shown in Figure 1 , the electronic device 200 can include a microphone array 220 and a computing device 240.

[0041] The microphone array 220 can be an audio acquisition device of the electronic device 200. The microphone array 220 can be configured to acquire local audio and output a microphone signal, that is, an electronic signal carrying audio information. The microphone array 220 can include M microphones 222 distributed in a preset array shape. The M is an integer greater than 1. The M microphones 222 can be uniformly distributed or non-uniformly distributed. The M microphones 222 can output a microphone signal. The M microphones 222 can output M microphone signals. Each microphone 222 corresponds to a microphone signal. The M microphone signals are collectively referred to as the microphone signal. In some embodiments, the M microphones 222 can be linearly distributed. In some embodiments, the M microphones 222 can also be distributed in other shapes of arrays, such as a circular array, a rectangular array, and the like. For the convenience of description, the following description will be described by taking the linear distribution of the M microphones 222 as an example. In some embodiments, M can be any integer greater than 1, such as 2, 3, 4, 5, or even more, and the like. In some embodiments, due to space limitations, M can be an integer greater than 1 and not greater than 5, such as in a product such as an earphone. When the electronic device 200 is an earphone, the distance between adjacent microphones 222 in the M microphones 222 can be between 20 mm and 40 mm. In some embodiments, the distance between adjacent microphones 222 can be smaller, such as between 10 mm and 20 mm.

[0042] In some embodiments, the microphone 222 can be a bone conduction microphone that directly acquires a human vibration signal. The bone conduction microphone can include a vibration sensor, such as an optical vibration sensor, an acceleration sensor, and the like. The vibration sensor can acquire a mechanical vibration signal (such as a signal generated by the vibration of the skin or bones when the user speaks) and convert the mechanical vibration signal into an electrical signal. The mechanical vibration signal referred to herein mainly refers to vibration transmitted through a solid. The bone conduction microphone contacts the skin or bones of the user through the vibration sensor or a vibration component connected to the vibration sensor, thereby acquiring the vibration signal generated by the bones or skin of the user when the user speaks, and converting the vibration signal into an electrical signal. In some embodiments, the vibration sensor can be a device sensitive to mechanical vibration and not sensitive to air vibration (that is, the response capability of the vibration sensor to mechanical vibration exceeds the response capability of the vibration sensor to air vibration). Since the bone conduction microphone can directly pick up the vibration signal of the sound generating part, the bone conduction microphone can reduce the influence of environmental noise.

[0043] In some embodiments, the microphone 222 can also be an air conduction microphone that directly acquires an air vibration signal. The air conduction microphone acquires the air vibration signal caused by the user when speaking and converts the air vibration signal into an electrical signal.

[0044] In some embodiments, the M microphones 222 can be M bone conduction microphones. In some embodiments, the M microphones 222 can also be M air conduction microphones. In some embodiments, the M microphones 222 can include both bone conduction microphones and air conduction microphones. Of course, the microphones 222 can also be other types of microphones. For example, optical microphones, microphones that receive myoelectric signals, etc.

[0045] The computing device 240 can be communicatively connected with the microphone array 220. The communicative connection refers to any form of connection that can directly or indirectly receive information. In some embodiments, the computing device 240 and the microphone array 220 can communicate data with each other through a wireless communicative connection; in some embodiments, the computing device 240 and the microphone array 220 can also communicate data with each other through a direct wired connection; in some embodiments, the computing device 240 can be directly connected with other circuits through a wired connection to establish an indirect connection with the microphone array 220, so as to communicate data with each other. In this specification, the direct wired connection between the computing device 240 and the microphone array 220 will be taken as an example for description.

[0046] The computing device 240 can be a hardware device with data information processing function. In some embodiments, the voice activity detection system can include the computing device 240. In some embodiments, the voice activity detection system can be applied to the computing device 240. That is, the voice activity detection system can run on the computing device 240. The voice activity detection system can include a hardware device with data information processing function and necessary programs for driving the hardware device to work. Of course, the voice activity detection system can also be only a hardware device with data processing function, or only a program running in the hardware device.

[0047] The voice activity detection system can store data or instructions for executing the voice activity detection method described in this specification, and can execute the data and / or instructions. When the voice activity detection system runs on the computing device 240, the voice activity detection system can obtain the microphone signals from the microphone array 220 based on the communicative connection, and execute the data or instructions of the voice activity detection method described in this specification to calculate the voice presence probability of the target voice signal existing in the microphone signals. The voice activity detection method is introduced in other parts of this specification. For example, the voice activity detection method is introduced in the description of Figures 3 to 8

[0048] As shown in Figure 1 , the computing device 240 can include at least one storage medium 243 and at least one processor 242. In some embodiments, the electronic device 200 can also include a communication port 245 and an internal communication bus 241.​

[0049] The internal communication bus 241 can connect different system components, including the storage medium 243, the processor 242, and the communication port 245.

[0050] The communication port 245 can be used for data communication between the computing device 240 and the outside world. For example, the computing device 240 can obtain the microphone signals from the microphone array 220 through the communication port 245.

[0051] The at least one storage medium 243 can include a data storage device. The data storage device can be a non-transitory storage medium or a transitory storage medium. For example, the data storage device can include one or more of a magnetic disk, a read-only memory (ROM), or a random access memory (RAM). When the voice activity detection system can be run on the computing device 240, the storage medium 243 can further include at least one instruction set stored in the data storage device for performing voice activity detection on the microphone signals. The instructions can be computer program code, which can include programs, routines, objects, components, data structures, procedures, modules, and the like that perform the voice activity detection methods provided in this specification.

[0052] The at least one processor 242 can be communicatively connected with the at least one storage medium 243 through an internal communication bus 241. The communicatively connected refers to any form of connection capable of directly or indirectly receiving information. The at least one processor 242 is configured to execute the at least one set of instructions described above. When the voice activity detection system can be run on the computing device 240, the at least one processor 242 reads the at least one set of instructions and performs the voice activity detection method provided in the specification according to the instructions of the at least one set of instructions. The processor 242 can perform all the steps included in the voice activity detection method. The processor 242 can be in the form of one or more processors, and in some embodiments, the processor 242 can include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of executing one or more functions, etc., or any combination thereof. For the purpose of illustration only, only one processor 242 is described in the computing device 240 in the specification. However, it should be noted that the computing device 240 in the specification can also include multiple processors 242, and therefore, the operations and / or method steps disclosed in the specification can be performed by one processor as described in the specification, or jointly performed by multiple processors. For example, if the processor 242 of the computing device 240 in the specification performs step A and step B, it should be understood that step A and step B can also be jointly or separately performed by two different processors 242 (for example, the first processor performs step A, and the second processor performs step B, or the first and second processors jointly perform steps A and B).

[0053] Figure 2A An exploded structural schematic diagram of an electronic device 200 is shown according to an embodiment of the specification. As shown, the electronic device 200 can include a microphone array 220, a computing device 240, a first housing 260, and a second housing 280. Figure 2A

[0054] ​The first housing 260 can be a mounting base of the microphone array 220. The microphone array 220 can be mounted inside the first housing 260. The shape of the first housing 260 can be adaptively designed according to the distributed shape of the microphone array 220, which is not limited in the present specification. The second housing 280 can be a mounting base of the computing device 240. The computing device 240 can be mounted inside the second housing 280. The shape of the second housing 280 can be adaptively designed according to the shape of the computing device 240, which is not limited in the present specification. When the electronic device 200 is an earphone, the second housing 280 can be connected with a wearing part. The second housing 280 can be connected with the first housing 260. As described above, the microphone array 220 can be electrically connected with the computing device 240. Specifically, the microphone array 220 can be electrically connected with the computing device 240 through the connection of the first housing 260 and the second housing 280.

[0055] In some embodiments, the first housing 260 can be fixedly connected with the second housing 280, such as integrally formed, welded, riveted, bonded, etc. In some embodiments, the first housing 260 can be detachably connected with the second housing 280. The computing device 240 can be communicatively connected with different microphone arrays 220. Specifically, the different microphone arrays 220 can be different in the number of microphones 222, different in the array shape, different in the distance between the microphones 222, different in the installation angle of the microphone array 220 in the first housing 260, different in the installation position of the microphone array 220 in the first housing 260, etc. The user can replace the corresponding microphone array 220 according to different application scenarios, so that the electronic device 200 is suitable for a wider range of scenarios. For example, when the user is close to the electronic device 200 in the application scenario, the user can replace the microphone array 220 with a smaller distance. For another example, when the user is close to the electronic device 200 in the application scenario, the user can replace the microphone array 220 with a larger distance and a larger number, etc.

[0056] The detachable connection can be any form of physical connection, such as threaded connection, buckle connection, magnetic attraction connection, etc. In some embodiments, the first housing 260 and the second housing 280 can be magnetically connected. That is, the first housing 260 and the second housing 280 are detachably connected through the adsorption force of a magnetic device.

[0057] Figure 2B A front view of a first housing 260 is shown according to an embodiment of the present specification; Figure 2C A top view of a first housing 260 is shown according to an embodiment of the present specification. As Figure 2B and Figure 2CAs shown, the first shell 260 can include a first interface 262. In some embodiments, the first shell 260 can further include a contact 266. In some embodiments, the first shell 260 can further include an angle sensor (not shown in Figure 2B and Figure 2C ).

[0058] The first interface 262 can be a mounting interface of the first shell 260 and the second shell 280. In some embodiments, the first interface 262 can be circular. The first interface 262 can be rotatably connected with the second shell 280. When the first shell 260 is mounted on the second shell 280, the first shell 260 can be rotated relative to the second shell 280 to adjust the angle of the first shell 260 relative to the second shell 280, thereby adjusting the angle of the microphone array 220.

[0059] A first magnetic device 263 can be provided on the first interface 262. The first magnetic device 263 can be provided on the first interface 262 close to the second shell 280. The first magnetic device 263 can generate a magnetic attraction force, thereby achieving detachable connection with the second shell 280. When the first shell 260 is close to the second shell 280, the first shell 260 and the second shell 280 are quickly connected by the attraction force. In some embodiments, after the first shell 260 is connected with the second shell 280, the first shell 260 can still be rotated relative to the second shell 280 to adjust the angle of the microphone array 220. Under the action of the attraction force, the connection between the first shell 260 and the second shell 280 can still be maintained when the first shell 260 is rotated relative to the second shell 280.

[0060] In some embodiments, a first positioning device (not shown in Figure 2B and Figure 2C ) can also be provided on the first interface 262. The first positioning device can be a positioning step protruding outward, or a positioning hole extending inward. The first positioning device can cooperate with the second shell 280 to achieve quick mounting of the first shell 260 and the second shell 280.

[0061] As shown in Figure 2B and Figure 2CAs shown, in some embodiments, the first housing 260 can further include contacts 266. The contacts 266 can be mounted at the first interface 262. The contacts 266 can protrude outwardly from the first interface 262. The contacts 266 can be elastically connected with the first interface 262. The contacts 266 can be communicatively connected with the M microphones 222 in the microphone array 220. The contacts 266 can be made of metal with elasticity to realize data transmission. When the first housing 260 is connected with the second housing 280, the microphone array 220 can be communicatively connected with the computing device 240 through the contacts 266. In some embodiments, the contacts 266 can be distributed in a circular manner. When the first housing 260 is connected with the second housing 280, the first housing 260 can rotate relative to the second housing 280, and the contacts 266 can also rotate relative to the second housing 280 and maintain the communicative connection with the computing device 240.

[0062] In some embodiments, an angle sensor (not shown in FIGS. 2A and 2B) can be further arranged on the first housing 260. The angle sensor can be communicatively connected with the contacts 266, thereby realizing the communicative connection with the computing device 240. The angle sensor can collect angle data of the first housing 260, thereby determining an angle at which the microphone array 220 is located, and providing reference data for subsequent calculation of the voice presence probability. Figure 2B and Figure 2C The angle sensor can be communicatively connected with the contacts 266, thereby realizing the communicative connection with the computing device 240. The angle sensor can collect angle data of the first housing 260, thereby determining an angle at which the microphone array 220 is located, and providing reference data for subsequent calculation of the voice presence probability.

[0063] Figure 2D FIG. 3 shows a front view of a second housing 280 according to an embodiment of the present specification; Figure 2E FIG. 4 shows a bottom view of the second housing 280 according to an embodiment of the present specification. As shown in FIG. 4, the second housing 280 can include a second interface 282. Figure 2D and Figure 2E As shown, the second housing 280 can include a second interface 282. In some embodiments, the second housing 280 can further include a guide rail 286.

[0064] The second interface 282 can be a mounting interface of the second housing 280 with the first housing 260. In some embodiments, the second interface 282 can be circular. The second interface 282 can be rotationally connected with the first interface 262 of the first housing 260. When the first housing 260 is mounted on the second housing 280, the first housing 260 can rotate relative to the second housing 280, thereby adjusting an angle of the first housing 260 relative to the second housing 280, and adjusting an angle of the microphone array 220.

[0065] The second interface 282 can be provided with a second magnetic device 283. The second magnetic device 283 can be provided on the second interface 282 close to the first shell 260. The second magnetic device 283 can generate a magnetic attraction force, thereby achieving detachable connection with the first interface 262. The second magnetic device 283 can be used in cooperation with the first magnetic device 263. When the first shell 260 is close to the second shell 280, the first shell 260 is quickly mounted on the second shell 280 through the attraction force between the second magnetic device 283 and the first magnetic device 263. When the first shell 260 is mounted on the second shell 280, the second magnetic device 283 is opposite to the first magnetic device 263. In some embodiments, after the first shell 260 is connected with the second shell 280, the first shell 260 can also rotate relative to the second shell 280 to adjust the angle of the microphone array 220. Under the action of the attraction force, the connection between the first shell 260 and the second shell 280 can still be maintained when the first shell 260 rotates relative to the second shell 280.

[0066] In some embodiments, the second interface 282 can also be provided with a second positioning device (not shown in FIGS. 7 and 8). The second positioning device can be a positioning step protruding outward or a positioning hole extending inward. The second positioning device can cooperate with the first positioning device of the first shell 260 to achieve quick mounting of the first shell 260 and the second shell 280. When the first positioning device is the positioning step, the second positioning device can be the positioning hole. When the first positioning device is the positioning hole, the second positioning device can be the positioning step. Figure 2D and Figure 2E In some embodiments, the second interface 282 can also be provided with a second positioning device (not shown in FIGS. 7 and 8). The second positioning device can be a positioning step protruding outward or a positioning hole extending inward. The second positioning device can cooperate with the first positioning device of the first shell 260 to achieve quick mounting of the first shell 260 and the second shell 280. When the first positioning device is the positioning step, the second positioning device can be the positioning hole. When the first positioning device is the positioning hole, the second positioning device can be the positioning step.

[0067] As Figure 2D and Figure 2EAs shown, in some embodiments, the second housing 280 may further include a guide rail 286. The guide rail 286 may be mounted at the second interface 282. The guide rail 286 may be communicatively connected to the computing device 240. The guide rail 286 may be made of metal to enable data transmission. When the first housing 260 is connected to the second housing 280, the contact 266 may contact the guide rail 286 to form a communication connection, thereby enabling the microphone array 220 to communicate with the computing device 240 for data transmission. As mentioned above, the contact 266 may be elastically connected to the first interface 262. Therefore, after the first housing 260 and the second housing 280 are connected, under the elastic force of the elastic connection, the contact 266 may fully contact the guide rail 286 to achieve a reliable communication connection. In some embodiments, the guide rail 286 may be circularly distributed. After the first housing 260 is connected to the second housing 280, when the first housing 260 rotates relative to the second housing 280, the contact 266 can also rotate relative to the guide rail 286 and maintain the communication connection with the guide rail 286.

[0068] Figure 3 A flowchart of a speech activity detection method P100 provided according to an embodiment of this specification is shown. Method P100 can calculate the probability of speech presence in the microphone signal containing a target speech signal. Specifically, processor 242 can execute method P100. Figure 3 As shown, method P100 may include:

[0069] S120: Acquire the microphone signals output by M microphones 222.

[0070] As previously mentioned, each microphone 222 can output a corresponding microphone signal. M microphones 222 correspond to M microphone signals. When calculating the probability of the presence of a target speech signal in the microphone signals, method P100 can perform the calculation based on all microphone signals among the M microphone signals, or it can perform the calculation based on a portion of the microphone signals. Therefore, the microphone signals can include either the M microphone signals corresponding to the M microphones 222 or a portion of the microphone signals. The following description in this specification will use the example of the microphone signals including the M microphone signals corresponding to the M microphones 222 as an example.

[0071] In some embodiments, the microphone signal can be a time-domain signal. For ease of description, we denote the microphone signal at time t as x(t). The microphone signal x(t) can be a signal vector composed of M microphone signals. In this case, the microphone signal x(t) can be expressed by the following formula:

[0072] x(t) = [x1(t), x2(t), ..., x M (t)]T Equation (1)

[0073] The microphone signal x(t) is a time-domain signal. In some embodiments, in step S120, the computing device 240 can also perform a spectral analysis on the microphone signal x(t). Specifically, the computing device 240 can perform a Fourier transform on the time-domain signal x(t) of the microphone signal to obtain a frequency-domain signal x f,t In the following description, the microphone signal x f,t will be described in the frequency domain. The microphone signal x f,t may be an M-dimensional signal vector composed of M microphone signals. At this time, the microphone signal x f,t may be represented as the following equation:

[0074] x f,t = [x 1,f,t , x 2,f,t , …, x M,f,t ] T Equation (2)

[0075] As mentioned previously, the microphone 222 can collect noise in the surrounding environment and output a noise signal, or can collect speech of a target user and output the target speech signal. When the target user does not emit speech, the microphone signal only contains the noise signal. When the target user emits speech, the microphone signal contains the target speech signal and the noise signal. The microphone signal x f,t may be represented as the following equation:

[0076] x f,t = Ps f,t + Qd f,t Equation (3)

[0077] where s f,t is the complex amplitude of the target speech signal. P is the target steering vector of the target speech signal. d f,t is the noise signal in the microphone signal x f,t . Q is the noise steering vector of the noise signal.

[0078] s f,t is the complex amplitude of the target speech signal. In some embodiments, there is one target speech signal source around the microphone 222. In some embodiments, there are L target speech signal sources around the microphone 222. At this time, s f,t may be an Lx1-dimensional vector. s f,t may be represented as the following equation:

[0079] s f,t = [s 1,f,t , s2,f,t ,…,s L,f,t ] T Equation (4)

[0080] The target steering vector P is an M x L matrix. The target steering vector P can be expressed as the following equation:

[0081]

[0082] where f0is the carrier frequency. d is the distance between adjacent microphones 222. c is the speed of sound. θ1, …, θ L are the incident angles between the L target speech signal sources and the microphones 222, respectively. In some embodiments, the target speech signal sources s f,t are usually distributed in a certain group of specific angle ranges. Therefore, θ1, …, θ L are known. The relative position relationship, such as the relative distance or the relative coordinates, of the M microphones 222 is pre-stored in the computing device 240. That is, the distance d between adjacent microphones 222 is pre-stored in the computing device 240. In other words, the target steering vector P can be pre-stored in the computing device 240.

[0083] The noise signal d f,t may be an M-dimensional signal vector collected by the M microphones 222. The noise signal d f,t may be expressed as the following equation: d f,t is the complex amplitude of the noise signal. In some embodiments, there are N noise signal sources around the microphone 222. At this time, d f,t may be an N x 1 vector. d f,t may be expressed as the following equation:

[0084] d f,t = [d 1,f,t , d 2,f,t , …, d N,f,t ] T Equation (5)

[0085] The noise steering vector Q is an M x N matrix. Due to the irregularity of the noise signal, the noise signal d f,t and the noise steering vector Q are unknown.

[0086] In some embodiments, the equation (3) can also be expressed as the following equation:

[0087]

[0088] where [P, Q] is the signal steering vector of the microphone signal x f,t .

[0089] As mentioned earlier, microphone 222 can collect the target speech signal s f,t It can also collect noise signals d f,t Target speech signal s f,t With noise signal d f,t These are two types of signals that are not interconnected. Therefore, the target speech signal s f,t The target subspace and noise signal d f,t The noise subspace it belongs to two unconnected subspaces. Therefore, when the microphone signal x f,t The target speech signal s does not exist in the middle. f,t At that time, microphone signal x f,t It contains only noise signal d f,t At this moment, the microphone signal x f,t The signal subspace where it is located and the target speech signal s f,t The target subspace belongs to two unconnected subspaces, and the linear correlation between the signal subspace and the target subspace is low or zero. When the microphone signal x f,t There is a target speech signal s in it. f,t At that time, the microphone signal x f,t It contains the target speech signal s f,t It also contains noise signal d f,t At this moment, the microphone signal x f,t The signal subspace where it is located and the target speech signal s f,t The target subspace belongs to two interconnected subspaces, and the signal subspace and the target subspace have a high linear correlation. Therefore, the computing device 240 can calculate the microphone signal x. f,t The signal subspace where it is located and the target speech signal s f,t The linear correlation of the target subspace determines the microphone signal x. f,t There is a target speech signal s in it. f,t The probability of speech presence. For ease of description, we define the signal subspace as span(U s We define the target subspace as span(U). t ).

[0090] like Figure 3 As shown, the method P100 may further include:

[0091] S140: Based on the microphone signal x f,t Determine the microphone signal x f,t The signal subspace span(U) constitutes s ).

[0092] In some embodiments, the signal subspace span(U) is determined.s ) of the signal subspace span(U s ) is determined. s In some embodiments, the computing device 240 can determine the basis matrix U f,t ) of the signal subspace span(U s ) based on a sample covariance matrix of the microphone signals x s . For convenience of description, we define the sample covariance matrix as M x . Figure 4 A flowchart of determining the basis matrix U s ) of the signal subspace span(U s ) according to an embodiment of the present specification is shown. Figure 4 The steps shown correspond to the computing device 240 determining the basis matrix U f,t ) of the signal subspace span(U x ) based on a sample covariance matrix M s of the microphone signals x s . As shown, the step S140 can include: Figure 4

[0093] S142: determining a sample covariance matrix M f,t of the microphone signals x f,t based on the microphone signals x x .

[0094] The sample covariance matrix M x may be represented as the following formula:

[0095] M x = x f,t x f,t H Formula (7)

[0096] S144: performing eigen decomposition on the sample covariance matrix M x to determine a plurality of eigenvectors of the sample covariance matrix M x .

[0097] Further, performing eigen decomposition on the sample covariance matrix M x to obtain a plurality of eigenvalues of the sample covariance matrix M x . The sample covariance matrix M x is an MxM dimensional matrix. The number of eigenvalues of the sample covariance matrix M x is M, and the M eigenvalues can be represented as the following formula:

[0098] λ1≥λ2≥…≥λ q = λ q+1 =…= λ M ​Formula (8)

[0099] Where 1≤q≤M.

[0100] M eigenvalues ​​correspond to M eigenvectors. For ease of description, we define the M eigenvectors as u1, u2, ..., u q ,u q+1 ,…,u M Among them, there is a one-to-one correspondence between the M eigenvalues ​​and the M eigenvectors.

[0101] S146: The matrix formed by at least some of the eigenvectors among the plurality of eigenvectors is taken as the signal subspace span(U). s The basis matrix U s .

[0102] In some embodiments, the computing device 240 may use a matrix composed of M feature vectors as the signal subspace span(U) s The basis matrix U s In some embodiments, the computing device 240 may use a matrix composed of a subset of the M feature vectors as the signal subspace span(U). s The basis matrix U s Because of λ q =λ q+1 =…=λ M Therefore, the computing device 240 can process q different eigenvalues ​​λ1, λ2, ..., λ q The corresponding q feature vectors u1, u2, ..., u q The matrix formed is used as the signal subspace span(U) s The basis matrix U s The signal subspace span(U) s The basis matrix U s Let M be a column full-rank matrix with a length of q.

[0103] In some embodiments, the computing device 240 may be based on microphone signal x f,t The signal guidance vector [P,Q] is used to determine the signal subspace span(U). s The basis matrix U s . Figure 5 Another method for determining the signal subspace span(U) according to an embodiment of this specification is shown. s The basis matrix U s The flowchart. Figure 5 The steps shown correspond to the computing device 240 based on the microphone signal x f,t The signal guidance vector [P,Q] is used to determine the signal subspace span(U). sThe basis matrix U s .like Figure 5 As shown, step S140 may include:

[0104] S148: Based on microphone signal x f,t The microphone signal x is determined by spatial estimation methods. f,t The azimuth angle of the signal source in the image is used to determine the microphone signal x. f,t The signal guidance vector [P,Q].

[0105] As mentioned earlier, the target guidance vector P can be pre-stored in the computing device 240, while the noise guidance vector Q is unknown. The computing device 240 can estimate the noise guidance vector Q based on a spatial estimation method, thereby determining the signal guidance vector [P,Q]. The spatial estimation method can include at least one of the following: the Direction of Arrival (DOA) estimation method and the spatial spectrum estimation method. The DOA estimation method can include, but is not limited to, linear spectrum estimation (such as the periodogram method), maximum likelihood spectrum estimation, maximum entropy method, MUSIC algorithm (Multiple Signal Classification algorithm), ESPRIT algorithm (Estimating signal parameters via rotational invariance techniques), etc.

[0106] S149: Determine the signal guidance vector [P,Q] as the signal subspace span(U s The basis matrix U s .

[0107] At this time, the signal subspace span(U) s The basis matrix U s It is an M×(L+N) dimensional column full-rank matrix.

[0108] like Figure 3 As shown, the method P100 may further include:

[0109] S160: Determine the target subspace span(U) formed by the target speech signal. t ).

[0110] In some embodiments, the target subspace span(U) is determined. t ) refers to determining the target subspace span(U t The basis matrix U tAs mentioned above, the computing device may pre-store the target guidance vector P. Step S160 may be: the computing device 240 determines the preset target guidance vector P corresponding to the target speech signal as the target subspace span(U). t The basis matrix U t The target subspace span(U) t The basis matrix U t It is an M×L dimensional full-rank matrix.

[0111] like Figure 3 As shown, the method P100 may further include:

[0112] S180: Based on signal subspace span(U s ) and target subspace span(U t The linear correlation of ) determines the microphone signal x. f,t The probability λ of the presence of the target speech signal is output.

[0113] In some embodiments, the computing device 240 can compute the signal subspace span(U) s ) and target subspace span(U t The volume correlation function is used to calculate the signal subspace span(U). s ) and target subspace span(U t The linear correlation of ). Figure 6 A flowchart illustrating a method for calculating the probability λ of speech presence according to an embodiment of this specification is shown. Figure 6 This shows step S180. For example... Figure 6 As shown, step S180 may include:

[0114] S182: Determine the signal subspace span(U s ) and target subspace span(U t The volume correlation function of ).

[0115] Signal subspace span(U) s ) and target subspace span(U t The volume correlation function can be the signal subspace span(U). s The basis matrix U s With the target subspace span(U) t The basis matrix U t Volume correlation function vcc(U s U t ).

[0116] When the signal subspace span(U) s The basis matrix Us Based on microphone signal x f,t The sampling covariance matrix M x For a given M×q-dimensional column full-rank matrix, the volume correlation function vcc(U) s U t This can be expressed as the following formula:

[0117]

[0118] Where, vol q (U s ) represents the M×q dimensional basis matrix U s The volume function in q dimensions. vol L (U t ) represents the M×L dimensional basis matrix U t The volume function in L dimensions. vol q (U s This can be expressed as the following formula:

[0119]

[0120] Where, γ s,i For U s The singular values ​​of γ, i = 1, ..., min{M, q}. s,1 ≥γ s,2 ≥…≥γ s,min{M,q} ≥0. vol q (U s ) can be considered as being caused by U s The volume of the polyhedron spanned by the column vectors.

[0121] vol L (U t This can be expressed as the following formula:

[0122]

[0123] Where, γ t,j For U t The singular values ​​of γ, j = 1, ..., min{M, L}. t,1 ≥γ t,2 ≥…≥γ t,min{M,L} ≥0. vol L (U t ) can be considered as being caused by U t The volume of the polyhedron spanned by the column vectors.

[0124] vol q+L ([U s U t The formula can be expressed as follows:

[0125]

[0126] in, θ p Indicated by U t Zhang Cheng's target subspace span(U) t ) and by U s Zhang Cheng's signal subspace span(U) s The main character in the space.

[0127] From formulas (9) and (12), it can be seen that the volume correlation function vcc(U) s U t In fact, it consists of two subspaces (target subspace span(U)). t ) and signal subspace span(U s The product of the spatial protagonists and the sine. Figure 7 A schematic diagram of a spatial protagonist provided according to an embodiment of this specification is shown. Figure 7 As shown, two intersecting planes (plane 1 and plane 2) in three-dimensional space share a common basis vector 3. Therefore, their first principal vector θ1 is 0. The second principal vector θ2 is equal to the angle between their respective other basis vectors (basis vectors 4 and 5). Therefore, the volume dependence function of plane 1 and plane 2 is 0.

[0128] When the signal subspace span(U) s The basis matrix U s When the signal guidance vector is [P,Q], the volume correlation function vcc(U) s U t The calculation of ) is similar to the method described above, and will not be repeated here.

[0129] Based on the above analysis, the volume correlation function vcc(U) was found to be... s U t It has the following property: when the basis matrix U s and basis matrix U t The two subspaces spanned respectively (signal subspace span(U)) s ) and target subspace span(U t When linearly correlated, i.e., the signal subspace span(U) s ) and target subspace span(U t If there are other shared basis vectors besides the zero vector, then in θ1, θ2, ..., θ min{q,L} There exists at least one spatial protagonist with a sine value of 0, at which point the volume correlation function vcc(U) s U t The value is 0. When the two subspaces are linearly independent, i.e., the signal subspace span(U) is equal to 0.s ) and target subspace span(U t ) have only zero vector, then all the principal angle sines in θ1, θ2, …, θ min{q,L} are greater than 0, and the volume correlation function vcc(U s , U t ) is greater than 0. In some embodiments, when the two subspaces are orthogonal, all the principal angle sines in θ1, θ2, …, θ min{q,L} are maximum 1, and the volume correlation function vcc(U s , U t ) reaches maximum 1.

[0130] Based on the above analysis, the volume correlation function vcc(U s , U t ) provides a measure of the linear correlation between the signal subspace span(U s ) and the target subspace span(U t ) spanned by the two basis matrices U s and U t respectively. Based on this, the problem of voice activity detection on the microphone signal x t can be transformed into the problem of the volume correlation function vcc(U s , U f,t ) of the two basis matrices U s and U t of the signal subspace span(U s ) and the target subspace span(U t ) based on the intrinsic geometric structure between the target subspace span(U s ) and the signal subspace span(U t ).

[0131] S184: Based on the volume correlation function vcc(U s , U t ), determine the linear correlation coefficient r of the signal subspace span(U s ) and the target subspace span(U t ).

[0132] From the definition and properties of the volume correlation function vcc(U s , U t ), the closer the value of the volume correlation function vcc(U s , U t ) is to 1, the lower the coincidence degree of the signal subspace span(U s ) and the target subspace span(U t ) is, and the closer the orthogonality is. The signal subspace span(Us ) and target subspace span(U t The lower the linear correlation coefficient of the microphone signal x, the stronger the signal. f,t The more likely the signal in the data comes from noise, the more these components should be suppressed. Conversely, when the volume correlation function vcc(U) is high, the signal should be suppressed. s U t The closer the value is to 0, the stronger the signal subspace span(U) is. s ) and target subspace span(U t The higher the overlap, the larger the signal subspace span(U) s ) and target subspace span(U t The higher the linear correlation coefficient of the microphone signal x, the stronger the signal x becomes. f,t The more likely the target speech signal is to dominate, the more likely these components can be preserved. Therefore, the linear correlation coefficient r and the volume correlation function vcc(U) s U t Negative correlation. Specifically, step S184 may include one of the following:

[0133] Determine the volume correlation function vcc(U) s U t If the value is greater than the first threshold 'a', determine the linear correlation coefficient t = 0; determine the volume correlation function vcc(U s U t If the value is less than the second threshold b, determine the linear correlation coefficient r = 1; determine the volume correlation function vcc(U s U t Between the first threshold a and the second threshold b, the linear correlation coefficient r is determined to be between 0 and 1, and the linear correlation coefficient r is the volume correlation function vcc(U s U t The negative correlation function of ), where the second threshold b is less than the first threshold a. Specifically, the linear correlation coefficient r can be expressed as the following formula:

[0134]

[0135] Wherein, f(vcc(U) s U t )) can be with vcc(U s U t A linearly negatively correlated function. In some embodiments, f(vcc(U) s U t This can be expressed as the following formula:

[0136]

[0137] S186: Output the linear correlation coefficient r as the probability λ of the speech presence.

[0138] In summary, in the speech activity detection system and method P100 provided in this specification, the computing device 240 can calculate the basis matrix U. s and U t Volume correlation function vcc(U s U t To calculate the signal subspace span(U) s ) and target subspace span(U t The linear correlation of ) . Volume correlation function vcc(U s U t The higher the value of ), the more it proves that the signal subspace span(U) is. s ) and target subspace span(U t The lower the linear correlation of ), the larger the signal subspace span(U) s ) and target subspace span(U t The lower the overlap between the microphone signal x and the microphone signal x, the better. f,t The lower the overlap with the target speech signal, the stronger the microphone signal x becomes. f,t The lower the probability λ of the presence of the target speech signal, the lower the probability λ of the speech signal. Volume correlation function vcc(U s U t The lower the value, the more it proves that the signal subspace span(U) is. s ) and target subspace span(U t The higher the linear correlation of ), the larger the signal subspace span(U) s ) and target subspace span(U t The higher the overlap between the microphone signal x and the signal x, the better the overlap between the microphone signal x and the signal x. f,t The higher the overlap with the target speech signal, the stronger the microphone signal x becomes. f,t The higher the probability λ of the presence of the target speech signal, the better. The speech activity detection system and method P100 provided in this specification can improve the accuracy of speech activity detection and the speech enhancement effect.

[0139] This specification also provides a voice enhancement system. The voice enhancement system can also be applied to electronic device 200. In some embodiments, the voice enhancement system may include a computing device 240. In some embodiments, the voice enhancement system can be applied to computing device 240. That is, the voice enhancement system can run on computing device 240. The voice enhancement system may include hardware devices with data processing capabilities and the necessary programs required to drive the hardware devices. Of course, the voice enhancement system may also be simply a hardware device with data processing capabilities, or simply a program running on the hardware device.

[0140] The voice enhancement system can store data or instructions for performing the voice enhancement method described in the present specification, and can execute the data and / or instructions. When the voice enhancement system is running on the computing device 240, the voice enhancement system can acquire the microphone signals from the microphone array 220 based on the communication connection, and execute the data or instructions of the voice enhancement method described in the present specification. The voice enhancement method is introduced in other parts of the present specification. For example, the voice enhancement method is introduced in the description of Figure 8 .

[0141] When the voice enhancement system is running on the computing device 240, the voice enhancement system is in communication connection with the microphone array 220. The storage medium 243 can further include at least one instruction set stored in the data storage device for performing voice enhancement computation on the microphone signals. The instructions are computer program codes, which can include programs, routines, objects, components, data structures, procedures, modules, etc. that perform the voice enhancement method provided in the present specification. The processor 242 can read the at least one instruction set and perform the voice enhancement method provided in the present specification according to the instructions of the at least one instruction set. The processor 242 can perform all steps included in the voice enhancement method.

[0142] Figure 8 A flowchart of a voice enhancement method P200 provided according to an embodiment of the present specification is shown. The method P200 can perform voice enhancement on the microphone signals. Specifically, the processor 242 can perform the method P200. As shown in Figure 8 , the method P200 can include:

[0143] S220: Acquire the microphone signals x f,t outputted by the M microphones.

[0144] As described in step S120, which will not be repeated here.

[0145] S240: Determine the voice presence probability λ that the target voice signal exists in the microphone signals x f,t based on the voice activity detection method P100.

[0146] S260: Determine the filter coefficient vector ω f,t corresponding to the microphone signals x f,t based on the voice presence probability λ.

[0147] The filter coefficient vector ω f,t may be an Mx1 dimensional vector. The filter coefficient vector ω f,t may be represented by the following formula:

[0148] ω f,t =[ω 1,f,t ,ω 2,f,t ,…,ω M,f,t ] H Formula (15)

[0149] Wherein, the filter coefficient corresponding to the m-th microphone 222 is ω m,f,t m = 1, 2, ..., M.

[0150] In some embodiments, the computing device 240 can determine the filter coefficient vector ω based on the MVDR method. f,t Specifically, step S260 may include: determining the microphone signal x based on the speech presence probability λ. f,t noise covariance matrix And based on the MVDR method and noise covariance matrix Determine the filter coefficient vector ω f,t Noise covariance matrix This can be expressed as the following formula:

[0151]

[0152] Filter coefficient vector ω f,t This can be expressed as the following formula:

[0153]

[0154] Where P is the preset target guidance vector.

[0155] In some embodiments, the computing device 240 can determine the filter coefficient vector ω based on a fixed-beam method. f,t Specifically, step S260 may include: using the speech presence probability λ as the microphone signal x f,t The target microphone signal x k,f,t The corresponding filter coefficient ω k,f,t Determine the microphone signal x f,t Mid-target microphone signal x k,f,t The filter coefficients for all other microphone signals are 0. Where k = 1, 2, ..., M. The filter coefficient vector ω f,t Including target microphone signal x k,f,t The corresponding filter coefficient ω k,t,t And a vector consisting of the filter coefficients corresponding to the remaining microphone signals. Filter coefficient vector ω f,t This can be expressed as the following formula:

[0156] ω f,t =[0,0,…,ω k,f,t =λ,…,0]H Equation (18)

[0157] In some embodiments, the target microphone signal x k,f,t may be the microphone signal x f,t with the highest signal-to-noise ratio. In some embodiments, the target microphone signal x k,f,t may be the microphone signal x f,t closest to the target speech signal source (such as a mouth). In some embodiments, k can be pre-stored in the computing device 240.

[0158] S280: merging the microphone signals x f,t based on the filter coefficient vector ω f,t to obtain a target audio signal y f,t and outputting.

[0159] The target audio signal y f,t may be represented by the following equation:

[0160] y f,t H x f,t Equation (21)

[0161] In summary, the speech activity detection system and method P100 and the speech enhancement system and method P200 provided by the present specification are used for a microphone array 220 composed of multiple microphones 222. The speech activity detection system and method P100 and the speech enhancement system and method P200 can obtain microphone signals x f,t collected by the microphone array 220. The microphone signals x f,t may include noise signals and may also include target speech signals. The target speech signals and the noise signals belong to two non-intersecting signals. The target subspace span(U t ) in which the target speech signals are located and the noise subspace in which the noise signals are located belong to two non-intersecting subspaces. When there is no target speech signal in the microphone signals x f,t , the microphone signals x f,t only contain noise signals. At this time, the signal subspace span(U f,t ) in which the microphone signals x s are located and the target subspace span(U t ) in which the target speech signals are located belong to two non-intersecting subspaces, and the linear correlation of the signal subspace span(U s ) and the target subspace span(U t ) is low. When there is a target speech signal in the microphone signals x f,t , the microphone signals x f,tIt contains both the target speech signal and the noise signal. At this time, the microphone signal x f,t The signal subspace span(U) s ) and the target subspace span(U) where the target speech signal is located t ) belongs to two interconnected subspaces, the signal subspace span(U) S ) and target subspace span(U t The linear correlation of the microphone signal x is relatively high. Therefore, the speech activity detection method P100 and system provided in this specification can calculate the microphone signal x. f,t The signal subspace span(U) s ) and the target subspace span(U) where the target speech signal is located t The linear correlation is used to determine the microphone signal x. f,t The probability λ of the presence of the target speech signal is given. The speech enhancement method P200 and system can calculate the filter coefficient vector ω based on the speech presence probability λ. f,t Thus, the microphone signal x f,t Speech enhancement is performed. In summary, the speech activity detection system and method P100 and the speech enhancement system and method P200 provided in this specification can improve the calculation accuracy of the speech presence probability λ, thereby improving the speech enhancement effect.

[0162] In another aspect of the disclosure, a non-transitory storage medium storing at least one set of instructions executable by a processor to perform speech activity detection is provided. When executed by the processor, the instructions direct the processor to implement the steps of the speech activity detection method P100 described in the present disclosure. In some possible implementations, various aspects of the present disclosure can also be implemented as a program product in the form of a computer readable medium having program code portions stored therein. When the program product is run on a computing device (such as the computing device 240), the program code portions cause the computing device to perform the steps of speech activity detection described in the present disclosure. The program product for implementing the above-described methods can include the program code portions on a portable compact disc read-only memory (CD-ROM) and can be run on the computing device. However, the program product of the present disclosure is not limited to this, and in the present disclosure, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system (such as the processor 242). The program product can take any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. More specific examples of the readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer readable storage medium can include a data signal carried in a baseband or as part of a carrier wave, in which readable program code is borne. Such a propagated data signal can take on many forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The readable storage medium can also be any readable medium that is not a storage medium that can be readable by a general purpose computing device. The readable storage medium can transmit, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, or any suitable combination of the above. The program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the computing device, partly on the computing device, as a stand-alone software package, partly on the computing device and partly on a remote computing device, or entirely on the remote computing device.

[0163] The above described embodiments of the disclosure have been described. Other embodiments are within the scope of the following claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still accomplish desirable results. Additionally, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.

[0164] In light of the above, those skilled in the art will appreciate that the foregoing detailed description of the present disclosure is susceptible to various modifications and / or revisions without departing from the spirit and scope of the present disclosure. Although specific embodiments have been illustrated and described herein, it should be appreciated that any arrangement which is calculated to achieve the same purpose can be substituted for the specific embodiments shown. This disclosure is intended to cover any adaptations or variations of the present disclosure. Therefore, it is intended that the application be protected by: the broadest interpretation of the appended claims to take into account unforeseen equivalents and alternatives based on current knowledge, or future knowledge.

[0165] In addition, certain terminology has been used to describe embodiments of the disclosure. For example, "one embodiment," "an embodiment," and / or "some embodiments” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosure. Therefore, it is understood that the use of "embodiment" or "one embodiment" or "an embodiment" or "some embodiments” in various places throughout this specification are not necessarily referring to the same embodiment. Furthermore, the particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0166] It should be understood that in the foregoing description of embodiments of the disclosure, various features are sometimes grouped together in a single embodiment, figure, or description of a figure for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various aspects, embodiments, and / or features. However, this should not be interpreted as a requirement that these features must be provided together in order to form an embodiment of the disclosure. In fact, some embodiments of the disclosure can provide only a subset of these features. In addition, streams of features from different embodiments can be combined to provide an embodiment of the disclosure. In some cases, features from one embodiment can be combined with features from another embodiment to provide an embodiment of the disclosure. In some cases, features from one embodiment can be combined with features from another embodiment to provide an embodiment of the disclosure.

[0167] Every patent, patent application, publication of a patent application, and other material, for example articles, books, specifications, publications, documents, things, or the like which can be cited in this document is either incorporated by reference in its entirety for all purposes or can be cited for its relevant art-disclosing

[0168] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the present specification. Other modifications that fall within the scope of the present specification can also be made. Accordingly, embodiments disclosed in the present specification are to be considered as illustrative and not restrictive, and not limiting the scope of the application. As such, the embodiments disclosed in the present specification are susceptible to alternative configurations and embodiments being employed as would be known to those skilled in the art. The embodiments disclosed in the present specification are not intended to limit the scope of the application but rather are presented as illustrative embodiments of the application.

Claims

1. A method for detecting speech activity, characterized in that, For M microphones distributed in a preset array shape, where M is an integer greater than 1, including: Obtain the microphone signals output by the M microphones; Based on the microphone signals, determine the signal subspace formed by the microphone signals; Determine the target subspace formed by the target speech signal; Determine the volume correlation function between the signal subspace and the target subspace; Based on the volume correlation function, the linear correlation coefficient between the signal subspace and the target subspace is determined; and The linear correlation coefficient is used as the speech presence probability of the target speech signal, and the speech presence probability is output.

2. The speech activity detection method as described in claim 1, characterized in that, Determining the signal subspace formed by the microphone signals based on the microphone signals includes: Based on the microphone signal, determine the sampling covariance matrix of the microphone signal; Perform eigenvalue decomposition on the sampling covariance matrix to determine multiple eigenvectors of the sampling covariance matrix; and The matrix formed by at least some of the eigenvectors among the plurality of eigenvectors shall be used as the basis matrix of the signal subspace.

3. The speech activity detection method as described in claim 1, characterized in that, Determining the signal subspace formed by the microphone signals based on the microphone signals includes: Based on the microphone signal, the azimuth angle of the signal source in the microphone signal is determined by a spatial estimation method, thereby determining the signal guidance vector of the microphone signal. The spatial estimation method includes at least one of a DOA estimation method and a spatial spectrum estimation method. The signal guiding vector is determined to be the basis matrix of the signal subspace.

4. The speech activity detection method as described in claim 1, characterized in that, The determination of the target subspace constituted by the target speech signal includes: The preset target guidance vector corresponding to the target speech signal is determined as the basis matrix of the target subspace.

5. The speech activity detection method as described in claim 1, characterized in that, The linear correlation coefficient is negatively correlated with the volume correlation function; Wherein, determining the linear correlation coefficient between the signal subspace and the target subspace based on the volume correlation function includes one of the following cases: If the volume correlation function is determined to be greater than a first threshold, the linear correlation coefficient is determined to be 0. The volume correlation function is determined to be less than a second threshold, and the linear correlation coefficient is determined to be 1, wherein the second threshold is less than the first threshold; and The volume correlation function is determined to be between the first threshold and the second threshold, the linear correlation coefficient is determined to be between 0 and 1, and the linear correlation coefficient is a negative correlation function of the volume correlation function.

6. A voice activity detection system, characterized in that, include: At least one storage medium storing at least one instruction set for voice activity detection; as well as At least one processor is communicatively connected to the at least one storage medium. When the voice activity detection system is running, the at least one processor reads the at least one instruction set and implements the voice activity detection method according to any one of claims 1-5.

7. A speech enhancement method, characterized in that, For M microphones distributed in a preset array shape, where M is an integer greater than 1, including: Obtain the microphone signals output by the M microphones; Based on the voice activity detection method according to any one of claims 1-5, determine the probability of the presence of the target voice signal in the microphone signal; The filter coefficient vector corresponding to the microphone signal is determined based on the probability of the presence of the speech; and The microphone signals are combined based on the filter coefficient vector to obtain the target audio signal and output it.

8. The speech enhancement method as described in claim 7, characterized in that, Determining the filter coefficient vector corresponding to the microphone signal based on the probability of speech presence includes: The noise covariance matrix of the microphone signal is determined based on the probability of the speech presence; and The filter coefficient vector is determined based on the MVDR method and the noise covariance matrix.

9. The speech enhancement method as described in claim 7, characterized in that, Determining the filter coefficient vector corresponding to the microphone signal based on the probability of speech presence includes: The probability of speech presence is used as the filtering coefficient corresponding to the target microphone signal in the microphone signal, wherein the target microphone signal includes the microphone signal with the highest signal-to-noise ratio among the microphone signals; and The filter coefficients for all microphone signals other than the target microphone signal in the microphone signal are determined to be 0. The filter coefficient vector includes a vector consisting of the filter coefficients corresponding to the target microphone signal and the filter coefficients corresponding to the other microphone signals.

10. A speech enhancement system, characterized in that, include: At least one storage medium storing at least one instruction set for speech enhancement; as well as At least one processor is communicatively connected to the at least one storage medium. When the speech enhancement system is running, the at least one processor reads the at least one instruction set and implements the speech enhancement method according to any one of claims 7-9.

Citation Information

Patent Citations

  • Voice equipment DOA estimation enhancement method and device

    CN108538306A

  • Voice activity detection method and system, and voice enhancement method and system

    CN116982112A

  • Microphone array, method to process signals from this microphone array and speech recognition method and system using the same

    EP1473964A2