METHOD FOR CALCULATING VOICE PRESENCE PROBABILITY, SYSTEM FOR CALCULATING VOICE PRESENCE PROBABILITY, METHOD FOR ENHANCEMENT OF VOICE, SYSTEM FOR ENHANCEMENT OF VOICE, AND EARPHONES

By iteratively optimizing speech presence and absence models using maximum likelihood estimation and EM algorithms, the method addresses the low accuracy in speech presence probability estimation in MVDR algorithms, leading to improved sound enhancement in devices with small microphone arrays.

JP7672737B2Active Publication Date: 2025-05-08ショックス·ヒアリング·ピーティーイー·リミテッド
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023542599
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-10-11
Publication Date
2025-05-08
Estimated Expiration
2041-10-11

AI Technical Summary

Technical Problem

Existing speech enhancement techniques using the Minimum Variance Distortionless Response (MVDR) algorithm suffer from low accuracy in estimating speech presence probability, leading to poor noise covariance matrix estimation and subsequently, ineffective sound enhancement, especially in devices with small numbers of microphones.

Method used

A method for calculating the probability of speech presence with higher accuracy using a system that iteratively optimizes speech presence and absence models based on maximum likelihood estimation and EM algorithms, determining the model parameters by comparing the entropy of the probabilities, and converging to improve the estimation of speech existence and noise covariance matrix.

Benefits of technology

The proposed method significantly enhances the accuracy of speech presence probability estimation and noise covariance matrix calculation, thereby improving the sound enhancement effect of the MVDR algorithm, even in devices with small microphone arrays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672737000294
    Figure 0007672737000294
  • Figure 0007672737000295
    Figure 0007672737000295
  • Figure 0007672737000296
    Figure 0007672737000296
Patent Text Reader

Abstract

The present specification provides a method for calculating the presence of voice probability, a system for calculating the presence of voice probability, a speech enhancement method, a speech enhancement system, and an earphone, which correct the presence of voice probability and the absence of voice probability in an iterative process by comparing the entropy of the presence of voice probability and the entropy of the absence of voice probability, thereby obtaining a faster convergence speed and a better convergence result, improving the estimation accuracy of the presence of voice probability and the noise covariance matrix, and further improving the speech enhancement effect of MVDR.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present specification relates to the technical field of processing of speech signals, and in particular to a method for calculating a voice presence probability, a system for calculating a voice presence probability, a method for speech enhancement, a system for speech enhancement and an earphone. [Background technology]

[0002] In the voice enhancement technology based on the beamforming algorithm, particularly in the adaptive beamforming algorithm of the Minimum Variance Distortionless Response (MVDR), how to obtain the noise covariance matrix, which is a parameter describing the noise statistical characteristic relationship between different microphones, is very important. The main method in the prior art is to calculate the noise covariance matrix based on the method of voice presence probability, estimate the voice presence probability by, for example, a voice activity detection method (VAD), and then calculate the noise covariance matrix. However, since the estimation accuracy of the voice presence probability in the prior art is not sufficient, the estimation accuracy of the noise covariance matrix is ​​low, and the voice enhancement effect of the MVDR algorithm is low. In particular, when the number of microphones is small, for example, less than five, the effect drops sharply. Therefore, the MVDR algorithm in the prior art is often used in microphone array devices with a large number of microphones and large spacing, such as mobile phones and smart speakers, and in devices with a small number of microphones and small spacing, such as earphones, the voice enhancement effect is low. Summary of the Invention [Problem to be solved by the invention]

[0003] Therefore, there is a need to provide a more accurate voice presence probability calculation method, voice presence probability calculation system, voice enhancement method, voice enhancement system and earphone. [Means for solving the problem]

[0004] The present specification provides a voice presence probability calculation method, a voice presence probability calculation system, a voice enhancement method, a voice enhancement system, and an earphone with higher accuracy.

[0005] In a first aspect, a method for calculating a voice presence probability according to the present specification is used for M microphones, where M is an integer greater than 1, distributed in a predetermined array shape, and includes the steps of acquiring microphone signals output from the M microphones, where the microphone signals follow a first model or a second model of a Gaussian distribution, one of the first model and the second model being a voice presence model and the other being an absence of voice model; iteratively optimizing the first model and the second model respectively based on maximum likelihood estimation and an EM algorithm until convergence, and in the iterative process, determining whether the voice presence model is the first model or the second model based on an entropy of a first probability when the microphone signal is the first model and an entropy of a second probability when the microphone signal is the second model, where the first probability and the second probability are complementary; and when the maximum likelihood estimation and the EM algorithm converge, outputting a probability that the microphone signal is the voice presence model as a voice presence probability of the microphone signal.

[0006] In some embodiments, a first variance of a Gaussian distribution corresponding to the first model comprises a product of a first parameter and a first spatial covariance matrix, and a second variance of a Gaussian distribution corresponding to the second model comprises a product of a second parameter and a second spatial covariance matrix.

[0007] In some embodiments, the step of iteratively optimizing the first model and the second model based on maximum likelihood estimation and an EM algorithm includes the steps of: constructing an objective function based on maximum likelihood estimation and an EM algorithm; determining optimization parameters including the first spatial covariance matrix and the second spatial covariance matrix; determining initial values ​​of the optimization parameters; iterating the optimization parameters based on the objective function and the initial values ​​of the optimization parameters a plurality of times until the objective function converges, wherein in the plurality of iterations, determining whether the voice presence model is the first model or the second model based on an entropy of the first probability and an entropy of the second probability; and outputting a convergence value of the optimization parameters and the first probability and the second probability corresponding to the convergence value.

[0008] In some embodiments, the step of determining whether the voice presence model is the first model or the second model based on the entropy of the first probability and the entropy of the second probability in the multiple iterations comprises the step of calculating, in any one of the multiple iterations, the entropy of the first probability and the entropy of the second probability and determining whether the voice presence model is the first model or the second model, wherein if it is determined that the entropy of the first probability is greater than the entropy of the second probability, it is determined that the voice presence model is the second model, or if it is determined that the entropy of the first probability is less than the entropy of the second probability, it is determined that the voice presence model is the first model.

[0009] In some embodiments, the step of determining whether the voice presence model is the first model or the second model based on the entropy of the first probability and the entropy of the second probability in the multiple iterations comprises the step of calculating the entropy of the first probability and the entropy of the second probability in a first iteration of the multiple iterations and determining whether the voice presence model is the first model or the second model, wherein if it is determined that the entropy of the first probability is greater than the entropy of the second probability, it is determined that the voice presence model is the second model, or if it is determined that the entropy of the first probability is less than the entropy of the second probability, it is determined that the voice presence model is the first model.

[0010] In some embodiments, the step of iterating the optimization parameters a plurality of times further comprises the steps of: in each iteration of the plurality of iterations, correcting the first probability and the second probability based on an entropy of the first probability and an entropy of the second probability, wherein if it is determined that the first model is the voice presence model and the entropy of the first probability is greater than the entropy of the second probability, replacing a value corresponding to the first probability with a value corresponding to the second probability, or if it is determined that the second model is the voice presence model and the entropy of the second probability is greater than the entropy of the first probability, replacing a value corresponding to the first probability with a value corresponding to the second probability; and updating the optimization parameters based on the corrected first probability and the second probability.

[0011] In some embodiments, the step of iterating the optimization parameters multiple times further includes a step of performing an invertible correction on the optimization parameters in each iteration of the multiple iterations, and when it is determined that the optimization parameters are non-invertible, a step of correcting the optimization parameters based on a deviation matrix including one of an identity matrix and a random matrix following a normal distribution or a uniform distribution.

[0012] In a second aspect, a system for computing voice presence probability according to the present specification includes at least one storage medium and at least one processor, wherein the at least one storage medium stores at least one instruction set for computing voice presence probability, and the at least one processor is communicatively connected to the at least one storage medium, and when the system for computing voice presence probability is executed, the at least one processor reads the at least one instruction set and performs the method for computing voice presence probability according to the first aspect of the present specification.

[0013] In a third aspect, a speech enhancement method according to the present specification is used for M microphones, where M is an integer greater than 1, distributed in a predetermined array shape, and includes the steps of acquiring microphone signals output from the M microphones, determining the speech presence probability of the microphone signals based on the speech presence probability calculation method described in the first aspect of the present specification, determining a noise covariance matrix of the microphone signals based on the speech presence probability, determining filter coefficients corresponding to the microphone signals based on an MVDR method and the noise spatial covariance matrix, and combining the microphone signals based on the filter coefficients to output a target audio signal.

[0014] In a fourth aspect, a voice enhancement system according to the present specification includes at least one storage medium and at least one processor, wherein the at least one storage medium stores at least one instruction set for voice enhancement, the at least one processor is communicatively connected to the at least one storage medium, and when the voice enhancement system is executed, the at least one processor reads the at least one instruction set and executes the voice enhancement method described in the third aspect of the present specification.

[0015] In a fifth aspect, an earphone according to the present specification includes a microphone array and a computing device, the microphone array including M microphones, where M is an integer greater than 1, distributed in a predetermined array shape, the computing device being communicatively connected to the microphone array in operation to perform the speech enhancement method according to the third aspect of the present specification.

[0016] In some embodiments, the M microphones are linearly distributed, M is less than or equal to 5, and the spacing between adjacent ones of the M microphones is between 20mm and 40mm.

[0017] In some embodiments, the earphone further includes a first housing and a second housing, the microphone array is attached to the first housing, the first housing includes a first connection port at which a first magnetic device is installed, the computing device is attached to the second housing, the second housing includes a second connection port at which a second magnetic device is installed, and the first housing and the second housing are removably connected by an adhesive force between the first magnetic device and the second magnetic device.

[0018] In some embodiments, the first housing further includes a contact point mounted on the first connection port and communicatively connected to the microphone array, and the second housing further includes a guide rail mounted on the second connection port and communicatively connected to the computing device, and when the first housing is connected to the second housing, the contact point contacts the guide rail, thereby communicatively connecting the microphone array to the computing device.

[0019] As can be seen from the above technical solutions, the method for calculating the sound presence probability, the system for calculating the sound presence probability, the method for enhancing the voice, the system for enhancing the voice, and the earphones according to the present specification are used in a microphone array consisting of multiple microphones. Each microphone in the microphone array can collect audio of multiple sound sources in a space and output corresponding microphone signals. The audio signals of each sound source follow a Gaussian distribution. The multiple microphone signals output from the multiple microphone array follow a joint Gaussian distribution. In order to obtain the voice presence probability in the multiple microphone signals, the voice presence probability calculation method, voice presence probability calculation system, voice enhancement method, voice enhancement system and earphone respectively obtain a voice presence model when voice is present in the multiple microphone signals and an absence of voice model when voice is not present, and perform multiple iterative optimization based on maximum likelihood estimation and EM algorithm, and correct the voice presence probability and the absence of voice probability according to the entropy of the voice presence probability and the entropy of the absence of voice probability in the iterative process, thereby calculating and determining model parameters of the voice presence model and the absence of voice probability, and when the maximum likelihood estimation and EM algorithm converge, the voice presence probability corresponding to the voice presence model can be obtained. The present invention relates to a speech presence probability calculation method, a speech presence probability calculation system, a speech enhancement method, a speech enhancement system, and an earphone, and more particularly, to a speech presence probability calculation method and a speech presence probability calculation system. The present invention relates to a speech presence probability calculation method and a speech presence probability calculation system.

[0020] Other functions of the voice presence probability calculation method, the voice presence probability calculation system, the voice enhancement method, the voice enhancement system, and the earphones according to the present specification are partially listed in the following description. According to the description, the following figures and exemplary description contents are obvious to those skilled in the art. The inventive step of the voice presence probability calculation method, the voice presence probability calculation system, the voice enhancement method, the voice enhancement system, and the earphones according to the present specification can be fully understood by practicing or using the methods, devices, and combinations described in the following detailed examples.

[0021] In order to more clearly describe the technical means in the embodiments of the present specification, the following briefly introduces drawings necessary for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present specification, and those skilled in the art can also obtain other drawings based on these drawings without creative efforts. [Brief description of the drawings]

[0022] [Figure 1] 1 shows a hardware schematic diagram of a voice presence probability calculation system according to an embodiment of the present specification. [Figure 2A] 1 shows a schematic exploded configuration diagram of an electronic device according to an embodiment of the present specification. [Figure 2B] 1 illustrates a front view of a first housing according to an embodiment of the present disclosure. [Figure 2C] FIG. 2 illustrates a top view of a first housing according to an embodiment of the present disclosure. [Figure 2D] 1 illustrates a front view of a second housing according to an embodiment of the present disclosure. [Figure 2E] FIG. 2 illustrates a bottom view of a second housing according to an embodiment of the present disclosure. [Diagram 3] 2 shows a flowchart of a method for calculating a voice presence probability according to an embodiment of the present specification; [Figure 4] 1 shows a flow chart of an iterative optimization according to an embodiment of the present specification. [Diagram 5] 1 shows a flow chart of multiple iterations according to an embodiment of the present disclosure. [Figure 6]1 shows a flow chart of another multiple iteration according to an embodiment of the present disclosure. [Figure 7] 1 shows a flowchart of a speech enhancement method according to an embodiment of the present specification. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0023] In order to enable those skilled in the art to implement and use the contents of this specification, the following describes specific applications and requirements of this specification. Various partial modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not intended to be limited to the embodiments shown, but should be accorded the broadest scope consistent with the claims.

[0024] The terms used herein are for the purpose of describing particular example embodiments only and are not limiting. For example, the singular forms "a", "one" and "the" as used herein may include the plural unless the context clearly dictates otherwise. As used herein, the terms "comprise", "comprise" and / or "comprising" refer to the presence of associated integers, steps, operations, elements and / or assemblies, but do not preclude the presence of one or more other features, integers, steps, operations, elements, assemblies and / or groups, or the addition of other features, integers, steps, operations, elements, assemblies and / or groups to the system / method.

[0025] These and other features of the present specification, the operation and function of the associated elements of structure, and the combination of parts and economies of manufacture can be significantly improved upon consideration of the following description. All drawings form part of this specification. It is to be expressly understood, however, that the drawings are for the purpose of illustration and description only and are not intended to limit the scope of the present specification. It is to be understood that the drawings are not drawn to scale.

[0026] As used herein, flowcharts illustrate operations performed by systems according to embodiments of the present disclosure. It should be clearly understood that the operations in the flowcharts do not have to be performed in sequential order. Operations may be performed in reverse order or simultaneously. Also, one or more other operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.

[0027] For the sake of convenience, the terms used in the specification are first explained as follows.

[0028] Minimum Variance Distortionless Response (MVDR) is an adaptive beamforming algorithm based on the maximum signal-to-interference-plus-noise ratio (SINR) criterion. The MVDR algorithm can adaptively minimize the power of the array output in the desired direction and maximize the signal-to-interference-plus-noise ratio. The goal is to minimize the variance of the recorded signal. If the noise signal and the desired signal are uncorrelated, the variance of the recorded signal is the sum of the variances of the desired signal and the noise signal. Therefore, the MVDR solution seeks to minimize this sum, thereby reducing the effect of the noise signal. The principle is to select appropriate filter coefficients to minimize the average power of the array output under the constraint that the desired signal is distortion-free.

[0029] Voice presence probability: The probability that the target voice signal is present in the current audio signal.

[0030] Gaussian distribution: Normal distribution, also known as Gaussian distribution. The normal distribution curve is bell-shaped, low at both ends, high in the middle, and symmetrical. Because the curve is bell-shaped, it is often called a bell-shaped curve. The random variable X has a mathematical expectation μ and a variance σ 2 If the normal distribution of N(μ,σ 2) The position of the probability density function is determined by the expected value μ of the normal distribution, and the width of the distribution is determined by its standard deviation σ. When μ=0 and σ=1, the normal distribution is the standard normal distribution.

[0031] 1 shows a hardware schematic diagram of a sound presence probability calculation system according to an embodiment of the present specification. The sound presence probability calculation system can be applied in an electronic device 200.

[0032] In some embodiments, the electronic device 200 may be a wireless earphone, a wired earphone, a smart wearable device, such as a device with audio processing capabilities, such as smart glasses, a smart helmet, or a smart watch. The electronic device 200 may be a mobile device, a tablet computer, a laptop, an automobile built-in device, or the like, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, or the like, or any combination thereof. For example, the smart mobile device may include a mobile phone, a personal digital assistant, a gaming device, a navigation device, an ultra-mobile personal computer (UMPC), or any combination thereof. In some embodiments, the smart home device may include a smart television, a desktop personal computer, or the like, or any combination thereof. In some embodiments, the automobile built-in device may include an in-car computer, an in-car television, or the like.

[0033] In this specification, the electronic device 200 is taken as an example of an earphone. The earphone may be a wireless earphone or a wired earphone. As shown in FIG. 1, the electronic device 200 may include a microphone array 220 and a computing device 240.

[0034] The microphone array 220 may be an audio collecting device of the electronic device 200. The microphone array 220 may be configured to capture local audio and output a microphone signal, which is an electronic signal carrying audio information. The microphone array 220 may include M microphones 222 distributed in a predetermined array shape. The M is an integer greater than 1. The M microphones 222 may be uniformly or non-uniformly distributed. The M microphones 222 may output microphone signals. The M microphones 222 may output M microphone signals. Each microphone 222 corresponds to one microphone signal. The M microphone signals are collectively referred to as the microphone signals. In some embodiments, the M microphones 222 may be linearly distributed. In some embodiments, the M microphones 222 may be distributed to be an array of other shapes, such as a circular array, a rectangular array, etc. For convenience of explanation, the following description will be given taking the case where the M microphones 222 are linearly distributed as an example. In some embodiments, M may be an integer greater than 1, such as 2, 3, 4, 5, or even greater. In some embodiments, M may be an integer greater than 1 and less than or equal to 5, due to space limitations, for example, in products such as earphones. When the electronic device 200 is an earphone, the spacing between adjacent microphones 222 among the M microphones 222 may be 20 mm to 40 mm. In some embodiments, the spacing between adjacent microphones 222 may be smaller, for example, 10 mm to 20 mm.

[0035] In some embodiments, the microphone 222 may be a bone conduction microphone that directly collects vibration signals from the human body. The bone conduction microphone may include a vibration sensor such as an optical vibration sensor, an acceleration sensor, etc. The vibration sensor can collect mechanical vibration signals (e.g., signals generated by vibrations of the skin or bones when a user speaks) and convert the mechanical vibration signals into electrical signals. In this specification, mechanical vibration signals mainly refer to vibrations that are transmitted through solid objects. The bone conduction microphone contacts the skin or bones of the user via the vibration sensor or a vibrating member connected to the vibration sensor, thereby collecting vibration signals generated by the bones or skin when the user produces a sound, and converting the vibration signals into electrical signals. In some embodiments, the vibration sensor may be a device that is sensitive to mechanical vibrations and not sensitive to air vibrations (i.e., the response capability of the vibration sensor to mechanical vibrations exceeds the response capability of the vibration sensor to air vibrations). The bone conduction microphone can directly pick up vibration signals from the sound generating site, thereby reducing the influence of environmental noise.

[0036] In some embodiments, the microphone 222 may be an air conduction microphone that directly collects air vibration signals generated when a user speaks and converts the air vibration signals into electrical signals.

[0037] In some embodiments, the M microphones 220 may be M bone conduction microphones. In some embodiments, the M microphones 220 may be M air conduction microphones. In some embodiments, the M microphones 220 may include bone conduction microphones or air conduction microphones. Of course, the microphones 222 may be other types of microphones, such as optical microphones, microphones that receive electromyographic signals, etc.

[0038] The computing device 240 may be communicatively connected to the microphone array 220. The communicative connection may be any form of connection capable of receiving information directly or indirectly. In some embodiments, the computing device 240 may communicate data with the microphone array 220 via a wireless communication connection, in some embodiments, the computing device 240 may be directly connected to the microphone array 220 via wires to communicate data with each other, and in some embodiments, the computing device 240 may be directly connected to other circuits via wires to establish an indirect connection with the microphone array 220 to communicate data with each other. In this specification, the case where the computing device 240 is directly connected to the microphone array 220 via wires is described as an example.

[0039] The computing device 240 may be a hardware device having the capability of processing data information. In some embodiments, the voice presence probability computation system may include the computing device 240. In some embodiments, the voice presence probability computation system may be applied to the computing device 240. That is, the voice presence probability computation system may be executed in the computing device 240. The voice presence probability computation system may include a hardware device having the capability of processing data information and a program required to drive the operation of the hardware device. Of course, the voice presence probability computation system may be only a hardware device having data processing capabilities, or only a program executed on the hardware device.

[0040] The voice presence probability calculation system stores data or instructions for performing the voice presence probability calculation method described herein, and can execute the data and / or instructions. When the voice presence probability calculation system is executed in the computing device 240, the voice presence probability calculation system can obtain the microphone signal from the microphone array 220 based on the communication connection, execute data or instructions of the voice presence probability calculation method described herein, and calculate the voice presence probability in the microphone signal. The voice presence probability calculation method is introduced in other parts of this specification. For example, the description of Figures 3 to 6 introduces the voice presence probability calculation method.

[0041] 1, the computing device 240 may include at least one storage medium 243 and at least one processor 242. In some embodiments, the electronic device 200 may further include a communication port 245 and an internal communication bus 241.

[0042] The internal communication bus 241 may connect to different system assemblies including a storage medium 243 , a processor 242 and a communication port 245 .

[0043] The communication port 245 may be used for data communication between the computing device 240 and the outside world. For example, the computing device 240 may obtain the microphone signals from the microphone array 220 via the communication port 245.

[0044] At least one storage medium 243 may include a data storage device. The data storage device may be a non-transitory storage medium or a transitory storage medium. For example, the data storage device may include one or more of a magnetic disk, a read-only storage medium (ROM), and a random access storage medium (RAM). When the voice presence probability calculation system is executed in the computing device 240, the storage medium 243 may include at least one set of instructions stored in the data storage device for calculating the voice presence probability of the microphone signal. The instructions are computer program code, and the computer program code may include programs, routines, objects, assemblies, data structures, processes, modules, etc., for performing the voice presence probability calculation method according to the present disclosure.

[0045] The at least one processor 242 may be communicatively connected to at least one storage medium 243 via an internal communication bus 241. The communication connection may be any form of connection capable of receiving information directly or indirectly. The at least one processor 242 executes the at least one instruction set. When the voice presence probability calculation system is executed in the computing device 240, the at least one processor 242 reads the at least one instruction set and executes the voice presence probability calculation method according to the present disclosure based on the instructions of the at least one instruction set. The processor 242 may execute all steps included in the voice presence probability calculation method. The processor 242 may be in the form of one or more processors, and in some embodiments, the processor 242 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physical processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuitry and processor capable of performing one or more functions, or any combination thereof. For purposes of illustration only, only one processor 242 is described herein in the computing device 240. However, the computing device 240 herein may include multiple processors 242. Thus, the operations and / or steps of the methods disclosed herein may be performed by one processor or jointly by multiple processors, as described herein.For example, where herein the processor 242 of the computing device 240 performs steps A and B, it should be understood that steps A and B may be performed jointly or separately by two different processors 242 (e.g., a first processor performs step A and a second processor performs step B, or the first and second processors both perform steps A and B).

[0046] 2A illustrates an exploded schematic diagram of an electronic device 200 according to an embodiment of the present disclosure. As shown in FIG. 2A, the electronic device 200 may include a microphone array 220, a computing device 240, a first housing 260, and a second housing 280.

[0047] The first housing 260 may be a mounting base for the microphone array 220. The microphone array 220 may be mounted inside the first housing 260. The shape of the first housing 260 may be adaptively designed according to the distribution shape of the microphone array 220, and this specification is not limited thereto. The second housing 280 may be a mounting base for the computing device 240. The computing device 240 may be mounted inside the second housing 280. The shape of the second housing 280 may be adaptively designed according to the shape of the computing device 240, and this specification is not limited thereto. If the electronic device 200 is an earphone, the second housing 280 may be connected to a wearing site. The second housing 280 may be connected to the first housing 260. As described above, the microphone array 220 may be electrically connected to the computing device 240. In particular, the microphone array 220 may be electrically connected to the computing device 240 by a connection between the first housing 260 and the second housing 280 .

[0048] In some embodiments, the first housing 260 may be fixedly connected to the second housing 280, for example, by integral molding, welding, riveting, or bonding. In some embodiments, the first housing 260 may be removably connected to the second housing 280. The computing device 240 may be communicatively connected to different microphone arrays 220. Specifically, the different microphone arrays 220 may have different numbers of microphones 222 for the microphone array 220, different array shapes, different intervals between the microphones 222, different mounting angles on the first housing 260, different mounting positions on the first housing 260, etc. A user may replace the corresponding microphone array 220 according to application scenarios to adapt the electronic device 200 to a wider range of scenarios. For example, when the distance between the user and the electronic device 200 is close in the application scenario, the user may replace the microphone array 220 with a microphone array 220 with a smaller interval. Also for example, if the application scenario involves a close distance between the user and the electronic device 200, the user may switch to a microphone array 220 with greater spacing and greater number.

[0049] The releasable connection may be any type of physical connection, such as a threaded connection, a snap-fit ​​connection, a magnetic attraction connection, etc. In some embodiments, the first housing 260 and the second housing 280 may be magnetically connected, i.e., the first housing 260 and the second housing 280 are releasably connected by the attraction force of a magnetic device.

[0050] FIG. 2B illustrates a front view of the first housing 260 according to an embodiment of the present disclosure. FIG. 2C illustrates a top view of the first housing 260 according to an embodiment of the present disclosure. As shown in FIGS. 2B and 2C, the first housing 260 may include a first connection port 262. In some embodiments, the first housing 260 may further include a contact point 266. In some embodiments, the first housing 260 may include an angle sensor (not shown in FIGS. 2B and 2C).

[0051] First connection port 262 may be an attachment connection port between first housing 260 and second housing 280. In some embodiments, first connection port 262 may be circular. First connection port 262 may be rotatably connected to second housing 280. When first housing 260 is attached to second housing 280, first housing 260 can rotate relative to second housing 280 to adjust the angle of first housing 260 relative to second housing 280, thereby adjusting the angle of microphone array 220.

[0052] A first magnetic device 263 may be installed in the first connection port 262. The first magnetic device 263 may be installed in a position adjacent to the second housing 280 of the first connection port 262. The first magnetic device 263 can generate a magnetic attraction force to realize a detachable connection with the second housing 280. When the first housing 260 approaches the second housing 280, the attraction force quickly connects the first housing 260 and the second housing 280. In some embodiments, after the first housing 260 and the second housing 280 are connected, the first housing 260 can further rotate relative to the second housing 280 to adjust the angle of the microphone array 220. When the first housing 260 rotates relative to the second housing 280 under the action of the attraction force, the connection between the first housing 260 and the second housing 280 can still be maintained.

[0053] In some embodiments, a first positioning device (not shown in FIGS. 2B and 2C ) may be further installed on the first connection port 262. The first positioning device may be a positioning step protruding outward, or a positioning hole extending inward. The first positioning device may be engaged with the second housing 280 to realize quick attachment of the first housing 260 and the second housing 280.

[0054] 2B and 2C, in some embodiments, the first housing 260 may further include contact points 266. The contact points 266 may be attached to the first connection port 262. The contact points 266 may protrude outward from the first connection port 262. The contact points 266 may be elastically connected to the first connection port 262. The contact points 266 may be communicatively connected to the M microphones 222 in the microphone array 220. The contact points 266 may be made of an elastic metal for transmitting data. When the first housing 260 and the second housing 280 are connected, the microphone array 220 may be communicatively connected to the computing device 240 through the contact points 266. In some embodiments, the contact points 266 may be distributed in a circular shape. After the first housing 260 and the second housing 280 are connected, if the first housing 260 is rotated relative to the second housing 280, the contact point 266 can rotate relative to the second housing 280 while maintaining a communication connection with the computing device 240.

[0055] In some embodiments, an angle sensor (not shown in FIGS. 2B and 2C ) may be further installed in the first housing 260. The angle sensor may be communicatively connected to the contact point 266 for communicatively connecting to the computing device 240. The angle sensor may collect angle data of the first housing 260 to determine the angle of the microphone array 220 and provide reference data for subsequent calculation of the voice presence probability.

[0056] 2D illustrates a front view of second housing 280 according to an embodiment of the present disclosure, and FIG. 2E illustrates a bottom view of second housing 280 according to an embodiment of the present disclosure. As shown in FIG. 2D and FIG. 2E, second housing 280 may include a second connection port 282. In some embodiments, second housing 280 may further include a guide rail 286.

[0057] Second connection port 282 may be an attachment connection port between second housing 280 and first housing 260. In some embodiments, second connection port 282 may be circular. Second connection port 282 may be rotatably connected to first connection port 262 of first housing 260. When first housing 260 is attached to second housing 280, first housing 260 can rotate relative to second housing 280 to adjust the angle of first housing 260 relative to second housing 280, thereby adjusting the angle of microphone array 220.

[0058] A second magnetic device 283 may be installed in the second connection port 282. The second magnetic device 283 may be installed in a position adjacent to the first housing 260 of the second connection port 282. The second magnetic device 283 can generate a magnetic attraction force to realize a detachable connection with the first connection port 262. The second magnetic device 283 may be used in conjunction with the first magnetic device 263. When the first housing 260 is adjacent to the second housing 280, the attraction force between the second magnetic device 283 and the first magnetic device 263 allows the first housing 260 to be quickly attached to the second housing 280. When the first housing 260 is attached to the second housing 280, the second magnetic device 283 and the first magnetic device 263 are positioned opposite each other. In some embodiments, after the first housing 260 and the second housing 280 are connected, the first housing 260 can further rotate relative to the second housing 280 to adjust the angle of the microphone array 220. When the first housing 260 rotates relative to the second housing 280 under the action of the above-mentioned suction force, the connection between the first housing 260 and the second housing 280 can still be maintained.

[0059] In some embodiments, a second positioning device (not shown in FIG. 2D and FIG. 2E) may be further installed on the second connection port 282. The second positioning device may be a positioning step protruding outward, or may be a positioning hole extending inward. The second positioning device may engage with the first positioning device of the first housing 260 to realize quick attachment between the first housing 260 and the second housing 280. When the first positioning device is the positioning step, the second positioning device may be the positioning hole. When the first positioning device is the positioning hole, the second positioning device may be the positioning step.

[0060] As shown in FIG. 2D and FIG. 2E, in some embodiments, the second housing 280 may further include a guide rail 286. The guide rail 286 may be attached to the second connection port 282. The guide rail 286 may be communicatively connected to the computing device 240. The guide rail 286 may be made of a metal material for data transmission. When the first housing 260 and the second housing 280 are connected, the contact point 266 contacts the guide rail 286 to form a communicative connection, thereby achieving a communicative connection between the microphone array 220 and the computing device 240 and transmitting data. As described above, the contact point 266 may be elastically connected to the first connection port 262. Therefore, after the first housing 260 and the second housing 280 are connected, the elastic force of the elastic connection can cause the contact point 266 to fully contact the guide rail 286 to achieve a reliable communicative connection. In some embodiments, the guide rail 286 may be distributed in a circular shape. After the first housing 260 and the second housing 280 are connected, when the first housing 260 rotates relative to the second housing 280, the contact point 266 can rotate relative to the guide rail 286 while maintaining a communication connection with the guide rail 286.

[0061] 3 shows a flowchart of a method P100 for calculating a voice presence probability according to an embodiment of the present specification. The method P100 can calculate the voice presence probability of the microphone signal. Specifically, the processor 242 can execute the method P100. As shown in FIG. 3, the method P100 can include the following steps S120 and S140.

[0062] In S120, microphone signals output from the M microphones 222 are acquired.

[0063] As described above, each microphone 222 can output a corresponding microphone signal. The M microphones 222 correspond to M microphone signals. When the method P100 calculates the voice presence probability, the method P100 may calculate based on all or some of the M microphone signals. Therefore, the microphone signals may include M microphone signals or some microphone signals corresponding to the M microphones 222. In the following description of this specification, a case where the microphone signals may include M microphone signals corresponding to the M microphones 222 will be described as an example.

[0064] As described above, the microphone 222 may collect noise in the surrounding environment or may collect the target voice of the target user. N For convenience of explanation, let us call the N signal sources s v Let us define it as (t). v (t) is N signal sources s1(t), ..., s N (t), where v=n or s+n. If v=n, then there are N signal sources s v (t) is all noise signals. If v=s+n, then there are N signal sources s v (t) consists of a noise signal and a target speech signal. v The sound field pattern of (t) is the far field pattern. N signal sources s v (t) can be regarded as a plane wave. For convenience of explanation, the microphone signal at time t is denoted as x(t). The microphone signal x(t) may be a signal vector consisting of M microphone signals. In this case, the microphone signal x(t) can be expressed by the following equation.

[0065]

number

[0066] In the formula, a v (θ) is the number of N signal sources s v (t) are the steering vectors θ1, …, θ N are N signal sources s1(t), ..., s N is the angle of incidence between (t) and microphone 222. v (θ) is θ1, …, θ N , and the distances d1, . . . , d M-1 The calculation device 240 stores in advance the relative positional relationships, for example, relative distances or relative coordinates, of the M microphones 222. That is, the calculation device 240 stores d1, ..., d M-1 is stored in advance.

[0067] The microphone signal x(t) is a time domain signal. In some embodiments, in step S120, the computing device 240 may further perform a spectrum analysis on the microphone signal. Specifically, the computing device 240 may perform a Fourier transform on the time domain signal x(t) of the microphone signal to obtain a frequency domain signal x of the microphone signal. f,t Hereinafter, the microphone signal x in the frequency domain may be obtained. f,t In this case, the microphone signal x f,t can be expressed by the following formula:

[0068]

number

[0069] During the ceremony,

number

number

number

number

[0070]

number

[0071] In some embodiments, the Gaussian distribution

number

number

number

number

number

number

number

number

[0072] As can be seen from equations (2) and (3), the microphone signal x f,t Also follows a Gaussian distribution. Specifically, the microphone signal x f,t x can be a Gaussian-distributed speech presence model or speech absence model. f,t can be expressed by the following formula:

[0073]

number

[0074] During the ceremony,

number

number

number

[0075] Above microphone signal x f,t The voice presence probability corresponding to the microphone signal x f,t For convenience of explanation, the microphone signal x f,t The corresponding voice presence probability is

number

number

number

number

number

[0076]

number

[0077]

number

number

number

[0078] For convenience of explanation, the first model is defined as the following equation.

[0079]

number

[0080] During the ceremony,

number

number

number

number

[0081] The second model is defined as follows:

[0082]

number

[0083] During the ceremony,

number

number

number

number

[0084]

number

[0085] In S140, the first model and the second model are iteratively optimized based on maximum likelihood estimation and EM algorithm respectively until convergence.

[0086] The computing device 240 iteratively optimizes the first model and the second model using an iterative optimization method to obtain a first variance of the first model.

number

number

number

number

number

number

[0087] First probability

number

number

number

number

number

number

number

number

number

[0088]

number

[0089] Above microphone signal x f,t The second probability corresponding to

number

[0090]

number

[0091] In S142, an objective function is constructed based on maximum likelihood estimation and the EM algorithm.

[0092] As mentioned before, the unknown parameters are the first variances of the first model.

number

number

number

number

number

number

[0093]

number

[0094] In S144, the optimization parameters are determined.

[0095] First parameter

number

number

[0096]

number

number

number

[0097]

number

[0098] Therefore, the optimization parameters are the first spatial covariance matrix

number

number

[0099] In S145, the initial values ​​of the optimization parameters are determined.

[0100] For convenience of explanation, the first spatial covariance matrix

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0101]

number

[0102] At S146, the optimization parameters are iterated multiple times based on the objective function and initial values ​​of the optimization parameters until the objective function converges.

[0103] As described above, the computing device 240 calculates the first probability in the multiple iterations.

number

number

number

number

[0104] In some embodiments, as shown in FIG. 5, the computing device 240 may, in any one of the plurality of iterations, determine the first probability

number

number

number

number

[0105] In S146-2, reversible correction is performed on the optimized parameters.

[0106] Specifically, step S146-2 may be to correct the optimization parameters by a deviation matrix if it is determined that the optimization parameters are irreversible. The deviation matrix may include one of an identity matrix and a random matrix following a normal distribution or a uniform distribution. As described above, the optimization parameters are calculated by a first spatial covariance matrix

number

number

number

number

number

number

number

number

number

number

[0107] Specifically, the computing device 240 calculates a first spatial covariance matrix

number

number

number

number

number

number

[0108] The first spatial covariance matrix

number

number

number

number

number

number

[0109]

number

[0110]

number

[0111] where Q is the deviation matrix, μ is the deviation coefficient, and in some embodiments μ=0.001.

[0112] The first spatial covariance matrix

number

number

[0113] In S146-3, a first parameter is calculated based on the formula (11) and the formula (12).

number

number

[0114] In S146-4, the first probability is calculated based on the formula (8) and the formula (9).

number

number

[0115] In S146-5, the first probability

number

number

number

number

[0116] The first spatial covariance matrix

number

number

[0117]

number

[0118]

number

[0119] In S146-6, it is determined whether the iterations should stop based on the objective function.

[0120] Step S146-6 may include step S146-7 or step S146-8.

[0121] At S146-7 it is decided to stop the iterations and the converged values ​​of the optimization parameters are output.

[0122] At S146-8 it is determined that the iterations do not stop and the next iteration continues.

[0123] As shown in FIG. 5, step S146 may further include step S146-9.

[0124] In S146-9, in any one of the multiple repetitions,

number

number

number

number

[0125] Step S146-9 may be performed during the iteration process or after the iteration is completed, and the first probability in any one of the multiple iterations is

number

number

number

number

number

number

[0126] Specifically, in step S146-9, the computing device 240 calculates the first probability in any one of the iterations.

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0127] In some embodiments, as shown in FIG. 6, the computing device 240 may calculate the first probability in a first iteration of the plurality of iterations.

number

number

number

number

number

number

[0128] In S146-10, in the first iteration of the multiple iterations,

number

number

number

number

[0129] Specifically, in step S146-1, the computing device 240 calculates a first parameter based on equations (11) and (12) in a first iteration:

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0130] As shown in FIG. 6, step S146 may include performing steps S146-11 to S146-16 in each iteration after the first iteration.

[0131] In S146-11, reversible correction is performed on the optimized parameters, as described in step S146-2 above, and so a detailed description is omitted here.

[0132] In S146-12, a first parameter is calculated based on the formula (11) and the formula (12).

number

number

[0133] In S146-13, the first probability is calculated based on the formula (8) and the formula (9).

number

number

[0134] In S146-14, the first probability

number

number

number

number

number

number

[0135] Specifically, step S146-14 is performed by the computing device 240 calculating the first probability

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0136] In S146-15, the corrected first probability

number

number

number

number

[0137] In steps S146-14 and S146-15, the entropy of the voice presence model is made smaller than the entropy of the voice absence model during each iteration to ensure that each iteration converges towards the goal, thereby accelerating the convergence speed.

[0138] In S146-16, it is determined whether the iterations should stop based on the objective function.

[0139] Step S146-16 may include step S146-17 or step S146-18.

[0140] In S146-17 it is decided to stop the iterations and the converged values ​​of the optimization parameters are output.

[0141] In S146-18 it is determined that the iterations do not stop and the next iteration continues.

[0142] As shown in FIG. 4, step S140 may further include step S148.

[0143] In S148, a convergence value of the optimization parameter and a first probability corresponding to the convergence value are calculated.

number

number

[0144] As described above, when the objective function converges, the calculation device 240 can output the value of the optimization parameter corresponding to when the objective function converges as the convergence value of the optimization parameter. At the same time, the calculation device 240 can output a first probability corresponding to the convergence value of the optimization parameter. rate

number

number

number

number

number

number

number

number

number

number

[0145] As shown in FIG. 3, the method P100 may further include step S160.

[0146] In S160, if the maximum likelihood estimation and the EM algorithm converge, the microphone signal x f,tis the voice presence model as the voice presence probability of the microphone signal.

[0147] As mentioned above, in step S140, the computing device 240 calculates the first probability

number

number

number

number

number

number

number

number

number

number

number

number

[0148] The computing device 240 calculates the voice presence probability

number

[0149] In view of the above, in the system and method P100 for calculating the presence of speech probability according to the present specification, the calculation device 240 calculates a first probability corresponding to the first model.

number

number

number

number

number

number

number

number

number

number

number

number

number

[0150] The present specification further provides a voice enhancement system. The voice enhancement system may be applied to the electronic device 200. In some embodiments, the voice enhancement system may include a computing device 240. In some embodiments, the voice enhancement system may be applied to the computing device 240. That is, the voice enhancement system may be executed in the computing device 240. The voice enhancement system may include a hardware device having a function of processing data information and a program required to drive the operation of the hardware device. Of course, the voice enhancement system may be only a hardware device having data processing capabilities, or only a program executed on the hardware device.

[0151] The voice enhancement system may store data or instructions and execute the data and / or instructions for performing the voice enhancement methods described herein. When the voice enhancement system is executed on the computing device 240, the voice enhancement system may obtain the microphone signals from the microphone array 220 based on the communication connection and execute data or instructions for the voice enhancement methods described herein. The voice enhancement methods are introduced elsewhere in this specification. For example, the description of FIG. 7 introduces the voice enhancement methods.

[0152] When the voice enhancement system is executed in the computing device 240, the voice enhancement system is communicatively connected to the microphone array 220. The storage medium 243 may further include at least one instruction set stored in the data storage device for performing MVDR-based voice enhancement calculations on the microphone signals. The instructions are computer program code, and the computer program code may include programs, routines, objects, assemblies, data structures, processes, modules, etc., for performing the voice enhancement method according to the present disclosure. The processor 242 may read the at least one instruction set and perform the voice enhancement method according to the present disclosure based on the instructions of the at least one instruction set. The processor 242 may perform all steps included in the voice enhancement method.

[0153] FIG. 7 shows a flowchart of a speech enhancement method P200 according to an embodiment of the present specification. The method P200 may perform speech enhancement on the microphone signal based on an MVDR method. Specifically, the processor 242 may execute the method P200. As shown in FIG. 7, the method P200 may include steps S220 to S290.

[0154] In S220, the microphone signals x output from the M microphones are f,t Get the.

[0155] This is as described in step S120, and the description will be omitted here.

[0156] In S240, based on the above-mentioned voice presence probability calculation method P100, the microphone signal x f,t Probability of voice presence

number

[0157] In S260, the probability of the presence of voice

number

number

[0158] Noise covariance matrix

number

[0159]

number

[0160] In S280, the MVDR method and the noise spatial covariance matrix

number

[0161] Filter coefficient ω f,t can be expressed by the following formula:

[0162]

number

[0163] During the ceremony,

number

number

number

number

[0164] In some embodiments, the filter coefficient ω f,t can be expressed by the following formula:

[0165]

number

[0166] During the ceremony,

number

number

number

number

number

[0167] In S290, the filter coefficient ω f,t Based on the microphone signal xf,t and combine them to produce the target audio signal y f,t Output.

[0168] Target audio signal y f,t can be expressed by the following formula:

[0169]

number

[0170] The computing device 240 calculates the target audio signal y f,t may be output to other electronic devices such as a remote calling device.

[0171] As described above, the system and method P100 for calculating the presence of voice probability, the system and method P200 for speech enhancement, and the electronic device 200 for calculating the presence of voice probability according to the present specification are used for a microphone array 220 consisting of a plurality of microphones 222. The system and method P100 for calculating the presence of voice probability, the system and method P200 for speech enhancement, and the electronic device 200 for calculating the presence of voice probability according to the present specification are used for a microphone array 220 consisting of a plurality of microphones 222. The system and method P100 for calculating the presence of voice probability, the system and method P200 for speech enhancement, and the electronic device 200 for calculating the presence of voice probability according to the present specification are used for a microphone array 220 consisting of a plurality of microphones 222. The system and method P100 for calculating the presence of voice probability, the system and method P200 for speech enhancement, and the electronic device 200 for calculating the presence of voice probability according to the present specification are used for a plurality of microphone signals. The above-mentioned system and method P100 for calculating the presence of voice probability, the system and method P200 for speech enhancement and the electronic device 200 for computing the presence of voice probability and the absence of voice probability are adjusted in the iterative process by comparing the entropy of the presence of voice probability and the entropy of the absence of voice probability, thereby obtaining a faster convergence speed and a better convergence result, improving the estimation accuracy of the presence of voice probability and the noise covariance matrix, and further improving the speech enhancement effect of MVDR.

[0172] Another aspect of the present specification provides a non-transitory storage medium storing at least one set of executable instructions for calculating a voice presence probability, which, when executed by a processor, instructs the processor to perform steps of a voice presence probability calculation method P100 described herein. In some possible embodiments, each aspect of the present specification may be realized in the form of a program product including a program code. When the program product is executed in a computing device (e.g., computing device 240), the program code is used to cause the computing device to perform the voice presence probability calculation method described herein. The program product implementing the method may be executed in a computing device using a portable compact disc read-only memory (CD-ROM) including the program code. However, the program product in the present specification is not limited thereto, and in the present specification, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system (e.g., processor 242). The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include an electrical connection having one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium may include a data signal propagated on baseband or as part of a carrier wave, and carrying readable program code. Such a propagated data signal may take various forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above.The readable storage medium may be any readable medium other than a readable storage medium, which may transmit, propagate, or transmit a program used by or in conjunction with an instruction execution system, apparatus, or device. The program code contained in the readable storage medium may be transmitted in any suitable medium, including, but not limited to, wireless, wired, optical cable, RF, or the like, or any suitable combination of the above. The program code for performing the operations herein may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and further including conventional process programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the computing device, partially on the computing device, or as a separate software package, partly on the computing device and partly on a remote computing device, or entirely on a remote computing device.

[0173] Certain embodiments of the present specification have been described above. Other embodiments are within the scope of the following claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the examples and still achieve desirable results. Also, the processes depicted in the figures do not necessarily require a particular order or sequential order to achieve desirable results. In some embodiments, multitasking and parallel processing may also be possible or advantageous.

[0174] From the above, after reading this detailed disclosure, it will be apparent to those skilled in the art that the above detailed disclosure is presented by way of example only and is not limiting. Although not explicitly described herein, those skilled in the art can understand that this specification is intended to include various reasonable changes, improvements and modifications to the embodiments. These changes, improvements and modifications are intended to be suggested by this specification and are within the spirit and scope of the exemplary embodiments of this specification.

[0175] In addition, some terms in this specification are used to describe embodiments of this specification. For example, "one embodiment," "embodiment," and / or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of this specification. Therefore, it is emphasized and understood that two or more references to "embodiment" or "one embodiment" or "alternative embodiments" in each part of this specification do not necessarily all refer to the same embodiment. Also, a particular feature, structure, or characteristic may be appropriately combined in one or more embodiments of this specification.

[0176] In the above description of the embodiments of the present specification, it should be understood that in order to simplify the specification, in order to facilitate understanding of one feature, the specification may group various features together in a single embodiment, drawing, or description thereof. However, this does not mean that the combination of these features is essential, and it is entirely possible for a person skilled in the art to extract some of the features and understand them as separate embodiments when reading the specification. That is, the embodiments in the present specification may be understood as a combination of multiple sub-embodiments. Also, the content of each sub-embodiment may be valid even when it has less than all the features of a single embodiment disclosed above.

[0177] All patents, patent applications, published patent applications, and other materials, such as documents, books, specifications, publications, documents, articles, etc., referenced herein are incorporated herein by reference in their entirety for all purposes, excluding any associated docket history, and any that may be inconsistent or contradictory with this specification or that may have a limiting effect on the broadest scope of the claims, whether now or later relating to this specification. For example, in the event of any inconsistency or discrepancy between the description, definition, and / or use of a term relating to any material incorporated herein and the description, definition, and / or use of a term relating to this specification, the term in this specification shall control.

[0178] Finally, it should be understood that the embodiments of the present application disclosed herein are illustrative of the principles of the embodiments of the present application. Other modified embodiments are also within the scope of the present application. Therefore, the embodiments disclosed herein are merely illustrative and not limiting. Those skilled in the art can implement the invention herein using alternative configurations based on the embodiments of the present application. Therefore, the embodiments of the present application are not limited to the embodiments described in detail in the application. [Explanation of symbols]

[0179] 200 Electronic Devices 220 Microphone Array 222 Microphone 240 Computing equipment 241 Internal Communication Bus 242 processors 243 Storage medium 245 communication port 260 First Housing 262 First Connection Port 263 First Magnetic Device 266 contact points 280 Second Housing 282 Second Connection Port 283 Second Magnetic Device 286 Guide Rail

Claims

1. for M microphones, where M is an integer greater than 1, distributed in a given array shape; obtaining microphone signals output from the M microphones, the microphone signals following a first model or a second model of a Gaussian distribution, one of the first model and the second model being a voice presence model and the other being an absence of voice model; iteratively optimizing the first model and the second model based on maximum likelihood estimation and EM algorithm respectively until convergence, and in the iterative process, determining whether the voice presence model is the first model or the second model based on an entropy of a first probability when the microphone signal is the first model and an entropy of a second probability when the microphone signal is the second model, where the first probability and the second probability are complementary; and if the maximum likelihood estimation and the EM algorithm converge, outputting the probability that the microphone signal is the voice presence model as the voice presence probability of the microphone signal.

2. a first variance of a Gaussian distribution corresponding to the first model includes a product of a first parameter and a first spatial covariance matrix; 2. The method of claim 1, wherein a second variance of a Gaussian distribution corresponding to the second model comprises a product of a second parameter and a second spatial covariance matrix.

3. Iteratively optimizing the first model and the second model based on maximum likelihood estimation and an EM algorithm, respectively, constructing an objective function based on maximum likelihood estimation and the EM algorithm; determining optimization parameters including the first spatial covariance matrix and the second spatial covariance matrix; determining initial values ​​of the optimization parameters; iterating the optimization parameters a number of times based on the objective function and initial values ​​of the optimization parameters until the objective function converges, and during the number of iterations determining whether the voice presence model is the first model or the second model based on an entropy of the first probability and an entropy of the second probability; outputting a convergence value of the optimization parameter and the first probability and the second probability corresponding to the convergence value; 3. The method of claim 2, further comprising:

4. determining whether the voice presence model is the first model or the second model based on an entropy of the first probability and an entropy of the second probability in the multiple iterations, calculating an entropy of the first probability and an entropy of the second probability in any one of the plurality of iterations and determining whether the voice presence model is the first model or the second model, determining that the voice presence model is the second model upon determining that the entropy of the first probability is greater than the entropy of the second probability; or 4. The method of claim 3, further comprising the step of determining that the voice presence model is the first model upon determining that the entropy of the first probability is less than the entropy of the second probability.

5. determining whether the voice presence model is the first model or the second model based on an entropy of the first probability and an entropy of the second probability in the multiple iterations, calculating an entropy of the first probability and an entropy of the second probability in a first iteration of the plurality of iterations and determining whether the voice presence model is the first model or the second model, determining that the voice presence model is the second model upon determining that the entropy of the first probability is greater than the entropy of the second probability; or 4. The method of claim 3, further comprising the step of determining that the voice presence model is the first model upon determining that the entropy of the first probability is less than the entropy of the second probability.

6. The step of iterating the optimization parameters a plurality of times includes, in each of the plurality of iterations, correcting the first probability and the second probability based on an entropy of the first probability and an entropy of the second probability, upon determining that the first model is the voice presence model and that the entropy of the first probability is greater than the entropy of the second probability, replacing the value corresponding to the first probability with the value corresponding to the second probability; or upon determining that the second model is the voice presence model and that the entropy of the second probability is greater than the entropy of the first probability, replacing the value corresponding to the first probability with the value corresponding to the second probability; updating the optimization parameters based on the corrected first probability and the corrected second probability; 6. The method of claim 5, further comprising:

7. The step of iterating the optimization parameters a plurality of times includes, in each of the plurality of iterations, 7. The method for calculating a voice presence probability according to claim 3, further comprising: a step of performing a reversible correction on the optimization parameter, and when it is determined that the optimization parameter is irreversible, the step of correcting the optimization parameter based on a deviation matrix including one of an identity matrix and a random matrix following a normal distribution or a uniform distribution.

8. 1. A system for computing a sound presence probability, comprising: at least one storage medium having stored thereon at least one set of instructions for the calculation of a voice presence probability; at least one processor communicatively coupled to the at least one storage medium; Including, When the voice presence probability calculation system is executed, the at least one processor reads the at least one instruction set and executes a voice presence probability calculation method, the voice presence probability calculation method comprising: obtaining microphone signals output from M microphones, the microphone signals following a first model or a second model of a Gaussian distribution, one of the first model and the second model being a voice presence model and the other being an absence of voice model; iteratively optimizing the first model and the second model based on maximum likelihood estimation and EM algorithm respectively until convergence, and in the iterative process, determining whether the voice presence model is the first model or the second model based on an entropy of a first probability when the microphone signal is the first model and an entropy of a second probability when the microphone signal is the second model, where the first probability and the second probability are complementary; if the maximum likelihood estimation and the EM algorithm converge, outputting the probability that the microphone signal is the voice presence model as a voice presence probability of the microphone signal; A system for calculating a voice presence probability, comprising:

9. for M microphones, where M is an integer greater than 1, distributed in a given array shape; acquiring microphone signals output from the M microphones; determining a voice presence probability of said microphone signal; determining a noise covariance matrix of the microphone signals based on the voice presence probability; determining filter coefficients corresponding to the microphone signals based on an MVDR method and the noise covariance matrix; combining the microphone signals based on the filter coefficients to output a target audio signal; Including, The microphone signal is according to a first model or a second model of a Gaussian distribution, one of the first model and the second model being a voice presence model and the other being an absence of voice model, and the step of determining the voice presence probability of the microphone signal comprises: iteratively optimizing the first model and the second model based on maximum likelihood estimation and EM algorithm respectively until convergence, and in the iterative process, determining whether the voice presence model is the first model or the second model based on an entropy of a first probability when the microphone signal is the first model and an entropy of a second probability when the microphone signal is the second model, where the first probability and the second probability are complementary; if the maximum likelihood estimation and the EM algorithm converge, outputting the probability that the microphone signal is the voice presence model as a voice presence probability of the microphone signal; A method for enhancing speech, comprising:

10. 1. A speech enhancement system, comprising: at least one storage medium having stored thereon at least one instruction set for speech enhancement; at least one processor communicatively coupled to the at least one storage medium; Including, When the speech enhancement system is executed, the at least one processor reads the at least one instruction set and executes a speech enhancement method, the speech enhancement method comprising: acquiring microphone signals output from M microphones; determining a voice presence probability of said microphone signal; determining a noise covariance matrix of the microphone signals based on the voice presence probability; determining filter coefficients corresponding to the microphone signals based on an MVDR method and the noise covariance matrix; combining the microphone signals based on the filter coefficients to output a target audio signal; Including, The microphone signal is according to a first model or a second model of a Gaussian distribution, one of the first model and the second model being a voice presence model and the other being an absence of voice model, and the step of determining the voice presence probability of the microphone signal comprises: iteratively optimizing the first model and the second model based on maximum likelihood estimation and EM algorithm respectively until convergence, and in the iterative process, determining whether the voice presence model is the first model or the second model based on an entropy of a first probability when the microphone signal is the first model and an entropy of a second probability when the microphone signal is the second model, where the first probability and the second probability are complementary; if the maximum likelihood estimation and the EM algorithm converge, outputting the probability that the microphone signal is the voice presence model as a voice presence probability of the microphone signal; A speech enhancement system comprising:

11. A microphone array including M microphones, M being an integer greater than 1, distributed in a predetermined array shape; a computing device, in operation, communicatively connected to the microphone array and configured to perform a speech enhancement method; Including, The speech enhancement method includes: acquiring microphone signals output from the M microphones; determining a voice presence probability of said microphone signal; determining a noise covariance matrix of the microphone signals based on the voice presence probability; determining filter coefficients corresponding to the microphone signals based on an MVDR method and the noise covariance matrix; combining the microphone signals based on the filter coefficients to output a target audio signal; Including, a first housing in which the microphone array is mounted; a second housing in which the computing device is mounted; and Further comprising: the first housing includes a first connection port in which a first magnetic device is installed; the second housing includes a second connection port in which a second magnetic device is installed; an adhesive force between the first magnetic device and the second magnetic device removably connects the first housing to the second housing; the first housing further includes a contact located at the first port and communicatively coupled to the microphone array; the second housing further includes a guide rail disposed in the second connection port and communicatively connected to the computing device; when the first housing is connected to the second housing, the contact points contact the guide rails, thereby communicatively connecting the microphone array to the computing device.

12. The earphone of claim 11, wherein the M microphones are linearly distributed, M is less than or equal to 5, and a distance between adjacent microphones among the M microphones is between 20 mm and 40 mm.

Citation Information

Patent Citations

  • Earphone with microphone

    JP1999164382A

  • Noise canceling system and audio apparatus

    JP2009141491A

  • Sound source separation device and method, and program

    JP2013054258A