A speech enhancement method and system
Patent Information
- Application Number
- CN202180088314.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-27
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2041-05-27
AI Technical Summary
在进行语音通话和语音信号采集等场景中,会存在环境噪声、他人语音等各种噪声信号干扰,导致采集的目标语音不是干净的语音信号,影响了语音信号的质量,导致听不清语音、通话质量不高等问题
Smart Images

Figure CN116724352B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to methods and systems for speech enhancement. Background Technology
[0002] With the development of speech processing technology, the requirements for the quality of speech signals are becoming increasingly stringent in fields such as communication and speech acquisition. In scenarios involving voice calls and speech signal acquisition, various noise signals, such as environmental noise and other people's voices, can interfere, resulting in a non-clean speech signal and affecting the quality of the acquired speech, leading to problems such as unclear speech and poor call quality.
[0003] Therefore, there is an urgent need to provide a speech enhancement method and system. Summary of the Invention
[0004] One embodiment of this specification provides a speech enhancement method. The method may include acquiring a first signal and a second signal of target speech, wherein the first signal is a signal of the target speech acquired based on a first location, and the second signal is a signal of the target speech acquired based on a second location. The method may further include processing the first signal and the second signal based on the target speech location, the first location, and the second location to determine a first coefficient; determining a plurality of parameters related to a plurality of sound source directions based on the first signal and the second signal, each parameter corresponding to the probability of emitting sound from a sound source direction to form the first signal and the second signal. The method may further include determining a second coefficient based on the plurality of parameters and the target speech location; and processing the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain a first output speech signal after speech enhancement corresponding to the target speech.
[0005] In some embodiments, processing the first signal and the second signal based on the target speech location, the first location, and the second location to determine the first coefficient may include: performing a differential operation on the first signal and the second signal based on the target speech location, the first location, and the second location to obtain a signal pointing in a first direction and a signal pointing in a second direction, wherein the signal pointing in the first direction and the signal pointing in the second direction contain different proportions of effective signal; determining a third signal corresponding to the effective signal based on the signal pointing in the first direction and the signal pointing in the second direction; and determining the first coefficient based on the third signal.
[0006] In some embodiments, determining the third signal corresponding to the valid signal may include: performing an adaptive differential operation on the signal pointing in the first direction and the signal pointing in the second direction to determine a fourth signal; and enhancing the low-frequency components in the fourth signal to obtain the third signal.
[0007] In some embodiments, the method may further include: updating the adaptive parameters of the adaptive differential operation based on the fourth signal, the signal pointing in the first direction, and the signal pointing in the second direction.
[0008] In some embodiments, processing the first signal and the second signal based on the target speech location, the first location, and the second location to determine the first coefficient may include: performing a difference operation on the first signal and the second signal based on the target speech location, the first location, and the second location to obtain a signal pointing in a first direction and a signal pointing in a second direction, wherein the signal pointing in the first direction and the signal pointing in the second direction contain different proportions of effective signal; determining the estimated signal-to-noise ratio of the target speech based on the signal pointing in the first direction and the signal pointing in the second direction; and determining the first coefficient based on the estimated signal-to-noise ratio.
[0009] In some embodiments, determining multiple parameters related to multiple sound source directions based on the first signal and the second signal may include: performing a differential operation on the first signal and the second signal based on each sound source direction, the first position, and the second position to determine parameters related to each sound source direction.
[0010] In some embodiments, determining the second coefficient based on the plurality of parameters and the target speech location may include: determining the direction of the synthesized sound source based on the plurality of parameters; and determining the second coefficient based on the direction of the synthesized sound source and the target speech location.
[0011] In some embodiments, determining the second coefficient based on the direction of the synthesized sound source and the location of the target speech may include: determining whether the location of the target speech is located in the direction of the synthesized sound source; setting the second coefficient to a first value in response to the location of the target speech being located in the direction of the synthesized sound source; and setting the second coefficient to a second value in response to the location of the target speech not being located in the direction of the synthesized sound source.
[0012] In some embodiments, determining the second coefficient based on the direction of the synthesized sound source and the location of the target speech may include: determining the second coefficient by means of a regression function based on the angle between the location of the target speech and the direction of the synthesized sound source.
[0013] In some embodiments, the method may further include smoothing the second coefficient based on a smoothing factor.
[0014] In some embodiments, the method may further include performing at least one of the following operations on the first signal and the second signal: framing the first signal and the second signal; windowing smoothing the first signal and the second signal; and converting the first signal and the second signal to the frequency domain.
[0015] In some embodiments, the method may further include: determining at least one target subband signal in the first output speech signal; and processing the at least one target subband signal based on a single-microphone filtering algorithm to obtain a second output speech signal.
[0016] In some embodiments, determining at least one target sub-band signal in the first output speech signal may include: acquiring a plurality of sub-band signals based on the first output speech signal; calculating the signal-to-noise ratio (SNR) of each of the sub-band signals; and determining the target sub-band signal based on the SNR of each of the sub-band signals.
[0017] In some embodiments, the method may further include: processing the first signal and / or the second signal based on a single-microphone filtering algorithm to determine a third coefficient; and processing the first output speech signal based on the third coefficient to obtain a third output speech signal.
[0018] In some embodiments, the method may further include: determining a fourth coefficient based on the energy difference between the first signal and the second signal; and processing the first signal and / or the second signal based on the first coefficient, the second coefficient, and the fourth coefficient to obtain a fourth output speech signal with speech enhancement corresponding to the target speech.
[0019] In some embodiments, determining the fourth coefficient based on the energy difference between the first signal and the second signal may include: obtaining a noise power spectral density based on a silent region in the first signal and the second signal; obtaining the energy difference based on a first power spectral density of the first signal, a second power spectral density of the second signal, and the noise power spectral density; and determining the fourth coefficient based on the energy difference and the noise power spectral density.
[0020] One embodiment of this specification provides a speech enhancement system. The system may include: at least one storage medium including a set of instructions; and at least one processor communicating with the at least one storage medium, wherein, when executing the set of instructions, the at least one processor can cause the system to: acquire a first signal and a second signal of target speech, the first signal being a signal of the target speech acquired based on a first location, and the second signal being a signal of the target speech acquired based on a second location; process the first signal and the second signal based on the target speech location, the first location, and the second location to determine a first coefficient; determine a plurality of parameters related to a plurality of sound source directions based on the first signal and the second signal, each parameter corresponding to the probability of emitting sound from a sound source direction to form the first signal and the second signal; determine a second coefficient based on the plurality of parameters and the target speech location; and process the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain a first output speech signal with speech enhancement corresponding to the target speech.
[0021] This specification provides a speech enhancement system according to one embodiment. The system may include an acquisition module, a processing module, and a generation module. The acquisition module may acquire a first signal and a second signal of target speech, wherein the first signal is a signal of the target speech acquired based on a first location, and the second signal is a signal of the target speech acquired based on a second location. The processing module may process the first signal and the second signal based on the target speech location, the first location, and the second location to determine a first coefficient; determine multiple parameters related to multiple sound source directions based on the first signal and the second signal, each parameter corresponding to the probability of emitting sound from a sound source direction to form the first signal and the second signal; and determine a second coefficient based on the multiple parameters and the target speech location. The generation module may process the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain a first output speech signal after speech enhancement corresponding to the target speech.
[0022] One embodiment of this specification provides a non-transitory computer-readable medium that may include executable instructions that, when executed by at least one processor, cause the at least one processor to perform the methods described in this specification.
[0023] Additional features will be set forth in part in the description which follows, and will become apparent to those skilled in the art upon consulting the following description and the accompanying drawings, or may be learned by the generation or operation of examples. The features of the invention can be realized and obtained by practice or use of various aspects of the methods, tools, and combinations set forth in the following detailed examples. Attached Figure Description
[0024] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:
[0025] Figure 1 These are schematic diagrams illustrating application scenarios of the speech enhancement system according to some embodiments of this specification;
[0026] Figure 2 These are schematic diagrams of exemplary hardware and / or software components of an exemplary computing device according to some embodiments of this specification;
[0027] Figure 3 These are schematic diagrams of exemplary hardware and / or software components of an exemplary mobile device according to some embodiments of this specification;
[0028] Figure 4 These are exemplary block diagrams of a speech enhancement system according to some embodiments of this specification;
[0029] Figure 5 This is an exemplary flowchart of a speech enhancement method according to some embodiments of this specification;
[0030] Figure 6 This is a schematic diagram of an exemplary dual microphone according to some embodiments of this specification;
[0031] Figure 7 These are schematic diagrams illustrating the filtering effect of the ANF algorithm at different noise angles according to some embodiments of this specification;
[0032] Figure 8 This is an exemplary flowchart of a method for determining a first coefficient according to some embodiments of this specification;
[0033] Figure 9 This is an exemplary flowchart of a method for determining a first coefficient according to some embodiments of this specification;
[0034] Figure 10 This is an exemplary flowchart of a method for determining a second coefficient according to some embodiments of this specification;
[0035] Figure 11 This is an exemplary flowchart of a single-microphone filtering method according to some embodiments of this specification;
[0036] Figure 12 This is an exemplary flowchart of a single-microphone filtering method according to some embodiments of this specification;
[0037] Figure 13This is an exemplary flowchart of a speech enhancement method according to some embodiments of this specification. Detailed Implementation
[0038] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. It should be understood that these exemplary embodiments are given merely to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0039] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements. The term "based on" means "at least partially based on." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment."
[0040] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0041] Figure 1 This is a schematic diagram illustrating application scenarios of a voice enhancement system according to some embodiments of this specification. The voice enhancement system 100 shown in the embodiments of this specification can be applied to various software, systems, platforms, and devices to achieve voice signal enhancement processing. For example, it can be applied to perform voice enhancement processing on user voice signals acquired by various software, systems, platforms, and devices, and it can also be applied to perform voice enhancement processing when making voice calls using devices (such as mobile phones, tablets, computers, headphones, etc.).
[0042] In voice call scenarios, various noise signals, such as environmental noise and other people's voices, can interfere, resulting in a non-clean voice signal being acquired. To improve the quality of voice calls, noise filtering and voice signal enhancement processing are needed to obtain a clean voice signal. This specification proposes a voice enhancement system and method that can enhance target voice signals in scenarios such as the aforementioned voice call.
[0043] like Figure 1 As shown, the voice enhancement system 100 may include a processing device 110, a acquisition device 120, a terminal 130, a storage device 140, and a network 150.
[0044] In some embodiments, the processing device 110 can process data and / or information obtained from other devices or system components. The processing device 110 can execute program instructions based on this data, information, and / or processing results to perform one or more functions described in this specification. For example, the processing device 110 can acquire and process a first signal and a second signal of target speech, and output an enhanced speech signal.
[0045] In some embodiments, processing device 110 may be a single processing device or a group of processing devices, such as a server or a group of servers. The group of processing devices may be centralized or distributed (e.g., processing device 110 may be a distributed system). In some embodiments, processing device 110 may be local or remote. For example, processing device 110 may access information and / or data in acquisition device 120, terminal 130, and storage device 140 via network 150. As another example, processing device 110 may be directly connected to acquisition device 120, terminal 130, and storage device 140 to access stored information and / or data. In some embodiments, processing device 110 may be implemented on a cloud platform. By way of example only, the cloud platform may include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, inter-cloud, multi-cloud, etc., or any combination thereof. In some embodiments, processing device 110 may be implemented in conjunction with this specification. Figure 2 This is implemented on the computing device shown. For example, the processing device 110 can be implemented on a computing device such as... Figure 2 Implemented on one or more components of the computing device 200 shown.
[0046] In some embodiments, the processing device 110 may include a processing engine 112. The processing engine 112 may process data and / or information related to speech enhancement to perform one or more methods or functions described herein. For example, the processing engine 112 may acquire a first signal and a second signal of target speech, wherein the first signal is a signal of the target speech acquired based on a first location, and the second signal is a signal of the target speech acquired based on a second location. In some embodiments, the processing engine 112 may process the first signal and / or the second signal to obtain a speech-enhanced output speech signal corresponding to the target speech.
[0047] In some embodiments, the processing engine 112 may include one or more processing engines (e.g., a single-chip processing engine or a multi-chip processor). By way of example only, the processing engine 112 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a graphics processing unit (GPU), a physical processing unit (PPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic device (PLD), a controller, a microcontroller unit, a reduced instruction set computer (RISC), a microprocessor, or any combination thereof. In some embodiments, the processing engine 112 may be integrated into the acquisition device 120 or the terminal 130.
[0048] In some embodiments, the acquisition device 120 can be used to acquire the speech signal of the target speech, such as acquiring a first signal and a second signal of the target speech. In some embodiments, the acquisition device 120 can be a single acquisition device or a group of acquisition devices consisting of multiple acquisition devices. In some embodiments, the acquisition device 120 can be a device (e.g., mobile phone, headset, walkie-talkie, tablet, computer, etc.) containing one or more microphones or other sound sensors (e.g., 120-1, 120-2, ..., 120-n). For example, the acquisition device 120 can include at least two microphones, which are spaced a certain distance apart. When the acquisition device 120 acquires user speech, the at least two microphones can simultaneously acquire sound from the user's mouth at different locations. The at least two microphones can include a first microphone and a second microphone. The first microphone can be located closer to the user's mouth, and the second microphone can be located farther from the user's mouth. The line connecting the second microphone and the first microphone can extend towards the location of the user's mouth.
[0049] The acquisition device 120 can convert the acquired speech into electrical signals and send them to the processing device 110 for processing. For example, the first and second microphones described above can convert the acquired user speech into a first signal and a second signal, respectively. The processing device 110 can perform speech enhancement processing based on the first and second signals.
[0050] In some embodiments, the acquisition device 120 can transmit information and / or data with the processing device 110, the terminal 130, and the storage device 140 via the network 150. In some embodiments, the acquisition device 120 can be directly connected to the processing device 110 or the storage device 140 to transmit information and / or data. For example, the acquisition device 120 and the processing device 110 can be different parts of the same electronic device (e.g., headphones, glasses, etc.) and connected by a metal wire.
[0051] In some embodiments, terminal 130 may be a terminal used by a user or other entity. For example, it may be a terminal used by the sound source (person or other entity) corresponding to the target speech, or it may be a terminal used by other users or entities conducting voice calls with the sound source (person or other entity) corresponding to the target speech.
[0052] In some embodiments, terminal 130 may include mobile device 130-1, tablet computer 130-2, laptop computer 130-3, etc., or any combination thereof. In some embodiments, mobile device 130-1 may include smart home devices, wearable devices, smart mobile devices, virtual reality devices, augmented reality devices, etc., or any combination thereof. In some embodiments, smart home devices may include smart lighting devices, smart appliance control devices, smart monitoring devices, smart TVs, smart cameras, walkie-talkies, etc., or any combination thereof. In some embodiments, wearable devices may include smart bracelets, smart shoes and socks, smart glasses, smart helmets, smartwatches, smart headphones, smart wearables, smart backpacks, smart accessories, etc., or any combination thereof. In some embodiments, smart mobile devices may include smartphones, personal digital assistants (PDAs), gaming devices, navigation devices, point-of-sale (POS) devices, etc., or any combination thereof. In some embodiments, virtual reality devices and / or augmented reality devices may include virtual reality helmets, virtual reality glasses, virtual reality goggles, augmented reality helmets, augmented reality glasses, augmented reality goggles, etc., or any combination thereof.
[0053] In some embodiments, terminal 130 can acquire / receive the voice signal of the target speech, such as a first signal and a second signal. In some embodiments, terminal 130 can acquire / receive the voice-enhanced output voice signal of the target speech. In some embodiments, terminal 130 can directly acquire / receive the voice signal of the target speech, such as the first signal and the second signal, from acquisition device 120 and storage device 140, or terminal 130 can acquire / receive the voice signal of the target speech, such as the first signal and the second signal, from acquisition device 120 and storage device 140 via network 150. In some embodiments, terminal 130 can directly acquire / receive the voice-enhanced output voice signal of the target speech from processing device 110 and storage device 140, or terminal 130 can acquire / receive the voice-enhanced output voice signal of the target speech from processing device 110 and storage device 140 via network 150.
[0054] In some embodiments, terminal 130 may send instructions to processing device 110, and processing device 110 may execute instructions from terminal 130. For example, terminal 130 may send one or more instructions to processing device 110 to implement a speech enhancement method for the target speech, so that processing device 110 performs one or more operations / steps of the speech enhancement method.
[0055] Storage device 140 may store data and / or information obtained from other devices or system components. For example, storage device 140 may store speech signals of a target speech, such as a first signal and a second signal, and may also store an enhanced output speech signal of the target speech. In some embodiments, storage device 140 may store data acquired from acquisition device 120. In some embodiments, storage device 140 may store data acquired from processing device 110. In some embodiments, storage device 140 may store data and / or instructions used by processing device 110 to perform or use in order to complete the exemplary methods described herein. In some embodiments, storage device 140 may include mass storage, removable storage, volatile read-write storage, read-only storage (ROM), etc., or any combination thereof. Exemplary mass storage may include a hard disk, optical disk, solid-state drive, etc. Exemplary removable storage may include a flash drive, floppy disk, optical disk, memory card, compact disk, magnetic tape, etc. Exemplary volatile read-only storage may include random access memory (RAM). Exemplary RAMs may include Dynamic RAM (DRAM), Double Data Rate Synchronous Dynamic RAM (DDR SDRAM), Static RAM (SRAM), Thyristor RAM (T-RAM), and Zero Capacitor RAM (Z-RAM), etc. Exemplary ROMs may include Mask ROM (MROM), Programmable ROM (PROM), Erasable Programmable ROM (PEROM), Electronically Erasable Programmable ROM (EEPROM), Optical Disc ROM (CD-ROM), and Digital Universal Disk ROM, etc. In some embodiments, the storage device 140 may be implemented on a cloud platform. By way of example only, the cloud platform may include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, internal cloud, multi-layer cloud, etc., or any combination thereof.
[0056] In some embodiments, storage device 140 may be connected to network 150 to communicate with one or more components of voice enhancement system 100 (e.g., processing device 110, acquisition device 120, terminal 130). One or more components of voice enhancement system 100 may access data or instructions stored in storage device 140 via network 150. In some embodiments, storage device 140 may be directly connected to or communicate with one or more components of voice enhancement system 100 (e.g., processing device 110, acquisition device 120, terminal 130). In some embodiments, storage device 140 may be part of processing device 110.
[0057] In some embodiments, one or more components of the speech enhancement system 100 (e.g., processing device 110, acquisition device 120, terminal 130) may have permission to access storage device 140. In some embodiments, one or more components of the speech enhancement system 100 may read and / or modify information related to the target speech when one or more conditions are met.
[0058] Network 150 can facilitate the exchange of information and / or data. In some embodiments, one or more components of the voice enhancement system 100 (e.g., processing device 110, acquisition device 120, terminal 130, and storage device 140) can send / receive information and / or data to / from other components in the voice enhancement system 100 via network 150. For example, processing device 110 can acquire a first signal and a second signal of the target voice from acquisition device 120 or storage device 140 via network 150, and terminal 130 can acquire the voice-enhanced output voice signal of the target voice from processing device 110 or storage device 140 via network 150. In some embodiments, network 150 can be any form of wired or wireless network or any combination thereof. By way of example only, network 150 may include cable networks, wired networks, fiber optic networks, telecommunications networks, intranets, the Internet, local area networks (LANs), wide area networks (WANs), wireless local area networks (WLANs), metropolitan area networks (MANs), public switched telephone networks (PSTNs), Bluetooth networks, cellular networks, near field communication (NFC) networks, Global System for Mobile Communications (GSM) networks, code division multiple access (CDMA) networks, time division multiple access (TDMA) networks, general packet radio service (GPRS) networks, enhanced data rate GSM evolution (EDGE) networks, wideband code division multiple access (WCDMA) networks, high-speed downlink packet access (HSDPA) networks, long-term evolution (LTE) networks, user datagram protocol (UDP) networks, transmission control protocol / Internet protocol (TCP / IP) networks, short message service (SMS) networks, wireless application protocol (WAP) networks, ultra-wideband (UWB) networks, infrared, etc., or any combination thereof. In some embodiments, voice enhancement system 100 may include one or more network access points. For example, the voice enhancement system 100 may include wired or wireless network access points, such as base stations and / or wireless access points 150-1, 150-2, ..., through which one or more components of the voice enhancement system 100 may be connected to the network 150 to exchange data and / or information.
[0059] Those skilled in the art will understand that when the elements or components of the speech enhancement system 100 are executed, the components can be executed via electrical signals and / or electromagnetic signals. For example, when the acquisition device 120 sends a first signal and a second signal of the target speech to the processing device 110, the acquisition device 120 can generate an encoded electrical signal. The acquisition device 120 can then send the electrical signal to an output port. If the acquisition device 120 communicates via a wired network or data transmission line, the output port can be physically connected to a cable that further transmits the electrical signal to the input port of the acquisition device 120. If the acquisition device 120 communicates via a wireless network, the output port of the acquisition device 120 can be one or more antennas that convert electrical signals into electromagnetic signals. Within an electronic device, such as the acquisition device 120 and / or the processing device 110, instructions and / or actions are performed via electrical signals when processing instructions, issuing instructions, and / or performing actions. For example, when processing device 110 retrieves or saves data from a storage medium (e.g., storage device 140), it can send electrical signals to a read / write device of the storage medium, which can read or write structured data in the storage medium. This structured data can be transmitted to the processor in the form of electrical signals via the bus of the electronic device. Here, an electrical signal can refer to a single electrical signal, a series of electrical signals, and / or at least two discontinuous electrical signals.
[0060] Figure 2 This is a schematic diagram of an exemplary computing device 200 according to some embodiments of this specification. In some embodiments, a processing device 110 may be implemented on the computing device 200. Figure 2 As shown, computing device 200 may include memory 210, processor 220, input / output (I / O) 230 and communication port 240.
[0061] The memory 210 can store data / information obtained from the acquisition device 120, terminal 130, storage device 140, or any other component of the voice enhancement system 100. In some embodiments, the memory 210 may include mass storage, removable storage, volatile read-write storage, read-only storage (ROM), etc., or any combination thereof. Exemplary mass storage may include a hard disk, optical disk, solid-state drive, etc. Exemplary removable storage may include a flash drive, floppy disk, optical disk, memory card, compact disk, magnetic tape, etc. Exemplary volatile read-only storage may include random access memory (RAM). Exemplary RAM may include dynamic RAM (DRAM), double-rate synchronous dynamic RAM (DDR SDRAM), static RAM (SRAM), thyristor RAM (T-RAM), and zero-capacitance RAM (Z-RAM), etc. Exemplary ROM may include mask ROM (MROM), programmable ROM (PROM), erasable programmable ROM (PEROM), electronically erasable programmable ROM (EEPROM), optical disc ROM (CD-ROM), and digital universal disk ROM, etc. In some embodiments, memory 210 may store one or more programs and / or instructions to perform the exemplary methods described herein. For example, memory 210 may store a program executable by processing device 110 to implement a speech enhancement method.
[0062] Processor 220 can execute computer instructions (program code) and perform the functions of processing device 110 according to the techniques described herein. Computer instructions may include, for example, routines, programs, objects, components, signals, data structures, procedures, modules, and functions that perform the specific functions described herein. For example, processor 220 can process data acquired from acquisition device 120, terminal 130, storage device 140, and / or any other component of speech enhancement system 100. For example, processor 220 can process a first signal and a second signal of target speech acquired from acquisition device 120 to obtain a speech-enhanced output speech signal. In some embodiments, the output speech signal may be stored in storage device 140, memory 210, etc. In some embodiments, the output speech signal may be output to a broadcasting device such as a speaker via I / O 230. In some embodiments, processor 220 can execute instructions obtained from terminal 130.
[0063] In some embodiments, processor 220 may include one or more hardware processors, such as microcontrollers, microprocessors, reduced instruction set computers (RISC), application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), central processing units (CPUs), graphics processing units (GPUs), physical processing units (PPUs), microcontroller units, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), advanced RISC machines (ARMs), programmable logic devices (PLDs), any circuit or processor capable of performing one or more functions, or any combination thereof.
[0064] For illustrative purposes only, only one processor is described in computing device 200. However, it should be noted that computing device 200 in this specification may also include multiple processors. Therefore, the operations and / or method steps performed by one processor as described in this specification may also be performed jointly or separately by multiple processors. For example, if, in this specification, the processors of computing device 200 simultaneously perform operation A and operation B, it should be understood that operation A and operation B may also be performed jointly or separately by two or more different processors in the computing device. For example, a first processor performs operation A, a second processor performs operation B, or a first processor and a second processor jointly perform operations A and B.
[0065] I / O 230 can input or output signals, data, and / or information. In some embodiments, I / O 230 enables a user to interact with processing device 110. In some embodiments, I / O 230 may include input devices and output devices. Exemplary input devices may include a keyboard, mouse, touchscreen, microphone, etc., or combinations thereof. Exemplary output devices may include display devices, speakers, printers, projectors, etc., or combinations thereof. Exemplary display devices may include liquid crystal displays (LCDs), light-emitting diode (LED) based displays, monitors, flat panel displays, curved screens, television equipment, cathode ray tubes (CRTs), etc., or combinations thereof.
[0066] Communication port 240 can be connected to a network (e.g., network 150) to facilitate data communication. Communication port 240 can establish a connection between processing device 110 and acquisition device 120, terminal 130, or storage device 140. This connection can be a wired connection, a wireless connection, or a combination of both to enable data transmission and reception. Wired connections can include cables, optical fibers, telephone lines, etc., or any combination thereof. Wireless connections can include Bluetooth, Wi-Fi, WiMax, WLAN, ZigBee, mobile networks (e.g., 3G, 4G, 5G, etc.), etc., or combinations thereof. In some embodiments, communication port 240 can be a standardized communication port, such as RS232, RS485, etc. In some embodiments, communication port 240 can be a specially designed communication port. For example, communication port 240 can be designed according to the voice signals that need to be transmitted.
[0067] Figure 3 These are schematic diagrams illustrating exemplary hardware and / or software components of an exemplary mobile device 300 on which a terminal 130 may be implemented, according to some embodiments of this specification. Figure 3 As shown, the mobile device 300 may include a communication unit 310, a display unit 320, a graphics processing unit (GPU) 330, a central processing unit (CPU) 340, an input / output unit 350, a memory 360, and a storage unit 370.
[0068] Central processing unit (CPU) 340 may include interface circuitry and processing circuitry similar to processor 220. In some embodiments, any other suitable components, including but not limited to a system bus or controller (not shown), may also be included within mobile device 300. In some embodiments, mobile operating system 362 (e.g., iOS) TM Andro TM Windows Phone TM One or more applications 364 may be loaded from memory 370 into memory 360 for execution by the central processing unit (CPU) 340. Application 364 may include a browser or any other suitable mobile application for receiving and presenting information related to the target speech and speech enhancement of the target speech from the speech enhancement system on the mobile device 300. Interaction of signals and / or data may be achieved through input / output devices 350 and provided via network 150 to the processing engine 112 and / or other components of the speech enhancement system 100.
[0069] To implement the various modules, units, and their functions described above, a computer hardware platform can be used as one or more components (e.g., Figure 1The hardware platform of the processing device 110 (modules) described herein. Since these hardware components, operating systems, and programming languages are common, it can be assumed that those skilled in the art are familiar with these technologies and can provide the information needed for route planning based on the technologies described herein. A computer with a user interface can be used as a personal computer (PC) or other type of workstation or terminal device. After proper programming, a computer with a user interface can be used as a processing device such as a server. It can be assumed that those skilled in the art are also familiar with this type of computer device's structure, programs, or general operation. Therefore, no additional explanation is provided with reference to the accompanying drawings.
[0070] Figure 4 This is an exemplary block diagram of a speech enhancement system according to some embodiments of this specification. In some embodiments, the speech enhancement system 100 may be implemented on a processing device 110. Figure 4 As shown, the processing device 110 may include an acquisition module 410, a processing module 420, and a generation module 430.
[0071] The acquisition module 410 can be used to acquire a first signal and a second signal of the target speech. In some embodiments, the target speech may include speech emitted by a target sound source. In some embodiments, the target speech signal can be acquired at different locations using different acquisition devices (e.g., different microphones). For example, the first signal may be the target speech signal acquired by a first microphone (or front microphone) based on a first location, and the second signal may be the target speech signal acquired by a second microphone (or rear microphone) based on a second location. In some embodiments, the acquisition module 410 can directly acquire the first signal and the second signal of the target speech from the different acquisition devices. In some embodiments, the first signal and the second signal may be stored in a storage device (e.g., storage device 140, memory 210, memory 370, external storage device, etc.). The acquisition module 410 can acquire the first signal and the second signal from the storage device.
[0072] Processing module 420 can be used to process the first signal and the second signal based on the target speech position, the first position, and the second position to determine a first coefficient. In some embodiments, processing module 420 can be used to determine the first coefficient based on an adaptive null-forming (ANF) algorithm. For example, processing module 420 can be used to perform a differential operation on the first signal and the second signal based on the target speech position, the first position, and the second position to obtain a signal pointing in a first direction and a signal pointing in a second direction. The signal pointing in the first direction and the signal pointing in the second direction contain different proportions of effective signal. Further, processing module 420 can be used to determine a fourth signal based on the signal pointing in the first direction and the signal pointing in the second direction. For example, processing module 420 can use a Wiener filtering algorithm to filter the signal pointing in the first direction and the signal pointing in the second direction using an adaptive filter to determine the fourth signal. In some embodiments, processing module 420 can enhance the low-frequency components in the fourth signal to obtain a third signal. Further, processing module 420 can be used to determine the first coefficient based on the third signal. For example, processing module 420 can determine the ratio of the third signal to the first signal or the second signal as the first coefficient. Optionally or additionally, the processing module 420 may update the adaptive parameters of the adaptive differential operation based on the fourth signal, the signal pointing in the first direction, and the signal pointing in the second direction.
[0073] In some embodiments, to determine the first coefficient, the processing module 420 may further determine an estimated signal-to-noise ratio (SNR) of the target speech based on the signal pointing in the first direction and the signal pointing in the second direction. For example, the estimated SNR may be the ratio of the signal pointing in the first direction to the signal pointing in the second direction. Further, the processing module 420 determines the first coefficient based on the estimated SNR.
[0074] The processing module 420 can also be used to determine multiple parameters related to multiple sound source directions based on the first signal and the second signal. In some embodiments, the multiple sound source directions may include preset sound source directions. In some embodiments, the processing module 420 may perform a difference operation on the first signal and the second signal based on each sound source direction, a first position, and a second position to determine parameters related to each sound source direction. In some embodiments, the parameters may include a likelihood function. Each parameter may correspond to the probability of emitting sound from a sound source direction to form the first signal and the second signal.
[0075] The processing module 420 can also be used to determine a second coefficient based on the plurality of parameters and the target speech position. In some embodiments, to determine the second coefficient, the processing module 420 can determine the direction of the synthesized sound source based on the plurality of parameters. The synthesized sound source can be considered as a virtual sound source formed by combining the target sound source and a noise source. As an example only, the processing module 420 can determine the parameter with the largest value among the plurality of parameters. The parameter with the largest value can represent the highest probability of emitting sound from its corresponding sound source direction to form the first signal and the second signal. Thus, the processing module 420 can determine that the sound source direction corresponding to the parameter with the largest value is the direction of the synthesized sound source. Further, the processing module 420 can determine the second coefficient based on the direction of the synthesized sound source and the target speech position. For example, the processing module 420 can determine whether the target speech position is located in the direction of the synthesized sound source, or whether the target speech position is within a certain angular range of the direction of the synthesized sound source. In response to the target speech position being located in the direction of the synthesized sound source or within a certain angular range of the direction of the synthesized sound source, the second coefficient is set to a first value. In response to the target speech location not being located in the direction of the synthesized sound source or not being within a certain angular range of the direction of the synthesized sound source, the second coefficient is set to a second value. Optionally or additionally, the processing module 420 may smooth the second coefficient based on a smoothing factor. For example, the processing module 420 may determine the second coefficient based on the angle between the target speech location and the direction of the synthesized sound source using a regression function.
[0076] The generation module 430 can be used to process the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain a first output speech signal with enhanced speech corresponding to the target speech. In some embodiments, the generation module 430 can be used to perform weighted processing on the first signal and / or the second signal based on the first coefficient and the second coefficient. For example, the generation module 430 can assign corresponding weights to the third signal obtained based on the first signal and the second signal according to the value of the first coefficient, and assign corresponding weights to the first signal or the third signal according to the value of the second coefficient. The generation module 430 can further process the weighted signals to obtain the first output speech signal with enhanced speech. For example, the first output speech signal can be the average of the weighted third signal and the first signal. Another example is that the first output speech signal can be the product of the weighted third signal and the first signal. Yet another example is that the first output speech signal can be the larger value between the weighted third signal and the first signal. Yet another example is that the generation module 430 can weight the third signal based on the first coefficient and then perform weighted processing again based on the second coefficient.
[0077] It should be noted that the above description of the processing device 110 and its modules is for convenience only and should not be construed as limiting this specification to the scope of the illustrated embodiments. It is understood that those skilled in the art, after understanding the principles of this system, can arbitrarily combine the various modules or construct subsystems connected to other modules without departing from these principles. For example, Figure 4 The acquisition module 410 and processing module 420 disclosed herein can be different modules within a single system, or a single module can implement the functions of two or more of the aforementioned modules. For example, the acquisition module 410 and processing module 420 can be two separate modules, or a single module can simultaneously perform the functions of acquiring and processing the target speech. Such variations are all within the scope of protection of this specification.
[0078] It should be understood that Figure 4 The systems and modules shown can be implemented in various ways. For example, in some embodiments, the systems and modules can be implemented by hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the methods and systems described above can be implemented using computer-executable instructions and / or included in processor control code, for example, on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The systems and modules of this specification can be implemented not only by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., but also by software, for example, executed by various types of processors, or by a combination of the aforementioned hardware circuits and software (e.g., firmware).
[0079] Figure 5 This is an exemplary flowchart of a speech enhancement method according to some embodiments of this specification. In some embodiments, method 500 may be executed by processing device 110, processing engine 112, or processor 220. For example, method 500 may be stored in a storage device (e.g., storage device 140 or storage unit of processing device 110) as a program or instructions, and executed by processing device 110, processing engine 112, processor 220, or... Figure 4 When the module shown executes a program or instructions, it can implement method 500. In some embodiments, method 500 may be accomplished using one or more additional operations / steps not described below, and / or not through one or more operations / steps discussed below. Additionally, as Figure 5The order of operations / steps shown is not restrictive.
[0080] In step 510, the processing device 110 (e.g., acquisition module 410) can acquire the first signal and the second signal of the target speech.
[0081] In some embodiments, the target speech may include speech emitted by a target sound source. The target sound source may be a user, a robot (e.g., an automated response robot, a robot that converts human input data such as text, gestures, etc., into speech signals for broadcast), or other biological entities and devices capable of emitting speech information. In some embodiments, the speech emitted by the target sound source may serve as a valid signal. In some embodiments, the target speech may also include useless or interfering noise signals, such as ambient noise or sounds from other sound sources besides the target sound source. Exemplary noise may include additive noise, white noise, multiplicative noise, etc., or any combination thereof. Additive noise refers to an independent noise signal unrelated to the speech signal; multiplicative noise refers to a noise signal proportional to the speech signal; and white noise refers to a noise signal with a constant power spectrum. In some embodiments, when the target speech includes noise signals, the target speech may be a synthesized signal of a valid signal and a noise signal. The synthesized signal may be equivalent to a speech signal emitted by a synthesized sound source consisting of a target sound source and a noise source.
[0082] In some embodiments, different acquisition devices (e.g., different microphones) can be used to acquire target speech signals at different locations. Taking a dual-microphone setup as an example... Figure 6 This is a schematic diagram of an exemplary dual microphone according to some embodiments of this specification. Figure 6As shown, the target sound source (such as the user's mouth) is located to the upper left of the dual microphones, and the angle formed by the direction of the target sound source pointing to the dual microphones (e.g., the direction of the target sound source pointing to the first microphone A) and the line connecting the dual microphones is θ. The first signal Sig 1 can be the signal of the target speech acquired by the first microphone A (or the front microphone) based on the first position, and the second signal Sig 2 can be the signal of the target speech acquired by the second microphone B (or the rear microphone) based on the second position. As an example only, the first position and the second position can be two positions with a distance d and different distances relative to the target sound source (such as the user's mouth). d can be set according to actual needs; for example, in a specific scenario, d can be set to not less than 0.5 cm or not less than 1 cm. In some embodiments, the first signal or the second signal can include an electrical signal generated by the acquisition device after receiving the target speech (or an electrical signal generated after further processing), which can reflect the positional information of the target speech relative to the acquisition device. In some embodiments, the first signal and the second signal can be the temporal representation of the target speech. For example, the processing device 110 can frame the signals acquired by the first microphone A and the second microphone B to obtain the first signal and the second signal respectively. Taking the acquisition of a first signal as an example, the processing device 110 can divide the signal acquired by the first microphone into multiple segments in the time domain (e.g., equally divided or overlappingly divided into multiple segments with a duration of 10-30 ms). Each segment can serve as a frame signal, and the first signal can include one or more of these frames. In some alternative embodiments, the first signal and the second signal can be the frequency domain representation of the target speech. For example, the processing device 110 can perform a Fast Fourier Transform (FFT) on the aforementioned one or more frames to obtain the first signal or the second signal. Optionally, before performing the FFT on the frame signal, the frame signal can be windowed and smoothed. Specifically, the processing device 110 can multiply the frame signal by a window function to periodically expand the frame signal and obtain a periodic continuous signal. Exemplary window functions can include rectangular windows, Hanning windows, flat-top windows, exponential windows, etc. The windowed and smoothed frame signal can be further subjected to an FFT to generate the first signal or the second signal.
[0083] In some embodiments, the difference between the first signal and the second signal may be related to the intensity, signal amplitude, phase difference, etc. of the target speech and noise signals at different acquisition locations.
[0084] In step 520, the processing device 110 (e.g., processing module 420) can process the first signal and the second signal based on the target speech location, the first location, and the second location to determine the first coefficient.
[0085] In some embodiments, the processing device 110 may determine the first coefficient based on an Adaptive Null-Forming (ANF) algorithm. The ANF algorithm may include two differential beamformers and an adaptive filter. In some embodiments, the processing device 110 may perform differential operations on the first signal and the second signal using the two differential beamformers based on the target speech position, the first position, and the second position to obtain a signal pointing in a first direction and a signal pointing in a second direction. For example, the processing device 110 may perform time-delay processing on the first signal and the second signal based on the target speech position, the first position, and the second position according to the differential microphone principle. Then, it may perform differential operations on the time-delayed first signal and the second signal to obtain a signal pointing in the first direction and a signal pointing in the second direction. In some embodiments, the signal pointing in the first direction is a signal pointing towards the target sound source, and the signal pointing in the second direction is a signal pointing in the opposite direction to the target sound source. The signals pointing in the first direction and the signals pointing in the second direction contain different proportions of effective signal. The effective signal refers to the speech emitted by the target sound source. For example, the signal pointing in the first direction may contain a larger proportion of effective signal (and / or a smaller proportion of noise signal). The signal pointing in the second direction may contain a small proportion of valid signal (and / or a large proportion of noise signal). In some embodiments, the signal pointing in the first direction and the signal pointing in the second direction may correspond to two directional microphones. Further, the processing device 110 may determine a third signal corresponding to the valid signal based on the signal pointing in the first direction and the signal pointing in the second direction. For example, the processing device 110 may perform an adaptive differential operation on the signal pointing in the first direction and the signal pointing in the second direction to determine a fourth signal. As an example only, during the adaptive differential operation, the processing device 110 may filter the signal pointing in the first direction and the signal pointing in the second direction using the Wiener filtering algorithm and the adaptive filter to determine the fourth signal. In some embodiments, during the adaptive differential operation, the processing device 110 may adjust the parameters of the adaptive filter so that the zero point of the cardiogram corresponding to the fourth signal points in the direction of noise. In some embodiments, the processing device 110 may enhance the low-frequency components in the fourth signal to obtain the third signal. Further, the processing device 110 may determine the first coefficient based on the third signal. For example, the first coefficient may be the ratio of the third signal to the first signal or the second signal. For more information on determining the first coefficient based on the third signal, please refer to [link to relevant documentation]. Figure 8 Its description will not be repeated here.
[0086] In some embodiments, to determine the first coefficient, the processing device 110 may determine an estimated signal-to-noise ratio (SNR) of the target speech based on the signal pointing in the first direction and the signal pointing in the second direction. For example, the estimated SNR may be the ratio between the signal pointing in the first direction and the signal pointing in the second direction. Further, the processing device 110 may determine the first coefficient based on the estimated SNR. In some embodiments, the processing device 110 may determine the first coefficient based on a mapping relationship between the estimated SNR and the first coefficient. The mapping relationship may take various forms, such as a mapping relationship database or a relational function. More information on determining the first coefficient based on the estimated SNR can be found at [link to relevant documentation]. Figure 9 Its description will not be repeated here.
[0087] In some embodiments, the first coefficient can reflect the impact of noise on the effective signal. Taking a first coefficient determined based on an estimated signal-to-noise ratio (SNR) as an example, the estimated SNR can be the ratio between the signal pointing in a first direction and the signal pointing in a second direction. The signal pointing in the first direction may contain a larger proportion of effective signal (and / or a smaller proportion of noise). The signal pointing in the second direction may contain a smaller proportion of effective signal (and / or a larger proportion of noise). Therefore, the magnitude of the noise signal can affect the value of the estimated SNR, thereby affecting the value of the first coefficient. For example, the larger the noise signal, the smaller the estimated SNR, and the value of the first coefficient determined based on the estimated SNR will change accordingly. Thus, the first coefficient can reflect the impact of noise on the effective signal.
[0088] In some embodiments, the first coefficient is related to the direction of the noise source. For example, the first coefficient may have a larger value when the noise source direction is close to the target sound source direction; and a smaller value when the noise source direction deviates from the target sound source direction by a large angle. The noise source direction is the direction of the noise source relative to the dual microphones, and the target sound source direction is the direction of the target sound source (such as the user's mouth) relative to the dual microphones. The processing device 110 can process the third signal corresponding to the valid signal according to the first coefficient. For example, the first coefficient may represent the weight of the third signal corresponding to the valid signal in the speech enhancement process. As an example only, when the first coefficient is "1", it means that the third signal can be completely preserved as part of the enhanced speech signal; when the first coefficient is "0", it means that the third signal is completely filtered out from the enhanced speech signal.
[0089] In some embodiments, when the angular difference between the noise source direction and the target sound source direction is large, the first coefficient can have a small value, and the third signal processed according to the first coefficient can be attenuated or removed; when the angular difference between the noise source direction and the target sound source direction is small, the first coefficient can have a large value, and the third signal processed according to the first coefficient can be retained as part of the enhanced speech signal. Therefore, when the angular difference between the noise source direction and the target sound source direction is large, the ANF algorithm can achieve better filtering results. Figure 7 This diagram illustrates the filtering effect of the ANF algorithm at different noise angles, based on some embodiments of this specification. The noise angle refers to the angle between the direction of the noise source and the direction of the target sound source. For example... Figure 7 As shown in Figure af, the filtering effect of the ANF algorithm is illustrated when the noise angle is 180°, 150°, 120°, 90°, 60°, and 30°, respectively. Figure 7 It can be seen that the ANF algorithm performs well when the noise angle is large (e.g., 180°, 150°, 120°, 90°), and poorly when the noise angle is small (e.g., 60°, 30°).
[0090] In step 530, the processing device 110 (e.g., processing module 420) can determine multiple parameters related to the directions of multiple sound sources based on the first signal and the second signal.
[0091] In some embodiments, the plurality of sound source directions may include preset sound source directions. For example, the plurality of sound source directions may have preset incident angles (e.g., 0°, 30°, 60°, 90°, 120°, 180°, etc.). The sound source directions can be selected and / or adjusted according to actual needs, and are not limited here. In some embodiments, the processing device 110 may perform differential operations on the first signal and the second signal based on each sound source direction, the first position, and the second position to determine parameters related to each sound source direction. For example, each of the plurality of sound source directions may correspond to a time delay, and the processing device 110 may perform differential operations on the first signal and the second signal based on the time delay. Further, the processing device 110 may calculate parameters related to the sound source directions based on the differential operations. In some embodiments, the parameters may include a likelihood function. For example, for each signal point in the current frame, the processing device 110 may calculate a likelihood function corresponding to each of the plurality of sound source directions. In some embodiments, the likelihood function may correspond to the probability of emitting sound from the sound source direction to form the first signal and the second signal. As an example only, the likelihood function value when the sound source direction is θ = 30° is 0.8 after normalization, which can be interpreted as an 80% probability that the sound emitted from the sound source direction of θ = 30° will form the first signal and the second signal.
[0092] In step 540, the processing device 110 (e.g., processing module 420) can determine a second coefficient based on the plurality of parameters and the target speech location.
[0093] In some embodiments, the processing device 110 may determine the direction of the synthesized sound source based on the plurality of parameters. The synthesized sound source can be considered as a virtual sound source formed by combining a target sound source and a noise source. That is, the signal generated by the target sound source and the noise source at the dual microphones (e.g., the first signal and the second signal) can be equivalently generated by the synthesized sound source at the dual microphones.
[0094] In some embodiments, to determine the direction of the synthesized sound source, the processing device 110 can determine the parameter with the largest value among the plurality of parameters. The parameter with the largest value may represent the highest probability of sound being emitted from its corresponding sound source direction to form the first signal and the second signal. Thus, the processing device 110 can determine that the sound source direction corresponding to the parameter with the largest value is the direction of the synthesized sound source. As another example, to determine the direction of the synthesized sound source, the processing device 110 can construct a plurality of directional microphones with poles pointing towards the plurality of sound source directions. The response of each of the plurality of directional microphones is a cardiogram. For ease of description, the cardiograms corresponding to the plurality of sound source directions can be called simulated cardiograms. The poles of each simulated cardiogram can point towards the corresponding sound source direction. Further, based on the first signal and the second signal, the processing device 110 can calculate a likelihood function corresponding to each of the plurality of sound source directions. The response of the likelihood function corresponding to the plurality of sound source directions can be a cardiogram. For ease of description, the cardiogram corresponding to the likelihood function can be called a synthesized cardiogram (or an actual cardiogram). The poles of the synthesized cardiogram point towards the direction of the synthesized sound source. The processing device 110 can determine the simulated cardioid whose poles are closest to those of the actual cardioid, and determine the direction of the sound source corresponding to the simulated cardioid as the direction of the synthesized sound source.
[0095] In some embodiments, the processing device 110 may determine the second coefficient based on the direction of the synthesized sound source and the target speech position. For example, the processing device 110 may determine whether the target speech position is located in the direction of the synthesized sound source, or whether the target speech position is within a certain angular range of the direction of the synthesized sound source. In response to the target speech position being located in the direction of the synthesized sound source or within a certain angular range of the direction of the synthesized sound source, the second coefficient is set to a first value. In response to the target speech position not being located in the direction of the synthesized sound source or not within a certain angular range of the direction of the synthesized sound source, the second coefficient is set to a second value. Optionally or additionally, the processing device 110 may smooth the second coefficient based on a smoothing factor. For example, the processing device 110 may determine the second coefficient using a regression function based on the angle between the target speech position and the direction of the synthesized sound source. More information on determining the second coefficient can be found in [link to relevant documentation]. Figure 10 Its description will not be repeated here.
[0096] In some embodiments, the second coefficient may reflect the direction of the synthesized sound source relative to the target sound source, thereby reducing or removing synthesized sound sources that are not in the direction of the target sound source and / or synthesized sound sources that deviate from the direction of the target sound source by a certain angle. In some embodiments, the second coefficient may be used to filter out noise where the angle difference between the direction of the noise source and the direction of the target sound source exceeds a certain threshold. For example, when the angle difference between the direction of the noise source and the direction of the target sound source exceeds a certain threshold, the second coefficient may have a small value; when the angle difference between the direction of the noise source and the direction of the target sound source is less than a certain threshold, the second coefficient may have a large value. The processing device 110 may process the first signal or the third signal according to the second coefficient. For example, the second coefficient may represent the weight of the first signal in the speech enhancement process. As an example only, when the second coefficient is "1", it means that the first signal can be completely preserved as part of the enhanced speech signal; when the second coefficient is "0", it means that the first signal is completely filtered out from the enhanced speech signal.
[0097] In step 550, the processing device 110 (e.g., generation module 430) can process the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain a first output speech signal with speech enhancement corresponding to the target speech.
[0098] In some embodiments, the processing device 110 can perform weighted processing on the first signal and / or the second signal based on the first coefficient and the second coefficient. Taking the first coefficient as an example, the processing device 110 can assign corresponding weights to the third signal obtained based on the first signal and the second signal according to the value of the first coefficient. For example, the processing device 110 can assign corresponding weights to the third signal according to the range of the first coefficient. Another example is that the processing device 110 can directly use the value of the first coefficient as the weight of the third signal. Yet another example is that when the value of the first coefficient is less than a preset first coefficient threshold, the processing device 110 can set the weight of the third signal to 0. Taking the second coefficient as an example, the processing device 110 can assign corresponding weights to the first signal or the third signal according to the value of the second coefficient. The processing device 110 can further process the weighted signals to obtain a first output speech signal after speech enhancement. For example, the first output speech signal can be the average of the weighted third signal and the first signal. Another example is that the first output speech signal can be the product of the weighted third signal and the first signal. Yet another example is that the first output speech signal can be the larger value between the weighted third signal and the first signal. For example, the generation module 430 can weight the third signal based on the first coefficient, and then perform a second weighting process based on the second coefficient.
[0099] It should be noted that the above description of the speech enhancement method 500 is for convenience only and should not be construed as limiting this specification to the scope of the embodiments described. It is understood that those skilled in the art, after understanding the principle of this method, can arbitrarily combine the various steps, or add or delete any steps, without departing from this principle.
[0100] In some embodiments, the speech enhancement method 500 may further include a single-microphone filtering process. For example, the processing device 110 may perform single-microphone filtering on the first output speech signal based on a single-microphone filtering algorithm. As another example, the processing device 110 may process the first signal and / or the second signal based on a single-microphone filtering algorithm to obtain a third coefficient, and then filter the first output speech signal based on the third coefficient. More information on the single-microphone filtering process can be found in [link to relevant documentation]. Figure 11 , Figure 12 Its description will not be repeated here.
[0101] In some embodiments, the processing device 110 may also perform speech enhancement processing based on a fourth coefficient. For example, the processing device 110 determines the fourth coefficient based on the energy difference between the first and second signals. It then processes the first and / or second signals based on any one or a combination of the first, second, and fourth coefficients to obtain the enhanced output speech signal. More information on speech enhancement based on the fourth coefficient can be found at [link to relevant documentation]. Figure 13 Its description will not be repeated here.
[0102] In some embodiments, the speech enhancement method 500 described above can be implemented on a first signal and / or a second signal obtained after preprocessing (e.g., framing, windowing smoothing, FFT transform, etc.). That is, the first output speech signal can be a single-frame speech signal. Therefore, the speech enhancement method 500 may also include a post-processing procedure. Exemplary post-processing may include inverse FFT transform, frame splicing, etc. After the post-processing procedure, the processing device 110 can obtain a continuous output speech signal.
[0103] Figure 8 This is an exemplary flowchart illustrating a method for determining a first coefficient according to some embodiments of this specification. In some embodiments, method 800 may be executed by processing device 110, processing engine 112, or processor 220. For example, method 800 may be stored as a program or instructions in a storage device (e.g., storage device 140 or storage unit of processing device 110), and executed by processing device 110, processing engine 112, processor 220, or... Figure 4When the module shown executes a program or instructions, it can implement method 800. In some embodiments, operation 520 described in method 500 can be implemented by method 800. In some embodiments, method 800 can be accomplished using one or more additional operations / steps not described below, and / or not by one or more operations / steps discussed below. Additionally, as Figure 8 The order of operations / steps shown is not restrictive.
[0104] In some embodiments, the processing device 110 may determine the first coefficient based on an Adaptive Null-Forming (ANF) algorithm. The ANF algorithm may include two differential beamformers and an adaptive filter. The two differential beamformers may perform differential processing on the first signal and the second signal to form a signal pointing in a first direction and a signal pointing in a second direction. The adaptive filter may adaptively filter the signal pointing in the first direction and the signal pointing in the second direction to obtain a third signal corresponding to the effective signal. Figure 8 As shown, method 800 may include:
[0105] Step 810: The processing device 110 can perform differential operations on the first signal and the second signal based on the target voice position, the first position and the second position to obtain a signal pointing to the first direction and a signal pointing to the second direction.
[0106] In some embodiments, the processing device 110 can perform time delay processing on the first signal and the second signal based on the target speech location, the first location, and the second location, according to the differential microphone principle. For example, as Figure 6 Given that the distance between the front microphone A and the rear microphone B is d, and the target sound source has an incident angle θ, the propagation time of the target sound source between the front microphone A and the rear microphone B can be expressed as:
[0107] τ=dcosθ / c,(1)
[0108] Where c is the speed of sound propagation. The propagation time τ can be used as the time delay between the first signal and the second signal. When θ = 0°, τ = d / c. According to the principle of differential microphones, signals pointing in the first direction and signals pointing in the second direction can be obtained:
[0109] x s (t)=sig1(t)-sig2(t-τ), (2)
[0110] x n (t)=sig2(t)-sig1(t-τ), (3)
[0111] Where t represents each time point of the current frame, sig1 represents the first signal, and sig2 represents the second signal.
[0112] According to the above formulas (2) and (3), by delaying the first signal sig1 and the second signal sig2 by time delay τ respectively and then performing differential division, the signal x pointing to the first direction can be obtained. s and the signal x pointing in the second direction n The signal x pointing in the first direction s This can correspond to a first directional microphone, whose response is a cardioid, with its poles pointing in the direction of the target sound source. The signal x pointing in the second direction... n This corresponds to a second directional microphone, whose response is a cardiogram with its zero point pointing in the direction of the target sound source.
[0113] In some embodiments, the signal x pointing in the first direction s and the signal x pointing in the second direction n It can contain different proportions of effective signals. For example, the signal x pointing in the first direction. s It may contain a large proportion of effective signal (and / or a small proportion of noise signal). The signal x pointing in the second direction n It may contain a small proportion of effective signal (and / or a large proportion of noise signal).
[0114] In step 820, the processing device 110 can perform adaptive differential operation on the signal pointing to the first direction and the signal pointing to the second direction to determine the fourth signal.
[0115] In some embodiments, the processing device 110 can filter the signal pointing in the first direction and the signal pointing in the second direction using the adaptive filter based on the Wiener filtering algorithm. The adaptive filter can be a Least Mean Square (LMS) filter. During filtering, the signal x pointing in the first direction... s The desired signal that can be used as the LMS filter is the signal x pointing in the second direction. n This can be used as reference noise for the LMS filter. Based on the desired signal and the reference noise, the processing device 110 can use the LMS filter to process the signal x pointing in the first direction. s and the signal x pointing in the second direction n Adaptive filtering (i.e., adaptive differential operation) is performed to determine the fourth signal. The fourth signal may be the signal after noise has been filtered out. In some embodiments, an exemplary process of adaptive differential operation may be shown in the following formula:
[0116] y = s-x n (4)
[0117] Where y represents the fourth signal (i.e., the output signal of the LMS filter), and w represents the adaptive parameters of the adaptive differential operation (i.e., the coefficients of the LMS filter).
[0118] In some embodiments, during the adaptive difference operation, the processing device 110 can base its operation on the fourth signal y and the signal x pointing in the first direction. s and the signal x pointing in the second direction n The adaptive parameter w of the adaptive differential operation is updated. For example, based on the first and second signals in each frame, the processing device 110 can acquire a signal x pointing in the first direction. s and the signal x pointing in the second direction n Furthermore, the processing device 110 can update the adaptive parameter w using gradient descent, causing the loss function of the adaptive difference operation (e.g., the mean squared error loss function) to gradually converge.
[0119] In step 830, the processing device 110 can enhance the low-frequency components in the fourth signal to obtain the third signal.
[0120] In some embodiments, the differential beamformer may have high-pass filtering characteristics. When differential processing of the first signal and the second signal is performed using the differential beamformer, low-frequency components in the first signal and the second signal may be attenuated. Accordingly, the low-frequency components in the fourth signal y obtained after adaptive differential operation are attenuated. In some embodiments, the processing device 110 may enhance the low-frequency components in the fourth signal y using a compensation filter. By way of example only, the compensation filter may be as shown in the following formula:
[0121]
[0122] Among them, W EQ This represents a compensation filter, where ω represents the frequency of the fourth signal y. c This represents the cutoff frequency of the high-pass filter. In some embodiments, ω is an example. c The possible values are:
[0123] ω c =0.5πc / d, (6)
[0124] Where c represents the speed of sound propagation and d represents the distance between the two microphones.
[0125] In some embodiments, the processing device 110 may be based on the compensation filter W EQThe fourth signal y is filtered to obtain the third signal. For example, the third signal can be the fourth signal y and the compensation filter W. EQ The product of.
[0126] In step 840, the processing device 110 can determine the first coefficient based on the third signal.
[0127] In some embodiments, the processing device 110 may determine the ratio of the third signal to the first signal or the second signal, and determine a first coefficient based on the ratio.
[0128] It should be noted that the above description of method 800 is for convenience only and should not be construed as limiting this specification to the scope of the illustrated embodiments. It is understood that those skilled in the art, after understanding the principle of this method, can arbitrarily combine the various steps, or add or delete any steps, without departing from this principle.
[0129] For illustrative purposes only, the operation of method 800 in the above embodiments involves processing the first and second signals in the time domain. It should be understood that one or more operations in method 800 can also be performed in the frequency domain. For example, time delay processing of the first and second signals in the time domain can also be equivalent to phase shifting the first and second signals in the frequency domain. In some embodiments, step 830 is not mandatory; that is, the fourth signal obtained in step 820 can be used directly as the third signal without low-frequency enhancement.
[0130] Figure 9 This is an exemplary flowchart illustrating a method for determining a first coefficient according to some embodiments of this specification. In some embodiments, method 900 may be executed by processing device 110, processing engine 112, or processor 220. For example, method 900 may be stored as a program or instructions in a storage device (e.g., storage device 140 or storage unit of processing device 110), and executed by processing device 110, processing engine 112, processor 220, or... Figure 4 When the module shown executes a program or instructions, it can implement method 900. In some embodiments, operation 520 described in method 500 can be implemented by method 900. In some embodiments, method 900 can be accomplished using one or more additional operations / steps not described below, and / or not by one or more operations / steps discussed below. Additionally, as Figure 9 The order of operations / steps shown is not restrictive. For example... Figure 9 As shown, method 900 may include:
[0131] In step 910, the processing device 110 can perform differential operations on the first signal and the second signal based on the target voice location, the first location, and the second location to obtain a signal pointing in the first direction and a signal pointing in the second direction. The signal pointing in the first direction and the signal pointing in the second direction contain different proportions of effective signal.
[0132] In some embodiments, it can be achieved by executing Figure 8 Step 810 is used to execute step 910, which will not be repeated here.
[0133] In step 920, the processing device 110 can determine the estimated signal-to-noise ratio of the target speech based on the signal pointing in the first direction and the signal pointing in the second direction.
[0134] In some embodiments, the signal pointing in the first direction may contain a large proportion of effective signal (and / or a small proportion of noise signal). The signal pointing in the second direction may contain a small proportion of effective signal (and / or a large proportion of noise signal). The estimated signal-to-noise ratio can be expressed as the ratio between the signal pointing in the first direction and the signal pointing in the second direction (i.e., x). s / x n In some embodiments, different estimated signal-to-noise ratios (SNRs) can correspond to different synthetic sound source incident angles θ. For example, a larger estimated SNR can correspond to a smaller synthetic sound source incident angle θ. In some embodiments, the synthetic sound source incident angle θ can reflect the influence of the noise signal on the effective signal. For example, when the noise signal has a significant influence on the effective signal (e.g., the angular difference between the noise source direction and the target sound source direction is large), the synthetic sound source incident angle θ can have a large value; when the noise signal has a minor influence on the effective signal (e.g., the angular difference between the noise source direction and the target sound source direction is small), the synthetic sound source incident angle θ can have a small value. Thus, the estimated SNR can reflect the direction of the synthetic sound source and further reflect the influence of the noise signal on the effective signal.
[0135] In step 930, the processing device 110 can determine the first coefficient based on the estimated signal-to-noise ratio.
[0136] In some embodiments, the processing device 110 may determine the first coefficient based on a mapping relationship between the estimated signal-to-noise ratio and the first coefficient. The mapping relationship may take various forms, such as a mapping database or a relational function.
[0137] In some embodiments, different noise source directions may correspond to different synthetic sound source incident angles θ, and correspondingly, different estimated signal-to-noise ratios (SNRs). That is, the estimated SNR may be related to the noise source direction (i.e., the degree of influence of the noise signal on the effective signal). Therefore, different first coefficients can be determined for different estimated SNRs. For example, when the estimated SNR is small, the corresponding synthetic sound source incident angle θ may have a large value, indicating a greater influence of the noise signal on the effective signal. Correspondingly, the third signal corresponding to the effective signal may contain a large proportion of noise. Therefore, the value of the first coefficient can be determined to weaken or remove the third signal. The processing device 110 can process the third signal corresponding to the effective signal according to the first coefficient. For example, the first coefficient may represent the weight of the third signal corresponding to the effective signal in the speech enhancement process. As an example only, when the first coefficient is "1", it means that the third signal can be completely preserved as part of the enhanced speech signal; when the first coefficient is "0", it means that the third signal is completely filtered out from the enhanced speech signal. In some embodiments, a mapping database between the estimated SNR and the first coefficient can be established. The processing device 110 can retrieve the database based on the estimated signal-to-noise ratio to determine the first coefficient.
[0138] In some embodiments, the processing device 110 may further determine the first coefficient based on a relationship function between the estimated signal-to-noise ratio and the first coefficient. For example, the relationship function may be as shown in the following formula:
[0139]
[0140] in, This represents the estimated signal-to-noise ratio.
[0141] It should be noted that the above description of method 900 is for convenience only and should not be construed as limiting this specification to the scope of the embodiments described. It is understood that those skilled in the art, after understanding the principle of this method, can arbitrarily combine the various steps, or add or delete any steps, without departing from this principle.
[0142] Figure 10 This is an exemplary flowchart illustrating a method for determining a second coefficient according to some embodiments of this specification. In some embodiments, method 1000 may be executed by processing device 110, processing engine 112, or processor 220. For example, method 1000 may be stored as a program or instructions in a storage device (e.g., storage device 140 or storage unit of processing device 110), and executed by processing device 110, processing engine 112, processor 220, or... Figure 4When the module shown executes a program or instructions, it can implement method 1000. In some embodiments, operations 530 and 540 described in method 500 can be implemented by method 1000. In some embodiments, method 1000 can be accomplished using one or more additional operations / steps not described below, and / or without using one or more operations / steps discussed below. Additionally, as Figure 10 The order of operations / steps shown is not restrictive. For example... Figure 10 As shown, method 1000 may include:
[0143] In step 1010, the processing device 110 can perform differential operations on the first signal and the second signal based on each sound source direction, the first position and the second position to determine the parameters related to each sound source direction.
[0144] In some embodiments, the multiple sound source directions may include preset sound source directions. The sound source directions can be selected and / or adjusted according to actual needs, and are not limited here. For example, the multiple sound sources may have preset incident angles θ = (θ1, θ2, ..., θ...). n (For example, 0°, 30°, 60°, 90°, 120°, 150°, 180°, etc.). The processing device 110 can perform differential operations on the first signal and the second signal based on each sound source direction, the first position, and the second position. For example, each of the plurality of sound source directions can correspond to a time delay, and the processing device 110 can construct a time delay combination τ = (τ1, τ2, ..., τ...) corresponding to the plurality of sound source directions. n Based on the aforementioned time delay combination, the processing device 110 can perform differential operations on the first signal and the second signal to construct multiple directional microphones. The poles of the multiple directional microphones can point to the directions of the multiple sound sources. The response of each of the multiple directional microphones is a cardiogram. For ease of description, the cardiograms corresponding to the multiple sound source directions can be called simulated cardiograms. The poles of each simulated cardiogram can point to the corresponding sound source direction.
[0145] In some embodiments, the processing device 110 may calculate parameters related to the sound source directions based on the differential operation. The parameters may include a likelihood function. For example, for each signal point in the current frame, the processing device 110 may calculate a likelihood function corresponding to each of the plurality of sound source directions. For example, the likelihood function may be expressed as follows:
[0146] LH i (f,t)=-|exp(-2θ i )sig1(f,t)-sig2(f,t)| 2 (8)
[0147] Among them, LH i (f,t) represents the likelihood function corresponding to frequency f at time t, and is the time-frequency domain table of the first signal when the source direction is θ. sig1(f,t) and sig2(f,t) are respectively the source direction when the source direction is θ. i The time-frequency domain expression of the first and second signals, exp(-2θ) i -2θ in sig1(f,t) i This indicates that the direction of the sound source is θ. i The phase difference between the sound source propagating to the second position and the first position.
[0148] In some embodiments, the likelihood function may correspond to the probability of sound being emitted from the sound source direction to form the first signal and the second signal. As an example only, a likelihood function value of 0.8 when the sound source direction is θ = 30° can represent an 80% probability of sound being emitted from the sound source direction at θ = 30° to form the first signal and the second signal. In some embodiments, the response of the likelihood function corresponding to multiple sound source directions may be a cardiogram. For ease of description, the cardiogram corresponding to the likelihood function may be called a synthetic cardiogram (or an actual cardiogram).
[0149] In step 1020, the processing device 110 can determine the direction of the synthesized sound source based on the multiple parameters.
[0150] In some embodiments, to determine the direction of the synthesized sound source, the processing device 110 may determine the parameter with the largest value among the plurality of parameters. For example, the processing device 110 may calculate the direction of the plurality of sound sources θ = (1,2,…). n The likelihood functions LH1(f,t), H2(f,t), ..., H for each sound source direction in the equation are: n (f,t). Likelihood function LH i (f,t) can correspond to the direction θ from the sound source. i The probability of emitting sound to form the first signal and the second signal. The likelihood function with the largest value represents the probability that the sound is emitted from the direction of the corresponding sound source to form the first signal and the second signal. Therefore, the processing device 110 can determine that the direction of the sound source corresponding to the direction of the largest likelihood function is the direction of the synthesized sound source. For example, if θ i The likelihood function value is maximum at 30°, and the processing device 110 can determine that the direction of the synthesized sound source is 30°.
[0151] In some embodiments, the processing device 110 can determine the direction of the synthesized sound source based on the plurality of directional microphones. According to the above embodiments, the plurality of directional microphones can be respectively aligned with a preset sound source direction θ = (1,2,…,…). nEach microphone's response is a simulated cardiogram. The poles of the simulated cardiogram point to the corresponding sound source direction. Further, the responses of the likelihood functions corresponding to multiple sound source directions can also be cardiograms (called synthetic cardiograms). The poles of the synthetic cardiogram point to the direction of the synthesized sound source. The processing device 110 can compare the synthetic cardiogram with the aforementioned multiple simulated cardiograms to determine the simulated cardiogram whose pole direction is closest to the actual cardiogram. The sound source direction corresponding to the simulated cardiogram can be determined as the direction of the synthesized sound source.
[0152] In step 1030, the processing device 110 can determine the second coefficient based on the direction of the synthesized sound source and the location of the target speech.
[0153] In some embodiments, to determine the second coefficient, the processing device 110 can determine whether the target speech location is located in the direction of the synthesized sound source. For example, the target speech location may be located on the extension line of the dual microphones. That is, the sound source direction corresponding to the target speech location is θ = 0°. The processing device 110 can determine whether the direction of the synthesized sound source is 0°. If the direction of the synthesized sound source is 0°, then it can be determined that the target speech location is located in the direction of the synthesized sound source. In some embodiments, the processing device 110 can determine whether the likelihood function with the largest value is located in the set where the target sound source is dominant. If the likelihood function with the largest value is located in the set where the target sound source is dominant, the processing device 110 can determine that the target speech location is located in the direction of the synthesized sound source. In some embodiments, the set where the target sound source is dominant can be represented by the following formula:
[0154]
[0155] Where LH0(f,t) represents the likelihood function value when the target speech location is in the direction of the synthesized sound source. Based on set Processing device 110 can determine a time-frequency point (f,t) such that the likelihood function corresponding to that time-frequency point (f,t) reaches its maximum value maxLH when θ=0°. i (f,t). At this time, the signal corresponding to this time frequency point (e.g., the first signal or the third signal) can be a signal emitted from the direction of the target sound source.
[0156] It should be noted that the set of target sound sources dominated by the above formula (8) is... This is merely an example. In formula (8), the target speech location is located on the extension of the line connecting the two microphones (i.e., θ = 0°), therefore the time-frequency point (f,t) determined by the above method reaches its maximum value at θ = 0°. Optionally or additionally, the target speech location may not be located on the extension of the line connecting the two microphones (i.e., θ ≠ 0°). For example, the angle between the target speech location and the extension of the line connecting the two microphones is 30 degrees. In this case, LH0(f,t) in formula (8) can be the likelihood function value at θ = 30°. That is, the time-frequency point (f,t) obtained from the set of dominant target sound sources should make the likelihood function reach its maximum value at θ = 30°.
[0157] In some embodiments, in response to the target speech location being located in the direction of the synthesized sound source, the processing device 110 can set the second coefficient to a first value (e.g., 1). In response to the target speech location not being located in the direction of the synthesized sound source, the processing device 110 can set the second coefficient to a second value (e.g., 0). In some embodiments, the processing device 110 can process the corresponding first signal or the third signal obtained after ANF filtering based on the second coefficient. For example, taking the first signal as an example, the second coefficient can be used as a weight for the first signal. For example, when the second coefficient is 1, it can indicate that the first signal is retained. Conversely, when the target speech location is not located in the direction of the synthesized sound source, the first signal corresponding to the direction of the synthesized sound source can be considered a noise signal. The processing device 110 can set the second coefficient to 0. Thus, the processing device 110 can filter out or reduce the noise signal corresponding to the target speech location not being located in the direction of the synthesized sound source based on the second coefficient.
[0158] In some embodiments, the second coefficients may form a masking matrix for filtering the input signal (e.g., a first signal or a third signal). For example, the masking matrix may be as shown in the following formula:
[0159]
[0160] The masking matrix M is a binary matrix that can directly remove input signals identified as noise. Therefore, processing speech signals based on the masking matrix M may cause problems such as spectral leakage and speech discontinuity. In some embodiments, the processing device 110 can smooth the second coefficient based on a smoothing factor. For example, based on the smoothing factor, the processing device 110 can smooth the second coefficient in the time domain. The time-domain smoothing process can be illustrated by the following formula:
[0161]
[0162] Here, α represents the smoothing factor, M(f,t-1) represents the masking matrix corresponding to the previous frame, and M(f,t) represents the masking matrix corresponding to the current frame. The smoothing factor α can be used to weight the masking matrices of the previous frame and the current frame to obtain a smoothed masking matrix corresponding to the current frame.
[0163] In some embodiments, the processing device 110 may also smooth the second coefficient in the frequency domain. For example, the processing device 110 may use a sliding Hamming window to smooth the second coefficient.
[0164] In some embodiments, the processing device 110 may determine the second coefficient based on the angle between the target speech location and the direction of the synthesized sound source using a regression function. For example, for each time-frequency point, the processing device 110 may calculate the angle relative to the plurality of sound source directions θ = (θ1, θ2, ..., θ...). n The likelihood functions LH1(f,t), LH2(f,t), ..., LH2(f,t), corresponding to each sound source direction in the equation are given. n (f,t). And determine the direction of the sound source corresponding to the largest likelihood function as the direction of the synthesized sound source. For example, if θ i The likelihood function value is highest at 30°, allowing the processing device 110 to determine the direction of the synthesized sound source as 30°. Further, the processing device 110 can determine the second coefficient using a regression function based on the angle between the target sound source direction and the synthesized sound source direction. For example, the processing device 110 can construct a regression function between the angle and the second coefficient. In some embodiments, the regression function may include a smooth regression function, such as a linear regression function. As an example only, the value of the regression function may decrease as the angle between the target sound source direction and the synthesized sound source direction increases. Thus, when the input signal is processed with the second coefficient as a weight, the input signal with a large angle between the target sound source direction and the synthesized sound source direction can be weakened or removed, thereby achieving noise removal.
[0165] It should be noted that the above description of method 1000 is for convenience only and should not be construed as limiting this specification to the scope of the embodiments described. It is understood that those skilled in the art, after understanding the principle of this method, can arbitrarily combine the various steps, or add or delete any steps, without departing from this principle.
[0166] According to some embodiments of this specification, the processing device 110 can acquire a target speech signal using dual microphones and filter the target speech signal based on a dual-microphone filtering algorithm. For example, when the angle difference between the noise signal and the effective signal is large (i.e., the angle difference between the direction of the noise source and the direction of the target sound source is large), the processing device 110 can filter based on a first coefficient to remove the noise signal. When the angle difference between the noise signal and the effective signal is small, the processing device 110 can filter based on a second coefficient. In this way, the processing device 110 can substantially filter out the noise signal in the target speech signal. In some embodiments, the first output speech signal acquired after the above dual-microphone filtering process may include residual noise. For example, in some frequency sub-bands (e.g., mid-high frequency sub-bands), the first output speech signal may include continuous noise on the amplitude spectrum. Therefore, in some embodiments, the processing device 110 can also perform post-filtering on the first output speech signal based on a single-microphone filtering algorithm.
[0167] Figure 11 This is an exemplary flowchart of a single-microphone filtering method according to some embodiments of this specification. In some embodiments, method 1100 may be executed by processing device 110, processing engine 112, or processor 220. For example, method 1100 may be stored in a storage device (e.g., storage device 140 or storage unit of processing device 110) as a program or instructions, and executed by processing device 110, processing engine 112, processor 220, or... Figure 4 When the module shown executes a program or instructions, it can implement method 1100. In some embodiments, method 1100 may be accomplished using one or more additional operations / steps not described below, and / or not through one or more operations / steps discussed below. Additionally, as Figure 11 The order of operations / steps shown is not restrictive. For example... Figure 11 As shown, method 1100 may include:
[0168] In step 1110, the processing device 110 (e.g., processing module 420) can determine at least one target subband signal in the first output speech signal.
[0169] In some embodiments, the processing device 110 may determine the at least one target sub-band signal based on the signal-to-noise ratio (SNR) of each sub-band signal in the first output speech signal. In some embodiments, the processing device 110 may acquire multiple sub-band signals based on the first output speech signal. For example, the processing device 110 may divide the first output speech signal into sub-bands based on signal frequency bands to acquire multiple sub-band signals. As an example only, the processing device 110 may divide the first output speech signal into sub-bands according to low-frequency, mid-frequency, or high-frequency band categories, or it may divide the first output speech signal into sub-bands according to a specific bandwidth (e.g., each 2kHz is a frequency band). As another example, the processing device 110 may divide the sub-bands based on the signal frequency point of the first output speech signal. The signal frequency point may refer to the value after the decimal point in the frequency value of the signal; for example, if the frequency value of the signal is 72.810, then the signal frequency point of the signal is 810. Subband division based on signal frequency points can be achieved by dividing the signal into subbands according to a specific signal frequency width. For example, signal frequencies 810–830 can be used as one subband, and signal frequencies 600–620 can be used as another subband. In some embodiments, the processing device 110 can obtain multiple subband signals by filtering, or it can use other algorithms or devices to perform subband division and obtain multiple subband signals; no limitation is made here.
[0170] Further, the processing device 110 can calculate the signal-to-noise ratio (SNR) of each sub-band signal. The signal-to-noise ratio (SNR) can refer to the ratio of speech signal energy to noise signal energy. Signal energy can be signal power, other energy data obtained based on signal power, etc. In some embodiments, a higher SNR indicates less noise in the speech signal. In some embodiments, the SNR of a sub-band signal can be the ratio of the energy of the clean speech signal (i.e., the effective signal) to the noise signal energy in the sub-band signal, or it can be the ratio of the energy of the noisy sub-band signal to the noise signal energy. In some embodiments, the processing device 110 can calculate the SNR of each sub-band signal using an SNR estimation algorithm. For example, for each sub-band signal, the processing device 110 can calculate the noise signal value in the sub-band signal based on a noise estimation algorithm. Exemplary noise estimation algorithms may include minimum tracking algorithms, time recursive averaging algorithms, etc., or combinations thereof. Further, the processing device 110 can calculate the SNR based on the original sub-band signal and the noise signal value. In some embodiments, the processing device 110 may use a trained signal-to-noise ratio (SNR) estimation model to calculate the SNR of each sub-band signal. Exemplary SNR estimation models may include, but are not limited to, any algorithm or model capable of feature extraction and / or classification, such as Multi-Layer Perception (MLP), Decision Tree (DT), Deep Neural Network (DNN), Support Vector Machine (SVM), and K-Nearest Neighbor (KNN). In some embodiments, the SNR estimation model can be obtained by training an initial model using training samples. Training samples may include speech signal samples (e.g., at least one historical speech signal, each containing noise), and label values for the speech signal samples (e.g., the SNR of historical speech signal v1 is 0.5, and the SNR of historical speech signal v2 is 0.6). The model processes the speech signal samples to obtain the predicted SNR. A loss function is constructed based on the predicted SNR and the corresponding label values of the training samples. The model parameters are adjusted based on the loss function to reduce the difference between the predicted target SNR and the label values. For example, model parameters can be updated or adjusted using methods such as gradient descent. This iterative training process is repeated multiple times. Training ends when the trained model meets preset conditions, yielding the trained signal-to-noise ratio (SNR) estimation model. These preset conditions could include the loss function result converging or falling below a preset threshold.
[0171] Further, the processing device 110 can determine the target sub-band signal based on the signal-to-noise ratio (SNR) of each of the sub-band signals. In some embodiments, the processing device 110 can determine the target sub-band signal based on a SNR threshold. For example, for each sub-band signal, the processing device 110 can determine whether the SNR of the sub-band signal is less than the SNR threshold. In response to the SNR of the sub-band signal being less than the SNR threshold, the processing device 110 can determine that the sub-band signal is the target sub-band signal.
[0172] In some embodiments, the processing device 110 may also determine the at least one target sub-band signal based on a preset sub-band range. For example, the preset sub-band range may be a preset frequency range. The preset frequency range may be determined based on empirical values. The empirical values may be empirical values obtained during the speech analysis processing. As an example only, if the speech analysis processing finds that speech signals in the frequency range of 3000-4000Hz typically contain a large proportion of noise, then the preset frequency range may at least include 3000-4000Hz.
[0173] In step 1120, the processing device 110 (e.g., generation module 430) can process the at least one target sub-band signal based on a single-microphone filtering algorithm to obtain a second output speech signal.
[0174] In some embodiments, the processing device 110 can process the at least one target sub-band signal based on a single-microphone filtering algorithm to filter out noise in the at least one target sub-band signal and obtain a noise-reduced second output speech signal. Exemplary single-microphone filtering algorithms may include spectral subtraction, Wiener filtering algorithm, minimum-controlled recursive averaging algorithm, speech generation model algorithm, or combinations thereof.
[0175] According to the above embodiments, the processing device 110 can further process the first output speech signal obtained by the dual-microphone filtering algorithm according to the single-microphone filtering algorithm. For example, the processing device 110 can filter a portion of the sub-band signal (e.g., a signal of a specific frequency) in the first output speech signal according to the single-microphone filtering algorithm, thereby reducing or filtering out noise signals in the first output speech signal, and realizing the correction and / or supplementation of the dual-microphone filtering algorithm.
[0176] It should be noted that the above description of method 1100 is for convenience only and should not be construed as limiting this specification to the scope of the embodiments described. It is understood that those skilled in the art, after understanding the principle of this method, can arbitrarily combine the various steps, or add or delete any steps, without departing from this principle. For example, step 1110 can be omitted. The processing device 110 can not only perform filtering processing on the target sub-band signal, but can also directly perform filtering processing on the entire first output speech signal. For another example, the processing device 110 can automatically detect noise signals in the first output speech signal based on an automatic noise detection algorithm, and filter the detected noise signals using a single-microphone filtering algorithm.
[0177] Figure 12 This is an exemplary flowchart of a single-microphone filtering method according to some embodiments of this specification. In some embodiments, method 1200 may be executed by processing device 110, processing engine 112, or processor 220. For example, method 1200 may be stored in a storage device (e.g., storage device 140 or storage unit of processing device 110) as a program or instructions, and executed by processing device 110, processing engine 112, processor 220, or... Figure 4 When the module shown executes a program or instructions, it can implement method 1200. In some embodiments, method 1200 may be accomplished using one or more additional operations / steps not described below, and / or not through one or more operations / steps discussed below. Additionally, as Figure 12 The order of operations / steps shown is not restrictive. For example... Figure 12 As shown, method 1200 may include:
[0178] In step 1210, the processing device 110 (e.g., processing module 420) can process the first signal and / or the second signal based on a single-microphone filtering algorithm to determine the third coefficient.
[0179] In some embodiments, the processing device 110 may determine the third coefficient based on either the first signal or the second signal. For example, the processing device 110 may determine the third coefficient based on the first signal, or it may determine the third coefficient based on the second signal. In some embodiments, the processing device 110 may determine the third coefficient based on both the first and second signals. For example, the processing device 110 may determine a first value of the third coefficient based on the first signal, determine a second value of the third coefficient based on the second signal, and then determine the third coefficient based on both the first and second values (e.g., by averaging, weighted summation, etc.).
[0180] Taking the first coefficient as an example, in some embodiments, the processing device 110 can process the first signal based on a single-microphone filtering algorithm. Exemplary single-microphone filtering algorithms may include spectral subtraction, Wiener filtering, recursive averaging with minimum control, speech generation model algorithms, or combinations thereof. As an example only, the processing device 110 can obtain the noise signal and the effective signal in the first signal based on the single-microphone filtering algorithm, and determine the signal-to-noise ratio (SNR) corresponding to the first signal based on at least two of the noise signal, the effective signal, and the first signal. The SNR corresponding to the first signal may include a priori SNR, posterior SNR, etc. The a priori SNR may be the energy ratio of the effective signal to the noise signal. The posterior SNR may be the energy ratio of the effective signal to the first signal. Further, the processing device 110 can determine the third coefficient based on the SNR corresponding to the first signal. For example, the processing device 110 can determine the gain coefficient corresponding to the single-microphone filtering algorithm based on the a priori SNR and / or the posterior SNR, and determine the third coefficient based on the gain coefficient. For example, the processing device 110 can directly use the gain coefficient as the third coefficient. For example, processing device 110 can determine the mapping relationship between the gain coefficient and the third coefficient, and determine the third coefficient based on the mapping relationship. Here, the gain coefficient can refer to the transfer function in a single-microphone filtering algorithm. The transfer function can filter a noisy speech signal to obtain a valid signal. For example, the transfer function can be in matrix form; by multiplying the transfer function by the noisy speech signal, the noise signal in the speech signal can be filtered out. Accordingly, the third coefficient can be used to remove noise from the speech signal.
[0181] Optionally or additionally, the processing device 110 can also obtain a smoothed signal-to-noise ratio (SNR) by weighting the prior and posterior SNRs using a smoothing factor based on a logistic regression algorithm (e.g., the sigmoid function). The gain coefficient corresponding to the single-microphone filtering algorithm is then determined as a third coefficient based on the smoothed SNR. Therefore, the third coefficient can have better smoothness, thereby avoiding strong musical noise when using the single-microphone filtering algorithm.
[0182] In step 1220, the processing device 110 (e.g., generation module 430) can process the first output speech signal based on the third coefficient to obtain the third output speech signal.
[0183] In some embodiments, the processing device 110 may multiply the third coefficient by the first output speech signal to obtain the third output speech signal. For example, according to step 1210, the third coefficient may be a gain coefficient obtained based on a single-microphone filtering algorithm. By multiplying the gain coefficient by the first output speech signal, noise signals in the first output speech signal can be filtered out.
[0184] It should be noted that the above description of method 1200 is for convenience only and should not be construed as limiting this specification to the scope of the embodiments described. It is understood that those skilled in the art, after understanding the principle of this method, can arbitrarily combine the various steps, or add or delete any steps, without departing from this principle.
[0185] Figure 13 This is an exemplary flowchart of a speech enhancement method according to some embodiments of this specification. In some embodiments, method 1300 may be executed by processing device 110, processing engine 112, or processor 220. For example, method 1300 may be stored in a storage device (e.g., storage device 140 or storage unit of processing device 110) as a program or instructions, and executed by processing device 110, processing engine 112, processor 220, or... Figure 4 When the module shown executes a program or instructions, it can implement method 1300. In some embodiments, method 1300 may be accomplished using one or more additional operations / steps not described below, and / or not through one or more operations / steps discussed below. Additionally, as Figure 13 The order of operations / steps shown is not restrictive. For example... Figure 13 As shown, method 1300 may include:
[0186] In step 1310, the processing device 110 can acquire the first signal and the second signal of the target speech.
[0187] In step 1320, the processing device 110 can process the first signal and the second signal based on the target speech location, the first location, and the second location to determine the first coefficient.
[0188] In step 1330, the processing device 110 can determine multiple parameters related to the directions of multiple sound sources based on the first signal and the second signal.
[0189] In step 1340, the processing device 110 can determine the second coefficient based on the plurality of parameters and the target speech location.
[0190] In some embodiments, it can be achieved by executing Figure 5 Steps 510-540 described herein are used to execute steps 1310-1340, and will not be repeated here.
[0191] In step 1350, the processing device 110 (e.g., processing module 420) can determine a fourth coefficient based on the energy difference between the first signal and the second signal.
[0192] In some embodiments, to determine the fourth coefficient, the processing device 110 can obtain the noise power spectral density based on silent intervals in the first and second signals. The silent interval can be a speech signal interval where no valid signal exists (i.e., the target sound source does not emit speech). Within the silent interval, since there is no speech from the target sound source, the first and second signals acquired by the two microphones contain only noise components. In some embodiments, the processing device 110 can determine the silent intervals in the first and second signals based on a Voice Activity Detection (VAD) algorithm. In some embodiments, the processing device 110 can determine one or more speech intervals in the first and second signals as silent intervals, respectively. For example, for each of the first and second signals, the processing device 110 can directly define a speech interval (e.g., 200ms, 300ms, etc.) at the beginning of that signal as a silent interval. Further, the processing device 110 can obtain the noise power density spectrum based on the silent intervals. In some embodiments, when the noise signal source is far from the dual microphones, the noise signals received by the dual microphones can be considered similar or identical. Therefore, the processing device 110 can obtain the noise power spectral density based on any corresponding silent interval of the first signal or the second signal. In some embodiments, the processing device 110 can obtain the noise power spectral density based on a periodogram algorithm. Optionally or additionally, the processing device 110 can transform the first signal and / or the second signal to the frequency domain based on an FFT transform, thereby obtaining the noise power spectral density in the frequency domain based on a periodogram algorithm.
[0193] Further, the processing device 110 can obtain the energy difference based on the first power spectral density of the first signal, the second power spectral density of the second signal, and the noise power spectral density. In some embodiments, the processing device 110 can determine the first power spectral density of the first signal and the second power spectral density of the second signal based on a periodogram algorithm. In some embodiments, the processing device 110 can obtain the energy difference based on a Power Level Difference (PLD) algorithm. In the PLD algorithm, it can be assumed that the distance between the two microphones is relatively far, thus the energy difference between the effective signal in the first signal and the effective signal in the second signal is large, and the noise signals in the first signal and the second signal are the same or similar. Therefore, the energy difference between the first signal and the second signal can be expressed as a function related to the effective signal in the first signal.
[0194] Further, the processing device 110 can determine the fourth coefficient based on the energy difference and the noise power spectral density. In some embodiments, the processing device 110 can determine the gain coefficient based on the PLD algorithm and determine the gain coefficient as the fourth coefficient.
[0195] In step 1360, the processing device 110 (e.g., generation module 430) can process the first signal and / or the second signal based on the first coefficient, the second coefficient, and the fourth coefficient to obtain a fourth output speech signal with enhanced speech corresponding to the target speech.
[0196] In some embodiments, the processing device 110 may perform gain compensation on the first signal and / or the second signal based on a fourth coefficient to obtain an estimated effective signal. For example, the estimated effective signal may be the product of the fourth coefficient and the first signal and / or the second signal. In some embodiments, the processing device 110 may perform weighted processing on the output signal (e.g., the third signal, the estimated effective signal) obtained based on the first signal and / or the second signal based on the first coefficient, the second coefficient, and the fourth coefficient. For example, the processing device 110 may perform weighted processing on the third signal, the first signal, and the estimated effective signal based on the first coefficient, the second coefficient, and the fourth coefficient, respectively, and determine a fourth output speech signal based on the weighted signal. For example, the fourth output speech signal may be the average value of the weighted signals. As another example, the fourth output speech signal may be a larger value in the weighted signal.
[0197] In some embodiments, the processing device 110 may process the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain the first output speech signal with speech enhancement corresponding to the target speech, and then process the first output speech signal based on the fourth coefficient to obtain the fourth output speech signal.
[0198] It should be noted that the above description of method 1300 is for convenience only and should not be construed as limiting this specification to the scope of the embodiments described. It is understood that those skilled in the art, after understanding the principle of this method, can arbitrarily combine the various steps, or add or delete any steps, without departing from this principle. In some embodiments, the processing device 110 can also determine the fourth coefficient based on the power difference between the first signal and the second signal. In some embodiments, the processing device 110 can also determine the fourth coefficient based on the amplitude difference between the first signal and the second signal.
[0199] The beneficial effects that the embodiments of this specification may bring include, but are not limited to: (1) Processing the target speech signal based on the ANF algorithm results in less damage to the target speech signal, and when the angle difference between the effective signal and the noise signal is large, the noise signal can be effectively filtered; (2) Processing the target speech signal based on the probability distribution algorithm can effectively filter the noise signal near the target sound source when the angle difference between the effective signal and the noise signal is small; (3) Processing the target speech signal by combining dual-microphone filtering and single-microphone filtering can effectively filter out the residual noise after dual-microphone filtering.
[0200] The basic concepts have been described above. It is clear that the above disclosure is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, various modifications, improvements, and corrections may be made to this specification by those skilled in the art. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.
[0201] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.
[0202] Furthermore, those skilled in the art will understand that various aspects of this specification can be described and illustrated in several patentable ways, including any new and useful combinations of processes, machines, products, or substances, or any new and useful improvements thereof. Accordingly, various aspects of this specification can be implemented entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. All of the above hardware or software may be referred to as a “data block,” “module,” “engine,” “unit,” “component,” or “system.” Furthermore, various aspects of this specification may be represented as a computer product located on one or more computer-readable media, including computer-readable program code.
[0203] Furthermore, unless expressly stated in the claims, the order of elements and sequences, the use of numbers and letters, or other names in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on an existing server or mobile device.
[0204] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.
[0205] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples by terms such as "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical data used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, the numerical data should take into account specified significant digits and employ general methods of digit reservation. Although the numerical ranges and data used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such numerical values are set as precisely as feasible.
Claims
1. A speech enhancement method, characterized in that, The method includes: Acquire a first signal and a second signal of the target speech, wherein the target speech includes speech emitted by a target sound source, the first signal is a signal of the target speech acquired based on a first position, and the second signal is a signal of the target speech acquired based on a second position; Based on the target speech location, the first location, and the second location, a differential operation is performed on the first signal and the second signal to obtain a signal pointing in a first direction and a signal pointing in a second direction. The signal pointing in the first direction and the signal pointing in the second direction contain different proportions of effective signals. The signal pointing in the first direction is a signal pointing in the direction of the target sound source, and the signal pointing in the second direction is a signal pointing in the opposite direction to the target sound source. Based on the signal pointing in the first direction and the signal pointing in the second direction, a third signal corresponding to the effective signal is determined, and based on the third signal, a first coefficient is determined. Based on the first signal and the second signal, multiple parameters related to the directions of multiple sound sources are determined, each parameter corresponding to the probability of emitting sound from a sound source direction to form the first signal and the second signal; Based on the aforementioned parameters, the direction of the synthesized sound source is determined; and based on the direction of the synthesized sound source and the target speech location, a second coefficient is determined; and Based on the first coefficient and the second coefficient, process the first signal and / or the second signal to obtain the first output speech signal after speech enhancement corresponding to the target speech; The step of processing the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain the first output speech signal after speech enhancement corresponding to the target speech includes: The weight of the third signal in the first output speech signal is determined based on the first coefficient, and the weight of the first signal or the third signal in the first output speech signal is determined based on the second coefficient.
2. The method as described in claim 1, characterized in that, The determination of the third signal corresponding to the valid signal includes: An adaptive differential operation is performed on the signal pointing in the first direction and the signal pointing in the second direction to determine the fourth signal; and The low-frequency components in the fourth signal are enhanced to obtain the third signal.
3. The method as described in claim 2, characterized in that, The method further includes: The adaptive parameters of the adaptive differential operation are updated based on the fourth signal, the signal pointing in the first direction, and the signal pointing in the second direction.
4. The method as described in claim 1, characterized in that, The determination of multiple parameters related to the directions of multiple sound sources based on the first signal and the second signal includes: Based on each sound source direction, the first position, and the second position, differential operations are performed on the first signal and the second signal to determine parameters related to each sound source direction.
5. The method as described in claim 1, characterized in that, Determining the second coefficient based on the direction of the synthesized sound source and the location of the target speech includes: Determine whether the target speech location is located in the direction of the synthesized sound source; In response to the target speech location being located in the direction of the synthesized sound source, the second coefficient is set to a first value; and In response to the fact that the target speech location is not located in the direction of the synthesized sound source, the second coefficient is set to a second value.
6. The method as described in claim 1, characterized in that, Determining the second coefficient based on the direction of the synthesized sound source and the location of the target speech includes: The second coefficient is determined by a regression function based on the angle between the target speech location and the direction of the synthesized sound source.
7. The method as described in claim 1, characterized in that, Also includes: The second coefficient is smoothed based on a smoothing factor.
8. The method as described in claim 1, characterized in that, The method further includes performing at least one of the following operations on the first signal and the second signal: The first signal and the second signal are framed; Windowing smoothing is applied to the first signal and the second signal; and Convert the first signal and the second signal to the frequency domain.
9. The method as described in claim 1, characterized in that, The method further includes: Determine at least one target subband signal in the first output speech signal; and Based on a single-microphone filtering algorithm, the at least one target sub-band signal is processed to obtain a second output speech signal.
10. The method as described in claim 9, characterized in that, Determining at least one target sub-band signal in the first output speech signal includes: Based on the first output voice signal, multiple sub-band signals are obtained; Calculate the signal-to-noise ratio of each of the sub-band signals; and The target sub-band signal is determined based on the signal-to-noise ratio of each sub-band signal.
11. The method as described in claim 1, characterized in that, The method further includes: Based on a single-microphone filtering algorithm, the first signal and / or the second signal are processed to determine the third coefficient; and Based on the third coefficient, the first output speech signal is processed to obtain the third output speech signal.
12. The method as described in claim 8, characterized in that, The method further includes: The fourth coefficient is determined based on the energy difference between the first signal and the second signal; and Based on the first coefficient, the second coefficient, and the fourth coefficient, the first signal and / or the second signal are processed to obtain the fourth output speech signal after speech enhancement corresponding to the target speech.
13. The method as described in claim 12, characterized in that, The determination of the fourth coefficient based on the energy difference between the first signal and the second signal includes: Based on the silent regions in the first and second signals, the noise power spectral density is obtained; The energy difference is obtained based on the first power spectral density of the first signal, the second power spectral density of the second signal, and the noise power spectral density; and The fourth coefficient is determined based on the energy difference and the noise power spectral density.
14. A speech enhancement system, characterized in that, The system includes: At least one storage medium including a set of instructions; and At least one processor communicating with at least one storage medium, wherein, when executing the set of instructions, the at least one processor causes the system to: Acquire a first signal and a second signal of the target speech, wherein the target speech includes speech emitted by a target sound source, the first signal is a signal of the target speech acquired based on a first position, and the second signal is a signal of the target speech acquired based on a second position; Based on the target speech location, the first location, and the second location, a differential operation is performed on the first signal and the second signal to obtain a signal pointing in a first direction and a signal pointing in a second direction. The signal pointing in the first direction and the signal pointing in the second direction contain different proportions of effective signals. The signal pointing in the first direction is a signal pointing in the direction of the target sound source, and the signal pointing in the second direction is a signal pointing in the opposite direction to the target sound source. Based on the signal pointing in the first direction and the signal pointing in the second direction, a third signal corresponding to the valid signal is determined, and based on the third signal, a first coefficient is determined; Based on the first signal and the second signal, multiple parameters related to the directions of multiple sound sources are determined, each parameter corresponding to the probability of emitting sound from a sound source direction to form the first signal and the second signal; Based on the aforementioned parameters, the direction of the synthesized sound source is determined; and based on the direction of the synthesized sound source and the target speech location, a second coefficient is determined; and Based on the first coefficient and the second coefficient, process the first signal and / or the second signal to obtain the first output speech signal after speech enhancement corresponding to the target speech; The step of processing the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain the first output speech signal after speech enhancement corresponding to the target speech includes: The weight of the third signal in the first output speech signal is determined based on the first coefficient, and the weight of the first signal or the third signal in the first output speech signal is determined based on the second coefficient.
15. The system as described in claim 14, characterized in that, In order to determine the third signal corresponding to the valid signal, the at least one processor causes the system to: An adaptive differential operation is performed on the signal pointing in the first direction and the signal pointing in the second direction to determine the fourth signal; and The low-frequency components in the fourth signal are enhanced to obtain the third signal.
16. The system as described in claim 15, characterized in that, The at least one processor further enables the system to: The adaptive parameters of the adaptive differential operation are updated based on the fourth signal, the signal pointing in the first direction, and the signal pointing in the second direction.
17. The system as claimed in claim 14, characterized in that, In order to determine multiple parameters related to the directions of multiple sound sources based on the first signal and the second signal, the at least one processor enables the system to: Based on each sound source direction, the first position, and the second position, differential operations are performed on the first signal and the second signal to determine parameters related to each sound source direction.
18. The system as claimed in claim 14, characterized in that, In order to determine the second coefficient based on the direction of the synthesized sound source and the location of the target speech, the at least one processor enables the system to: Determine whether the target speech location is located in the direction of the synthesized sound source; In response to the target speech location being located in the direction of the synthesized sound source, the second coefficient is set to a first value; and In response to the fact that the target speech location is not located in the direction of the synthesized sound source, the second coefficient is set to a second value.
19. The system as claimed in claim 14, characterized in that, In order to determine the second coefficient based on the direction of the synthesized sound source and the location of the target speech, the at least one processor enables the system to: The second coefficient is determined by a regression function based on the angle between the target speech location and the direction of the synthesized sound source.
20. The system as claimed in claim 14, characterized in that, The at least one processor further enables the system to: The second coefficient is smoothed based on a smoothing factor.
21. The system as described in claim 14, characterized in that, The at least one processor further causes the system to perform at least one of the following operations on the first signal and the second signal: The first signal and the second signal are framed; Windowing smoothing is applied to the first signal and the second signal; and Convert the first signal and the second signal to the frequency domain.
22. The system as described in claim 14, characterized in that, The at least one processor further enables the system to: Determine at least one target subband signal in the first output speech signal; and Based on a single-microphone filtering algorithm, the at least one target sub-band signal is processed to obtain a second output speech signal.
23. The system as described in claim 22, characterized in that, In order to determine at least one target subband signal in the first output speech signal, the at least one processor enables the system to: Based on the first output voice signal, multiple sub-band signals are obtained; Calculate the signal-to-noise ratio of each of the sub-band signals; and The target sub-band signal is determined based on the signal-to-noise ratio of each sub-band signal.
24. The system as claimed in claim 14, characterized in that, The at least one processor further enables the system to: Based on a single-microphone filtering algorithm, the first signal and / or the second signal are processed to determine the third coefficient; and Based on the third coefficient, the first output speech signal is processed to obtain the third output speech signal.
25. The system as claimed in claim 21, characterized in that, The at least one processor further enables the system to: The fourth coefficient is determined based on the energy difference between the first signal and the second signal; and Based on the first coefficient, the second coefficient, and the fourth coefficient, the first signal and / or the second signal are processed to obtain the fourth output speech signal after speech enhancement corresponding to the target speech.
26. The system as described in claim 25, characterized in that, In order to determine a fourth coefficient based on the energy difference between the first signal and the second signal, the at least one processor enables the system to: Based on the silent regions in the first and second signals, the noise power spectral density is obtained; The energy difference is obtained based on the first power spectral density of the first signal, the second power spectral density of the second signal, and the noise power spectral density; and The fourth coefficient is determined based on the energy difference and the noise power spectral density.
27. A speech enhancement system, characterized in that, include: An acquisition module is used to acquire a first signal and a second signal of target speech, wherein the target speech includes speech emitted by a target sound source, the first signal is a signal of the target speech acquired based on a first position, and the second signal is a signal of the target speech acquired based on a second position; Processing module, used for Based on the target speech location, the first location, and the second location, a differential operation is performed on the first signal and the second signal to obtain a signal pointing in a first direction and a signal pointing in a second direction. The signal pointing in the first direction and the signal pointing in the second direction contain different proportions of effective signals. The signal pointing in the first direction is a signal pointing in the direction of the target sound source, and the signal pointing in the second direction is a signal pointing in the opposite direction to the target sound source. Based on the signal pointing in the first direction and the signal pointing in the second direction, a third signal corresponding to the valid signal is determined, and based on the third signal, a first coefficient is determined; Based on the first signal and the second signal, determine multiple parameters related to the directions of multiple sound sources, each parameter corresponding to the probability of emitting sound from a sound source direction to form the first signal and the second signal; and Based on the aforementioned parameters, the direction of the synthesized sound source is determined; and based on the direction of the synthesized sound source and the target speech location, a second coefficient is determined; and The generation module is used to process the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain a first output speech signal after speech enhancement corresponding to the target speech; The step of processing the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain the first output speech signal after speech enhancement corresponding to the target speech includes: The weight of the third signal in the first output speech signal is determined based on the first coefficient, and the weight of the first signal or the third signal in the first output speech signal is determined based on the second coefficient.
28. A speech enhancement method, characterized in that, The method includes: Acquire a first signal and a second signal of the target speech, wherein the target speech includes speech emitted by a target sound source, the first signal is a signal of the target speech acquired based on a first position, and the second signal is a signal of the target speech acquired based on a second position; Based on the target speech location, the first location, and the second location, a differential operation is performed on the first signal and the second signal to obtain a signal pointing in a first direction and a signal pointing in a second direction. The signal pointing in the first direction and the signal pointing in the second direction contain different proportions of effective signal. The signal pointing in the first direction is a signal pointing in the direction of the target sound source, and the signal pointing in the second direction is a signal pointing in the opposite direction to the target sound source. Based on the signal pointing in the first direction and the signal pointing in the second direction, a third signal corresponding to the effective signal is determined, and the estimated signal-to-noise ratio of the target speech is determined; and based on the estimated signal-to-noise ratio, a first coefficient is determined. Based on the first signal and the second signal, multiple parameters related to the directions of multiple sound sources are determined, each parameter corresponding to the probability of emitting sound from a sound source direction to form the first signal and the second signal; Based on the aforementioned parameters, the direction of the synthesized sound source is determined; and based on the direction of the synthesized sound source and the target speech location, a second coefficient is determined; and Based on the first coefficient and the second coefficient, process the first signal and / or the second signal to obtain the first output speech signal after speech enhancement corresponding to the target speech; The step of processing the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain the first output speech signal after speech enhancement corresponding to the target speech includes: The weight of the third signal in the first output speech signal is determined based on the first coefficient, and the weight of the first signal or the third signal in the first output speech signal is determined based on the second coefficient.
29. The method as described in claim 28, characterized in that, The determination of multiple parameters related to the directions of multiple sound sources based on the first signal and the second signal includes: Based on each sound source direction, the first position, and the second position, differential operations are performed on the first signal and the second signal to determine parameters related to each sound source direction.
30. The method as described in claim 28, characterized in that, Determining the second coefficient based on the direction of the synthesized sound source and the location of the target speech includes: Determine whether the target speech location is located in the direction of the synthesized sound source; In response to the target speech location being located in the direction of the synthesized sound source, the second coefficient is set to a first value; and In response to the fact that the target speech location is not located in the direction of the synthesized sound source, the second coefficient is set to a second value.
31. The method as described in claim 28, characterized in that, Determining the second coefficient based on the direction of the synthesized sound source and the location of the target speech includes: The second coefficient is determined by a regression function based on the angle between the target speech location and the direction of the synthesized sound source.
32. The method as described in claim 28, characterized in that, Also includes: The second coefficient is smoothed based on a smoothing factor.
33. The method as described in claim 28, characterized in that, The method further includes performing at least one of the following operations on the first signal and the second signal: The first signal and the second signal are framed; Windowing smoothing is applied to the first signal and the second signal; and Convert the first signal and the second signal to the frequency domain.
34. The method as described in claim 28, characterized in that, The method further includes: Determine at least one target subband signal in the first output speech signal; and Based on a single-microphone filtering algorithm, the at least one target sub-band signal is processed to obtain a second output speech signal.
35. The method as described in claim 34, characterized in that, Determining at least one target sub-band signal in the first output speech signal includes: Based on the first output voice signal, multiple sub-band signals are obtained; Calculate the signal-to-noise ratio of each of the sub-band signals; and The target sub-band signal is determined based on the signal-to-noise ratio of each sub-band signal.
36. The method as described in claim 28, characterized in that, The method further includes: Based on a single-microphone filtering algorithm, the first signal and / or the second signal are processed to determine the third coefficient; and Based on the third coefficient, the first output speech signal is processed to obtain the third output speech signal.
37. The method as described in claim 33, characterized in that, The method further includes: The fourth coefficient is determined based on the energy difference between the first signal and the second signal; and Based on the first coefficient, the second coefficient, and the fourth coefficient, the first signal and / or the second signal are processed to obtain the fourth output speech signal after speech enhancement corresponding to the target speech.
38. The method as described in claim 37, characterized in that, The determination of the fourth coefficient based on the energy difference between the first signal and the second signal includes: Based on the silent regions in the first and second signals, the noise power spectral density is obtained; The energy difference is obtained based on the first power spectral density of the first signal, the second power spectral density of the second signal, and the noise power spectral density; and The fourth coefficient is determined based on the energy difference and the noise power spectral density.
39. A speech enhancement system, characterized in that, The system includes: At least one storage medium including a set of instructions; and At least one processor communicating with at least one storage medium, wherein, when executing the set of instructions, the at least one processor causes the system to: Acquire a first signal and a second signal of the target speech, wherein the target speech includes speech emitted by a target sound source, the first signal is a signal of the target speech acquired based on a first position, and the second signal is a signal of the target speech acquired based on a second position; Based on the target speech location, the first location, and the second location, a differential operation is performed on the first signal and the second signal to obtain a signal pointing in a first direction and a signal pointing in a second direction. The signal pointing in the first direction and the signal pointing in the second direction contain different proportions of effective signal. The signal pointing in the first direction is a signal pointing in the direction of the target sound source, and the signal pointing in the second direction is a signal pointing in the opposite direction to the target sound source. Based on the signal pointing in the first direction and the signal pointing in the second direction, a third signal corresponding to the effective signal is determined, and the estimated signal-to-noise ratio of the target speech is determined; and based on the estimated signal-to-noise ratio, a first coefficient is determined. Based on the first signal and the second signal, multiple parameters related to the directions of multiple sound sources are determined, each parameter corresponding to the probability of emitting sound from a sound source direction to form the first signal and the second signal; Based on the aforementioned parameters, the direction of the synthesized sound source is determined; and based on the direction of the synthesized sound source and the target speech location, a second coefficient is determined; and Based on the first coefficient and the second coefficient, process the first signal and / or the second signal to obtain the first output speech signal after speech enhancement corresponding to the target speech; The step of processing the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain the first output speech signal after speech enhancement corresponding to the target speech includes: The weight of the third signal in the first output speech signal is determined based on the first coefficient, and the weight of the first signal or the third signal in the first output speech signal is determined based on the second coefficient.
40. The system as described in claim 39, characterized in that, In order to determine multiple parameters related to the directions of multiple sound sources based on the first signal and the second signal, the at least one processor enables the system to: Based on each sound source direction, the first position, and the second position, differential operations are performed on the first signal and the second signal to determine parameters related to each sound source direction.
41. The system as described in claim 39, characterized in that, In order to determine the second coefficient based on the direction of the synthesized sound source and the location of the target speech, the at least one processor enables the system to: Determine whether the target speech location is located in the direction of the synthesized sound source; In response to the target speech location being located in the direction of the synthesized sound source, the second coefficient is set to a first value; and In response to the fact that the target speech location is not located in the direction of the synthesized sound source, the second coefficient is set to a second value.
42. The system as described in claim 39, characterized in that, In order to determine the second coefficient based on the direction of the synthesized sound source and the location of the target speech, the at least one processor enables the system to: The second coefficient is determined by a regression function based on the angle between the target speech location and the direction of the synthesized sound source.
43. The system as described in claim 39, characterized in that, The at least one processor further enables the system to: The second coefficient is smoothed based on a smoothing factor.
44. The system as described in claim 39, characterized in that, The at least one processor further causes the system to perform at least one of the following operations on the first signal and the second signal: The first signal and the second signal are framed; Windowing smoothing is applied to the first signal and the second signal; and Convert the first signal and the second signal to the frequency domain.
45. The system as described in claim 39, characterized in that, The at least one processor further enables the system to: Determine at least one target subband signal in the first output speech signal; and Based on a single-microphone filtering algorithm, the at least one target sub-band signal is processed to obtain a second output speech signal.
46. The system as described in claim 45, characterized in that, In order to determine at least one target subband signal in the first output speech signal, the at least one processor enables the system to: Based on the first output voice signal, multiple sub-band signals are obtained; Calculate the signal-to-noise ratio of each of the sub-band signals; and The target sub-band signal is determined based on the signal-to-noise ratio of each sub-band signal.
47. The system as described in claim 39, characterized in that, The at least one processor further enables the system to: Based on a single-microphone filtering algorithm, the first signal and / or the second signal are processed to determine the third coefficient; and Based on the third coefficient, the first output speech signal is processed to obtain the third output speech signal.
48. The system as described in claim 44, characterized in that, The at least one processor further enables the system to: The fourth coefficient is determined based on the energy difference between the first signal and the second signal; and Based on the first coefficient, the second coefficient, and the fourth coefficient, the first signal and / or the second signal are processed to obtain the fourth output speech signal after speech enhancement corresponding to the target speech.
49. The system as described in claim 48, characterized in that, In order to determine a fourth coefficient based on the energy difference between the first signal and the second signal, the at least one processor enables the system to: Based on the silent regions in the first and second signals, the noise power spectral density is obtained; The energy difference is obtained based on the first power spectral density of the first signal, the second power spectral density of the second signal, and the noise power spectral density; and The fourth coefficient is determined based on the energy difference and the noise power spectral density.
50. A speech enhancement system, characterized in that, include: An acquisition module is used to acquire a first signal and a second signal of target speech, wherein the target speech includes speech emitted by a target sound source, the first signal is a signal of the target speech acquired based on a first position, and the second signal is a signal of the target speech acquired based on a second position; Processing module, used for Based on the target speech location, the first location, and the second location, a differential operation is performed on the first signal and the second signal to obtain a signal pointing in a first direction and a signal pointing in a second direction. The signal pointing in the first direction and the signal pointing in the second direction contain different proportions of effective signal. The signal pointing in the first direction is a signal pointing in the direction of the target sound source, and the signal pointing in the second direction is a signal pointing in the opposite direction to the target sound source. Based on the signal pointing in the first direction and the signal pointing in the second direction, a third signal corresponding to the effective signal is determined, and the estimated signal-to-noise ratio of the target speech is determined; and based on the estimated signal-to-noise ratio, a first coefficient is determined. Based on the first signal and the second signal, determine multiple parameters related to the directions of multiple sound sources, each parameter corresponding to the probability of emitting sound from a sound source direction to form the first signal and the second signal; and Based on the aforementioned parameters, the direction of the synthesized sound source is determined; and based on the direction of the synthesized sound source and the target speech location, a second coefficient is determined; and The generation module is used to process the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain a first output speech signal after speech enhancement corresponding to the target speech; The step of processing the first signal and / or the second signal based on the first coefficient and the second coefficient to obtain the first output speech signal after speech enhancement corresponding to the target speech includes: The weight of the third signal in the first output speech signal is determined based on the first coefficient, and the weight of the first signal or the third signal in the first output speech signal is determined based on the second coefficient.
51. A non-transitory computer-readable medium, characterized in that, Includes executable instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any one of claims 1-13 and 28-38.
Citation Information
Patent Citations
Method and system for eliminating noise
CN101510426A
Noise reduction method and device, electronic equipment and readable storage medium
CN111063366A