A multi-speaker speech separation method and related apparatus

By using a microphone array and spatial separation model to separate the speech of multiple speakers, the problem of insufficient ability of voiceprint information separation methods to distinguish same-sex speakers is solved, and stable speech separation results are achieved.

CN119580759BActive Publication Date: 2026-02-06IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411836396.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2026-02-06
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

Existing multi-speaker speech separation methods based on voiceprint information have weak ability to distinguish same-sex speakers and cannot meet application requirements.

Method used

A microphone array consisting of two microphones placed at different locations is used to collect speech signals from multiple speakers. By determining the phase difference of the microphones and performing fixed beamforming processing, the speaker's speech time-frequency mask is obtained, and speech separation is performed using a spatial separation model and a Gaussian mixture model.

Benefits of technology

It achieves effective speaker differentiation, with good separation effect and stability, and is not affected by the speaker's gender, thus meeting application requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580759B_ABST
    Figure CN119580759B_ABST
Patent Text Reader

Abstract

The application discloses a multi-speaker voice separation method and a related device, and relates to the technical field of voice processing. The method comprises the following steps: acquiring multi-speaker voice signals collected by two microphone arrays for a first speaker and a second speaker located at different positions; determining the phase difference of the two microphones according to the multi-speaker voice signals, and performing fixed beamforming processing on the multi-speaker voice signals for a first area and fixed beamforming processing for a second area, wherein the first area is the area where the first speaker is located, and the second area is the area where the second speaker is located; determining the voice time-frequency mask of the two speakers at different positions according to the phase difference and the beamforming signals corresponding to the two areas respectively; and separating the voice signals of the first speaker and the second speaker from the signals of any channel of the multi-speaker voice signals according to the determined voice time-frequency mask. The multi-speaker voice separation method disclosed by the application has good separation effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, in particular to a multi-speaker speech separation method and related device. BACKGROUND

[0002] In many scenarios, multi-speaker speech is generated. In order to facilitate subsequent data processing and analysis, multi-speaker speech separation is usually required, that is, the speech of different speakers is separated from the multi-speaker speech.

[0003] Taking an enterprise management scenario as an example, enterprise employees usually need to wear a work card. With the rapid development of science and technology, intelligent electronic work cards, as a new type of management tool, have entered the management field of more and more enterprises and are popularized by more and more sales and service companies. The intelligent electronic work card most commonly requires a multi-speaker speech separation function, that is, a function of separating the speech of a person wearing a work card from the speech of a customer from recorded multi-speaker speech.

[0004] Current multi-speaker speech separation methods are mostly multi-speaker speech separation methods based on voiceprint information, that is, using voiceprint information to separate the speech of different speakers from multi-speaker speech. However, the voiceprint information has weak distinguishing ability for speakers of the same sex, which leads to certain application limitations of the multi-speaker speech separation method based on voiceprint information, and cannot meet the application requirements. SUMMARY

[0005] Therefore, the present application provides a multi-speaker speech separation method and related device to solve the problem that the current multi-speaker speech separation method based on voiceprint information has certain application limitations and cannot meet the application requirements. The technical solutions are as follows:

[0006] The first aspect of the present application provides a multi-speaker speech separation method, comprising:

[0007] obtaining a multi-speaker speech signal collected by a microphone array composed of two microphones arranged at different positions for a first speaker and a second speaker located at different positions;

[0008] determining a phase difference of the two microphones according to the multi-speaker speech signal, and respectively performing fixed beamforming processing for a first region and fixed beamforming processing for a second region on the multi-speaker speech signal to obtain beamforming signals corresponding to the first region and the second region respectively, wherein the first region is a region where the first speaker is located, and the second region is a region where the second speaker is located;

[0009] determine the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker according to the phase difference and the beamforming signals corresponding to the first region and the second region respectively;

[0010] separate the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker.

[0011] In a possible implementation, the determining the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker according to the phase difference and the beamforming signals corresponding to the first region and the second region respectively comprises:

[0012] obtaining two logarithmic power spectrums by respectively calculating the logarithmic power spectrum of the beamforming signals corresponding to the first region and the second region respectively;

[0013] determining the speech time-frequency mask corresponding to the first region and the second region according to the phase difference and the two logarithmic power spectrums by using a pre-trained spatial separation model, wherein the speech time-frequency mask corresponding to the first region is taken as the first speech time-frequency mask of the first speaker, and the speech time-frequency mask corresponding to the second region is taken as the first speech time-frequency mask of the second speaker;

[0014] The spatial separation model is trained by using training samples labeled with real speech time-frequency masks corresponding to two different regions, wherein the training samples comprise a microphone phase difference and two logarithmic power spectrums determined according to a training noisy mixed speech signal, and the training noisy mixed speech signal is a noisy speech signal of multiple speakers in the two different regions.

[0015] In a possible implementation, the multi-speaker speech separation method further comprises:

[0016] determining the second speech time-frequency mask of the first speaker and the second speech time-frequency mask of the second speaker by clustering the multi-speaker speech signal by using a Gaussian mixture model;

[0017] The separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker comprises:

[0018] According to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker, and simultaneously combining the second speech time-frequency mask of the first speaker and the second speech time-frequency mask of the second speaker, the speech signal of the first speaker and the speech signal of the second speaker are separated from the signal of any channel of the multi-speaker speech signal.

[0019] In a possible implementation, the separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker, and simultaneously combining the second speech time-frequency mask of the first speaker and the second speech time-frequency mask of the second speaker comprises:

[0020] fusing the first speech time-frequency mask of the first speaker with the second speech time-frequency mask of the first speaker to obtain a target speech time-frequency mask of the first speaker, and fusing the first speech time-frequency mask of the second speaker with the second speech time-frequency mask of the second speaker to obtain a target speech time-frequency mask of the second speaker;

[0021] separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker.

[0022] In a possible implementation, the first speech time-frequency mask of the first speaker, the first speech time-frequency mask of the second speaker, the second speech time-frequency mask of the first speaker, and the second speech time-frequency mask of the second speaker are frame-level speech time-frequency masks.

[0023] The fusing the first speech time-frequency mask of the first speaker with the second speech time-frequency mask of the first speaker to obtain the target speech time-frequency mask of the first speaker, and fusing the first speech time-frequency mask of the second speaker with the second speech time-frequency mask of the second speaker to obtain the target speech time-frequency mask of the second speaker comprises:

[0024] taking a maximum value of the first speech time-frequency mask of the first speaker and the second speech time-frequency mask of the first speaker at a frame level to obtain the target speech time-frequency mask of the first speaker;

[0025] taking a maximum value of the first speech time-frequency mask of the second speaker and the second speech time-frequency mask of the second speaker at a frame level to obtain the target speech time-frequency mask of the second speaker.

[0026] In a possible implementation, the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker are both frame-level speech time-frequency masks.

[0027] The separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker comprises:

[0028] According to the target speech time-frequency mask of the first speaker, speech segments and non-speech segments of each frame signal of the multi-speaker speech signal are distinguished to obtain a first discrimination result, and according to the target speech time-frequency mask of the second speaker, speech segments and non-speech segments of each frame signal of the multi-speaker speech signal are distinguished to obtain a second discrimination result.

[0029] The speech signal of the first speaker and the speech signal of the second speaker are separated from the signal of any channel of the multi-speaker speech according to the first discrimination result and the second discrimination result.

[0030] In a possible implementation, the distinguishing speech segments and non-speech segments of each frame signal of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker comprises:

[0031] For each frame signal of the multi-speaker speech signal, if the speech time-frequency mask corresponding to the frame signal in the target speech time-frequency mask of the first speaker is greater than a set threshold, the frame signal is determined as a speech segment, otherwise, the frame signal is determined as a non-speech segment.

[0032] The distinguishing speech segments and non-speech segments of each frame signal of the multi-speaker speech signal according to the target speech time-frequency mask of the second speaker comprises:

[0033] For each frame signal of the multi-speaker speech signal, if the speech time-frequency mask corresponding to the frame signal in the target speech time-frequency mask of the second speaker is greater than a set threshold, the frame signal is determined as a speech segment, otherwise, the frame signal is determined as a non-speech segment.

[0034] In a possible implementation, the first discrimination result and the second discrimination result are represented by a 0 / 1 sequence, for any frame signal of the multi-speaker speech signal, if the frame signal is a speech segment, the discrimination result of the frame signal is represented by 0, if the frame signal is a non-speech segment, the discrimination result of the frame signal is represented by 1.

[0035] The separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the first discrimination result and the second discrimination result comprises:

[0036] The signal of any channel of the multi-speaker speech signal is multiplied by the first discrimination result to obtain the speech signal of the first speaker, and multiplied by the second discrimination result to obtain the speech signal of the second speaker.

[0037] In a possible implementation, the multi-speaker speech separation method further comprises:

[0038] The speech signal of the first speaker and the speech signal of the second speaker are respectively denoised by using a single-channel post-filtering model to obtain the denoised speech signal of the first speaker and the denoised speech signal of the second speaker.

[0039] The second aspect of the application provides a multi-speaker speech separation device, comprising: a speech signal acquisition module, a phase difference determination module, a beamforming processing module, a first speech time-frequency mask determination module, and a speech separation module;

[0040] The speech signal acquisition module is configured to acquire a multi-speaker speech signal collected by a microphone array composed of two microphones arranged at different positions for a first speaker and a second speaker located at different positions.

[0041] The phase difference determination module is configured to determine a phase difference between the two microphones according to the multi-speaker speech signal.

[0042] The beamforming processing module is configured to perform fixed beamforming processing for a first region and fixed beamforming processing for a second region on the multi-speaker speech signal to obtain beamforming signals corresponding to the first region and the second region respectively, wherein the first region is a region where the first speaker is located, and the second region is a region where the second speaker is located.

[0043] The first speech time-frequency mask determination module is configured to determine a first speech time-frequency mask of the first speaker and a first speech time-frequency mask of the second speaker according to the phase difference and the beamforming signals corresponding to the first region and the second region respectively.

[0044] The multi-speaker speech separation module is configured to separate the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker.

[0045] The third aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:

[0046] The memory is configured to store a computer program;

[0047] The processor is configured to execute the computer program, so that the electronic device can implement the steps of the multi-speaker speech separation method according to any one of the preceding aspects.

[0048] The fourth aspect of the present application provides a computer storage medium, the storage medium carrying one or more computer programs, when the one or more computer programs are executed by an electronic device, the electronic device can implement the steps of the multi-speaker speech separation method according to any one of the preceding aspects.

[0049] The fifth aspect of the present application provides a computer program product, comprising computer readable instructions, when the computer readable instructions are executed on an electronic device, the electronic device implements the steps of the multi-speaker speech separation method according to any one of the preceding aspects.

[0050] By the above technical solution, the multi-speaker speech separation method provided by the present application first acquires multi-speaker speech signals collected by two microphone arrays for first and second speakers located at different positions, then determines the phase difference of the two microphones according to the multi-speaker speech signals, and respectively performs fixed beamforming processing for the first area (the area where the first speaker is located) and fixed beamforming processing for the second area (the area where the second speaker is located) on the multi-speaker speech signals, to obtain the beamforming signals corresponding to the first area and the second area respectively, then determines the speech time-frequency mask of the first speaker and the speech time-frequency mask of the second speaker according to the phase difference and the beamforming signals corresponding to the first area and the second area respectively, and finally separates the speech signals of the first speaker and the speech signals of the second speaker from the signals of any channel of the multi-speaker speech according to the speech time-frequency mask of the first speaker and the speech time-frequency mask of the second speaker. The multi-speaker speech separation method provided by the present application uses the spatial information of the signals collected by the two microphone arrays to realize speaker differentiation, which is not affected by the gender of the speaker, has good separation effect and stability, and can meet the application requirements. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only are the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of the provided drawings.

[0052] Figure 1 A schematic diagram of a system architecture related to the present application;

[0053] Figure 2 A schematic diagram of a hardware structure of a terminal provided by the embodiments of the present application;

[0054] Figure 3 A schematic diagram of a hardware structure of a server provided by the embodiments of the present application;

[0055] Figure 4 A schematic diagram of a flow of a multi-speaker speech separation method provided by the embodiments of the present application;

[0056] Figure 5 A schematic diagram of a first area and a second area provided by the embodiments of the present application;

[0057] Figure 6 A schematic diagram of a flow of determining a first speech time-frequency mask of a first speaker and a first speech time-frequency mask of a second speaker according to a phase difference and beamforming signals corresponding to the first area and the second area respectively provided by the embodiments of the present application;

[0058] Figure 7 A schematic diagram of a flow of separating speech signals of a first speaker and a second speaker from signals of any channel of multi-speaker speech signals according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker, and simultaneously combining a second speech time-frequency mask of the first speaker and a second speech time-frequency mask of the second speaker provided by the embodiments of the present application;

[0059] Figure 8 A schematic diagram of a structure of a multi-speaker speech separation device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0060] The embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. The terms used in the embodiment part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0061] The embodiments of the present application will be described below in conjunction with the drawings. Those skilled in the art can know that with the development of technology and the appearance of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0062] The terms "first", "second", and the like in the description and in the claims of the present application and above drawings are used for distinguishing between similar objects and not necessarily for describing a specific sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and are merely employed for descriptive purposes. Furthermore, the terms "comprise", "include", "have" and any variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, system, product or apparatus that comprises, includes or has a list of elements is not necessarily limited to those elements, but can include other elements not expressly listed or inherent to such process, method, system, product or apparatus.

[0063] In one possible implementation, as shown in Figure 1 The system architecture related to the present application can include a terminal 101 and a server 102, and the terminal 101 can interact with the server 102 through a network (wired network or wireless network). The server 102 can include one or more servers (as an example, one server is included in the server 102). Figure 1 The terminal 101 can obtain a multi-speaker speech signal, and transmit the obtained multi-speaker speech signal to the server 102 through the network. The server 102 can perform speech separation on the multi-speaker speech signal by using the multi-speaker speech separation method provided by the present application, and feed back the separation result to the terminal.

[0064] In another possible implementation, the system architecture related to the present application can include a terminal. The terminal has strong data processing capability. The terminal can obtain a multi-speaker speech signal, and perform speech separation on the multi-speaker speech signal by using the multi-speaker speech separation method provided by the present application.

[0065] Next, the product form of the terminal is described.

[0066] The terminal described above can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, a robot, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and the like, and the embodiments of the present application do not make any limitation on this.

[0067] Figure 2 An optional hardware structure schematic diagram of the terminal is shown.

[0068] Reference is made to Figure 2As shown, the terminal can include a radio frequency unit 210, a memory 220, an input unit 230, a display unit 240, a camera 250 (optional), an audio circuit 260 (optional), a speaker 261 (optional), a microphone 262 (optional), a headphone jack 263 (optional), a processor 270, an external interface 280, a power supply 290, and the like. Those skilled in the art can understand that Figure 2 The terminal is merely an example and does not constitute a limitation on the terminal, and can include more or fewer components than shown, or combine certain components, or different components.

[0069] The input unit 230 can be used to receive input digital or character information, and to generate key signal input related to user settings and function controls of the terminal. Specifically, the input unit 230 can include a touch screen 231 (optional) and / or other input devices 232. The touch screen 231 can collect user touch operations (such as user operations using a finger, a joint, a stylus, or any suitable object on or near the touch screen) and drive the corresponding connection device according to the pre-set program. The touch screen can detect user touch actions on the touch screen, convert the touch actions into touch signals and send them to the processor 270, and can receive commands from the processor 270 and execute them; the touch signal at least includes touch point coordinate information. The touch screen 231 can provide an input interface and an output interface between the terminal and the user. In addition, the touch screen can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch screen 231, the input unit 230 can also include other input devices. Specifically, the other input devices 232 can include one or more of a physical keyboard, function keys (such as volume control buttons, on-off buttons, etc.), trackballs, mice, joysticks, etc.

[0070] The display unit 240 can be used to display information input by the user or information provided to the user, various menus of the terminal, interactive interfaces, file displays, and / or playing of any kind of multimedia files.

[0071] The memory 220 can be used to store instructions and data. The memory 220 can mainly include a storage instruction area and a storage data area. The storage data area can store various data such as multimedia files, texts, etc.; the storage instruction area can store software units such as operating systems, applications, instructions required by at least one function, etc., or their subsets, expanded sets. It can also include a non-volatile random access memory; to provide the processor 270 with software and applications that include managing hardware, software, and data resources in a computing processing device, supporting control. Also used for the storage of multimedia files, and the storage of running programs and applications.

[0072] The processor 270 is the control center of the terminal, connects each part of the whole terminal by various interfaces and lines, executes various functions of the terminal and processes data by running or executing instructions stored in the memory 220 and calling data stored in the memory 220, thereby overall controlling the terminal. Optionally, the processor 270 can include one or more processing units; preferably, the processor 270 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and application program, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 270. In some embodiments, the processor, the memory, can be implemented on a single chip, and in some embodiments, they can also be implemented on independent chips respectively. The processor 270 can also be used to generate corresponding operation control signals to the corresponding components of the computing processing device, read and process the data in the software, especially read and process the data and programs in the memory 220, so that each functional module therein executes corresponding functions, thereby controlling the corresponding components to act according to the requirements of the instructions.

[0073] The memory 220 can be used to store software codes related to the multi-speaker speech separation method, and the processor 270 can execute the software codes in the memory 220, or can also schedule other units (such as the above-mentioned input unit 230 and the display unit 240) to realize corresponding functions.

[0074] The radio frequency unit 210 (optional) can be used for receiving and sending signals in the process of information or communication, for example, receiving the downlink information of the base station and processing by the processor 270; in addition, sending the uplink data to the base station. Generally, the radio frequency unit 210 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the radio frequency unit 210 can also communicate with network devices and other devices through wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0075] In the embodiments of the present application, the radio frequency unit 210 can send data to other devices, and can also receive data sent by other devices. It should be understood that the radio frequency unit 210 is optional, which can be replaced by other communication interfaces, for example, can be a network interface.

[0076] The terminal also includes a power supply 290 (such as a battery) for supplying power to each component. Preferably, the power supply can be logically connected to the processor 270 through a power management system, so as to realize the functions of managing charging, discharging, and power consumption management through the power management system.

[0077] The terminal also includes an external interface 280, which can be a standard Micro USB interface, or can be a multi-pin connector, and can be used for connecting the terminal with other devices for communication, or can be used for connecting a charger to charge the terminal.

[0078] Although not shown, the terminal can also include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, different function sensors, etc., which will not be described here.

[0079] Next, the product form of the above server is described.

[0080] Figure 3 A structural schematic diagram of the above server is provided, as shown in Figure 3As shown, the server can include a bus 301, a processor 302, a communication interface 303, and a memory 304. The processor 302, the memory 304, and the communication interface 303 communicate through the bus 301.

[0081] The bus 301 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the middle, but it does not mean that there is only one bus or one type of bus.

[0082] The processor 302 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.

[0083] The memory 304 can include a volatile memory, such as a random access memory (RAM). The memory 304 can also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a mechanical hard disk drive (HDD), or a solid state drive (SSD).

[0084] The memory 304 can be used to store software code related to the multi-speaker speech separation method, and the processor 302 can call the software code stored in the memory 304, or can schedule other units to realize the corresponding functions.

[0085] The processor (for example, the processor 270 and the processor 302) in the terminal and the server described above can be a hardware circuit (for example, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor, a microcontroller, or the like), or a combination of the hardware circuits. For example, the processor can be a hardware system with an instruction execution function, such as a CPU, a DSP, or the like, or a hardware system without an instruction execution function, such as an ASIC, an FPGA, or the like, or a combination of the hardware system without an instruction execution function and the hardware system with an instruction execution function.

[0086] Next, the multi-speaker speech separation method provided by the present application is introduced through the following examples.

[0087] Referring to Figure 4 , a flowchart of a multi-speaker speech separation method provided by an embodiment of the present application is shown, which can include:

[0088] Step S401: Obtain a multi-speaker speech signal collected by a microphone array composed of two microphones arranged at different positions for a first speaker and a second speaker located at different orientations.

[0089] In this embodiment, a dual-microphone array (a microphone array composed of two microphones arranged at different positions) is used to collect speech signals for a first speaker and a second speaker located at different orientations, and a multi-speaker speech signal in two channels is obtained.

[0090] For example, the application scenario is an enterprise management scenario, the microphone array is a microphone array arranged on a smart electronic badge, and the two microphones constituting the microphone array can be vertically arranged on the smart electronic badge. The microphone array on the smart electronic badge can collect the speech signals of the wearer of the smart electronic badge (the first speaker) and others (the second speaker), thereby obtaining a multi-speaker speech signal.

[0091] Step S402: Determine the phase difference of the two microphones according to the multi-speaker speech signal, and perform fixed beamforming processing for a first region and fixed beamforming processing for a second region on the multi-speaker speech signal, respectively, to obtain beamformed signals corresponding to the first region and the second region, respectively.

[0092] The first area is an area where the first speaker (for example, a wearer of an intelligent electronic badge) is located, and the second area is an area where the second speaker (for example, a non-wearer of an intelligent electronic badge) is located, as shown in Figure 5 FIG. 2 shows a schematic diagram of the first area and the second area.

[0093] The voice signals collected by the two microphones have a slight time difference, which causes a phase difference between the signals. In this embodiment, the "phase difference between the two microphones" refers to the phase difference between the signals collected by the two microphones.

[0094] The fixed beamforming processing of the multi-speaker voice signal for the first area can enhance the voice signal of the first area speaker and suppress the voice signal of the second area speaker. The fixed beamforming processing of the multi-speaker voice signal for the second area can enhance the voice signal of the second area speaker and suppress the voice signal of the first area speaker.

[0095] Step S403: determining a first voice time-frequency mask of the first speaker and a first voice time-frequency mask of the second speaker according to the phase difference and the beamformed signals corresponding to the first area and the second area, respectively.

[0096] The first voice time-frequency mask of the first speaker can represent the proportion of the voice of the first speaker in the multi-speaker voice signal, and the first voice time-frequency mask of the second speaker can represent the proportion of the voice of the second speaker in the multi-speaker voice signal.

[0097] In this embodiment, the voice time-frequency masks corresponding to the first area and the second area are predicted according to the phase difference and the beamformed signals corresponding to the first area and the second area, respectively. The voice time-frequency mask corresponding to the first area is taken as the first voice time-frequency mask of the first speaker, and the voice time-frequency mask corresponding to the second area is taken as the first voice time-frequency mask of the second speaker.

[0098] Step S404: separating the voice signal of the first speaker and the voice signal of the second speaker from the signal of any channel of the multi-speaker voice signal according to the first voice time-frequency mask of the first speaker and the first voice time-frequency mask of the second speaker.

[0099] The multi-speaker voice signal is a two-channel signal, and the voice signal of the first speaker and the voice signal of the second speaker can be separated from the signal of any channel according to the first voice time-frequency mask of the first speaker and the first voice time-frequency mask of the second speaker.

[0100] The multi-speaker speech separation method provided in this application first acquires multi-speaker speech signals collected by a two-microphone array targeting a first speaker and a second speaker located at different positions. Then, based on the multi-speaker speech signals, the phase difference between the two microphones is determined. Fixed beamforming processing is then applied to the multi-speaker speech signals for a first region and a second region, respectively, to obtain beamforming signals corresponding to the first and second regions. Next, based on the phase difference and the beamforming signals corresponding to the first and second regions, a first speech time-frequency mask for the first speaker and a second speaker are determined. Finally, based on the first and second speaker time-frequency masks, the speech signals of the first speaker and the second speaker are separated from the signal in any channel of the multi-speaker speech. The multi-speaker speech separation method provided in this application utilizes the spatial information of the signals collected by the two-microphone array to distinguish speakers. This method is unaffected by the speaker's gender, has good separation effect and stability, and can meet application requirements.

[0101] In another embodiment of this application, the specific implementation process of "step S403: determining the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker based on the phase difference and the beamforming signals corresponding to the first region and the second region respectively" in the above embodiment will be described.

[0102] In one possible implementation, such as Figure 6 As shown, the process of determining the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker based on the phase difference and the beamforming signals corresponding to the first region and the second region respectively may include:

[0103] Step S601: Calculate the logarithmic power spectrum of the beamforming signals corresponding to the first region and the second region respectively, and obtain two logarithmic power spectra.

[0104] The logarithmic power spectrum is obtained for the beamforming signal corresponding to the first region, and the logarithmic power spectrum is obtained for the beamforming signal corresponding to the second region, resulting in two logarithmic power spectra.

[0105] Step S602: Using the pre-trained spatial separation model, based on the phase difference and two log power spectra, determine the speech time-frequency masks corresponding to the first region and the second region respectively. The speech time-frequency mask corresponding to the first region is used as the first speech time-frequency mask of the first speaker, and the speech time-frequency mask corresponding to the second region is used as the first speech time-frequency mask of the second speaker.

[0106] The phase difference and the two log power spectrums are input into a pre-trained spatial separation model, and the spatial separation model predicts the speech time-frequency masks corresponding to the first region and the second region respectively according to the input phase difference and the two log power spectrums and outputs the speech time-frequency masks.

[0107] The spatial separation model in the embodiment is trained by using training samples labeled with real speech time-frequency masks corresponding to two different regions (for example, a region where a smart electronic badge wearer is located and a region where a smart electronic badge non-wearer is located), and the training samples include a microphone phase difference and two log power spectrums determined according to a training noisy mixed speech signal, and the training noisy mixed speech signal is a noisy speech signal of multiple speakers (for example, a smart electronic badge wearer and a smart electronic badge non-wearer) in the two different regions.

[0108] The training noisy mixed speech signal is obtained by: obtaining clean speech signals s of the multiple speakers in the two different regions, and obtaining a noise signal n (which can be obtained by simulation or by a microphone array); generating room impulse responses I of sound sources and the microphone array at different distances, in different room environments, and at different azimuth and elevation angles by using the image source method, convolving the room impulse responses I with the clean speech signals s, and superimposing the noise signal n, as follows:

[0109] (1)

[0110] After obtaining the noisy mixed speech signal y, the speech energy of the noisy mixed speech signal y can be adjusted so that the signal-to-noise ratio is within a range of 0 dB to 20 dB, that is, the noisy mixed speech signal y is adjusted to a noisy mixed speech signal y' with a signal-to-noise ratio within a range of 0 dB to 20 dB, and the noisy mixed speech signal y' with a signal-to-noise ratio within a range of 0 dB to 20 dB is used as the training noisy mixed speech signal.

[0111] After obtaining the training noisy mixed speech signal y', real speech time-frequency masks corresponding to the two different regions can be determined according to the clean speech signal s and the training noisy mixed speech signal y', and specifically, the short-time Fourier transform is performed on the clean speech signal s and the training noisy mixed speech signal y' to obtain a frequency domain amplitude spectrum S of the clean speech signal s and a frequency domain amplitude spectrum Y of the training noisy mixed speech signal y', and the real speech time-frequency masks corresponding to the two different regions are obtained according to the following formula:

[0112] (2)

[0113] It should be noted that the real speech time-frequency mask is calculated frame by frame, and when calculating the real speech time-frequency mask corresponding to one region, the frequency domain amplitude spectrum S of each frame of the clean speech signal s i is calculated, and the frequency domain amplitude spectrum Y of each frame of the training noisy mixed speech signal y iFor the speech signal of the speaker in the region, the frequency domain amplitude spectrum S of the frame signal is calculated i and the frequency domain amplitude spectrum Y of the corresponding frame of the training noisy mixed speech signal y' i The ratio S i / Y i , that is, mask i ′=S i / Y i If the frame signal s i is not the speech signal of the speaker in the region, the frequency domain amplitude spectrum S i of the frame signal s i is 0, that is, mask i ′ is 0, and a real speech time-frequency mask corresponding to the region is obtained through the above process, and a real speech time-frequency mask corresponding to another region is obtained in a similar manner.

[0114] After obtaining the training noisy mixed speech signal y', the phase difference IPD12 of the two microphones is determined according to the training noisy mixed speech signal y', and the training noisy mixed speech signal y' is subjected to fixed beamforming processing for two different regions to obtain beamforming signals corresponding to the two different regions respectively, and the logarithmic power spectrum of the beamforming signals corresponding to the two different regions respectively is obtained to obtain two logarithmic power spectrums LPS_fix1 and LPS_fix2.

[0115] After obtaining the phase difference IPD12 and the two logarithmic power spectrums LPS_fix1 and LPS_fix2, the phase difference IPD12 and the two logarithmic power spectrums LPS_fix1 and LPS_fix2 are taken as training samples, and the real speech time-frequency masks mask1' and mask2' corresponding to the two different regions respectively are taken as sample annotation information to obtain training data (training samples: LPS_fix1, LPS_fix2, IPD12; sample annotation information: mask1', mask2').

[0116] A large amount of training data can be obtained in the above manner, and then the spatial separation model is trained using the large amount of training data.

[0117] Specifically, the training process of the spatial separation model can include: obtaining training data; inputting a training sample in the training data (i.e., LPS_fix1, LPS_fix2, IPD12) into the spatial separation model, the spatial separation model predicting speech time-frequency masks mask1 and mask2 corresponding to two different regions according to the input IPD12, LPS_fix1 and LPS_fix2, after predicting the speech time-frequency masks mask1 and mask2 corresponding to the two different regions, determining a prediction loss of the spatial separation model according to mask1 and mask2 and sample label information (i.e., real speech time-frequency masks mask1' and mask2' corresponding to the two different regions) in the obtained training data, and updating parameters of the spatial separation model according to the prediction loss of the spatial separation model. Optionally, the prediction loss can adopt a mean square error loss, i.e., calculating a mean square error loss L1 for mask1 and mask1', calculating a mean square error loss L2 for mask2 and mask2', fusing L1 and L2, and then updating parameters of the spatial separation model according to the fused loss. The spatial separation model is iteratively trained in the above manner until a training end condition (such as reaching a preset training number or model convergence) is met.

[0118] The spatial separation model in this embodiment can but is not limited to adopting a feedforward sequential memory neural network (FSMN).

[0119] In another embodiment of the present application, the specific implementation process of "step S404: separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker" in the above embodiment is introduced.

[0120] There are many implementation manners for separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker, in a possible implementation manner, the first speech time-frequency mask of the first speaker can be taken as the target speech time-frequency mask of the first speaker, and the first speech video mask of the second speaker can be taken as the target speech time-frequency mask of the second speaker, and then the speech signal of the first speaker and the speech signal of the second speaker are separated from the signal of any channel of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker.

[0121] The implementation manner is to separate the speech signals of the first speaker and the second speaker from the signals of any channel of the multi-speaker speech signal according to the first speech time-frequency mask of the first speaker and the first speech video mask of the second speaker.

[0122] To improve the speech signal separation effect, in another possible implementation manner, the second speech time-frequency mask of the first speaker and the second speech time-frequency mask of the second speaker can be determined by clustering the multi-speaker speech signal by using a Gaussian mixture model, and then the speech signals of the first speaker and the second speaker can be separated from the signals of any channel of the multi-speaker speech signal according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker, and in combination with the second speech time-frequency mask of the first speaker and the second speech time-frequency mask of the second speaker. The Gaussian mixture model can be trained by using the training noisy mixed speech signal and the corresponding real speech time-frequency mask, and then the multi-speaker speech signal can be clustered by using the trained Gaussian mixture model to determine the second speech time-frequency mask of the first speaker and the second speech time-frequency mask of the second speaker.

[0123] Specifically, as shown in Figure 7 The process of separating the speech signals of the first speaker and the second speaker from the signals of any channel of the multi-speaker speech signal according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker, and in combination with the second speech time-frequency mask of the first speaker and the second speech time-frequency mask of the second speaker can include:

[0124] Step S701: fuse the first speech time-frequency mask of the first speaker with the second speech time-frequency mask of the first speaker to obtain a target speech time-frequency mask of the first speaker, and fuse the first speech time-frequency mask of the second speaker with the second speech time-frequency mask of the second speaker to obtain a target speech time-frequency mask of the second speaker.

[0125] Specifically, the process of fusing the first speech time-frequency mask of the first speaker with the second speech time-frequency mask of the first speaker to obtain the target speech time-frequency mask of the first speaker can include: taking the maximum value of the first speech time-frequency mask of the first speaker and the second speech time-frequency mask of the first speaker at the frame level to obtain the target speech time-frequency mask of the first speaker.

[0126] Specifically, the process of fusing the first speech time-frequency mask of the second speaker with the second speech time-frequency mask of the second speaker to obtain the target speech time-frequency mask of the second speaker can include: taking the maximum value of the first speech time-frequency mask of the second speaker and the second speech time-frequency mask of the second speaker at the frame level to obtain the target speech time-frequency mask of the second speaker.

[0127] If the first speech time-frequency mask of the first speaker is denoted as nn_mask1, the second speech time-frequency mask of the first speaker is denoted as CGMMmask1, the first speech time-frequency mask of the second speaker is denoted as nn_mask2, and the second speech time-frequency mask of the second speaker is denoted as CGMMmask2, the target speech time-frequency mask of the first speaker MASK1 and the target speech time-frequency mask of the second speaker MASK2 can be expressed as:

[0128] MASK1 = max (nn_mask1, CGMMmask1) (3)

[0129] MASK2 = max (nn_mask2, CGMMmask2) (4)

[0130] For example, the first speech time-frequency mask nn_mask1 of the first speaker is {nn_mask11, nn_mask12, nn_mask13, nn_mask14, nn_mask15,...}, the second speech time-frequency mask CGMMmask1 of the first speaker is {CGMMmask11, CGMMmask12, CGMMmask13, CGMMmask14, CGMMmask15,...}, and it is assumed that nn_mask11 > CGMMmask11, nn_mask12 < CGMMmask12, nn_mask13 > CGMMmask13, nn_mask14 < CGMMmask14, nn_mask15 > CGMMmask14,..., then the target speech time-frequency mask of the first speaker MASK1 = max (nn_mask1, CGMMmask1) = {nn_mask11, CGMMmask12, nn_mask13, CGMMmask14, nn_mask15,...}.

[0131] Step S702: Separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker.

[0132] Next, the specific implementation process of "separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker" is introduced.

[0133] For example, the first speech time-frequency mask nn_mask1 of the first speaker is {nn_mask11, nn_mask12, nn_mask13, nn_mask14, nn_mask15,...}, the second speech time-frequency mask CGMMmask1 of the first speaker is {CGMMmask11, CGMMmask12, CGMMmask13, CGMMmask14, CGMMmask15,...}, and it is assumed that nn_mask11 > CGMMmask11, nn_mask12 < CGMMmask12, nn_mask13 > CGMMmask13, nn_mask14 < CGMMmask14, nn_mask15 > CGMMmask14,..., then the target speech time-frequency mask of the first speaker MASK1 = max (nn_mask1, CGMMmask1) = {nn_mask11, CGMMmask12, nn_mask13, CGMMmask14, nn_mask15,...}. Figure 8As shown, the specific implementation process of separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker can include:

[0134] Step S801: According to the target speech time-frequency mask of the first speaker, the speech segments and non-speech segments of each frame signal of the multi-speaker speech signal are distinguished to obtain a first discrimination result, and according to the target speech time-frequency mask of the second speaker, the speech segments and non-speech segments of each frame signal of the multi-speaker speech signal are distinguished to obtain a second discrimination result.

[0135] Among them, the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker are both frame-level speech time-frequency masks.

[0136] Specifically, the process of distinguishing the speech segments and non-speech segments of each frame signal of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker to obtain the first discrimination result can include: for each frame signal of the multi-speaker speech signal, if the speech time-frequency mask corresponding to the frame signal in the target speech time-frequency mask of the first speaker is greater than a set threshold (such as 0.2), it is determined that the frame signal is a speech segment, otherwise, it is determined that the frame signal is a non-speech segment, so the first discrimination result composed of the discrimination results of each frame signal of the multi-speaker speech signal can be obtained.

[0137] Specifically, the process of distinguishing the speech segments and non-speech segments of each frame signal of the multi-speaker speech signal according to the target speech time-frequency mask of the second speaker to obtain the second discrimination result can include: for each frame signal of the multi-speaker speech signal, if the speech time-frequency mask corresponding to the frame signal in the target speech time-frequency mask of the second speaker is greater than a set threshold (such as 0.2), it is determined that the frame signal is a speech segment, otherwise, it is determined that the frame signal is a non-speech segment, so the second discrimination result composed of the discrimination results of each frame signal of the multi-speaker speech signal can be obtained.

[0138] Step S802: According to the first discrimination result and the second discrimination result, the speech signal of the first speaker and the speech signal of the second speaker are separated from the signal of any channel of the multi-speaker speech.

[0139] For the signal of any channel of the multi-speaker speech signal, the speech signal of the first speaker can be separated from the signal of the channel according to the first discrimination result, and the speech signal of the second speaker can be separated from the signal of the channel according to the second discrimination result.

[0140] In a possible implementation, for any frame signal of the multi-speaker voice signal, if the frame signal is a speech segment, the discrimination result of the frame signal can be represented by 1, if the frame signal is a non-speech segment, the discrimination result of the frame signal can be represented by 0, that is, the first discrimination result and the second discrimination result can be represented by a 0 / 1 sequence. For the signal of any channel of the multi-speaker voice signal, the signal of the channel can be multiplied by the first discrimination result (0 / 1 sequence) to obtain the voice signal of the first speaker, and the signal of the channel can be multiplied by the second discrimination result (0 / 1 sequence) to obtain the voice signal of the second speaker.

[0141] If the first discrimination result is denoted as VAD1, the second discrimination result is denoted as VAD2, and the signal of one channel of the multi-speaker voice signal is denoted as mic-sig1 (the signal collected by one microphone in the microphone array), the voice signal SIG1 of the first speaker and the voice signal SIG2 of the second speaker can be represented as:

[0142] (5)

[0143] (6)

[0144] Considering that the voice signal SIG1 of the first speaker and the voice signal SIG2 of the second speaker can be transcribed subsequently, in order to improve the transcription rate, in another embodiment of the present application, after the voice signal SIG1 of the first speaker and the voice signal SIG2 of the second speaker are obtained, the voice signal SIG1 of the first speaker and the voice signal SIG2 of the second speaker can be respectively denoised by using a single-channel post-filtering model to obtain the denoised voice signal of the first speaker and the denoised voice signal of the second speaker.

[0145] In a possible implementation, the single-channel post-filtering model can be a neural network denoising model, which is trained by using a training single-channel noisy voice signal and a corresponding single-channel clean voice signal, and the training target is to make the signal obtained by denoising the training single-channel noisy voice signal by using the neural network denoising model close to the corresponding single-channel clean voice signal of the training single-channel noisy voice signal.

[0146] The multi-speaker voice separation method provided in the embodiments of the present application uses the spatial information of the signals received by the two microphone arrays to distinguish the speakers, and can obtain more accurate separation results by combining the separation strategy based on the spatial separation model with the separation strategy based on Gaussian mixture model clustering. In addition, further post-filtering of the separation results can enhance the voice, thereby improving the transcription rate of subsequent voice transcription.

[0147] The multi-speaker speech separation method provided in the embodiments of the present application is introduced above, and the device corresponding to the multi-speaker speech separation method is introduced below.

[0148] Please refer to Figure 8 , Figure 8 The multi-speaker speech separation device provided in the embodiments of the present application has a structure diagram as shown in FIG. 8. The multi-speaker speech separation device can include a speech signal acquisition module 801, a phase difference determination module 802a, a beamforming processing module 802b, a first time-frequency mask determination module 803a, and a speaker speech separation module 804.

[0149] The speech signal acquisition module 801 is configured to acquire multi-speaker speech signals collected by a microphone array composed of two microphones arranged at different positions for a first speaker and a second speaker located at different orientations.

[0150] The phase difference determination module 802a is configured to determine the phase difference of the two microphones according to the multi-speaker speech signals.

[0151] The beamforming processing module 802b is configured to perform fixed beamforming processing for a first region and fixed beamforming processing for a second region on the multi-speaker speech signals respectively, to obtain beamforming signals corresponding to the first region and the second region respectively. The first region is the region where the first speaker is located, and the second region is the region where the second speaker is located.

[0152] The first time-frequency mask determination module 803a is configured to determine the first time-frequency mask of the first speaker and the first time-frequency mask of the second speaker according to the phase difference and the beamforming signals corresponding to the first region and the second region respectively.

[0153] The multi-speaker speech separation module 804 is configured to separate the speech signals of the first speaker and the second speaker from the signals of any channel of the multi-speaker speech signals according to the first time-frequency mask of the first speaker and the first time-frequency mask of the second speaker.

[0154] In a possible implementation, when the first time-frequency mask determination module 803a determines the first time-frequency mask of the first speaker and the first time-frequency mask of the second speaker according to the phase difference and the beamforming signals corresponding to the first region and the second region respectively, the first time-frequency mask determination module 803a is specifically configured to:

[0155] obtain two log power spectrums by respectively calculating the log power spectrums of the beamforming signals corresponding to the first region and the second region respectively;

[0156] The pre-trained spatial separation model is used to determine speech time-frequency masks corresponding to the first region and the second region based on the phase difference and the two log power spectrums, and the speech time-frequency mask corresponding to the first region is taken as the first speech time-frequency mask of the first speaker, and the speech time-frequency mask corresponding to the second region is taken as the first speech time-frequency mask of the second speaker.

[0157] The spatial separation model is trained by using training samples labeled with real speech time-frequency masks corresponding to two different regions, and the training samples include the microphone phase difference and the two log power spectrums determined according to a training noisy mixed speech signal, and the training noisy mixed speech signal is a noisy speech signal of multiple speakers in the two different regions.

[0158] In a possible implementation, the multi-speaker speech separation device can further include a second speech time-frequency mask determination module 803b.

[0159] The second speech time-frequency mask determination module 803b is configured to determine the second speech time-frequency mask of the first speaker and the second speech time-frequency mask of the second speaker by clustering the multi-speaker speech signal by using a Gaussian mixture model.

[0160] The multi-speaker speech separation module 804 is configured to, when separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker, specifically:

[0161] The multi-speaker speech separation module 804 is configured to, when separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker, specifically:

[0162] In a possible implementation, the multi-speaker speech separation module 804 can include a time-frequency mask fusion module and a separation module.

[0163] The time-frequency mask fusion module is configured to fuse the first speech time-frequency mask of the first speaker and the second speech time-frequency mask of the first speaker to obtain a target speech time-frequency mask of the first speaker, and fuse the first speech time-frequency mask of the second speaker and the second speech time-frequency mask of the second speaker to obtain a target speech time-frequency mask of the second speaker.

[0164] The separation module is configured to separate the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker.

[0165] In a possible implementation, the first speech time-frequency mask of the first speaker, the first speech time-frequency mask of the second speaker, the second speech time-frequency mask of the first speaker, and the second speech time-frequency mask of the second speaker are all frame-level speech time-frequency masks.

[0166] The time-frequency mask fusion module is specifically configured to:

[0167] maximizing the first speech time-frequency mask of the first speaker and the second speech time-frequency mask of the first speaker at the frame level to obtain the target speech time-frequency mask of the first speaker;

[0168] maximizing the first speech time-frequency mask of the second speaker and the second speech time-frequency mask of the second speaker at the frame level to obtain the target speech time-frequency mask of the second speaker.

[0169] In a possible implementation, the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker are both frame-level speech time-frequency masks.

[0170] The separation module is specifically configured to:

[0171] performing speech segment and non-speech segment discrimination on each frame signal of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker to obtain a first discrimination result, and performing speech segment and non-speech segment discrimination on each frame signal of the multi-speaker speech signal according to the target speech time-frequency mask of the second speaker to obtain a second discrimination result;

[0172] separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech according to the first discrimination result and the second discrimination result.

[0173] In a possible implementation, when performing speech segment and non-speech segment discrimination on each frame signal of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker, the separation module is specifically configured to:

[0174] For each frame signal of the multi-speaker speech signal, if the speech time-frequency mask corresponding to the frame signal in the target speech time-frequency mask of the first speaker is greater than a set threshold, it is determined that the frame signal is a speech segment, otherwise, it is determined that the frame signal is a non-speech segment.

[0175] In a possible implementation, when the separation module discriminates the speech segments and the non-speech segments of each frame signal of the multi-speaker speech signal according to the target speech time-frequency mask of the second speaker, the separation module is specifically configured to:

[0176] For each frame signal of the multi-speaker speech signal, if the speech time-frequency mask corresponding to the frame signal in the target speech time-frequency mask of the second speaker is greater than a set threshold, it is determined that the frame signal is a speech segment, otherwise, it is determined that the frame signal is a non-speech segment.

[0177] In a possible implementation, the first discrimination result and the second discrimination result are represented by a 0 / 1 sequence, and for any frame signal of the multi-speaker speech signal, if the frame signal is a speech segment, the discrimination result of the frame signal is represented by 0, and if the frame signal is a non-speech segment, the discrimination result of the frame signal is represented by 1.

[0178] When the separation module separates the speech signal of the first speaker and the speech signal of the second speaker from any channel signal of the multi-speaker speech according to the first discrimination result and the second discrimination result, the separation module is specifically configured to:

[0179] For any channel signal of the multi-speaker speech signal, the channel signal is multiplied by the first discrimination result to obtain the speech signal of the first speaker, and the channel signal is multiplied by the second discrimination result to obtain the speech signal of the second speaker.

[0180] In a possible implementation, the multi-speaker speech separation apparatus can further include a post-filtering module 805.

[0181] The post-filtering module 805 is configured to perform noise reduction on the speech signal of the first speaker and the speech signal of the second speaker respectively by using a single-channel post-filtering model, to obtain a noise-reduced speech signal of the first speaker and a noise-reduced speech signal of the second speaker.

[0182] The multi-speaker speech separation apparatus provided by the embodiments of the present application uses the spatial information of the signals collected by the two microphone arrays to realize speaker differentiation, the method is not affected by the gender of the speaker, has good separation effect and stability, and can meet the application requirements.

[0183] The embodiments of the present application further provide an electronic device, which can include at least one processor, at least one communication interface, at least one memory and at least one communication bus.

[0184] In the embodiments of the present application, the number of the processor, the communication interface, the memory and the communication bus is at least one, and the processor, the communication interface and the memory complete the communication with each other through the communication bus.

[0185] The processor can be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application, etc.

[0186] The memory can include a high-speed RAM memory, and can also include a non-volatile memory, such as at least one disk memory.

[0187] The memory stores a program, and the processor can invoke the program stored in the memory, and the program is used to implement the steps of the multi-speaker speech separation method provided by the above embodiments.

[0188] The embodiments of the present application also provide a computer storage medium, which carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the steps of the multi-speaker speech separation method provided by the above embodiments.

[0189] The embodiments of the present application also provide a computer program product, which includes computer readable instructions, and when the computer readable instructions run on an electronic device, the electronic device implements the steps of the multi-speaker speech separation method provided by the above embodiments.

[0190] In addition, it should be noted that the apparatus embodiments described above are only schematic, and the units described as separate units can or can not be physically separate, and the units displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to actual needs, part or all of the modules can be selected to achieve the purpose of the embodiments. In addition, the connection relationship between the modules in the apparatus embodiments provided by the present application indicates that there is a communication connection between them, which can be realized as one or more communication buses or signal lines.

[0191] Those skilled in the art can clearly understand that the application can be implemented by means of software plus necessary universal hardware, and of course can also be implemented by means of dedicated hardware including special integrated circuit, special CPU, special memory, special component, etc. Generally, any function completed by computer program can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the application, software program implementation is a better embodiment. Based on such understanding, the technical solution of the application or the part of the application which makes contribution to the prior art can be embodied in the form of software product, which is stored in readable storage medium, such as computer floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a plurality of instructions for making a computer device (which can be personal computer, training device or network device, etc.) execute the method described in various embodiments of the application.

[0192] In the above embodiments, the implementation can be achieved by software, hardware, firmware or any combination thereof, entirely or partially. When implemented by software, the implementation can be achieved in the form of a computer program product, entirely or partially.

[0193] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the flow or function described in the embodiments of the application is generated entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as training device, data center, etc. integrated with one or more available media sets. The available medium can be magnetic medium (such as floppy disk, hard disk, magnetic tape), optical medium (such as DVD) or semiconductor medium (such as solid state disk (SSD)) etc.

Claims

1. A multi-speaker speech separation method, characterized by, The method comprises: acquiring multi-speaker voice signals collected by a microphone array composed of two microphones arranged at different positions for a first speaker and a second speaker located at different positions; determining a phase difference of the two microphones according to the multi-speaker voice signals, and performing fixed beamforming processing for a first region and fixed beamforming processing for a second region on the multi-speaker voice signals respectively to obtain beamforming signals corresponding to the first region and the second region respectively, wherein the first region is a region where the first speaker is located, and the second region is a region where the second speaker is located; determining a first speech time-frequency mask of the first speaker and a first speech time-frequency mask of the second speaker according to the phase difference and the beamforming signals corresponding to the first region and the second region respectively by using a pre-trained spatial separation model, wherein the spatial separation model is trained by using training samples labeled with real speech time-frequency masks corresponding to two different regions, and the training samples comprise microphone phase differences and two log power spectrums determined according to training noisy mixed voice signals, and the training noisy mixed voice signals are noisy voice signals of multiple speakers in the two different regions; separating the speech signals of the first speaker and the second speaker from signals of any channel of the multi-speaker voice signals according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker.

2. The multi-speaker speech separation method of claim 1, wherein, The method further comprises: determining a second speech time-frequency mask of the first speaker and a second speech time-frequency mask of the second speaker by clustering the multi-speaker voice signals by using a Gaussian mixture model. The method further comprises:

3. The multi-speaker speech separation method of claim 1, wherein, determining a second speech time-frequency mask of the first speaker and a second speech time-frequency mask of the second speaker by clustering the multi-speaker voice signals by using a Gaussian mixture model. The method further comprises: determining a second speech time-frequency mask of the first speaker and a second speech time-frequency mask of the second speaker by clustering the multi-speaker voice signals by using a Gaussian mixture model. According to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker, and simultaneously combining the second speech time-frequency mask of the first speaker and the second speech time-frequency mask of the second speaker, the speech signal of the first speaker and the speech signal of the second speaker are separated from the signal of any channel of the multi-speaker speech signal.

4. The multi-speaker speech separation method of claim 3, wherein, The method according to the first speech time-frequency mask of the first speaker and the first speech time-frequency mask of the second speaker, and simultaneously combining the second speech time-frequency mask of the first speaker and the second speech time-frequency mask of the second speaker, the speech signal of the first speaker and the speech signal of the second speaker are separated from the signal of any channel of the multi-speaker speech signal, comprising: Fusing the first speech time-frequency mask of the first speaker with the second speech time-frequency mask of the first speaker to obtain the target speech time-frequency mask of the first speaker, and fusing the first speech time-frequency mask of the second speaker with the second speech time-frequency mask of the second speaker to obtain the target speech time-frequency mask of the second speaker; According to the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker, the speech signal of the first speaker and the speech signal of the second speaker are separated from the signal of any channel of the multi-speaker speech signal.

5. The multi-speaker speech separation method of claim 4, wherein, The first speech time-frequency mask of the first speaker, the first speech time-frequency mask of the second speaker, the second speech time-frequency mask of the first speaker and the second speech time-frequency mask of the second speaker are all frame-level speech time-frequency masks; The method of fusing the first speech time-frequency mask of the first speaker with the second speech time-frequency mask of the first speaker to obtain the target speech time-frequency mask of the first speaker, and fusing the first speech time-frequency mask of the second speaker with the second speech time-frequency mask of the second speaker to obtain the target speech time-frequency mask of the second speaker, comprising: Taking the maximum value of the first speech time-frequency mask of the first speaker and the second speech time-frequency mask of the first speaker at the frame level to obtain the target speech time-frequency mask of the first speaker; Taking the maximum value of the first speech time-frequency mask of the second speaker and the second speech time-frequency mask of the second speaker at the frame level to obtain the target speech time-frequency mask of the second speaker.

6. The multi-speaker speech separation method of claim 4, wherein, The target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker are both frame-level speech time-frequency masks; The method of separating the speech signal of the first speaker and the speech signal of the second speaker from the signal of any channel of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker and the target speech time-frequency mask of the second speaker, comprising: According to the target speech time-frequency mask of the first speaker, speech segments and non-speech segments of each frame signal of the multi-speaker speech signal are distinguished to obtain a first discrimination result, and according to the target speech time-frequency mask of the second speaker, speech segments and non-speech segments of each frame signal of the multi-speaker speech signal are distinguished to obtain a second discrimination result; According to the first discrimination result and the second discrimination result, the speech signal of the first speaker and the speech signal of the second speaker are separated from the signal of any channel of the multi-speaker speech.

7. The multi-speaker speech separation method of claim 6, wherein, The distinguishing of speech segments and non-speech segments of each frame signal of the multi-speaker speech signal according to the target speech time-frequency mask of the first speaker comprises: If the speech time-frequency mask corresponding to the frame signal in the target speech time-frequency mask of the first speaker is greater than a set threshold, it is determined that the frame signal is a speech segment, otherwise, it is determined that the frame signal is a non-speech segment; The distinguishing of speech segments and non-speech segments of each frame signal of the multi-speaker speech signal according to the target speech time-frequency mask of the second speaker comprises: If the speech time-frequency mask corresponding to the frame signal in the target speech time-frequency mask of the second speaker is greater than a set threshold, it is determined that the frame signal is a speech segment, otherwise, it is determined that the frame signal is a non-speech segment.

8. The multi-speaker speech separation method of claim 6, wherein, The first discrimination result and the second discrimination result are represented by a 0 / 1 sequence, and for any frame signal of the multi-speaker speech signal, if the frame signal is a speech segment, the discrimination result of the frame signal is represented by 0, and if the frame signal is a non-speech segment, the discrimination result of the frame signal is represented by 1; According to the first discrimination result and the second discrimination result, the speech signal of the first speaker and the speech signal of the second speaker are separated from the signal of any channel of the multi-speaker speech. For the signal of any channel of the multi-speaker speech signal, the signal of the channel is multiplied by the first discrimination result to obtain the speech signal of the first speaker, and the signal of the channel is multiplied by the second discrimination result to obtain the speech signal of the second speaker.

9. The multi-speaker speech separation method of any one of claims 1-8, wherein, Further comprising: The speech signal of the first speaker and the speech signal of the second speaker are respectively denoised by using a single-channel post-filtering model to obtain the denoised speech signal of the first speaker and the denoised speech signal of the second speaker.

10. A multi-speaker speech separation apparatus, characterized by, Comprise: A speech signal acquisition module, a phase difference determination module, a beamforming processing module, a first speech time-frequency mask determination module, and a speech separation module; The speech signal acquisition module is configured to acquire a multi-speaker speech signal collected by a microphone array composed of two microphones arranged at different positions for a first speaker and a second speaker located at different positions; The phase difference determination module is configured to determine a phase difference between the two microphones according to the multi-speaker speech signal. The beamforming processing module is configured to perform fixed beamforming processing for a first region and fixed beamforming processing for a second region on the multi-speaker voice signal respectively, to obtain beamformed signals corresponding to the first region and the second region respectively, wherein the first region is a region where the first speaker is located, and the second region is a region where the second speaker is located. The first voice time-frequency mask determination module is configured to determine a first voice time-frequency mask of the first speaker and a first voice time-frequency mask of the second speaker by using a pre-trained spatial separation model, according to the phase difference and the beamformed signals corresponding to the first region and the second region respectively, wherein the spatial separation model is trained by using training samples labeled with real voice time-frequency masks corresponding to two different regions, and the training samples include microphone phase differences and two logarithmic power spectrums determined according to training noisy mixed voice signals, and the training noisy mixed voice signals are noisy voice signals of multiple speakers in the two different regions. The multi-speaker voice separation module is configured to separate voice signals of the first speaker and voice signals of the second speaker from signals of any channel of the multi-speaker voice signal according to the first voice time-frequency mask of the first speaker and the first voice time-frequency mask of the second speaker.

11. An electronic device, comprising: An electronic device includes at least one processor and a memory connected to the processor, wherein: The memory is configured to store a computer program; The processor is configured to execute the computer program to enable the electronic device to implement steps of the multi-speaker voice separation method according to any one of claims 1-9.

12. A computer storage medium, characterized in that The storage medium carries one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement steps of the multi-speaker voice separation method according to any one of claims 1-9.

13. A computer program product, characterised in that, The computer readable instructions, when executed on an electronic device, enable the electronic device to implement steps of the multi-speaker voice separation method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Audio signal processing method, device, terminal and storage medium

    CN111128221A

  • Low-latency speech separation

    US20200322722A1