A multi-channel array multi-speaker speech separation method, electronic device, medium
By using a multi-channel microphone array for sound source localization and multi-layer convolutional neural network separation, the challenge of multi-speaker speech separation is solved, achieving efficient speech separation and enhancement effects, and is suitable for complex acoustic environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG LAB
- Filing Date
- 2023-10-17
- Publication Date
- 2026-07-24
AI Technical Summary
In real-world scenarios such as indoor meetings and roundtable discussions, separating the speech of each speaker from audio signals captured by multiple microphones presents a challenge because the signal may contain overlapping speech from multiple sound sources, background noise, and reverberation.
Sound source localization is performed using a multi-channel microphone array, the angular distance of the sound source is calculated, a sound source angular distance registration coefficient is introduced, and a multi-layer convolutional neural network is used for speech separation to separate the single-channel speech signal of the target speaker.
It effectively separates speech signals from multiple speakers, improving speech intelligibility and quality. It is suitable for multi-channel microphone arrays, does not depend on a specific target speaker or a fixed number of speakers, and has high accuracy and adaptability.
Smart Images

Figure CN117275506B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to array-based speech signal separation and enhancement, and more particularly to a multi-channel array multi-speaker speech separation method, electronic device, and medium. Background Technology
[0002] Multichannel speech separation refers to the task of separating individual speech sources from mixed audio signals captured by multiple microphones or sensors. The goal is to extract speech signals from the mixed speech, thereby improving the intelligibility and quality of each speech source.
[0003] In real-world scenarios such as indoor meetings and roundtable discussions, multiple microphones are typically used to capture audio signals from different spatial locations. However, the captured audio may contain overlapping speech from multiple sound sources, background noise, and reverberation, making it challenging to separate the speech of each speaker.
[0004] Therefore, there is an urgent need to propose a speech separation method for multi-channel arrays. Summary of the Invention
[0005] In view of this, the present invention proposes a multi-channel array multi-speaker speech separation method, electronic device, and medium.
[0006] In a first aspect, embodiments of the present invention provide a multi-channel array multi-speaker speech separation method, the method comprising:
[0007] Acquire mixed speech signals from multiple speakers using a multi-channel microphone array;
[0008] The desired sound source and other interfering sound sources are located in the mixed speech signal of multiple speakers in a multi-channel microphone array. The direction angle of the desired sound source and the direction angle of the interfering sound source are obtained. The vector difference between the direction angle of the desired sound source and the direction angle of the interfering sound source is used as the angular distance of the sound source.
[0009] Calibrate the sound source angle distance registration coefficient, and fit the sound source angle distance registration coefficient mapping relationship curve based on the sound source angle distance and the sound source angle distance registration coefficient;
[0010] Based on the mapping relationship curve between the sound source angular distance and the fitted sound source angular distance registration coefficient, the speech signal received by each acoustic sensor in the multi-channel microphone array in the desired direction is obtained.
[0011] The speech signals received by each acoustic sensor in the multi-channel microphone array in the desired direction are input into a pre-trained speech separation network to separate the single-channel speech signal of the target speaker.
[0012] Secondly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-described multi-channel array multi-speaker speech separation method.
[0013] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the above-described multi-channel array multi-speaker speech separation method.
[0014] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a multi-channel array multi-speaker speech separation method. By introducing a sound source angle-distance registration coefficient into the mixed speech signal acquired by the multi-channel array, the mixed speech signal from the microphone array after sound source angle-distance registration control is input into a speech separation network to obtain the separated single-channel speech signal of the target speaker. This provides a convenient way to separate the mixed speech signal of multiple speakers acquired by a multi-channel microphone array, effectively enhancing the recognition of single speaker content. Furthermore, the multi-channel multi-speaker speech separation system proposed in this invention does not rely on a specific target speaker and can handle a variable number of speakers. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating a multi-channel array multi-speaker speech separation method provided in an embodiment of the present invention;
[0017] Figure 2 This is a schematic diagram of the sound source angle and distance provided in an embodiment of the present invention;
[0018] Figure 3 This is a schematic diagram of the sound source angle distance registration coefficient mapping relationship curve provided in an embodiment of the present invention;
[0019] Figure 4 This is a schematic diagram of the structure and post-processing steps of the multi-channel, multi-layer convolutional neural network provided in an embodiment of the present invention;
[0020] Figure 5 This is a time-domain waveform diagram of one single channel of the array-mixed speech signal provided in an embodiment of the present invention;
[0021] Figure 6This is a time-domain waveform diagram of the single-channel speech signal of the separated target speaker provided in an embodiment of the present invention;
[0022] Figure 7 This is a time-domain waveform diagram of the reference speech signal of the target speaker provided in an embodiment of the present invention;
[0023] Figure 8 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0024] The present invention will be further described below with reference to specific embodiments and accompanying drawings. The description of the embodiments below is only for the purpose of helping to understand the present invention. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principle of the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
[0025] Speech signals acquired by microphone arrays have advantages over those acquired using a single acoustic sensor in terms of sound source localization, noise suppression, and reverberation cancellation. This invention proposes a multi-channel array method for multi-speaker speech separation. It combines the spatial location information generated by the microphone array from the captured multi-speaker speech signals, analyzes the spatial and temporal characteristics of the audio signals (i.e., sound waves from different sources arrive at each microphone with different time delays, phases, and amplitudes), and estimates the contribution of each speech source in the mixed signal. This method can effectively separate mixed speech signals from simultaneous conversations with multiple speakers in complex scenarios.
[0026] In a typical noisy environment, the signal received by the m-th acoustic sensor in the microphone array can be expressed as:
[0027]
[0028] Where N is the number of speakers, and M is the number of acoustic sensors in the multi-channel microphone array. p (t) represents the speech signal of the p-th speaker. V m (t) represents the noise received by the m-th acoustic sensor, and it is assumed that it is uncorrelated with the speech signal.
[0029] Furthermore, dividing the N speakers into one desired target speaker and the remaining c interfering speakers (c = N-1), the received signal of the m-th acoustic sensor in the multi-channel microphone array can be expressed as:
[0030]
[0031] Among them, S q (t) represents the speech signal of the q-th interfering speaker, where q = 1, 2, ..., c, and S... exp(t) represents the speech signal of the target speaker in the desired direction.
[0032] like Figure 1 As shown, this invention proposes a multi-channel array multi-speaker speech separation method, which includes the following steps:
[0033] Step S1: Acquire mixed speech signals from multiple speakers using a multi-channel microphone array.
[0034] Step S2: Locate the desired sound source and other interfering sound sources in the mixed speech signal of the multi-speaker microphone array, obtain the direction angle of the desired sound source and the direction angle of the interfering sound source, and use the vector difference between the direction angle of the desired sound source and the direction angle of the interfering sound source as the angular distance of the sound source.
[0035] Specifically, the desired sound source and other interfering sound sources are located in the mixed speech signal of multiple speakers from a multi-channel microphone array, and the direction angles θ of the desired sound source and the direction angles θ of the interfering sound sources are obtained. i i = 1, 2, ..., N-1, θ and θ i The numerical values range from 0 to 360 degrees.
[0036] In this example, using the positive X-axis counterclockwise direction as a reference, the desired sound source direction angle θ and the interfering sound source direction angle θ are compared. i The difference in values is denoted as Figure 2 This is a schematic diagram showing the angle and distance of the sound source.
[0037] Furthermore, in order to accurately describe the numerical value of the angular distance between the directions of each sound source in a multi-channel array multi-speaker speech separation system, the numerical value of the angular distance Δθ between the sound sources can be calculated by the following expression:
[0038] when When the angle is less than 180 degrees, Δθ = |θ - θ i |
[0039] when When the angle is greater than 180 degrees, Δθ = θ + (360 - θ) i ).
[0040] Step S3: Calibrate the sound source angle distance registration coefficient, and fit the sound source angle distance registration coefficient mapping relationship curve based on the sound source angle distance and the sound source angle distance registration coefficient.
[0041] To enhance the speech signal from the direction of the sound source and suppress interfering sound sources, a sound source angle-distance registration coefficient ω(Δθ) is introduced into the microphone array speech separation system. The magnitude of the registration coefficient and Δθ have a non-linear mapping relationship, which is represented by a piecewise function, such as... Figure 3 As shown.
[0042] If Δθ = 0, that is, when there is only one sound source in one direction, the relation ω(Δθ) = 1 is satisfied.
[0043] In this example, when When Δθ = π, the first sound source angular distance registration coefficient ω(Δθ) = c is set. When Δθ = π, the second sound source angular distance registration coefficient ω(Δθ) = b is set. Here, b and c are real numbers, ranging from (0,1).
[0044] Let the first part of the piecewise function be s1 = k1Δθ 2 +1, substitute the calibrated point coordinates (0,1) and Find the slope k1 of the first part of the piecewise function: The analytical expression of the first part of the piecewise function is:
[0045] Let the second part of the piecewise function be s² = k²Δθ. 2 +m, substitute the calibrated point coordinates (π, b) and Find the slope k2 of the second part of the piecewise function: constant term The analytical expression of the second part of the piecewise function is:
[0046] Therefore, the expression for the sound source angular distance registration coefficient mapping relationship curve is obtained as follows:
[0047]
[0048] Step S4: Based on the mapping relationship curve between the sound source angular distance and the fitted sound source angular distance registration coefficient, obtain the speech signal received by each acoustic sensor in the multi-channel microphone array in the desired direction.
[0049] Specifically, the numerical value Δθ of the sound source angle distance obtained in step S2 is substituted into the fitted sound source angle distance registration coefficient mapping curve to be assigned to the speech signal received by the multi-channel microphone array.
[0050] The smaller the L2 norm of the vector difference obtained from the angular distance between two sound sources, the greater its impact on effective speech separation. To enhance the speech signal in the desired direction and suppress the speech signal in the interference direction, we obtain the received signal of the m-th acoustic sensor in the desired direction of the multi-channel microphone array under the sound source angular distance registration coefficient mapping curve. It can be represented as:
[0051]
[0052] In the formula, S q(t) represents the speech signal of the q-th interfering speaker, q = 1, 2, ..., c, where c is the number of interfering speakers, c = N-1, S exp (t) represents the speech signal of the target speaker in the desired direction, V m (t) represents the noise received by the m-th acoustic sensor, where M is the number of acoustic sensors in the multi-channel microphone array.
[0053] Step S5: Input the speech signals received by each acoustic sensor in the multi-channel microphone array in the desired direction into the pre-trained speech separation network to separate the single-channel speech signal of the target speaker.
[0054] Specifically, the speech signals received by each acoustic sensor in the multi-channel microphone array in the desired direction are used as the speech signals to be separated. The speech signals to be separated are then windowed. In this example, the length of each window is set to a fixed duration (2 seconds) multiplied by the audio signal sampling frequency. Each window does not overlap, and if the length is insufficient at the end of the signal, it is padded with zeros.
[0055] Furthermore, the speech separation network is a multi-channel, multi-layer convolutional neural network (MMConv-Net), and the structure and post-processing steps of the speech separation network are as follows: Figure 4 As shown. Specifically, the input data of the speech separation network is a three-dimensional tensor of shape (M, Length, W), where M is the number of acoustic sensors in the multi-channel array, Length is the length of the window, and W is the number of windows.
[0056] The speech separation network consists of a first Long Short-Term Memory (LSTM) network layer, a second LSTM network layer, and a fully connected layer connected sequentially. The first LSTM network layer contains 3200 hidden units, initialized with a shape of (M, 3200). The second LSTM network layer contains 800 hidden units, initialized with a shape of (M, 800). The output signal shape of the fully connected layer is the same as the input data shape of the network. Finally, the output signal of the fully connected layer is averaged along the dimension of the number of acoustic sensors (M) in the array to obtain a single-channel output signal. The W window signals are then sequentially concatenated to obtain the final separated single-channel speech signal of the target speaker.
[0057] Furthermore, the training process of the speech separation network includes: acquiring training data, inputting the training dataset into the speech separation network (MMConv-Net) for training. In this embodiment, the training iterations are 200 times, and the root mean square backpropagation algorithm optimizer is used. A loss function is set, and when the training loss value remains stable and gradually decreases and converges after more than 100 iterations, the speech separation network training is complete, and the trained speech separation network is saved.
[0058] The loss function of the speech separation network is the average of the sum of squares of the training loss functions of the speech signals received by M sensors in the multi-channel microphone array, as expressed below:
[0059]
[0060] In the formula, L i Let i be the training loss function for the speech signal received by the i-th sensor in the multi-channel microphone array, where i = 1, 2, ..., M.
[0061] The training loss function for the speech signal received by the i-th sensor in a multi-channel microphone array is set based on the scale-invariant signal-to-noise ratio, and its expression is as follows:
[0062]
[0063] Among them, S m (t) target The expression is as follows:
[0064]
[0065] e noise The expression is as follows:
[0066]
[0067] In the above formula, S exp (t) represents the label corresponding to the pure single-person speech signal in the desired direction, i.e., the training data. It is the separation estimation result of the single-person speech signal in the desired direction obtained through the speech separation network during the training process.
[0068] The combined signal power of the single-speaker speech signal estimation result obtained through the speech separation network and the training data labels is ||S exp (t)|| 2 S represents the signal power of the label corresponding to the training data. m (t) target This is the ratio of the power of the two signals mentioned above, e noise This is the difference between the estimated single-person speech signal obtained by the speech separation network and the power ratio of this signal.
[0069] Furthermore, Figure 5 The diagram shows the time-domain waveform of the mixed speech signal from one of the single channels of the microphone array. Figure 6 The time-domain waveform of the single-channel speech signal of the target speaker after separation is shown. Figure 7 The time-domain waveform of the standard reference speech signal for this speaker is shown. Figure 6 and Figure 7 As can be seen from this, the microphone array multi-speaker speech separation system proposed in this invention can separate mixed speech signals and extract the speech signal of a single target speaker.
[0070] Preferably, in the presence of multiple sound sources in the environment, the multi-channel array multi-speaker speech separation method provided by this invention can automatically calculate the angular distances of all existing sound sources and sequentially separate and enhance each sound source. Based on the speech separation effect, by fine-tuning the second sound source angular distance registration coefficient b and the first sound source angular distance registration coefficient c, the mapping relationship curve of the sound source angular distance registration coefficient is adjusted, which can separate the speech signal of the target speaker applicable to more channels. Verification shows that the multi-channel array multi-speaker speech separation method provided in this example can separate up to 5 speakers speaking simultaneously.
[0071] Preferably, the microphone array verified in the examples of the present invention can be a circular array or a linear array; the circular array has a radius of 10 cm and can hold up to 6 acoustic sensors; the linear array can hold up to 4 acoustic sensors with a spacing of 80 mm.
[0072] In summary, this invention provides a multi-channel array multi-speaker speech separation method. This method can process each microphone channel, achieves high positioning accuracy, and demonstrates significant effectiveness in complex acoustic environments. By introducing sound source angle-distance registration coefficients into the mixed speech signals acquired by the multi-channel array, the mixed speech signals from the microphone array after sound source angle-distance registration control are input into the speech separation network MMConv-Net to obtain the separated single-channel speech signals of the target speaker. This provides a convenient method for separating mixed speech signals from multiple speakers acquired by multi-channel microphone arrays and effectively enhances the recognition of single-speaker speech content. Furthermore, the multi-channel multi-speaker speech separation system proposed in this invention does not rely on a specific target speaker and can handle a variable number of speakers, up to five speakers.
[0073] This specification also provides a computer-readable storage medium storing a computer program that can be used to perform the above-described data synchronization method.
[0074] This instruction manual also provides Figure 8 The diagram shows a schematic structural representation of the electronic device. Figure 8 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for the business operations. The processor reads the corresponding computer program from the non-volatile storage into memory and then runs it to achieve the aforementioned data synchronization method.
[0075] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0076] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0077] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0078] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0079] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.
[0080] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0081] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0082] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0083] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0084] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0085] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0086] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0087] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0088] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0089] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0090] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0091] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A multi-channel array multi-speaker speech separation method, characterized in that, The method includes: Acquire mixed speech signals from multiple speakers using a multi-channel microphone array; The desired sound source and other interfering sound sources are located in the mixed speech signal of multiple speakers in a multi-channel microphone array. The direction angle of the desired sound source and the direction angle of the interfering sound source are obtained. The vector difference between the direction angle of the desired sound source and the direction angle of the interfering sound source is used as the angular distance of the sound source. Calibrate the sound source angle distance registration coefficient, and fit the sound source angle distance registration coefficient mapping relationship curve based on the sound source angle distance and the sound source angle distance registration coefficient; Based on the mapping relationship curve between the sound source angular distance and the fitted sound source angular distance registration coefficient, the speech signal received by each acoustic sensor in the multi-channel microphone array in the desired direction is obtained; including: The received signal of the m-th acoustic sensor in the desired direction in a multi-channel microphone array Represented as: ; In the formula, This represents the speech signal of the q-th interfering speaker. c represents the number of interfering speakers, where c = N-1 and N is the total number of speakers. Indicates the angular distance registration coefficient of the sound source. The speech signal of the target speaker representing the desired direction. Let M be the noise received by the m-th acoustic sensor, where M is the number of acoustic sensors in the multi-channel microphone array. The speech signals received by each acoustic sensor in the multi-channel microphone array in the desired direction are input into a pre-trained speech separation network to separate the single-channel speech signal of the target speaker.
2. The multi-channel array multi-speaker speech separation method according to claim 1, characterized in that, The vector difference between the desired sound source direction angle and the interfering sound source direction angle is used as the sound source angular distance, including: Using the positive X-axis counterclockwise direction as a reference, calculate the desired sound source direction angle. Angle relative to the direction of the interfering sound source The difference is denoted as The magnitude of the angular distance from the sound source. The expression is: when When the angle is less than 180 degrees, the angular distance of the sound source ; when When the angle is greater than 180 degrees, the angular distance of the sound source .
3. The multi-channel array multi-speaker speech separation method according to claim 1, characterized in that, The sound source angular distance registration coefficient is calibrated, and the sound source angular distance registration coefficient mapping relationship curve is fitted based on the sound source angular distance and the sound source angular distance registration coefficient, including: When the angle of the sound source is far At that time, the registration coefficient of the sound source angle distance ; When the angle of the sound source is far When, set The corresponding first sound source angular distance registration coefficient According to coordinates and Fitting the angular distance of the sound source The curve showing the mapping relationship between the sound source angle and distance registration coefficients at that time; When the angle of the sound source is far When, set Corresponding second sound source angular distance registration coefficient According to coordinates and Fitting the angular distance of the sound source The curve showing the mapping relationship between the sound source angle and the registration coefficient at that time.
4. The multi-channel array multi-speaker speech separation method according to claim 1 or 3, characterized in that, The expression for the sound source angular distance registration coefficient mapping curve is: 。 5. The multi-channel array multi-speaker speech separation method according to claim 1, characterized in that, The speech signals received by each acoustic sensor in the multi-channel microphone array in the desired direction are input into a pre-trained speech separation network to separate the single-channel speech signal of the target speaker, including: The speech signals received by each acoustic sensor in the multi-channel microphone array in the desired direction are used as the speech signals to be separated. The speech signals to be separated are then windowed. The size of each window is the product of the preset duration and the sampling frequency of the audio signal. Each window does not overlap. If the length of the separated speech signal is insufficient at the end, it is padded with zeros. The speech signal to be separated is input into a pre-trained speech separation network. The output signal of the speech separation network is averaged along the dimension of the number of acoustic sensors to obtain a single-channel output signal. Finally, the speech signals corresponding to each window are spliced in sequence to separate the single-channel speech signal of the target speaker.
6. The multi-channel array multi-speaker speech separation method according to claim 1 or 5, characterized in that, The training process of the speech separation network includes: Acquire training data, input the training dataset into the speech separation network for training, set the number of iterations, use the root mean square backpropagation algorithm optimizer, set the loss function, and when the training loss value converges, save the trained speech separation network. The expression for the loss function is as follows: ; In the formula, Let be the training loss function for the speech signal received by the i-th sensor in the multi-channel microphone array. M is the number of acoustic sensors in the multi-channel microphone array.
7. The multi-channel array multi-speaker speech separation method according to claim 6, characterized in that, The expression for the training loss function of the speech signal received by the i-th sensor in a multi-channel microphone array is: ; ; ; In the formula, These are the labels corresponding to the training data. It is the separation estimation result of the single-person speech signal in the desired direction obtained through the speech separation network during the training process. The power of the combined signal of the single-speaker speech signal estimation result obtained through the speech separation network and the training data labels is given. The signal power of the label corresponding to the training data. This is the ratio of the combined signal power of the single-person speech signal estimation result obtained through the speech separation network and the training data labels to the signal power of the corresponding labels in the training data. This is the difference between the estimated single-person speech signal obtained by the speech separation network and the signal power ratio.
8. An electronic device comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the multi-channel array multi-speaker speech separation method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the multi-channel array multi-speaker speech separation method as described in any one of claims 1-7.