Speech Separation Method, System and Electronic Device Based on Multi-Speaker Speech Detection
By performing overlapping speech detection and guided speech separation on mixed speech, traditional speech separation technology solves the problems of high computational cost and low separation performance in real scenarios, and efficient speech separation with low computational volume is achieved.
Patent Information
- Application Number
- CN202211695785.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-12-28
AI Technical Summary
Traditional speech separation technology has high calculation cost when separating long speech in real scenes, low separation performance, and cannot ensure that the separated speech is arranged in a certain order.
By performing overlapping speech detection on mixed speech, the time interval between the multi-speaker overlapping speech segment and the single-speaker non-speaker non-speaker segment is determined, and the adjacent single-speaker non-speaker segment is input into the guided speech separation model as auxiliary speech segments, and only the multi-speaker overlapping speech segment is processed for separation.
The calculation amount of speech separation is reduced, the speech separation performance is improved, the arrangement problem is solved, and the accuracy and efficiency of speech separation is improved.
Smart Images

Figure CN115938386B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent voice, and particularly to a voice separation method, system and electronic device based on multi-speaker voice detection. Background Art
[0002] With the development of intelligent voice technology, voice separation plays an important role in the front-end processing of voice. Among them, voice separation refers to the technology of separating voices mixed by multiple sound sources. Traditional voice separation technology has made great progress in short conversations with a high overlap ratio. However, conversations in real scenarios usually partially overlap with a relatively low overlap rate, and in some scenarios of real scenarios, the conversations are usually relatively long (such as in a meeting scenario), and the effect of applying traditional voice separation tasks to voice separation in real scenarios will be reduced.
[0003] Traditional continuous voice separation technology will cut long-time voice into smaller paragraphs for separate processing, and connect the separated voices back to non-overlapping long voices through a splicing algorithm. Among them, long voice refers to continuous long-time voice, which contains complex scenarios of multiple speakers, multiple people speaking simultaneously, and the presence of noise.
[0004] In the process of implementing the present invention, the inventors found that there are at least the following problems in the related art:
[0005] Due to the permutation problem in traditional voice separation technology, that is, there are multiple permutation orders among multiple sound sources obtained by separating mixed voices, and it is impossible to ensure that the separated voices are arranged in a certain order. Therefore, it is necessary to use PIT (Permutation Invariant Training) for the voice separation model. At the same time, since the long voice is divided into independent paragraphs, in order to ensure the consistency of the voices in the front and back paragraphs, it is also necessary to use a splicing algorithm to splice the separated voices, resulting in a relatively high computational cost and relatively low voice separation performance. Summary of the Invention
[0006] In order to at least solve the problems of relatively high computational cost and relatively low separation performance in separating long voices in real scenarios in the prior art. In a first aspect, an embodiment of the present invention provides a voice separation method based on multi-speaker voice detection, including:
[0007] Performing overlapping voice detection on a mixed voice containing multiple speakers to obtain the respective time intervals corresponding to multi-speaker overlapping voice segments and single-speaker non-overlapping voice segments in the mixed voice;
[0008] Determine the single-speaker non-overlapping speech segments adjacent to the time interval of the multi-speaker overlapping speech segments as the auxiliary speech segments corresponding to the multi-speaker overlapping speech segments, and input the multi-speaker overlapping speech segments and the corresponding auxiliary speech segments into a guided speech separation model;
[0009] Multiple non-overlapping speeches separated by using the guided speech separation model.
[0010] In a second aspect, an embodiment of the present invention provides a speech separation system based on multi-speaker speech detection, including:
[0011] A speech overlap detection program module for performing overlap speech detection on a mixed speech containing multiple speakers to obtain the time intervals corresponding to the multi-speaker overlapping speech segments and the single-speaker non-overlapping speech segments in the mixed speech respectively;
[0012] An auxiliary speech determination program module for determining the single-speaker non-overlapping speech segments adjacent to the time interval of the multi-speaker overlapping speech segments as the auxiliary speech segments corresponding to the multi-speaker overlapping speech segments, and inputting the multi-speaker overlapping speech segments and the corresponding auxiliary speech segments into a guided speech separation model;
[0013] A speech separation program module for separating multiple non-overlapping speeches by using the guided speech separation model.
[0014] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the speech separation method based on multi-speaker speech detection according to any embodiment of the present invention.
[0015] In a fourth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the speech separation method based on multi-speaker speech detection according to any embodiment of the present invention are implemented.
[0016] The beneficial effects of the embodiments of the present invention are as follows: Determine the time intervals of the multi-person speech segments and the single-person speech segments in the mixed speech of multiple speakers, use the single-person speech segments as additional inputs to the guided speech separation model to assist in long speech separation and solve the permutation problem. The guided speech separation model only processes multi-person speech, and on the basis of relatively low speech separation computation, the performance of speech separation is improved. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0018] Figure 1 is a flowchart of a voice separation method based on multi-speaker voice detection provided by an embodiment of the present invention;
[0019] Figure 2 is a schematic diagram of a continuous voice separation framework of a voice separation method based on multi-speaker voice detection provided by an embodiment of the present invention;
[0020] Figure 3 is a schematic diagram of the proportion of window blocks with different overlap ratios of a voice separation method based on multi-speaker voice detection provided by an embodiment of the present invention;
[0021] Figure 4 is a schematic diagram of the evaluation of an overlapping voice detection model of a voice separation method based on multi-speaker voice detection provided by an embodiment of the present invention;
[0022] Figure 5 is a schematic diagram of model data of different splicing methods of a voice separation method based on multi-speaker voice detection provided by an embodiment of the present invention;
[0023] Figure 6 is a schematic diagram of data for predicting overlapping voice information of a voice separation method based on multi-speaker voice detection provided by an embodiment of the present invention;
[0024] Figure 7 is a schematic diagram of the word error rate evaluation of the continuous voice separation of LibriCSS by a voice separation method based on multi-speaker voice detection provided by an embodiment of the present invention;
[0025] Figure 8 is a schematic diagram of the structure of a voice separation system based on multi-speaker voice detection provided by an embodiment of the present invention;
[0026] Figure 9 is a schematic diagram of the structure of an embodiment of an electronic device for voice separation based on multi-speaker voice detection provided by an embodiment of the present invention. Detailed implementation manners
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] As Figure 1 shown in the flowchart of a voice separation method based on multi-speaker voice detection provided by an embodiment of the present invention, the method includes the following steps:
[0029] S11: Perform overlapping speech detection on the mixed speech containing multiple speakers to obtain the respective time intervals of the multi-speaker overlapping speech segments and the single-speaker non-overlapping speech segments in the mixed speech;
[0030] S12: Determine the single-speaker non-overlapping speech segment adjacent to the time interval of the multi-speaker overlapping speech segment as the auxiliary speech segment corresponding to the multi-speaker overlapping speech segment, and input the multi-speaker overlapping speech segment and the corresponding auxiliary speech segment into the guided voice separation model;
[0031] S13: Use the guided voice separation model to separate multiple non-overlapping voices.
[0032] In this embodiment, considering the problem of large computational complexity in traditional long-duration voice separation, a new method for processing long voices is proposed to reduce the computational complexity of voice separation and the computational power requirements for intelligent devices equipped with this method. Generally speaking, this method consists of two parts: overlapping speech detection and guided voice separation.
[0033] For step S11, an intelligent device equipped with this method receives the mixed speech containing multiple speakers. Following the basic assumption in CSS (Continuous speech separation), that is, at most C speakers speak simultaneously, where C is also the number of output channels. For example, there are three speakers (A, B, and C) in a room. At this time, in a time period, only two speakers speak simultaneously (A + B, A + C, B + C), then C = 2 at this time, and the output channels of the separated voices are 2, that is, two separated voices will be generated.
[0034] As an implementation, the performing overlapping speech detection on the mixed speech containing multiple speakers includes:
[0035] Performing overlapping speech detection on the mixed speech containing multiple speakers using a frame-level multi-person voice detection model.
[0036] In this embodiment, an OSD (Overlapping Speech Detection) model can be specifically used to identify overlapping and non-overlapping segments. Specifically, it is essentially a binary classifier that can determine the multi-speaker overlapping speech segments and single-speaker non-overlapping speech segments in the mixed speech with low computing power and high efficiency, and also knows the respective time intervals corresponding to the multi-speaker overlapping speech segments and single-speaker non-overlapping speech segments in the mixed speech in the mixed speech.
[0037] As for why the overlapping and non-overlapping segments in the mixed speech need to be separated, because this method finds that each non-overlapping speech segment and its subsequent overlapping segment next to it may contain the same speaker. For example, in the first scenario: during the process of A speaking to B, C's interest is suddenly aroused, and C interrupts. At the data level, assuming that in the speech from the 5th second to the 10th second, there is only the non-overlapping speech segment of single speaker A. At the 10th second, when C suddenly inserts, there is a multi-speaker overlapping speech segment where A and C are speaking simultaneously from the 10th second to the 13th second. In another scenario, in daily life, A asks a question to B and C. After a little thought, B and C answer simultaneously, but B is more talkative and is still answering when C finishes. At the data level, A asks a question, B and C think for a while and enter the next new long speech. During the process from the 0th second to the 7th second, there is a multi-speaker overlapping speech segment where B and C are speaking simultaneously. From the 7th second to the 10th second, there is only the non-overlapping speech segment of single speaker B.
[0038] In the first scenario, the speakers in the multi-speaker overlapping speech segment after the time interval of the single-speaker non-overlapping speech segment (the time interval from the 10th second to the 13th second after the 5th - 10th second) all include A; while in the second scenario, the speaker in the single-speaker non-overlapping speech segment after the time interval of the multi-speaker overlapping speech segment (the time interval from the 7th second to the 10th second after the 0th - 7th second) all includes B.
[0039] At the same time, this method further considers the long speech in the real scenario. As an implementation, after obtaining the respective time intervals corresponding to the multi-speaker overlapping speech segments and single-speaker non-overlapping speech segments in the mixed speech, the method further includes:
[0040] When the speech length of the mixed speech exceeds a preset length, the mixed speech of multiple speakers is windowed to obtain the mixed speech segments of each window.
[0041] Considering the voice separation of long voices, traditional CSS usually additionally uses PIT (Permutation Invariant Training). Based on the idea that each non-overlapping voice segment and its subsequent overlapping segment beside it may contain the same speaker, this method obtains a deterministic permutation order by adjusting the adjacent non-overlapping segments before the current overlapping segment, so that the shared speaker is always separated in the first output channel. For example, in the voice separation of the first scenario mentioned above, the voice of A will be separated in the first output channel, and the voice of C will be separated in the second output channel. In the voice separation of the second scenario mentioned, the voice of B will be separated in the first output channel, and the voice of C will be separated in the second output channel. In this way, additional calculations can be avoided, and at the same time, the efficiency of accurate voice splicing can be guaranteed after the long voice is separated.
[0042] For step S12, in order to further reduce the computational complexity of the voice separation of long voices, this method constructs a guided voice separation model. Different from the prior art that directly separates the mixed voice containing multiple speakers, the guided voice separation model of this method is the auxiliary voice segment corresponding to the overlapping voice segments of multiple speakers. That is to say, this method only needs to separate the overlapping voice segments of multiple speakers, and during the separation process, the auxiliary voice segment is also used as the auxiliary of the guided voice separation model.
[0043] As an implementation manner, within the mixed voice segment, the single-speaker non-overlapping voice segment adjacent to the time interval of the overlapping voice segments of multiple speakers before is selected and determined as the auxiliary voice segment corresponding to the overlapping voice segments of multiple speakers.
[0044] In this implementation manner, this judgment mode corresponds to the first scenario mentioned above. In the voice from the 5th second to the 10th second, only the single-speaker non-overlapping voice segment of A is used as the auxiliary voice segment (for the overlapping voice segments of multiple speakers of A and C from the 10th second to the 13th second), and the overlapping voice segments of multiple speakers and the corresponding auxiliary voice segments are input into the guided voice separation model.
[0045] As another implementation manner, within the mixed voice segment, when there is no adjacent single-speaker non-overlapping voice segment before the time interval of the overlapping voice segments of multiple speakers, the single-speaker non-overlapping voice segment adjacent to the time interval of the overlapping voice segments of multiple speakers after is selected and determined as the auxiliary voice segment corresponding to the overlapping voice segments of multiple speakers.
[0046] In this embodiment, the judgment mode corresponds to the second scenario exemplified above. Among the voices from the 7th second to the 10th second, only the non-overlapping speech segment of speaker B is used as the auxiliary speech segment for the multi-speaker overlapping speech segment (during the process from the 0th second to the 7th second, there is overlapping speech of speakers B and C), and the multi-speaker overlapping speech segment and the corresponding auxiliary speech segment are input into the guided speech separation model.
[0047] Generally speaking, considering the windowing of long speech to obtain the mixed speech segments of each window, this method uses the term "block" to represent the model input (multi-speaker overlapping speech segment and the corresponding auxiliary speech segment) selected by windowing from the long-form speech through a sliding window, and this window may contain overlapping and non-overlapping "segments". After identifying the non-overlapping single-speaker segments through the OSD model, this method constructs a guided speech separation model without permutation problems. The proposed model takes the mixed signal X of the multi-speaker overlapping speech segment and the adjacent non-overlapping speech segment C of a single speaker a as inputs and generates the corresponding separated speech. As mentioned above, this method ensures that the speaker in the first output channel is the same as the speaker in the single-speaker condition Ca, and the output permutation is deterministic. The structure of the frequency-domain SpeakerBeam is used as the guided speech separation model of this method, which includes a mixture encoder Enc Mix (·) for processing the mixture, an auxiliary network AuxNet(·) for processing the conditional segment, and a separator for generating the output. In order to generate multiple outputs, this method further increases the number of output heads extracting the output of the last layer of the model. The training process can be formulated as follows:
[0048]
[0049]
[0050] where X and C a are the input mixed speaker and the corresponding condition of speaker a respectively. and represent the separated speech of speaker a and speaker b respectively, while S a and S b are the corresponding reference signals. L is the loss function.
[0051] As an embodiment, the guided speech separation model includes: a speech encoder, an auxiliary network, and a separation module, where
[0052] the speech encoder is used to determine the multi-speaker hidden layer feature encoding of the multi-speaker overlapping speech segment;
[0053] The auxiliary network is used to determine the single-speaker hidden layer feature encoding of the auxiliary speech segment, where the single speaker corresponding to the auxiliary speech segment is among the multiple speakers corresponding to the overlapping speech segment;
[0054] The separation module is used to separate multiple non-overlapping voices according to the multiple-speaker hidden layer feature encoding and the single-speaker hidden layer feature encoding.
[0055] As Figure 2 shown, it is the overall structure diagram of this method, which includes speech detection, audio selection, windowing, selection of single speakers as assistance, Fourier transform, construction of the model, etc. In the windowing process of long speech, for each windowed block, the corresponding single-speaker segment is selected as a condition according to the rules listed above. If there is an adjacent single-speaker segment before the block, it is used as the auxiliary speech segment. When no adjacent single-speaker segment is found before the block, the non-overlapping segment of the last block (if any) is tried to be used as the auxiliary speech segment. If none of them are available, it is also possible to fallback to using the non-overlapping segment at the beginning of the current block as a condition.
[0056] For step S13, in the guided speech separation model, if the input is not long speech and not segmented, multiple non-overlapping voices can be directly separated. With the help of the OSD model, a new overlap-aware splicing method can be further proposed. In the inference stage of the guided speech separation model, the separation model only processes the multiple-speaker overlapping speech segment. Then, only the non-overlapping voices after separating the multiple-speaker overlapping speech segment are spliced with the single-speaker speech that is originally non-overlapping. Since non-overlapping segments are used as additional conditions, the permutation problem in the splicing process can also be avoided.
[0057] It can be seen from this implementation that by determining the time intervals of the multi-speaker speech segment and the single-speaker speech segment in the mixed speech of multiple speakers, using the single-speaker speech segment as an additional input to the guided speech separation model to assist in long speech separation and solving the permutation problem, the guided speech separation model only processes multi-speaker speech, and on the basis of relatively low speech separation computational complexity, the performance of speech separation is improved.
[0058] An experiment on this method is described. This method is trained using simulated reverberant meeting-style data based on LibriSpeech, and the generated data does not contain noise. The simulated data contains 30,000, 900, and 900 samples for training, development, and evaluation respectively. Each sample is 90 seconds long, contains 3 - 5 speakers, and the window overlap rate ranges from 50% to 80%. The reverberation time ranges from 100 ms to 500 ms. Additionally, the LibriCSS dataset is used to evaluate the separation performance in a real-world scenario, which contains 10 hours of audio recordings in a regular room. Each session in LibriCSS contains 8 speakers, and the overlap ratio varies from 0 to 40%. The recordings are first processed by the separation model. The separated signals are then processed by a hybrid automatic speech recognition (ASR) model to generate the corresponding transcripts.
[0059] This method uses a T-F (Time Frequency) masking method for speech separation. The size of the STFT (short-time Fourier transform) is 512 points, and the hop length is 256. The sliding window size of CSS is 3.2 seconds. All single-speaker conditions are padded or shortened to 1 second. The proportions of window blocks with different overlap rates in the simulated data are as Figure 3 shown, and all experiments are conducted using the ESPnet toolkit.
[0060] For the RNN (Recurrent Neural Network) model, the structure of the frequency-domain SpeakerBeam is adopted in the guided speech separation model. Both the baseline and the proposed models have 5 BLSTM (Bidirectional Long Short-Term Memory) blocks. Each block consists of a BLSTM layer, a linear projection layer, a global LayerNorm layer, and an activation layer with residual connections. The input size of the first block is 257, and the input of the remaining blocks is 256. The hidden dimension of the BLSTM block is 515 for the baseline model and 512 for the guided separation model. The auxiliary network of the guided separation model has the same configuration as that in the frequency-domain SpeakerBeam. As a result, the baseline model has 23.05M parameters, and the guided separation model has 23.01M parameters. We set the initial learning rate to 1e -4 and use the StepLR scheduler, where the learning rate decays by a factor of 0.98 every two epochs.
[0061] For the Conformer-based model, this configuration has 16 Conformer encoder layers with 4 attention heads, 256 attention dimensions, and 1024 FFN (feed forward network) dimensions. For the guided separation model, a cross-attention consistency structure is adopted. It consists of two independent consistency encoders for processing the input mixture and the condition respectively, followed by 8 cross-attention consistency blocks to obtain the separation mask. Each consistency encoder includes 8 consistency layers with 4 attention heads, 228 attention dimensions, and 512 FFN dimensions. Similarly, each cross-attention suppressor block consists of cross-attention consistency layers with 4 attention heads, 228 attention dimensions, and 512 FFN dimensions. As a result, the baseline model has 21.54M parameters and the guided separation model has 21.43M parameters. The learning rate is set to 2e -4 , and a warm-up learning rate scheduler with 20,000 warm-up steps is used.
[0062] During training, each mini-batch includes 8 long samples. All separation models are trained for 150 epochs using the Adam optimizer, and the early stopping period is set to 10. The OSD model parameters are 888.58K. The learning rate is 1e-3 and the batch size is 1.
[0063] As Figure 4 shows the performance of the OSD model. Since the model is not completely accurate in predicting overlapping segments and there may be some extremely short overlapping segments, this method performs a similar dilation-erosion post-processing smoothing strategy on the predicted overlapping segments with the kernel size set to 5 frames. In this way, extremely short overlapping segments can be eliminated.
[0064] As Figure 5 shows the results of the baseline model and the guided separation model of this method on the simulated dataset. Different splicing strategies when processing long-form speech are compared, where splicing windows with an overlap rate of 0% or 50% can be used. After generating the long-form separation signal, the long-form signal is divided into blocks using a 3.2s sliding window to calculate the window-level SNR (SIGNAL-NOISE RATIO). The average SNR is calculated based on the speaker overlap rate in each block.
[0065] The results show that both the guided BLSTM and the guided consistency model of this method outperform their baseline models, with an SNR improvement of more than 1 dB. In addition, when the traditional block-by-block splicing method is applied with 0% window overlap, the model of this method can still achieve a strong performance comparable to or even better than the baseline performance when using 50% window overlap. In addition, the effectiveness of the proposed overlap-aware splicing method is evaluated. It can be seen that the computational cost is greatly reduced while the overall performance only decreases slightly. In blocks with a low overlap rate (0–25%), the SNR performance is even better than the block-by-block splicing method, especially for the BLSTM-based model. This may be attributed to the energy leakage problem when processing almost single-speaker signals using the speech separation model. In addition, it is also observed that the proposed splicing method is insensitive to the window overlap rate, which can further reduce the computation.
[0066] As Figure 6 shows the performance of the guided separation model when using the predicted overlap information as a condition. Although better performance can be obtained by using oracle overlap information, it can be seen that in most cases, the SNR performance gap is less than 0.3 dB with a high overlap rate (>25%). This also demonstrates the performance of this method.
[0067] As Figure 7 shows the results of the baseline and the model proposed in this method on LibriCSS. In the proposed overlap-aware splicing method, this method only processes the overlapping speech segments and finds that all the overlapping segments predicted by the OSD model are shorter than the window length (3.2 s). Therefore, this method only needs one window to cover each overlapping segment. The window overlap rate of this method is always 0%. It can also be observed that the model of this method is superior to the baseline model when using a 50% window overlap rate in the block-by-block splicing method or applying the overlap-aware splicing method. Compared with the baseline model, the best performance of the model of this method achieves an absolute average WER (Word error rate) reduction of ~1%.
[0068] Overall, the framework proposed in this method trains the guided speech separation model by providing single-speaker segments to help the model determine the arrangement of the output, thus getting rid of the PIT method. In addition, an overlap-aware inference algorithm is introduced to generate the separated long speech with the help of the OSD model. The experimental results show that the framework of this method is superior to the traditional splicing CSS method, with an SNR improvement of more than 1 dB and a relative word error rate reduction of >4%. In addition, this method greatly reduces the computational cost and the model can maintain a strong performance.
[0069] As Figure 8The following is a schematic structural diagram of a voice separation system based on multi-speaker voice detection provided by an embodiment of the present invention. This system can execute the voice separation method based on multi-speaker voice detection described in any of the above embodiments and is configured in a terminal.
[0070] A voice separation system 10 based on multi-speaker voice detection provided in this embodiment includes: a voice overlap detection program module 11, an auxiliary voice determination program module 12, and a voice separation program module 13.
[0071] Among them, the voice overlap detection program module 11 is used to perform overlap voice detection on the mixed voice containing multiple speakers to obtain the respective time intervals of the multi-speaker overlap voice segments and the single-speaker non-overlap voice segments in the mixed voice; the auxiliary voice determination program module 12 is used to determine the single-speaker non-overlap voice segments adjacent to the time interval of the multi-speaker overlap voice segments as the auxiliary voice segments corresponding to the multi-speaker overlap voice segments, and input the multi-speaker overlap voice segments and the corresponding auxiliary voice segments into the guided voice separation model; the voice separation program module 13 is used to separate multiple non-overlapping voices by using the guided voice separation model.
[0072] The embodiment of the present invention also provides a non-volatile computer storage medium. The computer storage medium stores computer-executable instructions, and these computer-executable instructions can execute the voice separation method based on multi-speaker voice detection in any of the above method embodiments;
[0073] As an implementation manner, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as:
[0074] Perform overlap voice detection on the mixed voice containing multiple speakers to obtain the respective time intervals of the multi-speaker overlap voice segments and the single-speaker non-overlap voice segments in the mixed voice;
[0075] Determine the single-speaker non-overlap voice segments adjacent to the time interval of the multi-speaker overlap voice segments as the auxiliary voice segments corresponding to the multi-speaker overlap voice segments, and input the multi-speaker overlap voice segments and the corresponding auxiliary voice segments into the guided voice separation model;
[0076] Separate multiple non-overlapping voices by using the guided voice separation model.
[0077] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method in the embodiments of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium, and when executed by a processor, they perform the voice separation method based on multi-speaker voice detection in any of the above method embodiments.
[0078] Figure 9 FIG. is a schematic hardware structure diagram of an electronic device for the voice separation method based on multi-speaker voice detection provided in another embodiment of the present application, as Figure 9 shown, the device includes:
[0079] One or more processors 910 and a memory 920, Figure 9 Taking one processor 910 as an example. The device for the voice separation method based on multi-speaker voice detection may further include: an input device 930 and an output device 940.
[0080] The processor 910, the memory 920, the input device 930, and the output device 940 may be connected through a bus or other means, Figure 9 Taking the connection through a bus as an example.
[0081] The memory 920, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the voice separation method based on multi-speaker voice detection in the embodiments of the present application. The processor 910 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 920, that is, implements the voice separation method based on multi-speaker voice detection in the above method embodiments.
[0082] The memory 920 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data, etc. In addition, the memory 920 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 920 may optionally include a memory remotely set relative to the processor 910, and these remote memories may be connected to the mobile device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0083] The input device 930 can receive input digital or character information. The output device 940 may include a display device such as a display screen.
[0084] The one or more modules are stored in the memory 920 and, when executed by the one or more processors 910, perform the voice separation method based on multi-speaker voice detection in any of the above method embodiments.
[0085] The above product can execute the method provided in the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference may be made to the method provided in the embodiments of the present application.
[0086] The non-volatile computer-readable storage medium may include a storage program area and a storage data area. Among them, the storage program area may store an operating system and application programs required for at least one function; the storage data area may store data created according to the use of the device. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely provided with respect to the processor, and these remote memories may be connected to the device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0087] An embodiment of the present invention further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the voice separation method based on multi-speaker voice detection in any embodiment of the present invention.
[0088] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to:
[0089] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.
[0090] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristics of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, such as tablet computers.
[0091] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.
[0092] (4) Other electronic devices with data processing functions.
[0093] In this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also other elements not explicitly listed, or also include elements inherent in such a process, method, article or device. Without more limitations, the elements defined by the statement "comprising..." do not exclude the existence of additional identical elements in the process, method, article or device comprising the said elements.
[0094] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A voice separation method based on multi - speaker voice detection, comprising: Performing overlapping voice detection on a mixed voice containing multiple speakers to obtain the respective time intervals in the mixed voice for multi - speaker overlapping voice segments and single - speaker non - overlapping voice segments; Determining the single - speaker non - overlapping voice segments adjacent to the time interval of the multi - speaker overlapping voice segments as the auxiliary voice segments corresponding to the multi - speaker overlapping voice segments, and inputting the multi - speaker overlapping voice segments and the corresponding auxiliary voice segments into a guided voice separation model; wherein the single speaker corresponding to the auxiliary voice segment is the same as at least one speaker in the multi - speaker overlapping voice segments; Using the guided voice separation model to separate multiple non - overlapping voices; The guided voice separation model only performs separation operations within the time interval corresponding to the multi - speaker overlapping voice segments.
2. The method according to claim 1, wherein, After obtaining the respective time intervals in the mixed voice for the multi - speaker overlapping voice segments and the single - speaker non - overlapping voice segments, the method further comprises: When the voice length of the mixed voice exceeds a preset length, performing windowing processing on the mixed voice of the multiple speakers to obtain mixed voice segments for each window; Respectively determining the single - speaker non - overlapping voice segments adjacent to the time interval of the multi - speaker overlapping voice segments in each window of the mixed voice segments as the auxiliary voice segments corresponding to the multi - speaker overlapping voice segments, and sequentially inputting the multi - speaker overlapping voice segments and the corresponding auxiliary voice segments in each window of the mixed voice segments into the guided voice separation model; Using the guided voice separation model to separate multiple non - overlapping voice segments for each window, and splicing the multiple non - overlapping voice segments into multiple non - overlapping voices according to the sorting of each window.
3. The method according to claim 2, wherein The step of respectively determining the single - speaker non - overlapping voice segments adjacent to the time interval of the multi - speaker overlapping voice segments in each window of the mixed voice segments as the auxiliary voice segments corresponding to the multi - speaker overlapping voice segments includes: In the mixed voice segment, selecting the single - speaker non - overlapping voice segment adjacent before the time interval of the multi - speaker overlapping voice segment as the auxiliary voice segment corresponding to the multi - speaker overlapping voice segment.
4. The method according to claim 3, wherein In the mixed voice segment, when there is no adjacent single - speaker non - overlapping voice segment before the time interval of the multi - speaker overlapping voice segment, selecting the single - speaker non - overlapping voice segment adjacent after the time interval of the multi - speaker overlapping voice segment as the auxiliary voice segment corresponding to the multi - speaker overlapping voice segment.
5. The method according to claim 1, wherein The guided voice separation model includes: a voice encoder, an auxiliary network, and a separation module, wherein, The voice encoder is used to determine the multi - speaker hidden layer feature encoding of the multi - speaker overlapping voice segments; The auxiliary network is used to determine the single - speaker hidden layer feature encoding of the auxiliary voice segments, wherein the single speaker corresponding to the auxiliary voice segment is among the multi - speakers corresponding to the overlapping voice segments; The separation module is used to separate multiple non - overlapping voices according to the multi - speaker hidden layer feature encoding and the single - speaker hidden layer feature encoding.
6. The method according to claim 1, wherein The overlapping speech detection for the mixed speech containing multiple speakers includes: Performing overlapping speech detection on the mixed speech containing multiple speakers by using a multi-speaker speech detection model at the frame level.
7. A voice separation system based on multi-speaker speech detection, comprising: A voice overlapping detection program module, configured to perform overlapping speech detection on the mixed speech containing multiple speakers, so as to obtain respective time intervals corresponding to multi-speaker overlapping speech segments and single-speaker non-overlapping speech segments in the mixed speech; An auxiliary speech determination program module, configured to determine a single-speaker non-overlapping speech segment adjacent to the time interval of the multi-speaker overlapping speech segment as an auxiliary speech segment corresponding to the multi-speaker overlapping speech segment, and input the multi-speaker overlapping speech segment and the corresponding auxiliary speech segment into a guided voice separation model; wherein the single speaker corresponding to the auxiliary speech segment is the same as at least one speaker in the multi-speaker overlapping speech segment; A voice separation program module, configured to use the guided voice separation model to separate multiple non-overlapping voices; The guided voice separation model only performs separation operations within the time interval corresponding to the multi-speaker overlapping speech segment.
8. The system according to claim 7, wherein, The guided voice separation model includes: a voice encoder, an auxiliary network, and a separation module, wherein, The voice encoder is configured to determine the multi-speaker hidden layer feature encoding of the multi-speaker overlapping speech segment; The auxiliary network is configured to determine the single-speaker hidden layer feature encoding of the auxiliary speech segment, wherein the single speaker corresponding to the auxiliary speech segment is among the multi-speakers corresponding to the overlapping speech segment; The separation module is configured to separate multiple non-overlapping voices according to the multi-speaker hidden layer feature encoding and the single-speaker hidden layer feature encoding.
9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the method according to any one of claims 1-6.
10. A storage medium, on which a computer program is stored, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
End-to-end multi-speaker overlapping speech recognition
CN115485768A