A method, apparatus, device and storage medium for hybrid speech separation
By acquiring the amplitude and phase information of the mixed speech and using the target speech separation model to determine the time-frequency mask, the problem of low speech recognition performance in the existing technology is solved, achieving higher-precision speech separation and a better user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AGRICULTURAL BANK OF CHINA
- Filing Date
- 2023-03-10
- Publication Date
- 2026-05-26
AI Technical Summary
Existing speech separation methods have low speech recognition performance when dealing with noise aliasing, which cannot meet practical needs.
By acquiring the amplitude and phase information of the mixed speech, the target speech separation model is used for separation processing to determine the time-frequency mask, and the human voice and noisy speech are separated based on the mask and phase information.
It improves the accuracy and smoothness of speech separation, enhancing the user experience.
Smart Images

Figure CN116230001B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech processing technology, and in particular to a method, apparatus, device, and storage medium for separating mixed speech. Background Technology
[0002] With social development and advancements in science and technology, intelligent voice systems can save manpower and resources and more conveniently and quickly separate and recognize speech signals. However, in real life, speech signals inevitably mix with surrounding noise, which greatly reduces the overall speech recognition performance of the system.
[0003] Currently used speech separation methods include those based on statistical modeling, those based on computational scene visual analysis, those based on blind source separation, and those based on deep learning. However, existing speech separation methods extract limited features, have low accuracy, and achieve poor separation results, failing to meet practical needs. Summary of the Invention
[0004] This invention provides a hybrid speech separation method, apparatus, device, and storage medium to improve speech separation accuracy, optimize speech separation smoothness, and enhance user experience.
[0005] According to one aspect of the present invention, a method for separating mixed speech is provided. The method includes:
[0006] Obtain the mixed speech to be separated;
[0007] The mixed speech to be separated is decomposed and transformed to obtain the mixed amplitude information and mixed phase information corresponding to the mixed speech to be separated.
[0008] The mixed amplitude information is input into the target speech separation model for speech separation processing, and the time-frequency mask corresponding to the mixed speech to be separated is determined based on the output of the target speech separation model.
[0009] Based on the time-frequency mask and the mixed speech to be separated, determine the human voice amplitude information and noise amplitude information;
[0010] Based on the mixed phase information, the human voice amplitude information, and the noise amplitude information, the target human voice speech and the target noise speech are determined.
[0011] According to another aspect of the present invention, a mixed speech separation device is provided. The device includes:
[0012] The mixed speech acquisition module is used to acquire the mixed speech to be separated;
[0013] The mixed information acquisition module is used to decompose and transform the mixed speech to be separated to obtain the mixed amplitude information and mixed phase information corresponding to the mixed speech to be separated.
[0014] The time-frequency mask acquisition module is used to input the mixed amplitude information into the target speech separation model for speech separation processing, and determine the time-frequency mask corresponding to the mixed speech to be separated based on the output of the target speech separation model;
[0015] The amplitude information determination module is used to determine the human voice amplitude information and noise amplitude information based on the time-frequency mask and the mixed speech to be separated;
[0016] The target speech determination module is used to determine the target human voice speech and the target noise speech based on the mixed phase information, the human voice amplitude information and the noise amplitude information.
[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the mixed speech separation method according to any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the mixed speech separation method according to any embodiment of the present invention.
[0022] The technical solution of this invention involves acquiring mixed speech to be separated. The mixed speech is then decomposed and transformed to obtain mixed amplitude and mixed phase information. The mixed amplitude information is input into a target speech separation model for speech separation processing. Based on the output of the target speech separation model, a time-frequency mask corresponding to the mixed speech to be separated can be automatically determined. Based on the time-frequency mask and the mixed speech to be separated, human voice amplitude information and noise amplitude information are determined. Based on the mixed phase information, the human voice amplitude information, and the noise amplitude information, target human voice speech and target noise speech are determined, thereby improving speech separation accuracy, optimizing speech separation smoothness, and enhancing user experience.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of a hybrid speech separation method provided in Embodiment 1 of the present invention;
[0026] Figure 2 This is a flowchart of a hybrid speech separation method provided in Embodiment 2 of the present invention;
[0027] Figure 3 This is a schematic diagram of a generative adversarial network structure provided in Embodiment 2 of the present invention;
[0028] Figure 4 This is a schematic diagram of the generator structure provided in Embodiment 2 of the present invention;
[0029] Figure 5 This is a schematic diagram of the discriminator structure provided in Embodiment 2 of the present invention;
[0030] Figure 6 This is a structural diagram of a hybrid speech separation device provided in Embodiment 3 of the present invention;
[0031] Figure 7 This is a structural diagram of an electronic device that implements the hybrid speech separation method of Embodiment 4 of the present invention. Detailed Implementation
[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] Example 1
[0035] Figure 1 The flowchart illustrates a mixed speech separation method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where mixed speech needs to be separated into pure human voice and noise. This method can be executed by a mixed speech separation device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0036] S101. Obtain the mixed speech to be separated.
[0037] S102. The mixed speech to be separated is decomposed and transformed to obtain the mixed amplitude information and mixed phase information corresponding to the mixed speech to be separated.
[0038] The mixed speech to be separated can refer to any superimposed mixed speech signals. Mixed amplitude information can refer to the amplitude spectrum information in the mixed speech to be separated. Mixed phase information can refer to the phase spectrum information in the mixed speech to be separated.
[0039] Specifically, due to the short-term stability of speech signals, time-frequency domain analysis of speech signals can effectively combine the advantages of both the time and frequency domains. Converting speech signals to the time-frequency domain for analysis is a common method to enhance the sparsity of speech signals. This invention uses short-time Fourier transform (SFT) technology to separate the mixed speech signals to be separated into time frames. SFT coefficients are obtained by performing a SFT on each frame of the mixed speech signal to be separated, and modulo operations are performed based on these SFT coefficients to obtain the mixing amplitude and mixing phase information corresponding to the mixed speech signals to be separated.
[0040] S103. Input the mixed amplitude information into the target speech separation model for speech separation processing, and determine the time-frequency mask corresponding to the mixed speech to be separated based on the output of the target speech separation model.
[0041] The target speech separation model can be obtained by training a generative adversarial network. This model can be used to obtain time-frequency masks for various types of mixed speech. The time-frequency mask refers to the proportion of human voice amplitude information in the mixed speech.
[0042] Specifically, the obtained mixing amplitude information is input into the trained target speech separation model for speech separation processing. Based on the output of the target speech separation model, the time-frequency mask corresponding to the mixed speech to be separated can be determined.
[0043] For example, the step of inputting the mixed amplitude information into the target speech separation model for speech separation processing, and determining the time-frequency mask corresponding to the mixed speech to be separated based on the output of the target speech separation model, includes:
[0044] The mixed amplitude information is input into the target speech separation model for speech separation processing, and based on the output of the target speech separation model, predicted human voice amplitude information and predicted noise amplitude information are obtained; according to the predicted human voice amplitude information and the predicted noise amplitude information, the time-frequency mask corresponding to the mixed speech to be separated is determined.
[0045] The predicted human voice amplitude information can refer to the amplitude information corresponding to the predicted human voice amplitude spectrum obtained by performing speech separation processing on the mixed amplitude information and analyzing the output of the target speech separation model. The predicted noise amplitude information can refer to the amplitude information corresponding to the predicted noise amplitude spectrum obtained by performing speech separation processing on the mixed amplitude information and analyzing the output of the target speech separation model. The predicted human voice amplitude information and the predicted noise amplitude information can be used to determine the time-frequency mask corresponding to the mixed speech to be separated.
[0046] Specifically, the mixed amplitude information is input into the target speech separation model for speech separation processing. Based on the output of the target speech separation model, predicted human voice amplitude information and predicted noise amplitude information can be obtained. Based on the predicted human voice amplitude information and predicted noise amplitude information, the time-frequency mask corresponding to the mixed speech to be separated can be calculated.
[0047] For example, determining the time-frequency mask corresponding to the mixed speech to be separated based on the predicted human voice amplitude information and the predicted noise amplitude information includes: determining the summation result of the predicted human voice amplitude information and the predicted noise amplitude information; and determining the quotient result of the predicted human voice amplitude information and the summation result as the time-frequency mask corresponding to the mixed speech to be separated.
[0048] Specifically, the sum of the predicted human voice amplitude information and the predicted noise amplitude information is calculated. The quotient between the predicted human voice amplitude information and the sum is determined as the time-frequency mask corresponding to the mixed speech to be separated.
[0049] For example, the time-frequency mask corresponding to the mixed speech to be separated can be implemented in the following way:
[0050]
[0051] Where m is the time-frequency mask. It could refer to predicting the amplitude information of human voices. This could refer to information about the predicted noise amplitude.
[0052] S104. Based on the time-frequency mask and the mixed speech to be separated, determine the human voice amplitude information and noise amplitude information.
[0053] Here, voice amplitude information can refer to the voice amplitude signal obtained after separating the mixed speech to be separated. Noise amplitude information can refer to the noise amplitude signal obtained after separating the mixed speech to be separated.
[0054] Specifically, based on the time-frequency mask and the mixed speech to be separated, the amplitude information of human voice and the amplitude information of noise can be calculated separately.
[0055] For example, determining the human voice amplitude information and noise amplitude information based on the time-frequency mask and the mixed speech to be separated includes: determining the human voice amplitude information by multiplying the time-frequency mask and the mixed speech to be separated; and determining the noise amplitude information by the difference between the mixed speech to be separated and the human voice amplitude information.
[0056] Specifically, the product of the time-frequency mask and the mixed speech to be separated is calculated, and the product result is determined as the human voice amplitude information. The difference between the mixed speech to be separated and the human voice amplitude information is calculated, and the difference result is determined as the noise amplitude information.
[0057] For example, voice amplitude information and noise amplitude information can be implemented in the following way:
[0058]
[0059]
[0060] in, It could refer to the amplitude information of human voice. It can refer to noise amplitude information, m can refer to time-frequency mask, and X can refer to the mixed speech to be separated.
[0061] Optionally, in this embodiment of the invention, the predicted human voice amplitude information and the predicted noise amplitude information output by the target speech separation model can also be directly determined as human voice amplitude information and noise amplitude information, respectively. Preferably, determining a time-frequency mask by using the predicted human voice amplitude information and the predicted noise amplitude information, and determining the human voice amplitude information and noise amplitude information based on the time-frequency mask and the mixed speech to be separated, can further smooth the speech separation effect and improve the user experience.
[0062] S105. Based on the mixed phase information, the human voice amplitude information, and the noise amplitude information, determine the target human voice speech and the target noise speech.
[0063] Specifically, the target human voice speech can refer to the separated, pure human voice speech. The target noisy speech can refer to the noise other than the target human voice speech in the mixed speech to be separated.
[0064] Specifically, performing an inverse short-time Fourier transform on the amplitude information and mixed phase information of the human voice can yield the target human voice speech. Performing an inverse short-time Fourier transform on the amplitude information and mixed phase information can yield the target noisy speech. Performing an inverse transform on the amplitude information and mixed phase information can convert the amplitude information into audible speech, thereby further improving the user experience.
[0065] The technical solution of this invention involves acquiring mixed speech to be separated. The mixed speech is then decomposed and transformed to obtain mixed amplitude and mixed phase information. The mixed amplitude information is input into a target speech separation model for speech separation processing. Based on the output of the target speech separation model, a time-frequency mask corresponding to the mixed speech to be separated can be automatically determined. Based on the time-frequency mask and the mixed speech to be separated, human voice amplitude information and noise amplitude information are determined. Based on the mixed phase information, the human voice amplitude information, and the noise amplitude information, target human voice speech and target noise speech are determined, thereby improving speech separation accuracy, optimizing speech separation smoothness, and enhancing user experience.
[0066] Example 2
[0067] Figure 2 This is a flowchart of a hybrid speech separation method provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment discloses the training process of the target speech separation model. For example... Figure 2 As shown, the method includes:
[0068] S201. Obtain mixed amplitude information samples and the expected human voice amplitude information and expected noise amplitude information corresponding to the mixed amplitude information samples.
[0069] The mixed amplitude information sample can include a pure human voice amplitude information sample, a pure noise amplitude information sample, and an amplitude information sample after the human voice and noise are mixed. Specifically, the mixed amplitude information sample, as well as the expected human voice amplitude information and expected noise amplitude information corresponding to the mixed amplitude information sample, can be obtained by recording or synthesis.
[0070] S202. Input the mixed amplitude information sample into the generator and perform speech separation processing based on the speech dictionary. Obtain the model separation result according to the output of the generator.
[0071] It should be noted that the target speech separation model can be obtained by training a generative adversarial network. Figure 3 This is a schematic diagram of the generative adversarial network structure provided in an embodiment of the present invention. Figure 3 As shown, the generative adversarial network includes a generator and a discriminator. The generator generates prediction magnitude information, and the discriminator determines whether the prediction magnitude information meets the actual requirements.
[0072] The speech dictionary can refer to the model parameters of the target speech separation model. The initial speech dictionary can be obtained by practicing with clean human speech and clean noisy speech based on non-negative matrix factorization techniques. Figure 4 This is a schematic diagram of the generator structure provided in an embodiment of the present invention. Figure 4 As shown, the generator may include a nonnegative matrix factorization layer, a masking layer, and a reconstruction layer.
[0073] Specifically, the mixed amplitude information samples are input into the generator, which performs speech separation processing on the mixed amplitude information samples based on a training function and a speech dictionary. The model separation result can be obtained from the generator's output.
[0074] S203. Input the expected human voice amplitude information, the expected noise amplitude information, the model separation result, and the mixed amplitude information sample into the discriminator to determine the model discrimination result corresponding to the discriminator.
[0075] The model separation results include predicted human voice amplitude information and predicted noise amplitude information. Figure 5 This is a schematic diagram of the discriminator structure provided in an embodiment of the present invention. Figure 5 As shown, the discriminator may include an input layer, a convolutional layer, a feature layer, and a result discrimination layer.
[0076] Specifically, the expected human voice amplitude information, expected noise amplitude information, model separation results, and mixed amplitude information samples are input into the discriminator. The discriminator can determine the discrimination result of the model corresponding to the current training by comparison and judgment.
[0077] For example, determining the model discrimination result corresponding to the discriminator includes: splicing the predicted human voice amplitude information, the predicted noise amplitude information, and the mixed amplitude information sample to obtain spliced predicted amplitude information; inputting the spliced predicted amplitude information into the discriminator to judge the prediction result, and determining the model discrimination result based on the output of the discriminator.
[0078] Among them, the spliced predicted amplitude information can be obtained by splicing the predicted human voice amplitude information, the predicted noise amplitude information, and the mixed amplitude information samples.
[0079] Specifically, the predicted human voice amplitude information, the predicted noise amplitude information, and the mixed amplitude information samples are spliced together to obtain spliced predicted amplitude information. This spliced predicted amplitude information is then input into the discriminator for prediction and judgment, and based on the discriminator's output, a model-compliant judgment result is determined.
[0080] For example, determining the model discrimination result based on the output of the discriminator includes: calculating the comparison amplitude error between the splicing prediction amplitude information and the splicing expected amplitude information output by the discriminator; if the comparison amplitude error is less than or equal to a preset amplitude error, determining that the model discrimination result meets the preset termination condition; if the comparison amplitude error is greater than the preset amplitude error, determining that the model discrimination result does not meet the preset termination condition.
[0081] The spliced expected amplitude information is obtained by splicing the expected human voice amplitude information, the expected noise amplitude information, and the mixed amplitude information samples. The model discrimination result can include whether the preset termination condition is met or not. The preset amplitude error can be determined according to the actual situation and is not limited here.
[0082] Specifically, the comparison amplitude error between the splicing predicted amplitude information and the splicing expected amplitude information is calculated. This comparison amplitude error is then compared with a preset amplitude error. If the comparison amplitude error is less than or equal to the preset amplitude error, the model's judgment result is determined to meet the preset termination condition. If the comparison amplitude error is greater than the preset amplitude error, the model's judgment result is determined to not meet the preset termination condition.
[0083] It should be noted that the generator and discriminator are updated alternately using a batch gradient descent algorithm. While one is being updated, the other parameter remains fixed. The goal of generative adversarial networks is to minimize the divergence of the distribution difference between natural and generated speech. This invention incorporates additional mixed speech to be separated into the discriminator during training, thus improving the smoothness of the separated speech and consequently enhancing the separation performance.
[0084] S204. If the model discrimination result does not meet the preset termination condition, adjust the speech dictionary in the generator and continue training the generator; or, if the model discrimination result meets the preset termination condition, determine the trained generator as the target speech separation model.
[0085] Specifically, if the model's discrimination result does not meet the preset termination condition, it indicates that the model training is not yet complete. The speech dictionary in the generator needs to be adjusted, and the generator training needs to continue until the training result meets the preset termination condition. Alternatively, if the model's discrimination result meets the preset termination condition, it indicates that the generator's prediction result can reach the expected level, and the trained generator can be identified as the target speech separation model.
[0086] S205. Obtain the mixed speech to be separated.
[0087] S206. Perform decomposition and transformation processing on the mixed speech to be separated to obtain the mixed amplitude information and mixed phase information corresponding to the mixed speech to be separated.
[0088] S207. Input the mixed amplitude information into the target speech separation model for speech separation processing, and determine the time-frequency mask corresponding to the mixed speech to be separated based on the output of the target speech separation model.
[0089] S208. Based on the time-frequency mask and the mixed speech to be separated, determine the human voice amplitude information and noise amplitude information.
[0090] S209. Based on the mixed phase information, the human voice amplitude information, and the noise amplitude information, determine the target human voice speech and the target noise speech.
[0091] The technical solution of this invention involves acquiring mixed amplitude information samples and corresponding expected human voice amplitude information and expected noise amplitude information. The mixed amplitude information samples are input into a generator, and speech separation processing is performed based on a speech dictionary. A model separation result is obtained based on the generator's output. The expected human voice amplitude information, the expected noise amplitude information, the model separation result, and the mixed amplitude information samples are input into a discriminator to determine the model discrimination result corresponding to the discriminator. If the model discrimination result does not meet a preset termination condition, the speech dictionary in the generator is adjusted, and the generator continues to be trained; or, if the model discrimination result meets the preset termination condition, the trained generator is determined as the target speech separation model. By using mixed amplitude information samples and corresponding expected human voice amplitude information and expected noise amplitude information for model training, the accuracy of speech separation by the target speech separation model can be guaranteed, improving the user experience.
[0092] Example 3
[0093] Figure 6 This is a schematic diagram of a hybrid speech separation device provided in Embodiment 3 of the present invention. Figure 6 As shown, the device includes: a mixed speech acquisition module 301, a mixed information acquisition module 302, a time-frequency mask acquisition module 303, an amplitude information determination module 304, and a target speech determination module 305.
[0094] in,
[0095] The system includes a mixed speech acquisition module 301 for acquiring mixed speech to be separated; a mixed information acquisition module 302 for decomposing and transforming the mixed speech to be separated to obtain mixed amplitude information and mixed phase information corresponding to the mixed speech; a time-frequency mask acquisition module 303 for inputting the mixed amplitude information into a target speech separation model for speech separation processing, and determining the time-frequency mask corresponding to the mixed speech to be separated based on the output of the target speech separation model; an amplitude information determination module 304 for determining human voice amplitude information and noise amplitude information based on the time-frequency mask and the mixed speech to be separated; and a target speech determination module 305 for determining target human voice speech and target noise speech based on the mixed phase information, the human voice amplitude information, and the noise amplitude information.
[0096] The technical solution of this invention involves acquiring mixed speech to be separated. The mixed speech is then decomposed and transformed to obtain mixed amplitude and mixed phase information. The mixed amplitude information is input into a target speech separation model for speech separation processing. Based on the output of the target speech separation model, a time-frequency mask corresponding to the mixed speech to be separated can be automatically determined. Based on the time-frequency mask and the mixed speech to be separated, human voice amplitude information and noise amplitude information are determined. Based on the mixed phase information, the human voice amplitude information, and the noise amplitude information, target human voice speech and target noise speech are determined, thereby improving speech separation accuracy, optimizing speech separation smoothness, and enhancing user experience.
[0097] Optionally, the time-frequency mask acquisition module 303 includes a prediction amplitude information acquisition unit and a time-frequency mask determination unit. Wherein,
[0098] The prediction amplitude information acquisition unit is used to input the mixed amplitude information into the target speech separation model for speech separation processing, and obtain the predicted human voice amplitude information and the predicted noise amplitude information based on the output of the target speech separation model;
[0099] The time-frequency mask determination unit is used to determine the time-frequency mask corresponding to the mixed speech to be separated based on the predicted human voice amplitude information and the predicted noise amplitude information.
[0100] Optionally, the time-frequency mask determination unit may be specifically used to: determine the summation result of the predicted human voice amplitude information and the predicted noise amplitude information; and determine the quotient result of the predicted human voice amplitude information and the summation result as the time-frequency mask corresponding to the mixed speech to be separated.
[0101] Optionally, the target speech separation model is obtained based on training with a generative adversarial network, wherein the generative adversarial network includes a generator and a discriminator; the device includes a model training module.
[0102] The model training module includes: an amplitude information sample acquisition unit, a generator training unit, a discriminator discrimination unit, and a discrimination result determination unit.
[0103] An amplitude information sample acquisition unit is used to acquire mixed amplitude information samples and the expected human voice amplitude information and expected noise amplitude information corresponding to the mixed amplitude information samples;
[0104] The generator training unit is used to input the mixed amplitude information samples into the generator, perform speech separation processing based on the speech dictionary, and obtain the model separation result based on the output of the generator;
[0105] The discriminator discrimination unit is used to input the expected human voice amplitude information, the expected noise amplitude information, the model separation result and the mixed amplitude information sample into the discriminator, and determine the model discrimination result corresponding to the discriminator;
[0106] The discrimination result determination unit is used to adjust the speech dictionary in the generator and continue training the generator when the model discrimination result does not meet the preset termination condition, or to determine the trained generator as the target speech separation model when the model discrimination result meets the preset termination condition.
[0107] Optionally, the model separation result includes predicted human voice amplitude information and predicted noise amplitude information; the discriminator discrimination unit includes a predicted amplitude information splicing subunit and a model discrimination result determination subunit. Wherein,
[0108] The prediction amplitude information splicing subunit is used to splice the predicted human voice amplitude information, the predicted noise amplitude information and the mixed amplitude information sample to obtain spliced prediction amplitude information;
[0109] The model discrimination result determination subunit is used to input the splicing prediction amplitude information into the discriminator to judge the prediction result, and determine the model discrimination result based on the output of the discriminator.
[0110] Optionally, the model's discrimination results determine the sub-units, specifically used for:
[0111] Calculate the comparison amplitude error between the splicing prediction amplitude information and the splicing expected amplitude information output by the discriminator, wherein the splicing expected amplitude information is obtained by splicing the expected human voice amplitude information, the expected noise amplitude information and the mixed amplitude information sample;
[0112] If the comparison amplitude error is less than or equal to the preset amplitude error, the model discrimination result is determined to meet the preset termination condition.
[0113] If the comparison amplitude error is greater than the preset amplitude error, the model discrimination result is determined to be that the preset termination condition is not met.
[0114] Optionally, the amplitude information determination module 304 is specifically used for:
[0115] The product of the time-frequency mask and the mixed speech to be separated is determined as the human voice amplitude information;
[0116] The difference between the mixed speech to be separated and the human voice amplitude information is determined as the noise amplitude information.
[0117] The hybrid speech separation device provided in the embodiments of the present invention can execute the hybrid speech separation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0118] Example 4
[0119] Figure 7 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0120] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0121] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0122] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as method-mixed speech separation.
[0123] In some embodiments, the method of hybrid speech separation can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method of hybrid speech separation described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform method of hybrid speech separation by any other suitable means (e.g., by means of firmware).
[0124] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0125] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0126] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0127] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0128] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0129] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0130] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0131] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for separating mixed speech into pure human voice and noise, characterized in that, include: Obtain the mixed speech to be separated; The mixed speech to be separated is decomposed and transformed to obtain the mixed amplitude information and mixed phase information corresponding to the mixed speech to be separated. The mixed amplitude information is input into the target speech separation model for speech separation processing, and the time-frequency mask corresponding to the mixed speech to be separated is determined based on the output of the target speech separation model. Based on the time-frequency mask and the mixed speech to be separated, determine the human voice amplitude information and noise amplitude information; Based on the mixed phase information, the human voice amplitude information, and the noise amplitude information, the target human voice speech and the target noise speech are determined. The step of determining the human voice amplitude information and noise amplitude information based on the time-frequency mask and the mixed speech to be separated includes: The product of the time-frequency mask and the mixed speech to be separated is determined as the human voice amplitude information; The difference between the mixed speech to be separated and the human voice amplitude information is determined as the noise amplitude information.
2. The method according to claim 1, characterized in that, The step of inputting the mixed amplitude information into the target speech separation model for speech separation processing, and determining the time-frequency mask corresponding to the mixed speech to be separated based on the output of the target speech separation model, includes: The mixed amplitude information is input into the target speech separation model for speech separation processing, and based on the output of the target speech separation model, the predicted human voice amplitude information and the predicted noise amplitude information are obtained. Based on the predicted human voice amplitude information and the predicted noise amplitude information, the time-frequency mask corresponding to the mixed speech to be separated is determined.
3. The method according to claim 2, characterized in that, The step of determining the time-frequency mask corresponding to the mixed speech to be separated based on the predicted human voice amplitude information and the predicted noise amplitude information includes: Determine the summation result of the predicted human voice amplitude information and the predicted noise amplitude information; The quotient of the predicted human voice amplitude information and the summation result is determined as the time-frequency mask corresponding to the mixed speech to be separated.
4. The method according to claim 1, characterized in that, The target speech separation model is obtained based on a generative adversarial network (GAN), wherein the GAN includes a generator and a discriminator; prior to acquiring the mixed speech to be separated, the model includes: Obtain mixed amplitude information samples and the corresponding expected human voice amplitude information and expected noise amplitude information; The mixed amplitude information sample is input into the generator, and speech separation processing is performed based on the speech dictionary. The model separation result is obtained based on the output of the generator. The expected human voice amplitude information, the expected noise amplitude information, the model separation result, and the mixed amplitude information sample are input into the discriminator to determine the model discrimination result corresponding to the discriminator; If the model's discrimination result does not meet the preset termination condition, the speech dictionary in the generator is adjusted, and the generator continues to be trained; or, if the model's discrimination result meets the preset termination condition, the trained generator is determined as the target speech separation model.
5. The method according to claim 4, characterized in that, The model separation results include predicted human voice amplitude information and predicted noise amplitude information; Determining the model discrimination result corresponding to the discriminator includes: The predicted human voice amplitude information, the predicted noise amplitude information, and the mixed amplitude information samples are spliced together to obtain spliced predicted amplitude information. The splicing prediction amplitude information is input into the discriminator to judge the prediction result, and the model discrimination result is determined based on the output of the discriminator.
6. The method according to claim 5, characterized in that, The determination of the model discrimination result based on the output of the discriminator includes: Calculate the comparison amplitude error between the splicing prediction amplitude information and the splicing expected amplitude information output by the discriminator, wherein the splicing expected amplitude information is obtained by splicing the expected human voice amplitude information, the expected noise amplitude information and the mixed amplitude information sample; If the comparison amplitude error is less than or equal to the preset amplitude error, the model discrimination result is determined to meet the preset termination condition. If the comparison amplitude error is greater than the preset amplitude error, the model discrimination result is determined to be that the preset termination condition is not met.
7. A mixed speech separation device, suitable for separating mixed speech into pure human voice and noise, characterized in that, include: The mixed speech acquisition module is used to acquire the mixed speech to be separated; The mixed information acquisition module is used to decompose and transform the mixed speech to be separated to obtain the mixed amplitude information and mixed phase information corresponding to the mixed speech to be separated. The time-frequency mask acquisition module is used to input the mixed amplitude information into the target speech separation model for speech separation processing, and determine the time-frequency mask corresponding to the mixed speech to be separated based on the output of the target speech separation model; The amplitude information determination module is used to determine the human voice amplitude information and noise amplitude information based on the time-frequency mask and the mixed speech to be separated; The target speech determination module is used to determine the target human speech and the target noise speech based on the mixed phase information, the human voice amplitude information and the noise amplitude information; The amplitude information determination module is specifically used for: The product of the time-frequency mask and the mixed speech to be separated is determined as the human voice amplitude information; The difference between the mixed speech to be separated and the human voice amplitude information is determined as the noise amplitude information.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the mixed speech separation method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the mixed speech separation method according to any one of claims 1-6.