Speech enhancement network training, speech enhancement method, apparatus and electronic device
By separating sound sources and aligning features, target-enhanced speech information is generated and the network is trained, which solves the problem of poor adaptability of speech enhancement networks to real environments and achieves more natural and understandable speech enhancement effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
- Filing Date
- 2022-09-16
- Publication Date
- 2026-04-21
AI Technical Summary
While existing speech enhancement networks can reduce noise and remove reverberation during training, the enhanced speech information has poor adaptability to the real environment, resulting in unnatural speech information and poor speech enhancement effect.
By acquiring sample speech information, clean speech signals, original impulse response signals, and target sound source characteristic data, sound source separation and characteristic alignment processing are performed to generate target enhanced speech information, and a target speech enhancement network is trained.
It improves the adaptability of voice information to the real environment, ensures the naturalness and intelligibility of voice information, and significantly improves the voice enhancement effect.
Smart Images

Figure CN115631759B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of speech processing technology, and in particular to a speech enhancement network training, speech enhancement method, apparatus, and electronic device. Background Technology
[0002] In recent years, with the rapid development of artificial intelligence technology, speech enhancement technology based on neural networks has also been applied in various speech processing scenarios, such as audio and video conferencing and live streaming. In order to eliminate noise, reverberation and other speech signals in speech information, corresponding speech enhancement networks are often trained.
[0003] In related technologies, during the training of speech enhancement networks, impact response signals in the room are usually combined for speech enhancement, thereby achieving the effect of noise reduction and reverberation removal. However, in related technologies, the enhanced speech information after training by the speech enhancement network is unnatural and differs greatly from the speech information in a normal environment, resulting in poor speech enhancement effect. Summary of the Invention
[0004] This disclosure provides a speech enhancement network training method, device, and electronic device, which can significantly improve the adaptability of speech information to the real environment based on noise reduction and reverberation removal, ensuring the naturalness and intelligibility of the speech information, and effectively improving the speech enhancement effect. The technical solution of this disclosure is as follows:
[0005] According to a first aspect of the present disclosure, a method for training a speech enhancement network is provided, comprising:
[0006] Acquire sample speech information, the clean speech signal corresponding to the sample speech information, the original impulse response signal corresponding to the sample speech information, the first device frequency response corresponding to the clean speech signal, and the target sound source distribution data corresponding to the target sound source characteristics;
[0007] The original impact response signal is subjected to sound source separation processing to obtain the original direct source response signal and the original reflection source response signal;
[0008] Based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal, sound source characteristic alignment processing is performed to obtain the target impact response signal;
[0009] Based on the target impact response signal, the clean speech signal, and the frequency response of the first device, target enhanced speech information corresponding to the sample speech information is generated;
[0010] Based on the sample speech information and the target enhanced speech information, a preset neural network is trained to enhance speech, thereby obtaining a target speech enhancement network corresponding to the characteristics of the target sound source.
[0011] In an optional embodiment, the process of separating the sound source from the original impact response signal to obtain the original direct source response signal and the original reflection source response signal includes:
[0012] Based on the first sampling function, the original impact response signal is sampled and processed to obtain the first power information corresponding to the original impact response signal;
[0013] Based on the second sampling function, the original impact response signal is sampled to obtain the second power information corresponding to the original impact response signal; the unit sampling length of the second sampling function is greater than the unit sampling length corresponding to the first sampling function.
[0014] Based on the first power information and the second power information, construct power ratio data;
[0015] Based on the power ratio data, the original direct source response signal and the original reflection source response signal are determined from the original impact response information.
[0016] In an optional embodiment, the power ratio data is data with the time information corresponding to the original impact response information as the horizontal axis and the ratio data corresponding to the first power information and the second power information as the vertical axis; determining the original direct source response signal and the original reflection source response signal from the original impact response information based on the power ratio data includes:
[0017] From the time information, at least one target time point is determined. The at least one target time point is the time point in which the corresponding ratio data satisfies a first preset condition and the corresponding signal amplitude in the original impact response signal satisfies a second preset condition.
[0018] Based on the at least one target time point and the preset time range, at least one signal interception interval is constructed;
[0019] Based on the at least one signal interception interval, at least one significant source response signal is intercepted from the original impact response signal; the at least one significant source response signal is the impact response signal located in the at least one signal interception interval in the original impact response;
[0020] Based on the at least one significant source response signal, the original direct source response signal and the original reflected source response signal are determined.
[0021] In an optional embodiment, the original reflection source response signal includes a first reflection response signal and a second reflection response signal; the at least one significant source response signal includes a plurality of significant source response signals; determining the original direct source response signal and the original reflection source response signal based on the at least one significant source response signal includes:
[0022] The first significant source response signal is taken as the original direct source response signal, and the first significant source response signal is the earliest significant source response signal at the target time point among the plurality of significant source response signals;
[0023] The other significant source response signals are aggregated to obtain the first reflection response signal, wherein the other significant source response signals are significant source response signals other than the first significant source response signal among the plurality of significant source response signals;
[0024] The multiple significant source response signals are aggregated to obtain a cumulative significant source response signal;
[0025] The signal difference between the original impact response signal and the cumulative significant source response signal is used as the second reflection response signal.
[0026] In an optional embodiment, the original reflection source response signal includes a first reflection response signal; the target sound source distribution data includes a first preset power ratio, the first preset power ratio characterizing the energy ratio between the target direct source response signal corresponding to the target sound source characteristics and the third reflection response signal corresponding to the target sound source characteristics, wherein the third reflection response signal and the first reflection response signal are the same type of reflection response signal;
[0027] The step of aligning the sound source characteristics based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal to obtain the target impact response signal includes:
[0028] The first sound source parameter tuning is determined based on the first preset power ratio, the first reflection response signal, and the original direct source response signal;
[0029] Based on the first sound source parameter tuning, the sound source characteristics of the first reflection response signal and the original direct source response signal are adjusted to obtain the target impact response signal.
[0030] In an optional embodiment, the original reflection source response signal includes a first reflection response signal; the target sound source distribution data includes a second preset power ratio; the second preset power ratio characterizes the energy ratio between the target direct source response signal and the first intercepted reflection signal corresponding to the target sound source characteristics, the first intercepted reflection signal is a signal obtained by intercepting the third reflection response signal corresponding to the target sound source characteristics based on preset signal interception parameters; the third reflection response signal and the first reflection response signal are the same type of reflection response signal;
[0031] The step of aligning the sound source characteristics based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal to obtain the target impact response signal includes:
[0032] The second sound source tuning parameters are determined based on the second preset power ratio, the preset signal interception parameters, the first reflection response signal, and the original direct source response signal;
[0033] Based on the preset signal interception parameters, the first reflection response signal is intercepted to obtain the second intercepted reflection signal;
[0034] Based on the second sound source parameter tuning, the sound source characteristics of the second intercepted reflected signal and the original direct source response signal are adjusted to obtain the target impact response signal.
[0035] In an optional embodiment, the original reflection source response signal includes a first reflection response signal and a second reflection response signal; the target sound source distribution data includes a third preset power ratio and a fourth preset power ratio, wherein the third preset power ratio characterizes the energy ratio between the third reflection response signal corresponding to the target sound source characteristics and the fourth reflection response signal corresponding to the target sound source characteristics, the third reflection response signal and the first reflection response signal are the same type of reflection response signal; the fourth reflection response signal and the second reflection response signal are the same type of reflection response signal; the fourth preset power ratio characterizes the energy ratio between the target direct source response signal corresponding to the target sound source characteristics and the first aligned reflection signal;
[0036] The step of aligning the sound source characteristics based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal to obtain the target impact response signal includes:
[0037] The third sound source tuning parameters are determined based on the third preset power ratio, the first reflection response signal, and the second reflection response signal;
[0038] Based on the third sound source parameter tuning, the sound source characteristics of the first reflection response signal and the second reflection response signal are adjusted to obtain the first aligned reflection signal;
[0039] The fourth sound source parameter tuning is determined based on the fourth preset power ratio, the first aligned reflection signal, and the original direct source response signal;
[0040] Based on the fourth sound source parameter tuning, the sound source characteristics of the first aligned reflection signal and the original direct source response signal are adjusted to obtain the target impact response signal.
[0041] In an optional embodiment, the original reflection source response signal includes a first reflection response signal and a second reflection response signal; the target sound source distribution data includes: a fifth preset power ratio and a sixth preset power ratio, wherein the fifth preset power ratio characterizes the energy ratio between the first intercepted reflection signal and the third intercepted reflection signal, the first intercepted reflection signal is a signal obtained by intercepting the third reflection response signal corresponding to the target sound source characteristics based on preset signal interception parameters; the third intercepted reflection signal is a signal obtained by intercepting the fourth reflection response signal corresponding to the target sound source characteristics based on the preset signal interception parameters; the third reflection response signal and the first reflection response signal are the same type of reflection response signal; the fourth reflection response signal and the second reflection response signal are the same type of reflection response signal; the sixth preset power ratio characterizes the energy ratio between the target direct source response signal and the second aligned reflection signal corresponding to the target sound source characteristics;
[0042] The step of aligning the sound source characteristics based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal to obtain the target impact response signal includes:
[0043] The fifth sound source tuning parameters are determined based on the fifth preset power ratio, the preset signal interception parameters, the first reflection response signal, and the second reflection response signal.
[0044] Based on the preset signal interception parameters, the first reflection response signal is intercepted to obtain the second intercepted reflection signal;
[0045] Based on the preset signal interception parameters, the second reflection response signal is intercepted to obtain the fourth intercepted reflection signal;
[0046] Based on the fifth sound source parameter tuning, the sound source characteristics of the second intercepted reflection signal and the fourth intercepted reflection signal are adjusted to obtain the second aligned reflection signal;
[0047] The sixth sound source parameter tuning is determined based on the sixth preset power ratio, the second aligned reflection signal, and the original direct source response signal;
[0048] Based on the sixth sound source parameter tuning, the sound source characteristics of the second aligned reflection signal and the original direct source response signal are adjusted to obtain the target impact response signal.
[0049] In an optional embodiment, obtaining the sample voice information includes:
[0050] Acquire a preset noise signal and the corresponding frequency response of the second device;
[0051] The sample speech information is constructed based on the pure speech signal, the original impact response signal, the frequency response of the first device, the preset noise signal, and the frequency response of the second device.
[0052] According to a second aspect of the present disclosure, a speech enhancement method is provided, comprising:
[0053] Obtain the speech information to be enhanced;
[0054] The speech information to be enhanced is input into the target speech enhancement network for speech enhancement processing to obtain the enhanced speech information corresponding to the speech information to be enhanced.
[0055] The target speech enhancement network is obtained based on the speech enhancement network training method as described in any one of the first aspects.
[0056] According to a third aspect of the present disclosure, a speech enhancement network training apparatus is provided, comprising:
[0057] The information acquisition module is configured to acquire sample speech information, the clean speech signal corresponding to the sample speech information, the original impulse response signal corresponding to the sample speech information, the first device frequency response corresponding to the clean speech signal, and the target sound source distribution data corresponding to the target sound source characteristics.
[0058] The sound source separation processing module is configured to perform sound source separation processing on the original impact response signal to obtain the original direct source response signal and the original reflection source response signal;
[0059] The sound source characteristic alignment processing module is configured to perform sound source characteristic alignment processing based on the target sound source distribution data, the original direct source response signal and the original reflected source response signal to obtain the target impact response signal;
[0060] The enhanced speech information generation module is configured to generate target enhanced speech information corresponding to the sample speech information based on the target impulse response signal, the clean speech signal and the frequency response of the first device.
[0061] The speech enhancement training module is configured to perform speech enhancement training on a preset neural network based on the sample speech information and the target enhanced speech information, so as to obtain a target speech enhancement network corresponding to the characteristics of the target sound source.
[0062] In an optional embodiment, the sound source separation processing module includes:
[0063] The first sampling processing unit is configured to perform sampling processing on the original impact response signal based on a first sampling function to obtain the first power information corresponding to the original impact response signal.
[0064] The second sampling processing unit is configured to perform sampling processing on the original impact response signal based on a second sampling function to obtain second power information corresponding to the original impact response signal; the unit sampling length of the second sampling function is greater than the unit sampling length corresponding to the first sampling function.
[0065] The power ratio data construction unit is configured to construct power ratio data based on the first power information and the second power information;
[0066] The sound source separation processing unit is configured to perform the determination of the original direct source response signal and the original reflected source response signal from the original impact response information based on the power ratio data.
[0067] In an optional embodiment, the power ratio data is data with the time information corresponding to the original impact response information as the horizontal axis and the ratio data corresponding to the first power information and the second power information as the vertical axis; the sound source separation processing unit includes:
[0068] The target time point determination unit is configured to determine at least one target time point from the time information, wherein the at least one target time point is a time point in which the corresponding ratio data satisfies a first preset condition and the corresponding signal amplitude in the original impact response signal satisfies a second preset condition.
[0069] The signal interception interval construction unit is configured to construct at least one signal interception interval based on the at least one target time point and a preset time range;
[0070] A significant source response signal interception unit is configured to intercept at least one significant source response signal from the original impact response signal based on the at least one signal interception interval; the at least one significant source response signal is an impact response signal in the original impact response that is located within the at least one signal interception interval;
[0071] The response signal determination unit is configured to perform the determination of the original direct source response signal and the original reflection source response signal based on the at least one significant source response signal.
[0072] In an optional embodiment, the original reflection source response signal includes a first reflection response signal and a second reflection response signal; the at least one significant source response signal includes multiple significant source response signals; the response signal determination unit includes:
[0073] The original direct source response signal determination unit is configured to use the first significant source response signal as the original direct source response signal, wherein the first significant source response signal is the earliest significant source response signal at the target time point among the plurality of significant source response signals;
[0074] The first aggregation processing unit is configured to perform aggregation processing on other significant source response signals to obtain the first reflection response signal, wherein the other significant source response signals are significant source response signals other than the first significant source response signal among the plurality of significant source response signals.
[0075] The second aggregation processing unit is configured to perform aggregation processing on the plurality of significant source response signals to obtain a cumulative significant source response signal;
[0076] The second reflection response signal determination unit is configured to use the signal difference between the original impact response signal and the cumulative significant source response signal as the second reflection response signal.
[0077] In an optional embodiment, the original reflection source response signal includes a first reflection response signal; the target sound source distribution data includes a first preset power ratio, the first preset power ratio characterizing the energy ratio between the target direct source response signal corresponding to the target sound source characteristics and the third reflection response signal corresponding to the target sound source characteristics, wherein the third reflection response signal and the first reflection response signal are the same type of reflection response signal;
[0078] The sound source characteristic alignment processing module includes:
[0079] The first sound source parameter determination unit is configured to determine the first sound source parameters based on the first preset power ratio, the first reflection response signal, and the original direct source response signal.
[0080] The first sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the first reflection response signal and the original direct source response signal based on the first sound source parameter tuning to obtain the target impact response signal.
[0081] In an optional embodiment, the original reflection source response signal includes a first reflection response signal; the target sound source distribution data includes a second preset power ratio; the second preset power ratio characterizes the energy ratio between the target direct source response signal and the first intercepted reflection signal corresponding to the target sound source characteristics, the first intercepted reflection signal is a signal obtained by intercepting the third reflection response signal corresponding to the target sound source characteristics based on preset signal interception parameters; the third reflection response signal and the first reflection response signal are the same type of reflection response signal;
[0082] The sound source characteristic alignment processing module includes:
[0083] The second sound source parameter determination unit is configured to determine the second sound source parameter based on the second preset power ratio, the preset signal interception parameters, the first reflection response signal, and the original direct source response signal.
[0084] The first signal interception unit is configured to perform signal interception on the first reflection response signal based on the preset signal interception parameters to obtain a second intercepted reflection signal.
[0085] The second sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the second intercepted reflected signal and the original direct source response signal based on the second sound source parameter tuning to obtain the target impact response signal.
[0086] In an optional embodiment, the original reflection source response signal includes a first reflection response signal and a second reflection response signal; the target sound source distribution data includes: a third preset power ratio and a fourth preset power ratio, wherein the third preset power ratio characterizes the energy ratio between the third reflection response signal corresponding to the target sound source characteristic and the fourth reflection response signal corresponding to the target sound source characteristic, the third reflection response signal and the first reflection response signal are the same type of reflection response signal; the fourth reflection response signal and the second reflection response signal are the same type of reflection response signal; the fourth preset power ratio characterizes the energy ratio between the target direct source response signal corresponding to the target sound source characteristic and the first aligned reflection signal; the sound source characteristic alignment processing module includes:
[0087] The third sound source parameter determination unit is configured to determine the third sound source parameters based on the third preset power ratio, the first reflection response signal, and the second reflection response signal.
[0088] The third sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the first reflection response signal and the second reflection response signal based on the third sound source parameter adjustment to obtain the first aligned reflection signal;
[0089] The fourth sound source parameter determination unit is configured to determine the fourth sound source parameter based on the fourth preset power ratio, the first aligned reflection signal, and the original direct source response signal;
[0090] The fourth sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the first aligned reflection signal and the original direct source response signal based on the fourth sound source parameter adjustment to obtain the target impact response signal.
[0091] In an optional embodiment, the original reflection source response signal includes a first reflection response signal and a second reflection response signal; the target sound source distribution data includes: a fifth preset power ratio and a sixth preset power ratio, wherein the fifth preset power ratio characterizes the energy ratio between the first intercepted reflection signal and the third intercepted reflection signal, the first intercepted reflection signal is a signal after intercepting the third reflection response signal corresponding to the target sound source characteristic based on preset signal interception parameters; the third intercepted reflection signal is a signal after intercepting the fourth reflection response signal corresponding to the target sound source characteristic based on the preset signal interception parameters; the third reflection response signal and the first reflection response signal are the same type of reflection response signal; the fourth reflection response signal and the second reflection response signal are the same type of reflection response signal; the sixth preset power ratio characterizes the energy ratio between the target direct source response signal and the second aligned reflection signal corresponding to the target sound source characteristic; the sound source characteristic alignment processing module includes:
[0092] The fifth sound source parameter determination unit is configured to determine the fifth sound source parameters based on the fifth preset power ratio, the preset signal interception parameters, the first reflection response signal, and the second reflection response signal.
[0093] The second signal interception unit is configured to perform signal interception on the first reflection response signal based on the preset signal interception parameters to obtain a second intercepted reflection signal.
[0094] The third signal interception unit is configured to perform signal interception on the second reflection response signal based on the preset signal interception parameters to obtain the fourth intercepted reflection signal;
[0095] The fifth sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the second intercepted reflection signal and the fourth intercepted reflection signal based on the fifth sound source parameter adjustment to obtain the second aligned reflection signal;
[0096] The sixth sound source parameter determination unit is configured to determine the sixth sound source parameters based on the sixth preset power ratio, the second aligned reflection signal, and the original direct source response signal;
[0097] The sixth sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the second aligned reflection signal and the original direct source response signal based on the sixth sound source parameter adjustment to obtain the target impact response signal.
[0098] In an optional embodiment, the information acquisition module includes:
[0099] The data acquisition unit is configured to acquire a preset noise signal and a second device frequency response corresponding to the preset noise signal;
[0100] The sample speech information construction unit is configured to construct the sample speech information based on the clean speech signal, the original impulse response signal, the frequency response of the first device, the preset noise signal, and the frequency response of the second device.
[0101] According to a fourth aspect of the present disclosure, a voice enhancement device is provided, comprising:
[0102] The module for acquiring voice information to be enhanced is configured to acquire voice information to be enhanced.
[0103] The speech enhancement processing module is configured to input the speech information to be enhanced into a target speech enhancement network for speech enhancement processing, and obtain the enhanced speech information corresponding to the speech information to be enhanced.
[0104] The target speech enhancement network is obtained based on the speech enhancement network training method described in any one of the first aspects.
[0105] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method as described in any one of the first or second aspects above.
[0106] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method described in any one of the first or second aspects of the present disclosure.
[0107] According to a seventh aspect of the present disclosure, a computer program product including instructions is provided, which, when run on a computer, causes the computer to perform the method described in any one of the first or second aspects of the present disclosure.
[0108] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:
[0109] In the process of training the speech enhancement network, by performing sound source separation processing on the original impulse response signal and combining it with the target sound source distribution data corresponding to the target sound source effect, the original direct source response signal and the original reflected source response signal obtained by separation are aligned with the sound source characteristics. This can effectively improve the matching degree between the target impulse response signal used to construct the target enhanced speech information and the sound source distribution in the real environment. In this way, the speech enhancement effect of the target speech enhancement network trained based on sample speech information and target enhanced speech information can be guaranteed. On the basis of noise reduction and reverberation removal, the adaptability between speech information and the real environment is greatly improved, ensuring the naturalness and intelligibility of speech information.
[0110] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0111] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0112] Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment;
[0113] Figure 2 This is a flowchart illustrating a speech enhancement network training method according to an exemplary embodiment;
[0114] Figure 3 This is a flowchart illustrating an exemplary embodiment of performing sound source separation processing on the original impact response signal to obtain the original direct source response signal and the original reflected source response signal;
[0115] Figure 4 This is a schematic diagram of a raw impact response signal, first power information, and second power information provided according to an exemplary quantity;
[0116] Figure 5 This is a flowchart illustrating, according to an exemplary embodiment, a method for determining the original direct source response signal and the original reflected source response signal from the original impact response information based on power ratio data;
[0117] Figure 6 This is a flowchart illustrating, according to an exemplary embodiment, a method for determining an original direct source response signal and an original reflected source response signal based on at least one significant source response signal;
[0118] Figure 7 This is a schematic diagram illustrating a process of training a target speech enhancement network according to an exemplary embodiment;
[0119] Figure 8 This is a block diagram of a speech enhancement network training apparatus according to an exemplary embodiment;
[0120] Figure 9 This is a block diagram illustrating a speech enhancement device according to an exemplary embodiment;
[0121] Figure 10 This is a block diagram illustrating an electronic device for training a speech enhancement network or for speech enhancement, according to an exemplary embodiment.
[0122] Figure 11 This is a block diagram illustrating another electronic device for training a speech enhancement network or for speech enhancement, according to an exemplary embodiment. Detailed Implementation
[0123] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0124] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0125] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0126] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment, which may include a server 100 and a terminal 200.
[0127] In an optional embodiment, server 100 can be used to train a speech enhancement network; terminal 200 can be used to provide speech enhancement processing based on the speech enhancement network trained by the server.
[0128] In one specific embodiment, server 100 can be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.
[0129] In one specific embodiment, terminal 200 may be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices, or software running on the aforementioned electronic devices, such as applications. Optionally, the operating system running on the electronic device may include, but is not limited to, Android, iOS, Linux, and Windows.
[0130] In addition, it should be noted that, Figure 1 The example shown is merely one application environment provided by this disclosure. In practical applications, other application environments may also be included, such as training a speech enhancement network on a terminal.
[0131] In the embodiments described in this specification, the server 100 and the terminal 200 can be directly or indirectly connected via wired or wireless communication, and this disclosure does not impose any restrictions.
[0132] Figure 2 This is a flowchart illustrating a speech enhancement network training method according to an exemplary embodiment, such as... Figure 2 As shown, this method can be applied to electronic devices such as servers or terminals, and specifically, it may include the following steps:
[0133] In step S201, sample speech information, the clean speech signal corresponding to the sample speech information, the original impulse response signal corresponding to the sample speech information, the first device frequency response corresponding to the clean speech signal, and the target sound source distribution data corresponding to the target sound source characteristics are obtained.
[0134] In a specific embodiment, the aforementioned sample speech information can be speech information with noise and reverberation; specifically, the sample speech information can include multiple speech information; the clean speech signal corresponding to any speech information can be the clean speech signal in that speech information, such as a human voice. The original impulse response signal corresponding to the sample speech information can be the impulse response signal in the sample speech information (data with time as the horizontal axis and the signal amplitude of the impulse response signal as the vertical axis). Optionally, the original impulse response signal can be obtained in a real scene by combining frequency sweep signals; or it can be calculated by combining a preset impulse response signal construction algorithm. The first device frequency response corresponding to the clean speech signal can be the frequency response of the clean speech signal acquisition device and the playback device (frequency response describes the difference in the device's ability to process signals of different frequencies); the target sound source distribution data corresponding to the target sound source characteristics can characterize the energy distribution of the response signals of different sound sources in the impulse response signal corresponding to the target sound source characteristics. Specifically, the characteristics of the target sound source can reflect the naturalness and intelligibility of the speech information; specifically, different sound source effects correspond to different sound source distribution data, which can be set according to the requirements of the naturalness and intelligibility (sound source characteristics) of the speech information in actual applications, and different target sound source distribution data.
[0135] In an optional embodiment, the sample speech information can be collected in a real-world scenario; alternatively, the sample speech information can be generated by combining a clean speech signal from a real-world scenario, an original impulse response signal, a first device frequency response corresponding to the clean speech signal, a preset noise signal, and a second device frequency response corresponding to the preset noise signal; correspondingly, obtaining the sample speech information may include:
[0136] Acquire a preset noise signal and the corresponding frequency response of the second device;
[0137] Sample speech information is constructed based on the pure speech signal, the original impulse response signal, the frequency response of the first device, the preset noise signal, and the frequency response of the second device.
[0138] In one specific embodiment, the preset noise signal can be the noise signal in the sample voice information. Specifically, the preset noise signal can be set according to the actual application or collected in a real scene; the second device frequency response corresponding to the preset noise signal can be the frequency response of the preset noise signal acquisition device and the playback device.
[0139] In an optional embodiment, the above-described construction of sample speech information based on the clean speech signal, the original impulse response signal, the frequency response of the first device, the preset noise signal, and the frequency response of the second device can be combined with the following formula:
[0140] m = eq1*rir*S + eq2*n
[0141] Where m represents sample speech information (any speech information), eq1 represents the frequency response of the first device; rir represents the original impulse response signal; s represents the clean speech signal; eq2 represents the frequency response of the second device; n represents the preset noise signal; and * represents convolution.
[0142] In the above embodiments, combining clean speech signals, original impact response signals, frequency response of the first device, preset noise signals, and frequency response of the second device to construct sample speech information can effectively improve the effectiveness and generation efficiency of sample speech information and avoid the cost of collecting speech information from a large number of real-world scenarios.
[0143] In step S203, the original impact response signal is subjected to sound source separation processing to obtain the original direct source response signal and the original reflection source response signal.
[0144] In one specific embodiment, the original direct source response signal can be the direct source response signal (the response signal of the direct sound source) in the original impact response signal. The original reflection source response signal can be the reflection source response signal (the response signal of the reflected sound source) in the original impact response signal.
[0145] In an optional embodiment, such as Figure 3 As shown, the above-described source separation processing of the original impact response signal to obtain the original direct source response signal and the original reflection source response signal may include the following steps:
[0146] In step S301, the original impact response signal is sampled and processed based on the first sampling function to obtain the first power information corresponding to the original impact response signal;
[0147] In step S303, the original impact response signal is sampled and processed based on the second sampling function to obtain the second power information corresponding to the original impact response signal;
[0148] In step S305, power ratio data is constructed based on the first power information and the second power information;
[0149] In step S307, based on the power ratio data, the original direct source response signal and the original reflection source response signal are determined from the original impact response information.
[0150] In an optional embodiment, the first sampling function and the second sampling function can be two normalized Hamming window functions with different unit sampling lengths; specifically, the unit sampling length of the second sampling function is greater than the unit sampling length of the first sampling function. The unit sampling lengths of the first and second sampling functions can be set according to the actual application; for example, the unit sampling length of the second sampling function is 0.0003, and the unit sampling length of the first sampling function is 0.0002.
[0151] In a specific embodiment, the power information corresponding to any time point (sampling point) in the original impact response signal can be the square of the signal amplitude of the impact response signal at that time point; specifically, the power information can characterize the energy level of the signal; the aforementioned first power information can be information characterizing the energy level of the original impact response signal after smoothing based on the first sampling function; the aforementioned second power information can be information characterizing the energy level of the original impact response signal after smoothing based on the second sampling function.
[0152] In a specific embodiment, the above-mentioned sampling processing of the original impact response signal based on the first sampling function to obtain the first power information corresponding to the original impact response signal can be combined with the following formula:
[0153] P slow = (rir·rir)*H slow
[0154] Among them, P slow Indicates the first power information; H slow The first sampling function is represented by rir; the original impulse response signal is represented by rir. Hamming slow (n) represents the first original Hamming window function corresponding to the nth sampling point (the original Hamming window function corresponding to the first sampling function); H slow (n) represents the first sampling function corresponding to the nth sampling point; N slow =round(τ slow ·fs); N slow τ represents the unit sampling length (i.e., window length) of the first sampling function; slow The first sampling function is represented by the preset time constant; fs represents the preset sampling frequency; round() represents rounding the value to the nearest whole number.
[0155] In a specific embodiment, the sampling processing of the original impact response signal based on the second sampling function to obtain the second power information corresponding to the original impact response signal can be combined with the following formula:
[0156] P fast = (rir·rir)*Hfast
[0157] Among them, P fast Indicates the second power information; H fast The second sampling function is represented by rir; the original impact response signal is represented by rir. H fast (n) represents the first sampling function corresponding to the nth sampling point; Hamming fast (n) represents the second original Hamming window function corresponding to the nth sampling point (the original Hamming window function corresponding to the second sampling function); N fast =round(τ fast ·fs); N fast τ represents the unit sampling length (i.e., window length) of the second sampling function; fast The first sampling function is represented by the preset time constant; fs represents the preset sampling frequency; round() represents rounding the value to the nearest whole number.
[0158] In a specific embodiment, such as Figure 4 As shown, Figure 4 This is a schematic diagram of an original impact response signal, first power information, and second power information provided according to an exemplary quantity. The curve corresponding to 101 is the original impact response signal; the curve corresponding to 102 is the first power information; and the curve corresponding to 103 is the second power information. Specifically, in conjunction with... Figure 4 It can be seen that by combining two sampling functions with different unit sampling lengths, the original impulse response signal can be smoothed to different degrees, effectively avoiding the situation where small disturbances in the original impulse response signal are misjudged as reflection response signals.
[0159] In one specific embodiment, the power ratio data can characterize the relative magnitude relationship between the first power information and the second power information. In practical applications, audio is generally processed in the dB (decibels) domain. Optionally, the aforementioned power ratio data can be the ratio of the first power information and the second power information in the dB domain. Accordingly, the power ratio data constructed based on the first power information and the second power information can be combined with the following formula:
[0160]
[0161] Where R represents the power ratio data; P slow Indicates the first power information; P fast This indicates the second power information.
[0162] In a specific embodiment, the aforementioned power ratio data can be data with the time information corresponding to the original impact response information as the horizontal axis and the ratio data (ratio in the dB domain) corresponding to the first power information and the second power information as the vertical axis; correspondingly, as... Figure 5 As shown, determining the original direct source response signal and the original reflected source response signal from the original impact response information based on the power ratio data may include the following steps:
[0163] In step S501, at least one target time point is determined from the time information. The at least one target time point is the time point where the corresponding ratio data satisfies the first preset condition and the corresponding signal amplitude in the original impact response signal satisfies the second preset condition.
[0164] In step S503, at least one signal interception interval is constructed based on at least one target time point and a preset time range;
[0165] In step S505, based on at least one signal interception interval, at least one significant source response signal is intercepted from the original impact response signal;
[0166] In step S507, the original direct source response signal and the original reflection source response signal are determined based on at least one significant source response signal.
[0167] In a specific embodiment, the target time point can be the time point corresponding to the peak point in the original impact response signal. The power ratio data at a certain time point satisfying the first preset condition can include: the power ratio data at that time point is greater than a first preset threshold, and the power ratio data at that time point is a local maximum (the power ratio data corresponding to two adjacent time points are both smaller than the power ratio data at that time point); the signal amplitude corresponding to a certain time point in the original impact response signal satisfying the second preset condition can include: the signal amplitude (amplitude corresponding to the vertical axis) corresponding to that time point is greater than a second preset threshold. Specifically, the first preset threshold and the second preset threshold can be set according to the actual application.
[0168] In a specific embodiment, at least one target time point can be selected from the original impact response signal by combining the first and second preset conditions described above. Further, when at least one target time point is selected, for any target time point, a signal interception interval for the corresponding significant source response signal can be determined by combining a preset time range. Specifically, the preset time range can be a pre-set time length corresponding to a single significant source response signal. Optionally, a signal interception interval within the preset time range can be selected centered on the target time point. Specifically, combining... Figure 4 As shown, the time period corresponding to Tr can be a segment of a certain signal.
[0169] In a specific embodiment, the at least one significant source response signal is an impact response signal located within at least one signal truncation interval in the original impact response; the extraction of at least one significant source response signal from the original impact response signal based on at least one signal truncation interval can be combined with the following formula:
[0170]
[0171] in, This represents the response signal of a significant source; This represents the signal truncation parameter at time t; t p Indicates the target time point; T r The signal intercept interval represents the time length corresponding to the signal intercept interval; rir represents the original impact response signal.
[0172] In one specific embodiment, the original reflection source response signal may include a first reflection response signal and a second reflection response signal; and in cases where the at least one significant source response signal includes multiple significant source response signals; such as Figure 6 As shown, determining the original direct source response signal and the original reflected source response signal based on at least one significant source response signal may include the following steps:
[0173] In step S601, the first significant source response signal is taken as the original direct source response signal;
[0174] In step S603, the response signals from other significant sources are aggregated to obtain the first reflection response signal;
[0175] In step S605, multiple significant source response signals are aggregated to obtain a cumulative significant source response signal;
[0176] In step S607, the signal difference between the original impact response signal and the cumulative significant source response signal is used as the second reflection response signal.
[0177] In one specific embodiment, the first significant source response signal can be the earliest significant source response signal at the target time point among the multiple significant source response signals; the other significant source response signals can be significant source response signals other than the first significant source response signal among the multiple significant source response signals.
[0178] In one specific embodiment, the first reflection response signal can be an early reflection response signal in the original impact response signal; the second reflection response signal can be a late reflection response signal in the original impact response signal.
[0179] In the above embodiments, when at least one significant source response signal includes multiple significant source response signals, the live source response signal and the first reflection response signal can be separated by combining the time sequence of the multiple significant source response signals. Furthermore, the second reflection response signal can be separated by combining the signal difference between the original impulse response signal and the cumulative significant source response signal. This can greatly improve the accuracy and effectiveness of impulse response signal separation, thereby improving the subsequent speech enhancement processing effect.
[0180] In an optional embodiment, when the number of at least one significant source response signal is one, the aforementioned original impact response signal consists only of the direct source response signal.
[0181] In the above embodiments, combining two sampling functions with different unit sampling lengths to smooth the original impact response signal to different degrees can effectively avoid small disturbances in the original impact response signal being misjudged as reflected response signals, thereby improving the accuracy and effectiveness of filtering out significant source response signals in the original impact response signal. Furthermore, by combining the signal amplitude in the original impact response signal with the ratio data between the two power information points after different degrees of smoothing, the target time point corresponding to the peak point in the original impact response signal can be filtered, effectively improving the accuracy and effectiveness of at least one target time point. Based on this at least one target time point and a preset time range, at least one signal interception interval can be constructed, allowing for the rapid and accurate extraction of at least one significant source response signal from the original impact response signal. Combined with this at least one significant source response signal, the original direct source response signal and the original reflected source response signal can be accurately separated from the original impact response signal.
[0182] In step S205, sound source characteristic alignment processing is performed based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal to obtain the target impact response signal.
[0183] In one specific embodiment, the target impact response signal can be an impact response signal whose sound source distribution data is consistent with the sound source distribution data corresponding to the target sound source characteristics. In an optional embodiment, the aforementioned target sound source distribution data includes a first preset power ratio, which can characterize the energy ratio between the target direct source response signal corresponding to the target sound source characteristics and the third reflection response signal corresponding to the target sound source characteristics, wherein the third reflection response signal and the first reflection response signal are the same type of reflection response signal (early reflection signal); correspondingly, the aforementioned sound source characteristic alignment processing based on the target sound source distribution data, the original direct source response signal, and the original reflection source response signal to obtain the target impact response signal can include:
[0184] The first sound source parameters are determined based on the first preset power ratio, the first reflection response signal, and the original direct source response signal;
[0185] Based on the parameter tuning of the first sound source, the sound source characteristics of the first reflection response signal and the original direct source response signal are adjusted to obtain the target impact response signal.
[0186] In a specific embodiment, the first preset power ratio can be the power ratio between the target direct source response signal and the third reflection response signal; specifically, the power of the entire signal can be the sum of the squares of the signals; correspondingly, to ensure that the power ratio between the direct source response signal and the first reflection response signal in the original impact response signal is consistent with the first preset power ratio, the determination of the first sound source parameter tuning based on the first preset power ratio, the first reflection response signal, and the original direct source response signal can be combined with the following formula:
[0187]
[0188] Where g1 represents the first sound source parameter tuning; DER1 represents the first preset power ratio; rir d (t) represents the signal amplitude at time t in the original direct source response signal; rir e (t) represents the signal amplitude at time t in the first reflection response signal. Specifically, the first sound source parameter tuning can be used to adjust the original impact response signal into a target impact response signal in which the power ratio of the direct source response signal to the early reflection response signal is a first preset power ratio.
[0189] In a specific embodiment, the above-mentioned adjustment of the sound source characteristics of the first reflection response signal and the original direct source response signal based on the first sound source parameter tuning to obtain the target impact response signal can be combined with the following formula:
[0190] rir t (t)=rir d (t)+g1·rir e (t)
[0191] Among them, rir t (t) represents the signal amplitude at time t in the target impact response signal; g1 represents the parameter tuning of the first sound source; rir d (t) represents the signal amplitude at time t in the original direct source response signal; rir e (t) represents the signal amplitude at time t in the first reflection response signal.
[0192] In one specific embodiment, the more reflected response signal retained in the target impact response signal, the higher the naturalness and the lower the intelligibility of the speech information generated based on the target impact response signal. Optionally, the aforementioned first preset power ratio can be set according to actual application requirements, thereby controlling the naturalness and intelligibility of the speech information generated based on the target impact response signal.
[0193] In the above embodiments, the first preset power ratio, which represents the energy ratio between the target direct source response signal corresponding to the target sound source characteristics and the third reflection response signal corresponding to the target sound source characteristics, is used as the target sound source distribution data. This can retain more reflection response signals and effectively improve the naturalness of the speech information while ensuring the intelligibility of the speech information.
[0194] In an optional embodiment, the target sound source distribution data includes: a second preset power ratio; specifically, the second preset power ratio characterizes the energy ratio between the target direct source response signal and the first intercepted reflection signal corresponding to the target sound source characteristics; the first intercepted reflection signal can be a signal obtained by intercepting the third reflection response signal corresponding to the target sound source characteristics based on preset signal interception parameters; the third reflection response signal and the first reflection response signal are the same type of reflection response signal (early reflection response signal); correspondingly, the above-mentioned sound source characteristic alignment processing based on the target sound source distribution data, the original direct source response signal, and the original reflection source response signal to obtain the target impact response signal may include:
[0195] Based on the second preset power ratio, preset signal interception parameters, first reflection response signal, and original direct source response signal, the second sound source tuning parameters are determined; based on the preset signal interception parameters, the first reflection response signal is intercepted to obtain the second intercepted reflection signal; based on the second sound source tuning parameters, the sound source characteristics of the second intercepted reflection signal and the original direct source response signal are adjusted to obtain the target impact response signal.
[0196] In a specific embodiment, the second sound source parameter tuning can be used to adjust the original impact response signal to a target impact response signal whose power ratio is a second preset power ratio, where the power ratio of the direct source response signal and the truncated early reflection response signal (the signal after truncating the early reflection response signal from the original impact response signal based on preset signal truncation parameters) is equal to a second preset power ratio. Optionally, the determination of the second sound source parameter tuning based on the second preset power ratio, preset signal truncation parameters, the first reflection response signal, and the original direct source response signal can be combined with the formula:
[0197]
[0198] Where g2 represents the second sound source parameter tuning; DER2 represents the second preset power ratio; rir d(t) represents the signal amplitude at time t in the original direct source response signal; rir e (t) represents the signal amplitude at time t in the first reflection response signal; A(t) represents the preset signal truncation parameter; T1 and T2 are preset signal capture time points.
[0199] In one specific embodiment, the product of the preset signal interception parameter and the first reflection response signal can be used as the second intercepted reflection signal.
[0200] In a specific embodiment, the above-mentioned adjustment of the sound source characteristics of the second intercepted reflected signal and the original direct source response signal based on the second sound source parameter tuning, to obtain the target impact response signal, can be combined with the following formula:
[0201] rir t (t)=rir d (t)+g2·rir e (t)·A(t)
[0202] Among them, rir t (t) represents the signal amplitude at time t in the target impact response signal; g2 represents the parameter tuning of the second sound source; rir d (t) represents the signal amplitude at time t in the original direct source response signal; rir d (t)·A(t) represents the signal amplitude at time t in the second intercepted reflected signal.
[0203] In the above embodiments, the second preset power ratio, which characterizes the energy ratio between the target direct source response signal and the first intercepted reflection signal corresponding to the target sound source characteristics, is used as the target sound source distribution data. Only a portion of the reflection response signal can be retained, thereby effectively improving the intelligibility of the speech information while ensuring its naturalness.
[0204] In an optional embodiment, the target sound source distribution data may include: a third preset power ratio and a fourth preset power ratio. The third preset power ratio can characterize the energy ratio between the third reflection response signal and the fourth reflection response signal corresponding to the target sound source characteristics. The third reflection response signal and the first reflection response signal are the same type of reflection response signal. The fourth reflection response signal and the second reflection response signal are the same type of reflection response signal. The fourth preset power ratio characterizes the energy ratio between the target direct source response signal and the first aligned reflection signal corresponding to the target sound source characteristics. The first aligned reflection signal may be a signal after adjusting the sound source characteristics of the first reflection response signal and the second reflection response signal based on the third sound source parameter tuning corresponding to the third preset power ratio. Correspondingly, the above-mentioned sound source characteristic alignment processing based on the target sound source distribution data, the original direct source response signal, and the original reflection source response signal to obtain the target impact response signal may include:
[0205] Based on the third preset power ratio, the first reflection response signal, and the second reflection response signal, a third sound source tuning parameter is determined; based on the third sound source tuning parameter, the sound source characteristics of the first reflection response signal and the second reflection response signal are adjusted to obtain a first aligned reflection signal; based on the fourth preset power ratio, the first aligned reflection signal, and the original direct source response signal, a fourth sound source tuning parameter is determined; based on the fourth sound source tuning parameter, the sound source characteristics of the first aligned reflection signal and the original direct source response signal are adjusted to obtain a target impact response signal.
[0206] In a specific embodiment, the third sound source parameter tuning can be used to adjust the original impact response signal to a target impact response signal in which the power ratio of the early reflection response signal and the late reflection response signal is a third preset power ratio. Optionally, the determination of the third sound source parameter tuning based on the third preset power ratio, the first reflection response signal, and the second reflection response signal can be combined with the following formula:
[0207]
[0208] Where g3 represents the third sound source parameter tuning; ELR1 represents the third preset power ratio; rir e (t) represents the signal amplitude at time t in the first reflection response signal; rir l (t) represents the signal amplitude at time t in the second reflection response signal; A(t) represents the preset signal truncation parameter.
[0209] In a specific embodiment, the above-mentioned adjustment of the sound source characteristics of the first and second reflection response signals based on the third sound source parameter tuning to obtain the first aligned reflection signal can be combined with the following formula:
[0210] rir t-el(t)=rir e (t)+g3·rir l (t)
[0211] Among them, rir t-el (t) represents the signal amplitude at time t in the first aligned reflection signal; g3 represents the parameter tuning of the third sound source; rir e (t) represents the signal amplitude at time t in the first reflection response signal; rir l (t) represents the signal amplitude at time t in the second reflection response signal.
[0212] In a specific embodiment, the fourth sound source parameter tuning can be used to adjust the original direct source response signal to a target impact response signal where the power ratio of the direct source response signal to the first aligned reflection signal is a fourth preset power ratio. The determination of the fourth sound source parameter tuning based on the fourth preset power ratio, the first aligned reflection signal, and the original direct source response signal can be combined with the following formula:
[0213]
[0214] Where g4 represents the fourth sound source parameter tuning; DDR1 represents the fourth preset power ratio; rir t-el (t) represents the signal amplitude at time t in the first aligned reflected signal; rir d (t) represents the signal amplitude at time t in the original direct source response signal.
[0215] In a specific embodiment, the above-mentioned adjustment of the sound source characteristics of the first aligned reflection signal and the original direct source response signal based on the fourth sound source parameter tuning to obtain the target impact response signal can be combined with the following formula:
[0216] rir t (t)=rird(t)+g4·rir t-el (t)
[0217] Among them, rir t (t) represents the signal amplitude at time t in the target impact response signal; g4 represents the parameter tuning of the fourth sound source; rir d (t) represents the signal amplitude at time t in the original direct source response signal; rir t-el (t) represents the signal amplitude at time t in the first aligned reflected signal.
[0218] In the above embodiments, by combining the third preset power ratio, which characterizes the energy ratio between the third reflection response signal and the fourth reflection response signal corresponding to the target sound source characteristics, and the fourth preset power ratio, which characterizes the energy ratio between the third reflection response signal and the fourth reflection response signal corresponding to the target sound source characteristics, the sound source characteristics can be adjusted. This can retain more reflection response signals and, while ensuring the intelligibility of the speech information, more effectively improve the naturalness of the speech information.
[0219] In an optional embodiment, the target sound source distribution data may include: a fifth preset power ratio and a sixth preset power ratio, wherein the fifth preset power ratio represents the energy ratio between the first intercepted reflection signal and the third intercepted reflection signal, the first intercepted reflection signal is a signal obtained by intercepting the third reflection response signal corresponding to the target sound source characteristics based on preset signal interception parameters; the third intercepted reflection signal is a signal obtained by intercepting the fourth reflection response signal corresponding to the target sound source characteristics based on preset signal interception parameters; the third reflection response signal and the first reflection response signal are the same type of reflection response signal; the fourth reflection response signal and the second reflection response signal are the same type of reflection response signal (late reflection response signal); the sixth preset power ratio represents the energy ratio between the target direct source response signal and the second aligned reflection signal corresponding to the target sound source characteristics; the second aligned reflection signal is a signal obtained by adjusting the sound source characteristics of the second intercepted reflection signal and the fourth intercepted reflection signal based on the fourth sound source parameter adjustment corresponding to the fifth preset power ratio; correspondingly, the above-mentioned sound source characteristic alignment processing based on the target sound source distribution data, the original direct source response signal and the original reflection source response signal to obtain the target impact response signal may include:
[0220] Based on the fifth preset power ratio, preset signal interception parameters, the first reflection response signal, and the second reflection response signal, the fifth sound source tuning parameters are determined. Based on the preset signal interception parameters, the first reflection response signal is intercepted to obtain the second intercepted reflection signal. Based on the preset signal interception parameters, the second reflection response signal is intercepted to obtain the fourth intercepted reflection signal. Based on the fifth sound source tuning parameters, the sound source characteristics of the second and fourth intercepted reflection signals are adjusted to obtain the second aligned reflection signal. Based on the sixth preset power ratio, the second aligned reflection signal, and the original direct source response signal, the sixth sound source tuning parameters are determined. Based on the sixth sound source tuning parameters, the sound source characteristics of the second aligned reflection signal and the original direct source response signal are adjusted to obtain the target impact response signal.
[0221] In a specific embodiment, the fifth sound source parameter tuning can be used to adjust the original direct source response signal to a target impact response signal in which the power ratio of the truncated early reflection response signal to the truncated late reflection response signal is a fifth preset power ratio. The determination of the fifth sound source parameter tuning based on the fifth preset power ratio, preset signal truncation parameters, the first reflection response signal, and the second reflection response signal can be combined with the following formula:
[0222]
[0223] Where g5 represents the fifth sound source parameter tuning; ELR2 represents the fourth preset power ratio; rir e (t) represents the signal amplitude at time t in the first reflection response signal; rir l (t) represents the signal amplitude at time t in the second reflection response signal; A(t) represents the preset signal truncation parameter; T1 and T2 are preset signal capture time points.
[0224] In one specific embodiment, the above-mentioned method of intercepting the first reflection response signal based on preset signal interception parameters to obtain the second intercepted reflection signal may include using the product of the preset signal interception parameters and the first reflection response signal as the second intercepted reflection signal.
[0225] In one specific embodiment, the above-mentioned method of intercepting the second reflection response signal based on preset signal interception parameters to obtain the fourth intercepted reflection signal may include using the product of the preset signal interception parameters and the second reflection response signal as the fourth intercepted reflection signal.
[0226] In a specific embodiment, the above-mentioned adjustment of the sound source characteristics of the second and fourth intercepted reflection signals based on the fifth sound source parameter tuning, to obtain the second aligned reflection signal, can be combined with the following formula:
[0227] rir t-el (t)=rir e (t)·A(t)+g5·rir l (t)·A(t)
[0228] Among them, rir t-el (t) represents the signal amplitude at time t in the second aligned reflection signal; g5 represents the parameter tuning of the fifth sound source; rir e (t)·A(t) represents the signal amplitude at time t in the second intercepted reflected signal; rir l (t)·A(t) represents the signal amplitude at time t in the fourth reflection response signal.
[0229] In a specific embodiment, the sixth sound source parameter tuning can be used to adjust the original direct source response signal to a target impact response signal in which the power ratio of the direct source reflection response signal to the second aligned reflection signal is a sixth preset power ratio. Optionally, the determination of the sixth sound source parameter tuning based on the sixth preset power ratio, the second aligned reflection signal, and the original direct source response signal can be combined with the following formula:
[0230]
[0231] Where g6 represents the sixth sound source parameter tuning; DDR2 represents the sixth preset power ratio; rir t-el (t) represents the signal amplitude at time t in the second aligned reflection signal; rir d (t) represents the signal amplitude at time t in the original direct source response signal.
[0232] In a specific embodiment, the above-mentioned adjustment of the sound source characteristics of the second aligned reflection signal and the original direct source response signal based on the sixth sound source parameter tuning, to obtain the target impact response signal, can be combined with the following formula:
[0233] rir t (t)=rir d (t)+g6·rir t-el (t)
[0234] Among them, rir t (t) The signal amplitude at time t in the target impact response signal; g6 represents the parameter tuning of the sixth sound source; rir t-el (t) represents the signal amplitude at time t in the second aligned reflection signal; rir d (t) represents the signal amplitude at time t in the original direct source response signal.
[0235] In the above embodiments, by combining a fifth preset power ratio that can characterize the energy ratio between the first intercepted reflection signal and the third intercepted reflection signal, and a sixth preset power ratio that characterizes the energy ratio between the target direct source response signal and the second aligned reflection signal corresponding to the target sound source characteristics, the sound source characteristics can be adjusted. This can reduce the reflection response signal and, while ensuring the naturalness of the speech information, more effectively improve the intelligibility of the speech information.
[0236] In step S207, target enhanced speech information corresponding to the sample speech information is generated based on the target impact response signal, the clean speech signal and the frequency response of the first device;
[0237] In one specific embodiment, the target-enhanced speech information can be speech information that satisfies the sound source distribution corresponding to the characteristics of the target sound source, based on denoising and dereverberation of the sample speech information. Specifically, generating the target-enhanced speech information corresponding to the sample speech information based on the target impact response signal, the clean speech signal, and the frequency response of the first device can include performing convolution processing on the target impact response signal, the clean speech signal, and the frequency response of the first device to obtain the aforementioned target-enhanced speech information.
[0238] In step S209, based on the sample speech information and the target enhanced speech information, the preset neural network is trained to enhance speech, thereby obtaining the target speech enhancement network corresponding to the characteristics of the target sound source.
[0239] In one specific embodiment, the preset neural network can be a neural network to be trained. The above-described method of training the preset neural network for speech enhancement based on sample speech information and target enhanced speech information to obtain a target speech enhancement network corresponding to the characteristics of the target sound source can include: inputting sample speech information into the preset neural network for speech enhancement processing to obtain predicted enhanced speech information; determining speech loss information based on the predicted enhanced speech information and the target enhanced speech information; and training the preset neural network based on the speech loss information to obtain the target speech enhancement network.
[0240] In a specific embodiment, speech loss information can characterize the degree of difference between the predicted enhanced speech information and the target enhanced speech information; in the process of determining speech loss information based on the predicted enhanced speech information and the target enhanced speech information, a preset loss function can be used.
[0241] In a specific embodiment, training a preset neural network based on speech loss information to obtain a target speech enhancement network may include: updating the network parameters of the preset neural network according to the speech loss information; repeating the above-mentioned inputting sample speech information into the preset neural network for speech enhancement processing based on the updated preset neural network to obtain predicted enhanced speech information; and determining the training iteration operation of speech loss information based on the predicted enhanced speech information and the target enhanced speech information until a preset convergence condition is met, and using the preset neural network corresponding to the condition met as the target speech enhancement network.
[0242] In an optional embodiment, satisfying the preset convergence condition can be that the number of training iterations reaches a preset number of training iterations. Optionally, satisfying the preset convergence condition can also be that the speech loss information is less than a specified threshold. In the embodiments of this specification, the preset number of training iterations and the specified threshold can be preset in conjunction with the training speed and accuracy of the network in practical applications.
[0243] In a specific embodiment, such as Figure 7 As shown, Figure 7This is a schematic diagram illustrating a process for training a target speech enhancement network according to an exemplary embodiment. Specifically, the input information for a preset neural network, namely sample speech information, can be constructed by combining the frequency response of a first device, the original impulse response signal, the clean speech signal, the preset noise signal, and the frequency response of a second device. Additionally, the original direct source response signal, the first reflection response signal, and the second reflection response signal can be separated through source separation processing of the original impulse response signal. By aligning the source features of the original direct source response signal, the first reflection response signal, and the second reflection response signal, the target impulse response signal can be obtained. Next, the target enhanced speech information can be constructed by combining the target impulse response signal, the frequency response of the first device, and the clean speech signal. Then, the predicted enhanced speech information after enhancing the sample speech information using the preset neural network, and the aforementioned target enhanced speech information can be used to calculate speech loss information. The preset neural network is then trained using the speech loss information to obtain the target speech enhancement network.
[0244] As can be seen from the technical solutions provided in the embodiments of this specification above, in the process of training the speech enhancement network, by performing sound source separation processing on the original impulse response signal and combining it with the target sound source distribution data corresponding to the target sound source effect, the original direct source response signal and the original reflected source response signal obtained by separation are aligned with the sound source characteristics. This can effectively improve the matching degree between the target impulse response signal used to construct the target enhanced speech information and the sound source distribution in the real environment. In this way, the speech enhancement effect of the target speech enhancement network trained based on sample speech information and target enhanced speech information can be guaranteed. On the basis of noise reduction and reverberation removal, the adaptability between speech information and the real environment is greatly improved, ensuring the naturalness and intelligibility of speech information.
[0245] Based on the aforementioned target speech enhancement network, the following describes a speech enhancement method provided by an embodiment of this disclosure. This speech enhancement method can enable speech communication with electronic devices such as terminals or servers, and may include the following steps:
[0246] Obtain the speech information to be enhanced;
[0247] The speech information to be enhanced is input into the target speech enhancement network for speech enhancement processing to obtain the enhanced speech information corresponding to the speech information to be enhanced.
[0248] In one specific embodiment, the speech information to be enhanced can be speech information with noise and reverberation effects. Optionally, taking an audio-visual conferencing scenario as an example, the speech information to be enhanced is speech information collected from any party (user terminal) participating in the audio-visual conferencing; optionally, the speech information to be enhanced can be processed by combining it with a target speech enhancement network, and the enhanced speech information after speech enhancement processing can be sent to other parties (other user terminals) participating in the audio-visual conferencing. Optionally, taking a live streaming scenario as an example, the speech information to be enhanced is live audio information collected by the live streaming end; optionally, the speech information to be enhanced can be processed by combining it with a target speech enhancement network, and the enhanced speech information after speech enhancement processing can be sent to the viewer end.
[0249] As can be seen from the technical solutions provided in the embodiments of this specification above, the speech enhancement processing of the speech information to be enhanced by combining the target speech enhancement network in this specification can greatly improve the adaptability between the speech information and the real environment on the basis of noise reduction and reverberation removal, ensure the naturalness and intelligibility of the speech information, and greatly improve the speech enhancement effect.
[0250] Figure 8 This is a block diagram illustrating a speech enhancement network training apparatus according to an exemplary embodiment. (Refer to...) Figure 8 The device includes:
[0251] The information acquisition module 810 is configured to acquire sample speech information, the clean speech signal corresponding to the sample speech information, the original impulse response signal corresponding to the sample speech information, the first device frequency response corresponding to the clean speech signal, and the target sound source distribution data corresponding to the target sound source characteristics.
[0252] The sound source separation processing module 820 is configured to perform sound source separation processing on the original impact response signal to obtain the original direct source response signal and the original reflection source response signal;
[0253] The sound source characteristic alignment processing module 830 is configured to perform sound source characteristic alignment processing based on the target sound source distribution data, the original direct source response signal and the original reflected source response signal to obtain the target impact response signal;
[0254] The enhanced speech information generation module 840 is configured to generate target enhanced speech information corresponding to the sample speech information based on the target impulse response signal, the clean speech signal and the first device frequency response;
[0255] The speech enhancement training module 850 is configured to perform speech enhancement training on a preset neural network based on sample speech information and target enhanced speech information, so as to obtain a target speech enhancement network corresponding to the characteristics of the target sound source.
[0256] In an optional embodiment, the sound source separation processing module 820 includes:
[0257] The first sampling processing unit is configured to perform sampling processing on the original impact response signal based on the first sampling function to obtain the first power information corresponding to the original impact response signal.
[0258] The second sampling processing unit is configured to perform sampling processing on the original impulse response signal based on the second sampling function to obtain the second power information corresponding to the original impulse response signal; the unit sampling length of the second sampling function is greater than the unit sampling length corresponding to the first sampling function.
[0259] The power ratio data construction unit is configured to construct power ratio data based on the first power information and the second power information;
[0260] The sound source separation processing unit is configured to perform the determination of the original direct source response signal and the original reflected source response signal from the original impact response information based on the power ratio data.
[0261] In an optional embodiment, the power ratio data is data with the time information corresponding to the original impact response information as the horizontal axis and the ratio data corresponding to the first power information and the second power information as the vertical axis; the sound source separation processing unit includes:
[0262] The target time point determination unit is configured to determine at least one target time point from the time information. The at least one target time point is the time point in which the corresponding ratio data satisfies a first preset condition and the corresponding signal amplitude in the original impact response signal satisfies a second preset condition.
[0263] The signal interception interval construction unit is configured to construct at least one signal interception interval based on at least one target time point and a preset time range;
[0264] The significant source response signal interception unit is configured to perform an operation based on at least one signal interception interval to intercept at least one significant source response signal from the original impact response signal; the at least one significant source response signal is an impact response signal in the original impact response that is located in at least one signal interception interval;
[0265] The response signal determination unit is configured to perform the determination of the original direct source response signal and the original reflection source response signal based on at least one significant source response signal.
[0266] In an optional embodiment, the original reflection source response signal includes a first reflection response signal and a second reflection response signal; at least one salient source response signal includes multiple salient source response signals; the response signal determination unit includes:
[0267] The original direct source response signal determination unit is configured to use the first significant source response signal as the original direct source response signal, wherein the first significant source response signal is the earliest significant source response signal at the target time point among multiple significant source response signals;
[0268] The first aggregation processing unit is configured to perform aggregation processing on other significant source response signals to obtain a first reflection response signal, wherein the other significant source response signals are significant source response signals other than the first significant source response signal among a plurality of significant source response signals.
[0269] The second aggregation processing unit is configured to perform aggregation processing on multiple significant source response signals to obtain a cumulative significant source response signal;
[0270] The second reflection response signal determination unit is configured to use the signal difference between the original impact response signal and the cumulative significant source response signal as the second reflection response signal.
[0271] In an optional embodiment, the original reflection source response signal includes a first reflection response signal; the target sound source distribution data includes a first preset power ratio, the first preset power ratio characterizing the energy ratio between the target direct source response signal corresponding to the target sound source characteristics and the third reflection response signal corresponding to the target sound source characteristics, the third reflection response signal and the first reflection response signal are the same type of reflection response signal;
[0272] The sound source feature alignment processing module 830 includes:
[0273] The first sound source parameter determination unit is configured to determine the first sound source parameters based on the first preset power ratio, the first reflection response signal, and the original direct source response signal.
[0274] The first sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the first reflection response signal and the original direct source response signal based on the first sound source parameter tuning to obtain the target impact response signal.
[0275] In an optional embodiment, the original reflection source response signal includes a first reflection response signal; the target sound source distribution data includes a second preset power ratio; the second preset power ratio characterizes the energy ratio between the target direct source response signal and the first intercepted reflection signal corresponding to the target sound source characteristics, and the first intercepted reflection signal is a signal obtained by intercepting the third reflection response signal corresponding to the target sound source characteristics based on preset signal interception parameters; the third reflection response signal and the first reflection response signal are the same type of reflection response signal;
[0276] The sound source feature alignment processing module 830 includes:
[0277] The second sound source parameter determination unit is configured to determine the second sound source parameter based on the second preset power ratio, preset signal interception parameters, the first reflection response signal and the original direct source response signal.
[0278] The first signal interception unit is configured to perform signal interception on the first reflection response signal based on preset signal interception parameters to obtain the second intercepted reflection signal;
[0279] The second sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the second intercepted reflected signal and the original direct source response signal based on the second sound source parameter tuning to obtain the target impact response signal.
[0280] In an optional embodiment, the original reflection source response signal includes a first reflection response signal and a second reflection response signal; the target sound source distribution data includes: a third preset power ratio and a fourth preset power ratio, wherein the third preset power ratio characterizes the energy ratio between the third reflection response signal corresponding to the target sound source characteristics and the fourth reflection response signal corresponding to the target sound source characteristics, and the third reflection response signal and the first reflection response signal are the same type of reflection response signal; the fourth reflection response signal and the second reflection response signal are the same type of reflection response signal; the fourth preset power ratio characterizes the energy ratio between the target direct source response signal corresponding to the target sound source characteristics and the first aligned reflection signal; the sound source characteristic alignment processing module 830 includes:
[0281] The third sound source parameter determination unit is configured to determine the third sound source parameters based on the third preset power ratio, the first reflection response signal, and the second reflection response signal.
[0282] The third sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the first reflection response signal and the second reflection response signal based on the third sound source parameter tuning to obtain the first aligned reflection signal;
[0283] The fourth sound source parameter determination unit is configured to determine the fourth sound source parameters based on the fourth preset power ratio, the first aligned reflection signal, and the original direct source response signal;
[0284] The fourth sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the first aligned reflection signal and the original direct source response signal based on the fourth sound source parameter adjustment to obtain the target impact response signal.
[0285] In an optional embodiment, the original reflection source response signal includes a first reflection response signal and a second reflection response signal; the target sound source distribution data includes: a fifth preset power ratio and a sixth preset power ratio, wherein the fifth preset power ratio characterizes the energy ratio between the first intercepted reflection signal and the third intercepted reflection signal, the first intercepted reflection signal is a signal obtained by intercepting the third reflection response signal corresponding to the target sound source characteristics based on preset signal interception parameters; the third intercepted reflection signal is a signal obtained by intercepting the fourth reflection response signal corresponding to the target sound source characteristics based on preset signal interception parameters; the third reflection response signal and the first reflection response signal are the same type of reflection response signal; the fourth reflection response signal and the second reflection response signal are the same type of reflection response signal; the sixth preset power ratio characterizes the energy ratio between the target direct source response signal and the second aligned reflection signal corresponding to the target sound source characteristics; the sound source characteristic alignment processing module 830 includes:
[0286] The fifth sound source parameter determination unit is configured to determine the fifth sound source parameters based on the fifth preset power ratio, preset signal interception parameters, first reflection response signal and second reflection response signal;
[0287] The second signal interception unit is configured to perform signal interception on the first reflection response signal based on preset signal interception parameters to obtain the second intercepted reflection signal.
[0288] The third signal interception unit is configured to perform signal interception on the second reflection response signal based on preset signal interception parameters to obtain the fourth intercepted reflection signal;
[0289] The fifth sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the second intercepted reflection signal and the fourth intercepted reflection signal based on the fifth sound source parameter tuning to obtain the second aligned reflection signal;
[0290] The sixth sound source parameter determination unit is configured to determine the sixth sound source parameters based on the sixth preset power ratio, the second aligned reflection signal, and the original direct source response signal;
[0291] The sixth sound source characteristic adjustment unit is configured to perform sound source characteristic adjustment on the second aligned reflection signal and the original direct source response signal based on the sixth sound source parameter tuning to obtain the target impact response signal.
[0292] In an optional embodiment, the information acquisition module 810 includes:
[0293] The data acquisition unit is configured to acquire a preset noise signal and the corresponding second device frequency response;
[0294] The sample speech information construction unit is configured to construct sample speech information based on the clean speech signal, the original impulse response signal, the frequency response of the first device, the preset noise signal, and the frequency response of the second device.
[0295] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0296] Figure 9 This is a block diagram illustrating a speech enhancement device according to an exemplary embodiment. (Refer to...) Figure 9 The device includes:
[0297] The speech information acquisition module 910 is configured to acquire speech information to be enhanced.
[0298] The speech enhancement processing module 920 is configured to input the speech information to be enhanced into the target speech enhancement network for speech enhancement processing, and obtain the enhanced speech information corresponding to the speech information to be enhanced.
[0299] The target speech enhancement network is obtained based on the speech enhancement network training method described above.
[0300] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0301] Figure 10 This is a block diagram illustrating an electronic device for training a speech enhancement network or for speech enhancement, according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 9 As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a speech enhancement network training method or a speech enhancement method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.
[0302] Figure 11This is a block diagram illustrating another electronic device for training a speech enhancement network or for speech enhancement, according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 11 As shown, the electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a speech enhancement network training method or a speech enhancement method.
[0303] Those skilled in the art will understand that Figure 10 or Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0304] In an exemplary embodiment, an electronic device is also provided, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the speech enhancement network training method as described in the embodiments of this disclosure.
[0305] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the speech enhancement network training method of the present disclosure embodiments.
[0306] In an exemplary embodiment, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the speech enhancement network training method of the present disclosure embodiments.
[0307] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0308] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0309] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for training a speech enhancement network, characterized in that, include: Acquire sample speech information, the clean speech signal corresponding to the sample speech information, the original impulse response signal corresponding to the sample speech information, the first device frequency response corresponding to the clean speech signal, and the target sound source distribution data corresponding to the target sound source characteristics. The target sound source distribution data characterizes the energy distribution of the response signals of different sound sources in the impulse response signal corresponding to the target sound source characteristics. The original impact response signal is subjected to sound source separation processing to obtain the original direct source response signal and the original reflection source response signal; Based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal, sound source characteristic alignment processing is performed to obtain the target impact response signal. The target impact response signal is an impact response signal whose sound source distribution data is consistent with the sound source distribution data corresponding to the target sound source characteristics. The alignment processing refers to converting the original direct source response signal and the original reflected source response signal into an impact response signal whose sound source distribution data is consistent with the sound source distribution data corresponding to the target sound source characteristics. Based on the target impact response signal, the clean speech signal, and the frequency response of the first device, target enhanced speech information corresponding to the sample speech information is generated; Based on the sample speech information and the target enhanced speech information, a preset neural network is trained to enhance speech, thereby obtaining a target speech enhancement network corresponding to the characteristics of the target sound source.
2. The speech enhancement network training method according to claim 1, characterized in that, The process of separating the sound source from the original impact response signal to obtain the original direct source response signal and the original reflected source response signal includes: Based on the first sampling function, the original impact response signal is sampled and processed to obtain the first power information corresponding to the original impact response signal; Based on the second sampling function, the original impact response signal is sampled to obtain the second power information corresponding to the original impact response signal; the unit sampling length of the second sampling function is greater than the unit sampling length corresponding to the first sampling function. Based on the first power information and the second power information, construct power ratio data; Based on the power ratio data, the original direct source response signal and the original reflection source response signal are determined from the original impact response information.
3. The speech enhancement network training method according to claim 2, characterized in that, The power ratio data is data with the time information corresponding to the original impact response information as the horizontal axis and the ratio data corresponding to the first power information and the second power information as the vertical axis. The step of determining the original direct source response signal and the original reflection source response signal from the original impact response information based on the power ratio data includes: From the time information, at least one target time point is determined. The at least one target time point is the time point in which the corresponding ratio data satisfies a first preset condition and the corresponding signal amplitude in the original impact response signal satisfies a second preset condition. Based on the at least one target time point and the preset time range, at least one signal interception interval is constructed; Based on the at least one signal interception interval, at least one significant source response signal is intercepted from the original impact response signal; the at least one significant source response signal is the impact response signal located in the at least one signal interception interval in the original impact response; Based on the at least one significant source response signal, the original direct source response signal and the original reflected source response signal are determined.
4. The speech enhancement network training method according to claim 3, characterized in that, The original reflection source response signal includes a first reflection response signal and a second reflection response signal; the at least one significant source response signal includes multiple significant source response signals; the Determining the original direct source response signal and the original reflected source response signal based on the at least one significant source response signal includes: The first significant source response signal is taken as the original direct source response signal, and the first significant source response signal is the earliest significant source response signal at the target time point among the plurality of significant source response signals; The other significant source response signals are aggregated to obtain the first reflection response signal, wherein the other significant source response signals are significant source response signals other than the first significant source response signal among the plurality of significant source response signals; The multiple significant source response signals are aggregated to obtain a cumulative significant source response signal; The signal difference between the original impact response signal and the cumulative significant source response signal is used as the second reflection response signal.
5. The speech enhancement network training method according to claim 1, characterized in that, The original reflection source response signal includes a first reflection response signal; the target sound source distribution data includes a first preset power ratio, the first preset power ratio characterizing the energy ratio between the target direct source response signal corresponding to the target sound source characteristics and the third reflection response signal corresponding to the target sound source characteristics, the third reflection response signal and the first reflection response signal are the same type of reflection response signal; The step of aligning the sound source characteristics based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal to obtain the target impact response signal includes: The first sound source parameter tuning is determined based on the first preset power ratio, the first reflection response signal, and the original direct source response signal; Based on the first sound source parameter tuning, the sound source characteristics of the first reflection response signal and the original direct source response signal are adjusted to obtain the target impact response signal.
6. The speech enhancement network training method according to claim 1, characterized in that, The original reflection source response signal includes a first reflection response signal; The target sound source distribution data includes: a second preset power ratio; the second preset power ratio characterizes the energy ratio between the target direct source response signal and the first intercepted reflection signal corresponding to the target sound source characteristics, wherein the first intercepted reflection signal is a signal obtained by intercepting the third reflection response signal corresponding to the target sound source characteristics based on preset signal interception parameters; the third reflection response signal and the first reflection response signal are the same type of reflection response signal. The step of aligning the sound source characteristics based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal to obtain the target impact response signal includes: The second sound source tuning parameters are determined based on the second preset power ratio, the preset signal interception parameters, the first reflection response signal, and the original direct source response signal; Based on the preset signal interception parameters, the first reflection response signal is intercepted to obtain the second intercepted reflection signal; Based on the second sound source parameter tuning, the sound source characteristics of the second intercepted reflected signal and the original direct source response signal are adjusted to obtain the target impact response signal.
7. The speech enhancement network training method according to claim 1, characterized in that, The original reflection source response signal includes a first reflection response signal and a second reflection response signal; The target sound source distribution data includes: a third preset power ratio and a fourth preset power ratio. The third preset power ratio represents the energy ratio between the third reflection response signal and the fourth reflection response signal corresponding to the target sound source characteristics. The third reflection response signal and the first reflection response signal are the same type of reflection response signal. The fourth reflection response signal and the second reflection response signal are the same type of reflection response signal. The fourth preset power ratio represents the energy ratio between the target direct source response signal and the first aligned reflection signal corresponding to the target sound source characteristics. The step of aligning the sound source characteristics based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal to obtain the target impact response signal includes: The third sound source tuning parameters are determined based on the third preset power ratio, the first reflection response signal, and the second reflection response signal; Based on the third sound source parameter tuning, the sound source characteristics of the first reflection response signal and the second reflection response signal are adjusted to obtain the first aligned reflection signal; The fourth sound source parameter tuning is determined based on the fourth preset power ratio, the first aligned reflection signal, and the original direct source response signal; Based on the fourth sound source parameter tuning, the sound source characteristics of the first aligned reflection signal and the original direct source response signal are adjusted to obtain the target impact response signal.
8. The speech enhancement network training method according to claim 1, characterized in that, The original reflection source response signal includes a first reflection response signal and a second reflection response signal; The target sound source distribution data includes: a fifth preset power ratio and a sixth preset power ratio. The fifth preset power ratio represents the energy ratio between a first intercepted reflection signal and a third intercepted reflection signal. The first intercepted reflection signal is a signal obtained by intercepting the third reflection response signal corresponding to the target sound source characteristics based on preset signal interception parameters. The third intercepted reflection signal is a signal obtained by intercepting the fourth reflection response signal corresponding to the target sound source characteristics based on the preset signal interception parameters. The third reflection response signal and the first reflection response signal are the same type of reflection response signal. The fourth reflection response signal and the second reflection response signal are the same type of reflection response signal. The sixth preset power ratio represents the energy ratio between the target direct source response signal and the second aligned reflection signal corresponding to the target sound source characteristics. The step of aligning the sound source characteristics based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal to obtain the target impact response signal includes: The fifth sound source tuning parameters are determined based on the fifth preset power ratio, the preset signal interception parameters, the first reflection response signal, and the second reflection response signal. Based on the preset signal interception parameters, the first reflection response signal is intercepted to obtain the second intercepted reflection signal; Based on the preset signal interception parameters, the second reflection response signal is intercepted to obtain the fourth intercepted reflection signal; Based on the fifth sound source parameter tuning, the sound source characteristics of the second intercepted reflection signal and the fourth intercepted reflection signal are adjusted to obtain the second aligned reflection signal; The sixth sound source parameter tuning is determined based on the sixth preset power ratio, the second aligned reflection signal, and the original direct source response signal; Based on the sixth sound source parameter tuning, the sound source characteristics of the second aligned reflection signal and the original direct source response signal are adjusted to obtain the target impact response signal.
9. The speech enhancement network training method according to any one of claims 1 to 8, characterized in that, The acquisition of sample voice information includes: Acquire a preset noise signal and the corresponding frequency response of the second device; The sample speech information is constructed based on the pure speech signal, the original impact response signal, the frequency response of the first device, the preset noise signal, and the frequency response of the second device.
10. A speech enhancement method, characterized in that, include: Obtain the speech information to be enhanced; The speech information to be enhanced is input into the target speech enhancement network for speech enhancement processing to obtain the enhanced speech information corresponding to the speech information to be enhanced. The target speech enhancement network is obtained based on the speech enhancement network training method according to any one of claims 1 to 9.
11. A speech enhancement network training device, characterized in that, include: The information acquisition module is configured to acquire sample speech information, a clean speech signal corresponding to the sample speech information, an original impulse response signal corresponding to the sample speech information, a first device frequency response corresponding to the clean speech signal, and target sound source distribution data corresponding to the target sound source characteristics. The target sound source distribution data characterizes the energy distribution of the response signals of different sound sources in the impulse response signal corresponding to the target sound source characteristics. The sound source separation processing module is configured to perform sound source separation processing on the original impact response signal to obtain the original direct source response signal and the original reflection source response signal; The sound source characteristic alignment processing module is configured to perform sound source characteristic alignment processing based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal to obtain a target impact response signal. The target impact response signal is an impact response signal whose sound source distribution data is consistent with the sound source distribution data corresponding to the target sound source characteristic. The alignment processing refers to converting the original direct source response signal and the original reflected source response signal into an impact response signal whose sound source distribution data is consistent with the sound source distribution data corresponding to the target sound source characteristic. The enhanced speech information generation module is configured to generate target enhanced speech information corresponding to the sample speech information based on the target impulse response signal, the clean speech signal and the frequency response of the first device. The speech enhancement training module is configured to perform speech enhancement training on a preset neural network based on the sample speech information and the target enhanced speech information, so as to obtain a target speech enhancement network corresponding to the characteristics of the target sound source.
12. A voice enhancement device, characterized in that, include: The module for acquiring voice information to be enhanced is configured to acquire voice information to be enhanced. The speech enhancement processing module is configured to input the speech information to be enhanced into a target speech enhancement network for speech enhancement processing, and obtain the enhanced speech information corresponding to the speech information to be enhanced. The target speech enhancement network is obtained based on the speech enhancement network training method according to any one of claims 1 to 9.
13. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the speech enhancement network training method as described in any one of claims 1 to 9 or the speech enhancement method as described in claim 10.
14. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the speech enhancement network training method as described in any one of claims 1 to 9 or the speech enhancement method as described in claim 10.
15. A computer program product containing instructions, characterized in that, When the instructions are executed on a computer, the computer obtains the speech enhancement network training method as described in any one of claims 1 to 9 or the speech enhancement method as described in claim 10.
Citation Information
Patent Citations
Training method and device of speech enhancement model and speech enhancement method and device
CN113555031A
Speech enhancement model training method, speech enhancement model recognition method, electronic equipment and storage medium
CN114283795A