Speech processing method and apparatus, storage medium, and electronic device

CN116246645BActive Publication Date: 2026-09-04SHENZHEN GRANDSTREAM NETWORKS TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211474175.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2026-09-04
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

[0004]本申请提供一种语音处理方法、装置、存储介质及电子设备,用于缓解当前混响消除效率低的技术问题

Benefits of technology

[0038] Beneficial effects: This application provides a speech processing method, apparatus, storage medium, and electronic device. After determining the reverberant speech amplitude spectrum and reverberant speech phase spectrum of the reverberant speech signal, the reverberant speech features of the reverberant speech amplitude spectrum are input into the target de-reverberation network to obtain the de-reverberation ratio. Based on the de-reverberation ratio and the reverberant speech phase spectrum, the clean speech signal can be extracted from the reverberant speech signal. Since there is no need to measure the reverberation time in this reverberation elimination process, the reverberation elimination efficiency is effectively improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246645B_ABST
    Figure CN116246645B_ABST
Patent Text Reader

Abstract

The application provides a speech processing method and device, a storage medium and an electronic device. When a reverberation speech signal is detected, the reverberation speech signal is preprocessed to obtain a reverberation speech amplitude spectrum and a reverberation speech phase spectrum, then a reverberation speech feature is determined according to the reverberation speech amplitude spectrum, the reverberation speech feature is input into a target dereverberation network to obtain a dereverberation ratio of the reverberation speech amplitude spectrum, a dereverberation speech amplitude spectrum is determined based on the dereverberation ratio and the reverberation speech amplitude spectrum, and finally a pure speech signal in the reverberation speech signal is obtained according to the dereverberation speech amplitude spectrum and the reverberation speech phase spectrum. The reverberation speech feature of the reverberation speech amplitude spectrum of the reverberation speech signal is input into the target dereverberation network to obtain the dereverberation ratio, and the dereverberation ratio and the reverberation speech phase spectrum can be used to eliminate the reverberation in the reverberation speech signal to obtain the pure speech signal. Since the reverberation time does not need to be measured in the reverberation elimination process, the reverberation elimination efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a speech processing method, apparatus, storage medium and electronic device. Background Technology

[0002] Sound reverberation is a common phenomenon in daily life. A moderate amount of early reverberation can make the sound full, but excessive reverberation time can have serious negative effects and usually has a certain degree of adverse impact on the clarity of speech. For example, in a confined space, when the speaker is far away from the microphone, the speech picked up by the microphone usually contains more reverberation components, thus affecting the clarity of the speaker's speech. Therefore, reverberation cancellation of reverberant speech signals is of great significance.

[0003] Currently, spectral subtraction is commonly used to eliminate reverberation in speech signals. However, since spectral subtraction relies on the attenuation factor of the reverberation decay index in the speech signal to achieve reverberation elimination, and the value of the attenuation factor is closely related to the reverberation time of each frequency band of the speech signal, and the reverberation time of different frequency bands is different, the value of the attenuation factor in different frequency bands is also different. Therefore, a tedious and complicated reverberation time measurement process is required before eliminating reverberation, resulting in low reverberation elimination efficiency. Summary of the Invention

[0004] This application provides a speech processing method, apparatus, storage medium, and electronic device to alleviate the technical problem of low reverberation cancellation efficiency in current technologies.

[0005] To address the aforementioned technical problems, this application provides the following technical solution:

[0006] This application provides a speech processing method, including:

[0007] When a reverberant speech signal is detected, the reverberant speech signal is preprocessed to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum.

[0008] The reverberant speech features are determined based on the reverberant speech amplitude spectrum, and the reverberant speech features are input into the target dereverberation network to obtain the dereverberation ratio of the reverberant speech amplitude spectrum.

[0009] The dereverberation speech amplitude spectrum is determined based on the dereverberation ratio and the reverberation speech amplitude spectrum.

[0010] Based on the dereverberation speech amplitude spectrum and the reverberation speech phase spectrum, a clean speech signal with the reverberation signal filtered out is obtained from the reverberation speech signal.

[0011] The step of preprocessing the reverberant speech signal to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum when the reverberant speech signal is detected further includes:

[0012] Acquire the reverberation time measurement signal to be processed and the clean speech signal to be processed; wherein, the reverberation time measurement signal to be processed includes a frequency sweep signal with multiple frequency bands;

[0013] Based on the reverberation time measurement signal to be processed and the clean speech signal to be processed, determine the reference reverberation speech features and the reference de-reverberation ratio;

[0014] Based on the reference reverberant speech features and the reference dreverberation ratio, the initial dreverberation network is trained to obtain the target dreverberation network.

[0015] The step of determining the reference reverberant speech features and the reference dérevering ratio based on the reverberation time measurement signal to be processed and the clean speech signal to be processed includes:

[0016] The reverberation time measurement signal to be processed and the clean speech signal to be processed are input into the reverberation simulator, so that the reverberation simulator adds a simulated reverberation signal to the reverberation time measurement signal to be processed and the clean speech signal to be processed, thereby obtaining a simulated reverberation time measurement signal and a simulated reverberation speech signal.

[0017] The simulated reverberation time measurement signal is processed by impulse response integration to obtain the reverberation time corresponding to each frequency band of the simulated reverberation time measurement signal;

[0018] Based on the reverberation time and the simulated reverberant speech signal, the reference reverberant speech features and the reference de-reverberation ratio are obtained.

[0019] The step of obtaining the reference reverberant speech features and the reference déreverberation ratio based on the reverberation time and the simulated reverberant speech signal includes:

[0020] The simulated reverberant speech signal is preprocessed to obtain the simulated reverberant speech amplitude spectrum;

[0021] Calculate the simulated reverberant speech log energy spectrum of the simulated reverberant speech amplitude spectrum, and use the simulated reverberant speech log energy spectrum as the reference reverberant speech feature;

[0022] Based on the reverberation time, the simulated reverberation speech amplitude spectrum is subjected to spectrum subtraction processing to obtain the de-reverberation target speech amplitude spectrum;

[0023] The ratio of the target speech amplitude spectrum to the simulated reverberant speech amplitude spectrum is used as the reference de-reverberation ratio.

[0024] The step of preprocessing the reverberant speech signal to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum when a reverberant speech signal is detected includes:

[0025] When a reverberant speech signal is detected, the reverberant speech signal is subjected to frame segmentation, windowing, and fast Fourier transform processing to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum.

[0026] The step of determining reverberant speech features based on the reverberant speech amplitude spectrum and inputting the reverberant speech features into a target dreverberation network to obtain the dreverberation ratio of the reverberant speech amplitude spectrum includes:

[0027] Calculate the reverberant speech log energy spectrum of the reverberant speech amplitude spectrum, and use the reverberant speech log energy spectrum as the reverberant speech feature;

[0028] The reverberant speech features are input into a target dereverberation network to perform prediction processing on the reverberant speech features, thereby obtaining the dereverberation ratio of the reverberant speech amplitude spectrum.

[0029] The step of determining the dederea speech amplitude spectrum based on the dederea ratio and the reverberant speech amplitude spectrum includes:

[0030] Multiply the de-reverberation ratio and the reverberation speech amplitude spectrum to obtain the de-reverberation speech amplitude spectrum.

[0031] This application also provides a voice processing device, including:

[0032] The preprocessing module is used to preprocess the reverberant speech signal when a reverberant speech signal is detected, so as to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum.

[0033] The de-reverberation ratio acquisition module is used to determine the reverberant speech features based on the reverberant speech amplitude spectrum, and input the reverberant speech features into the target de-reverberation network to obtain the de-reverberation ratio of the reverberant speech amplitude spectrum;

[0034] The determining module is used to determine the dereverberation speech amplitude spectrum based on the dereverberation ratio and the reverberation speech amplitude spectrum;

[0035] The clean speech signal acquisition module is used to obtain the clean speech signal from the reverberant speech signal after the reverberation signal has been filtered out, based on the amplitude spectrum of the dereverberant speech and the phase spectrum of the reverberant speech.

[0036] This application also provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute the steps in the above-described speech processing method.

[0037] This application also provides an electronic device, including a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the steps in the above-described voice processing method.

[0038] Beneficial effects: This application provides a speech processing method, apparatus, storage medium, and electronic device. After determining the reverberant speech amplitude spectrum and reverberant speech phase spectrum of the reverberant speech signal, the reverberant speech features of the reverberant speech amplitude spectrum are input into the target de-reverberation network to obtain the de-reverberation ratio. Based on the de-reverberation ratio and the reverberant speech phase spectrum, the clean speech signal can be extracted from the reverberant speech signal. Since there is no need to measure the reverberation time in this reverberation elimination process, the reverberation elimination efficiency is effectively improved. Attached Figure Description

[0039] The technical solution and other beneficial effects of this application will become apparent from the following detailed description of specific embodiments in conjunction with the accompanying drawings.

[0040] Figure 1 This is a flowchart illustrating the speech processing method provided in an embodiment of this application.

[0041] Figure 2 This is a schematic diagram of a scenario for the speech processing method provided in the embodiments of this application.

[0042] Figure 3 This is a schematic diagram of the structure of the voice processing device provided in the embodiments of this application.

[0043] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application.

[0044] Figure 5 This is another structural schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0045] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0046] This application provides a voice processing method, apparatus, storage medium, and electronic device.

[0047] like Figure 1 As shown, Figure 1 This is a flowchart illustrating the speech processing method provided in an embodiment of this application. The specific process can be as follows:

[0048] S101. When a reverberant speech signal is detected, the reverberant speech signal is preprocessed to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum.

[0049] Among them, the reverberant speech signal is a speech signal mixed with reverberation signal, the reverberant speech amplitude spectrum represents the distribution of the amplitude of the reverberant speech signal with frequency, and the reverberant speech phase spectrum represents the distribution of the phase value of the reverberant speech signal with frequency.

[0050] Specifically, as sound waves propagate indoors, they are reflected by obstacles such as walls, ceilings, and floors. Each time they are reflected, some sound is absorbed by the obstacle. As a result, when the sound source stops emitting sound, the sound waves will undergo multiple reflections and absorptions and disappear indoors. From the perspective of human hearing, after the sound source stops emitting sound, there are still several sound waves mixed together for a period of time (i.e., the phenomenon of sound continuation that still exists after the sound source stops emitting sound indoors). This phenomenon is called reverberation, and the reverberation signal is the result of the accumulation of speech signals as they are continuously reflected by obstacles indoors.

[0051] Reverberant speech signals can be speech signals from various scenarios that may produce reverberation, such as... Figure 2 As shown, during a remote conference held indoors, the distance between the speaker 201 and the microphone 202 is relatively far. The speech emitted by the speaker 201 is reflected by the wall 2031, the ceiling 2032 and the table 2033 in sequence as it is transmitted to the microphone 202. This results in the speech signal received by the microphone 202 being mixed with reverberation signal, which is then regarded as a reverberant speech signal.

[0052] Optionally, preprocessing includes framing, windowing, and Fast Fourier Transform (FFT) processing. In practical applications, a speech processing module can be set up in the microphone or electronic device (e.g., a laptop computer used to transmit conference voice). When the microphone or electronic device receives a reverberant speech signal, the speech processing module detects the reverberant speech signal and performs framing, windowing, and FFT processing on it to obtain the reverberant speech amplitude spectrum and reverberant speech phase spectrum, so as to more accurately describe the characteristics of the reverberant speech signal.

[0053] Furthermore, prior to step S101, the following steps are also included:

[0054] Acquire the reverberation time measurement signal to be processed and the clean speech signal to be processed; wherein, the reverberation time measurement signal to be processed includes a frequency sweep signal with multiple frequency bands;

[0055] Based on the reverberation time measurement signal to be processed and the clean speech signal to be processed, determine the reference reverberation speech characteristics and the reference de-reverberation ratio;

[0056] The initial dereverberation network is trained based on the reference reverberation speech features and the reference dereverberation ratio to obtain the target dereverberation network.

[0057] The reverberation time measurement signal to be processed is used to measure the reverberation time of different frequency bands of the speech signal. The clean speech signal to be processed is a speech signal without reverberation signal. The initial dereverberation network / target dereverberation network is a machine learning network used to remove the reverberation signal mixed in with the speech signal. Optionally, the initial dereverberation network / target dereverberation network is a CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), or GRU (Gated Recurrent Unit).

[0058] Specifically, in this embodiment, the GRU network is used as the initial dereverberation network. A clean speech signal without reverberation and a reverberation time measurement signal with multiple frequency bands are pre-acquired. The reverberation time measurement signal and the clean speech signal are input into a reverberation simulator. The reverberation simulator adds a simulated signal with the same reverberation time to the reverberation time measurement signal and the clean speech signal, resulting in a simulated reverberation time measurement signal (containing the reverberation time measurement signal and the simulated reverberation signal) and a simulated reverberated speech signal (containing the clean speech signal and the simulated reverberation signal) with the same reverberation time.

[0059] Next, impulse response integration is performed on the simulated reverberation time measurement signal to obtain the reverberation time corresponding to each frequency band of the simulated reverberation time measurement signal. Then, based on the reverberation time and the simulated reverberation speech signal, reference reverberation speech features and reference de-reverberation ratio are obtained. The reference reverberation speech features can distinguish the reverberation components in the simulated reverberation speech signal, and the reference de-reverberation ratio is used to remove the reverberation signal. The reference reverberation speech features are used as samples and the reference de-reverberation ratio is used as targets to train the GRU network so that the trained GRU network has the ability to identify the de-reverberation ratio corresponding to each reverberation speech feature.

[0060] Optionally, the reverberation time of the simulated reverberation signal can be adjusted by controlling the reverberation simulator to obtain multiple sets of simulated reverberation time measurement signals and simulated reverberation speech signals, thereby obtaining multiple sets of reference reverberation speech features and reference dération ratios. In other words, multiple samples and targets can be obtained to train the GRU network, thereby improving the dération capability of the GRU network.

[0061] Furthermore, in this embodiment, when obtaining the reference reverberant speech features and the reference déreverberation ratio, the simulated reverberant speech signal is first preprocessed to obtain the simulated reverberant speech amplitude spectrum. This simulated reverberant speech amplitude spectrum is used to accurately describe the characteristics of the simulated reverberant speech signal. Then, the square value of each amplitude value in the simulated reverberant speech amplitude spectrum is calculated first, and then the logarithm is taken to obtain the simulated reverberant speech logarithmic energy spectrum. The simulated reverberant speech logarithmic energy spectrum is used as the reference reverberant speech feature. Optionally, the preprocessing includes frame segmentation, windowing, and FFT processing.

[0062] Next, the amplitude spectrum of the simulated reverberant speech is subjected to spectrum subtraction processing based on the reverberation time to obtain the target speech amplitude spectrum after dreverberation, and the ratio of the target speech amplitude spectrum to the simulated reverberant speech amplitude spectrum is used as the reference dreverberation ratio.

[0063] For example, if Ratio1 represents the reference de-reverberation ratio, X1 represents the target speech amplitude spectrum, and Y1 represents the simulated reverberant speech amplitude spectrum, then Ratio1 = X1 / Y1, where Ratio1 can be a numerical sequence, and each value in the numerical sequence is an amplitude ratio.

[0064] S102. Determine the reverberant speech features based on the reverberant speech amplitude spectrum, and input the reverberant speech features into the target dereverberation network to obtain the dereverberation ratio of the reverberant speech amplitude spectrum.

[0065] Reverberant speech features are used to distinguish the reverberant components in a reverberant speech signal. Specifically, since the target dereverberation network has the ability to identify the dereverberation ratio corresponding to each reverberant speech feature, the reverberant speech features that can distinguish the reverberant components in the reverberant speech signal are input into the target dereverberation network. The target dereverberation network then identifies the dereverberation ratio corresponding to the reverberant speech feature, so that the reverberant signal in the reverberant speech signal can be removed subsequently based on the dereverberation ratio.

[0066] In this embodiment, the square value of each amplitude value in the reverberant speech amplitude spectrum is first calculated and then the logarithm is taken to obtain the reverberant speech logarithmic energy spectrum. The reverberant speech logarithmic energy spectrum is used as the reverberant speech feature. Then, the reverberant speech feature is input into the target dereverberation network to perform prediction processing on the reverberant speech feature through the target dereverberation network to obtain the dereverberation ratio of the reverberant speech amplitude spectrum.

[0067] For example, let Ratio2 represent the dreverberation ratio. Input the reverberant speech features into the target dreverberation network so that the reverberant speech features can be predicted by the target dreverberation network to obtain Ratio2.

[0068] S103. Determine the dereverberation speech amplitude spectrum based on the dereverberation ratio and the reverberation speech amplitude spectrum.

[0069] Among them, the dereverberant speech amplitude spectrum is the speech amplitude spectrum after reverberation is removed from the reverberant speech signal.

[0070] Specifically, in order to remove reverberation from a reverberant speech signal to retain a clean speech signal (i.e., a speech signal without reverberation), it is necessary to first determine the amplitude spectrum of the dereverberant speech signal so as to determine the distribution of the amplitude of the clean speech signal in the reverberant speech signal as a function of frequency.

[0071] Since the reference dreaving ratio is the ratio of the target speech amplitude spectrum to the simulated reverberant speech amplitude spectrum, and the initial dreaving network is trained using the reference dreaving ratio as the training target to obtain the target dreaving network, the dreaving ratio output by the target dreaving network reflects the relationship between the dreaving speech amplitude spectrum and the reverberant speech amplitude spectrum. Therefore, the dreaving speech amplitude spectrum can be obtained based on the relationship between the dreaving ratio, the reverberant speech amplitude spectrum, and the dreaving speech amplitude spectrum.

[0072] In this embodiment, the de-reverberation ratio and the reverberant speech amplitude spectrum are multiplied to obtain the de-reverberation speech amplitude spectrum. For example, the de-reverberation ratio Ratio2 = X2 / Y2, and the reverberant speech amplitude spectrum is Y2, so the de-reverberation speech amplitude spectrum = Ratio2 * Y2 = X2.

[0073] S104. Based on the amplitude spectrum and phase spectrum of the dereverberated speech signal, obtain the pure speech signal from the reverberated speech signal after the reverberation signal has been filtered out.

[0074] The reverberant speech phase spectrum represents the speech characteristics of the original speech signal (i.e., the speech signal before reverberation is added after it is emitted from the sound source). In order to improve the restoration of the speech signal without reverberation components extracted from the reverberant speech signal with the original speech signal, it is necessary to ensure that the extracted speech signal is neither mixed with reverberation nor has any signal omission. Therefore, it is necessary to combine the dereverberant speech amplitude spectrum and the reverberant speech phase spectrum and process them accordingly to obtain a pure speech signal that is neither mixed with reverberation nor has any signal omission.

[0075] In this embodiment, the amplitude spectrum and phase spectrum of the dereverberated speech are subjected to time-frequency transformation processing to convert them from the frequency domain to the time domain, thereby obtaining a clean speech signal in the reverberated speech signal from which the reverberation signal has been filtered out. Due to the time-frequency transformation processing, the clean speech signal has energy density or intensity at different times and frequencies, which means that the original speech signal is restored better.

[0076] As described above, the speech processing method provided in this application, when detecting a reverberant speech signal, preprocesses the reverberant speech signal to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum. Then, it determines the reverberant speech features based on the reverberant speech amplitude spectrum and inputs these features into a target dreverberation network to obtain the dreverberation ratio of the reverberant speech amplitude spectrum. Based on the dreverberation ratio and the reverberant speech amplitude spectrum, it determines the dreverberant speech amplitude spectrum. Finally, based on the dreverberant speech amplitude spectrum and the reverberant speech phase spectrum, it obtains the clean speech signal from the reverberant speech signal, from which the reverberation signal has been filtered out. After determining the reverberant speech amplitude spectrum and the reverberant speech phase spectrum of the reverberant speech signal, inputting the reverberant speech features of the reverberant speech amplitude spectrum into the target dreverberation network yields the dreverberation ratio. Based on the dreverberation ratio and the reverberant speech phase spectrum, the clean speech signal can be extracted from the reverberant speech signal. Since there is no need to measure the reverberation time during this reverberation cancellation process, the reverberation cancellation efficiency is effectively improved.

[0077] Based on the methods described in the above embodiments, this embodiment will be further described from the perspective of a voice processing device.

[0078] Please see Figure 3 , Figure 3 This application provides a specific description of a speech processing device, which may include: a preprocessing module 10, a de-reverberation ratio acquisition module 20, a determination module 30, and a clean speech signal acquisition module 40, wherein:

[0079] (1) Preprocessing module 10

[0080] The preprocessing module 10 is used to preprocess the reverberant speech signal when a reverberant speech signal is detected, so as to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum.

[0081] Specifically, the preprocessing module 10 is used for:

[0082] When a reverberant speech signal is detected, it is processed by framing, windowing, and fast Fourier transform to obtain the reverberant speech amplitude spectrum and reverberant speech phase spectrum.

[0083] (2) De-reverberation ratio acquisition module 20

[0084] The de-reverberation ratio acquisition module 20 is used to determine the reverberant speech features based on the reverberant speech amplitude spectrum, and input the reverberant speech features into the target de-reverberation network to obtain the de-reverberation ratio of the reverberant speech amplitude spectrum.

[0085] Specifically, the de-reverberation ratio acquisition module 20 is used for:

[0086] Calculate the logarithmic energy spectrum of the reverberant speech amplitude spectrum and use the logarithmic energy spectrum of the reverberant speech as a feature of the reverberant speech.

[0087] The reverberant speech features are input into the target dereverberation network to predict the reverberant speech features and obtain the dereverberation ratio of the reverberant speech amplitude spectrum.

[0088] (3) Determine module 30

[0089] The determination module 30 is used to determine the dereverberation speech amplitude spectrum based on the dereverberation ratio and the reverberation speech amplitude spectrum.

[0090] Specifically, the determining module 30 is used for:

[0091] Multiply the de-reverberation ratio and the reverberation speech amplitude spectrum to obtain the de-reverberation speech amplitude spectrum.

[0092] (4) Pure voice signal acquisition module 40

[0093] The clean speech signal acquisition module 40 is used to obtain the clean speech signal from the reverberant speech signal after the reverberation signal has been filtered out, based on the amplitude spectrum of the dereverberant speech and the phase spectrum of the reverberant speech.

[0094] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.

[0095] As described above, the speech processing apparatus provided in this application, when detecting a reverberant speech signal, preprocesses the reverberant speech signal through the preprocessing module 10 to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum. Then, the de-reverberation ratio acquisition module 20 determines the reverberant speech features based on the reverberant speech amplitude spectrum and inputs the reverberant speech features into the target de-reverberation network to obtain the de-reverberation ratio of the reverberant speech amplitude spectrum. Then, the determination module 30 determines the de-reverberant speech amplitude spectrum based on the de-reverberation ratio and the reverberant speech amplitude spectrum. Finally, the clean speech signal acquisition module 40 obtains the clean speech signal from the reverberant speech signal after the reverberation signal has been filtered out based on the de-reverberant speech amplitude spectrum and the reverberant speech phase spectrum. After determining the reverberant speech amplitude spectrum and reverberant speech phase spectrum of the reverberant speech signal, the reverberant speech features of the reverberant speech amplitude spectrum are input into the target de-reverberation network to obtain the de-reverberation ratio. Based on the de-reverberation ratio and the reverberant speech phase spectrum, the clean speech signal can be extracted from the reverberant speech signal. Since there is no need to measure the reverberation time in this reverberation elimination process, the reverberation elimination efficiency is effectively improved.

[0096] Accordingly, embodiments of the present invention also provide a voice processing system, including any of the voice processing devices provided in the embodiments of the present invention, which can be integrated into an electronic device.

[0097] When a reverberant speech signal is detected, it is preprocessed to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum. The reverberant speech features are determined based on the reverberant speech amplitude spectrum and input into the target dereverberation network to obtain the dereverberation ratio of the reverberant speech amplitude spectrum. Based on the dereverberation ratio and the reverberant speech amplitude spectrum, the dereverberant speech amplitude spectrum is determined. Based on the dereverberant speech amplitude spectrum and the reverberant speech phase spectrum, the clean speech signal with the reverberant signal filtered out is obtained from the reverberant speech signal.

[0098] The specific implementation details of each of the above devices can be found in the preceding embodiments, and will not be repeated here.

[0099] Since the speech processing system can include any of the speech processing devices provided in the embodiments of the present invention, it can achieve the beneficial effects that any of the speech processing devices provided in the embodiments of the present invention can achieve, as detailed in the preceding embodiments, and will not be repeated here.

[0100] Additionally, this application also provides an electronic device, which may be a smartphone or a computer, etc. Figure 4 As shown, the electronic device 400 includes a processor 401 and a memory 402. The processor 401 and the memory 402 are electrically connected.

[0101] The processor 401 is the control center of the electronic device 400. It connects various parts of the electronic device through various interfaces and lines. By running or loading the application program stored in the memory 402 and calling the data stored in the memory 402, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole.

[0102] In this embodiment, the processor 401 in the electronic device 400 loads the instructions corresponding to the processes of one or more application programs into the memory 402 according to the following steps, and the processor 401 runs the application programs stored in the memory 402 to realize various functions:

[0103] When a reverberant speech signal is detected, the reverberant speech signal is preprocessed to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum.

[0104] The reverberant speech features are determined based on the reverberant speech amplitude spectrum, and then the reverberant speech features are input into the target dreverberation network to obtain the dreverberation ratio of the reverberant speech amplitude spectrum.

[0105] The dereverberation speech amplitude spectrum is determined based on the dereverberation ratio and the reverberation speech amplitude spectrum.

[0106] Based on the amplitude spectrum and phase spectrum of the dereverberated speech signal, the clean speech signal with the reverberation signal filtered out is obtained from the reverberated speech signal.

[0107] Figure 5 A specific structural block diagram of an electronic device provided in an embodiment of the present invention is shown. This electronic device can be used to implement the voice processing method provided in the above embodiments.

[0108] RF circuit 510 is used to receive and transmit electromagnetic waves, converting electromagnetic waves into electrical signals and vice versa, thereby enabling communication with communication networks or other devices. RF circuit 510 may include various existing circuit elements used to perform these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, Subscriber Identity Module (SIM) cards, memory, etc. RF circuit 510 can communicate with various networks such as the Internet, corporate intranets, and wireless networks, or communicate with other devices via wireless networks. The aforementioned wireless networks may include cellular telephone networks, wireless local area networks (WLANs), or metropolitan area networks (MANs). The aforementioned wireless networks may use various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messages, and any other suitable communication protocols, including those that have not yet been developed.

[0109] The memory 520 can be used to store software programs and modules. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, thereby realizing the function of storing 5G capability information. The memory 520 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 520 may further include memory remotely located relative to the processor 580, and these remote memories can be connected to the electronic device 500 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0110] The input unit 530 can be used to receive input digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, the input unit 530 may include a touch-sensitive surface 531 and other input devices 532. The touch-sensitive surface 531, also known as a touch display screen or touchpad, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch-sensitive surface 531), and drive the corresponding connection device according to a pre-set program. Optionally, the touch-sensitive surface 531 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 580, and can receive and execute commands from the processor 580. In addition, the touch-sensitive surface 531 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 531, the input unit 530 may also include other input devices 532. Specifically, other input devices 532 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0111] Display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of electronic device 500. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 540 may include display panel 541, optionally configured as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc. Further, touch-sensitive surface 531 may cover display panel 541. When touch-sensitive surface 531 detects a touch operation on or near it, it transmits the information to processor 580 to determine the type of touch event. Subsequently, processor 580 provides corresponding visual output on display panel 541 according to the type of touch event. Although in Figure 5 In this embodiment, the touch-sensitive surface 531 and the display panel 541 are implemented as two separate components to realize the input and output functions. However, in some embodiments, the touch-sensitive surface 531 and the display panel 541 can be integrated to realize the input and output functions.

[0112] The electronic device 500 may also include at least one sensor 550, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 541 according to the ambient light level, and the proximity sensor can turn off the display panel 541 and / or backlight when the electronic device 500 is moved to the ear. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometers, taps), etc. Other sensors that the electronic device 500 may also be equipped with, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0113] Audio circuitry 560, speaker 561, and microphone 562 provide an audio interface between the user and electronic device 500. Audio circuitry 560 converts received audio data into electrical signals and transmits them to speaker 561, where speaker 561 converts them into sound signals for output. Conversely, microphone 562 converts collected sound signals into electrical signals, which are then received by audio circuitry 560, converted back into audio data, and processed by processor 580. The audio data is then transmitted via RF circuitry 510 to, for example, another terminal, or output to memory 520 for further processing. Audio circuitry 560 may also include an earphone jack to facilitate communication between external headphones and electronic device 500.

[0114] Electronic device 500, through transmission module 570 (e.g., Wi-Fi module), can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 5 The transmission module 570 is shown, but it is understood that it is not a necessary component of the electronic device 500 and can be omitted as needed without changing the nature of the invention.

[0115] The processor 580 is the control center of the electronic device 500. It connects to various parts of the phone via various interfaces and lines, and performs various functions and processes data of the electronic device 500 by running or executing software programs and / or modules stored in the memory 520, and by calling data stored in the memory 520. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 580.

[0116] Electronic device 500 also includes a power supply 590 (such as a battery) for supplying power to various components. In some embodiments, the power supply may be logically connected to processor 580 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. The power supply 590 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0117] Although not shown, the electronic device 500 may also include a camera (such as a front-facing camera and a rear-facing camera), a Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the electronic device also includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors. One or more programs contain instructions for performing the following operations:

[0118] When a reverberant speech signal is detected, the reverberant speech signal is preprocessed to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum.

[0119] The reverberant speech features are determined based on the reverberant speech amplitude spectrum, and then the reverberant speech features are input into the target dreverberation network to obtain the dreverberation ratio of the reverberant speech amplitude spectrum.

[0120] The dereverberation speech amplitude spectrum is determined based on the dereverberation ratio and the reverberation speech amplitude spectrum.

[0121] Based on the amplitude spectrum and phase spectrum of the dereverberated speech signal, the clean speech signal with the reverberation signal filtered out is obtained from the reverberated speech signal.

[0122] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.

[0123] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, embodiments of the present invention provide a storage medium storing multiple instructions that can be loaded by a processor to execute the steps in any of the speech processing methods provided in the embodiments of the present invention.

[0124] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0125] Since the instructions stored in the storage medium can execute the steps of any of the speech processing methods provided in the embodiments of the present invention, the beneficial effects that any of the speech processing methods provided in the embodiments of the present invention can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0126] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0127] In summary, although the present application has disclosed the preferred embodiments as described above, the above preferred embodiments are not intended to limit the present application. Those skilled in the art can make various modifications and refinements without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be determined by the scope defined in the claims.

Claims

1. A speech processing method, characterized in that, include: When a reverberant speech signal is detected, the reverberant speech signal is preprocessed to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum; the preprocessing includes frame segmentation, windowing, and fast Fourier transform processing; The reverberant speech features are determined based on the reverberant speech amplitude spectrum, and the reverberant speech features are input into the target de-reverberation network to obtain the de-reverberation ratio of the reverberant speech amplitude spectrum. The dereverberation speech amplitude spectrum is determined based on the dereverberation ratio and the reverberation speech amplitude spectrum. Based on the dereverberation speech amplitude spectrum and the reverberation speech phase spectrum, a clean speech signal with the reverberation signal filtered out is obtained from the reverberation speech signal. The step of determining the dederea speech amplitude spectrum based on the dederea ratio and the reverberant speech amplitude spectrum includes: Multiply the de-reverberation ratio and the reverberation speech amplitude spectrum to obtain the de-reverberation speech amplitude spectrum; The step of preprocessing the reverberant speech signal to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum when the reverberant speech signal is detected further includes: Acquire the reverberation time measurement signal to be processed and the clean speech signal to be processed; wherein, the reverberation time measurement signal to be processed includes a frequency sweep signal with multiple frequency bands; The reverberation time measurement signal to be processed and the clean speech signal to be processed are input into the reverberation simulator. The reverberation simulator adds a simulated signal with the same reverberation time to the reverberation time measurement signal and the clean speech signal to be processed, so as to obtain a simulated reverberation time measurement signal and a simulated reverberation speech signal with the same reverberation time. The simulated reverberation time measurement signal is processed by impulse response integration to obtain the reverberation time corresponding to each frequency band of the simulated reverberation time measurement signal. Based on the reverberation time corresponding to each frequency band, the simulated reverberation speech signal is processed by spectrum subtraction to obtain the target speech amplitude spectrum after de-reverberation. The ratio of the target speech amplitude spectrum to the amplitude spectrum of the simulated reverberation speech signal is used as the reference de-reverberation ratio. The initial dereverberation network is trained based on the reference reverberation speech features and the reference dereverberation ratio to obtain the target dereverberation network.

2. The speech processing method according to claim 1, characterized in that, The step of determining the reference reverberant speech features and the reference dérevering ratio based on the reverberation time measurement signal to be processed and the clean speech signal to be processed includes: The reverberation time measurement signal to be processed and the clean speech signal to be processed are input into the reverberation simulator, so that the reverberation simulator adds a simulated reverberation signal to the reverberation time measurement signal to be processed and the clean speech signal to be processed, thereby obtaining a simulated reverberation time measurement signal and a simulated reverberation speech signal. The simulated reverberation time measurement signal is processed by impulse response integration to obtain the reverberation time corresponding to each frequency band of the simulated reverberation time measurement signal; Based on the reverberation time and the simulated reverberant speech signal, the reference reverberant speech features and the reference de-reverberation ratio are obtained.

3. The speech processing method according to claim 2, characterized in that, The step of obtaining the reference reverberation speech features and the reference déreverberation ratio based on the reverberation time and the simulated reverberation speech signal includes: The simulated reverberant speech signal is preprocessed to obtain the simulated reverberant speech amplitude spectrum; Calculate the simulated reverberant speech log energy spectrum of the simulated reverberant speech amplitude spectrum, and use the simulated reverberant speech log energy spectrum as the reference reverberant speech feature; Based on the reverberation time, the simulated reverberation speech amplitude spectrum is subjected to spectrum subtraction processing to obtain the de-reverberation target speech amplitude spectrum; The ratio of the target speech amplitude spectrum to the simulated reverberant speech amplitude spectrum is used as the reference de-reverberation ratio.

4. The speech processing method according to claim 3, characterized in that, The step of preprocessing the reverberant speech signal to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum when a reverberant speech signal is detected includes: When a reverberant speech signal is detected, the reverberant speech signal is subjected to frame segmentation, windowing, and fast Fourier transform processing to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum.

5. The speech processing method according to claim 4, characterized in that, The step of determining reverberant speech features based on the reverberant speech amplitude spectrum and inputting the reverberant speech features into a target dreverberation network to obtain the dreverberation ratio of the reverberant speech amplitude spectrum includes: Calculate the reverberant speech log energy spectrum of the reverberant speech amplitude spectrum, and use the reverberant speech log energy spectrum as the reverberant speech feature; The reverberant speech features are input into a target dereverberation network to perform prediction processing on the reverberant speech features, thereby obtaining the dereverberation ratio of the reverberant speech amplitude spectrum.

6. A voice processing device, characterized in that, include: The preprocessing module is used to preprocess the reverberant speech signal when a reverberant speech signal is detected, to obtain the reverberant speech amplitude spectrum and the reverberant speech phase spectrum; the preprocessing includes frame segmentation, windowing, and fast Fourier transform processing; The de-reverberation ratio acquisition module is used to determine the reverberant speech features based on the reverberant speech amplitude spectrum, and input the reverberant speech features into the target de-reverberation network to obtain the de-reverberation ratio of the reverberant speech amplitude spectrum; The determining module is used to determine the dereverberation speech amplitude spectrum based on the dereverberation ratio and the reverberation speech amplitude spectrum; A clean speech signal acquisition module is used to obtain a clean speech signal from the reverberant speech signal after the reverberation signal has been filtered out, based on the dereverberant speech amplitude spectrum and the reverberant speech phase spectrum. The determining module is further configured to multiply the de-reverberation ratio and the reverberant speech amplitude spectrum to obtain the de-reverberation speech amplitude spectrum. The preprocessing module is further used to acquire a reverberation time measurement signal to be processed and a clean speech signal to be processed. The reverberation time measurement signal to be processed includes a frequency sweep signal with multiple frequency bands. The reverberation time measurement signal and the clean speech signal to be processed are input into a reverberation simulator, which adds a simulated signal with the same reverberation time to both signals, resulting in a simulated reverberation time measurement signal and a simulated reverberation speech signal with the same reverberation time. The simulated reverberation time measurement signal undergoes impulse response integration processing to obtain the reverberation time corresponding to each frequency band. Based on the reverberation time corresponding to each frequency band, the simulated reverberation speech signal undergoes spectrum subtraction processing to obtain the de-reverberated target speech amplitude spectrum. The ratio of the target speech amplitude spectrum to the amplitude spectrum of the simulated reverberation speech signal is used as a reference de-reverberation ratio. Based on the reference reverberation speech features and the reference de-reverberation ratio, an initial de-reverberation network is trained to obtain the target de-reverberation network.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the speech processing method according to any one of claims 1 to 5.

8. An electronic device, characterized in that, The method includes a processor and a memory, the processor being electrically connected to the memory, the memory being used to store instructions and data, and the processor being used to execute the steps of the speech processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice signal dereverberation processing method and device, computer equipment and storage medium

    CN111489760A