Voice signal processing method, apparatus, device and storage medium
By processing voice signals in the frequency domain, calculating the gain values of each frequency point, and generating the target voice signal, the problem of instability in speech quality caused by cascade encoding is solved, and the quality and intelligibility of voice signals are improved.
Patent Information
- Application Number
- CN202110226589.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-03-01
AI Technical Summary
Cascade encoding leads to uneven speech signal damage, and the prior art uses the same enhancement amplitude compensation to lead to unstable speech quality.
By converting the voice signal from the time domain to the frequency domain, obtaining the power spectrum and phase information of each frequency point, calculating the frequency band gain value of individual frequency points, and generating a target voice signal that meets the voice playback conditions.
It realizes stable enhancement of voice signals, improves voice quality and intelligibility, and is suitable for uncascaded encoding and damaged voice signals.
Smart Images

Figure CN113707162B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to a method, device, equipment and storage medium for processing voice signals. Background Art
[0002] With the rapid development of mobile communication technology and Internet technology, various application programs with communication functions have emerged. Users can make voice calls through such application programs installed on the terminal. In order to enable voice docking between terminals located in different networks, there will be multiple encoding and decoding operations in the call link, that is, cascade encoding. However, the more times of cascade encoding, the more serious the damage to the voice signal, resulting in the inability of both parties in the voice call to clearly hear the speech content of the other party, that is, the voice intelligibility decreases.
[0003] The solutions of related technologies to solve the above problems are usually as follows: perform formant search on the voice signal after cascade encoding, then extract the formants of the damaged voice signal from the searched formants, and enhance such formants with the same enhancement amplitude to achieve compensation for the damaged voice signal.
[0004] However, after the voice signal undergoes cascade encoding, the damage degrees of voice signals at different frequencies are often inconsistent. The above solution uses the same enhancement amplitude, that is, the compensation for voice signals with different damage degrees is the same. This will make the enhancement effect of the damaged voice signal unstable and unable to effectively improve the voice quality. Summary of the Invention
[0005] Embodiments of the present application provide a method, device, equipment and storage medium for processing voice signals, which effectively improve the voice quality and thus enhance the voice intelligibility. The technical solution is as follows:
[0006] On the one hand, a method for processing voice signals is provided. The method includes:
[0007] Transform the voice signal to be processed from the time domain to the frequency domain, and obtain the first power spectrum and phase information of each frequency point in the frequency domain; wherein, the voice signal to be processed is an initial voice signal or a damaged voice signal. The initial voice signal refers to a voice signal that has not undergone cascade encoding processing, and the damaged voice signal refers to a voice signal obtained after the cascade encoding processing;
[0008] Obtain the frequency band gain value of each frequency point, and determine the second power spectrum of each frequency point based on the first power spectrum and the frequency band gain value of each frequency point;
[0009] Generate a target voice signal that meets the voice playback condition based on the phase information and the second power spectrum of each frequency point.
[0010] On the other hand, a voice signal processing device is provided, which includes:
[0011] An acquisition module, configured to transform the voice signal to be processed from the time domain to the frequency domain, and acquire the first power spectrum and phase information of each frequency point in the frequency domain; wherein, the voice signal to be processed is an initial voice signal or a damaged voice signal, the initial voice signal refers to a voice signal that has not undergone concatenated coding processing, and the damaged voice signal refers to a voice signal obtained after the concatenated coding processing;
[0012] A determination module, configured to acquire the frequency band gain value of each frequency point, and determine the second power spectrum of each frequency point based on the first power spectrum and the frequency band gain value of each frequency point;
[0013] A generation module, configured to generate a target voice signal meeting the voice playback condition based on the phase information and the second power spectrum of each frequency point.
[0014] In an optional implementation manner, in response to the voice signal to be processed being the damaged voice signal, the device further includes:
[0015] A processing module, configured to perform the concatenated coding processing on the initial voice signal before transforming the voice signal to be processed from the time domain to the frequency domain, to obtain the damaged voice signal.
[0016] In an optional implementation manner, the device further includes a training module, and the training module is configured to:
[0017] Acquire the third power spectrum of each frequency point of the voice sample in the frequency domain, where the third power spectrum is obtained by transforming the voice sample from the time domain to the frequency domain;
[0018] Input the third power spectrum corresponding to the voice sample into the initial neural network to obtain the predicted frequency band gain value corresponding to the third power spectrum;
[0019] Construct a loss function based on the predicted frequency band gain value and the target frequency band gain value of the voice sample;
[0020] Based on the loss function, continuously adjust the network parameters of the initial neural network until a preset condition is met, to obtain the target neural network;
[0021] Wherein, the target frequency band gain value is obtained based on the third power spectrum and the fourth power spectrum corresponding to the voice sample, and the fourth power spectrum is obtained by performing the concatenated coding processing on the voice sample and then transforming the voice sample from the time domain to the frequency domain.
[0022] In an optional implementation manner, the target frequency band gain value is the square root value of the ratio of the third power spectrum to the fourth power spectrum.
[0023] In an alternative implementation, the determining module is further configured to:
[0024] Input the first power spectrum of each frequency point into the first fully connected layer, and after the first fully connected layer extracts features from the first power spectrum of each frequency point, obtain a feature vector;
[0025] Input the feature vector into the gated recurrent unit layer, and after the update gate and reset gate in the gated recurrent unit layer, extract the correlation and valid information between the feature vectors to obtain an output vector;
[0026] Input the output vector into the second fully connected layer, and after the second fully connected layer integrates the output vector into the frequency band gain values of each frequency point.
[0027] In an alternative implementation, the obtaining module is further configured to:
[0028] Perform frame splitting processing and windowing processing on the to-be-processed speech signal in sequence;
[0029] Perform fast Fourier transform on the to-be-processed speech signal after frame splitting processing and windowing processing; based on the obtained transformation result, determine the first power spectrum and phase information of each frequency point in the frequency domain.
[0030] In an alternative implementation, the cascaded encoding process includes M times of encoding and decoding processes, where M is a positive integer greater than 1, and the processing module is further configured to:
[0031] Perform M times of encoding and decoding processes on the initial speech signal to obtain the damaged speech signal;
[0032] Wherein, the output of the previous encoding and decoding process is used as the input of the next encoding and decoding process; for any encoding and decoding process, the encoding and decoding process includes one encoding process and one decoding process, and the output of the encoding process is used as the input of the decoding process.
[0033] On the other hand, a computer device is provided, which includes a processor and a memory. The memory is used to store at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the speech signal processing method in the embodiments of the present application.
[0034] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the operations performed in the speech signal processing method in the embodiments of the present application.
[0035] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device executes the voice signal processing method provided in the above various alternative implementation manners.
[0036] For a voice signal to be processed, in the embodiments of the present application, first, the first power spectrum and phase information of each frequency point of such a voice signal in the frequency domain are obtained, and then by obtaining the frequency band gain value corresponding to each frequency point, the first power spectrum is enhanced to obtain the second power spectrum of each frequency point, and further, a target voice signal meeting the voice playback condition is generated according to the second power spectrum and phase information of each frequency point. Since this processing method enhances the power spectrum of each frequency point in a targeted manner, the enhancement effect of the voice signal is more stable, effectively improving the voice quality and further enhancing the voice intelligibility; moreover, regardless of whether the voice signal to be processed has been subjected to concatenated coding processing previously, this processing method can be used to enhance such a voice signal, and the applicable range is wide. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0038] Figure 1 is a schematic diagram of an implementation environment of a voice signal processing method provided in an embodiment of the present application;
[0039] Figure 2 is a flowchart of a voice signal processing method provided in an embodiment of the present application;
[0040] Figure 3 is a schematic diagram of a voice signal processing solution provided in an embodiment of the present application;
[0041] Figure 4 is a flowchart of another voice signal processing method provided in an embodiment of the present application;
[0042] Figure 5 is a flowchart of another voice signal processing method provided in an embodiment of the present application;
[0043] Figure 6 is a schematic diagram of another voice signal processing solution provided in an embodiment of the present application;
[0044] Figure 7 is a flowchart of another method for processing voice signals provided according to an embodiment of the present application;
[0045] Figure 8 is a flowchart of another method for processing voice signals provided according to an embodiment of the present application;
[0046] Figure 9 is a schematic structural diagram of a voice processing device provided according to an embodiment of the present application;
[0047] Figure 10 is a schematic structural diagram of a terminal provided according to an embodiment of the present application;
[0048] Figure 11 is a schematic structural diagram of a server provided according to an embodiment of the present application. Detailed implementation manners
[0049] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0050] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0051] In the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions. It should be understood that there is no logical or temporal dependence between "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms.
[0052] These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first power spectrum can be referred to as the second power spectrum, and similarly, the second power spectrum can also be referred to as the first power spectrum. The first power spectrum and the second power spectrum can both be power spectra, and in some cases, they can be separate and different power spectra.
[0053] Among them, "at least one" means one or more. For example, at least one frequency point can be one frequency point, two frequency points, three frequency points, etc., which are any integer greater than or equal to one. And "a plurality of" means two or more. For example, a plurality of frequency points can be two frequency points, three frequency points, etc., which are any integer greater than or equal to two.
[0054] Next, the technologies that may be used in the voice signal processing solution provided by the embodiments of the present application will be introduced.
[0055] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0056] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0057] Machine Learning (ML) is an interdisciplinary subject in multiple fields, involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or realizes human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make a computer intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0058] Next, the key terms or abbreviations that may be used in the voice signal processing solution provided by the embodiments of the present application will be introduced.
[0059] Voice over Internet Protocol (VoIP): A voice call technology that enables voice calls via the Internet Protocol (IP), i.e., communication via the Internet. Other informal names include IP telephony, Internet Telephony, Broadband Telephony, and Broadband Phone Service. VoIP can be used in many Internet access devices including VoIP phones, smartphones, and personal computers for making calls and sending text messages via cellular networks and wireless networks.
[0060] Fast Fourier Transform (FFT): A method for quickly calculating the discrete Fourier transform of a sequence or its inverse transform. Fourier analysis converts a signal from the original domain (usually the time domain or spatial domain) to a representation in the frequency domain or vice versa. Correspondingly, converting a signal from the frequency domain to the original domain is called the Inverse Fast Fourier Transform (IFFT).
[0061] Frequency domain: A coordinate system used to describe the characteristics of a signal in terms of frequency. In electronics, control systems engineering, and statistics, a frequency domain plot shows the amount of signal in each given frequency band within a frequency range. The frequency domain representation can also include the phase information of each sine curve to enable recombination of the frequency components to recover the original time signal.
[0062] Power spectrum: Short for power spectral density function, which is defined as the signal power per unit frequency band. It represents how the signal power changes with frequency, i.e., the distribution of signal power in the frequency domain. The power spectrum shows the relationship between signal power and frequency.
[0063] Gated Recurrent Unit (GRU): A variant of the Recurrent Neural Network (RNN) structure that not only effectively solves the problems of gradient vanishing and gradient explosion in traditional RNNs, but also learns the hidden state more quickly. Its structure is simpler than that of the Long Short-Term Memory (LSTM) network and has a faster training speed.
[0064] Concatenated coding: For a system with multiple encodings (at least two times), each level of encoding is regarded as a whole encoding, which is called a concatenated code. The concatenated code divides the encoding process into several levels to meet the requirements of channel error correction for the encoding length, obtaining error correction capabilities and high coding gains that are close to or even the same as those of long codes. Moreover, the associated increase in encoding and decoding complexity is not very large. That is to say, if a system includes multiple encodings, these multiple encodings are considered concatenated coding. In relevant transmission networks, concatenated codes can balance the coding gain performance and the encoding and decoding complexity, and thus are widely used. Concatenated codes can be implemented in the form of a combination of two or more encoding methods.
[0065] The following introduces the implementation environment related to the voice signal processing method provided in the embodiments of this application.
[0066] Refer to Figure 1 , Figure 1 FIG. is a schematic diagram of the implementation environment of the voice signal processing method provided in the embodiments of this application. The implementation environment includes: a terminal 101 and a server 102.
[0067] The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, which are not limited in this application. Optionally, the terminal 101 is a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. The terminal 101 can install and run an application program. Optionally, the application program is a social application program, an online meeting application program, or a voice call application program, etc. Schematically, the terminal 101 is a terminal used by a user, and a user account of the user is logged in to the application program running in the terminal 101. For example, a social application program is running on the terminal 101, and the social application program provides a voice call function, and users can make voice calls with each other through the social application program.
[0068] The server 102 can be an independent physical server, or can be a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The server 102 is used to provide background services for the application program running on the terminal 101.
[0069] Optionally, during the process of voice signal processing, server 102 undertakes the main computing work and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work and terminal 101 undertakes the main computing work; or, server 102 or terminal 101 can separately undertake the computing work.
[0070] Optionally, terminal 101 generally refers to one of multiple terminals, and only terminal 101 is used as an example in this embodiment. Those skilled in the art can know that the number of the above-mentioned terminals 101 can be more. For example, the above-mentioned terminal 101 is dozens or hundreds, or a larger number. At this time, the implementation environment of the above-mentioned voice signal processing method further includes other terminals. The embodiments of the present application do not limit the number and device type of the terminals.
[0071] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network. In some embodiments, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can also be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.
[0072] Schematically, the application scenarios of the voice signal processing method provided by the embodiments of the present application include but are not limited to the following several scenarios:
[0073] Scenario 1: Inter-fusion application scenario between different networks
[0074] With the popularization of VoIP services, the interoperable applications between different networks are increasing. For example, IP telephones over the Internet are interoperable with landline telephones over the Public Switched Telephone Network (PSTN), and IP telephones are interoperable with mobile phones on wireless networks, etc. Different networks use different voice encoding and decoding methods for voice. For example, the Global System for Mobile Communication (GSM) network uses AMR-NB encoding, landline telephones use G.711 encoding, and IP telephones use G.729 encoding, etc. Since the voice encoding formats supported by each network terminal are inconsistent, it will inevitably lead to cascaded encoding in the call link. In this scenario, by using the voice signal processing method provided in this application, the initial voice signal collected can be enhanced, or the damaged voice signal after cascaded encoding processing can be enhanced, so that the finally played voice signal is closer to the initial voice signal collected, which can effectively improve the voice quality in this application scenario and thus enhance the voice intelligibility.
[0075] Scenario 2: Multi-person voice call scenario
[0076] Currently, many application programs running on terminals can provide the function of multi-person voice calls, such as group audio and video calls and online audio and video conferences, etc. In this multi-person voice call scenario, since the voice signals of each calling party must go through a process of decoding and encoding once when they are sent to the mixing server for mixing processing after being collected, encoded, and compressed from each terminal, this also belongs to a kind of cascaded encoding processing, and the voice signals will be damaged accordingly, resulting in a decline in voice quality. By using the voice signal processing method provided in this application, the initial voice signals of each calling party can be enhanced, or the damaged voice signals after cascaded encoding processing can be enhanced, so that the finally played voice signal is closer to the initial voice signal collected, which can effectively improve the voice quality in this application scenario and thus enhance the voice intelligibility.
[0077] Scenario 3: Live broadcast scenario
[0078] With the development of Internet technology and the wide application of live broadcast services, people can watch live broadcasts through different types of terminals. Due to differences in the coding formats of different terminals, transcoding processing needs to be performed on the voice signals in the live broadcast, and this transcoding processing also belongs to a type of concatenated coding processing, resulting in damage to the voice signals in the live broadcast and affecting the live broadcast quality. By using the voice signal processing method provided in this application, the initial voice signal during the live broadcast process can be enhanced, or the damaged voice signal after concatenated coding processing can be enhanced, making the finally played voice signal closer to the initially collected voice signal, effectively improving the voice quality in this application scenario and further enhancing the voice intelligibility.
[0079] Embodiments of this application provide a voice signal processing method. Among them, the voice signal to be processed can be either an initial voice signal that has not undergone concatenated coding processing or a damaged voice signal that has undergone concatenated coding processing. Generally, since the human ear is sensitive to sound energy, the auditory sensations of voice signals in different frequency bands are quite different, and the subjective damage of concatenated coding processing to sound is most directly manifested in the damage to the frequency domain subbands. For example, after the initial voice signal undergoes multiple concatenated coding processes, the high-frequency part of the signal is significantly attenuated, resulting in a blurred sound perception and a low recognition rate of the sound. Embodiments of this application make full use of the above characteristics, use a target neural network to learn the damage problem caused by concatenated coding processing to the voice signal to be processed, and through this target neural network, obtain the frequency band gain values of each frequency point of the voice signal to be processed in the frequency domain, so as to enhance the voice signal in the frequency domain subbands and achieve the purpose of improving the voice quality to further enhance the voice intelligibility. For a more specific description, please refer to the following embodiments.
[0080] Figure 2 is a flowchart of a voice signal processing method provided according to an embodiment of this application. Among them, the execution subject of this voice signal processing method is a computer device. Schematically, this computer device is Figure 1 the terminal 101 or the server 102 in, and embodiments of this application do not limit this. The following refers to Figure 2 to illustrate by taking the application of the embodiments of this application to a terminal as an example. As Figure 2 shown, this voice signal processing method includes the following steps:
[0081] 201. Transform the voice signal to be processed from the time domain to the frequency domain, and obtain the first power spectrum and phase information of each frequency point in the frequency domain; among them, the voice signal to be processed is an initial voice signal or a damaged voice signal. The initial voice signal refers to a voice signal that has not undergone concatenated coding processing, and the damaged voice signal refers to a voice signal obtained after undergoing this concatenated coding processing.
[0082] In the embodiments of the present application, the voice signal to be processed may be the voice signal of a certain speaker, or may be the voice signal received in a certain scenario. For example, the terminal collects the voice signal of the speaker in real time through a microphone. For another example, the terminal receives a voice signal in a live broadcast scenario or an online meeting scenario. The embodiments of the present application do not limit the acquisition method of the voice signal to be processed.
[0083] Optionally, the cascade encoding process includes M times of encoding and decoding processes, where M is a positive integer greater than 1. Wherein, the output of the previous encoding and decoding process is used as the input of the next encoding and decoding process; for any encoding and decoding process, the encoding and decoding process includes one encoding process and one decoding process, and the output of the encoding process is used as the input of the decoding process.
[0084] Schematically, taking the scenario of mutual integration between different networks as an example, when the voice signal to be processed passes through the actual link, it needs to go through multiple encoding and decoding processes. For example, if an IP phone supporting G.729 communicates with a GSM mobile phone, the above cascade encoding process includes two encoding and decoding processes: G.729 encoding process + G.729 decoding process + AMR-NB encoding process + AMR-NB decoding process. The embodiments of the present application do not limit the number and type of cascade encoding processes.
[0085] Optionally, the terminal transforms the voice signal to be processed from the time domain to the frequency domain, and obtains the first power spectrum and phase information of each frequency point in the frequency domain, including but not limited to the following steps 2011 and step 2012.
[0086] 2011. Perform frame segmentation processing and windowing processing on the voice signal to be processed in sequence.
[0087] Among them, the voice signal to be processed is a series of ordered signals, which is non-stationary macroscopically but stationary microscopically, that is, the voice signal to be processed has short-term stationarity (for example, within 10 ms to 30 ms, the voice signal to be processed can be considered approximately unchanged). Based on this characteristic, the voice signal to be processed can be divided into some short segments for processing, where each short segment can be called a frame, that is, an audio frame. Schematically, the playing duration of an audio frame can be 16 ms, 46.64 ms, 128 ms, etc., and the embodiments of the present application do not limit this.
[0088] Optionally, when the terminal performs frame segmentation processing on the voice signal to be processed, in order to ensure the smoothness and continuity of the transition between adjacent audio frames, it is also necessary to ensure that there is an overlap between frames, where the overlapping part between two adjacent frames is called the frame shift.
[0089] Optionally, when the terminal performs windowing processing on the speech signal to be processed, an analysis window of 10 ms or 20 ms can be used. Among them, the window function can be selected from a Hanning window, a Hamming window, a rectangular window, etc., and the embodiments of the present application do not limit this. That is, after windowing, multiple analysis windows will be formed, and each time only the speech signal to be processed within one analysis window can be processed. Here, it should be understood that the windowing processing makes the speech signal to be processed periodic to reduce the speech energy leakage of the speech signal to be processed in the subsequent FFT.
[0090] 201. Perform FFT on the speech signal to be processed after frame segmentation processing and windowing processing; based on the obtained transformation result, determine the first power spectrum and phase information of each frequency point in the frequency domain.
[0091] Among them, for the speech signal to be processed after frame segmentation processing and windowing processing, the terminal performs N-point FFT (where N is a positive integer) on the speech signal to be processed according to the number of FFT points being N and the number of frequency points being K (where K is a positive integer), and obtains the FFT transformation result, that is, the spectrogram. Then, the terminal can calculate the power spectrum value of each frequency point according to the amplitude corresponding to each frequency point in the spectrogram, and obtain the phase information of each frequency point. In the embodiments of the present application, the power spectrum value of each frequency point is referred to as the first power spectrum.
[0092] Schematically, take the number of FFT points N as 256 points and K as 129. That is, the terminal performs 256-point FFT on a certain audio frame of the speech signal to be processed, and can obtain the power spectrum values of 129 frequency points. It should be noted that the number of FFT points and the number of frequency points can be set according to actual needs, and the embodiments of the present application do not limit this.
[0093] Through the above step 201, the terminal performs frequency domain conversion on the speech signal to be processed, and obtains the first power spectrum and phase information of each frequency point, providing a basis for obtaining the frequency band gain value of each frequency point subsequently.
[0094] 202. Obtain the frequency band gain value of each frequency point, and based on the first power spectrum and the frequency band gain value of each frequency point, determine the second power spectrum of each frequency point.
[0095] In the embodiments of the present application, the frequency band gain value is the enhancement amplitude required for the first power spectrum of each frequency point, and can also be understood as the gain value required to enhance the first power spectrum of each frequency point. Among them, enhancement can essentially be understood as enhancing the speech signal to be processed to improve the speech quality. The second power spectrum is obtained by enhancing the first power spectrum of each frequency point.
[0096] Schematically, the band gain value can also be used to measure the degree of damage to the speech signals corresponding to each frequency point, that is, the higher the band gain value of a certain frequency point, the more serious the degree of damage to the speech signal corresponding to that frequency point.
[0097] Optionally, the terminal uses the product of the first power spectrum of each frequency point and the band gain value as the second power spectrum of each frequency point. For example, the first power spectrum of a certain frequency point is 30 dB / Hz, and the terminal obtains the band gain value corresponding to this frequency point as 1.10, then the second power spectrum of this frequency point is 33 dB / Hz.
[0098] It should be noted that the above examples of the band gain value are only schematic. In some embodiments, the band gain value can also be expressed in percentage form. For example, the band gain value is 110% etc., and the embodiments of the present application do not limit this.
[0099] Optionally, the terminal obtains the band gain value of each frequency point through deep learning. Schematically, this step 202 can be replaced by the following steps 2021 and 2022.
[0100] 2021. Input the first power spectrum of each frequency point into the target neural network to obtain the band gain value of each frequency point; wherein, the target neural network includes a first fully connected layer, a gated recurrent unit (GRU) layer, and a second fully connected layer connected in sequence.
[0101] Among them, the target neural network is a neural network based on deep learning. Optionally, the target neural network adopts a four-layer network structure and has the following characteristics: the input layer of the target neural network adopts a fully connected layer, that is, the first fully connected layer; the output layer adopts a fully connected layer, that is, the second fully connected layer; the intermediate hidden layer has two layers and adopts a GRU layer, and the hidden layers are connected in sequence.
[0102] Schematically, in the embodiments of the present application, the number of neurons in the first fully connected layer and the GRU layer is set to 64; the number of neurons in the second fully connected layer is set to 129. In other embodiments, the number of neurons in the first fully connected layer and the second fully connected layer is set to 129, and the number of neurons in the GRU layer is set to 64. The embodiments of the present application do not limit this.
[0103] Schematically, in the embodiments of the present application, the activation function of the first fully connected layer is set to the tanh function; the activation functions of the GRU layer are set to the relu and sigmoid functions; the activation function of the second fully connected layer is set to the sigmoid function. The embodiments of the present application do not limit the types of activation functions of each layer of the network in the target neural network.
[0104] It should be noted that the structure of the target neural network can be flexibly adjusted according to the actual situation. Among them, the adjustment methods include but are not limited to: adjusting the connection method between layers, changing the feature input dimension, the number of neurons, the hidden layer type, and the activation function type of each layer, etc. The embodiments of the present application do not limit this.
[0105] The following elaborates on this step 2021 in detail, including the following steps 2021-1 to 2021-3.
[0106] 2021-1. Input the first power spectrum of each frequency point into the first fully connected layer. After the first fully connected layer extracts features from the first power spectrum of each frequency point, a feature vector is obtained.
[0107] Among them, the first fully connected layer serves as the input layer of the target neural network and can extract features from the first power spectrum of each frequency point through a feature extraction function. This feature vector is used to represent the power spectrum features of the first power spectrum.
[0108] 2021-2. Input the feature vector into the GRU layer. Through the update gate and reset gate in the GRU layer, the correlation and effective information between the feature vectors are extracted to obtain an output vector.
[0109] Among them, GRU is a commonly used gated recurrent neural network. Its input is the input at the current moment and the hidden state at the previous moment, that is, the output vector will be affected by the information at the current moment T and the information at the previous T-1 moments, where T is greater than 1. The GRU layer includes two gate functions: the update gate and the reset gate. Among them, the update gate is used to control the degree to which the state information at the previous moment is brought into the current state. The larger the value of the update gate, the more state information at the previous moment is brought in; the reset gate controls how much information from the previous state is written into the current output vector. The smaller the reset gate, the less information from the previous state is written in.
[0110] Since the speech signal belongs to time series features, after the terminal inputs the feature vectors of each frequency point into the GRU layer, the GRU layer can extract the correlation and effective information between the feature vectors of each frequency point, thereby obtaining an output vector. Schematically, the GRU layer combines the feature vector corresponding to the current frequency point with the output vector of the previous frequency point retained before, and through the processing of the update gate and the reset gate, generates an output vector for the current frequency point, and so on and iterates continuously.
[0111] Through the update gate and reset gate of the GRU layer, it can be determined which feature vectors can ultimately serve as the output vector of the GRU layer. These two gating mechanisms can preserve the information in the long-term sequence and will not be cleared over time or removed because they are not relevant to the prediction, ensuring the reliability of the target neural network.
[0112] In 2021, input the output vector into the second fully connected layer, and through the second fully connected layer, integrate the output vector into the band gain values of each frequency point.
[0113] Among them, the second fully connected layer serves as the output layer. Each neuron in this layer is fully connected to all neurons in the previous layer. Based on this fully connected method, the second fully connected layer can integrate the output vector output by the GRU layer, and finally obtain the band gain values of each frequency point.
[0114] In 2022, take the product of the first power spectrum of each frequency point and the band gain value of each frequency point as the second power spectrum of each frequency point.
[0115] Through the above steps 2021 and 2022, the terminal obtains the band gain values of each frequency point specifically through the target neural network based on deep learning, making the enhancement effect of the voice signal more stable.
[0116] In 203, generate a target voice signal that meets the voice playback conditions based on the phase information and the second power spectrum of each frequency point.
[0117] In the embodiments of the present application, the voice playback condition means that the voice quality of the voice signal reaches a preset requirement. The terminal performs an N-point IFFT based on the phase information and the second power spectrum of each frequency point with the IFFT point number being N and the frequency point number being K to obtain the IFFT transformation result, that is, generate the target voice signal. Since FFT and IFFT are two inverse transformation methods, the implementation method of FFT has been introduced in the above step 201, so the implementation method of IFFT will not be elaborated here.
[0118] Among them, this step 203 includes the following two situations:
[0119] Situation 1: When the terminal responds that the voice signal to be processed is an initial voice signal, this step 203 can be replaced by the following steps 2031 and 2032.
[0120] In 2031, generate an intermediate voice signal based on the phase information and the second power spectrum of each frequency point.
[0121] In 2032, perform concatenated coding processing on the intermediate voice signal to obtain the target voice signal.
[0122] Situation 2: When the terminal responds that the voice signal to be processed is a damaged voice signal, before the terminal executes the above step 201, execute the following step "perform concatenated coding processing on the initial voice signal to obtain the damaged voice signal", and then the terminal sequentially executes the above steps 201 to 203.
[0123] Optionally, the above step of "performing cascaded encoding processing on the initial voice signal to obtain a damaged voice signal" can also be replaced with "performing M times of encoding and decoding processing on the initial voice signal to obtain a damaged voice signal".
[0124] Optionally, before the terminal executes the above step 2021, the target neural network is trained with a large number of voice samples, and finally the above target neural network is obtained. Schematically, the above step 2021 further includes the training process of the target neural network, including the following steps 2021-4 to step 2021-7.
[0125] 2021-4. Obtain the third power spectrum of each frequency point of the voice sample in the frequency domain, and the third power spectrum is obtained by transforming the voice sample from the time domain to the frequency domain.
[0126] Among them, the voice samples include but are not limited to: the voice signal of a certain speaker, the voice signal in a certain video, and the voice signal collected in a certain specific scenario, etc. Usually, the number of voice samples is usually multiple. Training the target neural network with a large number of voice samples can make the finally trained target neural network have good universality and robustness.
[0127] In this step 2021-4, the terminal performs frame splitting processing, windowing processing, and FFT on the voice sample in sequence to obtain the FFT transformation result, realizes the frequency domain conversion of the voice sample, and obtains the third power spectrum of each frequency point of the voice sample in the frequency domain.
[0128] 2021-5. Input the third power spectrum corresponding to the voice sample into the initial neural network to obtain the predicted frequency band gain value corresponding to the third power spectrum.
[0129] Among them, the voice sample includes a voice signal marked with the target frequency band gain value corresponding to the third power spectrum. The terminal obtains the predicted frequency band gain value corresponding to the third power spectrum based on the network parameters of the initial neural network.
[0130] 2021-6. Construct a loss function based on the predicted frequency band gain value and the target frequency band gain value of the voice sample.
[0131] Among them, the target frequency band gain value is obtained based on the third power spectrum and the fourth power spectrum corresponding to the voice sample, and the fourth power spectrum is obtained by performing cascaded encoding processing on the voice sample and then transforming the voice sample from the time domain to the frequency domain.
[0132] Optionally, the ways for the terminal to construct the loss function include, but are not limited to: constructing the loss function by using the difference between the predicted band gain value and the target band gain value of the voice sample, constructing the loss function by using the ratio between the predicted band gain value and the target band gain value of the voice sample, constructing the loss function by using the product value between the predicted band gain value and the target band gain value of the voice sample, and so on.
[0133] In addition, the loss function in the embodiments of the present application can be various loss functions commonly used in neural network training, such as absolute value loss function, cosine similarity loss function, square loss function, cross entropy loss function, etc. The embodiments of the present application do not limit this.
[0134] Optionally, the target band gain value is the square root value of the ratio of the third power spectrum to the fourth power spectrum. Schematically, refer to the following formula (1):
[0135] Target band gain value = sqrt(E_org(i) / E_deg(i)) (1)
[0136] In the formula, sqrt is the square root function; E_org(i) represents the original voice power spectrum of the i-th frequency point of each audio frame after the voice sample undergoes FFT for frequency domain transformation, that is, the third power spectrum; E_deg(i) represents the degraded voice power spectrum of the i-th frequency point of each audio frame after the voice sample undergoes concatenated coding processing and FFT for frequency domain transformation, that is, the fourth power spectrum.
[0137] 2021-7. Based on this loss function, continuously adjust the network parameters of the initial neural network until the preset conditions are met to obtain the target neural network.
[0138] Among them, the preset condition is that the loss value (also called the error value) is less than the set threshold, and this set threshold can be set according to actual needs, such as setting according to the value accuracy of the target neural network. The present application does not limit this. In addition, in response to the loss function not meeting the preset conditions, adjust the network parameters of the current neural network, and then start executing from the above step 2021-4 again until the loss function meets the preset conditions and stop training to obtain the target neural network.
[0139] It should be noted that the training process of the above target neural network may also include other steps or other optional implementation manners. The present application does not limit this.
[0140] In addition, the target neural network in the embodiments of the present application is not limited to the above types. Any network based on machine learning or deep learning and for obtaining the band gain values of each frequency point can be used as the target neural network in the embodiments of the present application.
[0141] For the speech signal to be processed, in the embodiments of the present application, first, the first power spectrum and phase information of each frequency point of such a speech signal in the frequency domain are obtained. Then, by obtaining the frequency band gain value corresponding to each frequency point, the first power spectrum is enhanced to obtain the second power spectrum of each frequency point. Furthermore, a target speech signal meeting the speech playback condition is generated based on the second power spectrum and phase information of each frequency point. Since this processing method enhances the power spectrum of each frequency point in a targeted manner, the enhancement effect of the speech signal is more stable, effectively improving the speech quality and thus enhancing the speech intelligibility. Moreover, regardless of whether the speech signal to be processed has been previously cascade-coded, this processing method can be used to enhance such a speech signal, with a wide range of applications.
[0142] It should be noted that the above Figure 2 shown speech signal processing method covers two speech signal processing schemes. One is the speech signal processing method when the speech signal to be processed is an initial speech signal, and the other is the speech signal processing method when the speech signal to be processed is a damaged speech signal. Below, based on two specific embodiments, the two speech signal processing schemes provided by the present application will be respectively described schematically.
[0143] First, the speech signal processing scheme when the speech signal to be processed is an initial speech signal.
[0144] First, referring to Figure 3 , Figure 3 is a schematic diagram of a speech signal processing scheme provided by the embodiments of the present application. As Figure 3 shown, the terminal performs deep learning preprocessing on the initial speech signal, and then performs cascade coding processing on the initial speech signal after deep learning preprocessing to finally obtain the target speech signal. Among them, deep learning preprocessing is also a process in which the terminal obtains the frequency band gain value corresponding to the initial speech signal through a target neural network and then enhances the initial speech signal.
[0145] Next, referring to Figure 4 , Figure 4 is a flowchart of another speech signal processing method provided by the embodiments of the present application. Below, in combination with Figure 4 , this speech signal processing scheme will be elaborated in detail. As Figure 4As shown in the figure, first, the terminal performs FFT on the initial voice signal to obtain the power spectrum and phase information of each frequency point in the frequency domain of the initial voice signal. Then, the terminal inputs the power spectrum of each frequency point into the target neural network. After the target neural network processes the power spectrum, it obtains the frequency band gain value corresponding to each frequency point. Among them, the target neural network includes two fully connected layers and two GRU layers. Next, the terminal multiplies the power spectrum of each frequency point by the frequency band gain value corresponding to each frequency point to obtain the enhanced power spectrum of each frequency point. Finally, based on the phase information of each frequency point and the enhanced power spectrum of each frequency point, after IFFT to obtain the preprocessed voice signal, the terminal performs cascade coding processing on the preprocessed voice signal, and finally obtains the target voice signal. Among them, the preprocessed voice signal is also the intermediate voice signal shown in the above embodiment.
[0146] Finally, referring to Figure 5 , Figure 5 is a flowchart of another voice signal processing method provided according to an embodiment of the present application. As Figure 5 shown, the voice signal processing method includes the following steps 501 to step 504.
[0147] 501. Transform the initial voice signal from the time domain to the frequency domain, and obtain the first power spectrum and phase information of each frequency point in the frequency domain; where the initial voice signal refers to the voice signal that has not undergone cascade coding processing.
[0148] 502. Obtain the frequency band gain value of each frequency point, and determine the second power spectrum of each frequency point based on the first power spectrum and the frequency band gain value of each frequency point.
[0149] 503. Generate an intermediate voice signal based on the phase information and the second power spectrum of each frequency point.
[0150] 504. Perform cascade coding processing on the intermediate voice signal to obtain a target voice signal that meets the voice playback condition.
[0151] It should be noted that by performing deep learning preprocessing on the initial voice signal, a preprocessed voice signal can be obtained. After this preprocessed voice signal undergoes cascade coding processing, a voice signal closer to the initial voice signal can be obtained, that is, it can be restored to a better sound quality, effectively improving the voice quality in the cascade coding application scenario, and thus enhancing the voice intelligibility.
[0152] Second, the voice signal processing solution when the voice signal to be processed is a damaged voice signal.
[0153] First, referring to Figure 6 , Figure 6 is a schematic diagram of another voice signal processing solution provided according to an embodiment of the present application. As Figure 6As shown, the terminal first performs concatenated coding on the initial speech signal to obtain a damaged speech signal, which is also called a degraded speech signal; then the terminal performs deep learning repair on the damaged speech signal to finally obtain the target speech signal. Among them, deep learning repair is a process in which the terminal obtains the band gain value corresponding to the damaged speech signal through a target neural network and then enhances the damaged speech signal.
[0154] Next, referring to Figure 7 , Figure 7 is a flowchart of another speech signal processing method provided according to an embodiment of the present application. The following combines Figure 7 to elaborate on this speech signal processing solution in detail. As Figure 7 shown, first, the terminal performs concatenated coding on the initial speech signal to obtain a damaged speech signal; then, the terminal performs FFT on the damaged speech signal to obtain the power spectrum and phase information of each frequency point in the frequency domain of the damaged speech signal; then, the terminal inputs the power spectrum of each frequency point into the target neural network, and after the target neural network processes the power spectrum, it obtains the band gain value corresponding to each frequency point. Among them, the target neural network includes two fully connected layers and two GRU layers; then, the terminal multiplies the power spectrum of each frequency point by the band gain value corresponding to each frequency point to obtain the enhanced power spectrum of each frequency point; finally, the terminal performs IFFT based on the phase information of each frequency point and the enhanced power spectrum of each frequency point to obtain the target speech signal.
[0155] Finally, referring to Figure 8 , Figure 8 is a flowchart of another speech signal processing method provided according to an embodiment of the present application. As Figure 8 shown, the speech signal processing method includes the following steps 801 to step 804.
[0156] 801. Perform concatenated coding on the initial speech signal to obtain a damaged speech signal, where the initial speech signal refers to a speech signal that has not undergone concatenated coding, and the damaged speech signal refers to a speech signal obtained after concatenated coding.
[0157] 802. Transform the initial speech signal from the time domain to the frequency domain to obtain the first power spectrum and phase information of each frequency point in the frequency domain.
[0158] 803. Obtain the band gain value of each frequency point, and determine the second power spectrum of each frequency point based on the first power spectrum and the band gain value of each frequency point.
[0159] 804. Generate a target speech signal that meets the speech playback condition based on the phase information of each frequency point and the second power spectrum.
[0160] It should be noted that after deep learning repair of the damaged voice signal, the obtained voice signal is a voice signal closer to the initial voice signal, that is, it can be restored to a better sound quality, effectively improving the voice quality in the cascade coding application scenario, and further enhancing the voice intelligibility.
[0161] In summary, for the voice signal to be processed, the embodiment of the present application first obtains the first power spectrum and phase information of each frequency point of such voice signal in the frequency domain, and then realizes the enhancement of the first power spectrum by obtaining the frequency band gain value corresponding to each frequency point, obtains the second power spectrum of each frequency point, and further realizes the generation of the target voice signal meeting the voice playback condition according to the second power spectrum and phase information of each frequency point. Since this processing method enhances the power spectrum of each frequency point in a targeted manner, the enhancement effect of the voice signal is more stable, effectively improving the voice quality and further enhancing the voice intelligibility; moreover, regardless of whether the voice signal to be processed has been previously cascade-coded, this processing method can be used to enhance such voice signal, and the applicable range is wide.
[0162] Figure 9 It is a schematic structural diagram of a voice signal processing device provided by an embodiment of the present application. This device is used to execute the steps when the above voice signal processing method is executed. Refer to Figure 9 This voice signal processing device includes: an acquisition module 901, a determination module 902, and a generation module 903.
[0163] The acquisition module 901 is configured to transform the voice signal to be processed from the time domain to the frequency domain, and obtain the first power spectrum and phase information of each frequency point in the frequency domain; wherein, the voice signal to be processed is an initial voice signal or a damaged voice signal, the initial voice signal refers to a voice signal that has not been cascade-coded, and the damaged voice signal refers to a voice signal obtained after the cascade-coding process;
[0164] The determination module 902 is configured to obtain the frequency band gain value of each frequency point, and determine the second power spectrum of each frequency point based on the first power spectrum and the frequency band gain value of each frequency point;
[0165] The generation module 903 is configured to generate a target voice signal meeting the voice playback condition based on the phase information and the second power spectrum of each frequency point.
[0166] In an optional implementation manner, in response to the voice signal to be processed being an initial voice signal, the generation module 903 is further configured to:
[0167] Generate an intermediate voice signal based on the phase information and the second power spectrum of each frequency point;
[0168] Perform the cascaded encoding process on the intermediate speech signal to obtain the target speech signal.
[0169] In an optional implementation manner, in response to the speech signal to be processed being the damaged speech signal, the device further includes:
[0170] A processing module, configured to perform the cascaded encoding process on the initial speech signal before transforming the speech signal to be processed from the time domain to the frequency domain, to obtain the damaged speech signal.
[0171] In an optional implementation manner, the determining module 902 is configured to:
[0172] Input the first power spectrum of each frequency point into a target neural network to obtain the frequency band gain value of each frequency point; wherein, the target neural network includes a first fully connected layer, a gated recurrent unit layer, and a second fully connected layer connected in sequence;
[0173] Take the product of the first power spectrum of each frequency point and the frequency band gain value as the second power spectrum of each frequency point.
[0174] In an optional implementation manner, the device further includes a training module, and the training module is configured to:
[0175] Obtain the third power spectrum of each frequency point of a speech sample in the frequency domain, where the third power spectrum is obtained by transforming the speech sample from the time domain to the frequency domain;
[0176] Input the third power spectrum corresponding to the speech sample into an initial neural network to obtain the predicted frequency band gain value corresponding to the third power spectrum;
[0177] Construct a loss function based on the predicted frequency band gain value and the target frequency band gain value of the speech sample;
[0178] Based on the loss function, continuously adjust the network parameters of the initial neural network until a preset condition is met, to obtain the target neural network;
[0179] Wherein, the target frequency band gain value is obtained based on the third power spectrum and the fourth power spectrum corresponding to the speech sample, and the fourth power spectrum is obtained by performing the cascaded encoding process on the speech sample and then transforming the speech sample from the time domain to the frequency domain.
[0180] In an optional implementation manner, the target frequency band gain value is the square root value of the ratio of the third power spectrum to the fourth power spectrum.
[0181] In an optional implementation manner, the determining module 902 is further configured to:
[0182] Input the first power spectrum of each frequency point into the first fully connected layer. After the first fully connected layer extracts features from the first power spectrum of each frequency point, a feature vector is obtained;
[0183] Input the feature vector into the gated recurrent unit layer. Through the update gate and reset gate in the gated recurrent unit layer, extract the correlation and effective information between the feature vectors to obtain an output vector;
[0184] Input the output vector into the second fully connected layer. Through the second fully connected layer, integrate the output vector into the frequency band gain values of each frequency point.
[0185] In an optional implementation manner, the obtaining module 901 is further configured to:
[0186] Perform frame splitting processing and windowing processing on the to-be-processed speech signal in sequence;
[0187] Perform fast Fourier transform on the to-be-processed speech signal after frame splitting processing and windowing processing; Based on the obtained transformation result, determine the first power spectrum and phase information of each frequency point in the frequency domain.
[0188] In an optional implementation manner, the cascade encoding process includes M times of encoding and decoding processes, where M is a positive integer greater than 1. The processing module is further configured to:
[0189] Perform M times of encoding and decoding processes on the initial speech signal to obtain the damaged speech signal;
[0190] Wherein, the output of the previous encoding and decoding process is used as the input of the next encoding and decoding process; For any encoding and decoding process, the encoding and decoding process includes one encoding process and one decoding process, and the output of the encoding process is used as the input of the decoding process.
[0191] For the to-be-processed speech signal, in the embodiments of the present application, first obtain the first power spectrum and phase information of each frequency point of such speech signal in the frequency domain, and then by obtaining the frequency band gain values corresponding to each frequency point, enhance the first power spectrum to obtain the second power spectrum of each frequency point, and further generate a target speech signal that meets the speech playback condition according to the second power spectrum and phase information of each frequency point. Since this processing method enhances the power spectrum of each frequency point in a targeted manner, the enhancement effect of the speech signal is more stable, effectively improving the speech quality, and further enhancing the speech intelligibility; Moreover, regardless of whether the to-be-processed speech signal has been previously subjected to cascade encoding processing, this processing method can be used to enhance such speech signals, and the applicable range is wide.
[0192] It should be noted that: when the voice signal processing device provided in the above embodiment processes a voice signal, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the voice signal processing device provided in the above embodiment and the embodiment of the voice signal processing method belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.
[0193] The voice signal processing method provided in the embodiment of the present application, the computer device can be configured as a terminal or a server, that is, the method can be executed by the terminal as the execution subject, and can also be executed by the server as the execution subject. Of course, it can also be executed by the interaction between the terminal and the server. For example, the terminal sends the voice signal to be processed to the server and requests to obtain the target voice signal. The server processes the voice signal to be processed based on the received request, and feeds back the target voice signal to the terminal after obtaining it. It should be noted that the embodiment of the present application does not limit the interaction method between the terminal and the server.
[0194] In an exemplary embodiment, a computer device is also provided. Taking the computer device as a terminal as an example, Figure 10 Fig. shows a schematic structural diagram of a terminal 1000 provided in an exemplary embodiment of the present application. The terminal 1000 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The terminal 1000 may also be referred to by other names such as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0195] Generally, the terminal 1000 includes: a processor 1001 and a memory 1002.
[0196] The processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1001 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1001 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 may further include an AI (Artificial Intelligence) processor, which is used to process computational operations related to machine learning.
[0197] The memory 1002 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1002 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 is used to store at least one program code, and the at least one program code is used to be executed by the processor 1001 to implement the voice signal processing method provided in the method embodiments of the present application.
[0198] In some embodiments, the terminal 1000 may further optionally include: a peripheral device interface 1003 and at least one peripheral device. The processor 1001, the memory 1002, and the peripheral device interface 1003 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1003 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of the following: a radio frequency circuit 1004, a display screen 1005, a camera module 1006, an audio circuit 1007, and a power supply 1009.
[0199] The peripheral device interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1001 and the memory 1002. In some embodiments, the processor 1001, the memory 1002, and the peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1001, the memory 1002, and the peripheral device interface 1003 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0200] The radio frequency circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1004 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1004 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1004 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1004 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0201] The display screen 1005 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1005 is a touch display screen, the display screen 1005 also has the ability to collect touch signals on or above the surface of the display screen 1005. The touch signals can be input to the processor 1001 as control signals for processing. At this time, the display screen 1005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1005, which is disposed on the front panel of the terminal 1000; in other embodiments, there may be at least two display screens 1005, which are respectively disposed on different surfaces of the terminal 1000 or are in a foldable design; in other embodiments, the display screen 1005 may be a flexible display screen, which is disposed on a curved surface or a folding surface of the terminal 1000. Even further, the display screen 1005 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1005 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0202] The camera module 1006 is used to collect images or videos. Optionally, the camera module 1006 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to implement functions such as background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting functions or other fused shooting functions. In some embodiments, the camera module 1006 may further include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. A two-color-temperature flash refers to a combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0203] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1001 for processing, or input to the radio frequency circuit 1004 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1000. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1007 may further include a headphone jack.
[0204] The power supply 1009 is used to supply power to each component in the terminal 1000. The power supply 1009 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1009 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0205] In some embodiments, the terminal 1000 further includes one or more sensors 1010. The one or more sensors 1010 include but are not limited to: an acceleration sensor 1011, a gyroscope sensor 1012, a pressure sensor 1013, an optical sensor 1015, and a proximity sensor 1016.
[0206] The acceleration sensor 1011 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal 1000. For example, the acceleration sensor 1011 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1001 can control the display screen 1005 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1011. The acceleration sensor 1011 can also be used for collecting game or user's motion data.
[0207] The gyroscope sensor 1012 can detect the body direction and rotation angle of the terminal 1000. The gyroscope sensor 1012 can cooperate with the acceleration sensor 1011 to collect the 3D actions of the user on the terminal 1000. According to the data collected by the gyroscope sensor 1012, the processor 1001 can achieve the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0208] The pressure sensor 1013 can be disposed on the side frame of the terminal 1000 and / or the lower layer of the display screen 1005. When the pressure sensor 1013 is disposed on the side frame of the terminal 1000, it can detect the holding signal of the user for the terminal 1000, and the processor 1001 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1013. When the pressure sensor 1013 is disposed on the lower layer of the display screen 1005, the processor 1001 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1005. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0209] The optical sensor 1015 is used to collect the ambient light intensity. In one embodiment, the processor 1001 can control the display brightness of the display screen 1005 according to the ambient light intensity collected by the optical sensor 1015. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1005 is increased; when the ambient light intensity is low, the display brightness of the display screen 1005 is decreased. In another embodiment, the processor 1001 can also dynamically adjust the shooting parameters of the camera module 1006 according to the ambient light intensity collected by the optical sensor 1015.
[0210] The proximity sensor 1016, also known as the distance sensor, is usually disposed on the front panel of the terminal 1000. The proximity sensor 1016 is used to collect the distance between the user and the front of the terminal 1000. In one embodiment, when the proximity sensor 1016 detects that the distance between the user and the front of the terminal 1000 is gradually decreasing, the processor 1001 controls the display screen 1005 to switch from the lit state to the off state; when the proximity sensor 1016 detects that the distance between the user and the front of the terminal 1000 is gradually increasing, the processor 1001 controls the display screen 1005 to switch from the off state to the lit state.
[0211] Those skilled in the art can understand that Figure 10 the structure shown in does not limit the terminal 1000, and it may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.
[0212] Taking the computer device as a server as an example, Figure 11It is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1100 may vary greatly due to different configurations or performances, and can include one or more processors (Central Processing Units, CPUs) 1101 and one or more memories 1102. Among them, at least one computer program is stored in the memory 1102, and the at least one computer program is loaded and executed by the processor 1101 to implement the voice signal processing method provided by each of the above method embodiments. Of course, the server can also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input and output. The server can also include other components for implementing device functions, which will not be elaborated here.
[0213] An embodiment of the present application also provides a computer-readable storage medium, which is applied to a computer device. At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the computer device in the voice signal processing method of the above embodiment.
[0214] An embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer program code, and the computer program code is stored in a computer-readable storage medium. The processor of the computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device executes the voice signal processing method provided in the above various optional implementation manners.
[0215] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The described program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk, an optical disc, etc.
[0216] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for processing voice signals, characterized in that, The method includes: Transform the speech signal to be processed from the time domain to the frequency domain to obtain the first power spectrum and phase information of each frequency point in the frequency domain; wherein, the speech signal to be processed is an initial speech signal or a damaged speech signal, the initial speech signal refers to a speech signal that has not undergone concatenated coding processing, and the damaged speech signal refers to a speech signal obtained after the concatenated coding processing; Input the first power spectrum of each frequency point into the target neural network to obtain the frequency band gain value of each frequency point; wherein, the target neural network includes a first fully connected layer, a gated recurrent unit layer, and a second fully connected layer connected in sequence; Take the product of the first power spectrum and the frequency band gain value of each frequency point as the second power spectrum of each frequency point; based on the phase information and the second power spectrum of each frequency point, generate a target speech signal that meets the speech playback condition; The training process of the target neural network includes: obtaining the third power spectrum of each frequency point of the speech sample in the frequency domain, and the third power spectrum is obtained by transforming the speech sample from the time domain to the frequency domain; Input the third power spectrum corresponding to the speech sample into the initial neural network to obtain the predicted frequency band gain value corresponding to the third power spectrum; Based on the predicted frequency band gain value and the target frequency band gain value of the speech sample, construct a loss function, the target frequency band gain value is obtained based on the third power spectrum and the fourth power spectrum corresponding to the speech sample, and the fourth power spectrum is obtained by performing the concatenated coding processing on the speech sample and then transforming the speech sample from the time domain to the frequency domain; Based on the loss function, continuously adjust the network parameters of the initial neural network until a preset condition is met to obtain the target neural network.
2. The method according to claim 1, wherein In response to the speech signal to be processed being the initial speech signal, the generating a target speech signal that meets the speech playback condition based on the phase information and the second power spectrum of each frequency point includes: Generate an intermediate speech signal based on the phase information and the second power spectrum of each frequency point; Perform the concatenated coding processing on the intermediate speech signal to obtain the target speech signal.
3. The method according to claim 1, wherein In response to the speech signal to be processed being the damaged speech signal, the method further includes: Before transforming the speech signal to be processed from the time domain to the frequency domain, perform the concatenated coding processing on the initial speech signal to obtain the damaged speech signal.
4. The method according to claim 1, characterized in that, The target frequency band gain value is the square root value of the ratio of the third power spectrum to the fourth power spectrum.
5. The method according to any one of claims 1 to 4, characterized in that The inputting the first power spectrum of each frequency point into the target neural network to obtain the frequency band gain value of each frequency point includes: Input the first power spectrum of each frequency point into the first fully connected layer, and after the first fully connected layer extracts features from the first power spectrum of each frequency point, obtain a feature vector; Input the feature vector into the gated recurrent unit layer, and after the update gate and reset gate in the gated recurrent unit layer extract the correlation and effective information between the feature vectors, obtain an output vector; Input the output vector into the second fully-connected layer, and integrate the output vector into the band gain values of each frequency point through the second fully-connected layer.
6. The method according to claim 1, wherein The transformation of the speech signal to be processed from the time domain to the frequency domain to obtain the first power spectrum and phase information of each frequency point in the frequency domain includes: Perform frame segmentation processing and windowing processing on the speech signal to be processed in sequence; Perform a fast Fourier transform on the speech signal to be processed after frame segmentation processing and windowing processing; based on the obtained transformation result, determine the first power spectrum and phase information of each frequency point in the frequency domain.
7. The method according to claim 3, characterized in that The cascade coding process includes M times of encoding and decoding processes, where M is a positive integer greater than 1. The performing the cascade coding process on the initial speech signal to obtain the damaged speech signal includes: Perform M times of encoding and decoding processes on the initial speech signal to obtain the damaged speech signal; Among them, the output of the previous encoding and decoding process is used as the input of the next encoding and decoding process; for any encoding and decoding process, the encoding and decoding process includes one encoding process and one decoding process, and the output of the encoding process is used as the input of the decoding process.
8. A voice signal processing device, characterized in that, The device includes: An acquisition module, configured to transform a speech signal to be processed from the time domain to the frequency domain, and obtain the first power spectrum and phase information of each frequency point in the frequency domain; wherein, the speech signal to be processed is an initial speech signal or a damaged speech signal, the initial speech signal refers to a speech signal that has not undergone cascade coding processing, and the damaged speech signal refers to a speech signal obtained after undergoing the cascade coding process; A determination module, configured to input the first power spectrum of each frequency point into a target neural network to obtain the band gain value of each frequency point; wherein, the target neural network includes a first fully-connected layer, a gated recurrent unit layer, and a second fully-connected layer connected in sequence; use the product of the first power spectrum and the band gain value of each frequency point as the second power spectrum of each frequency point; A generation module, configured to generate a target speech signal meeting the speech playback condition based on the phase information and the second power spectrum of each frequency point; The device further includes a training module, and the training module is configured to: Obtain the third power spectrum of each frequency point of the speech sample in the frequency domain, and the third power spectrum is obtained by transforming the speech sample from the time domain to the frequency domain; Input the third power spectrum corresponding to the speech sample into an initial neural network to obtain the predicted band gain value corresponding to the third power spectrum; Construct a loss function based on the predicted band gain value and the target band gain value of the speech sample; Based on the loss function, continuously adjust the network parameters of the initial neural network until a preset condition is met to obtain the target neural network; Among them, the target band gain value is obtained based on the third power spectrum and the fourth power spectrum corresponding to the speech sample, and the fourth power spectrum is obtained by performing the cascade coding process on the speech sample and then transforming the speech sample from the time domain to the frequency domain.
9. The device according to claim 8, characterized in that In response to the speech signal to be processed being the initial speech signal, the generation module is further configured to: Generate an intermediate speech signal based on the phase information and the second power spectrum of each frequency point; Perform the cascade encoding process on the intermediate speech signal to obtain the target speech signal.
10. The device according to claim 8, characterized in that In response to the speech signal to be processed being the damaged speech signal, the apparatus further includes: A processing module, configured to perform the cascade encoding process on the initial speech signal to obtain the damaged speech signal before transforming the speech signal to be processed from the time domain to the frequency domain.
11. The device according to claim 8, characterized in that, The target frequency band gain value is the square root value of the ratio of the third power spectrum to the fourth power spectrum.
12. The device according to any one of claims 8 to 11, characterized in that, The determining module is configured to: Input the first power spectrum of each frequency point into the first fully-connected layer, and perform feature extraction on the first power spectrum of each frequency point through the first fully-connected layer to obtain a feature vector; Input the feature vector into the gated recurrent unit layer, and extract the correlation and effective information between the feature vectors through the update gate and the reset gate in the gated recurrent unit layer to obtain an output vector; Input the output vector into the second fully-connected layer, and integrate the output vector into the frequency band gain value of each frequency point through the second fully-connected layer.
13. The device according to claim 8, characterized in that, The obtaining module is configured to: Perform frame division processing and windowing processing on the speech signal to be processed in sequence; Perform a fast Fourier transform on the speech signal to be processed after frame division processing and windowing processing; based on the obtained transformation result, determine the first power spectrum and phase information of each frequency point in the frequency domain.
14. The device according to claim 10, characterized in that, The cascade encoding process includes M times of encoding and decoding processes, where M is a positive integer greater than 1, and the processing module is configured to: Perform M times of encoding and decoding processes on the initial speech signal to obtain the damaged speech signal; Wherein, the output of the previous encoding and decoding process is used as the input of the next encoding and decoding process; for any encoding and decoding process, the encoding and decoding process includes one encoding process and one decoding process, and the output of the encoding process is used as the input of the decoding process.
15. A computer device, characterized in that, The computer device includes a processor and a memory, and the memory is configured to store at least one computer program, and the at least one computer program is loaded and executed by the processor to perform the speech signal processing method according to any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the speech signal processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
System and method for encoding voice while suppressing acoustic background noise
CN1285945A