An audio generation method, apparatus, and electronic device

By constructing a neural network model of convolutional network and convolutional length and short-term memory network, combining the amplitude spectrum, DC component amplitude spectrum and phase spectrum of the audio signal, the problem that traditional audio super-resolution technology is difficult to effectively utilize frequency domain features, and achieving higher quality super-resolution audio generation.

CN114758672BActive Publication Date: 2025-06-13SHENZHEN GRANDSTREAM NETWORKS TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210329574.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-06-13
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

Traditional audio super-resolution technology is difficult to effectively utilize the frequency domain characteristics of speech signals, resulting in low quality of super-resolution audio generation.

Method used

A neural network model constructed from a convolutional network and a convolutional length and short-term memory network is used to obtain the amplitude spectrum, DC component amplitude spectrum and phase spectrum of the audio signal, predict and combine it to generate the target audio signal.

Benefits of technology

This method can effectively extract the temporal and spatial characteristics of the audio signal, improve the quality of super-resolution audio generation, and make up for the shortcomings of traditional methods in frequency domain feature processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114758672B_ABST
    Figure CN114758672B_ABST
Patent Text Reader

Abstract

The present application provides an audio generation method, apparatus, and electronic device. After obtaining the first amplitude spectrum, DC component amplitude spectrum, and phase spectrum of the to-be-predicted sampling signals of each audio frame in the to-be-expanded audio signal, the first amplitude spectrum is input into a trained audio prediction model to obtain the second amplitude spectrum corresponding to the to-be-predicted sampling signal. The audio prediction model is a neural network model constructed by a convolutional network and a convolutional long short-term memory network. Then, the target frequency spectrum is obtained by combining the DC component amplitude spectrum, the second amplitude spectrum, and the phase spectrum. Finally, the target audio signal is generated based on the target frequency spectrum. The present application predicts the frequency spectrum characteristics of the target audio signal based on the frequency spectrum characteristics of the to-be-expanded audio signal through the trained audio prediction model, making up for the defect that the current method is not applicable to frequency domain characteristics and improving the quality of super-resolution audio generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice signal processing, and in particular, to an audio generation method, apparatus, and electronic device. Background Art

[0002] With the development and maturity of mobile communication technology, people's requirements for the quality of voice in communication are getting higher and higher. In order to complement the high-frequency components missing in narrowband voice in traditional narrowband communication, audio super-resolution technology has emerged as the times require.

[0003] However, traditional audio super-resolution technology mainly applies the correlation between the high-frequency band and the low-frequency band of voice signals for bandwidth expansion. Due to limited technical methods, the expansion effect is often not very ideal and cannot achieve the effect of real broadband signals; the audio super-resolution technology that only uses a convolutional neural network (CNN) can only extract the spatial features of signals and cannot utilize the temporal characteristics of voice signals, and its algorithm effect is still limited; the method that combines a convolutional neural network (CNN) and a long short-term memory network (LSTM) mainly uses the time-domain sampling points of signals as features. Since the input data of the long short-term memory network is one-dimensional and not applicable to spatial sequence data, this method is not applicable to frequency-domain features. However, since the difference between narrowband voice signals and broadband voice signals is mainly reflected in the frequency band, it is difficult for a network trained using time-domain sampling points as features to learn the relationship between the low-frequency band and the high-frequency band of signals, resulting in low quality of super-resolution audio generation.

[0004] Therefore, there is a need to provide an audio generation method to make up for the defect that the current method is not applicable to frequency-domain features and improve the quality of super-resolution audio generation. Summary of the Invention

[0005] The present application provides an audio generation method, apparatus, and electronic device for making up for the defect that the current method is not applicable to frequency-domain features and improving the quality of super-resolution audio generation.

[0006] To solve the above technical problems, the present application provides the following technical solutions:

[0007] The present application provides an audio generation method, including:

[0008] Obtaining a first amplitude spectrum, a DC component amplitude spectrum, and a phase spectrum of a to-be-predicted sampling signal of each audio frame in the to-be-expanded audio signal;

[0009] Inputting the first amplitude spectrum into a trained audio prediction model to output a second amplitude spectrum corresponding to the to-be-predicted sampling signal, where the audio prediction model is a neural network model constructed by a convolutional network and a convolutional long short-term memory network;

[0010] Combining the DC component amplitude spectrum, the second amplitude spectrum, and the phase spectrum to obtain a target spectrum;

[0011] Generating a target audio signal according to the target spectrum.

[0012] Correspondingly, the present application further provides an audio generation device, including:

[0013] A first acquisition unit, configured to acquire the first amplitude spectrum, the DC component amplitude spectrum, and the phase spectrum of the to-be-predicted sampling signals of each audio frame in the to-be-expanded audio signal;

[0014] A prediction unit, configured to input the first amplitude spectrum into a trained audio prediction model, and output the second amplitude spectrum corresponding to the to-be-predicted sampling signal, where the audio prediction model is a neural network model constructed by a convolutional network and a convolutional long short-term memory network;

[0015] A target spectrum determination unit, configured to combine the DC component amplitude spectrum, the second amplitude spectrum, and the phase spectrum to obtain a target spectrum;

[0016] A generation unit, configured to generate a target audio signal according to the target spectrum.

[0017] Meanwhile, the present application provides an electronic device, which includes a processor and a memory. The memory is used to store a computer program, and the processor is used to run the computer program in the memory to execute the steps in the above audio generation method.

[0018] In addition, the present application provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the above audio generation method.

[0019] In addition, the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above audio generation method.

[0020] Beneficial effects: By combining a convolutional network and a convolutional long short-term memory network to construct a neural network model, compared with the traditional long short-term memory network, the convolutional long short-term memory network adds a convolutional operation, enabling it to not only have the ability to model time series but also capture basic spatial features through convolutional operations in multi-dimensional data. Therefore, the time features and spatial features of the audio signal to be extended can be extracted through the neural network model proposed in this application for training to obtain a trained audio prediction model. Then, based on the spectral features of the audio signal to be extended, the spectral features of the target audio signal can be predicted through this trained audio prediction model, and thus the target signal corresponding to the audio signal to be extended can be obtained, making up for the defect that the current method is not applicable to frequency domain features and improving the quality of super-resolution audio generation. Brief Description of the Drawings

[0021] The following will combine the drawings and describe the specific embodiments of the present application in detail, making the technical solutions and other beneficial effects of the present application obvious.

[0022] Figure 1 It is a schematic flowchart of the audio generation method provided by an embodiment of the present application.

[0023] Figure 2 It is a schematic framework diagram of the application stage of the audio prediction model provided by an embodiment of the present application.

[0024] Figure 3 It is a schematic diagram of the effect of a low-resolution audio signal becoming a super-resolution audio signal provided by an embodiment of the present application.

[0025] Figure 4 It is a schematic diagram of the network structure of the audio prediction model provided by an embodiment of the present application.

[0026] Figure 5 It is a schematic diagram of the structure of the long short-term memory network provided by an embodiment of the present application.

[0027] Figure 6 It is a schematic framework diagram of the training stage of the audio prediction model provided by an embodiment of the present application.

[0028] Figure 7 It is a schematic diagram of the structure of the audio generation device provided by an embodiment of the present application.

[0029] Figure 8 It is a schematic diagram of the structure of the electronic device provided by an embodiment of the present application. Detailed Description of the Embodiments

[0030] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0031] The terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.

[0032] In the present application, the audio signal to be extended may be a low-resolution audio signal, that is, a narrowband signal, with a sampling rate of 8 kHz and a bandwidth generally of 300 Hz - 3.4 kHz.

[0033] In the present application, an audio frame is a signal obtained by overlapping and framing and windowing the audio signal to be extended, and the signal of each frame of the audio signal to be extended is the audio frame.

[0034] In the present application, the signal to be predicted and sampled is a signal obtained by performing time-domain interpolation upsampling processing on each audio frame in the audio signal to be extended.

[0035] In the present application, the magnitude spectrum of the DC component is the modulus value of the first point after time-frequency domain conversion. Specifically, assuming that the sampling frequency is fs and the number of sampling points is N, after converting the time-domain signal to the frequency domain through time-frequency domain conversion, the first point corresponds to the DC component, and the curve composed of the first point and its modulus value is the magnitude spectrum of the DC component.

[0036] In the present application, the phase spectrum refers to the curve of the phase changing with frequency, representing the phase that each frequency component has at the time origin.

[0037] In the present application, the target audio signal may be a signal obtained by performing super-resolution processing on the audio signal to be extended, and the target audio signal complements the missing high-frequency components in the audio signal to be extended.

[0038] The present application provides an audio generation method, device, and electronic device.

[0039] Please refer to Figure 1 , Figure 1It is a schematic flowchart of the audio generation method provided by an embodiment of the present application. This method is applicable to conference scenarios such as voice calls, Internet Protocol Telephony (VoIP), and web conferences. In these scenarios, traditional narrowband communication is mostly used, which can only transmit narrowband signals. The narrowband signals transmitted lack high-frequency components, making the voice unnatural, sounding dull, and affecting the sound quality and listening experience. To improve the sound quality, audio super-resolution technology (i.e., audio bandwidth expansion technology) is used to perform super-resolution processing on low-sampling-rate audio data to accurately complement the missing high-frequency components as much as possible, making the audio sound louder and clearer with a better listening experience.

[0040] It should be noted that the present application places no restrictions on the sampling rates of low-resolution speech signals and high-resolution speech signals. For example, the present application can be applied to tasks where the sampling rate of the audio signal to be expanded is 8 kHz and the target sampling rates are 16 kHz, 24 kHz, and 48 kHz, or can also be applied to tasks where the sampling rate of the audio signal to be expanded is 16 kHz and the target sampling rate is 48 kHz, etc.

[0041] The audio generation method in the present application will be introduced in detail below. This method at least includes the following steps:

[0042] S101: Obtain the first amplitude spectrum, DC component amplitude spectrum, and phase spectrum of the to-be-predicted sampling signals of each audio frame in the audio signal to be expanded.

[0043] In one embodiment, before predicting the spectral characteristics of the audio signal to be expanded, it is necessary to first obtain the audio signal to be expanded and its corresponding information. The specific steps include: obtaining the audio signal to be expanded and the target sampling rate; preprocessing the audio signal to be expanded to obtain each audio frame of the audio signal to be expanded; resampling each audio frame according to the target sampling rate to obtain the to-be-predicted sampling signal; and performing feature extraction processing on the to-be-predicted sampling signal to obtain the first amplitude spectrum, DC component amplitude spectrum, and phase spectrum of the to-be-predicted sampling signal. Among them, the audio signal to be expanded can be a low-resolution audio signal; the target sampling rate is the sampling rate corresponding to the target audio signal (i.e., the high-resolution audio signal corresponding to this low-resolution audio signal) set manually or by the system; the preprocessing includes overlapping framing and windowing processing.

[0044] Specifically, as Figure 2 shown, Figure 2Schematic diagram of the framework in the application stage of the audio prediction model provided by the embodiments of the present application. After obtaining the audio signal to be extended (low-resolution audio signal), overlapping framing and windowing processing are performed on it to obtain each audio frame x(t) of the audio signal to be extended (low-resolution audio signal). Then, time-domain interpolation upsampling processing is performed on each frame of the signal x(t) according to the target sampling rate to make the sampling rate of the audio frame x(t) consistent with the sampling rate of the target audio signal, obtaining the sampling signal x to be predicted i (t). Finally, feature extraction processing is performed on the sampling signal x to be predicted i (t) to obtain the amplitude spectrum and phase spectrum of the sampling signal x to be predicted i (t), where the amplitude spectrum includes the first amplitude spectrum and the DC-separated amplitude spectrum.

[0045] In one embodiment, the specific steps for obtaining the phase spectrum and amplitude spectrum include: converting the sampling signal to be predicted to the frequency domain according to the preset time-frequency domain conversion condition to obtain the frequency spectrum of the sampling signal to be predicted; determining the phase spectrum of the sampling signal to be predicted according to the frequency spectrum of the sampling signal to be predicted; performing amplitude spectrum extraction processing on the sampling signal to be predicted to obtain the first amplitude spectrum and the DC component amplitude spectrum of the sampling signal to be predicted. Among them, the preset time-frequency domain conversion condition may be the short-time Fourier transform. Specifically, the sampling signal x to be predicted in the time domain i (t) is converted to the frequency domain through the short-time Fourier transform to obtain the frequency spectrum X i (ω) of the sampling signal to be predicted. Then, the phase spectrum is obtained according to the frequency spectrum, and finally, amplitude spectrum extraction processing is performed on the sampling signal x to be predicted i (t) to obtain the first amplitude spectrum and the DC component amplitude spectrum of the sampling signal to be predicted. The detailed steps will be described below.

[0046] In one embodiment, the specific steps for performing amplitude spectrum extraction processing on the sampling signal to be predicted to obtain the first amplitude spectrum and the DC component amplitude spectrum of the sampling signal to be predicted include: performing filtering processing on the sampling signal to be predicted to obtain the first filtered signal; converting the first filtered signal to the frequency domain according to the preset time-frequency domain conversion condition to obtain the frequency spectrum of the first filtered signal; determining the first amplitude spectrum and the DC component amplitude spectrum of the sampling signal to be predicted according to the frequency spectrum of the first filtered signal. Among them, the filtering processing may be low-pass filtering. The mirror frequency spectrum generated during the aforementioned interpolation resampling processing is filtered out through low-pass filtering to obtain the first filtered signal x 1 (t). Then, the first filtered signal x 1 (t) is converted to the frequency domain through the short-time Fourier transform to obtain the frequency spectrum X 1 (ω) of the first filtered signal. Then, the frequency spectrum X 1(w)Take the absolute value to obtain the amplitude spectrum of the first filtered signal. Since the amplitude spectrum is symmetric about the vertical axis, taking the symmetric half of this amplitude spectrum can obtain the first amplitude spectrum of the sampling signal x to be predicted (the curve formed by the amplitudes of each point except the first point in the spectrum X i (t) of the first filtered signal) and the DC component amplitude spectrum (the modulus value of the first point in the spectrum X 1 (w) of the first filtered signal). 1 (w)).

[0047] It should be noted that due to the short-time stationary characteristics of the speech signal, the short-time Fourier transform is used to convert the time-domain signal to the frequency domain.

[0048] S102: Input the first amplitude spectrum into the trained audio prediction model, and output the second amplitude spectrum corresponding to the sampling signal to be predicted. The audio prediction model is a neural network model constructed by a convolutional network and a convolutional long short-term memory network.

[0049] Traditional audio super-resolution techniques mainly use the correlation between the high-frequency band and the low-frequency band of the speech signal for bandwidth expansion. The expansion methods mainly include codebook mapping method, bandwidth replication method, linear prediction analysis method, method based on Gaussian mixture model (GMM), method based on hidden Markov model (HMM), etc. Because of the limited technical methods, the expansion effect is often not very ideal and cannot reach the effect of real wideband signals.

[0050] In recent years, with the continuous in-depth research, many audio super-resolution algorithms based on neural networks have emerged. Their main idea is to draw on the super-resolution algorithm of images and use a deep neural network to predict the high-frequency band signal corresponding to the low-frequency band signal. The effect is as Figure 3As shown. However, different from images, voice signals are sequential signals, and there is a great correlation between the front and back of the voice within a period of time, that is, the voice signal at the current time is related to the signals within a previous period of time. The convolutional neural network (CNN) used in image algorithms can only extract the spatial features of signals and cannot utilize the temporal characteristics of voice signals. Therefore, the audio super-resolution algorithm that only uses the convolutional neural network (CNN) still has limited effects. In order to utilize the temporal characteristics of voice signals, a method combining the convolutional neural network (CNN) and the long short-term memory network (LSTM) has emerged. The LSTM is a variant of the recurrent convolutional neural network (RNN). The RNN is a neural network structure that can process sequential data. Using the RNN network can consider the information of multiple previous frames when processing the current frame of voice, making the current output result not only related to the current input but also related to the inputs within a previous period of time. Therefore, as a variant of the RNN, the LSTM has a long-term "memory" function and avoids the problem of gradient disappearance. However, the current algorithms combining CNN and LSTM mainly use the time-domain sampling points of signals as features. Since the input data of the LSTM is one-dimensional and not applicable to spatial sequential data, this method is not applicable to frequency-domain features, and the algorithm delay and computational complexity of this method are relatively large.

[0051] Based on the above problems, the present application proposes an audio prediction model constructed by a convolutional network (CNN) and a convolutional long short-term memory network (ConvLSTM), and trains the audio prediction model to obtain a trained audio prediction model. The specific steps include: obtaining a training set and an initial audio prediction model. The initial audio prediction model includes a downsampling module, a bottleneck layer module, an upsampling module, and an output module. Among them, the downsampling module is composed of N downsampling layers with the same structure but different parameters, and the upsampling module is composed of N upsampling layers with the same structure but different parameters. Training the initial audio prediction model according to the training set to obtain a trained audio prediction model.

[0052] The network model proposed in the present application is as Figure 4 shown, Figure 4Schematic diagram of the network structure of the audio prediction model provided by the embodiments of the present application. The audio prediction model mainly includes four modules, namely a downsampling module, a bottleneck layer module, an upsampling module, and an output module. Among them, the downsampling module consists of N downsampling layers with the same structure but different parameters. The downsampling layer includes a convolutional layer, a pooling layer, and a convolutional long short-term memory (ConvLSTM) layer; the upsampling module consists of N upsampling layers with the same structure but different parameters. The upsampling layer includes a convolutional layer, a transposed convolutional layer, and a convolutional long short-term memory (ConvLSTM) layer. In addition, there is a 1×1 convolution in each upsampling layer and downsampling layer, whose function is to reduce the dimension of the feature channels before performing ConvLSTM, thereby reducing the number of network parameters. The bottleneck layer includes a convolutional layer, a pooling layer, and a ConvLSTM layer. The output module includes a convolutional layer and a transposed convolutional layer.

[0053] In one embodiment, the outputs of the downsampling layers in the downsampling module are connected to the inputs of the upsampling layers in the upsampling module. Specifically, the audio prediction model also introduces a residual network, that is, the output of each downsampling layer in the downsampling module is connected to the input of the corresponding upsampling layer in the upsampling module. The model does not limit the number of layers of the downsampling module and the upsampling module, that is, it does not limit the size of the number of layers N, but the number of layers in the two modules needs to be the same and the structures are symmetric.

[0054] The convolutional long short-term memory network (ConvLSTM) used in the model is an improvement of the traditional long short-term memory network (LSTM). The structure of the traditional long short-term memory network (LSTM) unit is as Figure 5 shown. LSTM is a variant of the recurrent neural network (RNN). Its essence is a fully connected network, which can also be called FC-LSTM. Its function is to have a "memory" function. LSTM adds the preservation of the state to the computing unit to improve the memory of the network, that is, the network will consider past information. Its output depends not only on the current input but also on previous information. The current output is jointly determined by previous information and the current input.

[0055] The formulas of LSTM are shown in Formulas 1 to 5, where σ is the gate, i t is the input gate, f t is the forget gate, c t is the memory cell, o t is the output gate, h t is the hidden state therein, W and b are the weights to be learned by the network, x t is the input sequence data, and ⊙ is the Hadamard product. The weights within the LSTM computing unit are shared, and each layer of LSTM shares a set of weights.

[0056]

[0057]

[0058]

[0059]

[0060]

[0061] ConvLSTM is a convolutional long short-term memory network. Compared with the traditional LSTM network, ConvLSTM adds a convolutional operation. LSTM can only extract temporal features, while after replacing the fully connected layer with a convolutional operation, ConvLSTM can extract both spatio-temporal characteristics. It not only has the ability of temporal modeling but also can extract spatial features. It captures the underlying spatial features by performing convolutional operations on multi-dimensional data. Its formulas are shown in Formulas 6 to 10, where * represents the convolutional operation.

[0062]

[0063]

[0064]

[0065]

[0066]

[0067] It can be seen that compared with LSTM, ConvLSTM uses a convolutional operation instead of matrix multiplication. After adding the convolutional operation, it can not only obtain the temporal relationship of the signal but also extract spatial features like a convolutional layer.

[0068] In one embodiment, based on the above analysis, the audio prediction model provided by the present application can simultaneously extract the temporal features and spatial features of the audio signal for training. Therefore, before training the audio prediction model, it is necessary to obtain a training set containing temporal features and spatial features, and its specific steps include: obtaining a first speech signal, a first sampling rate, and a second sampling rate; preprocessing the first speech signal to obtain each first audio frame of the first speech signal; resampling each first audio frame according to the first sampling rate to obtain a first sampled signal; resampling the first sampled signal according to the second sampling rate to obtain a second sampled signal; filtering the second sampled signal to obtain a second filtered signal; respectively converting the first speech signal and the second filtered signal to the frequency domain according to preset time-frequency domain conversion conditions to obtain a first training spectrum corresponding to the first speech signal and a second training spectrum corresponding to the second filtered signal; determining a training set according to the first training spectrum and the second training spectrum. Among them, the first speech signal can be an original high-resolution speech signal for training; the preprocessing includes overlapping framing and windowing; the preset time-frequency domain conversion condition can be a short-time Fourier transform.

[0069] As Figure 6 shown, Figure 6Schematic diagram of the framework in the training phase of the audio prediction model provided by the embodiments of the present application. Since the speech signal has short-term stationarity and the subsequent calculations are based on the frequency-domain characteristics of the signal, it is necessary to perform overlapping framing and windowing on the collected first speech signal (i.e., the original high-resolution speech signal) to obtain the first audio frame. Among them, the purpose of windowing is to prevent spectral leakage; then, according to the first sampling rate M (M is an integer greater than 1), downsampling is performed on each frame of high-resolution speech signal (i.e., the first audio frame) to obtain the first sampled signal; then, time-domain interpolation upsampling is performed according to the second sampling rate P (P is an integer greater than 1) to obtain the second sampled signal. The purpose of time-domain interpolation upsampling is to use the mirror high-frequency components obtained by upsampling to obtain the phase spectrum of the high-frequency components in the application phase. In order to keep the form of the data the same during training and actual application, although the phase spectrum is not required during the training phase, time-domain interpolation upsampling also needs to be performed on the training data; then, low-pass filtering is performed on the second sampled signal to filter out the mirror high-frequency components obtained by time-domain interpolation upsampling, and a second filtered signal (i.e., a low-resolution speech signal) with only low-frequency components and missing high-frequency components is obtained; then, short-time Fourier transform is performed on the first speech signal (i.e., the original high-resolution speech signal) and the second filtered signal (i.e., the low-resolution speech signal) respectively to obtain the spectral characteristics (i.e., the first training spectrum) of the corresponding first speech signal (i.e., the original high-resolution speech signal) and the spectral characteristics (i.e., the second training spectrum) of the second filtered signal (i.e., the low-resolution speech signal); due to the symmetric characteristics of the spectrum, only half of the spectrum is retained, the absolute value of half of the first training spectrum is obtained to get the amplitude spectrum of the first speech signal (i.e., the original high-resolution speech signal), denoted as label; the absolute value of half of the second training spectrum is obtained to get the amplitude spectrum of the second filtered signal (i.e., the low-resolution speech signal), denoted as data; the above operations of sampling, filtering, time-frequency domain transformation, and extracting the amplitude spectrum are repeated for each first audio frame to obtain the training set used for model training; finally, the initial audio prediction model is trained using the made training set, and the weight parameters obtained by training are saved to obtain the trained audio prediction model.

[0070] The first amplitude spectrum is input into the trained audio prediction model, and the second amplitude spectrum corresponding to the currently to-be-predicted sampled signal (i.e., the amplitude spectrum of the high-resolution signal corresponding to the currently to-be-predicted sampled signal) is predicted using the weight parameters of the trained audio prediction model.

[0071] S103: Combine the DC component amplitude spectrum, the second amplitude spectrum, and the phase spectrum to obtain the target spectrum.

[0072] Since the spectrum reflects the distribution of the amplitude and phase of a signal with respect to frequency, after obtaining the DC component amplitude spectrum, the second amplitude spectrum, and the phase spectrum of the sampling signal to be predicted according to the foregoing steps, combining them can yield the target spectrum Y(w).

[0073] S104: Generate a target audio signal based on the target spectrum.

[0074] In one embodiment, the specific steps for generating the target audio signal include: converting the target spectrum to the time domain according to preset inverse time-frequency domain conversion conditions to obtain a target audio frame corresponding to the target spectrum; performing time-domain synthesis processing on the target audio frame to obtain the target audio signal. Among them, the preset inverse time-frequency domain conversion conditions may be the inverse short-time Fourier transform.

[0075] Specifically, after obtaining the target spectrum Y(w), performing an inverse short-time Fourier transform on it can obtain the target audio frame signal y(t) in the time domain; then performing time-domain synthesis processing on the target audio frame signal, including windowing each target audio frame signal y(t) and overlapping and adding it to the previous frame audio frame signal y(t - 1) to obtain the target audio signal (i.e., the target high-resolution audio signal after expanding the audio signal to be expanded). In the process of synthesizing the target audio signal, the information of multiple previous frames can be used, improving the quality of super-resolution audio generation.

[0076] This application provides a super-resolution audio generation method and system based on a neural network. This method combines convolution and convolutional long short-term memory network (ConvLSTM), can simultaneously extract the temporal and spatial features of the audio signal for training, and then predict the missing high-frequency components based on the spectral features in the low-resolution speech signal, improving the quality of speech under limited bandwidth conditions, solving the problems of low resolution and poor sound quality of narrowband signals, and the algorithm delay of this method is small, which can support streaming processing and meet the real-time requirements.

[0077] Based on the content of the above embodiments, an audio generation device is provided in an embodiment of this application. This device can be set in an audio acquisition terminal. This audio generation device is used to execute the audio generation method provided in the above method embodiments. Specifically, please refer to Figure 7 , and this device includes:

[0078] A first acquisition unit 701, configured to acquire the first amplitude spectrum, the DC component amplitude spectrum, and the phase spectrum of the sampling signal to be predicted for each audio frame in the audio signal to be expanded;

[0079] A prediction unit 702, configured to input the first amplitude spectrum into a trained audio prediction model, and output a second amplitude spectrum corresponding to the to-be-predicted sampling signal, where the audio prediction model is a neural network model constructed by a convolutional network and a convolutional long short-term memory network;

[0080] A target spectrum determination unit 703, configured to obtain a target spectrum by combining the DC component amplitude spectrum, the second amplitude spectrum, and the phase spectrum;

[0081] A generation unit 704, configured to generate a target audio signal according to the target spectrum.

[0082] In one embodiment, the first acquisition unit 701 includes:

[0083] A second acquisition unit, configured to acquire a to-be-expanded audio signal and a target sampling rate;

[0084] A first preprocessing unit, configured to preprocess the to-be-expanded audio signal to obtain respective audio frames of the to-be-expanded audio signal;

[0085] A first sampling unit, configured to perform resampling processing on each audio frame according to the target sampling rate to obtain a to-be-predicted sampling signal;

[0086] A feature extraction unit, configured to perform feature extraction processing on the to-be-predicted sampling signal to obtain a first amplitude spectrum, a DC component amplitude spectrum, and a phase spectrum of the to-be-predicted sampling signal.

[0087] In one embodiment, the feature extraction unit includes:

[0088] A first time-frequency domain conversion unit, configured to convert the to-be-predicted sampling signal to the frequency domain according to preset time-frequency domain conversion conditions to obtain a spectrum of the to-be-predicted sampling signal;

[0089] A phase spectrum determination unit, configured to determine a phase spectrum of the to-be-predicted sampling signal according to the spectrum of the to-be-predicted sampling signal;

[0090] An amplitude spectrum extraction unit, configured to perform amplitude spectrum extraction processing on the to-be-predicted sampling signal to obtain a first amplitude spectrum and a DC component amplitude spectrum of the to-be-predicted sampling signal.

[0091] In one embodiment, the amplitude spectrum extraction unit includes:

[0092] A first filtering unit, configured to perform filtering processing on the to-be-predicted sampling signal to obtain a first filtered signal;

[0093] A second time-frequency domain conversion unit, configured to convert the first filtered signal to the frequency domain according to the preset time-frequency domain conversion conditions to obtain a spectrum of the first filtered signal;

[0094] An amplitude spectrum determination subunit, configured to determine a first amplitude spectrum and a DC component amplitude spectrum of the to-be-predicted sampling signal according to the spectrum of the first filtered signal.

[0095] In one embodiment, the audio generation device further includes:

[0096] A third acquisition unit, configured to acquire a training set and an initial audio prediction model, where the initial audio prediction model includes a downsampling module, a bottleneck layer module, an upsampling module, and an output module; wherein, the downsampling module is composed of N downsampling layers with the same structure but different parameters, and the upsampling module is composed of N upsampling layers with the same structure but different parameters;

[0097] A model training unit, configured to train the initial audio prediction model according to the training set to obtain a trained audio prediction model.

[0098] In one embodiment, the third acquisition unit includes:

[0099] A fourth acquisition unit, configured to acquire a first speech signal, a first sampling rate, and a second sampling rate;

[0100] A second preprocessing unit, configured to preprocess the first speech signal to obtain respective first audio frames of the first speech signal;

[0101] A second sampling unit, configured to resample each of the first audio frames according to the first sampling rate to obtain a first sampling signal;

[0102] A third sampling unit, configured to resample the first sampling signal according to the second sampling rate to obtain a second sampling signal;

[0103] A second filtering unit, configured to filter the second sampling signal to obtain a second filtered signal;

[0104] A third time-frequency domain conversion unit, configured to convert the first speech signal and the second filtered signal into the frequency domain respectively according to preset time-frequency domain conversion conditions to obtain a first training spectrum corresponding to the first speech signal and a second training spectrum corresponding to the second filtered signal;

[0105] A training set determination unit, configured to determine a training set according to the first training spectrum and the second training spectrum.

[0106] In one embodiment, the output of each downsampling layer in the downsampling module of the initial audio prediction model is connected to the input of each upsampling layer in the upsampling module.

[0107] In one embodiment, the audio generation unit 704 includes:

[0108] A frequency-time domain conversion unit, configured to convert the target frequency spectrum into the time domain according to a preset inverse time-frequency domain conversion condition, so as to obtain a target audio frame corresponding to the target frequency spectrum;

[0109] An audio frame synthesis unit, configured to perform time-domain synthesis processing on the target audio frame to obtain a target audio signal.

[0110] The audio generation device according to the embodiment of the present application can be used to execute the technical solutions of the foregoing method embodiments. The implementation principles and technical effects are similar, and will not be elaborated here.

[0111] Different from the current technology, the audio generation device provided by the present application is provided with a prediction unit. The prediction unit predicts the frequency spectrum feature of the target audio signal based on the frequency spectrum feature of the audio signal to be extended, and then obtains the target signal corresponding to the audio signal to be extended, making up for the defect that the current method is not applicable to frequency domain features and improving the quality of super-resolution audio generation.

[0112] Correspondingly, the embodiment of the present application further provides an electronic device, which includes a server or a terminal, etc.

[0113] As Figure 8 shown, the electronic device may include a processor 801 with one or more processing cores, a wireless (WiFi, Wireless Fidelity) module 802, a memory 803 with one or more computer-readable storage media, an audio circuit 804, a display unit 805, an input unit 806, a sensor 807, a power supply 808, and a radio frequency (RF, RadioFrequency) circuit 809 and other components. Those skilled in the art can understand that Figure 8 the structure of the electronic device shown in

[0114] The processor 801 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 803, and calling data stored in the memory 803, it executes various functions of the electronic device and processes data, thereby monitoring the entire electronic device. In one embodiment, the processor 801 may include one or more processing cores; preferably, the processor 801 may integrate an application processor and a modulation and demodulation processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modulation and demodulation processor mainly processes wireless communication. It can be understood that the above modulation and demodulation processor may not be integrated into the processor 801.

[0115] WiFi belongs to short-range wireless transmission technology. The wireless module 802 of the electronic device can help users send and receive emails, browse the web, and access streaming media, etc., and it provides users with wireless broadband Internet access. Although Figure 8 the wireless module 802 is shown, it can be understood that it does not belong to the essential components of the terminal and can be completely omitted within the scope of not changing the essence of the invention according to needs.

[0116] The memory 803 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the computer programs and modules stored in the memory 803. The memory 803 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the terminal (such as audio data, phone book, etc.). In addition, the memory 803 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 803 can also include a memory controller to provide access to the memory 803 by the processor 801 and the input unit 806.

[0117] The audio circuit 804 includes a speaker, and the speaker can provide an audio interface between the user and the electronic device. The audio circuit 804 can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output; on the other hand, the speaker converts the collected sound signal into an electrical signal, which is received by the audio circuit 804 and then converted into audio data. After the audio data is output and processed by the processor 801, it is sent to another electronic device, for example, through the radio frequency circuit 809, or the audio data is output to the memory 803 for further processing. The audio circuit 804 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device.

[0118] The display unit 805 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the terminal. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 805 can include a display panel. In one embodiment, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, a touch-sensitive surface can cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it is transmitted to the processor 801 to determine the type of touch event. Subsequently, the processor 801 provides a corresponding visual output on the display panel according to the type of touch event. Although in Figure 8 the touch-sensitive surface and the display panel are implemented as two independent components to achieve input and output functions, in some embodiments, the touch-sensitive surface and the display panel can be integrated to achieve input and output functions.

[0119] The input unit 806 can be used to receive input numerical or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls. Specifically, in a specific embodiment, the input unit 806 can include a touch-sensitive surface and other input devices. The touch-sensitive surface, also known as a touch display screen or a touchpad, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch-sensitive surface), and drive the corresponding connection device according to a pre-set program. In one embodiment, the touch-sensitive surface can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 801, and can receive and execute commands sent by the processor 801. In addition, the touch-sensitive surface can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface, the input unit 806 can also include other input devices. Specifically, the other input devices can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.

[0120] The electronic device may further include at least one sensor 807, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel according to the brightness of the ambient light. The proximity sensor can turn off the display panel and / or the backlight when the terminal is moved to the ear. As a kind of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in each direction (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used in applications for identifying the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. As for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the electronic device may also be configured with, they will not be elaborated here.

[0121] The electronic device further includes a power source 808 (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the processor 801 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power source 808 may also include any components such as one or more DC or AC power sources, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0122] The radio frequency circuit 809 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink information from the base station, it is handed over to one or more processors 801 for processing. Additionally, data related to the uplink is sent to the base station. Generally, the radio frequency circuit 809 includes, but is not limited to, antennas, at least one amplifier, a tuner, one or more oscillators, a subscriber identity module (SIM) card, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. Moreover, the radio frequency circuit 809 can also communicate with the network and other devices via wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0123] Although not shown, the electronic device may further include a camera, a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 801 in the electronic device will, according to the following instructions, load the executable files corresponding to the processes of one or more application programs into the memory 803, and the processor 801 will run the application programs stored in the memory 803 to achieve the following functions:

[0124] Obtain the first amplitude spectrum, the DC component amplitude spectrum, and the phase spectrum of the to-be-predicted sampling signals of each audio frame in the to-be-expanded audio signal;

[0125] Input the first amplitude spectrum into the trained audio prediction model to output the second amplitude spectrum corresponding to the to-be-predicted sampling signal. The audio prediction model is a neural network model constructed by a convolutional network and a convolutional long short-term memory network;

[0126] Combine the DC component amplitude spectrum, the second amplitude spectrum, and the phase spectrum to obtain the target spectrum;

[0127] Generate a target audio signal according to the target spectrum.

[0128] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not described in detail in a certain embodiment, reference may be made to the detailed description above, and details will not be repeated here.

[0129] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0130] Therefore, an embodiment of the present application provides a computer-readable storage medium, which stores multiple instructions that can be loaded by a processor to implement the functions of the above audio generation method.

[0131] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, and details will not be repeated here.

[0132] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0133] Since the instructions stored in the computer-readable storage medium can execute the steps in any one of the methods provided by the embodiments of the present application, the beneficial effects that can be achieved by any one of the methods provided by the embodiments of the present application can be realized. For details, reference may be made to the previous embodiments, and details will not be repeated here.

[0134] At the same time, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various optional implementation manners.

[0135] The above has introduced in detail the audio generation method, device, and electronic device provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. An audio generation method, characterized in that, comprising: obtaining a first amplitude spectrum, a DC component amplitude spectrum, and a phase spectrum of a to-be-predicted sampling signal of each audio frame in a to-be-expanded audio signal; the first amplitude spectrum includes a symmetric half part of a curve formed by amplitudes of each point except the first point in the spectrum of a first filtering signal; the DC component amplitude spectrum includes a modulus value of the first point in the spectrum of the first filtering signal; the first filtering signal is obtained based on the to-be-predicted sampling signal; inputting the first amplitude spectrum into a trained audio prediction model to output a second amplitude spectrum corresponding to the to-be-predicted sampling signal, where the audio prediction model is a neural network model constructed by a convolutional network and a convolutional long short-term memory network; combining the DC component amplitude spectrum, the second amplitude spectrum, and the phase spectrum to obtain a target spectrum; generating a target audio signal according to the target spectrum.

2. The audio generation method according to claim 1, characterized in that, the step of obtaining a first amplitude spectrum, a DC component amplitude spectrum, and a phase spectrum of a to-be-predicted sampling signal of each audio frame in the to-be-expanded audio signal includes: obtaining the to-be-expanded audio signal and a target sampling rate; preprocessing the to-be-expanded audio signal to obtain each audio frame of the to-be-expanded audio signal; performing resampling processing on each audio frame according to the target sampling rate to obtain a to-be-predicted sampling signal; performing feature extraction processing on the to-be-predicted sampling signal to obtain a first amplitude spectrum, a DC component amplitude spectrum, and a phase spectrum of the to-be-predicted sampling signal.

3. The audio generation method according to claim 2, characterized in that, the step of performing feature extraction processing on the to-be-predicted sampling signal to obtain a first amplitude spectrum, a DC component amplitude spectrum, and a phase spectrum of the to-be-predicted sampling signal includes: converting the to-be-predicted sampling signal to the frequency domain according to preset time-frequency domain conversion conditions to obtain a spectrum of the to-be-predicted sampling signal; determining a phase spectrum of the to-be-predicted sampling signal according to the spectrum of the to-be-predicted sampling signal; performing amplitude spectrum extraction processing on the to-be-predicted sampling signal to obtain a first amplitude spectrum and a DC component amplitude spectrum of the to-be-predicted sampling signal.

4. The audio generation method according to claim 3, characterized in that, the step of performing amplitude spectrum extraction processing on the to-be-predicted sampling signal to obtain a first amplitude spectrum and a DC component amplitude spectrum of the to-be-predicted sampling signal includes: performing filtering processing on the to-be-predicted sampling signal to obtain a first filtering signal; converting the first filtering signal to the frequency domain according to the preset time-frequency domain conversion conditions to obtain a spectrum of the first filtering signal; determining a first amplitude spectrum and a DC component amplitude spectrum of the to-be-predicted sampling signal according to the spectrum of the first filtering signal.

5. The audio generation method according to claim 1, characterized in that, Before the step of inputting the first amplitude spectrum into the trained audio prediction model to output the second amplitude spectrum corresponding to the to-be-predicted sampling signal, where the audio prediction model is a neural network model constructed by a convolutional network and a convolutional long short-term memory network, the method further includes: Obtaining a training set and an initial audio prediction model, where the initial audio prediction model includes a downsampling module, a bottleneck layer module, an upsampling module, and an output module; wherein, the downsampling module consists of N downsampling layers with the same structure but different parameters, and the upsampling module consists of N upsampling layers with the same structure but different parameters; Training the initial audio prediction model according to the training set to obtain the trained audio prediction model.

6. The audio generation method according to claim 5, wherein, the step of obtaining a training set and an initial audio prediction model includes: Obtaining a first speech signal, a first sampling rate, and a second sampling rate; Preprocessing the first speech signal to obtain each first audio frame of the first speech signal; Performing resampling processing on each first audio frame according to the first sampling rate to obtain a first sampling signal; Performing resampling processing on the first sampling signal according to the second sampling rate to obtain a second sampling signal; Performing filtering processing on the second sampling signal to obtain a second filtered signal; Converting the first speech signal and the second filtered signal to the frequency domain respectively according to preset time-frequency domain conversion conditions to obtain a first training spectrum corresponding to the first speech signal and a second training spectrum corresponding to the second filtered signal; Determining the training set according to the first training spectrum and the second training spectrum.

7. The audio generation method according to claim 5, wherein, The output of each downsampling layer in the downsampling module is connected to the input of each upsampling layer in the upsampling module.

8. The audio generation method according to claim 1, wherein, the step of generating a target audio signal according to the target spectrum includes: Converting the target spectrum to the time domain according to preset inverse time-frequency domain conversion conditions to obtain target audio frames corresponding to the target spectrum; Performing time-domain synthesis processing on the target audio frames to obtain a target audio signal.

9. An audio generation device, wherein, it includes: A first acquisition unit, configured to acquire a first amplitude spectrum, a DC component amplitude spectrum, and a phase spectrum of to-be-predicted sampling signals of each audio frame in a to-be-expanded audio signal; the first amplitude spectrum includes a symmetric half part of a curve formed by amplitudes of each point except the first point in the spectrum of the first filtered signal; the DC component amplitude spectrum includes the modulus value of the first point in the spectrum of the first filtered signal; the first filtered signal is obtained based on the to-be-predicted sampling signal; A prediction unit, configured to input the first amplitude spectrum into the trained audio prediction model to output a second amplitude spectrum corresponding to the to-be-predicted sampling signal, where the audio prediction model is a neural network model constructed by a convolutional network and a convolutional long short-term memory network; A target spectrum determination unit, configured to combine the DC component amplitude spectrum, the second amplitude spectrum, and the phase spectrum to obtain a target spectrum; A generation unit, configured to generate a target audio signal according to the target spectrum.

10. An electronic device, characterized in that, it includes a processor and a memory, the memory is used to store a computer program, and the processor is used to run the computer program in the memory to execute the steps in the audio generation method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Apparatus and method for generating a bandwidth extended signal

    CN102105931A

  • Super-resolution audio generation method and device

    CN111508508A