Audio beautification method, device, equipment and storage medium based on self-attention
By extracting audio energy, content and timbre through the self-attention audio beautification method and combining position embedding and attention mechanism, the problem of low audio beautification in the existing technology is solved, and high-quality audio beautification and pitch calibration are achieved.
Patent Information
- Application Number
- CN202310614023.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-26
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-05-26
AI Technical Summary
The existing technology has a low level of audio beautification. The existing model can only extract local features and has a small receptive field. The quality of the synthesized audio is not high, resulting in a low level of audio beautification.
A self-attention-based audio beautification method is adopted to extract the energy, content and timbre of the audio through the audio model. Position embedding and attention mechanism are used to encode and decode audio features. Pitch calibration is performed in combination with a pitch predictor to achieve audio beautification.
It improves the audio beautification level, removes noise and interference, ensures that the timbre characteristics remain unchanged, and makes pitch calibration more accurate, thus achieving the beautification of pitch and timbre in the audio.
Smart Images

Figure CN116612782B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing technology, and in particular to a self-attention-based audio beautification method, device, equipment and storage medium. Background Art
[0002] With the continuous development of technology, digital music has set off a wave of crazes online. However, due to the lack of skills of most ordinary people, the sound produced is less than satisfactory. Therefore, beautifying audio is extremely important.
[0003] In existing technology, beautifying raw audio involves two steps: pitch correction and timbre enhancement. Existing models are primarily based on generative algorithms (CVAEs), which can only extract local features and have a small receptive field. The timbre is simply altered through linear processing, which does not significantly improve the sound quality. The pitch-corrected and timbre-enhanced audio is then fused together using a synthesizer, but the synthesized audio quality is low, resulting in a low degree of audio enhancement. Summary of the Invention
[0004] The embodiments of the present invention provide a self-attention-based audio beautification method, apparatus, device, and storage medium to solve the problem of low audio beautification in the prior art.
[0005] A self-attention-based audio beautification method, comprising:
[0006] Get at least one audio to be processed;
[0007] Acquire an audio model, and extract content from all the audios to be processed using a content encoder in the audio model to obtain audio content corresponding to each of the audios to be processed;
[0008] Performing timbre extraction on all the audios to be processed by using a timbre encoder in the audio model to obtain an audio timbre corresponding to each of the audios to be processed;
[0009] Extracting energy from all the audios to be processed by an energy encoder in the audio model to obtain audio energy corresponding to each of the audios to be processed;
[0010] Performing position embedding on the audio content, the audio timbre, and the audio energy to obtain audio features;
[0011] Encoding the audio features by the encoding end of the audio model to obtain encoding features;
[0012] Acquire standard audio features and audio pitch, and decode the standard audio features, the encoding features, and the audio pitch through a decoding end of the audio model to obtain beautified audio.
[0013] An audio beautification device based on self-attention, comprising:
[0014] An audio acquisition module, configured to acquire at least one audio to be processed;
[0015] An audio content module is configured to obtain an audio model, and extract content from all the audios to be processed using a content encoder in the audio model to obtain audio content corresponding to each of the audios to be processed;
[0016] An audio timbre module, configured to extract timbre from all the audios to be processed using a timbre encoder in the audio model to obtain an audio timbre corresponding to each of the audios to be processed;
[0017] An audio energy module, configured to extract energy from all the audios to be processed using an energy encoder in the audio model to obtain audio energy corresponding to each of the audios to be processed;
[0018] An audio feature module, configured to perform position embedding on the audio content, the audio timbre, and the audio energy to obtain audio features;
[0019] An audio encoding module, configured to encode the audio features through an encoding end of the audio model to obtain encoding features;
[0020] The audio decoding module is used to obtain standard audio features and audio pitch, and decode the standard audio features, the encoding features and the audio pitch through the decoding end of the audio model to obtain beautified audio.
[0021] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the self-attention-based audio beautification method is implemented.
[0022] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned self-attention-based audio beautification method.
[0023] The present invention provides an audio beautification method, device, equipment and storage medium based on self-attention. The method extracts the energy, content and timbre of the audio to be processed respectively through an audio model, thereby realizing the extraction of audio energy, audio content and audio timbre in the audio, realizing the elimination of noise and interfering noise in the audio, and further improving the degree of audio beautification by adding energy features and invisible representation. By adding a position vector and an attention mechanism in the audio model, the position of timbre improvement is clarified. The pitch of the audio to be processed is predicted by a pitch predictor, making the pitch calibration more accurate, thereby keeping the timbre features unchanged while changing the pitch. The standard audio features, coding features and audio pitch are decoded by the decoding end in the audio model, thereby realizing the beautification of the pitch and timbre in the audio, and further realizing the determination of the beautified audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0025] Figure 1 Schematic diagram of an application environment of an audio beautification method based on self-attention in one embodiment of the present invention;
[0026] Figure 2 is a flow chart of an audio beautification method based on self-attention in one embodiment of the present invention;
[0027] Figure 3 is a flowchart of step S20 of the audio beautification method based on self-attention in one embodiment of the present invention;
[0028] Figure 4 is a flowchart of step S60 of the audio beautification method based on self-attention in one embodiment of the present invention;
[0029] Figure 5 is a principle block diagram of an audio beautification device based on self-attention in one embodiment of the present invention;
[0030] Figure 6 is a schematic diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0032] The audio beautification method based on self-attention provided by the embodiment of the present invention can be applied as follows: Figure 1 Specifically, the self-attention-based audio beautification method is applied in a self-attention-based audio beautification device, which includes: Figure 1 The client and server shown in the figure communicate with each other through the network to solve the problem of low audio beautification in the prior art. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The client, also known as the user end, refers to a program that corresponds to the server and provides classified services to customers. The client can be installed on, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices.
[0033] In one embodiment, if Figure 2 As shown in the figure, a self-attention-based audio beautification method is provided, which is applied in Figure 1 The server in the example is used as an example, and the steps are as follows:
[0034] S10: Obtain at least one audio to be processed.
[0035] Understandably, the audio to be processed can be a song sung by an ordinary person, a conversation between different people, or a person's recitation converted into two-dimensional frequency domain data, namely a Mel-spectrogram. The audio to be processed can be collected from various databases or pre-prepared data sent from a client to a database. For example, a Mel-spectrogram of a song sung by an ordinary person or a recitation of a poem can be used.
[0036] S20: Acquire an audio model, and extract content from all the audios to be processed using a content encoder in the audio model to obtain audio content corresponding to each of the audios to be processed.
[0037] Understandably, audio content refers to the content of the audio being processed, such as song lyrics or words from a poem. The content encoder is an encoder used for content extraction, built based on the Conformer model. The audio model is a model for beautifying the audio, and is an improvement on the NSVB model.
[0038] Specifically, after obtaining the audio to be processed, an audio model is obtained, and all the audio to be processed is input into the audio model. The content encoder in the audio model is used to extract the content of the audio to be processed. That is, the downsampling layer in the content encoder samples the beautified singing voice. That is, the downsampling layer in the content encoder performs data enhancement, convolution pooling, and linear transformation to obtain downsampling features. The downsampling features are processed by the attention layer in the content encoder. That is, the downsampling features are processed by the first feedback-forward layer and then by the multi-head attention mechanism to obtain attention features. The attention features are convolved by the convolution layer in the content encoder. That is, first, a gating mechanism consisting of point-by-point convolution and linear gating units is used, followed by a one-dimensional depth-separated convolution, and then a normalization layer is added for normalization to obtain the audio content.
[0039] S30, performing timbre extraction on all the audios to be processed by using a timbre encoder in the audio model to obtain an audio timbre corresponding to each of the audios to be processed.
[0040] Understandably, timbre refers to the fact that different sounds always have unique characteristics in terms of waveform, just as different objects have different vibration characteristics. Audio timbre refers to the sound quality of the audio being processed.
[0041] Specifically, after obtaining the audio to be processed, the timbre of all the audio to be processed is extracted using the timbre encoder in the audio model. Specifically, multiple matrices representing the timbre of the audio to be processed are extracted from the mel-spectrogram along the time domain. These multiple matrices are fully connected to obtain a connection matrix, and several matrices that are continuous in the time domain are selected from the connection matrix to obtain the audio timbre. Alternatively, the open source API resemblyzer8 Python package can be used as a timbre encoder to perform timbre extraction, and the timbre extraction method is not limited in this embodiment.
[0042] S40: Extracting energy from all the audios to be processed by an energy encoder in the audio model to obtain audio energy corresponding to each audio to be processed.
[0043] It can be understood that audio energy is the energy of each frame of audio in the time domain, which is the square of the amplitude. The preset energy encoder is an encoder set in advance for extracting energy from audio.
[0044] Specifically, after obtaining the audio to be processed, the energy encoder in the audio model is used to extract energy from all the audio to be processed. That is, the audio to be processed is first sampled, that is, the sampling frequency and sampling interval are determined. The sampling points of the audio to be processed are taken at the sampling interval and sampling frequency. The sampling points are then Fourier transformed to obtain the amplitude spectrum. The amplitude spectrum is then squared to obtain the energy spectrum. Based on the energy spectrum, the audio energy of the audio to be processed is determined. In this way, the audio energy corresponding to each audio to be processed can be obtained through the above method.
[0045] S50: Perform position embedding on the audio content, the audio timbre, and the audio energy to obtain audio features.
[0046] S60: Encoding the audio features through the encoding end of the audio model to obtain encoding features.
[0047] It can be understood that the coding feature is obtained by encoding the audio feature at the encoding end of the audio model. The audio feature is obtained by embedding and concatenating the audio content, audio timbre, and audio energy position.
[0048] Specifically, after obtaining the audio content, audio timbre, and audio energy, the audio content, audio timbre, and audio energy are positionally embedded. That is, the audio content, audio timbre, and audio energy are positionally embedded in time sequence to obtain position vectors. The three position vectors are then connected to obtain audio features. Furthermore, the audio features are input into the encoding end of the audio model, where they are encoded. Specifically, attention is first processed through a self-attention mechanism to obtain an attention sequence. The attention sequence is then subjected to feature extraction through a one-dimensional convolutional layer to obtain approximate mean and logarithmic scale standard deviation parameters, which are determined as encoded features.
[0049] S70: Obtain standard audio features and audio pitch, and decode the standard audio features, the encoding features, and the audio pitch through the decoding end of the audio model to obtain beautified audio.
[0050] Understandably, beautified audio is achieved by improving pitch and timbre. Audio pitch is obtained by feature extraction using a pitch predictor. Standard audio features are specified audio data, such as a professional singer's performance, while the processed audio is an amateur singer's performance. The decoder is built based on the WaveNet model, consisting of a preset number (e.g., 4) of wavelet network layers and one-dimensional convolutional layers.
[0051] Specifically, after obtaining the encoded features, the attention features are mapped using a latent mapping algorithm, namely, the mapping function M, which converts the latent variable from qφ(za|xa,ca) to qφ(zp|xp,cp). A latent variable representing the amateur timbre za is extracted from qφ(za|xa,ca) and mapped to M(za) to obtain the mapping features. The pitch of the processed audio is predicted by a pitch predictor to obtain the audio pitch. The audio pitch, the mapping features, and the standard audio features are fused to obtain the fused features. The fused features are then decoded by the decoder, that is, through a preset number of wavelet network layers and one-dimensional convolutional layers to obtain the beautified audio.
[0052] In an embodiment of the present invention, a self-attention-based audio beautification method is provided. The method extracts the energy, content, and timbre of the audio to be processed respectively through an audio model, thereby realizing the extraction of audio energy, audio content, and audio timbre in the audio, realizing the elimination of noise and interfering noise in the audio, and further improving the degree of audio beautification by adding energy features and invisible representation. By adding a position vector and an attention mechanism in the audio model, the position of timbre improvement is clarified. The pitch of the audio to be processed is predicted by a pitch predictor, making the pitch calibration more accurate, thereby keeping the timbre features unchanged while changing the pitch. The standard audio features, coding features, and audio pitch are decoded at the decoding end in the audio model, thereby realizing the beautification of the pitch and timbre in the audio, and further realizing the determination of the beautified audio.
[0053] In one embodiment, step S10, that is, before obtaining at least one audio to be processed, includes:
[0054] S101: Acquire at least one initial audio, and perform audio processing on all the initial audio to obtain Fourier spectra corresponding to the initial audio.
[0055] S102: Perform spectrum conversion on all the Fourier spectra to obtain at least one audio to be processed.
[0056] Understandably, the Fourier spectrum is obtained by processing the original audio. The processed audio is the Mel-spectrogram of the original audio. The original audio can be collected from different databases or sent to the database from the client, for example, the audio of ordinary people singing songs or reciting poetry.
[0057] Specifically, all the initial audio is obtained, and audio processing is performed on all the initial audio, that is, the initial audio is framed in a fixed time period (such as 25 milliseconds), that is, the initial audio is divided into audio signals of frames, and a plurality of framing units are obtained. Wherein, in order to avoid excessive changes in adjacent framing units, an overlapping area is provided between two adjacent framing units. Then, windowing is performed on the audio signal of each frame, that is, by multiplying each framing unit by a window function so that the left and right ends of each framing unit have continuity, and a windowed audio signal is obtained. Discrete Fourier transform is performed on the windowed audio signal, that is, the windowed audio signal is converted into energy in the frequency domain, and the frequency spectrum distributed in different time windows on the time axis is obtained, and a Fourier spectrum can be obtained. Furthermore, the Fourier spectrum is low-pass filtered using a Mel filter, converting the linear natural spectrum into a Mel spectrum that reflects the characteristics of human hearing. A Mel (scaled) filter bank is then used to transform the linear spectrum group into a Mel spectrum group to simulate the linear frequency perception relationship of the human ear. This yields a Mel spectrum graph corresponding to each initial audio source, which is then determined as the audio to be processed. Low resolution preserves more pitch information, while high resolution preserves more timbre information.
[0058] The embodiment of the present invention performs audio processing on the initial audio to convert the one-dimensional, difficult-to-process time series signal into two-dimensional frequency domain data that is easy to process and richer in information, thereby determining the Mel-spectrogram corresponding to each initial audio, thereby facilitating the subsequent extraction of energy, content, and timbre in the audio.
[0059] In one embodiment, if Figure 3 As shown, in step S20, that is, extracting content from all the audios to be processed by the content encoder in the audio model to obtain audio content corresponding to each audio to be processed, including:
[0060] S201 : Sampling the audio to be processed by a downsampling layer in the content encoder to obtain downsampling features.
[0061] It can be understood that the downsampled features are obtained by sampling the audio to be processed. The content encoder is built based on the comformer and is used to extract the content in the audio.
[0062] Specifically, after obtaining the audio to be processed, the trained and tested audio model is retrieved and all the audio to be processed is input into the audio model. The content of the audio to be processed is extracted through the content encoder in the audio model. In other words, the audio to be processed is sampled through the downsampling layer in the content encoder, which also performs data enhancement on the audio to be processed. The enhanced audio to be processed is then convolutionally pooled to obtain a convolutional pooling vector. The convolutional pooling vector is then linearly transformed, and the overfitting problem that occurs during the process is reduced through the dropout layer to obtain the downsampled features.
[0063] S202: Perform attention processing on the downsampled features through an attention layer in the content encoder to obtain attention features.
[0064] It can be understood that the attention feature is obtained by performing attention processing on the downsampled features.
[0065] Specifically, after obtaining the downsampled features, they are preprocessed through the first feedforward network layer. This involves normalizing them through a normalization layer, performing linear transformations through two linear layers, and performing activation processing via a swish function between the two linear transformations. A dropout layer is then used to reduce overfitting problems that occur during the process. Finally, a residual summation operation is performed to obtain the network vector. The network vector is then subjected to attention processing through an attention layer, using a multi-head attention mechanism. Relative sinusoidal position encoding is used to achieve better performance and greater robustness on variable-length inputs. The dropout layer also reduces some neural units to obtain the attention features.
[0066] S203: Perform convolution processing on the attention feature through the convolution layer in the content encoder to obtain audio content.
[0067] Specifically, after obtaining the attention feature, the attention feature is convolved through the convolution layer in the content encoder. That is, it is first normalized through a normalization layer and then processed through a gating mechanism consisting of point-by-point convolution and a linear gated unit (GLU) to obtain a convolution vector. The convolution vector is then depth-wise convolved through a one-dimensional depth-wise separation convolution layer, that is, the convolution vector is depth-wise convolved and point-by-point convolution is performed to obtain a depth-wise convolution vector. The depth-wise convolution vector is then normalized through a normalization layer and activated through a swish function to obtain a normalized vector. Finally, a point-by-point convolution is performed to reduce overfitting problems that occur during the convolution process and a dropout layer. Finally, the convolution vector is processed through a second feedforward network, mapped to the original dimension through two linear transformations, and then through a dropout layer to reduce overfitting problems that occur during the process. Finally, the residual summation operation is performed. After multiple iterations of the above process, the audio content can be obtained.
[0068] In this embodiment of the present invention, the downsampling layer samples the audio to be processed, thereby acquiring downsampled features. The attention layer performs attention processing on the downsampled features, using relative sinusoidal position encoding to achieve better performance and greater robustness on variable-length inputs. The convolutional layer convolves the attention features to extract the audio content, and the dropout layer performs random deactivation to enhance generalization.
[0069] In one embodiment, before step S20, that is, before obtaining the audio model, the following steps are included:
[0070] S204 , obtaining a sample training data set, where the sample training data set includes at least one sample training data; one sample training data corresponds to one sample label.
[0071] Understandably, sample training data is a spectrogram of various audio data. For example, sample training data can be a song sung by a singer, a conversation between different people, or a spectrogram of a single person reciting aloud. Each sample training data item corresponds to a sample label, which is used to characterize the beautified audio of each sample training data item. Sample training data can be collected from different databases or sent from the client to the server. A sample training dataset is then constructed based on all the acquired sample training data.
[0072] S205: Obtain a preset training model, and predict the sample training data using the preset training model to obtain a prediction label.
[0073] It is understood that the preset training model is a model set in advance for predicting sample training data, which is improved based on the NSVB model. The predicted label is the one predicted by the preset training model on the sample training data, that is, the beautified audio.
[0074] Specifically, after obtaining a sample training dataset, data augmentation is performed on all sample training data. The augmented data is then divided into a training set and a test set. The training set is used to train the pre-set training model, while the test set is used to test the trained pre-set training model. The content, timbre, and energy of the sample training data in the training set are extracted to obtain sample content, sample timbre, and sample energy. Positional embedding is then performed on the sample content, sample timbre, and sample energy, respectively, and these are concatenated according to the position vector to obtain sample audio features. This is then processed through an attention mechanism and convolution with a one-dimensional convolutional neural network to obtain sample encoding features. The sample encoding features are then mapped using a latent mapping algorithm. Specifically, the latent variable qφ(za|xa,ca) is converted to qφ(zp|xp,cp) using a mapping function M, resulting in a mapping vector. The predicted pitch and sample standard audio are obtained, and the mapping vector, predicted pitch, and sample standard audio are input into the decoder. The decoding end decodes the mapping vector, predicted pitch, and sample standard audio. In other words, the decoding end fuses the mapping vector, predicted pitch, and sample standard audio, and decodes the fused audio. In other words, the fused audio is processed through multiple wavelet network layers and one-dimensional convolutional layers to obtain the predicted label. For example, professional pitch, calibrated amateur content vector, and amateur timbre are mixed to obtain a new condition, and a new beautified mel-spectrogram is generated with the mapping vector after mapping by the audio model encoder. The new beautified mel-spectrogram is then converted to audio by the vocoder to obtain the beautified audio.
[0075] S206: Determine a prediction loss value of the preset training model according to the sample label and the prediction label corresponding to the same sample training data.
[0076] It can be understood that the prediction loss value is generated in the process of predicting the predicted label of the sample training data, and is used to represent the difference between the sample label and the predicted label.
[0077] Specifically, after obtaining the predicted label, all sample labels corresponding to the sample training data are arranged according to the order of the sample training data in the sample training data set, and then the predicted label associated with the sample training data is compared with the sample label of the sample training data with the same sequence; that is, according to the sorting of the sample training data, the sample label corresponding to the first sample training data is compared with the predicted label corresponding to the first sample training data, and the loss value between the sample label and the predicted label is determined by the loss function; and then the sample label corresponding to the second sample training data is compared with the predicted label corresponding to the second sample training data, until all sample labels and predicted labels are compared, and the predicted loss value of the preset training model can be determined.
[0078] S207, when the predicted loss value does not reach the preset convergence condition, iteratively update the initial parameters in the preset training model until the predicted loss value reaches the convergence condition, and record the preset training model after convergence as the audio model.
[0079] Understandably, the convergence condition may be that the predicted loss value is less than a set threshold, or that the predicted loss value is very small and will not decrease after 500 calculations, and the training is stopped.
[0080] Specifically, after determining the predicted loss value of the preset training model, when the predicted loss value does not reach the preset convergence condition, the initial parameters of the preset training model are adjusted according to the predicted loss value, and all sample training data are re-input into the preset training model after adjusting the initial parameters, and the predicted loss value corresponding to the preset training model with the adjusted initial parameters is obtained, and when the predicted loss value does not reach the preset convergence condition, the initial parameters of the preset training model are adjusted again according to the predicted loss value, so that the predicted loss value of the preset training model with the adjusted initial parameters reaches the convergence condition. In this way, the result output by the preset training model can continuously approach the accurate result, making the prediction accuracy higher and higher, until the predicted loss values of all sample training data reach the preset convergence condition, and the converged preset training model is recorded as the audio model.
[0081] Furthermore, the audio model is tested for performance using sample training data from the test set to obtain evaluation metrics, generalization ability, overfitting, and underfitting. Based on these metrics, generalization ability, overfitting, and underfitting, the audio model's performance is judged to be meeting the requirements. If the performance does not meet the requirements, the audio model is retrained until the performance meets the requirements, thereby obtaining the audio model.
[0082] This embodiment of the present invention trains a preset training model using sample training data from the training and test sets. A loss function is then used to determine the predicted loss between the predicted label and the sample label. The initial parameters of the preset training model are adjusted based on the predicted loss until the model converges, thereby confirming the audio model and ensuring a high degree of prediction accuracy.
[0083] In one embodiment, in step S50, position embedding is performed on the audio content, the audio timbre, and the audio energy to obtain audio features, including:
[0084] S501 , position embedding is performed on the audio content, the audio timbre, and the audio energy to obtain a content position vector corresponding to the audio content, a timbre position vector corresponding to the audio timbre, and an energy position vector corresponding to the audio energy.
[0085] S502 : Connect the audio content, the audio timbre, and the audio energy through the content position vector, the timbre position vector, and the energy position vector to obtain audio features.
[0086] Understandably, the content position vector is the position of each audio content. The timbre position vector is the position of each audio timbre. The energy position vector is the position of each audio energy. The audio feature is obtained by concatenating all audio content, audio energy, and audio timbre after position embedding.
[0087] Specifically, after obtaining the audio content, audio timbre and audio energy respectively, the audio content, audio timbre and audio energy are input into the position embedding layer, and the audio content, audio timbre and audio energy are positionally embedded respectively through the position embedding layer. That is, a position vector can be set for each audio content, each audio timbre and each audio energy according to the time sequence of the audio to be processed, which is used to facilitate the subsequent determination of the timbre change position, and the content position vector corresponding to the audio content, the timbre position vector corresponding to the audio timbre and the energy position vector corresponding to the audio energy can be obtained.
[0088] Furthermore, the audio content, audio timbre, and audio energy corresponding to the same audio to be processed are concatenated according to the content position vector, timbre position vector, and energy position vector, respectively. Specifically, all audio content is concatenated according to the content position vector of each audio segment, all audio content is concatenated according to the timbre position vector of each audio segment, and all audio content is concatenated according to the energy position vector of each audio segment. The concatenated audio content, audio timbre, and audio energy are then concatenated in parallel to obtain audio features.
[0089] By introducing position vectors, the present invention makes it possible to clearly identify the audio beautification parts during subsequent audio enhancement, specifically the locations of the timbre improvements. Connecting the various parts of the same audio to be processed facilitates subsequent timbre improvements and improves the robustness of the model.
[0090] In one embodiment, if Figure 4 As shown, in step S60, the audio features are encoded by the encoding end of the audio model to obtain encoded features, including:
[0091] S601, performing attention processing on the audio features through the attention layer of the encoding end in the audio model to obtain an attention sequence.
[0092] It can be understood that the attention sequence is obtained by performing attention processing and normalization on the audio features.
[0093] Specifically, after obtaining the audio model, all audio features are processed with attention through multiple attention mechanisms, that is, the Q vector, K vector, and V vector in the audio features are calculated through multiple attention mechanisms, that is, the correlation score between the Q vector and the K vector in the input vector is calculated using the dot product method, that is, the dot product is calculated with each input vector in Q and each input vector in K, and the correlation score between the Q vector and the K vector is normalized. Then, through the softmax function, the score vector between the input vectors is converted into a probability distribution between [0, 1], and according to the probability distribution between the input vectors, it is multiplied by the corresponding Values vector to obtain the attention result. Finally, the attention results of different groups are spliced together to obtain a spliced vector. The spliced vector is then subjected to residual summation and normalization to obtain an attention sequence.
[0094] S602: Perform convolution processing on the attention sequence through the convolution layer of the encoding end in the audio model to obtain encoding features.
[0095] It can be understood that the attention feature is obtained by convolving the attention sequence through the convolution layer of the encoding end. The convolution layer is a one-dimensional CNN (convolutional neural network).
[0096] Specifically, after obtaining the attention sequence, the attention sequence is convolved through the convolution layer, that is, the feature extraction of the attention sequence is performed with the preset convolution kernel and stride to obtain the convolution feature. The convolution feature is then downsampled through the pooling layer, that is, the convolution feature is reduced in dimension from high latitude to low latitude to obtain the pooled feature. Finally, the pooled feature is predicted through the fully connected layer, that is, the pooled feature is predicted based on the weight of each neuron feedback to obtain the prediction result. The residuals of all prediction results are summed and normalized to obtain the attention feature.
[0097] The embodiment of the present invention performs attention processing on audio features through an attention layer, thereby determining the attention sequence and, in turn, determining the location of the timbre modification. The attention sequence is convolved through a convolutional layer, thereby determining the approximate mean and logarithmic scale standard deviation parameters and, in turn, determining the encoding features. Convolution through a convolutional neural network allows for the extraction of local window features and better sequence features, thereby improving the robustness of the model.
[0098] In one embodiment, before step S70, that is, before obtaining the audio pitch, the following steps are included:
[0099] S701: Obtain a pitch predictor, and perform one-dimensional convolution processing on the audio to be processed through a first convolution layer in the pitch predictor to obtain a first convolution feature.
[0100] S702: Perform convolution processing on the first convolution feature through the second convolution layer in the pitch predictor to obtain a second convolution feature.
[0101] S703: Perform one-dimensional convolution processing on the second convolution feature through the third convolution layer in the pitch predictor to obtain the audio pitch.
[0102] Understandably, the pitch predictor is composed of three stacked one-dimensional convolutions. The first convolution feature is obtained by extracting the pitch feature of the processed audio. The second convolution feature is obtained by performing feature convolution on the first convolution feature. The audio pitch is improved by the pitch predictor.
[0103] Specifically, all phonemes in the audio to be processed are determined, for example, the RNNT (RNN-Transducer) model can be adopted, which is not limited here, and then the time of each phoneme in the current audio to be processed is calculated, and the phonemes and pitches are aligned by time, that is, the pitch of each time in the audio to be processed is converted into the pitch of each phoneme. The pitch of each phoneme is calibrated by a pitch predictor, that is, the pitch in the audio to be processed is calibrated and predicted by three stacked one-dimensional convolutional networks, that is, the audio to be processed is subjected to one-dimensional convolution processing by the first convolution layer to obtain a first convolution feature. The first convolution feature is subjected to convolution processing by the second convolution layer to obtain a second convolution feature. The second convolution feature is subjected to one-dimensional convolution processing by the third convolution layer to obtain the audio pitch.
[0104] The embodiment of the present invention predicts the pitch in the audio to be processed through a pitch predictor, thereby achieving determination of the audio pitch and calibration of the pitch in the audio to be processed, thereby improving the accuracy of the pitch calibration.
[0105] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0106] In one embodiment, a self-attention-based audio beautification device is provided, which corresponds to the self-attention-based audio beautification method in the above embodiment. Figure 5 As shown, the self-attention-based audio beautification device includes an audio acquisition module 11, an audio content module 12, an audio timbre module 13, an audio energy module 14, an audio feature module 15, an audio encoding module 16, and an audio decoding module 17. The functional modules are described in detail as follows:
[0107] An audio acquisition module 11 is configured to acquire at least one audio to be processed;
[0108] The audio content module 12 is configured to obtain an audio model, and extract content from all the audios to be processed using a content encoder in the audio model to obtain audio content corresponding to each of the audios to be processed;
[0109] An audio timbre module 13 is configured to extract timbre from all the audios to be processed using a timbre encoder in the audio model to obtain an audio timbre corresponding to each of the audios to be processed;
[0110] An audio energy module 14 is configured to extract energy from all the audios to be processed using an energy encoder in the audio model to obtain audio energy corresponding to each audio to be processed;
[0111] An audio feature module 15 is configured to perform position embedding on the audio content, the audio timbre, and the audio energy to obtain audio features;
[0112] An audio encoding module 16, configured to encode the audio features through an encoding end of the audio model to obtain encoding features;
[0113] The audio decoding module 17 is used to obtain standard audio features and audio pitch, and decode the standard audio features, the encoding features and the audio pitch through the decoding end of the audio model to obtain beautified audio.
[0114] In one embodiment, the audio content module 12 includes:
[0115] a downsampling unit, configured to sample the audio to be processed through a downsampling layer in the content encoder to obtain downsampling features;
[0116] an attention unit, configured to perform attention processing on the downsampled features through an attention layer in the content encoder to obtain attention features;
[0117] A convolution unit is used to convolve the attention feature through a convolution layer in the content encoder to obtain audio content.
[0118] In one embodiment, the audio acquisition module 11 includes:
[0119] an audio processing unit, configured to obtain at least one initial audio, and perform audio processing on all the initial audios to obtain Fourier spectra corresponding to the initial audios;
[0120] The spectrum conversion unit is used to perform spectrum conversion on all the Fourier spectra to obtain at least one audio to be processed.
[0121] In one embodiment, the audio content module 12 further includes:
[0122] A sample acquisition unit is used to acquire a sample training data set, wherein the sample training data set includes at least one sample training data; one sample training data corresponds to one sample label;
[0123] A content prediction unit, configured to obtain a preset training model, and predict the sample training data using the preset training model to obtain a prediction label;
[0124] A loss prediction unit, configured to determine a predicted loss value of the preset training model based on the sample label and the predicted label corresponding to the same sample training data;
[0125] A model convergence unit is used to iteratively update the initial parameters in the preset training model when the predicted loss value does not reach the preset convergence condition, until the predicted loss value reaches the convergence condition, and record the preset training model after convergence as an audio model.
[0126] In one embodiment, the audio feature module 15 includes:
[0127] a position embedding unit, configured to perform position embedding on the audio content, the audio timbre, and the audio energy to obtain a content position vector corresponding to the audio content, a timbre position vector corresponding to the audio timbre, and an energy position vector corresponding to the audio energy;
[0128] A vector connection unit is used to connect the audio content, the audio timbre and the audio energy through the content position vector, the timbre position vector and the energy position vector to obtain audio features.
[0129] In one embodiment, the audio encoding module 16 includes:
[0130] An attention sequence unit, configured to perform attention processing on the audio features through an attention layer of an encoder in the audio model to obtain an attention sequence;
[0131] The encoding feature unit is used to convolve the attention sequence through the convolution layer of the encoding end in the audio model to obtain the encoding feature.
[0132] In one embodiment, the audio decoding module 17 includes:
[0133] a first convolution unit, configured to obtain a pitch predictor, and perform one-dimensional convolution processing on the audio to be processed through a first convolution layer in the pitch predictor to obtain a first convolution feature;
[0134] a second convolution unit, configured to perform convolution processing on the first convolution feature through a second convolution layer in the pitch predictor to obtain a second convolution feature;
[0135] The third convolution unit is configured to perform one-dimensional convolution processing on the second convolution feature through the third convolution layer in the pitch predictor to obtain the audio pitch.
[0136] The specific limitations of the self-attention-based audio beautification device can be found in the limitations of the self-attention-based audio beautification method above and will not be repeated here. Each module in the self-attention-based audio beautification device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0137] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data used in the audio beautification method based on self-attention in the above-mentioned embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for audio beautification based on self-attention is implemented.
[0138] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned self-attention-based audio beautification method when executing the computer program.
[0139] In one embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the self-attention-based audio beautification method is implemented.
[0140] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0141] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0142] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. An audio beautification method based on self-attention, characterized in that: include: Get at least one audio to be processed; Acquire an audio model, and extract content from all the audios to be processed using a content encoder in the audio model to obtain audio content corresponding to each of the audios to be processed; Performing timbre extraction on all the audios to be processed by using a timbre encoder in the audio model to obtain an audio timbre corresponding to each of the audios to be processed; Extracting energy from all the audios to be processed by an energy encoder in the audio model to obtain audio energy corresponding to each of the audios to be processed; Performing position embedding and then splicing on the audio content, the audio timbre, and the audio energy to obtain audio features; Encoding the audio features by the encoding end of the audio model to obtain encoding features; Obtaining standard audio features and audio pitch, and decoding the standard audio features, the encoding features, and the audio pitch through the decoding end of the audio model to obtain beautified audio; wherein the audio pitch is obtained by processing the audio to be processed; The encoding process of the audio features by the encoding end of the audio model to obtain the encoded features includes: Performing attention processing on the audio features through a self-attention mechanism at the encoding end of the audio model to obtain an attention sequence; The attention sequence is subjected to feature extraction by a convolutional layer at the encoding end in the audio model to obtain encoding features.
2. The audio beautification method based on self-attention according to claim 1, characterized in that The extracting content of all the audios to be processed by the content encoder in the audio model to obtain audio content corresponding to each of the audios to be processed includes: Sampling the audio to be processed by a downsampling layer in the content encoder to obtain downsampling features; Performing attention processing on the downsampled features through an attention layer in the content encoder to obtain attention features; The attention feature is convolved by a convolution layer in the content encoder to obtain audio content.
3. The audio beautification method based on self-attention according to claim 1, characterized in that The performing position embedding on the audio content, the audio timbre, and the audio energy to obtain audio features includes: Performing position embedding on the audio content, the audio timbre, and the audio energy to obtain a content position vector corresponding to the audio content, a timbre position vector corresponding to the audio timbre, and an energy position vector corresponding to the audio energy; The audio content, the audio timbre and the audio energy are connected through the content position vector, the timbre position vector and the energy position vector to obtain audio features.
4. The audio beautification method based on self-attention according to claim 1, characterized in that Before obtaining the audio pitch, the following steps are included: Obtain a pitch predictor, and perform one-dimensional convolution processing on the audio to be processed through a first convolution layer in the pitch predictor to obtain a first convolution feature; Convolutionally processing the first convolutional feature by a second convolutional layer in the pitch predictor to obtain a second convolutional feature; Performing one-dimensional convolution processing on the second convolution feature through the third convolution layer in the pitch predictor to obtain the audio pitch.
5. The audio beautification method based on self-attention according to claim 1, wherein Before obtaining the audio model, the following steps are included: Obtain a sample training data set, wherein the sample training data set includes at least one sample training data; one sample training data corresponds to one sample label; Obtain a preset training model, and predict the sample training data using the preset training model to obtain a prediction label; Determining a prediction loss value of the preset training model based on the sample label and the prediction label corresponding to the same sample training data; When the predicted loss value does not reach the preset convergence condition, the initial parameters in the preset training model are iteratively updated until the predicted loss value reaches the convergence condition, and the preset training model after convergence is recorded as the audio model.
6. The audio beautification method based on self-attention according to claim 1, characterized in that Before obtaining at least one audio to be processed, the method includes: Acquire at least one initial audio, and perform audio processing on all the initial audios to obtain Fourier spectra corresponding to the initial audios; Perform spectrum conversion on all the Fourier spectra to obtain at least one audio to be processed.
7. An audio beautification device based on self-attention, characterized in that include: An audio acquisition module, configured to acquire at least one audio to be processed; An audio content module is configured to obtain an audio model, and extract content from all the audios to be processed using a content encoder in the audio model to obtain audio content corresponding to each of the audios to be processed; An audio timbre module, configured to extract timbre from all the audios to be processed using a timbre encoder in the audio model to obtain an audio timbre corresponding to each of the audios to be processed; an audio energy module, configured to extract energy from all the audios to be processed using an energy encoder in the audio model to obtain audio energy corresponding to each of the audios to be processed; An audio feature module, configured to perform position embedding and then splicing of the audio content, the audio timbre, and the audio energy to obtain audio features; An audio encoding module, configured to encode the audio features through an encoding end of the audio model to obtain encoding features; An audio decoding module, configured to obtain standard audio features and audio pitch, and decode the standard audio features, the encoding features, and the audio pitch through a decoding end of the audio model to obtain beautified audio; wherein the audio pitch is obtained by processing the audio to be processed; The audio encoding module is further used to: Performing attention processing on the audio features through a self-attention mechanism at the encoding end of the audio model to obtain an attention sequence; The attention sequence is subjected to feature extraction by a convolutional layer at the encoding end in the audio model to obtain encoding features.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the self-attention-based audio beautification method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the self-attention-based audio beautification method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Speech synthesis model, model training method and speech synthesis method
CN113920977A
Voice conversion method, voice conversion device, electronic equipment and storage medium
CN115294995A