Audio style conversion method and device, equipment and medium

Through the encoder-decoder architecture of a generative adversarial network, the problem of inefficient audio style conversion technology is solved, and efficient and accurate audio style conversion in the medical and financial fields is achieved to generate realistic audio to meet diagnostic and analysis needs.

CN120236598APending Publication Date: 2025-07-01PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510441880.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing audio style conversion technology is inefficient in the fields of medical health and financial technology, and it is difficult to accurately distinguish normal audio from lesion audio, unable to effectively remove environmental noise interference, and it is difficult to uniformly convert while ensuring the complete and accurate voice content, resulting in limited diagnostic accuracy and risk warning effectiveness.

Method used

Using an encoder-decoder architecture based on a generative adversarial network, an adversarial training generator and potential discriminator generates potential features that are independent of the target category to realize audio style conversion without parallel speech, transcription or time alignment processes, and a global conditional mechanism is used to perform one-time audio conversion.

Benefits of technology

Improves the efficiency and accuracy of audio style conversion, generates more realistic and natural audio, avoids the need for alignment and speaker-independent automatic speech recognition networks, and meets the complex needs of medical diagnosis and financial analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236598A_ABST
    Figure CN120236598A_ABST
Patent Text Reader

Abstract

The invention relates to the voice processing field, the financial science and technology field and the medical health field, and discloses an audio style conversion method and device, equipment and a medium, and the method comprises the steps: extracting the acoustic features of a target audio, generating a style embedding tensor according to the acoustic features through a reference encoder in a preset network architecture, generating potential features according to the acoustic features by using an encoder in the network architecture, minimizing a loss function of the encoder by adversarial training according to the potential features, generating optimized potential features according to the acoustic features by using the optimized encoder, and performing audio reconstruction by combining a style embedding tensor to obtain reconstructed audio. And judging whether the encoder is optimized or not according to the reconstruction loss values of the reconstructed audio and the test audio, and after judging that the encoder is optimized, performing audio style conversion on the to-be-processed audio based on the network architecture to obtain converted audio. And the audio style conversion efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of speech processing, financial technology and medical health, and in particular to an audio style conversion method, device, equipment and medium. Background Art

[0002] In the fields of healthcare and financial technology, the application potential of audio style transfer technology has gradually become prominent, but existing technologies still have many shortcomings in meeting the complex needs of specific fields. In the field of healthcare, the analysis and processing of medical audio covers a variety of complex audio data such as heart sounds and lung sounds. These audios present extremely complex characteristics due to individual differences among patients, the diversity of disease types, and the complexity of the collection environment. Existing audio style transfer methods are difficult to accurately distinguish the styles of normal audio from those of pathological audio, and thus cannot provide doctors with clear clues to the characteristics of the lesions. For example, when identifying heart murmurs, due to the subtle differences in audio styles between different murmurs, existing conversion technologies are difficult to extract key features, resulting in limited accuracy in early diagnosis of the disease. At the same time, for the interference of environmental noise, existing technologies lack effective removal and conversion mechanisms, making the audio analysis results susceptible to noise and deviations.

[0003] In the field of financial technology, audio style transfer technology is mainly used in scenarios such as customer identity authentication, risk assessment, and financial information broadcasting. Taking customer service voice interaction as an example, the audio generated by customers in different emotional states, accent differences, and call environments poses a huge challenge to audio style transfer. Existing technologies make it difficult to uniformly convert audio of different styles while ensuring the integrity and accuracy of the voice content to meet the needs of subsequent intelligent analysis. In addition, in the voice prompts of financial transaction risk warnings, it is necessary to convert audio styles with strong recognition according to different risk levels, but the existing methods are insufficient in achieving precise control and differentiation of styles, which greatly reduces the effectiveness of risk warnings.

[0004] At present, existing audio style transfer technologies mainly rely on statistical methods such as Gaussian mixture models, which require complex feature extraction processes and a large amount of parallel time-aligned speech data. In recent years, some studies have overcome the need for parallel data by combining attribute labels and acoustic features for local conditioning, but there is still the limitation of relying on a specific vocabulary during training, and only target speakers that have appeared in the training phase can be converted. Some studies have attempted to overcome the limitations of speech conversion, but they mainly rely on the accuracy of automatic speech recognition systems and require intermediate speech transcription. Speech conversion cannot be completed in one go, which reduces conversion efficiency and has limitations. Summary of the invention

[0005] The present invention provides an audio style conversion method, apparatus, device and medium to solve the problems of low efficiency and high limitations of existing audio style conversion methods in the current market.

[0006] In a first aspect, an audio style conversion method is provided, including:

[0007] Performing feature extraction on a pre-acquired target audio to obtain acoustic features, and using a reference encoder of a network architecture to generate a style embedding tensor according to the acoustic features;

[0008] Using an encoder of the network architecture to generate latent features according to the acoustic features;

[0009] According to the latent features, adversarially training to minimize the loss function of the encoder and the loss function of a latent discriminator of the network architecture, and using the optimized encoder to generate optimized latent features according to the acoustic features;

[0010] Using a decoder of the network architecture to perform audio reconstruction according to the optimized latent features and the style embedding tensor to obtain a reconstructed audio, and outputting a test reconstructed audio according to the optimized latent features;

[0011] Judging whether the encoder is optimized according to the reconstruction loss value between the test reconstructed audio and the reconstructed audio;

[0012] If the encoder is not optimized, returning to the step of using the encoder of the network architecture to generate latent features according to the acoustic features;

[0013] If the encoder is optimized, performing audio style conversion on the audio to be processed based on the network architecture to obtain a converted audio.

[0014] In a second aspect, an audio style conversion apparatus is provided, including:

[0015] A feature extraction module, configured to perform feature extraction on a pre-acquired target audio to obtain acoustic features, and use a reference encoder of a network architecture to generate a style embedding tensor according to the acoustic features, and use an encoder of the network architecture to generate latent features according to the acoustic features;

[0016] An adversarial training module, configured to adversarially train according to the latent features to minimize the loss function of the encoder and the loss function of a latent discriminator of the network architecture, and use the optimized encoder to generate optimized latent features according to the acoustic features;

[0017] An audio reconstruction module, configured to use the decoder of the network architecture to perform audio reconstruction based on the optimized latent features and the style embedding tensor to obtain a reconstructed audio, and output a test reconstructed audio according to the optimized latent features;

[0018] An optimization judgment module, configured to judge whether the encoder is optimized according to the reconstruction loss value between the test reconstructed audio and the reconstructed audio. If the encoder is not optimized, it jumps back to the feature extraction module to execute using the encoder of the network architecture to generate latent features according to the acoustic features. Or, if the encoder is optimized, it jumps to the audio conversion module to execute audio style conversion;

[0019] An audio conversion module, configured to perform audio style conversion on the audio to be processed based on the network architecture to obtain a converted audio.

[0020] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above audio style conversion method are implemented.

[0021] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above audio style conversion method are implemented.

[0022] In the solution implemented by the above audio style conversion method, device, computer device, and storage medium, acoustic features can be obtained by extracting features from a pre-acquired target audio, and a style embedding tensor can be generated based on the acoustic features by a reference encoder of a network architecture. A latent feature can be generated based on the acoustic features by an encoder of the network architecture. According to the latent feature, adversarial training is performed to minimize the loss function of the encoder and the loss function of a latent discriminator of the network architecture. An optimized latent feature is generated based on the acoustic features by the optimized encoder. An audio is reconstructed by a decoder of the network architecture based on the optimized latent feature and the style embedding tensor to obtain a reconstructed audio, and a test reconstructed audio is output based on the optimized latent feature. It is determined whether the encoder is optimized based on a reconstruction loss value between the test reconstructed audio and the reconstructed audio. If the encoder is not optimized, the step of generating a latent feature based on the acoustic features by the encoder of the network architecture is returned. If the encoder is optimized, audio style conversion is performed on an audio to be processed based on the network architecture to obtain a converted audio. In this solution, in view of the problem of audio style conversion, an audio conversion network architecture based on an encoder-decoder is proposed, and a fine-tuning solution based on a generative adversarial network is adopted. The generator consists of an encoder-decoder network, and the latent discriminator attempts to distinguish between the generated audio and the target audio. Through adversarial training, the generator can generate more realistic and natural audio, and the latent discriminator continuously improves its discrimination ability to optimize the performance of the encoder. After training is completed, the acoustic features of the target audio are input into the reference encoder, and the acoustic features of the source audio are input into the basic encoder. The output of the decoder is the converted audio sequence, realizing audio style conversion. This method can run without relying on parallel speech, transcription, or time alignment processes, adopts a global conditional mechanism, makes it vocabulary-independent, and can convert audio styles without being restricted by the target identity; it can perform one-time audio conversion without intermediate speech representations, thereby avoiding the need for speech alignment and speaker-independent automatic speech recognition networks. The efficiency and accuracy of audio style conversion are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a schematic diagram of an application environment of an audio style conversion method in an embodiment of the present invention;

[0025] Figure 2 It is a schematic flowchart of an audio style conversion method in an embodiment of the present invention;

[0026] Figure 3 It is a schematic structural diagram of an audio style conversion device in an embodiment of the present invention;

[0027] Figure 4 It is a schematic structural diagram of a computer device in an embodiment of the present invention;

[0028] Figure 5 It is another schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0030] The audio style conversion method provided by the embodiments of the present invention can be applied to an application environment such as Figure 1 . Among them, the client extracts features from the pre-acquired target audio to obtain acoustic features, and uses the reference encoder of the network architecture to generate a style embedding tensor according to the acoustic features, uses the encoder of the network architecture to generate latent features according to the acoustic features, and according to the latent features, performs adversarial training to minimize the loss function of the encoder and the loss function of the latent discriminator of the network architecture. Use the optimized encoder to generate optimized latent features according to the acoustic features, use the decoder of the network architecture to perform audio reconstruction according to the optimized latent features and the style embedding tensor to obtain a reconstructed audio, and output a test reconstructed audio according to the optimized latent features. Determine whether the encoder is optimized according to the reconstruction loss value between the test reconstructed audio and the reconstructed audio. If the encoder is not optimized, return to the step of using the encoder of the network architecture to generate latent features according to the acoustic features. If the encoder is optimized, perform audio style conversion on the audio to be processed based on the network architecture to obtain a converted audio. The efficiency of audio style conversion is improved. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail through specific embodiments below.

[0031] Please refer to Figure 2 as described Figure 2A flowchart of the audio style conversion method provided by an embodiment of the present invention includes the following steps:

[0032] S1. Extract features from the pre-acquired target audio to obtain acoustic features, and generate a style embedding tensor according to the acoustic features by using a reference encoder of a network architecture.

[0033] Example: In the field of medical and health, the target audio may be audio such as user heart sounds and lung sounds. The audio of the user's heart sounds can be obtained through devices such as electronic stethoscopes and synchronous heart sound and electrocardiogram detectors, and the lung sounds can be obtained through devices such as electronic stethoscopes, lung sound sensors, and wearable lung sound detectors.

[0034] In financial business, the target audio may be the voice input by a user to a voice customer service in a financial software, and the target audio can be obtained through intelligent terminal devices such as mobile phones and tablets.

[0035] In an embodiment of the present invention, the extraction of features from the pre-acquired target audio to obtain acoustic features refers to extracting the Mel frequency cepstral coefficients and Mel spectrogram of the target audio.

[0036] Specifically, the Mel frequency cepstral coefficients are a kind of characteristic parameters widely used in the fields of speech signal processing and speech processing. It is proposed based on the auditory characteristics of the human ear. The Mel frequency scale is a frequency scale that simulates the characteristic that the human ear has higher resolution in the low-frequency part and lower resolution in the high-frequency part.

[0037] Specifically, the Mel spectrogram is a two-dimensional visual representation of an audio. Its horizontal axis usually represents time, and the vertical axis represents Mel frequency. The gray value or color value of each pixel (or point) in the image represents the energy intensity (after logarithmic processing, etc.) corresponding to the time-Mel frequency point. It provides an intuitive way to observe the energy distribution of the audio signal in two dimensions of time and Mel frequency.

[0038] In an embodiment of the present invention, the extraction of features from the pre-acquired target audio to obtain acoustic features includes:

[0039] Perform pre-emphasis processing on the target audio to obtain preprocessed audio;

[0040] Perform frame segmentation processing on the preprocessed audio to obtain frame-segmented audio;

[0041] Perform window function processing on the frame-segmented audio to obtain an audio time-domain signal;

[0042] Perform fast Fourier transform on the audio time-domain signal to obtain a frequency-domain signal;

[0043] Perform a linear frequency transformation on the frequency-domain signal to obtain a Mel-frequency signal;

[0044] Perform a discrete pre-transformation on the Mel-frequency signal to obtain Mel-frequency cepstral coefficients;

[0045] Extract the Mel spectrogram of the target audio, and aggregate the Mel-frequency cepstral coefficients and the Mel spectrogram to obtain acoustic features.

[0046] Specifically, the pre-emphasis processing of the target audio is to enhance the energy of the high-frequency part of the target audio by using a first-order high-pass filter.

[0047] Specifically, the frame segmentation processing of the preprocessed audio refers to dividing the preprocessed audio into multiple frames according to a set frame length and frame shift.

[0048] Specifically, the window function processing of the framed audio refers to multiplying each frame signal in the framed audio by a window function, and a commonly used window function is the Hamming window.

[0049] Specifically, the fast Fourier transform of the audio time-domain signal is to convert the time-domain signal into a frequency-domain signal by using the fast Fourier transform algorithm.

[0050] In the embodiment of the present invention, the performing a linear frequency transformation on the frequency-domain signal to obtain a Mel-frequency signal includes:

[0051] Perform a linear frequency transformation on the frequency-domain signal by using the following formula:

[0052]

[0053] where Mel i represents the i-th Mel frequency in the Mel-frequency signal, and f i represents the i-th linear frequency in the frequency-domain signal.

[0054] In the embodiment of the present invention, the performing a discrete pre-transformation on the Mel-frequency signal to obtain Mel-frequency cepstral coefficients includes:

[0055] Perform a discrete pre-transformation on the Mel-frequency signal by using the following formula:

[0056]

[0057] where X k represents the k-th frequency-domain coefficient in the Mel-frequency cepstral coefficients, α k is a normalization factor, and when k is 0, α k takes when k is from 1 to N - 1, α kTake Mel i represents the i-th Mel frequency in the Mel frequency signal, and N is the length of the Mel frequency cepstral coefficients.

[0058] Specifically, the discrete pre-transformation of the Mel frequency signal to obtain Mel frequency cepstral coefficients includes:

[0059] Convert each sample in the Mel frequency signal into a frequency domain coefficient;

[0060] Multiply each sample by a preset cosine function based on the index corresponding to each frequency domain coefficient to obtain a product calculation result;

[0061] Sum the product calculation results, and multiply the sum result by a preset normalization factor to obtain Mel frequency cepstral coefficients.

[0062] In an embodiment of the present invention, the extraction of the Mel spectrogram of the target audio is performed by sampling, framing, and window function processing on the target audio, performing a short-time Fourier transform on the processing result of the window function to obtain a frequency domain signal, calculating the power spectrum of the frequency domain signal and then converting it to the Mel frequency scale, and using a logarithmic function for compression and then outputting and arranging it into a matrix to obtain the Mel spectrogram.

[0063] In an embodiment of the present invention, the pre-established network architecture is an encoder-decoder based architecture, which includes an encoder, a reference encoder, and a decoder.

[0064] Specifically, the encoder and the reference encoder adopt a combination design of a gated convolutional layer and a bidirectional LSTM (Long Short-Term Memory) network. Among them, the gated convolutional layer can capture the spectral relationship between the input acoustic feature sequences while retaining the time features. The bidirectional LSTM network is used to model the time characteristics of these acoustic sequences.

[0065] Specifically, the reference encoder has a similar design to the encoder. The difference is that a unidirectional LSTM network is used instead of the bidirectional LSTM, and a global average pooling layer is added to the unidirectional LSTM network to capture the global style features of the input audio while ignoring the local speech features. The global average pooling layer ensures that the learned style embedding is independent of local features (such as speech content).

[0066] Specifically, the LSTM network is a special type of recurrent neural network (RNN). It is mainly used to process and predict time series data and can effectively solve the problems of gradient disappearance and gradient explosion in traditional recurrent neural networks, so that it can learn and process long sequence data well.

[0067] In an embodiment of the present invention, the reference encoder of the network architecture is used to generate a style embedding tensor according to the acoustic features. This is achieved by pre-training the reference encoder for a classification task, and using the reference encoder to learn the mapping from the global style features of the input audio sequence to a fixed-length vector, thereby obtaining the style embedding tensor.

[0068] In an embodiment of the present invention, by extracting features from the pre-acquired target audio to obtain acoustic features, the efficiency of generating the style embedding tensor is improved. Then, the reference encoder of the network architecture is used to generate the style embedding tensor according to the acoustic features, which improves the efficiency of subsequent audio reconstruction.

[0069] S2. Use the encoder of the network architecture to generate latent features according to the acoustic features.

[0070] In an embodiment of the present invention, when using the encoder of the network architecture to generate latent features according to the acoustic features, the encoder is used to map the acoustic features to a latent space to obtain the latent features.

[0071] S3. According to the latent features, perform adversarial training to minimize the loss function of the encoder and the loss function of the latent discriminator of the network architecture, and use the optimized encoder to generate optimized latent features according to the acoustic features.

[0072] In an embodiment of the present invention, the step of performing adversarial training to minimize the loss function of the encoder and the loss function of the latent discriminator of the network architecture according to the latent features is an adversarial training scheme based on the latent discriminator. The goal of the latent discriminator is to predict the target category of the input audio based on its encoded representation, while the goal of the encoder is to deceive the latent discriminator by maximizing the negative log-likelihood of the target category. In this way, the latent discriminator tries to correctly predict the target category of the input audio, while the encoder tries to generate a representation that is irrelevant to the target category to deceive the latent discriminator. This adversarial training mechanism helps the encoder learn more discriminative feature representations and avoid overfitting to specific target categories.

[0073] Furthermore, during the training process, the latent discriminator and the encoder are alternately optimized. Specifically, first, fix the parameters of the encoder and update the parameters of the latent discriminator to minimize the loss function of the latent discriminator. Then, fix the parameters of the latent discriminator and update the parameters of the encoder to minimize the loss function of the encoder. This alternating training method enables the latent discriminator and the encoder to make progress together, and finally enables the encoder to generate a latent representation that is irrelevant to the target category.

[0074] Example illustration: In medical software, when a user needs to convert the voice style of a voice dialogue assistant, the converted audio needs to be independent of the original audio style. Therefore, it is necessary to train the encoder to generate the ability of latent representations independent of the target category.

[0075] In an embodiment of the present invention, when the latent discriminator and the encoder are alternately optimized, the gradient descent algorithm is used to update the parameters of the encoder and the latent discriminator.

[0076] Specifically, the gradient descent algorithm is an iterative method for optimizing the objective function, widely used in machine learning and deep learning to train model parameters to minimize the loss function. The gradient is a vector pointing in the direction of the fastest increase in the value of the objective function. For a multi-dimensional parameter space, each component of the gradient is the partial derivative of the objective function with respect to each parameter. The magnitude and direction of the gradient determine the direction and step size of parameter updates.

[0077] Specifically, the acoustic features are propagated forward through the encoder to obtain latent features. At the same time, the decoder reconstructs the input data according to the latent features and the style embedding tensor to obtain the reconstructed audio. After calculating the loss value between the reconstructed audio and the real audio, the encoder parameters are updated using the gradient descent algorithm along the negative gradient direction of the loss function.

[0078] Specifically, the forward propagation refers to the process in which data is passed from the input layer of the neural network through the hidden layer to the output layer. In this process, the input data is processed by weighted summation and activation functions in each layer, and finally the output result is obtained.

[0079] In an embodiment of the present invention, the adversarial training to minimize the loss function of the encoder and the loss function of the latent discriminator of the network architecture according to the latent features includes:

[0080] Using a preset latent discriminator to perform target category prediction based on the latent features to obtain a first predicted category;

[0081] Calculating a first loss value between the first predicted category and a preset real category using the loss function of the latent discriminator;

[0082] Keeping the parameters of the encoder unchanged, adjusting the parameters of the latent discriminator based on the first loss value to minimize the loss function of the latent discriminator to obtain an optimized latent discriminator; obtaining an optimized latent discriminator;

[0083] Using the optimized latent discriminator to perform target category prediction based on the latent features to obtain a second predicted category;

[0084] Calculate a second loss value between the second predicted category and the true category using the loss value function of the encoder;

[0085] Keep the latent discriminator parameters unchanged, and adjust the parameters of the encoder based on the second loss value to minimize the loss function of the encoder, obtaining an optimized encoder.

[0086] In the embodiments of the present invention, after the encoder is optimized, it can generate latent features independent of the target category, improving the feature representation ability.

[0087] In addition, the present invention also proposes a fine-tuning scheme based on adversarial training. Since the training process of adversarial training is unstable, gradient penalty can be combined for training.

[0088] Specifically, a decoder is used as the generator. For the latent discriminator, a network composed of two-dimensional convolutional layers is designed to distinguish between the true input audio feature sequence and the sequence generated by the model. The output of the latent discriminator network is a scalar representing the authenticity of the input feature sequence. The larger the scalar value, the more authentic. The training objective of the latent discriminator is to maximize the adversarial loss, while the objective of the generator is to simultaneously minimize the adversarial loss and the reconstruction loss to deceive the latent discriminator.

[0089] In the embodiments of the present invention, by adversarial training based on the latent features, the loss functions of the encoder and the latent discriminator of the network architecture are minimized, improving the feature representation ability of the encoder. Using the optimized encoder to generate optimized latent features based on the acoustic features improves the feature accuracy of the optimized latent features.

[0090] S4. Use the decoder of the network architecture to perform audio reconstruction based on the optimized latent features and the style embedding tensor to obtain a reconstructed audio, and output a test reconstructed audio based on the optimized latent features.

[0091] In the embodiments of the present invention, the use of the decoder of the network architecture to perform audio reconstruction based on the optimized latent features and the style embedding tensor to obtain a reconstructed audio is to use the decoder to output an audio sequence based on the optimized latent features and the style embedding tensor. This audio sequence contains the local speech features of the target audio and the global style features of the reconstructed audio.

[0092] Example: In financial business, the reconstructed audio can be the audio obtained after style conversion of the audio input by the user. By performing style conversion on the audio input by the user, the voice customer service of financial software can better understand the user's intention.

[0093] In an embodiment of the present invention, the step of outputting the test reconstructed audio according to the optimized latent features includes:

[0094] Using the latent discriminator to output a test target category according to the optimized latent features;

[0095] Using the reference encoder to output a test style embedding tensor according to the test target category;

[0096] Using the decoder to reconstruct the test audio according to the test style embedding tensor and the optimized latent features to obtain the test reconstructed audio.

[0097] In an embodiment of the present invention, using the reference encoder to output a test style embedding tensor according to the test target category is to test whether the correlation between the output accuracy of the encoder and the target category meets the requirements.

[0098] In an embodiment of the present invention, by using the decoder of the network architecture to perform audio reconstruction according to the optimized latent features and the style embedding tensor, the reconstructed audio is obtained, the efficiency of outputting the test reconstructed audio is improved, and the test reconstructed audio is output according to the optimized latent features, thereby improving the efficiency of subsequent determination of whether the encoder is optimized.

[0099] S5. Determine whether the encoder is optimized according to the reconstruction loss value between the test reconstructed audio and the reconstructed audio.

[0100] In an embodiment of the present invention, determining whether the encoder is optimized according to the reconstruction loss value between the test reconstructed audio and the reconstructed audio is to determine whether the encoder is optimized according to the magnitude relationship between the reconstruction loss value and a preset loss value threshold.

[0101] In an embodiment of the present invention, determining whether the encoder is optimized according to the reconstruction loss value between the test reconstructed audio and the reconstructed audio includes:

[0102] Calculating the mean absolute error between the test reconstructed audio and the reconstructed audio;

[0103] Calculating the Pearson correlation coefficient between the test reconstructed audio and the reconstructed audio;

[0104] Using the following formula to calculate the reconstruction loss value according to the mean absolute error and the Pearson correlation coefficient:

[0105] L recon =MAE+(1 - r)

[0106] where L recon is the reconstruction loss value, MAE is the mean absolute error, and r is the Pearson correlation coefficient;

[0107] If the reconstruction loss value is greater than a preset loss value threshold, it is confirmed that the encoder is not optimized yet;

[0108] If the reconstruction loss value is less than or equal to the loss value threshold, it is confirmed that the encoder is optimized.

[0109] In the embodiments of the present invention, the mean absolute error is a commonly used metric for measuring the performance of a prediction model, especially in regression problems. It represents the average of the absolute errors between the predicted values and the actual values.

[0110] In the embodiments of the present invention, the Pearson correlation coefficient is a statistical metric for measuring the degree of linear correlation between two variables. The value range of the Pearson correlation coefficient is from -1 to 1, where 1 represents perfect correlation.

[0111] If the encoder is not optimized yet, return to S2, and use the encoder of the network architecture to generate latent features according to the acoustic features.

[0112] In the embodiments of the present invention, when it is determined that the encoder is not optimized yet, it indicates that the optimization result of the encoder does not meet the requirements. Therefore, it is necessary to return to the step of using the encoder of the network architecture to generate latent features according to the acoustic features for the next round of encoder optimization.

[0113] If the encoder is optimized, execute S6, perform audio style conversion on the audio to be processed based on the network architecture to obtain the converted audio.

[0114] Example illustration: In the field of medical technology, the audio to be processed can be audio such as heart sounds and lung sounds of users obtained in real time by medical software in a medical diagnosis scenario. Or, in the field of financial technology, the audio to be processed can be audio input by users obtained in real time by financial software.

[0115] In the embodiments of the present invention, performing audio style conversion on the audio to be processed based on the network architecture to obtain the converted audio includes:

[0116] Use the optimized encoder in the network architecture to output the latent features of the audio to be processed according to the acoustic features of the audio to be processed;

[0117] Use the reference encoder in the network architecture to output the audio style embedding tensor of the audio to be processed according to the acoustic features of the audio to be processed;

[0118] Fuse the latent features of the audio to be processed and the audio style embedding tensor based on the channel dimension to obtain the fused features, and use the decoder to perform style conversion on the audio to be processed according to the fused features to obtain the converted audio.

[0119] As can be seen, in the above solution, based on the problem of audio style conversion, a fully differentiable encoder-decoder based audio conversion network architecture is proposed, and a fine-tuning scheme based on a generative adversarial network is adopted. The generator consists of an encoder-decoder network, while the latent discriminator attempts to distinguish between the generated audio and the target audio. Through adversarial training, the generator can generate more realistic and natural audio, and the latent discriminator continuously improves its discrimination ability to optimize the performance of the encoder. After training is completed, the acoustic features of the target audio are input into the reference encoder, and the acoustic features of the source audio are input into the base encoder. The output of the decoder is the converted audio sequence, realizing the style conversion of the audio. This method can operate without relying on parallel speech, transcription, or time alignment processes. It adopts a global conditioning mechanism, making it vocabulary-independent and capable of converting audio styles without being restricted by the target identity; it can perform one-shot audio conversion without intermediate speech representations, thus avoiding the need for speech alignment and speaker-independent automatic speech recognition networks.

[0120] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0121] In one embodiment, an audio style conversion device is provided, which corresponds one-to-one with the audio style conversion method in the above embodiment. As Figure 3 described, the audio style conversion device includes a feature extraction module 101, an adversarial training module 102, an audio reconstruction module 103, an optimization judgment module 104, and an audio conversion module 105. The detailed description of each functional module is as follows:

[0122] The feature extraction module 101 is configured to extract features from the pre-acquired target audio to obtain acoustic features, generate a style embedding tensor according to the acoustic features by using the reference encoder of the network architecture, and generate latent features according to the acoustic features by using the encoder of the network architecture;

[0123] The adversarial training module 102 is configured to, according to the latent features, perform adversarial training to minimize the loss function of the encoder and the loss function of the latent discriminator of the network architecture, and generate optimized latent features according to the acoustic features by using the optimized encoder;

[0124] The audio reconstruction module 103 is configured to perform audio reconstruction according to the optimized latent features and the style embedding tensor by using the decoder of the network architecture to obtain the reconstructed audio, and output the test reconstructed audio according to the optimized latent features;

[0125] The optimization judgment module 104 is used to judge whether the encoder is optimized according to the reconstruction loss value between the test reconstructed audio and the reconstructed audio. If the encoder is not optimized, it jumps back to the feature extraction module to execute generating the latent feature according to the acoustic feature by the encoder of the network architecture. Or, if the encoder is optimized, it jumps to the audio conversion module to execute audio style conversion;

[0126] The audio conversion module 105 is used to perform audio style conversion on the audio to be processed based on the network architecture to obtain the converted audio.

[0127] In one embodiment, the feature extraction module 101 is specifically used for:

[0128] Perform pre-emphasis processing on the target audio to obtain the preprocessed audio;

[0129] Perform frame division processing on the preprocessed audio to obtain the framed audio;

[0130] Perform window function processing on the framed audio to obtain the audio time-domain signal;

[0131] Perform fast Fourier transform on the audio time-domain signal to obtain the frequency-domain signal;

[0132] Perform linear frequency transformation on the frequency-domain signal to obtain the Mel frequency signal;

[0133] Perform discrete pre-transformation on the Mel frequency signal to obtain the Mel frequency cepstral coefficients;

[0134] Extract the Mel spectrogram of the target audio, and summarize the Mel frequency cepstral coefficients and the Mel spectrogram to obtain the acoustic feature.

[0135] In one embodiment, the feature extraction module 101 is specifically used for:

[0136] Convert each sample in the Mel frequency signal into a frequency-domain coefficient;

[0137] Multiply each sample by a preset cosine function based on the index corresponding to each frequency-domain coefficient to obtain the product calculation result;

[0138] Sum the product calculation results, and multiply the sum result by a preset normalization factor to obtain the Mel frequency cepstral coefficients.

[0139] In one embodiment, the adversarial training module 102 is specifically used for:

[0140] Use a preset latent discriminator to perform target class prediction according to the latent feature to obtain the first predicted class;

[0141] Calculate a first loss value of the first predicted category and a preset true category by using the loss function of the latent discriminator;

[0142] Keep the parameters of the encoder unchanged, and adjust the parameters of the latent discriminator based on the first loss value to minimize the loss function of the latent discriminator, so as to obtain an optimized latent discriminator; obtain an optimized latent discriminator;

[0143] Use the optimized latent discriminator to perform target category prediction according to the latent features to obtain a second predicted category;

[0144] Calculate a second loss value of the second predicted category and the true category by using the loss value function of the encoder;

[0145] Keep the parameters of the latent discriminator unchanged, and adjust the parameters of the encoder based on the second loss value to minimize the loss function of the encoder, so as to obtain an optimized encoder.

[0146] In one embodiment, the audio reconstruction module 103 is specifically configured to:

[0147] Use the latent discriminator to output a test target category according to the optimized latent features;

[0148] Use the reference encoder to output a test style embedding tensor according to the test target category;

[0149] Use the decoder to reconstruct a test audio according to the test style embedding tensor and the optimized latent features to obtain a test reconstructed audio.

[0150] In one embodiment, the optimization judgment module 104 is specifically configured to:

[0151] Calculate the mean absolute error between the test reconstructed audio and the reconstructed audio;

[0152] Calculate the Pearson correlation coefficient between the test reconstructed audio and the reconstructed audio;

[0153] Calculate the reconstruction loss value according to the mean absolute error and the Pearson correlation coefficient;

[0154] If the reconstruction loss value is greater than a preset loss value threshold, it is confirmed that the encoder is not optimized yet;

[0155] If the reconstruction loss value is less than or equal to the loss value threshold, it is confirmed that the encoder is optimized.

[0156] In one embodiment, the audio conversion module 105 is specifically configured to:

[0157] Output the potential features of the audio to be processed according to the acoustic features of the audio to be processed by using the optimized encoder in the network architecture;

[0158] Output the style embedding tensor of the audio to be processed according to the acoustic features of the audio to be processed by using the reference encoder in the network architecture;

[0159] Fuse the potential features of the audio to be processed and the style embedding tensor of the audio to be processed based on the channel dimension to obtain a fused feature, and use the decoder to perform style conversion on the audio to be processed according to the fused feature to obtain a converted audio.

[0160] The present invention provides an audio style conversion device. Based on the problem of audio style conversion, a fully differentiable encoder-decoder based audio conversion network architecture is proposed, and a fine-tuning scheme based on a generative adversarial network is adopted. The generator consists of an encoder-decoder network, and the latent discriminator attempts to distinguish the generated audio from the target audio. Through adversarial training, the generator can generate more realistic and natural audio, and the latent discriminator continuously improves its discrimination ability to optimize the performance of the encoder. After training is completed, the acoustic features of the target audio are input into the reference encoder, and at the same time, the acoustic features of the source audio are input into the base encoder. The output of the decoder is the converted audio sequence, realizing the style conversion of the audio. This method can run without relying on parallel speech, transcription or time alignment processes, adopts a global conditional mechanism, makes it vocabulary-independent, and can convert the audio style without being restricted by the target identity; it can perform one-time audio conversion without intermediate speech representation, thus avoiding the need for speech alignment and speaker-independent automatic speech recognition networks.

[0161] For the specific limitations of the audio style conversion device, reference can be made to the limitations of the method for audio style conversion in the above text, which will not be elaborated here. Each module in the above audio style conversion device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of the processor, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0162] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4As described above. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an audio style conversion method.

[0163] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5 As described above. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of an audio style conversion method

[0164] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0165] Extract features from the pre-acquired target audio to obtain acoustic features, and use the reference encoder of the network architecture to generate a style embedding tensor according to the acoustic features;

[0166] Use the encoder of the network architecture to generate latent features according to the acoustic features;

[0167] According to the latent features, perform adversarial training to minimize the loss function of the encoder and the loss function of the latent discriminator of the network architecture, and use the optimized encoder to generate optimized latent features according to the acoustic features;

[0168] Use the decoder of the network architecture to perform audio reconstruction according to the optimized latent features and the style embedding tensor to obtain a reconstructed audio, and output a test reconstructed audio according to the optimized latent features;

[0169] Judge whether the encoder is optimized according to the reconstruction loss value between the test reconstructed audio and the reconstructed audio;

[0170] If the encoder is not optimized yet, return to the step of generating latent features from the acoustic features using the encoder of the network architecture;

[0171] If the encoder is optimized, perform audio style conversion on the audio to be processed based on the network architecture to obtain a converted audio.

[0172] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0173] Extract features from a pre-acquired target audio to obtain acoustic features, and generate a style embedding tensor from the acoustic features using a reference encoder of a network architecture;

[0174] Generate latent features from the acoustic features using the encoder of the network architecture;

[0175] According to the latent features, perform adversarial training to minimize the loss function of the encoder and the loss function of a latent discriminator of the network architecture, and use the optimized encoder to generate optimized latent features from the acoustic features;

[0176] Use a decoder of the network architecture to perform audio reconstruction based on the optimized latent features and the style embedding tensor to obtain a reconstructed audio, and output a test reconstructed audio according to the optimized latent features;

[0177] Judge whether the encoder is optimized according to the reconstruction loss value between the test reconstructed audio and the reconstructed audio;

[0178] If the encoder is not optimized yet, return to the step of generating latent features from the acoustic features using the encoder of the network architecture;

[0179] If the encoder is optimized, perform audio style conversion on the audio to be processed based on the network architecture to obtain a converted audio.

[0180] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can implement, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0181] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0182] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0183] Finally, it should be noted that if non-company software tools or components appear in the application embodiments, they are only used for illustrative introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

[0184] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. An audio style transfer method, characterized in that: include: Extract features from the pre-acquired target audio to obtain acoustic features, and generate a style embedding tensor based on the acoustic features using a reference encoder of the network architecture; generating latent features according to the acoustic features using an encoder of the network architecture; According to the latent features, adversarial training minimizes the loss function of the encoder and the loss function of the latent discriminator of the network architecture, and uses the optimized encoder to generate optimized latent features according to the acoustic features; Using the decoder of the network architecture to perform audio reconstruction according to the optimized latent features and the style embedding tensor to obtain reconstructed audio, and outputting a test reconstructed audio according to the optimized latent features; Determining whether the encoder is optimized according to the reconstruction loss value of the test reconstructed audio and the reconstructed audio; If the encoder is not optimized, returning to the step of generating potential features according to the acoustic features using the encoder of the network architecture; If the encoder optimization is completed, audio style conversion is performed on the audio to be processed based on the network architecture to obtain converted audio.

2. The audio style transfer method according to claim 1, characterized in that The step of extracting features from the pre-acquired target audio to obtain acoustic features includes: Performing pre-emphasis processing on the target audio to obtain pre-processed audio; Performing frame processing on the preprocessed audio to obtain framed audio; Performing window function processing on the framed audio to obtain an audio time domain signal; Performing a fast Fourier transform on the audio time domain signal to obtain a frequency domain signal; Performing a linear frequency transform on the frequency domain signal to obtain a Mel frequency signal; Performing a discrete pre-transformation on the Mel-frequency signal to obtain Mel-frequency cepstrum coefficients; The Mel-frequency spectrogram of the target audio is extracted, and the Mel-frequency cepstral coefficients and the Mel-frequency spectrogram are summarized to obtain acoustic features.

3. The audio style transfer method according to claim 2, characterized in that The step of performing a discrete pre-transformation on the Mel-frequency signal to obtain Mel-frequency cepstrum coefficients includes: Convert each sample in the Mel frequency signal into a frequency domain coefficient; Multiply each sample by a preset cosine function based on the index corresponding to each frequency domain coefficient to obtain a product calculation result; The product calculation results are summed, and the summed result is multiplied by a preset normalization factor to obtain Mel-frequency cepstrum coefficients.

4. The audio style transfer method according to claim 1, characterized in that The adversarial training minimizes the loss function of the encoder and the loss function of the potential discriminator of the network architecture according to the potential features, including: Using a preset latent discriminator to predict the target category according to the latent features, to obtain a first predicted category; Calculating a first loss value between the first predicted category and a preset true category using the loss function of the potential discriminator; Keeping the parameters of the encoder unchanged, adjusting the parameters of the potential discriminator based on the first loss value to minimize the loss function of the potential discriminator, and obtaining an optimized potential discriminator; obtaining an optimized potential discriminator; Using the optimized latent discriminator to predict the target category according to the latent features, a second predicted category is obtained; Calculating a second loss value between the second predicted category and the true category using the loss value function of the encoder; The parameters of the potential discriminator are kept unchanged, and the parameters of the encoder are adjusted based on the second loss value to minimize the loss function of the encoder to obtain an optimized encoder.

5. The audio style transfer method according to claim 1, characterized in that The step of outputting a test reconstructed audio according to the optimized potential feature comprises: Outputting a test target category according to the optimized latent features using the latent discriminator; Outputting a test style embedding tensor according to the test target category using the reference encoder; The decoder is used to reconstruct the test audio according to the test style embedding tensor and the optimized latent features to obtain the test reconstructed audio.

6. The audio style transfer method according to claim 1, characterized in that: The step of judging whether the encoder is optimized according to the reconstruction loss value of the test reconstructed audio and the reconstructed audio includes: Calculating a mean absolute error between the test reconstructed audio and the reconstructed audio; Calculating the Pearson correlation coefficient between the test reconstructed audio and the reconstructed audio; Calculate the reconstruction loss value according to the mean absolute error and the Pearson correlation coefficient; If the reconstruction loss value is greater than a preset loss value threshold, it is confirmed that the encoder is not optimized; If the reconstruction loss value is less than or equal to the loss value threshold, it is confirmed that the encoder optimization is completed.

7. The audio style transfer method according to claim 1, characterized in that: The performing audio style conversion on the processed audio based on the network architecture to obtain converted audio includes: Outputting potential features of the audio to be processed according to the acoustic features of the audio to be processed using the optimized encoder in the network architecture; Outputting a style embedding tensor of the audio to be processed according to the acoustic features of the audio to be processed using a reference encoder in the network architecture; Based on the channel dimension, the potential features of the audio to be processed are fused with the style embedding tensor of the audio to be processed to obtain a fused feature, and the decoder is used to perform style conversion on the audio to be processed according to the fused feature to obtain a converted audio.

8. An audio style conversion device, characterized in that: include: A feature extraction module, used to extract features from the pre-acquired target audio to obtain acoustic features, generate a style embedding tensor based on the acoustic features using a reference encoder of the network architecture, and generate a latent feature based on the acoustic features using an encoder of the network architecture; An adversarial training module, configured to perform adversarial training to minimize the loss function of the encoder and the loss function of the potential discriminator of the network architecture according to the potential features, and to generate optimized potential features according to the acoustic features using the optimized encoder; An audio reconstruction module, configured to use a decoder of the network architecture to perform audio reconstruction according to the optimized latent features and the style embedding tensor to obtain reconstructed audio, and output a test reconstructed audio according to the optimized latent features; An optimization judgment module, used for judging whether the encoder is optimized according to the reconstruction loss value of the test reconstructed audio and the reconstructed audio, and if the encoder is not optimized, jumping back to the feature extraction module to execute the encoder of the network architecture to generate potential features according to the acoustic features, or, if the encoder is optimized, jumping to the audio conversion module to execute audio style conversion; The audio conversion module is used to perform audio style conversion on the audio to be processed based on the network architecture to obtain converted audio.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the audio style transfer method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the audio style transfer method according to any one of claims 1 to 7 are implemented.