Audio style conversion method and device based on diffusion model, equipment and medium

Through the audio style conversion method based on the diffusion model, the problem of style conversion deviation in cross-domain adaptation in the existing technology is solved, and high-quality speech conversion in the financial and medical fields is achieved, meeting the needs of professional accuracy and emotional adaptation.

CN120708637APending Publication Date: 2025-09-26PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511055908.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing speech style conversion technologies find it difficult to balance retaining the key pathological features of the original speech and the conversion style in cross-domain adaptation, resulting in a reduced user experience in applications in the medical and financial fields.

Method used

An audio style transfer method based on a diffusion model is adopted to generate target converted audio through feature extraction, spectrum generation, graph conversion and fundamental frequency adjustment, ensuring that the style transfer does not deviate from the core of the original melody.

Benefits of technology

It achieves professional accuracy and emotional adaptation of customer service voice in the financial field, and retains key pathological tone features of patient voice in medical scenarios, improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708637A_ABST
    Figure CN120708637A_ABST
Patent Text Reader

Abstract

The invention discloses an audio style conversion method and device based on a diffusion model, equipment and a medium. The method comprises the following steps: performing feature extraction on a received audio to be converted and a target style audio to obtain an initial spectrogram and an initial fundamental frequency of the audio to be converted and a target style feature of the target style audio; performing spectrum generation through a preset diffusion model according to the initial spectrogram, the initial fundamental frequency and the target style characteristics, and obtaining a target spectrogram with the target style characteristics; performing spectrum conversion on the target spectrogram through a preset audio vocoder to generate a preliminary audio waveform; and performing fundamental frequency adjustment on the initial audio waveform and the initial fundamental frequency through a preset filtering vocoder to generate a target conversion audio. The method and the device can be applied to audio style conversion of businesses such as financial insurance and medical health care, so as to solve the problem that the audio style cannot be accurately converted in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing and can be applied to fields such as finance, insurance, and healthcare. In particular, it relates to a method, device, equipment, and medium for audio style conversion based on a diffusion model. Background Art

[0002] With the rapid development of science and technology, the application of speech processing technology in the medical and financial fields is becoming more and more extensive. From the communication of medical conditions in telemedicine to the intelligent voice interaction of financial customer service, all rely on accurate speech feature analysis and style adaptation. However, the speech style conversion technology in the existing technology mostly focuses on the imitation of surface features, and has significant limitations in cross-domain adaptation. For example, in medical scenarios, when conducting personalized singing rehabilitation training for patients with speech disorders, the existing technology is difficult to balance the preservation of the key pathological features of the original voice and the conversion style; in the financial field, the converted customer service voice often loses the original emotional warmth or professional characteristics, and either appears mechanical and stereotyped, reducing the user experience, or has a mixed style. This has led to the inability of existing technologies to accurately convert audio styles, restricting its application in fields with high precision requirements. Summary of the Invention

[0003] The embodiments of the present invention provide a method, apparatus, device, and medium for audio style conversion based on a diffusion model, aiming to solve the problem that the existing technology cannot accurately convert audio styles.

[0004] In a first aspect, an embodiment of the present invention provides an audio style conversion method based on a diffusion model, which includes: extracting features from received audio to be converted and target style audio to obtain an initial spectrum graph and an initial fundamental frequency of the audio to be converted, as well as target style features of the target style audio; generating a spectrum through a preset diffusion model based on the initial spectrum graph, the initial fundamental frequency and the target style features to obtain a target spectrum graph with the target style features; performing spectrum conversion on the target spectrum graph through a preset audio vocoder to generate a preliminary audio waveform; and adjusting the fundamental frequency of the preliminary audio waveform and the initial fundamental frequency through a preset filtering vocoder to generate target converted audio.

[0005] In the second aspect, an embodiment of the present invention also provides an audio style conversion device based on a diffusion model, which includes: an extraction unit, used to extract features from the received audio to be converted and the target style audio, and obtain the initial spectrum graph and initial fundamental frequency of the audio to be converted, as well as the target style characteristics of the target style audio; an acquisition unit, used to generate a spectrum through a preset diffusion model according to the initial spectrum graph, the initial fundamental frequency and the target style characteristics, and obtain a target spectrum graph with the target style characteristics; a conversion unit, used to perform spectrum conversion on the target spectrum graph through a preset audio vocoder to generate a preliminary audio waveform; a generation unit, used to perform fundamental frequency adjustment on the preliminary audio waveform and the initial fundamental frequency through a preset filter vocoder to generate the target converted audio.

[0006] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.

[0007] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the above method can be implemented.

[0008] Embodiments of the present invention provide a method, apparatus, device, and medium for audio style conversion based on a diffusion model. The method comprises: extracting features from received audio to be converted and target style audio to obtain an initial spectrogram and initial fundamental frequency of the audio to be converted, as well as target style features of the target style audio; generating a spectrum using a preset diffusion model based on the initial spectrogram, the initial fundamental frequency, and the target style features to obtain a target spectrogram having the target style features; performing spectrum conversion on the target spectrogram using a preset audio vocoder to generate a preliminary audio waveform; and performing fundamental frequency adjustment on the preliminary audio waveform and the initial fundamental frequency using a preset filter vocoder to generate target converted audio. The embodiment of the present invention extracts features from the received audio to be converted and the target style audio to obtain an initial spectrogram, an initial fundamental frequency and target style features, thereby providing a clear voice benchmark and style target for subsequent conversion, avoiding conversion deviation caused by feature mixing, and then using a preset diffusion model to generate a target spectrogram containing target style features to ensure that the style migration does not deviate from the original melody core. Finally, a preset audio vocoder and a preset filter vocoder are used to convert and adjust the fundamental frequency to generate the target converted audio, fundamentally avoiding the pitch deviation caused by style conversion, thereby achieving high-quality audio conversion, so that it can not only meet the requirements of professional accuracy and style adaptation of financial customer service voice, but also support the retention of key pathological tone features when standardizing patient voice recording in medical scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0010] Figure 1 A schematic diagram of a flow chart of an audio style conversion method based on a diffusion model provided by an embodiment of the present invention;

[0011] Figure 2 A schematic diagram of a sub-process of an audio style conversion method based on a diffusion model provided by an embodiment of the present invention;

[0012] Figure 3 A schematic diagram of a sub-process of an audio style conversion method based on a diffusion model provided by an embodiment of the present invention;

[0013] Figure 4 A schematic diagram of a sub-process of an audio style conversion method based on a diffusion model provided by an embodiment of the present invention;

[0014] Figure 5 A schematic diagram of a sub-process of an audio style conversion method based on a diffusion model provided by an embodiment of the present invention;

[0015] Figure 6 A schematic diagram of a sub-process of an audio style conversion method based on a diffusion model provided by an embodiment of the present invention;

[0016] Figure 7 A schematic diagram of a sub-process of an audio style conversion method based on a diffusion model provided by an embodiment of the present invention;

[0017] Figure 8 A schematic block diagram of an audio style conversion apparatus based on a diffusion model provided by an embodiment of the present invention;

[0018] Figure 9 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0020] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0021] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0022] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0023] See also Figure 1 , Figure 1 This figure illustrates a flow chart of a diffusion-based audio style conversion method according to an embodiment of the present invention. The diffusion-based audio style conversion method in this embodiment can be applied in the financial services and healthcare sectors, for example, in voice interaction optimization for intelligent financial customer service and voice diagnosis assistance for telemedicine. This method accurately and precisely converts audio style while balancing information accuracy, emotional integrity, and scenario-specific flexibility, providing more reliable technical support for cross-domain voice interaction.

[0024] Figure 1 FIG. 1 is a flow chart of an audio style conversion method based on a diffusion model provided by an embodiment of the present invention. As shown in the figure, the method includes the following steps S110-S140.

[0025] S110 , performing feature extraction on the received audio to be converted and the target style audio to obtain an initial spectrogram and an initial fundamental frequency of the audio to be converted, and target style features of the target style audio.

[0026] In this embodiment, the audio is a user-input sound clip containing rich information, such as a singing clip, a voice broadcast of financial news in the financial field, or remote voice diagnosis instructions in the medical field. The spectrogram is a tool for visualizing audio signals in time and frequency. In this embodiment, it is a Mel spectrogram, which converts the audio signal into a two-dimensional image. It displays the intensity of each frequency component of the audio signal at different time points. The fundamental frequency refers to the lowest frequency component in the audio signal, which determines the pitch of the sound. The target style features are the unique style characteristics of the target style audio, such as timbre, rhythm, and prosody. Feature extraction is performed on the received audio to be converted and the target style audio. Specifically, the audio signal of the audio to be converted can be converted into a frequency domain spectrogram using algorithms such as the short-time Fourier transform (STFT). The initial fundamental frequency can be extracted using the autocorrelation method. The rhythm, prosody, timbre, and other features of the target style audio can then be extracted using a pre-trained deep learning model (such as an LSTM-based model) and represented as a feature vector, thereby obtaining the target style features. By extracting representative features from the audio to be converted and the target style audio, basic data is provided for subsequent audio style conversion.

[0027] In one embodiment, if Figure 2 As shown, the step S110 also includes steps S111-S112.

[0028] S111, dividing the target style audio into frames according to a preset time interval and extracting its spectral features;

[0029] S112: extracting the spectral features through a preset style encoder to obtain the target style features.

[0030] In this embodiment, the preset interval determines the time interval between frames. This interval can be set according to the audio sampling rate, analysis accuracy requirements, etc., and is not limited to this. The spectral feature is a way to describe the characteristics of the audio signal in the frequency domain. The target style audio is framed according to the preset interval and its spectral features are extracted. Specifically, if it is a financial news broadcast audio in the financial field, it can be framed according to the preset interval (such as 20 milliseconds). For example, if the total length of the audio is 1 minute and the sampling rate is 44.1kHz, then after framing, approximately 3000 frames of audio data will be obtained. Each frame of audio signal is converted into a frequency domain representation using a Fourier transform algorithm (such as a fast Fourier transform, FFT) to obtain the spectrum of the frame. Then, various spectral features can be extracted from the spectrum, such as Mel-frequency cepstral coefficients (MFCCs) and linear prediction coding coefficients (LPCCs). These spectral features can effectively capture the frequency characteristics of the audio signal and provide a basis for subsequent style encoding. The spectral features are extracted through a preset style encoder to obtain the target style features. Specifically, the preset style encoder is a deep learning model trained based on the audio filling task. The trained model distills multi-level target style features from the spectral features, wherein the target style features include: surface features: spectral texture (such as the proportion of high-frequency energy); middle-level features: breath rhythm (such as ventilation intervals, breath intensity changes); deep-level features: emotional tendencies (such as emotional tension corresponding to the amplitude of spectral fluctuations). The target style features are obtained by inputting the spectral features into the preset style encoder. This is to facilitate obtaining the target style features of the target style audio (such as slow and clear medical teaching speech).

[0031] S120 , generating a spectrum using a preset diffusion model according to the initial spectrum graph, the initial fundamental frequency, and the target style feature, to obtain a target spectrum graph having the target style feature.

[0032] In this embodiment, the preset diffusion model is a trained generative model. In this embodiment, the diffusion model is a diffusion model based on the DiT architecture. A spectrum is generated based on the initial spectrogram, the initial fundamental frequency, and the target style features using a preset diffusion model. Specifically, the initial spectrogram, the initial fundamental frequency, and the target style features may be dimensionally aligned, and key associations may be strengthened using an attention mechanism (e.g., binding the style features to the accent positions in the spectrogram) to obtain the fused features. The preset diffusion model uses random noise as a starting point and combines the fused features to perform multiple rounds of denoising iterations (e.g., 100-200 steps). In each round, the initial fundamental frequency is referenced to constrain the pitch direction (to avoid melodic deviation), and the target style features are injected to adjust spectral details (e.g., optimizing the low-frequency energy distribution based on the breath characteristics of the target style). After the iterations, a spectrogram is generated that retains the original content structure (e.g., keywords and pitch framework) while also possessing the target style characteristics (e.g., tone and rhythm). For example, if applied to financial news broadcast audio in the financial field, the generated target spectrogram may retain the pitch information corresponding to the initial fundamental frequency of the original audio, ensuring the comprehensibility of the broadcast content, while also incorporating a light and lively style feature, making the generated financial news broadcast audio more attractive and interesting. Spectrum generation is performed through a preset diffusion model to obtain a spectrum that retains the basic characteristics of the initial spectrum and incorporates the target style features, providing a solid data foundation for subsequent audio conversion.

[0033] In one embodiment, if Figure 3 As shown, the step S120 also includes steps S1201-S1202 before the step S120.

[0034] S1201: Collect training audio, and perform preset audio filling training and preset loop training on the initial diffusion model according to the training spectrogram of the training audio;

[0035] S1201: Determine the preset diffusion model according to the training result of the initial diffusion model.

[0036] In this embodiment, the training audio serves as the fundamental data for constructing and training the diffusion model. Its quality, diversity, and representativeness directly impact the model's learning and generalization capabilities. Specifically, the training audio can be collected from a variety of sources, such as professional audio databases. Based on the training spectrogram of the training audio, the initial diffusion model undergoes preset audio filling training and preset loop training. Specifically, the training audio is framed according to a preset frame length and frame shift, where the frame length and frame shift can be set based on the specific application scenario and are not limited. A short-time Fourier transform is performed on each frame of the audio signal to obtain the spectrum of that frame. The spectra of all frames are then concatenated in chronological order to obtain a complete training spectrogram. The preset audio filling training allows the diffusion model to learn how to fill in the complete audio spectrum based on partial audio information or noise. The training spectrogram can be randomly masked, i.e., a region is randomly selected and set to zero or noise is added to simulate incomplete input. The masked spectrogram is used as input, and the original spectrogram is used as the target. The diffusion model is trained to generate a complete spectrogram based on the incomplete input spectrogram for training. The preset cyclic training is a training that enables the diffusion model to better learn the long-term dependencies and style features of the audio spectrum. The training spectrum graph can be used as input, and multiple forward diffusion and reverse diffusion cyclic trainings can be performed through the diffusion model. The preset diffusion model is determined based on the training results of the initial diffusion model. Specifically, the evaluation indicators can be used to measure the performance of the model during the training process and the quality of the generated spectrum graph, wherein the evaluation indicators include: spectrum distortion, auditory quality, and style consistency. According to the results of the evaluation indicators, the model with the best performance in the training process is selected as the preset diffusion model. If the performance of the model does not meet the requirements, the structure, parameters or training methods of the model can be adjusted, and then retrained and evaluated until a satisfactory model is obtained. By training the initial diffusion model, the purity of the converted speech can be improved while avoiding pitch deviation and melody distortion, so as to facilitate subsequent accurate audio conversion.

[0037] In one embodiment, if Figure 4 As shown, the step S1201 also includes steps S12011-S12012.

[0038] S12011. Randomly masking a training spectrogram of a target training audio according to a preset masking range, wherein the training audio includes the target training audio;

[0039] S12012: Control the initial diffusion model to predict the spectrum characteristics of the masked area based on the unmasked area in the training spectrum graph to perform the preset audio filling training.

[0040] In this embodiment, the target training audio refers to a representative training audio that can cover the audio features and styles to be learned. The preset masking range refers to the size of the area to be masked in the training spectrogram. For example, the masking range can be set to 20%, that is, 20% of the audio is randomly masked. Among them, the masking range can be adjusted according to the specific task requirements and is not limited to this. The training spectrogram of the target training audio is randomly masked according to the preset masking range. Specifically, it is assumed that the training spectrogram is a two-dimensional matrix, in which the rows represent frequency and the columns represent time. If the preset masking range is 30% in the time dimension and 20% in the frequency dimension, the starting column and the ending column of a time period, as well as the starting row and the ending row of a frequency segment are randomly selected, and the matrix elements in these areas are set to zero. The initial diffusion model is controlled to predict the spectral characteristics of the masked area based on the unmasked area in the training spectrogram. Specifically, the training spectrogram after random masking is used as the input of the initial diffusion model. At this point, the model can see the unmasked areas in the spectrogram, while the information in the masked areas is hidden. The model's encoder extracts features from the unmasked areas, and the generator predicts the spectral features of the masked areas based on these features. The goal of preset audio filling training is to make the spectral features of the masked areas predicted by the model as close as possible to the true spectral features of the corresponding areas in the original spectrogram. A loss function can be defined to measure the difference between the predicted results and the true results, such as the mean squared error loss function. During the training process, the loss function is minimized by continuously adjusting the model's parameters, thereby improving the model's predictive ability. The initial diffusion model is trained on randomly masked training spectrograms to enhance the model's generalization ability without relying on parallel data.

[0041] In one embodiment, if Figure 5 As shown, the step S1201 also includes steps S12013-S12015.

[0042] S12013, generating a to-be-restored spectrogram by combining the training source speech and the target training audio with the initial diffusion model, wherein the training audio includes the training source speech, and the to-be-restored spectrogram has the style characteristics of the target training audio;

[0043] S12014, performing spectrum restoration on the spectrum graph to be restored using the initial diffusion model to generate a restored spectrum graph;

[0044] S12015: Reversely update the model parameters of the initial diffusion model according to the loss between the spectrum to be restored and the restored spectrum, so as to perform the preset cyclic training.

[0045] In this embodiment, the training source speech is the source speech used for training. For example, audio of doctors' quick and concise medical instructions and slow and clear medical teaching speech collected from channels such as hospitals and medical equipment control systems, or audio of serious and formal financial news broadcasts and lively and lively financial talk shows collected from channels such as financial news websites and television financial programs. The training source speech and the target training audio are combined through the initial diffusion model to generate a to-be-restored spectrogram. Specifically, the spectrogram of the training source speech and the spectrogram of the target training audio are used as inputs to the initial diffusion model. The model's task is to perform style conversion on the spectrogram of the training source speech based on the stylistic features of the target training audio to generate a to-be-restored spectrogram. The generated to-be-restored spectrogram is then used as input to the initial diffusion model for spectrogram restoration. The model's goal is to restore the to-be-restored spectrogram back to the original training source speech spectrogram. The diffusion model generator gradually restores the to-be-restored spectrogram through a reverse process based on the to-be-restored spectrogram. During this process, the model utilizes the distribution patterns of the audio features it has learned to attempt to remove the influence of the target stylistic features and recover the original training source speech information. The model parameters of the initial diffusion model are updated in reverse according to the loss between the spectrum graph to be restored and the restored spectrum graph. Specifically, a loss function is defined to measure the difference between the spectrum graph to be restored and the restored spectrum graph, for example, a mean square error loss function, a structural similarity index (SSIM) loss function, etc. According to the calculated loss value, the gradient of the model parameters is calculated using a backpropagation algorithm, and the model parameters of the initial diffusion model are updated using an optimization algorithm (such as a stochastic gradient descent algorithm, an Adam optimization algorithm, etc.) to minimize the loss function, thereby improving the performance of the model. In this case, it is necessary to repeat the above-mentioned process of generating the spectrum graph to be restored, spectrum restoration, and parameter updating, and perform preset loop training. In each loop, the model will continue to learn and adjust, gradually improving the accuracy of style conversion and restoration. The end condition of the loop training can be determined by setting the number of loops or observing the performance of the model on the validation set. By using the initial diffusion model to realize the conversion of the training source speech to the target training audio style features, the model performance is continuously optimized, thereby avoiding pitch deviation and melody distortion.

[0046] S130 , performing spectrum conversion on the target spectrum through a preset audio vocoder to generate a preliminary audio waveform.

[0047] In this embodiment, the audio waveform is a visual representation of the audio signal in the time domain, which depicts the curve of the amplitude change of the audio signal over time. The audio vocoder is an algorithm that converts spectral information into a time-domain audio waveform. In this embodiment, the preset audio vocoder is a HiFiGAN vocoder, which can convert the spectral characteristics of the audio into a high-quality audible speech waveform. The target spectrogram is subjected to a spectrum conversion by a preset audio vocoder. Specifically, the HiFiGAN vocoder performs spectrum conversion based on the frequency and energy information in the target spectrogram through a generative adversarial network (GAN) architecture to generate a series of sine waves, and these sine waves are superimposed and synthesized to generate a preliminary audio waveform. The preliminary audio waveform is generated by performing spectrum conversion according to the target spectrogram, so that the audio to be converted can be further converted to facilitate the subsequent synthesis of the final audio.

[0048] In one embodiment, if Figure 6 As shown, the step S130 also includes steps S131-S132.

[0049] S131, converting the format of the target spectrogram into a preset compatible format of the preset audio vocoder;

[0050] S132: The target spectrogram after format conversion is passed through a generator network of a preset audio vocoder to generate the preliminary audio waveform.

[0051] In this embodiment, the preset compatible format is a feature format compatible with the vocoder. The format of the target spectrogram is converted into the preset compatible format of the preset audio vocoder. Specifically, the target spectrogram output by the diffusion model is converted into a feature format compatible with the vocoder (such as a linear spectrogram), and then the edges of the spectrogram are smoothed to reduce the sound quality loss caused by inter-frame mutations. The processed target spectrogram is input into the preset HiFiGAN vocoder, and the spectral features are converted into a preliminary audio waveform through its generator network (including a residual module and an upsampling layer), and the discriminator of the vocoder is used to perform real-time quality evaluation on the preliminary waveform. If there is distortion (such as high-frequency noise), the generation parameters are adjusted by feedback until the sound quality standards are met. The preliminary audio waveform is generated by inputting the format-converted target spectrogram into the preset audio vocoder to facilitate speech conversion.

[0052] S140 , performing base frequency adjustment on the preliminary audio waveform and the initial base frequency through a preset filtering vocoder to generate a target converted audio.

[0053] In this embodiment, the preset filter vocoder is a model for adjusting the audio fundamental frequency. In this embodiment, the preset filter vocoder is a SiFiGAN (Sinusoidal Frequency Generative Adversarial Network) filter vocoder, which is an audio synthesis model based on a generative adversarial network (GAN). The preliminary audio waveform and the initial fundamental frequency are subjected to fundamental frequency adjustment by a preset filter vocoder. Specifically, the initial fundamental frequency after preprocessing the preliminary audio waveform is input into the preset SiFiGAN filter vocoder. The vocoder uses a source-filter structure to adjust the pitch component of the preliminary audio waveform based on the initial fundamental frequency to ensure that the melody is consistent with the audio to be converted, and performs loudness equalization processing on the adjusted waveform to keep the volume of the target conversion audio stable to generate the target conversion audio. By adjusting the preliminary audio waveform and the initial fundamental frequency in combination according to the budget filter vocoder, it is possible to simulate the rich fundamental frequency changes in human speech, such as the ups and downs of intonation, the expression of stress, etc., so that the generated speech is more natural, smooth and friendly. If applied in the financial field, it will help improve customers' experience when communicating with customer service and increase their favorability and trust in financial institutions.

[0054] In one embodiment, if Figure 7 As shown, the step S140 also includes steps S141-S142.

[0055] S141, normalizing the initial fundamental frequency according to a preset sound wave range;

[0056] S142: Adjust the preliminary audio waveform according to the processed initial fundamental frequency and the preset filter vocoder to generate the converted audio.

[0057] In the present embodiment, the preset sound wave range refers to the effective value range of the fundamental frequency set in advance according to the specific application scenario and requirements. The initial fundamental frequency is normalized according to the preset sound wave range. Specifically, the extracted initial fundamental frequency (F0) is normalized so that it matches the preset sound wave range of the filter vocoder, and the abnormal values ​​in the fundamental frequency (such as mutations caused by broken sounds) are eliminated to retain the natural pitch trajectory. According to the processed initial fundamental frequency and the preset filter vocoder, the preliminary audio waveform is adjusted to generate the converted audio. Specifically, the processed initial fundamental frequency is input into the preset filter vocoder as a control parameter. The preset filter vocoder adjusts the parameters of each filter, such as center frequency, bandwidth, etc., according to the value of the fundamental frequency. Then, the preliminary audio waveform is input into the adjusted filter vocoder, and the filter vocoder performs filtering processing on the audio waveform, changes the spectral distribution of the audio, and matches it with the processed initial fundamental frequency. Finally, the adjusted audio waveform is output, i.e., the converted audio. By adjusting the preliminary audio waveform, the fundamental frequency of the converted audio is made consistent with the processed initial fundamental frequency, thereby achieving the change of the audio pitch.

[0058] Figure 8 FIG is a schematic block diagram of an audio style conversion device 200 based on a diffusion model provided by an embodiment of the present invention. Figure 8 As shown, corresponding to the above audio style conversion method based on the diffusion model, the present invention also provides an audio style conversion device based on the diffusion model. The audio style conversion device based on the diffusion model includes a unit for executing the above audio style conversion method based on the diffusion model. The device can be configured in a terminal such as a desktop computer, tablet computer, laptop computer, etc. Specifically, please refer to Figure 8 The audio style conversion device based on the diffusion model includes an extraction unit 210, an acquisition unit 220, a conversion unit 230 and a generation unit 240.

[0059] The extraction unit 210 is configured to perform feature extraction on the received audio to be converted and the target style audio, and obtain an initial spectrogram and an initial fundamental frequency of the audio to be converted, as well as target style features of the target style audio.

[0060] In one embodiment, the extraction unit 210 includes a framing unit and an extraction sub-unit.

[0061] A framing unit, configured to perform frame processing on the target style audio according to a preset interval time and extract its spectral features;

[0062] The extraction subunit is configured to extract the spectral features using a preset style encoder to obtain the target style features.

[0063] The acquisition unit 220 is configured to generate a spectrum using a preset diffusion model according to the initial spectrum graph, the initial fundamental frequency, and the target style feature, to acquire a target spectrum graph having the target style feature.

[0064] In one embodiment, the acquisition unit 220 includes a collection unit and a training unit.

[0065] An acquisition unit, configured to acquire training audio, and perform preset audio filling training and preset loop training on the initial diffusion model according to a training spectrum of the training audio;

[0066] A training unit is configured to determine the preset diffusion model according to a training result of the initial diffusion model.

[0067] In one embodiment, the acquisition unit 220 includes a shielding unit and a prediction unit.

[0068] a masking unit, configured to randomly mask a training spectrogram of a target training audio according to a preset masking range, wherein the training audio includes the target training audio;

[0069] A prediction unit is used to control the initial diffusion model to predict the spectrum characteristics of the masked area according to the unmasked area in the training spectrum graph, so as to perform the preset audio filling training.

[0070] In one embodiment, the acquisition unit 220 includes a spectrum generating unit, a restoring unit, and an updating unit.

[0071] a spectrum generating unit, configured to generate a to-be-restored spectrogram by combining a training source speech with the target training audio through the initial diffusion model, wherein the training audio includes the training source speech, and the to-be-restored spectrogram has the style characteristics of the target training audio;

[0072] a restoration unit, configured to perform spectrum restoration on the spectrum graph to be restored using the initial diffusion model to generate a restored spectrum graph;

[0073] An updating unit is configured to reversely update the model parameters of the initial diffusion model according to the loss between the spectrum graph to be restored and the restored spectrum graph, so as to perform the preset cyclic training.

[0074] The conversion unit 230 is configured to perform spectrum conversion on the target spectrum through a preset audio vocoder to generate a preliminary audio waveform.

[0075] In one embodiment, the conversion unit 230 includes a conversion sub-unit and a waveform generation unit.

[0076] a conversion subunit, configured to convert the format of the target spectrogram into a preset compatible format of the preset audio vocoder;

[0077] The waveform generating unit is used to generate the preliminary audio waveform by passing the target spectrogram after format conversion through a generator network of a preset audio vocoder.

[0078] The generating unit 240 is configured to perform base frequency adjustment on the preliminary audio waveform and the initial base frequency through a preset filtering vocoder to generate a target converted audio.

[0079] In one embodiment, the generating unit 240 includes a normalizing unit and an adjusting unit.

[0080] a normalizing unit, configured to normalize the initial fundamental frequency according to a preset sound wave range;

[0081] The adjustment unit is configured to adjust the preliminary audio waveform according to the processed initial fundamental frequency and the preset filter vocoder to generate the converted audio.

[0082] It should be noted that those skilled in the art will clearly understand that the specific implementation process of the above-mentioned audio style conversion device 200 based on the diffusion model and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of the description, it will not be repeated here.

[0083] The above-mentioned audio style conversion device based on the diffusion model can be implemented in the form of a computer program. The computer program can be used in Figure 9 Runs on the computer device shown.

[0084] See also Figure 9 , Figure 9 This is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 can be a terminal or a server. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, personal digital assistant, wearable device, or other electronic device with communication capabilities. The server can be a standalone server or a server cluster consisting of multiple servers.

[0085] See Figure 9 The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .

[0086] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can enable the processor 502 to perform a composite laser printing method.

[0087] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0088] The internal memory 504 provides an environment for running the computer program 5032 in the non-volatile storage medium 503 . When the computer program 5032 is executed by the processor 502 , the processor 502 can execute an audio style conversion method based on a diffusion model.

[0089] The network interface 505 is used to communicate with other devices over the network. Figure 9 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0090] The processor 502 is configured to run a computer program 5032 stored in the memory to implement the steps of the above method.

[0091] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0092] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0093] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the steps of the above method.

[0094] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0095] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0096] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0097] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0098] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.

[0099] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. An audio style transfer method based on a diffusion model, characterized in that: include: Performing feature extraction on the received audio to be converted and the target style audio to obtain an initial spectrogram and an initial fundamental frequency of the audio to be converted, as well as target style features of the target style audio; Generate a spectrum using a preset diffusion model based on the initial spectrum graph, the initial fundamental frequency, and the target style feature to obtain a target spectrum graph having the target style feature; The target spectrogram is converted into a spectrum by a preset audio vocoder to generate a preliminary audio waveform; The preliminary audio waveform and the initial base frequency are subjected to base frequency adjustment by a preset filter vocoder to generate a target converted audio.

2. The method according to claim 1, characterized in that The steps of extracting features of the target style audio and obtaining target style features include: Frame processing is performed on the target style audio according to a preset interval time, and its spectrum features are extracted; The spectral features are subjected to feature extraction by a preset style encoder to obtain the target style features.

3. The method according to claim 1, characterized in that Before the step of generating a spectrum using a preset diffusion model according to the initial spectrum graph, the initial fundamental frequency, and the target style feature, the method includes: Collect training audio, and perform preset audio filling training and preset loop training on the initial diffusion model according to the training spectrum of the training audio; The preset diffusion model is determined according to the training result of the initial diffusion model.

4. The method according to claim 3, characterized in that The step of performing preset audio filling training on the initial diffusion model according to the training spectrogram of the training audio includes: Randomly masking a training spectrogram of a target training audio according to a preset masking range, wherein the training audio includes the target training audio; The initial diffusion model is controlled to predict the spectrum characteristics of the masked area according to the unmasked area in the training spectrogram to perform the preset audio filling training.

5. The method according to claim 3, characterized in that The steps of performing a preset loop training on the initial diffusion model according to the training spectrogram of the training audio include: generating a to-be-restored spectrogram by combining a training source speech and the target training audio with the initial diffusion model, wherein the training audio includes the training source speech, and the to-be-restored spectrogram has the style characteristics of the target training audio; Performing spectrum restoration on the spectrum graph to be restored using the initial diffusion model to generate a restored spectrum graph; The model parameters of the initial diffusion model are reversely updated according to the loss between the to-be-restored spectrum graph and the restored spectrum graph to perform the preset cyclic training.

6. The method according to claim 1, characterized in that The step of converting the target spectrogram through a preset audio vocoder to generate a preliminary audio waveform includes: Converting the target spectrogram into a format compatible with the preset audio vocoder; The target spectrogram after format conversion is passed through a generator network of a preset audio vocoder to generate the preliminary audio waveform.

7. The method according to claim 1, characterized in that The step of adjusting the base frequency of the preliminary audio waveform and the initial base frequency through a preset filter vocoder to generate a target converted audio includes: Normalizing the initial fundamental frequency according to a preset sound wave range; The preliminary audio waveform is adjusted according to the processed initial fundamental frequency and the preset filter vocoder to generate the converted audio.

8. An audio style conversion device based on a diffusion model, characterized in that: include: an extraction unit, configured to perform feature extraction on the received audio to be converted and the target style audio, to obtain an initial spectrogram and an initial fundamental frequency of the audio to be converted, and target style features of the target style audio; an acquiring unit, configured to generate a spectrum using a preset diffusion model according to the initial spectrum graph, the initial fundamental frequency, and the target style feature, to acquire a target spectrum graph having the target style feature; a conversion unit, configured to convert the target spectrogram into a spectrum through a preset audio vocoder to generate a preliminary audio waveform; The generating unit is configured to adjust the base frequency of the preliminary audio waveform and the initial base frequency through a preset filtering vocoder to generate a target converted audio.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 7 can be implemented.