Audio conversion method and system based on diffusion transformation, terminal and storage medium

By using an audio conversion method based on diffusion transform, the problems of difficulty in zero-sample conversion and sound quality loss in existing technologies are solved, achieving efficient, zero-sample vocal conversion and improving audio conversion quality and efficiency.

CN121789700APending Publication Date: 2026-04-03SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing singing conversion technology requires a large amount of training data from the target speaker, cannot achieve zero-sample conversion, and the conversion quality deteriorates and the sound quality is severely lost in scenarios with few or no samples.

Method used

An audio conversion method based on diffusion transform is adopted. By acquiring reference audio and initial audio, resampling and standardizing are performed, multi-dimensional features are extracted and fused, a vocoder is used for feature conversion, reverb processing is added, and high-quality audio is output.

Benefits of technology

It achieves zero-sample vocal conversion, improves the quality and efficiency of audio conversion, lowers the barrier to entry, and expands application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789700A_ABST
    Figure CN121789700A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio processing, and discloses an audio conversion method and system based on diffusion transformation, a terminal and a storage medium, and the method comprises the steps: carrying out the resampling of an initial audio, and carrying out the standardization processing, and obtaining a screened target audio; carrying out feature extraction on the target audio and carrying out multi-dimensional feature fusion processing to obtain a fusion feature; adaptive adjustment is performed on the fusion features according to the reference audio to obtain feature blocks, feature conversion is performed after the feature blocks are overlapped, and an audio waveform is output; and performing compression processing on the dynamic range of the audio waveform, adding reverberation according to an audio scene, and outputting a final audio. According to the invention, a complete processing flow from audio input to high-quality singing sound conversion output is realized, the continuity of the audio is ensured, and the quality of the converted audio is improved at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to an audio conversion method, system, terminal, and computer-readable storage medium based on diffusion transformation. Background Technology

[0002] Voice conversion for singing, an important branch of speech conversion technology, aims to convert the voice of a source singer into the vocal characteristics of a target singer, while preserving the lyrics, melody, rhythm, and emotional expression. With the rapid development of artificial intelligence technology, voice conversion for singing technology is showing broad application prospects in music production, online entertainment, and content creation.

[0003] However, most existing singing conversion technologies require a large amount of training data from the target speaker (usually several hours or even tens of hours of audio data). For new speakers, retraining or fine-tuning is required, making it impossible to achieve true zero-sample conversion. In scenarios with few or zero samples, the conversion quality of existing technologies drops significantly, easily leading to problems such as inaccurate voice features and loss of sound quality.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] The main objective of this invention is to provide an audio conversion method, system, terminal, and computer-readable storage medium based on diffusion transformation, aiming to solve the problems of low zero-sample performance, poor real-time performance, and low audio quality in the prior art.

[0006] To achieve the above objectives, the present invention provides an audio conversion method based on diffusion transform, the audio conversion method based on diffusion transform comprising the following steps: Obtain reference audio and initial audio, resample the initial audio, perform standardization processing, and obtain the filtered target audio; Feature extraction is performed on the target audio, and the extracted multi-dimensional features are fused to obtain fused features; The fusion features are adaptively adjusted based on the reference audio to obtain multiple feature blocks. Each feature block is overlapped and then a vocoder is used for feature conversion to output an audio waveform. The dynamic range of the audio waveform is compressed, and reverb is added to the processed audio according to the audio context input by the user, and the final audio is output.

[0007] Optionally, the audio conversion method based on diffusion transform, wherein obtaining the reference audio and the initial audio, resampling the initial audio, and performing normalization processing to obtain the filtered target audio specifically includes: The system obtains reference audio and source audio input by the user, inputs the source audio into the audio preprocessing module for format detection and quality assessment, and outputs initial audio that conforms to the rules. Based on the spectrogram of the reference audio, the initial audio is resampled, and the resampled result is input into a pre-trained model for separation processing, outputting human voice audio and accompaniment audio; The sampling rate, bit depth, and number of channels of the human voice audio are standardized to obtain the target audio.

[0008] Optionally, in the aforementioned audio conversion method based on diffusion transform, the multi-dimensional features include: acoustic features, semantic features, and pitch features; The step of extracting features from the target audio and fusing the extracted multi-dimensional features to obtain fused features specifically includes: The target audio is subjected to feature extraction to obtain the acoustic features, wherein the target audio is a first sampling rate; The target audio is resampled to a second sampling rate, and features are extracted from the target audio at the second sampling rate to obtain the semantic features; The target audio is resampled to a third sampling rate, and features are extracted from the target audio at the third sampling rate to obtain the pitch features; The acoustic features, semantic features, and pitch features are time-aligned to obtain fused features.

[0009] Optionally, the audio conversion method based on diffusion transform, wherein the step of adaptively adjusting the fusion features according to the reference audio to obtain multiple feature blocks, overlapping each feature block, and then performing feature conversion processing using a vocoder to output an audio waveform, specifically includes: The fused audio and the reference audio are length aligned using a length adjuster, and the aligned long audio is then cut into multiple feature blocks of equal length. All the feature blocks are overlapped using a preset overlap method, the overlapped areas are smoothed, and the processed superimposed audio is sent to the vocoder. Using the vocoder, the spectral characteristics of the superimposed audio are converted into a time-domain waveform. The time-domain waveform is then subjected to noise reduction and volume normalization to output a standard format audio waveform.

[0010] Optionally, the audio conversion method based on diffusion transform, wherein the step of using a preset overlap method to overlap all the feature blocks, smoothing the overlapped regions, and sending the processed superimposed audio to a vocoder, specifically includes: According to the time sequence, determine the frame order of all the feature blocks, and stack all the feature blocks according to the frame order; The audio features of all the feature blocks are aligned using a cross-correlation algorithm, and cross-fade-in and cross-fade-out processing is performed on all the overlapping regions to smooth the audio features of the overlapping regions. All overlapping regions are weighted using a window function, and the feature values ​​corresponding to each overlapping region are linearly interpolated according to all weights to generate the corresponding transition frame. All the aforementioned transition frames are used to replace the corresponding overlapping regions to output superimposed audio, and the superimposed audio is sent to a vocoder.

[0011] Optionally, the audio conversion method based on diffusion transform, wherein compressing the dynamic range of the audio waveform and adding reverb to the processed audio according to the user-input audio context to output the final audio, specifically includes: The audio waveform is input into a dynamic compression model, which calculates the instantaneous energy of the audio waveform and sets the compression threshold and compression ratio of the audio waveform. When the instantaneous energy of the audio waveform exceeds the compression threshold, the peak value of the audio waveform is reduced according to the compression ratio, and a window function is used to smooth the transition of the compressed audio waveform, and finally the compressed audio of the audio waveform is output. The system acquires the audio context input by the user, sets the reverberation time and pre-delay based on the audio context, and simulates the sound wave reflection of the audio context using a delay line. Based on the reverberation time and the pre-delay, the sound wave reflections and the compressed audio are linearly reverberated through convolution operations to output the final audio.

[0012] Optionally, the audio conversion method based on diffusion transform, wherein the dynamic range of the audio waveform is compressed, reverb is added to the processed audio according to the user-input audio context, and the final audio is output, further includes: The signal peak of the final audio is monitored in real time. If the signal peak of the final audio exceeds the distortion preset value range, the phase distortion of the final audio is corrected and the repaired audio is output. Set a compensation volume threshold. If the volume of the final audio exceeds the preset volume range, then compensate the volume according to the compression ratio and output the repaired audio. A monitoring window is set up, and the final audio is monitored in real time according to the size of the monitoring window to obtain the maximum amplitude value within the monitoring window. If the maximum amplitude value of the monitoring window is abnormal, spline interpolation is performed on the final audio to repair it, and the repaired audio is output.

[0013] Furthermore, to achieve the above objectives, the present invention also provides an audio conversion system based on diffusion transform, wherein the audio conversion system based on diffusion transform includes: The audio preprocessing module is used to acquire reference audio and initial audio, resample the initial audio, perform normalization processing, and obtain the filtered target audio. The feature extraction module is used to extract features from the target audio and fuse the extracted multi-dimensional features to obtain fused features. The diffusion conversion module is used to adaptively adjust the fusion features according to the reference audio to obtain multiple feature blocks, overlap each feature block and then use a vocoder to perform feature conversion processing to output an audio waveform; The post-processing module is used to compress the dynamic range of the audio waveform, add reverb to the processed audio according to the audio context input by the user, and output the final audio.

[0014] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and an audio conversion program based on diffusion transformation stored in the memory and executable on the processor, wherein when the audio conversion program based on diffusion transformation is executed by the processor, it implements the steps of the audio conversion method based on diffusion transformation as described above.

[0015] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an audio conversion program based on diffusion transformation, which, when executed by a processor, implements the steps of the audio conversion method based on diffusion transformation as described above.

[0016] In this invention, a reference audio and an initial audio are acquired. The initial audio is resampled and then standardized to obtain a filtered target audio. Feature extraction is performed on the target audio, and the extracted multi-dimensional features are fused to obtain fused features. The fused features are adaptively adjusted based on the reference audio to obtain multiple feature blocks. Each feature block is overlapped and then converted using a vocoder to output an audio waveform. The dynamic range of the audio waveform is compressed, and reverb is added to the processed audio based on the user-input audio context to output the final audio. This invention achieves a complete processing flow from audio input to high-quality vocal conversion output, ensuring audio continuity while improving the quality of the converted audio. Attached Figure Description

[0017] Figure 1 This is a flowchart of a preferred embodiment of the audio conversion method based on diffusion transformation of the present invention; Figure 2 This is an overall flowchart of a preferred embodiment of the audio conversion method based on diffusion transformation of the present invention; Figure 3 This is a structural diagram of a preferred embodiment of the audio conversion system based on diffusion transformation of the present invention; Figure 4 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0019] Existing singing voice conversion technologies require a large amount of training data from the target speaker, making true zero-sample singing voice conversion impossible. Furthermore, high-quality singing voice conversion models suffer from high inference latency, failing to meet real-time application requirements. Additionally, sound quality distortion easily occurs during conversion, leading to inaccurate pitch in the converted audio. To address these issues, this invention discloses an audio conversion method based on diffusion transform. Through technological and application innovations, this invention significantly improves the quality and efficiency of singing voice conversion, lowers the barrier to entry, expands application scenarios, and lays a solid foundation for the efficient application of singing voice conversion technology.

[0020] The preferred embodiment of the audio conversion method based on diffusion transform of the present invention, such as... Figure 1 As shown, the audio conversion method based on diffusion transform includes the following steps: Step S10: Obtain reference audio and initial audio, resample the initial audio, and perform standardization processing to obtain the filtered target audio.

[0021] Among them, such as Figure 2 As shown, the audio preprocessing module is responsible for performing preliminary processing on the input audio to prepare for subsequent feature extraction and conversion. This module includes a human voice separation submodule, which realizes the function of extracting pure human voice from mixed audio.

[0022] Specifically, the system obtains the reference audio and source audio input by the user, inputs the source audio into the audio preprocessing module for format detection and quality assessment, and outputs the initial audio that conforms to the rules. Based on the spectrogram of the reference audio, the initial audio is resampled, and the resampled result is input into a pre-trained model for separation processing, outputting human voice audio and accompaniment audio; The sampling rate, bit depth, and number of channels of the human voice audio are standardized to obtain the target audio.

[0023] Specifically, for the input audio, the audio preprocessing module supports loading multiple audio formats and further performs format detection and quality assessment on the loaded audio files to achieve initial separation of high-quality audio and improve the efficiency of audio processing.

[0024] Furthermore, the spectrogram of the audio is sampled at intervals to reflect the characteristics of the audio, thereby realizing the resampling operation of the initial audio. A pre-trained model is then used to separate the vocals and accompaniment from the resampled audio, and the vocal audio is standardized through various preprocessing steps.

[0025] Step S20: Extract features from the target audio and fuse the extracted multi-dimensional features to obtain fused features.

[0026] Among them, such as Figure 2 As shown, the multi-dimensional features include acoustic features, semantic features, and pitch features. The feature extraction module is responsible for extracting multi-dimensional features from the preprocessed audio, including semantic features, acoustic features, and pitch features, to provide a comprehensive feature representation for subsequent diffusion conversion.

[0027] Specifically, feature extraction is performed on the target audio to obtain the acoustic features, wherein the target audio is a first sampling rate; The target audio is resampled to a second sampling rate, and features are extracted from the target audio at the second sampling rate to obtain the semantic features; The target audio is resampled to a third sampling rate, and features are extracted from the target audio at the third sampling rate to obtain the pitch features; The acoustic features, semantic features, and pitch features are time-aligned to obtain fused features.

[0028] In the embodiments disclosed in this invention, different frequencies are used to extract corresponding features for different modal features; semantic features of audio at specific frequencies can be extracted using an automatic speech recognition model; multidimensional Mel spectrograms can be generated for the extracted acoustic features to capture audio spectral features; fundamental frequency information can be extracted for pitch features to support pitch analysis and adjustment; finally, for features of different dimensions, alignment of each modal feature is ensured on the time axis, and multidimensional features are fused into a unified feature representation for further audio optimization processing by subsequent modules.

[0029] This invention achieves comprehensive sound modeling by combining semantic, pitch, and spectral features, improves the integrity of audio features, integrates a complete audio processing pipeline, and realizes end-to-end automated processing.

[0030] Step S30: Adaptively adjust the fusion features according to the reference audio to obtain multiple feature blocks. After overlapping each feature block, use a vocoder to perform feature conversion processing and output the audio waveform.

[0031] Among them, such as Figure 2 As shown, the diffusion conversion module adopts an optimized diffusion model architecture, which can achieve high-quality vocal conversion and supports real-time processing and long audio conversion.

[0032] Specifically, the fused audio and the reference audio are length aligned using a length adjuster, and the aligned long audio is cut into multiple feature blocks of equal length. All the feature blocks are overlapped using a preset overlap method, the overlapped areas are smoothed, and the processed superimposed audio is sent to the vocoder. Using the vocoder, the spectral characteristics of the superimposed audio are converted into a time-domain waveform. The time-domain waveform is then subjected to noise reduction and volume normalization to output a standard format audio waveform.

[0033] In one of the embodiments disclosed in this invention, a threshold range can be set, multiple quantizers are used to adaptively adjust the audio length, a length adjuster is used to align the source audio and reference audio features, and the long audio is divided into 30-second blocks for processing; then a diffusion model is used to progressively denoise and transform the adjusted audio, and a 16-frame overlapping region is used between blocks for smooth transition to eliminate splicing marks, and finally a vocoder is used to convert the transformed features into an audio waveform.

[0034] It is important to note that this invention receives source vocal audio and reference audio of the target singer, with the reference audio duration not exceeding 10 seconds and not used for model training. The target timbre embedding is extracted from the reference audio, and then, based on a conditional flow matching diffusion model, a converted vocal version of the target singer's timbre is generated using the semantic features of the source vocal audio and the target timbre embedding as conditions. The model does not use any audio data of the target singer during the training phase, which significantly improves the convenience of audio conversion and achieves the effect of zero-sample vocal conversion.

[0035] This invention uses a block processing mechanism to divide long audio into 30-second blocks for processing, supporting audio of any length; then it uses crossfading technology to use 16 overlapping frames for smooth transition, eliminating splicing marks and improving audio fluency; finally, it uses an asynchronous processing architecture to improve system throughput.

[0036] Furthermore, the reference audio only requires an audio segment of no more than 10 seconds. Based on the reference audio, the fused audio is processed, and high-quality conversion can be achieved with only a second of reference audio. This improves zero-sample performance and ensures that the similarity between the converted audio and the target sound is no less than 90%, significantly improving the convenience and accuracy of audio conversion and enhancing the user experience.

[0037] Further, the frame order of all the feature blocks is determined according to the time sequence, and all the feature blocks are stacked according to the frame order; The audio features of all the feature blocks are aligned using a cross-correlation algorithm, and cross-fade-in and cross-fade-out processing is performed on all the overlapping regions to smooth the audio features of the overlapping regions. All overlapping regions are weighted using a window function, and the feature values ​​corresponding to each overlapping region are linearly interpolated according to all weights to generate the corresponding transition frame. All the aforementioned transition frames are used to replace the corresponding overlapping regions to output superimposed audio, and the superimposed audio is sent to a vocoder.

[0038] In one embodiment of the invention, multiple audio feature blocks of 16 frames in length are stacked, and smoothing is performed at the stacked area to improve audio quality and naturalness. First, each audio feature sequence is divided into 16 frames, with a 50% overlap between adjacent frames (including but not limited to this). The overlapping portion is smoothly transitioned using linear interpolation or a cosine window function to avoid abrupt changes. Then, dynamic time warping or cross-correlation algorithms are used to align the time axes of different audio features to ensure feature synchronization in the overlapping area. After alignment, cross-fade-in / fade-out or overlap-addition algorithms are applied to the overlapping area to smooth the feature transition. In the overlapping area, each feature value is interpolated (e.g., linear or spline interpolation) to generate transition frames, and a window function (e.g., Hamming window) is used to weight the overlapping portion to reduce splicing artifacts.

[0039] Step S40: Compress the dynamic range of the audio waveform, add reverb to the processed audio according to the audio context input by the user, and output the final audio.

[0040] Among them, such as Figure 2 As shown, the post-processing module performs professional audio processing on the converted audio to improve audio quality and naturalness, including functions such as dynamic compression, reverb addition, and volume adjustment.

[0041] Specifically, the audio waveform is input into a dynamic compression model, which calculates the instantaneous energy of the audio waveform and sets the compression threshold and compression ratio of the audio waveform. When the instantaneous energy of the audio waveform exceeds the compression threshold, the peak value of the audio waveform is reduced according to the compression ratio, and a window function is used to smooth the transition of the compressed audio waveform, and finally the compressed audio of the audio waveform is output. The system acquires the audio context input by the user, sets the reverberation time and pre-delay based on the audio context, and simulates the sound wave reflection of the audio context using a delay line. Based on the reverberation time and the pre-delay, the sound wave reflections and the compressed audio are linearly reverberated through convolution operations to output the final audio.

[0042] Depending on the audio scenario, corresponding adjustment thresholds and compression ratios can be set to improve the dynamic range of the audio. Then, the reverb effect time can be adjusted to enhance the sense of space. Finally, the volume is adaptively normalized to prevent clipping, and equalization is performed using selectable frequencies to improve the quality of the audio output. By calculating the pitch difference between the source and reference audio, automatic pitch matching is achieved to realize the effect of adaptive pitch adjustment.

[0043] Specifically, the tonic pitch distribution of the source vocalization and the target reference audio are calculated separately. Based on the semitone difference between the peak values ​​of the two tonic pitches, the pitch offset is determined. During the vocalization conversion process, the source F0 sequence is shifted according to the offset and used as a conditional input to achieve timbre migration for pitch adaptation. Furthermore, the user can also choose to enable or disable pitch offset. When enabled, the offset is limited to ±3 semitones.

[0044] Furthermore, the signal peak of the final audio is monitored in real time. If the signal peak of the final audio exceeds the distortion preset value range, the phase distortion of the final audio is corrected and the repaired audio is output. Set a compensation volume threshold. If the volume of the final audio exceeds the preset volume range, then compensate the volume according to the compression ratio and output the repaired audio. A monitoring window is set up, and the final audio is monitored in real time according to the size of the monitoring window to obtain the maximum amplitude value within the monitoring window. If the maximum amplitude value of the monitoring window is abnormal, spline interpolation is performed on the final audio to repair it, and the repaired audio is output.

[0045] In the embodiments disclosed in this invention, the automatic volume adjustment mechanism can monitor the peak value of the audio signal in real time and trigger adjustment when it exceeds a preset threshold (such as 0.9). A compressor algorithm is used to set a -20dB threshold and a 3:1 compression ratio, and automatically boost the low volume part by 3-6dB (different compensations are set for different audio scenarios) to ensure overall loudness balance. Finally, a soft limiter is used to prevent signal clipping and maintain waveform integrity.

[0046] Furthermore, to prevent clipping, a 100ms continuous monitoring window can be set to evaluate the maximum amplitude value within the window in real time. When a clipping risk is detected, the gain is automatically reduced by 2-4dB, and cubic spline interpolation is used to repair the waveform that has been slightly clipped. The volume status is updated every millisecond to ensure timely response.

[0047] Furthermore, harmonic analysis is performed on the output audio, the balance of high, mid and low frequencies is automatically adjusted, phase distortion is detected and corrected to maintain audio clarity, and a noise threshold of -40dB is set to eliminate background noise interference.

[0048] Furthermore, in the vocal conversion service, multiple concurrent requests share the same model parameters, and each request maintains only an independent intermediate computation cache to reduce memory usage. When the conditional diffusion model is jointly compressed using quantized perceptual training and knowledge distillation, the subjective sound quality score can be maintained at a high level with high accuracy. Based on the above methods, the system can also return the conversion result of the target singer's timbre in real time after the user uploads a humming audio of less than 5 seconds, and supports sliding adjustment of "timbre fidelity" and "pitch transfer intensity". The system embeds an inaudible digital watermark in the output audio, which can be used for permission tracking and authorization verification.

[0049] This invention realizes a complete processing flow from audio input to high-quality vocal conversion output, ensuring audio continuity while improving the quality of the converted audio.

[0050] Furthermore, such as Figure 3 As shown, based on the above-described audio conversion method based on diffusion transform, the present invention also provides an audio conversion system based on diffusion transform, wherein the audio conversion system based on diffusion transform includes: The audio preprocessing module 51 is used to acquire reference audio and initial audio, resample the initial audio, perform standardization processing, and obtain the filtered target audio. Feature extraction module 52 is used to extract features from the target audio and fuse the extracted multi-dimensional features to obtain fused features; The diffusion conversion module 53 is used to adaptively adjust the fusion features according to the reference audio to obtain multiple feature blocks, overlap each feature block and then use a vocoder to perform feature conversion processing to output an audio waveform. The post-processing module 54 is used to compress the dynamic range of the audio waveform, add reverb to the processed audio according to the audio context input by the user, and output the final audio.

[0051] Furthermore, such as Figure 4 As shown, based on the above-described audio conversion method and system based on diffusion transformation, the present invention also provides a terminal, which includes a processor 10, a memory 20, and a display 30. Figure 4 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0052] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage units. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores an audio conversion program 40 based on diffusion transformation, which can be executed by the processor 10 to implement the audio conversion method based on diffusion transformation in this application.

[0053] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the audio conversion method based on diffusion transformation.

[0054] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.

[0055] In one embodiment, when the processor 10 executes the diffusion-based audio conversion program 40 in the memory 20, the following steps are performed: Obtain reference audio and initial audio, resample the initial audio, perform standardization processing, and obtain the filtered target audio; Feature extraction is performed on the target audio, and the extracted multi-dimensional features are fused to obtain fused features; The fusion features are adaptively adjusted based on the reference audio to obtain multiple feature blocks. Each feature block is overlapped and then a vocoder is used for feature conversion to output an audio waveform. The dynamic range of the audio waveform is compressed, and reverb is added to the processed audio according to the audio context input by the user, and the final audio is output.

[0056] The process of obtaining reference audio and initial audio, resampling the initial audio, and performing standardization to obtain the filtered target audio specifically includes: The system obtains reference audio and source audio input by the user, inputs the source audio into the audio preprocessing module for format detection and quality assessment, and outputs initial audio that conforms to the rules. Based on the spectrogram of the reference audio, the initial audio is resampled, and the resampled result is input into a pre-trained model for separation processing, outputting human voice audio and accompaniment audio; The sampling rate, bit depth, and number of channels of the human voice audio are standardized to obtain the target audio.

[0057] The multi-dimensional features include: acoustic features, semantic features, and pitch features; The step of extracting features from the target audio and fusing the extracted multi-dimensional features to obtain fused features specifically includes: The target audio is subjected to feature extraction to obtain the acoustic features, wherein the target audio is a first sampling rate; The target audio is resampled to a second sampling rate, and features are extracted from the target audio at the second sampling rate to obtain the semantic features; The target audio is resampled to a third sampling rate, and features are extracted from the target audio at the third sampling rate to obtain the pitch features; The acoustic features, semantic features, and pitch features are time-aligned to obtain fused features.

[0058] Specifically, the step of adaptively adjusting the fusion features based on the reference audio to obtain multiple feature blocks, overlapping each feature block, performing feature conversion processing using a vocoder, and outputting an audio waveform includes: The fused audio and the reference audio are length aligned using a length adjuster, and the aligned long audio is then cut into multiple feature blocks of equal length. All the feature blocks are overlapped using a preset overlap method, the overlapped areas are smoothed, and the processed superimposed audio is sent to the vocoder. Using the vocoder, the spectral characteristics of the superimposed audio are converted into a time-domain waveform. The time-domain waveform is then subjected to noise reduction and volume normalization to output a standard format audio waveform.

[0059] Specifically, the step of using a preset overlap method to overlap all the feature blocks, smoothing the overlapped areas, and sending the processed superimposed audio to a vocoder includes: According to the time sequence, determine the frame order of all the feature blocks, and stack all the feature blocks according to the frame order; The audio features of all the feature blocks are aligned using a cross-correlation algorithm, and cross-fade-in and cross-fade-out processing is performed on all the overlapping regions to smooth the audio features of the overlapping regions. All overlapping regions are weighted using a window function, and the feature values ​​corresponding to each overlapping region are linearly interpolated according to all weights to generate the corresponding transition frame. All the aforementioned transition frames are used to replace the corresponding overlapping regions to output superimposed audio, and the superimposed audio is sent to a vocoder.

[0060] Specifically, the process of compressing the dynamic range of the audio waveform, adding reverb to the processed audio based on the user-input audio context, and outputting the final audio includes: The audio waveform is input into a dynamic compression model, which calculates the instantaneous energy of the audio waveform and sets the compression threshold and compression ratio of the audio waveform. When the instantaneous energy of the audio waveform exceeds the compression threshold, the peak value of the audio waveform is reduced according to the compression ratio, and a window function is used to smooth the transition of the compressed audio waveform, and finally the compressed audio of the audio waveform is output. The system acquires the audio context input by the user, sets the reverberation time and pre-delay based on the audio context, and simulates the sound wave reflection of the audio context using a delay line. Based on the reverberation time and the pre-delay, the sound wave reflections and the compressed audio are linearly reverberated through convolution operations to output the final audio.

[0061] The process includes compressing the dynamic range of the audio waveform, adding reverb to the processed audio based on the user-input audio context, and outputting the final audio. This is followed by: The signal peak of the final audio is monitored in real time. If the signal peak of the final audio exceeds the distortion preset value range, the phase distortion of the final audio is corrected and the repaired audio is output. Set a compensation volume threshold. If the volume of the final audio exceeds the preset volume range, then compensate the volume according to the compression ratio and output the repaired audio. A monitoring window is set up, and the final audio is monitored in real time according to the size of the monitoring window to obtain the maximum amplitude value within the monitoring window. If the maximum amplitude value of the monitoring window is abnormal, spline interpolation is performed on the final audio to repair it, and the repaired audio is output.

[0062] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an audio conversion program based on diffusion transformation, which, when executed by a processor, implements the steps of the audio conversion method based on diffusion transformation as described above.

[0063] In summary, this invention provides an audio conversion method and related equipment based on diffusion transform. The method includes: acquiring reference audio and initial audio; resampling the initial audio and performing standardization processing to obtain a filtered target audio; extracting features from the target audio and fusing the extracted multi-dimensional features to obtain fused features; adaptively adjusting the fused features according to the reference audio to obtain multiple feature blocks; overlapping each feature block and performing feature conversion processing using a vocoder to output an audio waveform; compressing the dynamic range of the audio waveform; adding reverb to the processed audio according to the audio context input by the user; and outputting the final audio. This invention realizes a complete processing flow from audio input to high-quality vocal conversion output, ensuring audio continuity while improving the quality of the converted audio.

[0064] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.

[0065] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0066] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. An audio conversion method based on diffusion transform, characterized in that, The audio conversion method based on diffusion transform includes: Obtain reference audio and initial audio, resample the initial audio, perform standardization processing, and obtain the filtered target audio; Feature extraction is performed on the target audio, and the extracted multi-dimensional features are fused to obtain fused features; The fusion features are adaptively adjusted based on the reference audio to obtain multiple feature blocks. Each feature block is overlapped and then a vocoder is used for feature conversion to output an audio waveform. The dynamic range of the audio waveform is compressed, and reverb is added to the processed audio according to the audio context input by the user, and the final audio is output.

2. The audio conversion method based on diffusion transform according to claim 1, characterized in that, The process of obtaining reference audio and initial audio, resampling the initial audio, and performing standardization to obtain the filtered target audio specifically includes: The system obtains reference audio and source audio input by the user, inputs the source audio into the audio preprocessing module for format detection and quality assessment, and outputs initial audio that conforms to the rules. Based on the spectrogram of the reference audio, the initial audio is resampled, and the resampled result is input into a pre-trained model for separation processing, outputting human voice audio and accompaniment audio; The sampling rate, bit depth, and number of channels of the human voice audio are standardized to obtain the target audio.

3. The audio conversion method based on diffusion transform according to claim 1, characterized in that, The multi-dimensional features include: acoustic features, semantic features, and pitch features; The step of extracting features from the target audio and fusing the extracted multi-dimensional features to obtain fused features specifically includes: The target audio is subjected to feature extraction to obtain the acoustic features, wherein the target audio is a first sampling rate; The target audio is resampled to a second sampling rate, and features are extracted from the target audio at the second sampling rate to obtain the semantic features; The target audio is resampled to a third sampling rate, and features are extracted from the target audio at the third sampling rate to obtain the pitch features; The acoustic features, semantic features, and pitch features are time-aligned to obtain fused features.

4. The audio conversion method based on diffusion transform according to claim 1, characterized in that, The step of adaptively adjusting the fused features based on the reference audio to obtain multiple feature blocks, overlapping each feature block, performing feature conversion processing using a vocoder, and outputting an audio waveform specifically includes: The fused audio and the reference audio are length aligned using a length adjuster, and the aligned long audio is then cut into multiple feature blocks of equal length. All the feature blocks are overlapped using a preset overlap method, the overlapped areas are smoothed, and the processed superimposed audio is sent to the vocoder. Using the vocoder, the spectral characteristics of the superimposed audio are converted into a time-domain waveform. The time-domain waveform is then subjected to noise reduction and volume normalization to output a standard format audio waveform.

5. The audio conversion method based on diffusion transform according to claim 4, characterized in that, The process of using a preset overlap method to overlap all the feature blocks, smoothing the overlapped areas, and sending the processed superimposed audio to a vocoder specifically includes: According to the time sequence, determine the frame order of all the feature blocks, and stack all the feature blocks according to the frame order; The audio features of all the feature blocks are aligned using a cross-correlation algorithm, and cross-fade-in and cross-fade-out processing is performed on all the overlapping regions to smooth the audio features of the overlapping regions. All overlapping regions are weighted using a window function, and the feature values ​​corresponding to each overlapping region are linearly interpolated according to all weights to generate the corresponding transition frame. All the aforementioned transition frames are used to replace the corresponding overlapping regions to output superimposed audio, and the superimposed audio is sent to a vocoder.

6. The audio conversion method based on diffusion transform according to claim 1, characterized in that, The process of compressing the dynamic range of the audio waveform, adding reverb to the processed audio based on the user-input audio context, and outputting the final audio specifically includes: The audio waveform is input into a dynamic compression model, which calculates the instantaneous energy of the audio waveform and sets the compression threshold and compression ratio of the audio waveform. When the instantaneous energy of the audio waveform exceeds the compression threshold, the peak value of the audio waveform is reduced according to the compression ratio, and a window function is used to smooth the transition of the compressed audio waveform, and finally the compressed audio of the audio waveform is output. The system acquires the audio context input by the user, sets the reverberation time and pre-delay based on the audio context, and simulates the sound wave reflection of the audio context using a delay line. Based on the reverberation time and the pre-delay, the sound wave reflections and the compressed audio are linearly reverberated through convolution operations to output the final audio.

7. The audio conversion method based on diffusion transform according to claim 6, characterized in that, The process of compressing the dynamic range of the audio waveform, adding reverb to the processed audio based on the user-input audio context, and outputting the final audio also includes: The signal peak of the final audio is monitored in real time. If the signal peak of the final audio exceeds the distortion preset value range, the phase distortion of the final audio is corrected and the repaired audio is output. Set a compensation volume threshold. If the volume of the final audio exceeds the preset volume range, then compensate the volume according to the compression ratio and output the repaired audio. A monitoring window is set up, and the final audio is monitored in real time according to the size of the monitoring window to obtain the maximum amplitude value within the monitoring window. If the maximum amplitude value of the monitoring window is abnormal, spline interpolation is performed on the final audio to repair it, and the repaired audio is output.

8. An audio conversion system based on diffusion transform, characterized in that, The diffusion-based audio conversion system is used to implement the diffusion-based audio conversion method as described in any one of claims 1-7, wherein the diffusion-based audio conversion system comprises: The audio preprocessing module is used to acquire reference audio and initial audio, resample the initial audio, perform normalization processing, and obtain the filtered target audio. The feature extraction module is used to extract features from the target audio and fuse the extracted multi-dimensional features to obtain fused features. The diffusion conversion module is used to adaptively adjust the fusion features according to the reference audio to obtain multiple feature blocks, overlap each feature block and then use a vocoder to perform feature conversion processing to output an audio waveform; The post-processing module is used to compress the dynamic range of the audio waveform, add reverb to the processed audio according to the audio context input by the user, and output the final audio.

9. A terminal, characterized in that, The terminal includes: a memory, a processor, and an audio conversion program based on diffusion transformation stored in the memory and executable on the processor, wherein when the audio conversion program based on diffusion transformation is executed by the processor, it implements the steps of the audio conversion method based on diffusion transformation as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an audio conversion program based on diffusion transformation, which, when executed by a processor, implements the steps of the audio conversion method based on diffusion transformation as described in any one of claims 1-7.