Tone conversion method and system based on machine learning algorithm

Through the tone conversion method based on machine learning algorithm, the source audio signal is decomposed and the target tone characteristics are fused, combined with anti-time frequency conversion and dynamic range adjustment, the problem of unnatural traditional tone conversion is solved, and more natural tone conversion and voice content retention is achieved.

CN120236602AActive Publication Date: 2025-07-01CHENGDU HAOXILI TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510334911.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-01
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Traditional tone conversion methods do not have natural tone, poorly retain voice content of the original audio, or poorly perform when dealing with complex tone attributes.

Method used

The tone conversion method based on machine learning algorithm is adopted, and the source audio signal and target tone description are obtained, the source audio signal and target tone description are obtained, the spectral characteristics are extracted, and the feature fusion is used to fusion network is used to generate the target audio signal, and dynamic range adjustment and post-processing are performed to improve the tone conversion quality.

Benefits of technology

A more natural tone conversion is achieved, retaining the voice content of the original audio, and improving the effect of tone conversion in complex situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236602A_ABST
    Figure CN120236602A_ABST
Patent Text Reader

Abstract

The invention provides a timbre conversion method and system based on a machine learning algorithm, and the method comprises the steps: firstly obtaining a source audio signal and a target timbre description, decomposing the source audio signal into a plurality of audio frames, extracting a spectrum feature, converting the target timbre description into a multi-dimensional timbre feature vector, and carrying out the recognition of the multi-dimensional timbre feature vector; and feature fusion is carried out through a timbre fusion network to obtain converted spectrum features, and a target audio signal is generated through inverse time-frequency transformation. Besides, the invention also relates to the steps of dynamic range adjustment of the converted spectrum features, post-processing of the target audio signal and the like, so that the quality and the effect of tone conversion are improved, and the problem of tone conversion under various complex conditions is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of machine learning, and specifically, to a timbre conversion method and system based on a machine learning algorithm. Background Art

[0002] In the field of audio processing, timbre conversion is a technology of great significance. Different timbres can convey different emotions, information, and styles. For example, in speech synthesis, users may hope to convert a plain speech timbre into a more emotional, magnetic, or softer timbre; in music production, it is also often necessary to convert the timbre of one instrument into that of another to obtain a unique musical effect. Traditional timbre conversion methods often have many limitations, such as the converted timbre not being natural enough, the speech content of the original audio being poorly retained, or the effect being unsatisfactory when dealing with complex timbre attributes. Summary of the Invention

[0003] In view of this, the purpose of the present application is to provide a timbre conversion method and system based on a machine learning algorithm.

[0004] According to a first aspect of the present application, a timbre conversion method based on a machine learning algorithm is provided, and the method includes: Obtain a source audio signal and a target timbre description, where the source audio signal contains the original timbre features to be converted, and the target timbre description is used to define the desired output timbre attributes; Decompose the source audio signal into multiple audio frames, and perform time-frequency transformation processing on each audio frame to extract the spectral features of the source audio signal; Convert the target timbre description into a multi-dimensional timbre feature vector, where the multi-dimensional timbre feature vector contains quantization parameters corresponding to the timbre attributes; Input the spectral features and the multi-dimensional timbre feature vector into a timbre fusion network for feature fusion to generate converted spectral features containing target timbre attributes; Perform inverse time-frequency transformation processing on the converted spectral features to generate a target audio signal, and the timbre attributes of the target audio signal are consistent with the desired timbre attributes in the target timbre description.

[0005] In a possible implementation manner of the first aspect, the step of decomposing the source audio signal into multiple audio frames, and performing time-frequency transformation processing on each audio frame to extract the spectral features of the source audio signal includes: Perform frame division processing on the source audio signal to obtain multiple audio frames with a fixed duration, and each audio frame is processed by a window function to reduce spectral leakage; Perform Fourier transform on each windowed audio frame to obtain the corresponding magnitude spectrum and phase spectrum; Extract Mel-frequency cepstral coefficients from the magnitude spectrum as the spectral features, and retain the phase spectrum as the reconstruction parameter.

[0006] In a possible implementation manner of the first aspect, the target timbre description includes a reference audio segment or a text description input by the user; When the target timbre description is a reference audio segment, perform fundamental frequency extraction and formant analysis on the reference audio segment to generate the multi-dimensional timbre feature vector; When the target timbre description is a text description, call a pre-trained timbre attribute parsing model to map the text description to the quantization parameters in the multi-dimensional timbre feature vector.

[0007] In a possible implementation manner of the first aspect, the timbre fusion network includes an encoder and a decoder. The encoder consists of multiple convolutional layers and is used to compress the spectral features into latent space features; The decoder consists of multiple transposed convolutional layers and is used to channel-concatenate the latent space features with the multi-dimensional timbre feature vector and then reconstruct them into the transformed spectral features; Wherein, the encoder and the decoder transfer low-frequency information through skip connections to retain the speech content of the source audio signal.

[0008] In a possible implementation manner of the first aspect, the training process of the timbre fusion network includes: Collect multiple groups of training audio data with different timbres, and label corresponding timbre attribute tags for each group of training audio data; Input the spectral features of the training audio into the encoder to obtain latent features, and convert the labeled timbre attribute tags into conditional vectors; Input the latent features and the conditional vectors into the decoder to generate reconstructed spectral features; Calculate the mean square error loss between the reconstructed spectral features and the target spectral features, and jointly optimize the network parameters of the encoder and the decoder.

[0009] In a possible implementation manner of the first aspect, the method further includes a step of dynamically adjusting the dynamic range of the transformed spectral features: According to the fundamental frequency trajectory and energy distribution of the source audio signal, adjust the ratio of the harmonic components and noise components of the transformed spectral features; Combine the adjusted spectral features with the retained phase spectrum to generate an intermediate signal that satisfies time-domain continuity; Perform linear prediction analysis on the intermediate signal to correct the spectral envelope to match the formant bandwidth parameters in the target timbre description.

[0010] In a possible implementation of the first aspect, the dynamic range adjustment is implemented by a differentiable signal processing module, and the differentiable signal processing module includes: A harmonic enhancement unit for enhancing the high-frequency harmonic energy according to the brightness parameter in the multi-dimensional timbre feature vector; A noise suppression unit for reducing the amplitude of the aperiodic noise according to the smoothness parameter in the multi-dimensional timbre feature vector; The outputs of the harmonic enhancement unit and the noise suppression unit are fused by weighted summation, and the weights are dynamically controlled by the emotion category parameter in the timbre attribute.

[0011] In a possible implementation of the first aspect, the method further includes the step of post-processing the target audio signal: Extracting the original rhythm feature and intonation contour of the source audio signal, and performing time alignment between the original rhythm feature and the spectrum of the target audio signal; Adjusting the fundamental frequency curve of the target audio signal according to the intonation contour so that it maintains the same speech rate and stress pattern as the source audio signal; Inputting the adjusted fundamental frequency curve into a vocoder to generate the final target audio signal.

[0012] In a possible implementation of the first aspect, when the text description input by the user contains multiple conflicting timbre attributes, the following processing is performed: Calculating the weight coefficient of each timbre attribute through an attention mechanism, and the weight coefficient is determined based on the compatibility between the original timbre feature of the source audio signal and the target timbre attribute; Combining the weighted timbre attribute parameters into a unified multi-dimensional timbre feature vector, and generating a corresponding conflict resolution log for the user to confirm.

[0013] According to the second aspect of the present application, there is provided a timbre conversion system based on a machine learning algorithm. The timbre conversion system based on a machine learning algorithm includes a machine-readable storage medium and a processor. The machine-readable storage medium stores machine-executable instructions. When the processor executes the machine-executable instructions, the timbre conversion system based on a machine learning algorithm implements the foregoing timbre conversion method based on a machine learning algorithm.

[0014] According to the third aspect of the present application, there is provided a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed, the foregoing timbre conversion method based on a machine learning algorithm is implemented.

[0015] According to any one of the above aspects, the technical effect of the present application is: First, obtain the source audio signal and the target timbre description. Decompose the source audio signal into multiple audio frames and extract spectral features. Convert the target timbre description into a multi-dimensional timbre feature vector. Then, perform feature fusion through a timbre fusion network to obtain the transformed spectral features, and generate the target audio signal through inverse time-frequency transformation. In addition, steps such as dynamic range adjustment of the transformed spectral features and post-processing of the target audio signal are also involved to improve the quality and effect of timbre conversion and solve the timbre conversion problem in various complex situations. Brief Description of the Drawings

[0016] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0017] Figure 1 Flowchart of the timbre conversion method based on machine learning algorithms provided by the embodiments of the present application.

[0018] Figure 2 Shows the component structure diagram of the timbre conversion system based on machine learning algorithms provided by the embodiments of the present application. Detailed Description of the Embodiments

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application only serve the purpose of illustration and description and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present application show the operations implemented according to some embodiments of the embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art can add multiple other operations to the flowchart or delete multiple operations from the flowchart under the guidance of the content of the present application.

[0020] Figure 1 Shows the flowchart of the timbre conversion method and system based on machine learning algorithms provided by the embodiments of the present application. It should be understood that in other embodiments, the order of some steps of the timbre conversion method based on machine learning algorithms in this embodiment can be shared according to actual needs, or some of the steps can also be omitted or maintained. The detailed steps of the timbre conversion method based on machine learning algorithms include: Step S110: Obtain a source audio signal and a target timbre description. The source audio signal contains the original timbre features to be converted, and the target timbre description is used to define the timbre attributes expected to be output.

[0021] In an actual audio processing scenario, assume that a speech synthesis project is in progress. The source audio signal can be the voice recording of a male announcer reading a news release. This voice has the unique timbre features of this male announcer, such as a relatively low and steady tone, a relatively clear pronunciation method, and there may be some slight nasal sounds in the voice. The target timbre description is the timbre attribute expected to be obtained according to the project requirements. For example, the target timbre description may be to convert the voice of this male announcer into the timbre of a young female. Such a timbre may have a higher pitch, a clearer sound quality, and a softer sound texture, etc. The target timbre description can also be defined in various ways. If there is a reference audio clip, which is the voice of a young female reading the same news release, then this reference audio clip can be used as the target timbre description; or it can be a text description directly input by the user, such as "young female timbre, higher pitch, clear and soft voice".

[0022] Step S120: Decompose the source audio signal into multiple audio frames, and perform time-frequency transformation processing on each audio frame to extract the spectral features of the source audio signal.

[0023] For the above-mentioned source audio signal of the male announcer, first perform frame division processing. The entire audio is divided into multiple audio frames of a fixed duration according to certain rules. For example, set the duration of each audio frame to 20 milliseconds, so that the entire source audio signal is divided into a series of 20-millisecond segments. After frame division, in order to reduce spectral leakage, a window function needs to be applied to each audio frame. Here, the Hann window function can be used, which can effectively reduce the spectral leakage phenomenon.

[0024] Next, perform Fourier transform on each windowed audio frame. Fourier transform is a powerful tool for converting a time-domain signal into a frequency-domain signal. Through Fourier transform, the amplitude spectrum and phase spectrum corresponding to each audio frame can be obtained. The amplitude spectrum reflects the energy distribution of the signal at different frequencies, while the phase spectrum contains the phase information of the signal. Extract the Mel-frequency cepstral coefficients from the amplitude spectrum as spectral features. Mel-frequency cepstral coefficients are features widely used in speech processing, and they can well represent the spectral characteristics of speech. At the same time, retain the phase spectrum as a reconstruction parameter because the phase spectrum is very important for restoring the time-domain characteristics of the audio signal in the subsequent inverse transformation process.

[0025] Step S130: Convert the target timbre description into a multi-dimensional timbre feature vector, where the multi-dimensional timbre feature vector contains quantization parameters corresponding to the timbre attributes.

[0026] If the target timbre description is a reference audio clip, such as the reference audio clip of a young female reading a news release mentioned earlier. First, perform fundamental frequency extraction and formant analysis on this reference audio clip. The fundamental frequency is the most basic frequency component in speech, which determines the pitch. A specific algorithm, such as the autocorrelation function method, is used to extract the fundamental frequency of the reference audio clip. The formant refers to some frequency regions where the energy is relatively concentrated in the sound spectrum, which reflects the shape and characteristics of the vocal tract. Methods such as linear predictive coding (LPC) are used to analyze the formants. Based on information such as the fundamental frequency and formants, a multi-dimensional timbre feature vector is generated. The quantization parameters in this multi-dimensional timbre feature vector can accurately describe various attributes of the target timbre, such as the specific value of the pitch, the frequency and bandwidth of the formants, etc.

[0027] If the target timbre description is a text description input by the user, such as "young female timbre, high pitch, clear and soft voice". Call a pre-trained timbre attribute parsing model, which is trained on a large number of timbre samples and corresponding text descriptions. Input the text description input by the user into this model, and the model will map the timbre attributes in the text to the quantization parameters in the multi-dimensional timbre feature vector according to the knowledge it has learned. For example, the attribute "high pitch" will be mapped to a quantization parameter with a relatively large fundamental frequency value, and "clear and soft voice" may be mapped to some quantization parameters related to spectral smoothness, harmonic structure, etc.

[0028] Step S140: Input the spectral features and the multi-dimensional timbre feature vector into a timbre fusion network for feature fusion to generate transformed spectral features containing target timbre attributes.

[0029] The timbre fusion network includes an encoder and a decoder. The encoder consists of multiple convolutional layers. For example, 3 convolutional layers can be set. The spectral features of the source audio signal extracted previously are fed into this encoder. The first convolutional layer may use a 3×3 convolutional kernel with a stride of 1 to perform preliminary compression and feature extraction on the spectral features. As it passes through multiple convolutional layers, the spectral features will be gradually compressed into latent space features.

[0030] The decoder consists of multiple transposed convolution layers, which are assumed to be set to 3 transposed convolution layers. In this process, the latent space features are concatenated with the multi-dimensional timbre feature vectors along the channel dimension. For example, if the number of channels of the latent space features is 16 and the number of channels of the multi-dimensional timbre feature vectors is 8, then the number of channels of the new features obtained after concatenation is 24. Then, through these transposed convolution layers, the concatenated features are reconstructed into the transformed spectral features. In this process, the encoder and the decoder transfer low-frequency information through skip connections to retain the speech content of the source audio signal. For example, the low-frequency information output by the first convolutional layer of the encoder will be directly transferred to the corresponding layer of the decoder through the skip connection, so as to ensure that the speech content in the original audio will not be lost during the timbre conversion process. For example, the pronunciation of words and the rhythm of speech in the original audio can be well retained.

[0031] Step S150: Perform an inverse time-frequency transformation process on the transformed spectral features to generate a target audio signal, and the timbre attribute of the target audio signal is consistent with the desired timbre attribute in the target timbre description.

[0032] After obtaining the transformed spectral features through the previous steps, now an inverse time-frequency transformation process needs to be performed on them. The inverse time-frequency transformation is the inverse process of the previous time-frequency transformation, which converts the signal in the frequency domain back to the time domain signal. In this process, using the phase spectrum retained before, the transformed spectral features are combined with the phase spectrum, and through operations such as inverse Fourier transform, a target audio signal is generated, and the timbre attribute of this target audio signal is consistent with the desired timbre attribute in the target timbre description. For example, if the target timbre description is the timbre of a young female, then the generated target audio signal should have the characteristics of a young female's timbre, such as a higher pitch, a clear and soft sound quality, etc.

[0033] Step S160: A step of performing dynamic range adjustment on the transformed spectral features.

[0034] According to the fundamental frequency trajectory and energy distribution of the source audio signal, adjust the ratio of the harmonic components and noise components of the transformed spectral features. For the example of converting a male announcer's voice to a young female's timbre mentioned above, the source audio signal of the male announcer has its specific fundamental frequency trajectory and energy distribution. The fundamental frequency trajectory reflects the change of pitch over time, and the energy distribution represents the energy magnitude at different frequencies. When converting to a young female's timbre, it may be necessary to adjust the ratio of the harmonic components and noise components according to the characteristics of a young female's timbre. For example, a young female's voice may have relatively richer harmonic components and less noise components.

[0035] Combine the adjusted spectral features with the retained phase spectrum to generate an intermediate signal that satisfies time-domain continuity. This step is to ensure that the generated signal is continuous in the time domain and will not exhibit sudden jumps or unnatural phenomena.

[0036] Perform linear prediction analysis on the intermediate signal, and correct the spectral envelope to match the formant bandwidth parameters in the target timbre description. The formant bandwidth is an important feature of timbre. For the timbre of a young female, its formant bandwidth may be different from that of a male announcer. Through linear prediction analysis, the spectral envelope can be corrected so that the formant bandwidth of the generated target audio signal can match the parameters in the target timbre description, thus making the timbre more accurately converted into the desired young female timbre.

[0037] This dynamic range adjustment is achieved through a differentiable signal processing module, which includes a harmonic enhancement unit and a noise suppression unit. The harmonic enhancement unit enhances the high-frequency harmonic energy according to the brightness parameter in the multi-dimensional timbre feature vector. For example, if the target timbre description requires a brighter sound, then the brightness parameter will be higher, and the harmonic enhancement unit will enhance the high-frequency harmonic energy accordingly, making the sound sound brighter. The noise suppression unit reduces the amplitude of the aperiodic noise according to the smoothness parameter in the multi-dimensional timbre feature vector. If the target timbre requires a smoother sound, then the smoothness parameter will prompt the noise suppression unit to reduce the noise amplitude, making the sound purer. The outputs of the harmonic enhancement unit and the noise suppression unit are fused by weighted summation, and the weights are dynamically controlled by the emotion category parameter in the timbre attribute. For example, if the target timbre is a lively young female timbre, the emotion category parameter may make it more inclined to enhance the harmonic energy during fusion to reflect the lively feeling.

[0038] Step S170: A step of post-processing the target audio signal.

[0039] Extract the original rhythm features and intonation contours of the source audio signal, and time-align the original rhythm features with the spectrum of the target audio signal. For the source audio signal of a male announcer, his reading has certain rhythm features, such as the pause time between each word, the stress distribution in the sentence, etc. The intonation contour reflects the ups and downs of the pitch. After converting the voice of the male announcer into the timbre of a young female, these original rhythm features need to be time-aligned with the spectrum of the target audio signal. This is like applying the original rhythm framework to the new timbre, so that the new timbre can be consistent with the original sound in terms of rhythm.

[0040] Adjust the fundamental frequency curve of the target audio signal according to the intonation contour to keep the same speech rate and stress pattern as the source audio signal. For example, when a male announcer reads a news script, certain important words are stressed and the speech rate of the sentence follows a certain pattern. In the target audio signal with the converted young female voice, the fundamental frequency curve should be adjusted according to the previously extracted intonation contour to ensure that the speech rate and stress pattern are the same as those of the source audio signal. The purpose of this is to make the converted voice unambiguous in expressing the content and sound more natural and fluent.

[0041] Input the adjusted fundamental frequency curve into the vocoder to generate the final target audio signal. A vocoder is a device that converts parameters such as the fundamental frequency curve into audible audio signals. Through the processing of the vocoder, the quality of the target audio signal can be further optimized, so that the finally generated target audio signal not only has the desired young female voice, but also can maintain important characteristics such as a similar rhythm, speech rate and stress pattern to the source audio signal.

[0042] Step S180: Processing steps when the text description input by the user contains multiple conflicting timbre attributes.

[0043] Suppose the text description input by the user is "both a deep tone and a clear timbre", which is a description containing conflicting timbre attributes. Calculate the weight coefficient of each timbre attribute through the attention mechanism, and this weight coefficient is determined based on the compatibility between the original timbre characteristics of the source audio signal and the target timbre attributes. For the case where the source audio signal is a male announcer, the deep tone has a certain compatibility with the original timbre because the male announcer's tone is relatively deep; while the clear timbre has a lower compatibility with the original timbre. The attention mechanism will calculate the weight coefficient of each timbre attribute according to this compatibility.

[0044] Merge the weighted timbre attribute parameters into a unified multi-dimensional timbre feature vector and generate a corresponding conflict resolution log for the user to confirm. For example, according to the calculated weight coefficients, the deep tone attribute may be given a higher weight, and the clear timbre attribute is given a lower weight. Then these weighted timbre attribute parameters are merged into a unified multi-dimensional timbre feature vector. At the same time, a conflict resolution log is generated, which will detail information such as the original requirements of each timbre attribute, the calculated weight coefficients, and the final merging process for the user to view and confirm to ensure that the result of the timbre conversion meets the user's expectations.

[0045] Figure 2The voice conversion system 100 based on a machine learning algorithm shown in the figure includes: a processor 1001 and a memory 1003. Among them, the processor 1001 and the memory 1003 are connected, such as through a bus 1002. Optionally, the voice conversion system 100 based on a machine learning algorithm may further include a transceiver 1004, and the transceiver 1004 may be used for data interaction between this server and other servers, such as data sending and / or data receiving, etc. It should be noted that in actual scheduling, the transceiver 1004 is not limited to one, and the structure of the voice conversion system 100 based on a machine learning algorithm does not constitute a limitation on the embodiments of the present application.

[0046] The processor 1001 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, digital signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of the present application. The processor 1001 may also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0047] The bus 1002 may include a path for transmitting information between the above components. The bus 1002 may be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard structure) bus, etc. The bus 1002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 2 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0048] The memory 1003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store program code and can be read by a computer, which is not limited herein.

[0049] The memory 1003 is used to store the program code for implementing the embodiments of this application and is controlled by the processor 1001 for execution. The processor 1001 is used to execute the program code stored in the memory 1003 to implement the steps shown in the foregoing method embodiments.

[0050] The embodiments of this application provide a computer-readable storage medium, on which program code is stored. When the program code is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0051] It should be understood that although the flowchart of the embodiments of this application represents each operation step through directed connection lines, the execution order of these steps is not limited to the order covered by the directed connection lines. Unless otherwise clearly stated in this document, in some implementation scenarios of the embodiments of this application, the implementation steps in each flowchart can be executed in other orders based on requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages according to the actual implementation scenario. Some or all of these sub-steps or stages can be executed in the same stage, and each sub-step or stage among these sub-steps or stages can also be executed in different stages respectively. In scenarios with different execution stages, the execution order of these sub-steps or stages can be flexibly configured based on requirements, and the embodiments of this application do not limit this.

[0052] The above are only optional implementation manners of some implementation scenarios of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the technical concept of the solution of this application, adopting other similar implementation means based on the technical idea of this application also belongs to the protection scope of the embodiments of this application.

Claims

1. A timbre conversion method based on a machine learning algorithm, characterized in that: The method comprises: Obtaining a source audio signal and a target timbre description, wherein the source audio signal contains original timbre features to be converted, and the target timbre description is used to define timbre properties of a desired output; Decomposing the source audio signal into a plurality of audio frames, and performing time-frequency transformation processing on each of the audio frames to extract frequency spectrum features of the source audio signal; Converting the target timbre description into a multi-dimensional timbre feature vector, wherein the multi-dimensional timbre feature vector includes a quantization parameter corresponding to the timbre attribute; Inputting the spectrum feature and the multi-dimensional timbre feature vector into a timbre fusion network for feature fusion to generate a converted spectrum feature containing target timbre attributes; An inverse time-frequency transform process is performed on the converted spectral features to generate a target audio signal, wherein the timbre attribute of the target audio signal is consistent with the expected timbre attribute in the target timbre description.

2. The method for tone conversion based on machine learning algorithm according to claim 1, characterized in that: The step of decomposing the source audio signal into a plurality of audio frames, performing time-frequency transformation processing on each of the audio frames, and extracting the frequency spectrum features of the source audio signal comprises: Performing frame processing on the source audio signal to obtain a plurality of audio frames of fixed duration, wherein each of the audio frames is processed by a windowing function to reduce spectrum leakage; Perform Fourier transform on each windowed audio frame to obtain the corresponding amplitude spectrum and phase spectrum; Mel-frequency cepstrum coefficients are extracted from the amplitude spectrum as the spectrum features, and the phase spectrum is retained as a reconstruction parameter.

3. The method for tone conversion based on machine learning algorithm according to claim 1, characterized in that: The target timbre description includes a reference audio clip or a text description input by a user; When the target timbre description is a reference audio segment, performing fundamental frequency extraction and formant analysis on the reference audio segment to generate the multi-dimensional timbre feature vector; When the target timbre description is a text description, a pre-trained timbre attribute analysis model is called to map the text description into a quantization parameter in the multi-dimensional timbre feature vector.

4. The method for tone conversion based on machine learning algorithm according to claim 1, characterized in that: The timbre fusion network includes an encoder and a decoder, wherein the encoder is composed of a plurality of convolutional layers and is used to compress the spectral features into latent space features; The decoder is composed of a plurality of deconvolution layers, and is used for performing channel concatenation on the latent space features and the multi-dimensional timbre feature vectors and reconstructing them into the converted spectrum features; The encoder and the decoder transmit low-frequency information through a skip connection to retain the speech content of the source audio signal.

5. The method for tone conversion based on machine learning algorithm according to claim 4, characterized in that: The training process of the timbre fusion network includes: Collect multiple sets of training audio data with different timbres, and label each set of training audio data with corresponding timbre attribute labels; Inputting the spectral features of the training audio into the encoder to obtain latent features, and converting the marked timbre attribute labels into conditional vectors; Inputting the latent feature and the conditional vector into the decoder to generate a reconstructed spectrum feature; The mean square error loss between the reconstructed spectrum feature and the target spectrum feature is calculated, and the network parameters of the encoder and the decoder are jointly optimized.

6. The method for tone conversion based on machine learning algorithm according to claim 1, characterized in that: The method further comprises the step of adjusting the dynamic range of the converted spectral features: According to the fundamental frequency trajectory and energy distribution of the source audio signal, adjusting the ratio of the harmonic component and the noise component of the converted frequency spectrum feature; Combining the adjusted frequency spectrum characteristics with the retained phase spectrum to generate an intermediate signal satisfying time domain continuity; A linear prediction analysis is performed on the intermediate signal, and the spectral envelope is modified to match the formant bandwidth parameter in the target timbre description.

7. The method for tone conversion based on machine learning algorithm according to claim 6, characterized in that: The dynamic range adjustment is implemented by a differentiable signal processing module, and the differentiable signal processing module includes: A harmonic enhancement unit, used for enhancing high-frequency harmonic energy according to the brightness parameter in the multi-dimensional timbre feature vector; A noise suppression unit, configured to reduce the amplitude of the non-periodic noise according to a smoothness parameter in the multi-dimensional timbre feature vector; The outputs of the harmonic enhancement unit and the noise suppression unit are fused by weighted summation, and the weight is dynamically controlled by the emotion category parameter in the timbre attribute.

8. The method for tone conversion based on machine learning algorithm according to claim 1, characterized in that: The method further comprises the step of post-processing the target audio signal: Extracting original rhythm features and intonation contours of the source audio signal, and time-aligning the original rhythm features with the frequency spectrum of the target audio signal; Adjusting the fundamental frequency curve of the target audio signal according to the intonation contour so as to maintain the same speech speed and stress pattern as the source audio signal; The adjusted fundamental frequency curve is input into the vocoder to generate the final target audio signal.

9. The method for tone conversion based on machine learning algorithm according to claim 3, characterized in that: When the text description input by the user contains multiple conflicting timbre attributes, the following processing is performed: Calculating a weight coefficient of each timbre attribute by an attention mechanism, wherein the weight coefficient is determined based on the compatibility between the original timbre feature of the source audio signal and the target timbre attribute; The weighted timbre attribute parameters are merged into a unified multi-dimensional timbre feature vector, and a corresponding conflict resolution log is generated for user confirmation.

10. A timbre conversion system based on a machine learning algorithm, characterized in that: It includes a processor and a computer-readable storage medium, wherein the computer-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by the processor, the timbre conversion method based on a machine learning algorithm as described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Speech synthesis model training method, speech synthesis method and related device

    CN114187891A

  • Voice tone conversion method and device, model training method and device, equipment and medium

    CN114360557A

  • Voice conversion method and device, electronic equipment and storage medium

    CN116312617A

  • Audio synthesis method, computer equipment, storage medium and program product

    CN116486778A

  • Voice conversion method and device, computer equipment and storage medium

    CN117672254A