Singing conversion methods, systems, terminals and storage media

CN117153176BActive Publication Date: 2026-08-14BEIJING UNISOUND INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]本发明实施例的目的在于提供一种歌唱转换方法、系统、终端及存储介质,旨在解决现有的歌唱转换效果较差的问题

Benefits of technology

[0037]本发明实施例,通过输入边界和参考边界对参考基频进行采样,能有效地利用参考歌曲的音素边界直接对输入语音中相同音素边界的基频进行控制,无需进行对输入语音进行调域处理,通过训练后的神经网络声码器对采样基频和语音输入特征进行歌唱转换,以达到以参考歌曲的基频作为条件,自动实现歌唱转换,提高了歌唱转换效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117153176B_ABST
    Figure CN117153176B_ABST
Patent Text Reader

Abstract

This invention provides a singing conversion method, system, terminal, and storage medium. The method includes: extracting the phoneme boundaries of input speech and a reference song to obtain input boundaries and reference boundaries, and acquiring the speech input features of the input speech and the reference fundamental frequency of the reference song; sampling the reference fundamental frequency based on the input boundaries and the reference boundaries to obtain a sampled fundamental frequency, and inputting the sampled fundamental frequency and the speech input features into a trained neural network vocoder; and performing singing conversion on the sampled fundamental frequency and the speech input features using the trained neural network vocoder to obtain the singing output speech. This invention eliminates the need for modulation processing of the input speech, and by using a trained neural network vocoder to perform singing conversion on the sampled fundamental frequency and speech input features, it automatically achieves singing conversion using the fundamental frequency of the reference song as a condition, thus improving the singing conversion effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and in particular to a singing conversion method, system, terminal, and storage medium. Background Technology

[0002] Currently, because ordinary users cannot sing professional songs, they hope to input voice commands and songs by professional singers, and output their own voice to sing songs by professional singers. Therefore, singing conversion methods are receiving increasing attention.

[0003] In existing singing conversion processes, the input speech is usually directly processed by modulation to achieve the singing conversion effect. However, because the input speech is directly processed by modulation, the singing conversion effect is poor. Summary of the Invention

[0004] The purpose of this invention is to provide a singing conversion method, system, terminal, and storage medium, aiming to solve the problem of poor singing conversion effects in existing systems.

[0005] The present invention is implemented as follows: a singing conversion method, the method comprising:

[0006] The phoneme boundaries of the input speech and the reference song are extracted separately to obtain the input boundary and the reference boundary, and the speech input features of the input speech and the reference fundamental frequency of the reference song are obtained respectively.

[0007] The reference fundamental frequency is sampled according to the input boundary and the reference boundary to obtain the sampled fundamental frequency, and the sampled fundamental frequency and the speech input features are input into the trained neural network vocoder.

[0008] The trained neural network vocoder performs singing conversion on the sampled fundamental frequency and the speech input features to obtain the singing output speech.

[0009] Preferably, sampling the reference fundamental frequency based on the input boundary and the reference boundary includes:

[0010] Perform duration matching on the phoneme durations in the input boundary and the reference boundary;

[0011] Based on the duration matching result, the phoneme duration of the input boundary, the phoneme duration of the reference boundary, and the reference fundamental frequency corresponding to the phoneme duration of the reference boundary are combined to obtain a fundamental frequency sampling group;

[0012] For each fundamental frequency sampling group, the corresponding reference fundamental frequency is upsampled or downsampled according to the phoneme duration of the input boundary to obtain the sampling fundamental frequency.

[0013] Preferably, before inputting the sampled fundamental frequency and the speech input features into the trained neural network vocoder, the method further includes:

[0014] The sample fundamental frequency and sample speech are input into the generator in the neural network vocoder, and speech is generated based on the sample fundamental frequency and sample speech by the generator to obtain the generated speech;

[0015] The generated speech is input into the multi-subband multi-scale discriminator in the neural network vocoder for speech discrimination, and the speech discrimination result is obtained.

[0016] Obtain the fundamental frequency of the sample and the singing standard audio corresponding to the sample speech, and determine the model loss based on the singing standard audio and the speech discrimination result;

[0017] The parameters of the neural network vocoder are updated based on the model loss until the neural network vocoder converges, resulting in the trained neural network vocoder.

[0018] Preferably, the step of generating speech based on the sample fundamental frequency and the sample speech using the generator includes:

[0019] The sample fundamental frequency and the sample speech are convolved according to each residual network in the generator, and the convolved sample fundamental frequency and the sample speech are upsampled according to the upsampling network in the generator.

[0020] The generator filters the upsampled sample fundamental frequency and the sample speech using an adaptive filter in the generator to obtain the generated speech.

[0021] Preferably, the step of extracting the phoneme boundaries of the input speech and the reference song respectively includes:

[0022] Phoneme alignment is performed on the input speech and the reference song respectively to obtain a first alignment result and a second alignment result;

[0023] Based on the first alignment result and the second alignment result, the alignment results of each phoneme are merged to obtain the input boundary and the reference boundary.

[0024] Preferably, the step of obtaining the speech input features of the input speech and the reference fundamental frequency of the reference song respectively includes:

[0025] The input speech is pre-emphasized, framed, and windowed sequentially, and the short-time analysis windows in the windowing result are subjected to Fast Welsh Leaf Transform to obtain the spectrum;

[0026] Each spectrum is input into a Mel filter bank for filtering to obtain the speech input features, and the reference fundamental frequency of the reference song is obtained according to the difference function.

[0027] Another objective of this invention is to provide a singing conversion system, the system comprising:

[0028] The boundary extraction module is used to extract the phoneme boundaries of the input speech and the reference song respectively, to obtain the input boundary and the reference boundary, and to obtain the speech input features of the input speech and the reference fundamental frequency of the reference song respectively.

[0029] The sampling module is used to sample the reference fundamental frequency according to the input boundary and the reference boundary to obtain the sampled fundamental frequency, and input the sampled fundamental frequency and the speech input features into the trained neural network vocoder;

[0030] The singing conversion module is used to perform singing conversion on the sampled fundamental frequency and the speech input features based on the trained neural network vocoder to obtain the singing output speech.

[0031] Preferably, the sampling module is further used for:

[0032] Perform duration matching on the phoneme durations in the input boundary and the reference boundary;

[0033] Based on the duration matching result, the phoneme duration of the input boundary, the phoneme duration of the reference boundary, and the reference fundamental frequency corresponding to the phoneme duration of the reference boundary are combined to obtain a fundamental frequency sampling group;

[0034] For each fundamental frequency sampling group, the corresponding reference fundamental frequency is upsampled or downsampled according to the phoneme duration of the input boundary to obtain the sampling fundamental frequency.

[0035] Another objective of this invention is to provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described above.

[0036] Another objective of this invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0037] In this embodiment of the invention, the reference fundamental frequency is sampled by the input boundary and the reference boundary. This can effectively utilize the phoneme boundaries of the reference song to directly control the fundamental frequency of the same phoneme boundary in the input speech without the need for modulation processing of the input speech. The trained neural network vocoder performs singing conversion on the sampled fundamental frequency and speech input features to achieve automatic singing conversion based on the fundamental frequency of the reference song, thereby improving the singing conversion effect. Attached Figure Description

[0038] Figure 1 This is a flowchart of the singing conversion method provided in the first embodiment of the present invention;

[0039] Figure 2 This is a flowchart of the singing conversion method provided in the second embodiment of the present invention;

[0040] Figure 3 This is a schematic diagram of the singing conversion system provided in the third embodiment of the present invention;

[0041] Figure 4 This is an implementation block diagram of the singing conversion system provided in the third embodiment of the present invention;

[0042] Figure 5 This is a schematic diagram of the structure of the terminal device provided in the fourth embodiment of the present invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0044] To illustrate the technical solution described in this invention, specific embodiments are described below.

[0045] Example 1

[0046] Please see Figure 1 This is a flowchart of a singing conversion method provided in the first embodiment of the present invention. This singing conversion method can be applied to any terminal device or system, and includes the following steps:

[0047] Step S10: Extract the phoneme boundaries of the input speech and the reference song respectively to obtain the input boundary and the reference boundary, and obtain the speech input features of the input speech and the reference fundamental frequency of the reference song respectively;

[0048] Among them, the HMM algorithm can be used to extract phoneme boundaries to obtain alignment results, and the alignment results of each phoneme can be merged to obtain phoneme boundary information. The reference fundamental frequency can be obtained by signal processing algorithm, and the speech input feature can be obtained by Mel spectrum.

[0049] Optionally, in this step, extracting the phoneme boundaries of the input speech and the reference song respectively includes:

[0050] The input speech and the reference song are phoneme aligned respectively to obtain a first alignment result and a second alignment result. Based on the first alignment result and the second alignment result, the alignment results of each phoneme are merged to obtain the input boundary and the reference boundary. The input boundary and the reference boundary store the duration of each phoneme. In this step, by merging the alignment results of each phoneme, the duration of each phoneme can be effectively obtained.

[0051] Step S20: Sample the reference fundamental frequency according to the input boundary and the reference boundary to obtain the sampled fundamental frequency, and input the sampled fundamental frequency and the speech input features into the trained neural network vocoder;

[0052] In this system, the phoneme sequences of the input and reference boundaries are identical, but the duration of each phoneme differs. By sampling the reference fundamental frequency, a sampling fundamental frequency with the same length as the input speech can be effectively obtained. The neural network vocoder recovers the speech from the Mel spectrum and is widely used in speech synthesis. For a neural network vocoder with controllable fundamental frequency, the main input includes not only the Mel spectrum but also the corresponding fundamental frequency. This ensures that the fundamental frequency of the output music signal is the same as that of the input speech, guaranteeing that the output music signal will not be out of tune and improving the accuracy of subsequent singing conversion.

[0053] Optionally, in this step, sampling the reference fundamental frequency based on the input boundary and the reference boundary includes:

[0054] Perform duration matching on the phoneme durations in the input boundary and the reference boundary;

[0055] Based on the duration matching result, the phoneme duration of the input boundary, the phoneme duration of the reference boundary, and the reference fundamental frequency corresponding to the phoneme duration of the reference boundary are combined to obtain a fundamental frequency sampling group;

[0056] For each fundamental frequency sampling group, the corresponding reference fundamental frequency is upsampled or downsampled according to the phoneme duration of the input boundary to obtain the sampling fundamental frequency;

[0057] Specifically, by matching the phoneme durations in the input boundary and the reference boundary to determine the phoneme boundaries with the same phoneme duration, the phoneme durations of the input boundary, the reference boundary, and the reference fundamental frequency corresponding to the phoneme durations of the reference boundary are combined to automatically generate fundamental frequency sampling groups. For each fundamental frequency sampling group, the corresponding reference fundamental frequency is upsampled or downsampled based on the phoneme duration of the input boundary. This effectively utilizes the phoneme boundaries of the reference song to directly control the fundamental frequency of the same phoneme boundary in the input speech.

[0058] Furthermore, before inputting the sampled fundamental frequency and the speech input features into the trained neural network vocoder, the method further includes:

[0059] The sample fundamental frequency and sample speech are input into the generator in the neural network vocoder, and speech is generated based on the sample fundamental frequency and sample speech by the generator to obtain the generated speech;

[0060] The generated speech is input into the multi-subband multi-scale discriminator in the neural network vocoder for speech discrimination, and the speech discrimination result is obtained.

[0061] Obtain the fundamental frequency of the sample and the singing standard audio corresponding to the sample speech, and determine the model loss based on the singing standard audio and the speech discrimination result;

[0062] The parameters of the neural network vocoder are updated according to the model loss until the neural network vocoder converges, thus obtaining the trained neural network vocoder.

[0063] The generator can be equipped with an adaptive filter, which effectively solves speech glitches and sustained articulation problems, improving speech quality. Through global and local multi-scale discriminators, it effectively improves the reconstruction of high and low frequencies, enhancing their quality.

[0064] Optionally, the step of generating speech based on the sample fundamental frequency and the sample speech using the generator includes:

[0065] The sample fundamental frequency and the sample speech are convolved according to each residual network in the generator, and the convolved sample fundamental frequency and the sample speech are upsampled according to the upsampling network in the generator.

[0066] The generator filters the upsampled sample fundamental frequency and the sample speech using an adaptive filter in the generator to obtain the generated speech.

[0067] Step S30: Perform singing conversion on the sampled fundamental frequency and the speech input features based on the trained neural network vocoder to obtain the singing output speech;

[0068] The system requires users to input voice data, which is then automatically combined with the singer's interpretation of the song's rhythm and key to perform a style conversion. Based on the converted voice output, it can automatically enable ordinary users to sing like professional singers.

[0069] In this embodiment, the reference fundamental frequency is sampled by the input boundary and the reference boundary. This can effectively utilize the phoneme boundaries of the reference song to directly control the fundamental frequency of the same phoneme boundary in the input speech without the need for modulation processing of the input speech. The trained neural network vocoder performs singing conversion on the sampled fundamental frequency and speech input features to achieve automatic singing conversion based on the fundamental frequency of the reference song, thereby improving the singing conversion effect.

[0070] Example 2

[0071] Please see Figure 2 This is a flowchart of a singing conversion method provided in the second embodiment of the present invention. This embodiment further refines step S10 in the first embodiment, including the following steps:

[0072] Step S11: The input speech is pre-emphasized, framed, and windowed sequentially, and each short-time analysis window in the windowing result is subjected to Fast Welsh Leaf Transform to obtain the spectrum;

[0073] The purpose of pre-emphasis is to boost the high-frequency components, flatten the signal spectrum, and maintain it across the entire frequency band from low to high, allowing for spectrum calculation with the same signal. Simultaneously, it aims to eliminate the effects of the vocal tract and lips during pronunciation, compensate for the high-frequency components suppressed by the speech system, and highlight the high-frequency formants.

[0074] Step S12: Input each spectrum into the Mel filter group for filtering to obtain the speech input features, and obtain the reference fundamental frequency of the reference song according to the difference function;

[0075] The extraction of the fundamental frequency can be broadly divided into time-domain and frequency-domain methods. The time-domain method takes the sound waveform as input, and its basic principle is to find the minimum positive period of the waveform. Of course, the periodicity of a real signal can only be approximated. The frequency-domain method first performs a Fourier transform on the signal to obtain the spectrum (only the amplitude spectrum is taken, and the phase spectrum is discarded). There will be peaks in the spectrum at integer multiples of the fundamental frequency; the basic principle of the frequency-domain method is to find the greatest common divisor of these peak frequencies.

[0076] This embodiment can automatically acquire voice input features and reference fundamental frequency, improving the efficiency of singing conversion.

[0077] Example 3

[0078] Please see Figure 3 This is a structural schematic diagram of the singing conversion system 100 provided in the third embodiment of the present invention, comprising, wherein:

[0079] The boundary extraction module 10 is used to extract the phoneme boundaries of the input speech and the reference song respectively, to obtain the input boundary and the reference boundary, and to obtain the speech input features of the input speech and the reference fundamental frequency of the reference song respectively.

[0080] Optionally, the boundary extraction module 10 is further configured to: perform phoneme alignment on the input speech and the reference song respectively to obtain a first alignment result and a second alignment result;

[0081] Based on the first alignment result and the second alignment result, the alignment results of each phoneme are merged to obtain the input boundary and the reference boundary.

[0082] Furthermore, the boundary extraction module 10 is also used to: sequentially pre-emphasize, frame, and window the input speech, and perform fast Welsh Leaf transform on each short-time analysis window in the windowing result to obtain the spectrum;

[0083] Each spectrum is input into a Mel filter bank for filtering to obtain the speech input features, and the reference fundamental frequency of the reference song is obtained according to the difference function.

[0084] The sampling module 11 is used to sample the reference fundamental frequency according to the input boundary and the reference boundary to obtain the sampled fundamental frequency, and input the sampled fundamental frequency and the speech input features into the trained neural network vocoder.

[0085] Optionally, the sampling module 11 is further configured to: perform duration matching on the phoneme durations in the input boundary and the reference boundary;

[0086] Based on the duration matching result, the phoneme duration of the input boundary, the phoneme duration of the reference boundary, and the reference fundamental frequency corresponding to the phoneme duration of the reference boundary are combined to obtain a fundamental frequency sampling group;

[0087] For each fundamental frequency sampling group, the corresponding reference fundamental frequency is upsampled or downsampled according to the phoneme duration of the input boundary to obtain the sampling fundamental frequency.

[0088] Furthermore, the sampling module 11 is also used to: input the sample fundamental frequency and sample speech into the generator in the neural network vocoder, and generate speech based on the sample fundamental frequency and sample speech to obtain generated speech;

[0089] The generated speech is input into the multi-subband multi-scale discriminator in the neural network vocoder for speech discrimination, and the speech discrimination result is obtained.

[0090] Obtain the fundamental frequency of the sample and the singing standard audio corresponding to the sample speech, and determine the model loss based on the singing standard audio and the speech discrimination result;

[0091] The parameters of the neural network vocoder are updated based on the model loss until the neural network vocoder converges, resulting in the trained neural network vocoder.

[0092] Furthermore, the sampling module 11 is also used to: perform convolution processing on the sample fundamental frequency and the sample speech according to each residual network in the generator, and perform upsampling on the convolution-processed sample fundamental frequency and the sample speech according to upsampling and other networks in the generator;

[0093] The generator filters the upsampled sample fundamental frequency and the sample speech using an adaptive filter in the generator to obtain the generated speech.

[0094] The singing conversion module 12 is used to perform singing conversion on the sampled fundamental frequency and the speech input features based on the trained neural network vocoder to obtain the singing output speech.

[0095] Please see Figure 4 The process involves extracting the phoneme boundaries and corresponding fundamental frequency (F0) of a reference song, while simultaneously extracting the acoustic features and phoneme boundaries of the input speech. The phoneme sequences on both sides are identical, but the duration of each phoneme differs. The F0 of the reference song is upsampled and downsampled according to the phoneme duration of the input speech to obtain a sampled fundamental frequency with the same length as the input speech. Then, the sampled fundamental frequency and the acoustic features of the input speech are input into a fundamental frequency-controllable neural network vocoder to obtain the song's timbre from the input speech.

[0096] In this embodiment, the reference fundamental frequency is sampled by the input boundary and the reference boundary. This can effectively utilize the phoneme boundaries of the reference song to directly control the fundamental frequency of the same phoneme boundary in the input speech without the need for modulation processing of the input speech. The trained neural network vocoder performs singing conversion on the sampled fundamental frequency and speech input features to achieve automatic singing conversion based on the fundamental frequency of the reference song, thereby improving the singing conversion effect.

[0097] Example 4

[0098] Figure 5 This is a structural block diagram of a terminal device 2 provided in the fourth embodiment of this application. For example... Figure 5As shown, the terminal device 2 in this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program for a singing conversion method. When the processor 20 executes the computer program 22, it implements the steps in each embodiment of the singing conversion method described above.

[0099] For example, the computer program 22 may be divided into one or more modules, which are stored in the memory 21 and executed by the processor 20 to complete this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, the processor 20 and the memory 21.

[0100] The processor 20 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0101] The memory 21 can be an internal storage unit of the terminal device 2, such as a hard drive or memory of the terminal device 2. The memory 21 can also be an external storage device of the terminal device 2, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal device 2. Furthermore, the memory 21 can include both internal and external storage units of the terminal device 2. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 can also be used to temporarily store data that has been output or will be output.

[0102] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0103] If an integrated module is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. This computer-readable storage medium can be non-volatile or volatile. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the contents of a computer-readable storage medium may be appropriately added to or subtracted from the contents as required by the legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, a computer-readable storage medium may not include electrical carrier signals and telecommunication signals.

[0104] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A singing conversion method, characterized in that, The method includes: The phoneme boundaries of the input speech and the reference song are extracted separately to obtain the input boundary and the reference boundary, and the speech input features of the input speech and the reference fundamental frequency of the reference song are obtained respectively. The reference fundamental frequency is sampled according to the input boundary and the reference boundary to obtain the sampled fundamental frequency, and the sampled fundamental frequency and the speech input features are input into the trained neural network vocoder. The trained neural network vocoder performs singing conversion on the sampled fundamental frequency and the speech input features to obtain the singing output speech. The step of sampling the reference fundamental frequency based on the input boundary and the reference boundary includes: performing duration matching on the phoneme durations in the input boundary and the reference boundary; Based on the duration matching result, the phoneme duration of the input boundary, the phoneme duration of the reference boundary, and the reference fundamental frequency corresponding to the phoneme duration of the reference boundary are combined to obtain a fundamental frequency sampling group; For each fundamental frequency sampling group, the corresponding reference fundamental frequency is upsampled or downsampled according to the phoneme duration of the input boundary to obtain the sampling fundamental frequency.

2. The singing conversion method as described in claim 1, characterized in that, Before inputting the sampled fundamental frequency and the speech input features into the trained neural network vocoder, the method further includes: The sample fundamental frequency and sample speech are input into the generator in the neural network vocoder, and speech is generated based on the sample fundamental frequency and sample speech by the generator to obtain the generated speech; The generated speech is input into the multi-subband multi-scale discriminator in the neural network vocoder for speech discrimination, and the speech discrimination result is obtained. Obtain the fundamental frequency of the sample and the singing standard audio corresponding to the sample speech, and determine the model loss based on the singing standard audio and the speech discrimination result; The parameters of the neural network vocoder are updated based on the model loss until the neural network vocoder converges, resulting in the trained neural network vocoder.

3. The singing conversion method as described in claim 2, characterized in that, The step of generating speech based on the sample fundamental frequency and the sample speech using the generator includes: The sample fundamental frequency and the sample speech are convolved according to each residual network in the generator, and the convolved sample fundamental frequency and the sample speech are upsampled according to the upsampling network in the generator. The generator filters the upsampled sample fundamental frequency and the sample speech using an adaptive filter in the generator to obtain the generated speech.

4. The singing conversion method as described in claim 1, characterized in that, The step of extracting the phoneme boundaries of the input speech and the reference song respectively includes: Phoneme alignment is performed on the input speech and the reference song respectively to obtain a first alignment result and a second alignment result; Based on the first alignment result and the second alignment result, the alignment results of each phoneme are merged to obtain the input boundary and the reference boundary.

5. The singing conversion method as described in any one of claims 1 to 4, characterized in that, The steps of obtaining the speech input features of the input speech and the reference fundamental frequency of the reference song respectively include: The input speech is pre-emphasized, framed, and windowed sequentially, and each short-time analysis window in the windowing result is subjected to Fast Fourier Transform to obtain the spectrum; Each spectrum is input into a Mel filter bank for filtering to obtain the speech input features, and the reference fundamental frequency of the reference song is obtained according to the difference function.

6. A singing conversion system, characterized in that, The system includes: The boundary extraction module is used to extract the phoneme boundaries of the input speech and the reference song respectively, to obtain the input boundary and the reference boundary, and to obtain the speech input features of the input speech and the reference fundamental frequency of the reference song respectively. The sampling module is used to sample the reference fundamental frequency according to the input boundary and the reference boundary to obtain the sampled fundamental frequency, and input the sampled fundamental frequency and the speech input features into the trained neural network vocoder; The step of sampling the reference fundamental frequency based on the input boundary and the reference boundary includes: performing duration matching on the phoneme durations in the input boundary and the reference boundary; combining the phoneme durations of the input boundary, the phoneme durations of the reference boundary, and the reference fundamental frequency corresponding to the phoneme durations of the reference boundary according to the duration matching result to obtain a fundamental frequency sampling group; and for each fundamental frequency sampling group, upsampling and downsampling the corresponding reference fundamental frequency according to the phoneme durations of the input boundary to obtain the sampled fundamental frequency. The singing conversion module is used to perform singing conversion on the sampled fundamental frequency and the speech input features based on the trained neural network vocoder to obtain the singing output speech.

7. The singing conversion system as described in claim 6, characterized in that, The sampling module is also used for: Perform duration matching on the phoneme durations in the input boundary and the reference boundary; Based on the duration matching result, the phoneme duration of the input boundary, the phoneme duration of the reference boundary, and the reference fundamental frequency corresponding to the phoneme duration of the reference boundary are combined to obtain a fundamental frequency sampling group; For each fundamental frequency sampling group, the corresponding reference fundamental frequency is upsampled or downsampled according to the phoneme duration of the input boundary to obtain the sampling fundamental frequency.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 5.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.