Song sound high correction method, device and equipment based on fundamental frequency modulation and medium

By extracting speech content sequences and timbre feature vectors, and combining fundamental frequency coding and conditional diffusion techniques, the problem of timbre preservation during pitch correction in existing technologies has been solved, achieving high-precision and natural pitch correction results.

CN122116874APending Publication Date: 2026-05-29HANGZHOU QUWEI SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU QUWEI SCI & TECH
Filing Date
2026-03-17
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve a natural, high-fidelity, and timbre-preserving effect during pitch correction, especially when making significant pitch adjustments, which can easily introduce mechanical sounds, electronic tones, and spectral distortion.

Method used

By extracting speech content sequences and timbre feature vectors from the original singing audio, combining them with the target fundamental frequency curve for fundamental frequency encoding, generating the target Mel spectrum using characteristic linear modulation and conditional diffusion, and finally reconstructing the audio waveform, the original timbre characteristics are maintained while adjusting the pitch.

Benefits of technology

It achieves high-precision and high-naturalness pitch correction, avoids timbre deviation, and improves the intelligence and versatility of vocal pitch processing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116874A_ABST
    Figure CN122116874A_ABST
Patent Text Reader

Abstract

The application provides a pitch correction method and device based on fundamental frequency modulation, equipment and medium, relates to the technical field of speech processing, and comprises the following steps: extracting a speech content sequence from an original song audio to be corrected; obtaining a timbre feature vector of a target speaker corresponding to the original song audio; performing fundamental frequency coding on a target fundamental frequency curve of a target song corresponding to the song audio to obtain a fundamental frequency condition feature of the target song; performing first feature linear modulation on the speech content sequence and the fundamental frequency condition feature to obtain a pitch mean condition; performing condition diffusion according to the pitch mean condition, the timbre feature vector and the fundamental frequency condition feature to generate a target mel-frequency spectrum; and performing audio waveform reconstruction according to the target mel-frequency spectrum to obtain a target song audio after pitch correction. The application can improve the pitch correction effect of the song.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and more specifically, to a method, apparatus, device, and medium for correcting singing pitch based on fundamental frequency modulation. Background Technology

[0002] In the field of vocal processing technology, pitch accuracy is a crucial element in improving the quality of vocal performances. Traditional automatic pitch adjustment technology, represented by Auto-Tune, is mainly based on digital signal processing, detecting the fundamental frequency of the audio and forcibly aligning it to the standard pitch. However, this type of method produces a noticeable mechanical and electronic sound when processing passages with large pitch deviations, especially severely damaging expressive elements such as vibrato and glissando, resulting in the processed vocals losing their natural qualities.

[0003] Another type of time-domain processing method, such as PSOLA (Pitch Synchronous Overlap and Add) and its variants, achieves pitch adjustment through pitch synchronization. However, when making large adjustments exceeding half a pitch interval, it introduces problems such as time-domain aliasing and spectral distortion, resulting in a tearing or metallic sound in the audio. Although deep learning-based speech conversion technology has made progress in recent years, most of it focuses on timbre transfer and lacks precise control over pitch.

[0004] Furthermore, in traditional methods, pitch and timbre information are tightly coupled, meaning that modifying pitch alters the timbre characteristics, making it difficult to maintain the original singer's timbre while adjusting pitch. Therefore, achieving natural, high-fidelity pitch correction that preserves timbre has become a pressing technical challenge in this field. Summary of the Invention

[0005] This application provides a method, apparatus, device, and medium for vocal pitch correction based on fundamental frequency modulation, which can improve the vocal pitch correction effect.

[0006] In a first aspect, embodiments of this application provide a method for correcting the pitch of singing voices based on fundamental frequency modulation, the method comprising: Extract the speech content sequence from the original singing audio to be corrected; Obtain the timbre feature vector of the target speaker corresponding to the original singing audio; The fundamental frequency curve of the target song corresponding to the original singing audio is encoded to obtain the fundamental frequency conditional features of the target song; The speech content sequence and the fundamental frequency condition feature are subjected to first feature linear modulation to obtain the pitch mean condition; Based on the pitch mean condition, the timbre feature vector, and the fundamental frequency condition feature, conditional diffusion is performed to generate the target Mel spectrum; The target Mel spectrum is used to reconstruct the audio waveform, and the pitch-corrected target singing audio is obtained.

[0007] Optionally, the step of extracting the speech content sequence from the original singing audio to be corrected includes: Log-Melbourne spectrum extraction is performed on the original singing audio to obtain log-Melbourne spectrum features; The log-Mel spectrum features are encoded to obtain the speech content sequence.

[0008] Optionally, obtaining the timbre feature vector of the target speaker corresponding to the original singing audio further includes: Speaker features are extracted from the original singing audio to obtain the timbre feature vector; or, Speaker features are extracted from the reference audio of the target speaker to obtain the timbre feature vector.

[0009] Optionally, the step of fundamental frequency encoding the fundamental frequency curve of the target song corresponding to the original singing audio to obtain the fundamental frequency conditional features of the target song includes: Quantize and encode multiple consecutive fundamental frequency values ​​in the target fundamental frequency curve to obtain a fundamental frequency level sequence with multiple discrete levels; The fundamental frequency range sequence is subjected to fundamental frequency embedding to obtain a fundamental frequency embedding vector sequence; Feature extraction is performed on the fundamental frequency embedding vector sequence to obtain the fundamental frequency conditional features.

[0010] Optionally, the fundamental frequency condition features include: hidden layer condition features and Mel dimension condition features; The step of performing first feature linear modulation on the speech content sequence and the fundamental frequency conditional features to obtain the pitch mean condition includes: The speech content sequence and the hidden layer conditional features are linearly modulated using the first feature to obtain the pitch mean condition; The step of generating the target Mel spectrum by performing conditional diffusion based on the pitch mean condition, the timbre feature vector, and the fundamental frequency condition feature includes: The target Mel spectrum is generated by performing conditional diffusion based on the pitch mean condition, the timbre feature vector, and the Mel dimension conditional feature.

[0011] Optionally, the step of performing first feature linear modulation on the speech content sequence and the hidden layer conditional features to obtain the pitch mean condition includes... The hidden layer conditional features are downsampled to the same temporal resolution as the speech content sequence to obtain the downsampled conditional features. The speech content sequence and the downsampled conditional features are linearly modulated using a first feature to obtain the first modulated features. The first modulated feature is semantically encoded to obtain the pitch mean condition.

[0012] Optionally, the step of generating the target Mel spectrum by performing conditional diffusion based on the pitch mean condition, the timbre feature vector, and the Mel dimension conditional feature includes: Upsampling is performed on the pitch mean condition to obtain the upsampling condition features; The upsampled conditional features and the Mel dimension conditional features are subjected to second feature linear modulation to obtain the second modulated features; Conditional diffusion decoding is performed based on the second modulated feature, the timbre feature vector, and the Mel dimension conditional feature to obtain the target Mel spectrum.

[0013] Secondly, embodiments of this application also provide a vocal pitch correction device based on fundamental frequency modulation, the device comprising: The extraction module is used to extract the speech content sequence from the original singing audio to be corrected; and to obtain the timbre feature vector of the target speaker corresponding to the original singing audio. The encoding module is used to perform fundamental frequency encoding on the target fundamental frequency curve of the target song corresponding to the singing audio, so as to obtain the fundamental frequency conditional features of the target song; The modulation module is used to perform first feature linear modulation on the speech content sequence and the fundamental frequency condition feature to obtain the pitch mean condition; The diffusion module is used to perform conditional diffusion based on the pitch mean condition, the timbre feature vector, and the fundamental frequency condition feature to generate the target Mel spectrum; The reconstruction module is used to reconstruct the audio waveform based on the target Mel spectrum to obtain the target singing audio after pitch correction.

[0014] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a bus, wherein the memory stores program instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the program instructions to perform the steps of the vocal pitch correction method based on fundamental frequency modulation as described in any of the first aspects.

[0015] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the vocal pitch correction method based on fundamental frequency modulation as described in any of the first aspects.

[0016] The method, apparatus, device, and medium for vocal pitch correction based on fundamental frequency modulation provided in this application first extract a speech content sequence representing the semantic content of the speech from the original vocal audio to be corrected, preserving the original lyrics pronunciation information for subsequent generation; simultaneously, it obtains the timbre feature vector of the target speaker corresponding to the original vocal audio, to constrain the timbre features to remain unchanged during the pitch correction process; and fundamental frequency encoding is performed on the target fundamental frequency curve of the target song to obtain fundamental frequency conditional features containing pitch information, which serve as the basis for pitch control; subsequently, the first feature linear modulation is performed on the speech content sequence and the fundamental frequency conditional features to obtain the pitch mean condition; then, conditional diffusion generation is performed based on the pitch mean condition, the timbre feature vector, and the fundamental frequency conditional features to generate a target Mel spectrum that conforms to the target pitch and maintains the original timbre; finally, the audio waveform is reconstructed based on the target Mel spectrum to obtain the pitch-corrected target vocal audio, which can accurately correct the pitch while maintaining the original singer's timbre characteristics. This method effectively decouples pitch and timbre, allowing for precise adjustment of vocal accuracy while fully preserving the original singer's timbre characteristics. This fundamentally avoids the timbre shift problem caused by pitch changes in traditional methods. By performing feature linear modulation on the speech content sequence and fundamental frequency conditional features, target pitch information can be injected into the content representation, ensuring that the generation process strictly follows the direction of the target fundamental frequency curve. This enables refined control over various pitch-related tasks such as out-of-tune singing and vibrato, overcoming the mechanical feel, electronic sound, and spectral distortion and tearing issues caused by time-domain processing in traditional signal processing methods when handling large pitch adjustments. It not only achieves high-precision and high-naturalness pitch correction but also significantly improves the intelligence level and versatility of vocal pitch processing technology. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic flowchart illustrating a vocal pitch correction method based on fundamental frequency modulation provided in this application embodiment; Figure 2 A flowchart illustrating the process of obtaining a speech content sequence in a fundamental frequency modulation-based vocal pitch correction method provided in this application embodiment; Figure 3 A flowchart illustrating the fundamental frequency condition characteristics in a fundamental frequency modulation-based vocal pitch correction method provided in this application embodiment; Figure 4A schematic diagram of the process for obtaining the pitch mean condition in a fundamental frequency modulation-based singing pitch correction method provided in an embodiment of this application; Figure 5 A schematic diagram illustrating the process of generating the target Mel spectrum in a fundamental frequency modulation-based vocal pitch correction method provided in this application embodiment; Figure 6 A schematic diagram of a vocal pitch correction device based on fundamental frequency modulation provided in this application embodiment; Figure 7 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0020] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0021] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0022] Before providing a detailed explanation of this application, let's first introduce its application scenarios.

[0023] In the professional music production field, even experienced singers inevitably experience pitch deviations during vocal recording, requiring pitch correction in post-production. Recording engineers typically sift through vast amounts of recording material to select the most emotionally expressive segments that may have individual pitch issues, then use professional software to correct each note individually—a process highly reliant on manual operation and extremely time-consuming. In scenarios geared towards ordinary users, such as short video creation, online karaoke, and live streaming, user-recorded covers often exhibit more pronounced pitch problems. Existing automatic pitch correction tools either produce noticeable mechanical electronic sounds that disrupt the naturalness of the sound or perform poorly with complex melodies, failing to meet the demands of ordinary users for quick, high-quality output. In professional audio production fields such as film and television dubbing, animation production, and game character voice acting, it is often necessary to adjust the pitch of voice actors' recorded vocals according to a predetermined melody line, or to convert the singing to different keys and modes while maintaining the original timbre. Existing technologies often introduce timbre distortion when making significant pitch adjustments, resulting in inconsistencies in the character's voice.

[0024] Based on this, this application provides a method, apparatus, device, and medium for vocal pitch correction based on fundamental frequency modulation. First, a speech content sequence representing the semantic content of the speech is extracted from the original vocal audio to be corrected, preserving the original lyrics pronunciation information for subsequent generation. Simultaneously, the timbre feature vector of the target speaker corresponding to the original vocal audio is obtained to constrain the timbre features to remain unchanged during pitch correction. The target fundamental frequency curve of the target song is then fundamentally encoded to obtain fundamental frequency conditional features containing pitch information, serving as the basis for pitch control. Subsequently, the speech content sequence and fundamental frequency conditional features are linearly modulated using a first feature to obtain the pitch mean condition. Then, conditional diffusion generation is performed based on the pitch mean condition, timbre feature vector, and fundamental frequency conditional features to generate a target Mel spectrum that conforms to the target pitch and maintains the original timbre. Finally, audio waveform reconstruction is performed based on the target Mel spectrum to obtain the pitch-corrected target vocal audio. This method can accurately correct pitch while maintaining the original singer's timbre characteristics, achieving a high-precision, high-naturalness pitch correction effect, and significantly improving the intelligence and versatility of vocal pitch processing technology.

[0025] The following description, in conjunction with the accompanying drawings, illustrates several embodiments. The vocal pitch correction method based on fundamental frequency modulation provided in this application can be implemented by a computer device pre-installed with a preset vocal pitch correction algorithm or software based on fundamental frequency modulation, through the execution of the algorithm or software. The computer device can be, for example, a server or a terminal, and the terminal can be a user computer.

[0026] Figure 1 A flowchart illustrating a vocal pitch correction method based on fundamental frequency modulation, as provided in this application embodiment, is shown below. Figure 1 As shown, this method for correcting vocal pitch based on fundamental frequency modulation includes: S101, extract the speech content sequence from the original singing audio to be corrected.

[0027] Among them, the original singing audio is a human singing audio with pitch deviation that needs to be pitch corrected. It can include pure human voice audio or singing audio with background music, reverberation and other environmental sound effects. The speech content sequence is a feature sequence that represents the semantic and pronunciation information in the original singing audio. The speech content sequence is independent of the pitch and timbre information of the audio and only retains the content features of the singing.

[0028] Specifically, after preprocessing the original singing audio, a preset speech content extraction model is used to encode its features, converting the continuous audio signal into a discrete feature sequence, thereby obtaining a speech content sequence that can represent the semantics and pronunciation of the singing.

[0029] In one possible implementation, the original vocal audio is first resampled to 16kHz to meet the input requirements of a preset speech content encoder. The resampled audio is then input into a preset speech content extraction model, outputting a discrete token sequence as the speech content sequence. The frame rate of the speech content sequence can be set to 50 frames per second, meaning one token corresponds to every 20 milliseconds, and each token encodes the speech semantics and pronunciation information within a corresponding time window.

[0030] S102, obtain the timbre feature vector of the target speaker corresponding to the original singing audio.

[0031] In the process of vocal pitch correction, it is necessary to maintain the singer's original timbre characteristics while adjusting the pitch. To this end, it is necessary to extract feature vectors that can characterize the singer's timbre.

[0032] In this context, the target speaker is the singer of the original singing audio, and the timbre feature vector is a feature vector that represents the timbre characteristics of the target speaker. It is used to keep the timbre of the singer unchanged during the pitch correction process and avoid timbre distortion and characteristic shift in the corrected singing.

[0033] Specifically, the audio data of the target speaker is extracted and vectorized using a preset speaker feature extraction model to obtain a fixed-dimensional timbre feature vector. This vector, after preprocessing such as normalization, is used as a conditional input for the subsequent generation process to ensure that the generated corrected audio is consistent with the timbre features of the original singing.

[0034] S103, perform fundamental frequency encoding on the target fundamental frequency curve of the target song corresponding to the original singing audio to obtain the fundamental frequency conditional features of the target song.

[0035] Among them, the target fundamental frequency curve is the standard pitch curve corresponding to the target song, which is a continuous fundamental frequency value sequence aligned with the time sequence of the original singing audio; the fundamental frequency condition feature is the feature obtained after encoding the target fundamental frequency curve, which can convert continuous pitch information into a feature form suitable for model processing.

[0036] Specifically, the first step is to obtain the target fundamental frequency curve of the target song. In one possible implementation, the target fundamental frequency curve can be generated based on the song's standard score or MIDI (Musical Instrument Digital Interface) file, with each moment corresponding to a standard pitch value; alternatively, the correct pitch reference can be extracted from a reference version of the original vocal audio through manual annotation or an automatic pitch detection algorithm. Secondly, the obtained target fundamental frequency curve undergoes fundamental frequency encoding processing, converting the time-series fundamental frequency information into a high-dimensional feature vector to obtain the fundamental frequency conditional features of the target song. These fundamental frequency conditional features characterize the standard pitch patterns of the target song.

[0037] It should be noted that there is no strict time dependency between the execution of S101 to S103 above, and they can be flexibly adjusted or executed in parallel according to the actual application scenario.

[0038] Specifically, S101 is used to extract the semantic content information of the original singing audio, S102 is used to obtain the timbre feature vector of the target speaker, and S103 is used to encode the target fundamental frequency curve to obtain the fundamental frequency conditional feature. These three types of features correspond to the content constraints, timbre constraints, and pitch constraints in the generation process, respectively, and are independent of each other. Therefore, in actual implementation, S101 can be executed first, followed by S102 and then S103, or three parallel processing threads can be started simultaneously to extract the three types of features respectively, as long as the speech content sequence, timbre feature vector, and fundamental frequency conditional feature are all ready before entering S104 for the first feature linear modulation.

[0039] S104, perform first feature linear modulation on the speech content sequence and fundamental frequency conditional features to obtain the pitch mean condition.

[0040] Among them, the first feature-wise linear modulation (FiLM) is a feature modulation method that injects the pitch information of the fundamental frequency conditional features into the speech content sequence, so that the speech content sequence carries the constraint information of the target pitch; the pitch mean condition is the feature obtained after the first feature-wise linear modulation, which is the mean constraint condition of the subsequent conditional diffusion generation process.

[0041] Specifically, the speech content sequence and fundamental frequency conditional features are input into a predefined linear modulation model. In one possible implementation, a scaling factor is generated based on the fundamental frequency conditional features. and offset factor An affine transformation is performed on the speech content sequence. The scaling factor is... The offset factor is used to control the amplification or reduction of each feature dimension in the speech content sequence. The mathematical expression for the affine transformation, used to control the translation amount of each feature dimension, is as follows: ,in, This is a feature representation of a speech content sequence. The conditions for the average pitch value obtained after modulation.

[0042] By using linear modulation of the first feature, the fundamental frequency conditional feature is embedded into the semantic representation of the speech content sequence, so that the output pitch mean condition not only retains the original semantic content information, but also carries the guidance information of the target pitch.

[0043] S105 generates the target Mel spectrum by performing conditional diffusion based on the pitch mean condition, timbre feature vector, and fundamental frequency condition feature.

[0044] After obtaining the pitch mean condition, timbre feature vector, and fundamental frequency condition feature, a generative framework based on the conditional diffusion model is adopted to gradually denoise from random noise to generate the target Mel spectrum that meets all conditional constraints.

[0045] Among them, the target Mel spectrum is a Mel spectrum feature that matches the standard pitch of the target song and the timbre of the target speaker, and its dimensions are adapted to the needs of subsequent audio waveform reconstruction.

[0046] Specifically, a diffusion generation architecture based on flow matching or conditional flow model (CFM) is adopted, enabling the learning of a deterministic transformation from noise distribution to data distribution, resulting in higher sampling efficiency compared to traditional denoising diffusion probability models. Pitch mean condition, timbre feature vector, and fundamental frequency condition feature are input. Through multi-condition fusion and progressive generation of the diffusion model, the initial noise features are gradually transformed into Mel spectrum features that conform to pitch and timbre constraints. Finally, a Mel spectrum that matches the target pitch, preserves the target timbre, and is consistent with the original semantic content is generated—the target Mel spectrum. The target Mel spectrum matches the standard pitch of the target song while retaining the unique timbre of the target speaker.

[0047] S106, reconstruct the audio waveform based on the target Mel spectrum to obtain the target singing audio after pitch correction.

[0048] Among them, audio waveform reconstruction is the process of converting Mel spectrum features into audio waveform signals; the target singing audio is the final singing audio after pitch correction. The target singing audio eliminates the pitch deviation of the original singing audio and maintains the content information and timbre characteristics of the original singing, possessing high naturalness and high fidelity acoustic characteristics.

[0049] Specifically, a neural network-based vocoder model is used to reconstruct audio waveforms from the Mel spectrum. This application employs a neural network vocoder based on the HiFi-GAN architecture, which is trained using a generative adversarial network (GAN) and is capable of reconstructing audio waveforms with a 24kHz sampling rate from the Mel spectrum. It should be noted that the audio sampling rate can be configured according to the actual application scenario (such as online playback or professional production) to meet the audio usage needs of different scenarios.

[0050] In one possible implementation, the HiFi-GAN vocoder generator employs a fully convolutional neural network structure, consisting of multiple upsampling modules and a multi-receptive field fusion (MRF) module cascaded together.

[0051] During the inference phase, the target Mel spectrum is input into the generator of the HiFi-GAN vocoder. After forward propagation calculation, it outputs a 24kHz sampling rate audio waveform of the corresponding time length. This vocoder supports real-time streaming inference, meaning it can synthesize audio block by block after receiving Mel spectrum frames of sufficient length, meeting the low latency requirements of online audio processing scenarios.

[0052] In the above manner, the target Mel spectrum is converted into a time-domain waveform, which is the target vocal audio after pitch correction. Its pitch has been adjusted to the correct position specified by the target fundamental frequency curve, while fully preserving the original singer's timbre characteristics and singing details, achieving natural and high-quality pitch correction without mechanical feel.

[0053] In this embodiment, effective decoupling of pitch and timbre can be achieved, allowing for precise adjustment of vocal accuracy while fully preserving the original singer's timbre characteristics. This fundamentally avoids the timbre shift problem caused by pitch changes in traditional methods. By performing feature linear modulation on the speech content sequence and fundamental frequency condition features, target pitch information can be injected into the content representation, ensuring that the generation process strictly follows the direction of the target fundamental frequency curve. This achieves refined control over various pitch-related tasks such as out-of-tune singing and vibrato, overcoming the sound quality degradation problems caused by mechanical sounds, electronic sounds, and spectral distortion and tearing caused by time-domain processing when traditional signal processing methods handle large pitch adjustments. This not only achieves high-precision and high-naturalness pitch correction effects but also significantly improves the intelligence level and versatility of vocal pitch processing technology.

[0054] In the above Figure 1 Based on the corresponding embodiments, in order to more clearly demonstrate the process of obtaining the speech content sequence, this application also provides a possible implementation of obtaining the speech content sequence in the fundamental frequency modulation-based vocal pitch correction method. Figure 2 This is a schematic flowchart illustrating the process of obtaining a speech content sequence in a fundamental frequency modulation-based vocal pitch correction method provided in an embodiment of this application. Figure 2 As shown, S101 above, which extracts the speech content sequence from the original singing audio to be corrected, includes: S210, Log-Mel spectrum extraction is performed on the original singing audio to obtain the log-Mel spectrum features.

[0055] Among them, the log-Mel spectrum feature is a frequency domain acoustic feature designed based on the characteristics of human hearing, which can accurately reproduce the frequency distribution and acoustic details of the original singing audio.

[0056] Specifically, the original vocal audio is resampled to 16kHz, and then the audio signal is processed by frame segmentation, windowing, and short-time Fourier transform to obtain a linear spectrum. This linear spectrum is then filtered using a Mel filter bank of 128 elements. Finally, the logarithm of the filtered energy is taken to obtain a 128-dimensional log-Mel spectral feature. Each frame of audio corresponds to a 128-dimensional feature vector, and the feature vectors of all frames are arranged in chronological order to form a log-Mel spectral feature sequence.

[0057] S220, the log-Mel spectrum features are encoded to obtain the speech content sequence.

[0058] Specifically, the log-Mel spectrum features are input into a pre-defined speech content encoder based on a convolutional neural network and vector quantization. After processing by the encoder, the log-Mel spectrum features are converted into a discrete token sequence, i.e., a speech content sequence. The frame rate of this sequence is reduced to 50 frames per second, i.e., one token every 20 milliseconds. These tokens encode the speech semantics and pronunciation information within the corresponding time window, thereby achieving decoupling from timbre and pitch information.

[0059] In this embodiment, by performing log-Mel spectrum extraction and speech content encoding on the original singing audio, the continuous audio signal is converted into a discrete speech content sequence. This effectively filters out speaker-related details such as pitch and timbre, achieving decoupling of content from pitch and timbre. It can obtain a pure content representation while preserving the original lyrics semantics and pronunciation structure, ensuring that the content remains unchanged during subsequent pitch correction and avoiding content distortion or semantic confusion caused by pitch adjustment.

[0060] In the above Figure 1Based on the corresponding embodiments, to more clearly demonstrate the process of obtaining the timbre feature vector, this application also provides a possible implementation of obtaining the timbre feature vector in a fundamental frequency modulation-based singing pitch correction method. Optionally, in the above S102, obtaining the timbre feature vector of the target speaker corresponding to the original singing audio includes: S310, Speaker features are extracted from the original singing audio to obtain the timbre feature vector.

[0061] Specifically, a speaker feature extraction model based on deep neural networks is used to extract fixed-dimensional timbre feature vectors from the original singing audio. This application uses the CAM++ speaker verification model as the speaker feature extractor.

[0062] In one possible implementation, the original singing audio is resampled to 16kHz and then input into the CAM++ model. Finally, a 192-dimensional speaker embedding vector is output as a timbre feature vector. The timbre feature vector encodes the timbre features in the original singing audio, including formant position, vocal tract characteristics, vocal habits and other voiceprint information.

[0063] After obtaining the 192-dimensional timbre feature vector, it needs to be normalized to a unit length to eliminate the influence of amplitude differences on subsequent conditional injection. Optionally, the normalized vector can be further mapped to a feature dimension adapted to the generative model through a linear projection layer, and then used as conditional information input into the subsequent generation module.

[0064] By using the above method, feature vectors representing the timbre characteristics of the singer are extracted from the original singing audio, providing timbre constraints for maintaining the original singer's timbre while adjusting the pitch.

[0065] S320: Extract speaker features from the reference audio of the target speaker to obtain a timbre feature vector.

[0066] In certain applications, it is necessary to convert the timbre of the original singing audio into the timbre of a specific target singer. For example, when a user wants to convert their cover song audio into the timbre of a professional singer, it is necessary to extract the timbre features of that professional singer as a generation condition.

[0067] Specifically, the implementation of S320 is basically the same as that of S310, the difference being the source of the input audio. S310 uses the original singing audio itself as input to extract the timbre features of the original singer; while S320 uses the reference audio of the target speaker as input to extract the timbre features of the target singer. The reference audio can be a short a cappella recording of the target singer, or any speech or singing segment that can clearly characterize the timbre of the target singer. The reference audio is input into the same CAM++ speaker feature extraction model as S310, which also extracts a 192-dimensional speaker embedding vector. After normalization and linear projection, the timbre feature vector of the target speaker is obtained.

[0068] It should be noted that S310 and S320 are two parallel optional implementations, and either one can be executed. When the application requirement is to maintain the original singer's timbre, S310 is executed to extract timbre features from the original vocal audio; when the application requirement is to convert the vocals to the timbre of a specific target singer, S320 is executed to extract timbre features from the target singer's reference audio. Both methods yield a fixed-dimensional timbre feature vector, which serves as the timbre conditional input in the subsequent conditional diffusion generation process.

[0069] In this embodiment, by setting two parallel optional implementation methods—extracting from the original vocal audio to preserve the original singer's timbre, or extracting from the target singer's reference audio to achieve timbre conversion—this application can flexibly adapt to various application scenarios: it can be used for pitch correction that preserves the original singer's personal characteristics, and it can also be used for style transfer that converts ordinary user covers into professional singer timbre. By using this method, timbre characteristics are preserved or controllably transformed during pitch correction, providing reliable timbre constraints for subsequently generating corrected audio that matches the target pitch and has a natural timbre, effectively solving the technical problem of timbre distortion accompanying pitch adjustment in traditional methods.

[0070] In the above Figure 1 Based on the corresponding embodiments, in order to more clearly demonstrate the process of obtaining the speech content sequence, this application also provides a possible implementation of obtaining the fundamental frequency condition features of the speech content sequence in the fundamental frequency modulation-based singing pitch correction method. Figure 3 This is a flowchart illustrating the fundamental frequency condition characteristics in a vocal pitch correction method based on fundamental frequency modulation, provided as an embodiment of this application. Figure 3 As shown, in S103 above, the fundamental frequency curve of the target song corresponding to the original singing audio is encoded to obtain the fundamental frequency conditional features of the target song, including: S410 quantizes and encodes multiple consecutive fundamental frequency values ​​in the target fundamental frequency curve to obtain a fundamental frequency range sequence with multiple discrete ranges.

[0071] Among them, quantization encoding is a process of mapping continuously distributed fundamental frequency values ​​to preset discrete levels. The fundamental frequency level sequence is a discrete pitch representation sequence aligned with the timing of the original singing audio, which can accurately correspond to the target pitch state of each audio frame.

[0072] Specifically, the target fundamental frequency curve of the target song is first obtained. Optionally, a deep learning-based fundamental frequency estimation algorithm (such as RMVPE) is used to extract the fundamental frequency curve, thereby accurately estimating the fundamental frequency of human voices in various complex musical environments (such as those with background music, reverberation, etc.) and correctly identifying silent segments (F0=0) and vocal segments. Then, quantization encoding mapping is performed on different types of audio frames to convert continuous fundamental frequency values ​​into integer codes of corresponding discrete levels, ultimately obtaining a temporally continuous and level-uniform fundamental frequency level sequence.

[0073] In one possible implementation, 256 discrete encoding levels are preset, where level 0 is assigned to silent frames, corresponding to the case where F0=0; levels 1 to 255 are assigned to frames with sound, which are evenly divided in the logarithmic frequency space corresponding to the fundamental frequency range of human voice from 50Hz to 1100Hz, covering about 4.5 octaves; for frames with sound, the fundamental frequency value is first converted to logarithmic space, normalized to the range of [0,1] and then mapped to the integer level of 1-255, and silent frames are directly encoded as level 0.

[0074] Using the above method, the fundamental frequency value at each moment in the target fundamental frequency curve is quantized and encoded to obtain a discrete frequency range sequence corresponding to the original singing audio time frame, i.e., the fundamental frequency range sequence. Each element in the sequence is an integer between 0 and 255, corresponding to silence or different pitch levels.

[0075] S420 performs baseband embedding on the baseband position sequence to obtain the baseband embedding vector sequence.

[0076] Specifically, the discrete baseband position sequence is input into a preset baseband embedding layer. The embedding layer performs vectorization mapping on each discrete position, generating a unique fixed-dimensional dense feature vector for each position, and finally obtaining a baseband embedding vector sequence that is time-aligned with the target baseband curve.

[0077] In one possible implementation, the fundamental frequency embedding layer outputs a 256-dimensional fundamental frequency embedding vector for each gear, thereby mapping discrete gear information to a continuous vector space and ensuring a complete representation of pitch information.

[0078] S430 extracts features from the fundamental frequency embedding vector sequence to obtain fundamental frequency conditional features.

[0079] Specifically, the fundamental frequency embedding vector sequence is input into a preset fundamental frequency encoder network. The encoder network performs deep feature extraction and context association modeling on the temporally sequenced embedding vectors, and finally outputs fundamental frequency conditional features that meet the dimensional requirements of subsequent processing.

[0080] In one possible implementation, the fundamental frequency encoder network consists of one-dimensional convolutional layers and residual blocks. After feature extraction, it outputs two fundamental frequency conditional features of different dimensions: a 512-dimensional hidden layer conditional feature adapted for the encoder stage processing, and an 80-dimensional Mel-spectrum conditional feature adapted for Mel-spectrum processing. The hidden layer conditional feature maintains the same temporal resolution as the input and is used for feature linear modulation in the encoder stage. The Mel-dimensional conditional feature undergoes additional dimensionality reduction projection, maintaining a temporal resolution consistent with the subsequently generated Mel-spectrum, and is used for feature linear modulation and conditional input in the decoder stage. These two output dimensions adapt to the conditional injection requirements of different stages, providing control information for subsequent pitch conditional modulation.

[0081] In this embodiment, the continuous fundamental frequency is converted into discrete pitch levels through quantization encoding, realizing the pitch representation of the entire range; the pitch information is converted into the feature form of the adaptation model through embedding and feature extraction, providing reliable support for the subsequent injection of pitch conditions and ensuring the accuracy of pitch control.

[0082] In the above Figure 1 Based on the corresponding embodiments, optionally, the fundamental frequency conditional features include: hidden layer conditional features and Mel dimension conditional features. In S104 above, the speech content sequence and the fundamental frequency conditional features are linearly modulated using the first feature to obtain the pitch mean condition, which includes: S510 performs first feature linear modulation on the speech content sequence and hidden layer conditional features to obtain the pitch mean condition.

[0083] Among them, the hidden layer conditional features are the feature dimensions in the fundamental frequency conditional features that have the same temporal resolution as the speech content sequence, provided by the 512-dimensional hidden layer conditional features output by the fundamental frequency encoder. The implementation method of the first feature linear modulation is as described in S104 above, that is, a scaling factor is generated based on the conditional information. and offset factor Affine transformation is performed on the input features.

[0084] Specifically, the speech content sequence and hidden layer conditional features are input into a feature-linear modulation model. In one possible implementation, a scaling factor is first generated based on the hidden layer conditional features. Sequence and offset factor sequence, and The dimension of the feature is consistent with the feature dimension of the speech content sequence; then the speech content sequence is compared with... Element-wise multiplication is used to scale and modulate the features; then the scaled features are multiplied by... Element-wise addition is used to achieve feature offset modulation. Through this affine transformation process, the target pitch information carried by the hidden layer conditional features is embedded into the semantic representation of the speech content sequence.

[0085] Through linear modulation of the first feature, the pitch information of the fundamental frequency conditional feature is embedded into the semantic representation of the speech content sequence, effectively realizing the fusion of content information and pitch information. This ensures that the output pitch mean condition not only fully preserves the semantic and pronunciation information of the original singing voice, but also provides pitch guidance for subsequent conditional diffusion generation, guaranteeing the processing requirements of unchanged content and accurate pitch during pitch correction.

[0086] Optionally, in S105 above, conditional diffusion is performed based on the pitch mean condition, timbre feature vector, and fundamental frequency condition feature to generate the target Mel spectrum, which includes: S520 generates the target Mel spectrum by performing conditional diffusion based on the pitch mean condition, timbre feature vector, and Mel dimension conditional features.

[0087] Among them, the Mel dimension conditional feature is the feature dimension in the fundamental frequency conditional feature that has the same time resolution as the Mel spectrum to be generated. It is provided by the 80-dimensional Mel dimension conditional feature output by the fundamental frequency encoder and is used for feature linear modulation and conditional input in the decoder stage.

[0088] Specifically, after adapting the pitch mean condition, timbre feature vector, and Mel dimension condition features for feature dimension and aligning them temporally, they are jointly input into a diffusion generation architecture based on flow matching or conditional flow models. The model achieves deep fusion of multiple condition features through cross-attention and self-attention mechanisms. Starting from initial random noise, it gradually maps to the target data distribution through deterministic transformation. Under the global pitch constraint of the pitch mean condition, the fine pitch guidance of the Mel dimension condition features, and the fixed timbre constraint of the timbre feature vector, the target Mel spectrum with temporal alignment, accurate pitch, and timbre matching is finally generated.

[0089] In this embodiment, by distinguishing the fundamental frequency conditional features into hidden layer conditional features and Mel dimension conditional features, different requirements of the encoder and decoder stages are met respectively, realizing hierarchical and refined control of fundamental frequency information. In the encoder stage, the first feature is linearly modulated by the speech content sequence and the hidden layer conditional features, so that the target pitch information is effectively embedded in the content representation to form the pitch mean condition. In the decoder stage, the pitch mean condition, timbre feature vector and Mel dimension conditional features are fused, and the target Mel spectrum is generated by the conditional diffusion decoder. This significantly improves the accuracy and robustness of pitch control, ensures the effective preservation of timbre features during pitch adjustment, and provides key technical support for high-quality and natural singing pitch correction.

[0090] Based on the embodiment corresponding to S510 above, in order to more clearly demonstrate the process of obtaining the pitch mean condition, this application also provides a possible implementation of obtaining the pitch mean condition in a fundamental frequency modulation-based singing pitch correction method. Figure 4 This is a schematic flowchart illustrating the process of obtaining the pitch mean condition in a fundamental frequency modulation-based vocal pitch correction method provided in an embodiment of this application. Figure 4 As shown, in S510 above, the first feature linear modulation is performed on the speech content sequence and the hidden layer conditional features to obtain the pitch mean condition, which includes: S610, the hidden layer conditional features are downsampled to the same temporal resolution as the speech content sequence to obtain the downsampled conditional features.

[0091] In practical applications, although the speech content sequence and the hidden layer conditional features have the same temporal resolution, their frame rate definitions may differ, or the original frame rate of the hidden layer conditional features may be higher than that of the speech content sequence. Therefore, temporal alignment processing is required to ensure that the feature linear modulation is accurately matched in the frame-by-frame correspondence.

[0092] Specifically, the hidden layer conditional features are downsampled to ensure their temporal resolution matches that of the speech content sequence. In one possible implementation, the speech content sequence has a frame rate of 50 frames per second (one frame every 20 milliseconds), and the original frame rate of the hidden layer conditional features is also 50 frames per second; they are naturally aligned and can be used directly. If the hidden layer conditional feature frame rate is higher due to differences in model design, its frame rate is adjusted through interpolation, pooling, etc., downsampled to 50 frames per second to maintain consistency with the temporal resolution of the speech content sequence. This achieves precise alignment of the two types of features in the temporal dimension, avoiding modulation deviations caused by resolution differences. The downsampling operation ensures that the length of the hidden layer conditional feature sequence is exactly the same as the number of tokens in the speech content sequence, establishing a correspondence for subsequent frame-by-frame modulation.

[0093] S620 performs first feature linear modulation on the speech content sequence and the downsampled conditional features to obtain the first modulated features.

[0094] After temporal alignment is completed, the speech content sequence and downsampled conditional features are input into the feature linear modulation model. The implementation of the first feature linear modulation is as described in S104 and S510 above, that is, a scaling factor is generated based on the conditional information. and offset factor Affine transformation is performed on the input features.

[0095] Specifically, a scaling factor is first generated based on the downsampled conditional features. Sequence and offset factor sequence, and The dimension of the feature is consistent with that of the speech content sequence. Through the affine transformation process, the target pitch information carried by the downsampled conditional features is embedded into the semantic representation of each frame of the speech content sequence, so that the output first modulated feature retains the semantic information of the original frame and carries the target pitch guidance information at that moment, thus realizing frame-level fusion of content and pitch.

[0096] S630, semantically encode the first modulated feature to obtain the pitch mean condition.

[0097] Specifically, the first modulated feature is input into a preset encoder network for semantic encoding. The preset encoder network is capable of feature extraction and context association of the modulated feature.

[0098] In one possible implementation, the encoder network employs a Transformer-based architecture. After semantic encoding by the encoder network, the first modulated feature is converted into pitch mean conditions. Compared to the input first modulated feature, the pitch mean conditions have stronger semantic expressive power and temporal modeling capabilities. They not only fully preserve the semantic and pronunciation information of the original singing voice but also carry the target pitch guidance information deepened by the encoder, providing more robust and richer mean constraints for subsequent conditional diffusion generation.

[0099] In this embodiment, downsampling is used to achieve precise alignment of feature timing, avoiding timing deviations in the modulation process; feature linear modulation is used to fuse pitch and content information, preserving the original semantics while injecting pitch constraints; and semantic encoding is used to enhance the output pitch mean condition, which provides a highly reliable mean constraint for subsequent diffusion generation, ensuring the accuracy of pitch correction and the integrity of the content.

[0100] Based on the embodiment corresponding to S510 above, in order to more clearly demonstrate the process of generating the target Mel spectrum, this application also provides a possible implementation of generating the target Mel spectrum in a fundamental frequency modulation-based vocal pitch correction method. Figure 5 This is a schematic diagram illustrating the process of generating the target Mel spectrum in a fundamental frequency modulation-based vocal pitch correction method provided in an embodiment of this application. Figure 5 As shown, S520 above performs conditional diffusion based on the pitch mean condition, timbre feature vector, and Mel dimension conditional features to generate the target Mel spectrum, which includes: S710 upsamples the pitch mean condition to obtain the upsampled condition features.

[0101] Specifically, the pitch mean condition is upsampled to ensure that its temporal resolution is consistent with the Mel dimension condition features and the Mel spectrum to be generated.

[0102] In one possible implementation, the original frame rate for the pitch mean condition is 50 frames per second (one frame every 20 milliseconds), while the frame rate for the Mel spectrum is set to 100 frames per second (one frame every 10 milliseconds), a factor of two. The pitch mean condition is upsampled to 100 frames per second using transposed convolution or linear interpolation to obtain the upsampled conditional features.

[0103] The upsampling operation ensures that the length of the conditional feature sequence is exactly the same as the number of Mel spectrum frames to be generated, thus establishing an accurate temporal correspondence for subsequent frame-by-frame modulation and conditional input.

[0104] S720 performs second feature linear modulation on the upsampled conditional features and the Mel dimension conditional features to obtain the second modulated features.

[0105] Specifically, the upsampling conditional features are used as the basic features to be modulated, and the Mel-dimensional conditional features are used as the modulation conditions, which are then input into a preset feature linear modulation model. The implementation of the second feature linear modulation follows the aforementioned feature linear modulation principle, that is, a scaling factor is generated based on the Mel-dimensional conditional features. and offset factor Affine transformation is performed on the upsampling conditional features.

[0106] In one possible implementation, the scaling factor is first generated based on the Mel dimension conditional features. Sequence and offset factor sequence, and The dimension of the upsampled conditional feature is kept consistent with the feature dimension of the upsampled conditional feature; then the upsampled conditional feature is... Element-wise multiplication is used to scale and modulate the features; then the scaled features are multiplied by... Element-by-element addition achieves feature offset modulation.

[0107] Through this affine transformation process, the pitch information carried by the Mel dimension conditional features is embedded into the upsampled conditional features, achieving deep embedding of pitch information and further strengthening the pitch constraint properties of the features, thus obtaining the second modulated features.

[0108] S730 performs conditional diffusion decoding based on the second modulated features, timbre feature vector, and Mel dimension conditional features to obtain the target Mel spectrum.

[0109] Specifically, the second modulated feature is used as the mean condition and input into the conditional diffusion decoder along with the timbre feature vector and the Mel dimension conditional feature.

[0110] In one possible implementation, the decoder employs a Transformer-based DiT (DiffusionTransformer) architecture, implemented by stacking multiple Transformer blocks. The decoder's input includes the following information: the noisy Mel spectrum corresponding to the current denoising step, the time-step embedding dynamically calculated from the current denoising step, the second modulated feature, the timbre feature vector, the Mel dimension conditional feature, and the effective region mask determined based on the actual length of the input audio.

[0111] The temporal step embedding is dynamically calculated based on the current denoising step number and is used to convey the noise level information at the current stage, thereby enabling the adoption of corresponding processing strategies according to different denoising stages. The effective region mask is determined by the actual length of the input audio and is used to distinguish between the effective audio region and the padding region. The padding region is masked when calculating the loss and attention weights to avoid interference from invalid information in the generation process and ensure that the model only models and generates the effective audio portion.

[0112] The diffusion decoder starts with standard Gaussian noise and iteratively denoises by solving Ordinary Differential Equations (ODEs). In each iteration, the noisy Mel spectrum of the current step, the time-step embedding, and various conditional features are input into a pre-defined velocity field estimator network to obtain the velocity field prediction for the current step. Then, the Mel spectrum is updated based on the velocity field and a pre-defined step size to obtain the noisy Mel spectrum for the next step. In one possible implementation, the Euler method is used as the ODE solver, and cosine scheduling is employed for time-step scheduling to improve generation quality. The sampling step count is set to 10 steps, achieving a good balance between generation quality and computational efficiency. After multiple iterations of denoising, the final Mel spectrum that matches the target pitch, preserves the target timbre, and is consistent with the original semantic content is generated; this is the target Mel spectrum.

[0113] In this embodiment, upsampling is used to align the feature timing with the Mel spectrum dimension, avoiding feature mismatch during the decoding stage. The second feature linear modulation further strengthens the pitch constraint, and the input of Mel dimension conditional features forms multiple pitch guidance. Based on diffusion decoding and fusion of multi-dimensional feature constraints, the accuracy of the target Mel spectrum pitch and the consistency of timbre are achieved while ensuring generation efficiency, providing a high-quality feature foundation for subsequent high-fidelity audio reconstruction.

[0114] The following describes a vocal pitch correction device and electronic device based on fundamental frequency modulation provided in this application, the specific implementation process and technical effects of which are described above and will not be repeated below.

[0115] Figure 6 A schematic diagram of a vocal pitch correction device based on fundamental frequency modulation provided in this application embodiment is shown below. Figure 6 As shown, the fundamental frequency modulation-based singing pitch correction device includes: The extraction module 1000 is used to extract the speech content sequence from the original singing audio to be corrected; and to obtain the timbre feature vector of the target speaker corresponding to the original singing audio.

[0116] Encoding module 2000 is used to encode the fundamental frequency curve of the target song corresponding to the original singing audio, so as to obtain the fundamental frequency conditional features of the target song.

[0117] The modulation module 3000 is used to perform first feature linear modulation on the speech content sequence and fundamental frequency condition features to obtain the pitch mean condition.

[0118] The diffusion module 4000 is used to perform conditional diffusion based on pitch mean conditions, timbre feature vectors, and fundamental frequency condition features to generate the target Mel spectrum.

[0119] The reconstruction module 5000 is used to reconstruct the audio waveform based on the target Mel spectrum to obtain the target singing audio after pitch correction.

[0120] Optionally, the extraction module 1000 is specifically used to extract the log-Mel spectrum from the original singing audio to obtain log-Mel spectrum features.

[0121] Optionally, the encoding module 5000 is also used to encode the speech content of the log-Mel spectrum features to obtain a speech content sequence.

[0122] Optionally, the extraction module 1000 is specifically used to extract speaker features from the original singing audio to obtain the timbre feature vector; or, to extract speaker features from the reference audio of the target speaker to obtain the timbre feature vector.

[0123] Optionally, the encoding module 2000 is specifically used to quantize and encode multiple consecutive fundamental frequency values ​​in the target fundamental frequency curve to obtain a fundamental frequency level sequence with multiple discrete levels; and to perform fundamental frequency embedding on the fundamental frequency level sequence to obtain a fundamental frequency embedding vector sequence.

[0124] Optionally, the extraction module 1000 is also used to extract features from the fundamental frequency embedding vector sequence to obtain fundamental frequency conditional features.

[0125] Optionally, the modulation module 3000 is specifically used for fundamental frequency conditional features, including hidden layer conditional features and Mel dimension conditional features. The speech content sequence and hidden layer conditional features are linearly modulated using the first feature to obtain the pitch mean condition.

[0126] Optionally, the diffusion module 4000 is specifically used to perform conditional diffusion based on the pitch mean condition, timbre feature vector, and Mel dimension conditional features to generate the target Mel spectrum.

[0127] Optionally, the modulation module 3000 is specifically used to downsample the hidden layer conditional features to the same temporal resolution as the speech content sequence to obtain downsampled conditional features; and to perform first feature linear modulation on the speech content sequence and the downsampled conditional features to obtain first modulated features.

[0128] Optionally, the encoding module 2000 is also used to perform semantic encoding on the first modulated feature to obtain the pitch mean condition.

[0129] Optionally, the modulation module 3000 is specifically used to upsample the pitch mean condition to obtain the upsampled condition feature; and to perform second feature linear modulation on the upsampled condition feature and the Mel dimension condition feature to obtain the second modulated feature.

[0130] Optionally, the diffusion module 4000 is specifically used to perform conditional diffusion decoding based on the second modulated features, timbre feature vector, and Mel dimension conditional features to obtain the target Mel spectrum.

[0131] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more digital signal processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).

[0132] Figure 7 This is a schematic diagram of an electronic device provided in an embodiment of this application. The device may be a computing device or a server with computing processing capabilities.

[0133] The electronic device 10 includes a processor 11, a storage medium 12, and a bus 13. The storage medium 12 stores program instructions executable by the processor 11. When the electronic device 10 is executed, the processor 11 communicates with the storage medium 12 via the bus 13, and the processor 11 executes the program instructions to perform the above-described method embodiment. The specific implementation and technical effects are similar and will not be described in detail here.

[0134] Optionally, this application also provides a program product, such as a computer-readable storage medium, including a program that, when executed by a processor, performs the above-described method embodiments.

[0135] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0136] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0137] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units.

[0138] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0139] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for correcting the pitch of singing voices based on fundamental frequency modulation, characterized in that, The method includes: Extract the speech content sequence from the original singing audio to be corrected; Obtain the timbre feature vector of the target speaker corresponding to the original singing audio; The fundamental frequency curve of the target song corresponding to the original singing audio is encoded to obtain the fundamental frequency conditional features of the target song; The speech content sequence and the fundamental frequency condition feature are subjected to first feature linear modulation to obtain the pitch mean condition; Based on the pitch mean condition, the timbre feature vector, and the fundamental frequency condition feature, conditional diffusion is performed to generate the target Mel spectrum; The target Mel spectrum is used to reconstruct the audio waveform, and the pitch-corrected target singing audio is obtained.

2. The method according to claim 1, characterized in that, The extraction of speech content sequences from the original singing audio to be corrected includes: Log-Melbourne spectrum extraction is performed on the original singing audio to obtain log-Melbourne spectrum features; The log-Mel spectrum features are encoded to obtain the speech content sequence.

3. The method according to claim 1, characterized in that, The step of obtaining the timbre feature vector of the target speaker corresponding to the original singing audio also includes: Speaker features are extracted from the original singing audio to obtain the timbre feature vector; or, Speaker features are extracted from the reference audio of the target speaker to obtain the timbre feature vector.

4. The method according to claim 1, characterized in that, The step of encoding the fundamental frequency curve of the target song corresponding to the original singing audio to obtain the fundamental frequency conditional features of the target song includes: Quantize and encode multiple consecutive fundamental frequency values ​​in the target fundamental frequency curve to obtain a fundamental frequency level sequence with multiple discrete levels; The fundamental frequency range sequence is subjected to fundamental frequency embedding to obtain a fundamental frequency embedding vector sequence; Feature extraction is performed on the fundamental frequency embedding vector sequence to obtain the fundamental frequency conditional features.

5. The method according to claim 1, characterized in that, The fundamental frequency condition features include: hidden layer condition features and Mel dimension condition features; The step of performing first feature linear modulation on the speech content sequence and the fundamental frequency conditional features to obtain the pitch mean condition includes: The speech content sequence and the hidden layer conditional features are linearly modulated using the first feature to obtain the pitch mean condition; The step of generating the target Mel spectrum by performing conditional diffusion based on the pitch mean condition, the timbre feature vector, and the fundamental frequency condition feature includes: The target Mel spectrum is generated by performing conditional diffusion based on the pitch mean condition, the timbre feature vector, and the Mel dimension conditional feature.

6. The method according to claim 5, characterized in that, The first feature linear modulation is performed on the speech content sequence and the hidden layer conditional features to obtain the pitch mean condition, including... The hidden layer conditional features are downsampled to the same temporal resolution as the speech content sequence to obtain the downsampled conditional features. The speech content sequence and the downsampled conditional features are linearly modulated using a first feature to obtain the first modulated features. The first modulated feature is semantically encoded to obtain the pitch mean condition.

7. The method according to claim 5, characterized in that, The step of generating the target Mel spectrum by performing conditional diffusion based on the pitch mean condition, the timbre feature vector, and the Mel dimension conditional feature includes: Upsampling is performed on the pitch mean condition to obtain the upsampling condition features; The upsampled conditional features and the Mel dimension conditional features are subjected to second feature linear modulation to obtain the second modulated features; Conditional diffusion decoding is performed based on the second modulated feature, the timbre feature vector, and the Mel dimension conditional feature to obtain the target Mel spectrum.

8. A vocal pitch correction device based on fundamental frequency modulation, characterized in that, The device includes: The extraction module is used to extract the speech content sequence from the original singing audio to be corrected; and to obtain the timbre feature vector of the target speaker corresponding to the original singing audio. The encoding module is used to perform fundamental frequency encoding on the target fundamental frequency curve of the target song corresponding to the singing audio, so as to obtain the fundamental frequency conditional features of the target song; The modulation module is used to perform first feature linear modulation on the speech content sequence and the fundamental frequency condition feature to obtain the pitch mean condition; The diffusion module is used to perform conditional diffusion based on the pitch mean condition, the timbre feature vector, and the fundamental frequency condition feature to generate the target Mel spectrum; The reconstruction module is used to reconstruct the audio waveform based on the target Mel spectrum to obtain the target singing audio after pitch correction.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores program instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. The processor executes the program instructions to perform the steps of the fundamental frequency modulation-based vocal pitch correction method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, performs the vocal pitch correction method based on fundamental frequency modulation as described in any one of claims 1-7.