Audio correction method and device, terminal equipment and computer program product
By acquiring the singer's historical performance audio and using a predictive model to adjust the pitch correction range, the problem that the pitch correction range cannot match the singing level in real time in existing technologies is solved, and dynamic pitch correction is realized to improve singing level and pitch correction accuracy.
Patent Information
- Application Number
- CN202510760324.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-10-28
AI Technical Summary
Existing vocal editing technologies use fixed templates and cannot adjust the degree of vocal correction according to the user's real-time singing level. This results in distorted vocals that lack personal characteristics and slows down the improvement of singing skills.
By acquiring the singer's historical performance audio, training suggestions are generated using a performance level prediction model. The target pitch correction range is dynamically adjusted based on the number of training sessions completed, and pitch correction is performed based on the baseline audio to ensure that the pitch correction range matches the singer's real-time performance level.
It improved the accuracy of pitch correction, reduced reliance on vocal tuning, and promoted the improvement of singers' performance levels.
Smart Images

Figure CN120853594A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of audio processing technology, and in particular relates to an audio correction method, apparatus, terminal equipment and computer program product. Background Technology
[0002] With the development of mobile internet, karaoke has gradually become an important form of leisure and entertainment. The application of vocal enhancement technology allows users to enjoy a better singing experience when singing karaoke, and even those with average singing skills can achieve satisfactory results.
[0003] However, current vocal editing technologies often use fixed templates, adjusting the user's voice to the template pitch while they sing. This method easily distorts the vocals, making the edited voice lack the user's personal characteristics. Furthermore, because current vocal editing technologies use fixed templates, they often don't consider the user's real-time singing level. When a user's singing shows significant improvement, the editing is often applied within a fixed range without adjustment. This leads to the user becoming increasingly reliant on the editing, slowing down their progress.
[0004] Currently, no effective solution has been proposed to address the problem that the pitch correction amplitude cannot be adjusted in real time in related technologies, resulting in a mismatch between the pitch correction amplitude and the singer's real-time singing level. Summary of the Invention
[0005] This application provides an audio correction method, apparatus, terminal device, and computer program product to at least solve the problem in the related art where the pitch correction amplitude cannot be adjusted in real time, resulting in a mismatch between the pitch correction amplitude and the singer's real-time singing level.
[0006] In a first aspect, embodiments of this application provide an audio correction method, comprising: acquiring an audio to be corrected and a reference audio; acquiring historical performance audio of the singer of the audio to be corrected; inputting the historical performance audio into a performance level prediction model to obtain training suggestions for the singer output by the performance level prediction model; acquiring the number of times the singer has completed training based on the training suggestions; determining a target pitch correction amplitude based on the number of training completions, wherein the target pitch correction amplitude decreases as the number of training completions increases; and performing pitch correction processing on the audio to be corrected according to the target pitch correction amplitude based on the reference audio to obtain corrected audio.
[0007] In some embodiments, inputting the historical singing audio into a singing level prediction model to obtain training suggestions for the singer output by the singing level prediction model includes: inputting the historical singing audio into the singing level prediction model to obtain a first pitch deviation curve and a first rhythm deviation curve of the singer with respect to the historical singing audio output by the singing level prediction model; determining the historical pitch deviation and historical rhythm deviation of the singer with respect to the historical singing audio based on the first pitch deviation curve and the first rhythm deviation curve; generating a pitch training piece when the historical pitch deviation is greater than a preset first threshold; generating a rhythm training piece when the historical rhythm deviation is greater than a preset second threshold; and generating the training suggestions based on the pitch training piece and the rhythm training piece.
[0008] In some embodiments, the target pitch correction amplitude and the number of training iterations are expressed according to the following mathematical expression: A n =A0·(1-β·n); where, A n To complete the target pitch correction amplitude after n training iterations, β is a preset decrease rate, A0 is a preset initial pitch correction amplitude, and n is the number of training iterations.
[0009] In some embodiments, performing pitch correction processing on the audio to be corrected according to the target pitch correction amplitude based on the reference audio to obtain corrected audio includes: extracting features from the audio to be corrected to obtain a first pitch feature vector, a first rhythm feature vector, and a timbre feature vector; correcting the first pitch feature vector according to the target pitch correction amplitude and the reference pitch of the reference audio to generate a second pitch feature vector; correcting the first rhythm feature vector according to the target pitch correction amplitude and the reference rhythm of the reference audio to generate a second rhythm feature vector; and fusing the second pitch feature vector, the second rhythm feature vector, and the timbre feature vector to obtain corrected audio.
[0010] In some embodiments, feature extraction of the audio to be corrected to obtain a first pitch feature vector, a first rhythm feature vector, and a timbre feature vector includes: performing a short-time Fourier transform on the audio to be corrected to obtain a first spectrum; performing noise reduction on the first spectrum to obtain a second spectrum; performing feature extraction on the second spectrum to obtain the first pitch feature vector, the first rhythm feature vector, and MFCC features; and performing timbre extraction on the MFCC features to obtain the timbre feature vector.
[0011] In some embodiments, feature extraction of the second spectrum to obtain the first pitch feature vector and the first rhythm feature vector includes: using a fundamental frequency extraction algorithm to extract features from the second spectrum to obtain the first pitch feature vector; and using a beat detection algorithm to extract features from the spectrum to obtain the first rhythm feature vector.
[0012] In some embodiments, after fusing the second pitch feature vector, the second rhythm feature vector, and the timbre feature vector to obtain the corrected audio, the method further includes: extracting features from the audio to be corrected to obtain a volume change rate and a vibrato frequency; determining a first score for the audio to be corrected based on the reference pitch and the first pitch feature vector; determining a second score for the audio to be corrected based on the reference rhythm and the first rhythm feature vector; determining a third score for the audio to be corrected based on the volume change rate and the vibrato frequency; determining a total score for the audio to be corrected based on the first score, the second score, and the third score; and determining a second pitch deviation curve between the audio to be corrected and the reference audio based on the reference pitch and the first pitch feature vector.
[0013] Secondly, embodiments of this application provide an audio correction device, comprising: a first acquisition module, configured to acquire an audio to be corrected and a reference audio; and to acquire historical singing audio of the singer of the audio to be corrected; an input module, configured to input the historical singing audio into a singing level prediction model to obtain training suggestions for the singer output by the singing level prediction model; a second acquisition module, configured to acquire the number of times the singer has completed training according to the training suggestions; a determination module, configured to determine a target pitch correction amplitude based on the number of training completions, wherein the target pitch correction amplitude decreases as the number of training completions increases; and a pitch correction module, configured to perform pitch correction processing on the audio to be corrected according to the target pitch correction amplitude based on the reference audio to obtain corrected audio.
[0014] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the audio correction method of any of the first aspects described above.
[0015] Fourthly, embodiments of this application provide a computer program product, including a computer program, which, when run, causes the audio correction method described in any one of the first aspects to be executed.
[0016] Compared to related technologies, the audio correction method, apparatus, terminal device, and computer program product provided in this application obtain the audio to be corrected, the reference audio, and the singer's historical singing audio of the audio to be corrected. The historical singing audio is then input into a singing level prediction model to obtain training suggestions for the singer output by the singing level prediction model. Subsequently, the number of times the singer has completed training based on the training suggestions can be obtained, and the target pitch correction amplitude is determined based on the number of training completions, wherein the target pitch correction amplitude decreases as the number of training completions increases. Finally, based on the reference audio, the audio to be corrected is processed according to the target pitch correction amplitude to obtain the corrected audio. In this way, by inputting the singer's historical performance audio into the singing level prediction model, training suggestions are obtained, and the number of times the singer completes training according to the suggestions is acquired. This allows the singer's real-time singing level to be determined, and the pitch correction amplitude is adjusted based on the number of training completions. The target pitch correction amplitude can decrease as the number of training completions increases, ensuring the target amplitude matches the singer's actual singing level. Pitch correction based on this target amplitude improves accuracy and allows the singer to gradually reduce their reliance on pitch correction as their singing level improves, thereby accelerating the rate of improvement. This application solves the problem in related technologies where the pitch correction amplitude cannot be adjusted in real time, leading to a mismatch between the pitch correction amplitude and the singer's real-time singing level. It achieves the technical effect of dynamically adjusting the pitch correction amplitude to match the singer's real-time singing level, thereby improving pitch correction accuracy.
[0017] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of an audio correction method according to an embodiment of this application;
[0020] Figure 2 This is a flowchart of a singing scoring method according to an embodiment of this application;
[0021] Figure 3 This is a schematic diagram of the structure of an audio correction device according to an embodiment of this application;
[0022] Figure 4This is a schematic diagram of the structure of a terminal device according to an embodiment of this application. Detailed Implementation
[0023] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0024] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0025] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0026] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0027] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0028] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0029] With the development of mobile internet, karaoke has gradually become an important form of leisure and entertainment. The application of vocal enhancement technology allows users to enjoy a better singing experience when singing karaoke, and even those with average singing skills can achieve satisfactory results.
[0030] However, current vocal editing technologies often use fixed templates, adjusting the user's voice to the template pitch while they sing. This method easily distorts the vocals, making the edited voice lack the user's personal characteristics. Furthermore, because current vocal editing technologies use fixed templates, they often don't consider the user's real-time singing level. When a user's singing shows significant improvement, the editing is often applied within a fixed range without adjustment. This leads to the user becoming increasingly reliant on the editing, slowing down their progress.
[0031] Currently, no effective solution has been proposed to address the problem that the pitch correction amplitude cannot be adjusted in real time in related technologies, resulting in a mismatch between the pitch correction amplitude and the singer's real-time singing level.
[0032] In view of this, embodiments of this application provide an audio correction method. This method involves acquiring the audio to be corrected, a reference audio, and historical performance audio of the singer performing the audio to be corrected. The historical performance audio is then input into a performance level prediction model to obtain training suggestions for the singer output by the model. Subsequently, the number of times the singer completes training based on the training suggestions is obtained, and a target pitch correction amplitude is determined based on the number of training completions. The target pitch correction amplitude decreases as the number of training completions increases. Finally, based on the reference audio, the audio to be corrected is processed according to the target pitch correction amplitude to obtain the corrected audio. In this way, by inputting the singer's historical performance audio into the singing level prediction model, training suggestions are obtained, and the number of times the singer completes training according to the suggestions is acquired. This allows the singer's real-time singing level to be determined, and the pitch correction amplitude is adjusted based on the number of training completions. The target pitch correction amplitude can decrease as the number of training completions increases, ensuring the target amplitude matches the singer's actual singing level. Pitch correction based on this target amplitude improves accuracy and allows the singer to gradually reduce their reliance on pitch correction as their singing level improves, thereby accelerating the rate of improvement. This application solves the problem in related technologies where the pitch correction amplitude cannot be adjusted in real time, leading to a mismatch between the pitch correction amplitude and the singer's real-time singing level. It achieves the technical effect of dynamically adjusting the pitch correction amplitude to match the singer's real-time singing level, thereby improving pitch correction accuracy.
[0033] The following will combine Figure 1For an explanation of an audio correction method provided in one embodiment of this application, please refer to [link to relevant documentation]. Figure 1 , Figure 1 This is a flowchart of an audio correction method according to an embodiment of this application, such as... Figure 1 As shown, the method includes:
[0034] Step S101: Obtain the audio to be corrected and the reference audio.
[0035] In this embodiment, an audio acquisition device such as a microphone can be used to acquire the audio to be corrected at a fixed frequency (e.g., a sampling rate of 44.1 kHz). The buffer size can be set to 10 ms to ensure low latency.
[0036] For example, when the audio correction method provided in this application is applied to karaoke software, KTV embedded systems, audio recording software, or augmented reality (AR) / virtual reality (VR) devices, the microphone in the terminal device with the karaoke software or audio recording software installed can be used to collect the audio to be corrected at a sampling rate of 44.1 kHz. Alternatively, a wireless or wired microphone connected to the KTV embedded system can be used to collect the audio to be corrected at a sampling rate of 44.1 kHz. Or, the microphone built into the AR / VR device can be used to collect the audio to be corrected at a sampling rate of 44.1 kHz. Furthermore, when the audio correction method provided in this application is applied to a live streaming platform (e.g., an OBS plugin or browser extension), the WebRTC or OBS audio interface can be used to collect the audio to be corrected at a sampling rate of 44.1 kHz.
[0037] In this embodiment, the reference audio can be the original audio of the song sung by the singer, or audio of another person singing a song corresponding to the song sung by the singer, as specified by the user, or electronically generated audio.
[0038] Step S102: Obtain the historical performance audio of the singer whose audio needs to be corrected.
[0039] Step S103: Input the historical singing audio into the singing level prediction model to obtain the singer's training suggestions output by the singing level prediction model.
[0040] In this embodiment, a singing level prediction model can be obtained by training a Transformer model. This singing level prediction model can be stored in the cloud or on a server. The model can record the singer's singing history, including historical audio recordings, and generate a first pitch deviation curve and a first rhythm deviation curve based on these historical audio recordings.
[0041] Specifically, historical singing audio can be input into the singing level prediction model to obtain the singer's first pitch deviation curve and first rhythm deviation curve related to the historical singing audio, output by the singing level prediction model. These first pitch deviation curves and first rhythm deviation curves can characterize the pitch or rhythm deviation between a historical singing audio and its corresponding historical reference audio; alternatively, they can characterize the singer's pitch or rhythm deviation rate over time.
[0042] Based on the first pitch deviation curve and the first rhythm deviation curve, the historical pitch deviation and historical rhythm deviation of the singer regarding the historical singing audio can be determined; if the historical pitch deviation is greater than a preset first threshold, a pitch training song can be generated; if the historical rhythm deviation is greater than a preset second threshold, a rhythm training song can be generated; and training suggestions can be generated based on the pitch training song and the rhythm training song.
[0043] In this embodiment, training suggestions may include training tracks. For example, if the first pitch deviation curve shows that the singer's historical pitch deviation is greater than a first threshold (e.g., it can be set to 10%), a pitch training track with a preset beat count (e.g., 120 BPM) can be generated; if the first rhythm deviation curve shows that the singer's historical rhythm deviation is greater than a second threshold (e.g., it can be set to 5%), a rhythm training track can be generated; if the first pitch deviation curve and the first rhythm deviation curve show that the singer's high note failure rate is greater than 20%, a high note training track can be generated.
[0044] In this embodiment, the singing level prediction model can also generate or predict the singer's pitch or rhythm improvement curve over time. Specifically, the singing level prediction model can predict the extent of the singer's pitch or rhythm improvement after completing a preset number of training sessions according to training recommendations, thereby predicting the change in the singer's singing ability with the number of training sessions and encouraging the singer to hone their singing skills.
[0045] Step S104: Obtain the number of times the singer completed the training according to the training recommendations.
[0046] Step S105: Determine the target pitch correction amplitude based on the number of training sessions completed, wherein the target pitch correction amplitude decreases as the number of training sessions completed increases.
[0047] In this embodiment, the training suggestions may include any combination of pitch training pieces, rhythm training pieces, or treble training pieces. Once a singer completes a performance of a pitch training piece, rhythm training piece, or treble training piece, it is considered that the singer has completed a training session based on the training suggestions.
[0048] In this embodiment, personalized training suggestions are generated for the singer, and the singer's real-time singing level can be determined by obtaining the number of times the singer completes training according to the suggestions. Therefore, the target pitch correction amplitude determined based on the number of training sessions completed by the singer can match the singer's real-time singing level, thereby ensuring that subsequent pitch correction also matches the singer's real-time singing level and improving the accuracy of pitch correction. Furthermore, since the target pitch correction amplitude decreases as the number of training sessions increases, the pitch correction amplitude can be gradually reduced after the singer completes training according to the suggestions. This reduces the singer's reliance on pitch correction, improves the singer's pitch accuracy, rhythm, and high-note ability, and thus accelerates the singer's progress in singing level.
[0049] In addition, multiple modes can be preset, allowing singers to choose or using a default mode to adjust the pitch correction. For example, three modes can be set: Full Correction, Fine Adjustment, and Natural. Full Correction is suitable for beginners, with a fixed target pitch correction of 100%, automatically adjusting the pitch and rhythm of the audio to be corrected in real time to achieve the original effect while preserving the singer's personal vocal characteristics. Fine Adjustment is suitable for intermediate users, making smaller adjustments to pitch and rhythm only when the pitch or rhythm deviation is large (e.g., greater than 10%) (e.g., the target pitch correction can be less than 80%). Natural mode minimizes pitch correction, providing correction only at key syllables, encouraging singers to perform independently.
[0050] It should be noted that a 100% pitch correction margin can mean adjusting the pitch, rhythm, and other parameters of the audio to be corrected to be completely consistent with the reference pitch, rhythm, and other parameters of the reference audio. If the pitch correction margin is x%, then the pitch, rhythm, and other parameters of the audio to be corrected can be corrected with an accuracy of x% (for example, if the audio to be corrected contains 100 pitch correction nodes, then 100*x% of the pitch correction nodes can be corrected to make the pitch, rhythm, and other parameters of the corrected pitch correction nodes consistent with the reference pitch, rhythm, and other parameters of the reference audio).
[0051] By setting multiple modes, each corresponding to different pitch correction amplitude and pitch correction node settings, the applicability of the audio correction method provided in this application embodiment can be improved. Singers of different singing levels can obtain pitch correction processing adapted to their level, thereby encouraging singers to sing and improving the speed of their singing level improvement.
[0052] In one embodiment, the target pitch correction amplitude and the number of training sessions can be represented by the following mathematical expression: A n = a0·(1-β·n); where a nThe target pitch correction amplitude after n training sessions is defined by β, the preset reduction rate is defined by β, the preset initial pitch correction amplitude is defined by A0 (which can be fixed at 100%), and n is the number of training sessions. β can be set to 0.02; for example, with β = 0.02, the pitch correction amplitude decreases by 10% every 5 training sessions.
[0053] Step S106: Based on the reference audio, perform audio correction processing on the audio to be corrected according to the target correction amplitude to obtain the corrected audio.
[0054] In this embodiment, only parameters such as pitch and rhythm in the audio to be corrected can be adjusted, without adjusting the timbre parameters, thereby improving the singing quality of the corrected voice while preserving the singer's personal timbre characteristics.
[0055] Specifically, step S106 may include the following steps:
[0056] Step 1: Extract features from the audio to be corrected to obtain the first pitch feature vector, the first rhythm feature vector, and the timbre feature vector.
[0057] In this embodiment, the frequency and time domain features of the audio to be corrected can be extracted using the Short-Time Fourier Transform (STFT) and Mel-Frequency Cepstral Coefficient (MFCC). These features may include, for example, pitch, tone, rhythm, resonance, and articulation.
[0058] Specifically, feature extraction of the audio to be corrected to obtain the first pitch feature vector, the first rhythm feature vector, and the timbre feature vector may include: performing a short-time Fourier transform on the audio to be corrected to obtain a first spectrum; performing noise reduction on the first spectrum to obtain a second spectrum; performing feature extraction on the second spectrum to obtain the first pitch feature vector, the first rhythm feature vector, and MFCC features; and performing timbre extraction on the MFCC features to obtain a timbre feature vector.
[0059] In this embodiment, the window size of the short-time Fourier transform can be set to 1024, and the step size can be set to 256.
[0060] In addition, a Time-Domain Audio Separation Network (TasNet) can be used to separate background noise from the first spectrum, or, in the case of multiple backing vocals, the backing vocals can be separated from the first spectrum to obtain a second spectrum containing only the lead singer's (i.e., the singer's) voice.
[0061] In one embodiment, when there are multiple singers, the voice corresponding to each singer can be separated from the first spectrum to obtain multiple second spectra, where each second spectrum corresponds to one singer. Subsequently, multi-channel analysis can be performed, and audio correction operations can be carried out on each second spectrum to complete the pitch correction operation for each singer.
[0062] In this embodiment, feature extraction of the second spectrum to obtain the first pitch feature vector and the first rhythm feature vector may include: using a fundamental frequency extraction algorithm (Yin algorithm) to extract features of the second spectrum to obtain the first pitch feature vector; using a beat detection algorithm to extract features of the spectrum to obtain the first rhythm feature vector; and extracting features of the second spectrum based on the spectral peaks of the second spectrum to obtain the first resonance feature vector.
[0063] In this embodiment, the second spectrum can be filtered to obtain MFCC features (13 dimensions). The calculation formula for the MFCC features is as follows:
[0064]
[0065] Where S(i) is the output of the Mel filter bank, N is the number of Mel filter banks, and K is the Mel frequency cepstral coefficient index (the first 13 can be taken, corresponding to 13 dimensions).
[0066] In this embodiment, by using STFT and MFCC, various frequency and time domain features (including pitch features, rhythm features, resonance features, and timbre features) of the audio to be corrected can be extracted. By separating the various features of the audio to be corrected, only a few features can be selected for correction according to actual needs, while retaining the timbre features. This improves the singing quality of the corrected voice while preserving the singer's personal timbre features.
[0067] Step 2: Correct the first pitch feature vector based on the target pitch correction amplitude and the reference pitch of the reference audio, and generate the second pitch feature vector.
[0068] In this embodiment, the first pitch feature vector can be corrected in real time based on the target pitch correction amplitude and the reference pitch of the reference audio (e.g., the original vocal audio or the original song audio), combined with the Pitch-Synchronous Overlap-Add (PSOLA) algorithm and the WaveNet generative model, thereby correcting the pitch deviation.
[0069] Specifically, the PSOLA algorithm can be used to correct the first pitch feature vector so that the corrected first pitch feature vector is consistent with the reference pitch of the reference audio; the WaveNet generative model can be used to generate a natural transition and generate the second pitch feature vector.
[0070] As an example, in the process of correcting the first pitch feature vector using the PSOLA algorithm, the pitch adjustment factor can be calculated: α = f target / f source Then, time axis stretching is performed: t new =t old *α, and finally perform overlapping addition: Where P is the pitch period and ω is the window function.
[0071] Step 3: Based on the target pitch correction amplitude and the reference rhythm of the reference audio, correct the first rhythm feature vector to generate the second rhythm feature vector.
[0072] In this embodiment, the first rhythm feature vector can be corrected using the Dynamic Time Warping (DTW) algorithm based on the target pitch correction amplitude and the reference rhythm of the reference audio, thereby correcting the rhythm deviation in real time.
[0073] Specifically, in the process of correcting the first rhythmic feature vector using the DTW algorithm, the distance matrix can be calculated as: D(i,j)=d(x i x j )+min[D(i-1,j),D(i,j-);where,d(x i ,y j )=|x i -y j |) represents the distance between the beat of the audio to be corrected and the reference beat of the reference audio, and xi,,yj represents the time points of the beat of the audio to be corrected and the reference beat of the reference audio; then, time warping (i.e., time axis alignment) can be performed after alignment: t aligned =interp(t user , t ref D).
[0074] In steps 2 and 3 above, a variational autoencoder (VAE) can be used to separate pitch and timbre, adjusting only the pitch. The loss function of the VAE can be expressed as: Where MSE is the reconstruction error and KL is the divergence between the latent distribution and the standard normal distribution.
[0075] Step 4: The second pitch feature vector, the second rhythm feature vector, and the timbre feature vector are fused to obtain the corrected audio.
[0076] In this embodiment, after correcting the pitch and rhythm of the audio to be corrected to generate a second pitch feature vector and a second rhythm feature vector, they can be fused together with the timbre feature vector to obtain the corrected audio. Since the timbre feature vector is a feature that retains the singer's personal timbre separated from the first spectrum, fusing it with the second pitch feature vector and the second rhythm feature vector can improve the singing quality of the corrected voice while preserving the singer's personal timbre features, thereby reducing vocal distortion during vocal correction and improving the naturalness of the corrected voice.
[0077] In addition, in one embodiment, after step 4, the method may further include: performing resonance enhancement on the corrected audio based on the first resonance feature vector to obtain enhanced audio.
[0078] In this embodiment, an equalizer (EQ) can be applied to enhance mid-frequency or high-frequency resonance, further improving the singing quality of the corrected vocals.
[0079] Through the above steps S101 to S106, by acquiring the audio to be corrected, the reference audio, and the singer's historical performance audio, the historical performance audio is input into the performance level prediction model to obtain the singer's training suggestions output by the performance level prediction model. Subsequently, the number of times the singer completes training according to the training suggestions can be obtained, and the target pitch correction range is determined based on the number of training completions, wherein the target pitch correction range decreases as the number of training completions increases. Finally, based on the reference audio, the audio to be corrected is processed according to the target pitch correction range to obtain the corrected audio. In this way, by inputting the singer's historical performance audio into the singing level prediction model, training suggestions are obtained, and the number of times the singer completes training according to the suggestions is acquired. This allows the singer's real-time singing level to be determined, and the pitch correction amplitude is adjusted based on the number of training completions. The target pitch correction amplitude can decrease as the number of training completions increases, ensuring the target amplitude matches the singer's actual singing level. Pitch correction based on this target amplitude improves accuracy and allows the singer to gradually reduce their reliance on pitch correction as their singing level improves, thereby accelerating the rate of improvement. This application solves the problem in related technologies where the pitch correction amplitude cannot be adjusted in real time, leading to a mismatch between the pitch correction amplitude and the singer's real-time singing level. It achieves the technical effect of dynamically adjusting the pitch correction amplitude to match the singer's real-time singing level, thereby improving pitch correction accuracy.
[0080] In addition to correcting audio, the audio correction method provided in this application embodiment can also provide a singing score function.
[0081] The following will combine Figure 2 An exemplary flow of a singing scoring method provided in one embodiment of this application will be described. Please refer to [link to documentation]. Figure 2 , Figure 2 This is a flowchart of a singing scoring method according to an embodiment of this application, such as... Figure 2 As shown, the singing scoring method includes the following steps:
[0082] Step S201: Perform feature extraction on the audio to be corrected to obtain the volume change rate and vibrato rate.
[0083] Step S202: Determine the first score of the audio to be corrected based on the reference pitch and the first pitch feature vector.
[0084] In this embodiment, the mean square error of the deviation between the reference pitch and the first pitch feature vector can be calculated to determine the first score (also known as the pitch score):
[0085]
[0086] Where, p user p represents the singer's pitch obtained from the first pitch feature vector. ref T represents the reference pitch, and T represents the duration.
[0087] Step S203: Determine the second score of the audio to be corrected based on the baseline rhythm and the first rhythm feature vector.
[0088] In this embodiment, the DTW matching degree can be calculated based on the baseline rhythm and the first rhythm feature vector to determine the second score (also known as the rhythm score):
[0089]
[0090] Among them, D DTW For the cumulative distance of DTW, D max This represents the maximum possible distance.
[0091] Step S204: Determine the third score of the audio to be corrected based on the volume change rate and the vibrato frequency.
[0092] In this embodiment, the third rating (also known as the sentiment rating) can be calculated using the following mathematical expression:
[0093]
[0094] Where Var(RMS) is the volume variance, which can be calculated from the volume change rate; ftremolo The frequency of the vibrato.
[0095] It should be noted that "50", "30", and "20" in the above mathematical expression are preset weighted scores, that is, the weight of the first score is 50, the weight of the second score is 30, and the weight of the third score is 20. The above weights can be adjusted according to the song type of the audio to be corrected (for example, using a CNN classifier to identify the song type (e.g., pop or rock)).
[0096] Step S205: Determine the total score of the audio to be corrected based on the first score, the second score, and the third score.
[0097] In this embodiment, by setting pitch score, rhythm score and emotion score, the total score of the audio to be corrected is determined by combining the three scores. This can solve the problem of the relatively simple singing score in related technologies, and thus more accurately evaluate the singer's singing.
[0098] Furthermore, after step S205, the method may further include: determining a second pitch deviation curve between the audio to be corrected and the reference audio based on the reference pitch and the first pitch feature vector.
[0099] By obtaining the reference pitch and the first pitch feature vector, a second pitch deviation curve can be plotted between the audio to be corrected and the reference audio as the song's time axis changes. The second pitch deviation curve and the total score of the audio to be corrected are then visually fed back to the singer, allowing the singer to more intuitively understand their real-time singing level and the gap between themselves and the reference audio.
[0100] Alternatively, a second rhythm deviation curve between the audio to be corrected and the reference audio can be determined based on the baseline rhythm and the first rhythm feature vector. Then, similar to the step of generating training suggestions, improvement suggestions are generated based on the second pitch deviation curve or the second rhythm deviation curve, allowing the singer to specifically improve their pitch or rhythm.
[0101] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0102] Corresponding to the audio correction method described in the above embodiments, Figure 3 A schematic diagram of an audio correction device according to an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown.
[0103] See Figure 3The audio correction device 3 includes: a first acquisition module 30, used to acquire the audio to be corrected and a reference audio; and to acquire the historical singing audio of the singer of the audio to be corrected; an input module 31, used to input the historical singing audio into a singing level prediction model to obtain the singer's training suggestions output by the singing level prediction model; a second acquisition module 32, used to acquire the number of times the singer has completed training according to the training suggestions; a determination module 33, used to determine the target pitch correction amplitude based on the number of times training has been completed, wherein the target pitch correction amplitude decreases as the number of times training has been completed increases; and a pitch correction module 34, used to perform pitch correction processing on the audio to be corrected according to the target pitch correction amplitude based on the reference audio to obtain the corrected audio.
[0104] In one embodiment, the input module 31 is further configured to input historical singing audio into the singing level prediction model to obtain the first pitch deviation curve and the first rhythm deviation curve of the singer with respect to the historical singing audio output by the singing level prediction model; based on the first pitch deviation curve and the first rhythm deviation curve, determine the historical pitch deviation and historical rhythm deviation of the singer with respect to the historical singing audio; generate a pitch training song if the historical pitch deviation is greater than a preset first threshold; generate a rhythm training song if the historical rhythm deviation is greater than a preset second threshold; and generate training suggestions based on the pitch training song and the rhythm training song.
[0105] In one embodiment, the target pitch correction amplitude and the number of training sessions completed are expressed by the following mathematical expression: A n =A0·(1-β·n); where, A n To complete the target pitch correction amplitude after n training iterations, β is the preset decrease rate, A0 is the preset initial pitch correction amplitude, and n is the number of training iterations.
[0106] In one embodiment, the pitch correction module 34 is further configured to extract features from the audio to be corrected, obtaining a first pitch feature vector, a first rhythm feature vector, and a timbre feature vector; correct the first pitch feature vector according to the target pitch correction amplitude and the reference pitch of the reference audio, generating a second pitch feature vector; correct the first rhythm feature vector according to the target pitch correction amplitude and the reference rhythm of the reference audio, generating a second rhythm feature vector; and fuse the second pitch feature vector, the second rhythm feature vector, and the timbre feature vector to obtain the corrected audio.
[0107] In one embodiment, the audio correction module 34 is further configured to perform short-time Fourier transform processing on the audio to be corrected to obtain a first spectrum; perform noise reduction processing on the first spectrum to obtain a second spectrum; perform feature extraction on the second spectrum to obtain a first pitch feature vector, a first rhythm feature vector, and MFCC features; and perform timbre extraction processing on the MFCC features to obtain a timbre feature vector.
[0108] In one embodiment, the pitch correction module 34 is further configured to use a fundamental frequency extraction algorithm to extract features from the second spectrum to obtain a first pitch feature vector; and use a beat detection algorithm to extract features from the spectrum to obtain a first rhythm feature vector.
[0109] In one embodiment, the audio correction device 3 further includes a scoring module for extracting features from the audio to be corrected to obtain the volume change rate and vibrato frequency; determining a first score for the audio to be corrected based on a reference pitch and a first pitch feature vector; determining a second score for the audio to be corrected based on a reference rhythm and a first rhythm feature vector; determining a third score for the audio to be corrected based on the volume change rate and vibrato frequency; determining a total score and a first pitch deviation curve for the audio to be corrected based on the first score, the second score, and the third score; and determining a second pitch deviation curve between the audio to be corrected and the reference audio based on the reference pitch and the first pitch feature vector.
[0110] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0111] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0112] Figure 4 This is a schematic diagram of the structure of a terminal device according to an embodiment of this application. Figure 4 As shown, the terminal device 4 includes: at least one processor 40 ( Figure 4 (Only one is shown) a processor, a memory 41, and a computer program 42 stored in the memory 41 and executable on at least one processor 40, which, when executing the computer program 42, implements the steps in any of the above-described audio correction method embodiments.
[0113] Terminal device 4 can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. Terminal device 4 may include, but is not limited to, processor 40 and memory 41. Those skilled in the art will understand that... Figure 4 This is merely an example of terminal device 4 and does not constitute a limitation on terminal device 4. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0114] The processor 40 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0115] In some embodiments, memory 41 may be an internal storage unit of terminal device 4, such as a hard disk or memory of terminal device 4. In other embodiments, memory 41 may be an external storage device of terminal device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on terminal device 4. In other embodiments, memory 41 may include both internal and external storage units of terminal device 4. Memory 41 is used to store operating system, applications, bootloader, data, and other programs, such as the program code of computer program 42. Memory 41 may also be used to temporarily store data that has been output or will be output.
[0116] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various audio correction method embodiments above.
[0117] This application provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the various audio correction method embodiments.
[0118] This application implements all or part of the processes in the methods of the above embodiments, which can be accomplished by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to an audio correction device or terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, such as a USB flash drive, a portable hard drive, a magnetic disk, or an optical disk.
[0119] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0120] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0121] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0122] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0123] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. An audio correction method, characterized in that, include: Obtain the audio to be corrected and the baseline audio; Obtain the historical performance audio of the singer whose audio needs to be corrected; The historical singing audio is input into the singing level prediction model to obtain the singing level prediction model's output training suggestions for the singer; The number of times the singer completed the training according to the training recommendations is obtained; Based on the number of training sessions completed, a target pitch correction amplitude is determined, wherein the target pitch correction amplitude decreases as the number of training sessions completed increases; Based on the reference audio, the audio to be corrected is processed according to the target correction amplitude to obtain the corrected audio.
2. The method according to claim 1, characterized in that, Inputting the historical singing audio into the singing level prediction model yields training suggestions for the singer output by the model, including: The historical singing audio is input into the singing level prediction model to obtain the first pitch deviation curve and the first rhythm deviation curve of the singer with respect to the historical singing audio, which are output by the singing level prediction model. Based on the first pitch deviation curve and the first rhythm deviation curve, determine the singer's historical pitch deviation and historical rhythm deviation with respect to the historical singing audio. If the historical pitch deviation is greater than a preset first threshold, a pitch training track is generated; If the historical rhythm deviation is greater than a preset second threshold, a rhythm training song is generated. The training suggestions are generated based on the pitch training piece and the rhythm training piece.
3. The method according to claim 1 or 2, characterized in that, The target pitch correction amplitude and the number of training sessions completed are expressed by the following mathematical expression: A n =A0·(1-β·n); Among them, A n To complete the target pitch correction amplitude after n training iterations, β is a preset decrease rate, A0 is a preset initial pitch correction amplitude, and n is the number of training iterations.
4. The method according to claim 1 or 2, characterized in that, Based on the reference audio, the audio to be corrected is processed according to the target correction amplitude to obtain the corrected audio, including: Feature extraction is performed on the audio to be corrected to obtain a first pitch feature vector, a first rhythm feature vector, and a timbre feature vector; Based on the target pitch correction amplitude and the reference pitch of the reference audio, the first pitch feature vector is corrected to generate the second pitch feature vector; Based on the target pitch correction amplitude and the reference rhythm of the reference audio, the first rhythm feature vector is corrected to generate the second rhythm feature vector; The second pitch feature vector, the second rhythm feature vector, and the timbre feature vector are fused to obtain the corrected audio.
5. The method according to claim 4, characterized in that, Feature extraction is performed on the audio to be corrected to obtain a first pitch feature vector, a first rhythm feature vector, and a timbre feature vector, including: The audio to be corrected is subjected to a short-time Fourier transform to obtain a first spectrum; The first spectrum is denoised to obtain the second spectrum; Feature extraction is performed on the second spectrum to obtain the first pitch feature vector, the first rhythm feature vector, and MFCC features; The MFCC features are subjected to timbre extraction processing to obtain the timbre feature vector.
6. The method according to claim 5, characterized in that, Feature extraction is performed on the second spectrum to obtain the first pitch feature vector and the first rhythm feature vector, including: The second spectrum is used to extract features using a fundamental frequency extraction algorithm to obtain the first pitch feature vector; The first rhythm feature vector is obtained by extracting features from the spectrum using a beat detection algorithm.
7. The method according to claim 4, characterized in that, After fusing the second pitch feature vector, the second rhythm feature vector, and the timbre feature vector to obtain the corrected audio, the method further includes: Feature extraction is performed on the audio to be corrected to obtain the volume change rate and vibrato rate; Based on the reference pitch and the first pitch feature vector, a first score is determined for the audio to be corrected; Based on the baseline rhythm and the first rhythm feature vector, a second score is determined for the audio to be corrected; A third score for the audio to be corrected is determined based on the volume change rate and the vibrato frequency. Based on the first score, the second score, and the third score, the total score of the audio to be corrected is determined; Based on the reference pitch and the first pitch feature vector, a second pitch deviation curve between the audio to be corrected and the reference audio is determined.
8. An audio correction device, characterized in that, include: The first acquisition module is used to acquire the audio to be corrected and the reference audio. And obtain the historical performance audio of the singer of the audio to be corrected; The input module is used to input the historical singing audio into the singing level prediction model to obtain the training suggestions for the singer output by the singing level prediction model; The second acquisition module is used to acquire the number of times the singer has completed training according to the training suggestions; A determining module is used to determine a target pitch correction amplitude based on the number of training sessions completed, wherein the target pitch correction amplitude decreases as the number of training sessions completed increases; The audio correction module is used to perform audio correction processing on the audio to be corrected based on the reference audio and according to the target audio correction amplitude, so as to obtain the corrected audio.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the audio correction method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, Includes a computer program, which, when run, causes the audio correction method as described in any one of claims 1 to 7 to be performed.